VLDB 2026 Research / reviewers in the wild / expert
Peng Gao 0007
dblp:29/5999-7
· DBLP profile ↗
98ranked-venue papers
11as first author
92since 2021 · last 2026
0009-0005-7881-712XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 89 · 11 first-author · 84 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 4 first-author · 44 since 2021Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image GenerationabstractDiffusion Transformers (DiTs) are a powerful yet underexplored class of generative models compared to U-Net-based diffusion architectures. We propose TIDE—Temporal-aware sparse autoencoders for Interpretable Diffusion transformErs—a framework designed to extract sparse, interpretable activation features across timesteps in DiTs. TIDE effectively captures temporally-varying representations and reveals that DiTs naturally learn hierarchical semantics (e.g., 3D structure, object class, and fine-grained concepts) during large-scale pretraining. Experiments show that TIDE enhances interpretability and controllability while maintaining reasonable generation quality, enabling applications such as safe image editing and style transfer. Victor Shea-Jay Huang, Le Zhuo, Yi Xin 0003, Zhaokai Wang, Fu-Yun Wang, Yuchi Wang, Renrui Zhang, Peng Gao 0007, Hongsheng Li 0001 |
AAAI | 8 |
| 2026 | Lumina-mGPT: Flexible Photorealistic Autoregressive Text-to-Image Generation
Yi Xin 0003, Shitian Zhao, Le Zhuo, Weifeng Lin, Xinyue Li 0001, Guangtao Zhai, Xiaohong Liu 0001, Hongsheng Li 0001, Yu Qiao 0001, Peng Gao 0007 |
Int. J. Comput. Vis. | 12 |
| 2025 | LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR UnderstandingabstractRecently, Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have shown promise in instruction following and image understanding. While these models are powerful, they have not yet been developed to comprehend the more challenging 3D geometric and physical scenes, especially when it comes to the sparse outdoor LiDAR data. In this paper, we introduce LiDAR-LLM, which takes raw LiDAR data as input and harnesses the remarkable reasoning capabilities of LLMs to gain a comprehensive understanding of outdoor 3D scenes. The central insight of our LiDAR-LLM is the reformulation of 3D outdoor scene cognition as a language modeling problem, encompassing tasks such as 3D captioning, 3D grounding, 3D question answering, etc. Specifically, due to the scarcity of 3D LiDAR-text pairing data, we introduce a three-stage training strategy and generate relevant datasets, progressively aligning the 3D modality with the language embedding of LLM. Furthermore, we design a Position-Aware Transformer (PAT) to connect the 3D encoder with the LLM, which effectively bridges the modality gap and enhances the LLM's spatial orientation comprehension of visual features. Our experiments demonstrate that LiDAR-LLM effectively comprehends a wide range of instructions related to 3D scenes, achieving a 40.9 BLEU-1 score on the 3D captioning dataset, a Grounded Captioning accuracy of 63.1%, and a BEV mIoU of 14.3%. Senqiao Yang, Jiaming Liu 0003, Renrui Zhang, Mingjie Pan, Xiaoqi Li 0009, Peng Gao 0007, Hongsheng Li 0001, Yandong Guo, Shanghang Zhang |
AAAI | 8 |
| 2025 | EfficientQAT: Efficient Quantization-Aware Training for Large Language ModelsabstractMengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang, Peng Gao, Kaipeng Zhang, Ping Luo. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Mengzhao Chen, Wenqi Shao, Peng Xu 0035, Jiahao Wang 0005, Peng Gao 0007, Kaipeng Zhang, Ping Luo 0002 |
ACL (1) | 5 |
| 2025 | Clinically Interpretable Geometric Constraints for Optic Cup and Disc SegmentationabstractWe present GCN-CIS, a segmentation framework that embeds clinically interpretable geometric constraints into deep networks for optic cup and disc delineation in fundus images. A Prior Boundary Attention Module (PBAM) sharpens boundary predictions, while an ISNT-guided geometric loss enforces anatomically plausible regional ordering. Evaluated on an internal glaucoma screening set and three public benchmarks (REFUGE, ORIGA, Drishti-GS), GCN-CIS yields consistent gains in cup/disc Dice and reduces vCDR error compared to strong baselines. The method improves clinical reliability of segmentation outputs and facilitates efficient vCDR estimation for screening workflows. Code will be released. Bailiang Zhao, Peng Gao 0007, Menghan Hu |
BIBM | 5 |
| 2025 | Adaptive Markup Language Generation for Contextually-Grounded Visual Document UnderstandingabstractVisual Document Understanding has become essential with the increase of text-rich visual content. This field poses significant challenges due to the need for effective integration of visual perception and textual comprehension, particularly across diverse document types with complex lay-outs. Moreover, existing fine-tuning datasets for this domain often fall short in providing the detailed contextual information for robust understanding, leading to hallucinations and limited comprehension of spatial relationships among visual elements. To address these challenges, we propose an innovative pipeline that utilizes adaptive generation of markup languages, such as Markdown, JSON, HTML, and TiKZ, to build highly structured document representations and deliver contextually-grounded responses. We intro-duce two fine-grained structured datasets: DocMark-Pile, comprising approximately 3.8M pretraining data pairs for document parsing, and DocMark-Instruct, featuring 624k fine-tuning data annotations for grounded instruction following. Extensive experiments demonstrate that our pro-posed model significantly outperforms existing state-of-the-art MLLMs across a range of visual document understanding benchmarks, facilitating advanced reasoning and comprehension capabilities in complex visual scenarios. Our code and models are released at https://github.com/Euphoria16/DocMark. Han Xiao 0010, Yina Xie, Guanxin Tan, Ke Wang 0036, Aojun Zhou, Hao Li 0069, Hao Shao, Peng Gao 0007, Yafei Wen, Xiaoxin Chen 0001, Shuai Ren 0002, Hongsheng Li 0001 |
CVPR | 11 |
| 2025 | Let's Verify and Reinforce Image Generation Step by StepabstractChain-of-Thought (CoT) reasoning has been extensively explored in large models to tackle complex understanding tasks. However, it remains an open question whether such strategies can be applied to verifying and reinforcing image generation scenarios. In this paper, we provide the first comprehensive investigation in the potential of CoT reasoning to enhance autoregressive image generation. We focus on three techniques: scaling test-time computation for verification, aligning model preferences with Direct Preference Optimization (DPO), and integrating these techniques for complementary effects. Our results demonstrate that these approaches can be effectively adapted and combined to significantly improve image generation performance. Furthermore, given the pivotal role of reward models in our findings, we propose the Potential Assessment Reward Model (PARM) specialized for autoregressive image generation. PARM adaptively assesses each generation step through a potential assessment mechanism, merging the strengths of existing reward models. Using our investigated reasoning strategies, we enhance a baseline model, Show-o, to achieve superior results, with a significant +24% improvement on the GenEval benchmark, surpassing Stable Diffusion 3 by +15%. We hope our study provides unique insights and paves a new path for integrating CoT reasoning with autoregressive image generation. Code is released at https://github.com/ZiyuGuo99/Image-Generation-CoT. Renrui Zhang, Chengzhuo Tong, Zhizheng Zhao, Haoquan Zhang, Manyuan Zhang, Jiaming Liu 0003, Peng Gao 0007, Hongsheng Li 0001 |
CVPR | 8 |
| 2025 | Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer Process
Kainan Yan, Shitian Zhao, Jie Wen 0001, Junjun He, Peng Gao 0007 |
ICCV | 8 |
| 2025 | VisualCloze: A Universal Image Generation Framework via Visual in-Context Learning
Ruoyi Du, Juncheng Yan, Le Zhuo, Peng Gao 0007, Zhanyu Ma, Ming-Ming Cheng |
ICCV | 6 |
| 2025 | Lumina-Image 2.0: a Unified and Efficient Image Generative FrameworkabstractWe introduce Lumina-Image 2.0, an advanced text-to-image generation framework that achieves significant progress compared to previous work, Lumina-Next. Lumina-Image 2.0 is built upon two key principles: (1) Unification - it adopts a unified architecture (Unified Next-DiT) that treats text and image tokens as a joint sequence, enabling natural cross-modal interactions and allowing seamless task expansion. Besides, since high-quality captioners can provide semantically well-aligned text-image training pairs, we introduce a unified captioning system, Unified Captioner (UniCap), specifically designed for T2I generation tasks. UniCap excels at generating comprehensive and accurate captions, accelerating convergence and enhancing prompt adherence. (2) Efficiency - to improve the efficiency of our proposed model, we develop multi-stage progressive training strategies and introduce inference acceleration techniques without compromising image quality. Extensive evaluations on academic benchmarks and public text-to-image arenas show that Lumina-Image 2.0 delivers strong performances even with only 2.6B parameters, highlighting its scalability and design efficiency. We have released our training details, code, and models at https://github.com/Alpha-VLLM/Lumina-Image-2.0. Le Zhuo, Yi Xin 0003, Ruoyi Du, Zhen Li 0026, Yiting Lu, Xinyue Li 0001, Will Beddow, Erwann Millon, Victor Perez 0005, Wenhai Wang, Yu Qiao 0001, Bo Zhang 0069, Xiaohong Liu 0001, Hongsheng Li 0001, Chang Xu 0002, Peng Gao 0007 |
ICCV | 20 |
| 2025 | Spatial Preference Rewarding for MLLMs Spatial Understanding
Han Qiu 0008, Peng Gao 0007, Lewei Lu, Xiaoqin Zhang 0002, Ling Shao 0001, Shijian Lu |
ICCV | 2 |
| 2025 | TAR3D: Creating High-Quality 3D Assets Via Next-Part PredictionabstractWe present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabilities of the next-token prediction paradigm to conditional 3D object generation. To achieve this, the 3D VQ-VAE first encodes a wide range of 3D shapes into a compact triplane latent space and utilizes a set of discrete representations from a trainable codebook to reconstruct fine-grained geometries under the supervision of query point occupancy. Then, the 3D GPT, equipped with a custom triplane position embedding called TriPE, predicts the codebook index sequence with prefilling prompt tokens in an autoregressive manner so that the composition of 3D geometries can be modeled part by part. Extensive experiments on ShapeNet and Objaverse demonstrate that TAR3D can achieve superior generation quality over existing methods in text-to-3D and image-to-3D tasks Xuying Zhang, Yangguang Li 0001, Renrui Zhang, Kai Wang 0001, Wanli Ouyang, Zhiwei Xiong, Peng Gao 0007, Qibin Hou, Ming-Ming Cheng |
ICCV | 9 |
| 2025 | From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection TuningabstractRecent text-to-image diffusion models achieve impressive visual quality through extensive scaling of training data and model parameters, yet they often struggle with complex scenes and fine-grained details. Inspired by the self-reflection capabilities emergent in large language models, we propose ReflectionFlow, an inference-time framework enabling diffusion models to iteratively reflect upon and refine their outputs. ReflectionFlow introduces three complementary inference-time scaling axes: (1) noise-level scaling to optimize latent initialization; (2) prompt-level scaling for precise semantic guidance; and most notably, (3) reflection-level scaling, which explicitly provides actionable reflections to iteratively assess and correct previous generations. To facilitate reflection-level scaling, we construct GenRef, a large-scale dataset comprising 1 million triplets, each containing a reflection, a flawed image, and an enhanced image. Leveraging this dataset, we efficiently perform reflection tuning on state-of-the-art diffusion transformer, FLUX.1-dev, by jointly modeling multimodal inputs within a unified framework. Experimental results show that ReflectionFlow significantly outperforms naive noise-level scaling methods, offering a scalable and compute-efficient solution toward higher-quality image synthesis on challenging tasks. Le Zhuo, Liangbing Zhao, Sayak Paul, Yue Liao, Renrui Zhang, Yi Xin 0003, Peng Gao 0007, Mohamed Elhoseiny 0001, Hongsheng Li 0001 |
ICCV | 7 |
| 2025 | Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution GenerationabstractSora unveils the potential of scaling Diffusion Transformer (DiT) for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lacks sufficient implementation details. In this paper, we introduce the Lumina-T2X family -- a series of Flow-based Large Diffusion Transformers (Flag-DiT) equipped with zero-initialized attention, as a simple and scalable generative framework that can be adapted to various modalities, e.g., transforming noise into images, videos, multi-view 3D objects, or audio clips conditioned on text instructions. By tokenizing the latent spatial-temporal space and incorporating learnable placeholders such as |[nextline]| and |[nextframe]| tokens, Lumina-T2X seamlessly unifies the representations of different modalities across various spatial-temporal resolutions. Advanced techniques like RoPE, KQ-Norm, and flow matching enhance the stability, flexibility, and scalability of Flag-DiT, enabling models of Lumina-T2X to scale up to 7 billion parameters and extend the context window to 128K tokens. This is particularly beneficial for creating ultra-high-definition images with our Lumina-T2I model and long 720p videos with our Lumina-T2V model. Remarkably, Lumina-T2I, powered by a 5-billion-parameter Flag-DiT, requires only 35% of the training computational costs of a 600-million-parameter naive DiT (PixArt-alpha), indicating that increasing the number of parameters significantly accelerates convergence of generative models without compromising visual quality. Our further comprehensive analysis underscores Lumina-T2X's preliminary capability in resolution extrapolation, high-resolution editing, generating consistent 3D views, and synthesizing videos with seamless transitions. All code and checkpoints of Lumina-T2X are released at https://github.com/Alpha-VLLM/Lumina-T2X to further foster creativity, transparency, and diversity in the generative AI community. Peng Gao 0007, Le Zhuo, Ruoyi Du, Longtian Qiu, Rongjie Huang 0001, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang 0001, Tianshuo Yang, Weicai Ye, Tong He 0001, Jingwen He, Junjun He, Yu Qiao 0001, Hongsheng Li 0001 |
ICLR | 1 |
| 2025 | MMSearch: Unveiling the Potential of Large Models as Multi-modal Search EnginesabstractThe advent of Large Language Models (LLMs) has paved the way for AI search engines, e.g., SearchGPT, showcasing a new paradigm in human-internet interaction. However, most current AI search engines are limited to text-only settings, neglecting the multimodal user queries and the text-image interleaved nature of website information. Recently, Large Multimodal Models (LMMs) have made impressive strides. Yet, whether they can function as AI search engines remains under-explored, leaving the potential of LMMs in multimodal search an open question. To this end, we first design a delicate pipeline, MMSearch-Engine, to empower any LMMs with multimodal search capabilities. On top of this, we introduce MMSearch, a comprehensive evaluation benchmark to assess the multimodal search performance of LMMs. The curated dataset contains 300 manually collected instances spanning 14 subfields, which involves no overlap with the current LMMs' training data, ensuring the correct answer can only be obtained within searching. By using MMSearch-Engine, the LMMs are evaluated by performing three individual tasks (requery, rerank, and summarization), and one challenging end-to-end task with a complete searching process. We conduct extensive experiments on closed-source and open-source LMMs. Among all tested models, GPT-4o with MMSearch-Engine achieves the best results, which surpasses the commercial product, Perplexity Pro, in the end-to-end task, demonstrating the effectiveness of our proposed pipeline. We further present error analysis to unveil current LMMs still struggle to fully grasp the multimodal search tasks, and conduct ablation study to indicate the potential of scaling test-time computation for AI search engine. We hope MMSearch may provide unique insights to guide the future development of multimodal AI search engine. Dongzhi Jiang, Renrui Zhang, Yanmin Wu, Jiayi Lei, Pengshuo Qiu, Pan Lu, Guanglu Song, Peng Gao 0007, Yu Liu 0015, Chunyuan Li, Hongsheng Li 0001 |
ICLR | 10 |
| 2025 | Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You WantabstractIn this paper, we present the Draw-and-Understand framework, exploring how to integrate visual prompting understanding capabilities into Multimodal Large Language Models (MLLMs). Visual prompts allow users to interact through multi-modal instructions, enhancing the models' interactivity and fine-grained image comprehension. In this framework, we propose a general architecture adaptable to different pre-trained MLLMs, enabling it to recognize various types of visual prompts (such as points, bounding boxes, and free-form shapes) alongside language understanding. Additionally, we introduce MDVP-Instruct-Data, a multi-domain dataset featuring 1.2 million image-visual prompt-text triplets, including natural images, document images, scene text images, mobile/web screenshots, and remote sensing images. Building on this dataset, we introduce MDVP-Bench, a challenging benchmark designed to evaluate a model's ability to understand visual prompting instructions. The experimental results demonstrate that our framework can be easily and effectively applied to various MLLMs, such as SPHINX-X and LLaVA. After training with MDVP-Instruct-Data and image-level instruction datasets, our models exhibit impressive multimodal interaction capabilities and pixel-level understanding, while maintaining their image-level visual perception performance. Weifeng Lin, Ruichuan An, Peng Gao 0007, Bocheng Zou, Yulin Luo, Siyuan Huang 0004, Shanghang Zhang, Hongsheng Li 0001 |
ICLR | 4 |
| 2025 | PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language InstructionsabstractThis paper presents a versatile image-to-image visual assistant, PixWizard, designed for image generation, manipulation, and translation based on free-from language instructions. To this end, we tackle a variety of vision tasks into a unified image-text-to-image generation framework and curate an Omni Pixel-to-Pixel Instruction-Tuning Dataset. By constructing detailed instruction templates in natural language, we comprehensively include a large set of diverse vision tasks such as text-to-image generation, image restoration, image grounding, dense image prediction, image editing, controllable generation, inpainting/outpainting, and more. Furthermore, we adopt Diffusion Transformers (DiT) as our foundation model and extend its capabilities with a flexible any resolution mechanism, enabling the model to dynamically process images based on the aspect ratio of the input, closely aligning with human perceptual processes. The model also incorporates structure-aware and semantic-aware guidance to facilitate effective fusion of information from the input image. Our experiments demonstrate that PixWizard not only shows impressive generative and understanding abilities for images with diverse resolutions but also exhibits generalization capabilities with unseen tasks and human instructions. Weifeng Lin, Renrui Zhang, Le Zhuo, Shitian Zhao, Siyuan Huang 0004, Junlin Xie, Peng Gao 0007, Hongsheng Li 0001 |
ICLR | 8 |
| 2025 | MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data EngineabstractMulti-modal Large Language Models (MLLMs) have recently showcased superior proficiency in general visual scenarios. However, we identify their mathematical capabilities remain under-explored with three areas to be improved: visual encoding of math diagrams, diagram-language alignment, and chain-of-thought (CoT) reasoning. This draws forth an urgent demand for an effective training paradigm and a large-scale, comprehensive dataset with detailed CoT rationales, which is challenging to collect and costly to annotate manually. To tackle this issue, we propose MAVIS, a MAthematical VISual instruction tuning pipeline for MLLMs, featuring an automatic data engine to efficiently create mathematical visual datasets.
We design the data generation process to be entirely independent of human intervention or GPT API usage, while ensuring the diagram-caption correspondence, question-answer correctness, and CoT reasoning quality. With this approach, we curate two datasets, MAVIS-Caption (558K diagram-caption pairs) and MAVIS-Instruct (834K visual math problems with CoT rationales), and propose four progressive stages for training MLLMs from scratch.
First, we utilize MAVIS-Caption to fine-tune a math-specific vision encoder (CLIP-Math) through contrastive learning, tailored for improved diagram visual encoding. Second, we also leverage MAVIS-Caption to align the CLIP-Math with a large language model (LLM) by a projection layer, enhancing vision-language alignment in mathematical domains. Third, we adopt MAVIS-Instruct to perform the instruction tuning for robust problem-solving skills, and term the resulting model as MAVIS-7B. Fourth, we apply Direct Preference Optimization (DPO) to enhance the CoT capabilities of our model, further refining its step-wise reasoning performance.
On various mathematical benchmarks, our MAVIS-7B achieves leading results among open-source MLLMs, e.g., surpassing other 7B models by +9.3% and the second-best LLaVA-NeXT (110B) by +6.9%, demonstrating the effectiveness of our method. Renrui Zhang, Dongzhi Jiang, Chengzhuo Tong, Jiaming Liu 0003, Aojun Zhou, Shanghang Zhang, Peng Gao 0007, Hongsheng Li 0001 |
ICLR | 10 |
| 2025 | MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and EfficiencyabstractAnswering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LMMs) still lacks a systematic assessment and in-depth investigation. In this paper, we introduce MME-CoT, a specialized benchmark evaluating the CoT reasoning performance of LMMs, spanning six domains: math, science, OCR, logic, space-time, and general scenes. As the first comprehensive study in this area, we propose a thorough evaluation suite incorporating three novel metrics that assess the reasoning quality, robustness, and efficiency at a fine-grained level. Leveraging curated high-quality data and a unique evaluation strategy, we conduct an in-depth analysis of state-of-the-art LMMs, uncovering several key insights: 1) Models with reflection mechanism demonstrate a superior CoT quality, with Kimi k1.5 outperforming GPT-4o and demonstrating the highest quality results; 2) CoT prompting often degrades LMM performance on perception-heavy tasks, suggesting a potentially harmful overthinking behavior; and 3) Although the CoT quality is high, LMMs with reflection exhibit significant inefficiency in both normal response and self-correction phases. We hope MME-CoT serves as a foundation for advancing multimodal reasoning in LMMs. Dongzhi Jiang, Renrui Zhang, Xinyan Chen 0001, Liuhui Wang, Jianhan Jin, Claire Guo, Bo Zhang 0069, Chaoyou Fu, Peng Gao 0007, Hongsheng Li 0001 |
ICML | 13 |
| 2025 | UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language ModelsabstractPrevious studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified formulation. Specifically, we constructed a dataset labeled with manipulation-related key attributes, comprising 900 articulated objects from 19 categories and 600 tools from 12 categories. Furthermore, we leverage MLLMs to infer object-centric representations for manipulation tasks, including affordance recognition and reasoning about 3D motion constraints. Comprehensive experiments in both simulation and real-world settings indicate that UniAff significantly improves the generalization of robotic manipulation for tools and articulated objects. We hope that UniAff will serve as a general baseline for unified robotic manipulation tasks in the future. Images, videos, dataset and code are published on the project website at:https://sites.google.com/view/uni-aff/home. Qiaojun Yu, Siyuan Huang 0004, Xibin Yuan, Zhengkai Jiang 0001, Ce Hao, Xin Li 0110, Haonan Chang, Junbo Wang 0004, Liu Liu 0012, Hongsheng Li 0001, Peng Gao 0007, Cewu Lu |
ICRA | 11 |
| 2025 | SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment ManipulationabstractAutomating garment manipulation poses a significant challenge for assistive robotics due to the diverse and de-formable nature of garments. Traditional approaches typically require separate models for each garment type, which limits scalability and adaptability. In contrast, this paper presents a unified approach using vision-language models (VLMs) to improve keypoint prediction across various garment categories. By interpreting both visual and semantic information, our model enables robots to manage different garment states with a single model. We created a large-scale synthetic dataset using advanced simulation techniques, allowing scalable training without extensive real-world data. Experimental results indicate that the VLM-based method significantly enhances keypoint detection accuracy and task success rates, providing a more flexible and general solution for robotic garment manipulation. In addition, this research also underscores the potential of VLMs to unify various garment manipulation tasks within a single framework, paving the way for broader applications in home automation and assistive robotics in the future. Xin Li 0110, Siyuan Huang 0004, Qiaojun Yu, Zhengkai Jiang 0001, Ce Hao, Yimeng Zhu, Hongsheng Li 0001, Peng Gao 0007, Cewu Lu |
IROS | 8 |
| 2025 | MMCD: Multi-Modal Collaborative Decision-Making for Connected Autonomy with Knowledge DistillationabstractAutonomous systems have advanced significantly, but challenges persist in accident-prone environments where robust decision-making is crucial. A single vehicle’s limited sensor range and obstructed views increase the likelihood of accidents. Multi-vehicle connected systems and multi-modal approaches, leveraging RGB images and LiDAR point clouds, have emerged as promising solutions. However, existing methods often assume the availability of all data modalities and connected vehicles during both training and testing, which is impractical due to potential sensor failures or missing connected vehicles. To address these challenges, we introduce a novel framework MMCD (Multi-Modal Collaborative Decision-making) for connected autonomy. Our framework fuses multi-modal observations from ego and collaborative vehicles to enhance decision-making under challenging conditions. To ensure robust performance when certain data modalities are unavailable during testing, we propose an approach based on cross-modal knowledge distillation with a teacher-student model structure. The teacher model is trained with multiple data modalities, while the student model is designed to operate effectively with reduced modalities. In experiments on connected autonomous driving with ground vehicles and aerial-ground vehicles collaboration, our method improves driving safety by up to 20.7%, surpassing the best-existing baseline in detecting potential accidents and making safe driving decisions. More information can be found on our website https://ruiiu.github.io/mmcd. Rui Liu 0040, Zikang Wang, Peng Gao 0007, Pratap Tokekar, Ming C. Lin |
IROS | 3 |
| 2025 | EnerVerse: Envisioning Embodied Future Space for Robotics ManipulationabstractWe introduce EnerVerse, a generative robotics foundation model that constructs and interprets embodied spaces. EnerVerse employs a chunk-wise autoregressive video diffusion framework to predict future embodied spaces from instructions, enhanced by a sparse context memory for long-term reasoning. To model the 3D robotics world, we adopt a multi-view video representation, providing rich perspectives to address challenges like motion ambiguity and 3D grounding. Additionally, EnerVerse-D, a data engine pipeline combining generative modeling with 4D Gaussian Splatting, forms a self-reinforcing data loop to reduce the sim-to-real gap. Leveraging these innovations, EnerVerse translates 4D world representations into physical actions via a policy head (EnerVerse-A), achieving state-of-the-art performance in both simulation and real-world tasks. For efficiency, EnerVerse-A reuses features from the first denoising step and predicts action chunks, achieving about 280 ms per 8-step action chunk on a single RTX 4090. Further video demos, dataset samples could be found in our project page. Siyuan Huang 0004, Liliang Chen, Shengcong Chen, Yue Liao, Zhengkai Jiang 0001, Peng Gao 0007, Hongsheng Li 0001, Maoqing Yao, Guanghui Ren |
NeurIPS | 8 |
| 2025 | CAML: Collaborative Auxiliary Modality Learning for Multi-Agent SystemsabstractMulti-modal learning has emerged as a key technique for improving performance across domains such as autonomous driving, robotics, and reasoning. However, in certain scenarios, particularly in resource-constrained environments, some modalities available during training may be absent during inference. While existing frameworks effectively utilize multiple data sources during training and enable inference with reduced modalities, they are primarily designed for single-agent settings. This poses a critical limitation in dynamic environments such as connected autonomous vehicles (CAV), where incomplete data coverage can lead to decision-making blind spots. Conversely, some works explore multi-agent collaboration but without addressing missing modality at test time. To overcome these limitations, we propose Collaborative Auxiliary Modality Learning (CAML), a novel multi-modal multi-agent framework that enables agents to collaborate and share multi-modal data during training, while allowing inference with reduced modalities during testing. Experimental results in collaborative decision-making for CAV in accident-prone scenarios demonstrate that CAML achieves up to a 58.1% improvement in accident detection. Additionally, we validate CAML on real-world aerial-ground robot data for collaborative semantic segmentation, achieving up to a 10.6% improvement in mIoU. Rui Liu 0040, Peng Gao 0007, Pratap Tokekar, Ming C. Lin |
NeurIPS | 3 |
| 2025 | 3DAxisPrompt: Promoting the 3D grounding and reasoning in GPT-4o
Dingning Liu, Cheng Wang 0026, Peng Gao 0007, Renrui Zhang, Xinzhu Ma, Zhihui Wang 0001 |
Neurocomputing | 3 |
| 2025 | LVLM-EHub: A Comprehensive Evaluation Benchmark for Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have recently played a dominant role in multimodal vision-language learning. Despite the great success, it lacks a holistic evaluation of their efficacy. This paper presents a comprehensive evaluation of publicly available large multimodal models by building an LVLM evaluation Hub (LVLM-eHub). Our LVLM-eHub consists of 13 representative LVLMs such as InstructBLIP and LLaVA, which are thoroughly evaluated by a quantitative capability evaluation and an online arena platform. The former evaluates five categories of multimodal capabilities of LVLMs such as visual question answering and object hallucination on 42 in-domain text-related visual benchmarks, while the latter provides the user-level evaluation of LVLMs in an open-world question-answering scenario. The study investigates how specific features of LVLMs such as model configurations, modality alignment mechanisms, and training data affect the multimodal understanding. By conducting a comprehensive comparison of these features on quantitative and arena evaluation, our study uncovers several innovative findings, which establish a fundamental framework for the development and evaluation of innovative strategies aimed at enhancing multimodal techniques. Our LVLM-eHub is available at https://github.com/OpenGVLab/Multi-Modality-Arena. Peng Xu 0035, Wenqi Shao, Kaipeng Zhang, Peng Gao 0007, Meng Lei, Fanqing Meng, Siyuan Huang 0004, Yu Qiao 0001, Ping Luo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have made significant strides in various multimodal tasks. Notably, GPT4V, Claude, Gemini, and others showcase exceptional multimodal capabilities, marked by profound comprehension and reasoning skills. This study introduces a comprehensive and efficient evaluation framework, TinyLVLM-eHub, to assess LVLMs’ performance, including proprietary models. TinyLVLM-eHub covers six key multimodal capabilities, such as visual perception, knowledge acquisition, reasoning, commonsense understanding, object hallucination, and embodied intelligence. The benchmark, utilizing 2.1K image-text pairs, provides a user-friendly and accessible platform for LVLM evaluation. The evaluation employs the ChatGPT Ensemble Evaluation (CEE) method, which improves alignment with human evaluation compared to word-matching approaches. Results reveal that closed-source API models like GPT4V and GeminiPro-V excel in most capabilities compared to previous open-source LVLMs, though they show some vulnerability in object hallucination. This evaluation underscores areas for LVLM improvement in real-world applications and serves as a foundational assessment for future multimodal advancements. Wenqi Shao, Meng Lei, Yutao Hu 0002, Peng Gao 0007, Peng Xu 0035, Kaipeng Zhang, Fanqing Meng, Siyuan Huang 0004, Hongsheng Li 0001, Yu Qiao 0001, Ping Luo 0002 |
IEEE Trans. Big Data | 4 |
| 2024 | Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationabstractRecently, video object segmentation (VOS) referred by multi-modal signals, e.g., language and audio, has evoked increasing attention in both industry and academia. It is challenging for exploring the semantic alignment within modalities and the visual correspondence across frames. However, existing methods adopt separate network architectures for different modalities, and neglect the inter-frame temporal interaction with references. In this paper, we propose MUTR, a Multi-modal Unified Temporal transformer for Referring video object segmentation. With a unified framework for the first time, MUTR adopts a DETR-style transformer and is capable of segmenting video objects designated by either text or audio reference. Specifically, we introduce two strategies to fully explore the temporal relations between videos and multi-modal signals. Firstly, for low-level temporal aggregation before the transformer, we enable the multi-modal references to capture multi-scale visual cues from consecutive video frames. This effectively endows the text or audio signals with temporal knowledge and boosts the semantic alignment between modalities. Secondly, for high-level temporal interaction after the transformer, we conduct inter-frame feature communication for different object embeddings, contributing to better object-wise correspondence for tracking along the video. On Ref-YouTube-VOS and AVSBench datasets with respective text and audio references, MUTR achieves +4.2% and +8.7% J&F improvements to state-of-the-art methods, demonstrating our significance for unified multi-modal VOS. Code is released at https://github.com/OpenGVLab/MUTR. Shilin Yan, Renrui Zhang, Wei Zhang 0016, Hongyang Li 0001, Yu Qiao 0001, Hao Dong 0003, Zhongjiang He, Peng Gao 0007 |
AAAI | 10 |
| 2024 | Masked AutoDecoder is Effective Multi-Task Vision GeneralistabstractInspired by the success of general-purpose models in NLP, recent studies attempt to unify different vision tasks in the same sequence format and employ autoregressive Transformers for sequence prediction. They apply uni-directional attention to capture sequential dependencies and generate task sequences recursively. However, such autoregressive Transformers may not fit vision tasks well, as vision task sequences usually lack the sequential dependencies typically observed in natural languages. In this work, we design Masked AutoDecoder (MAD), an effective multitask vision generalist. MAD consists of two core designs. First, we develop a parallel decoding framework that introduces bi-directional attention to capture contextual dependencies comprehensively and decode vision task sequences in parallel. Second, we design a masked sequence modeling approach that learns rich task contexts by masking and reconstructing task sequences. In this way, MAD handles all the tasks by a single network branch and a simple cross-entropy loss with minimal task-specific designs. Extensive experiments demonstrate the great potential of MAD as a new paradigm for unifying various vision tasks. MAD achieves superior performance and inference efficiency compared to autoregressive counterparts while obtaining competitive accuracy with task-specific models. Code will be released at https://github.com/hanqiu-hq/MAD. Han Qiu 0008, Jiaxing Huang 0001, Peng Gao 0007, Lewei Lu, Xiaoqin Zhang 0002, Shijian Lu |
CVPR | 3 |
| 2024 | Digital Life Project: Autonomous 3D Characters with Social IntelligenceabstractIn this work, we present Digital Life Project, a framework utilizing language as the universal medium to build autonomous 3D characters, who are capable of engaging in social interactions and expressing with articulated body motions, thereby simulating life in a digital environment. Our framework comprises two primary components: 1) SocioMind: a meticulously crafted digital brain that models personalities with systematic few-shot exemplars, incorporates a reflection process based on psychology principles, and emulates autonomy by initiating dialogue topics; 2) MoMat-MoGen: a text-driven motion synthesis paradigm for controlling the character's digital body. It integrates motion matching, a proven industry technique to ensure motion quality, with cutting-edge advancements in motion generation for diversity. Extensive experiments demonstrate that each module achieves state-of-the-art performance in its respective domain. Collectively, they enable virtual characters to initiate and sustain dialogues autonomously, while evolving their socio-psychological states. Concurrently, these characters can perform contextually relevant bodily movements. Additionally, an extension of DLP enables a virtual character to recognize and appropriately respond to human players' actions. Zhongang Cai, Zhongfei Qing, Zhengyu Lin, Haiyi Mei, Ruisi Wang, Wanqi Yin, Liang Pan, Xiangyu Fan 0002, Peng Gao 0007, Zhitao Yang, Yang Gao 0042, Jiaqi Li 0007, Tianxiang Ren, Yukun Wei, Xiaogang Wang 0001, Chen Change Loy, Lei Yang 0059, Ziwei Liu 0002 |
CVPR | 14 |
| 2024 | OneLLM: One Framework to Align All Modalities with LanguageabstractMultimodal large language models (MLLMs) have gained significant attention due to their strong multimodal understanding capability. However, existing works rely heavily on modality-specific encoders, which usually differ in architecture and are limited to common modalities. In this paper, we present OneLLM, an MLLM that aligns eight modalities to language using a unified framework. We achieve this through a unified multimodal encoder and a progressive multimodal alignment pipeline. In detail, we first train an image projection module to connect a vision encoder with LLM. Then, we build a universal projection module (UPM) by mixing multiple image projection modules and dynamic routing. Finally, we progressively align more modalities to LLM with the UPM. To fully leverage the potential of OneLLM in following instructions, we also curated a comprehensive multimodal instruction dataset, including 2M items from image, audio, video, point cloud, depth/normal map, IMU and fMRI brain activity. OneLLM is evaluated on 25 diverse benchmarks, encompassing tasks such as multimodal captioning, question answering and reasoning, where it delivers excellent performance. Code, data, model and online demo are available at https://github.com/csuhan/OneLLM. Jiaming Han, Kaixiong Gong, Jiaqi Wang 0003, Kaipeng Zhang, Dahua Lin, Yu Qiao 0001, Peng Gao 0007, Xiangyu Yue 0001 |
CVPR | 8 |
| 2024 | No Time to Train: Empowering Non-Parametric Networks for Few-Shot 3D Scene SegmentationabstractTo reduce the reliance on large-scale datasets, recent works in 3D segmentation resort to few-shot learning. current 3D few-shot segmentation methods first pre-train models on ‘seen’ classes, and then evaluate their generalization performance on ‘unseen’ classes. However, the prior pre-training stage not only introduces excessive time over-head but also incurs a significant domain gap on ‘un-seen’ classes. To tackle these issues, we propose a Non-parametric Network for few-shot 3D Segmentation, Seg-NN, and its Parametric variant, Seg-PN. Without training, Seg-NN extracts dense representations by hand-crafted filters and achieves comparable performance to existing parametric models. Due to the elimination of pre-training, Seg-NN can alleviate the domain gap issue and save a substantial amount of time. Based on Seg-NN, Seg-PN only requires training a lightweight QUEry-Support Transferring (QUEST) module, which enhances the interaction between the support set and query set. Experiments suggest that Seg-PN outperforms previous state-of-the-art method by +4.19% and +7.71% mloU on S3DIS and ScanNet datasets respectively, while reducing training time by -90%, indicating its effectiveness and efficiency. Code is available here. Renrui Zhang, Bowei He, Jiaming Liu 0003, Han Xiao 0010, Chaoyou Fu, Hao Dong 0003, Peng Gao 0007 |
CVPR | 9 |
| 2024 | SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models
Renrui Zhang, Peng Gao 0007, Longtian Qiu, Han Xiao 0010, Han Qiu 0010, Wenqi Shao, Keqin Chen, Jiaming Han, Siyuan Huang 0004, Xuming He 0001, Yu Qiao 0001, Hongsheng Li 0001 |
ECCV (62) | 4 |
| 2024 | Any2Point: Empowering Any-Modality Large Models for Efficient 3D Understanding
Ray Zhang 0002, Jiaming Liu 0003, Zoey Guo, Bin Zhao 0001, Zhigang Wang 0002, Peng Gao 0007, Hongsheng Li 0001, Dong Wang 0028, Xuelong Li 0001 |
ECCV (36) | 7 |
| 2024 | SpatialFormer: Towards Generalizable Vision Transformers with Explicit Spatial Understanding
Han Xiao 0010, Wenzhao Zheng, Sicheng Zuo, Peng Gao 0007, Jie Zhou 0001, Jiwen Lu |
ECCV (13) | 4 |
| 2024 | MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Renrui Zhang, Dongzhi Jiang, Haokun Lin, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang 0001, Yu Qiao 0001, Peng Gao 0007, Hongsheng Li 0001 |
ECCV (8) | 11 |
| 2024 | OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsabstractLarge language models (LLMs) have revolutionized natural language processing tasks. However, their practical deployment is hindered by their immense memory and computation requirements. Although recent post-training quantization (PTQ) methods are effective in reducing memory footprint and improving the computational efficiency of LLM, they hand-craft quantization parameters, leading to low performance, especially in extremely low-bit quantization. To tackle this issue, we introduce an Omnidirectionally calibrated Quantization ($\textbf{OmniQuant}$) technique for LLMs, which achieves good performance in diverse quantization settings while maintaining the computational efficiency of PTQ by efficiently optimizing various quantization parameters. OmniQuant comprises two innovative components including Learnable Weight Clipping (LWC) and Learnable Equivalent Transformation (LET). LWC modulates the extreme values of weights by optimizing the clipping threshold. Meanwhile, LET tackles activation outliers by shifting the challenge of quantization from activations to weights. Operating within a differentiable framework using block-wise error minimization, OmniQuant can optimize the quantization process efficiently for both weight-only and weight-activation quantization. For instance, the LLaMA-2 model family size 7-70B can be processed with OmniQuant on a single A100-40G GPU within 1-16 hours using 128 samples. Extensive experiments validate OmniQuant's superior performance across diverse quantization configurations such as W4A4 (4-bit weight, 4-bit activation), W6A6, W4A16, W3A16, and W2A16. Additionally, OmniQuant demonstrates effectiveness in instruction-tuned models and delivers notable improvements in inference speed and memory reduction on real devices. Codes are available at
\url{https://github.com/OpenGVLab/OmniQuant}. Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang 0004, Peng Xu 0035, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao 0007, Yu Qiao 0001, Ping Luo 0002 |
ICLR | 8 |
| 2024 | BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity AllocationabstractLarge language models (LLMs) have demonstrated outstanding performance in various tasks, such as text summarization, text question-answering, and etc. While their performance is impressive, the computational footprint due to their vast number of parameters can be prohibitive. Existing solutions such as SparseGPT and Wanda attempt to alleviate this issue through weight pruning. However, their layer-wise approach results in significant perturbation to the model's output and requires meticulous hyperparameter tuning, such as the pruning rate, which can adversely affect overall model performance. To address this, this paper introduces a novel LLM pruning technique dubbed blockwise parameter-efficient sparsity allocation (BESA) by applying a blockwise reconstruction loss. In contrast to the typical layer-wise pruning techniques, BESA is characterized by two distinctive attributes: i) it targets the overall pruning error with respect to individual transformer blocks, and ii) it allocates layer-specific sparsity in a differentiable manner, both of which ensure reduced performance degradation after pruning. Our experiments show that BESA achieves state-of-the-art performance, efficiently pruning LLMs like LLaMA1, and LLaMA2 with 7B to 70B parameters on a single A100 GPU in just five hours. Code is available at [here](https://github.com/LinkAnonymous/BESA). Peng Xu 0035, Wenqi Shao, Mengzhao Chen, Shitao Tang, Kaipeng Zhang, Peng Gao 0007, Fengwei An, Yu Qiao 0001, Ping Luo 0002 |
ICLR | 6 |
| 2024 | Personalize Segment Anything Model with One ShotabstractDriven by large-data pre-training, Segment Anything Model (SAM) has been demonstrated as a powerful promptable framework, revolutionizing the segmentation field. Despite the generality, customizing SAM for specific visual concepts without man-powered prompting is under-explored, e.g., automatically segmenting your pet dog in numerous images. In this paper, we introduce a training-free Personalization approach for SAM, termed PerSAM. Given only one-shot data, i.e., a single image with a reference mask, we first obtain a positive-negative location prior for the target concept in new images. Then, aided by target visual semantics, we empower SAM for personalized object segmentation via two proposed techniques: target-guided attention and target-semantic prompting. In this way, we can effectively customize the general-purpose SAM for private use without any training. To further alleviate the ambiguity of segmentation scales, we present an efficient one-shot fine-tuning variant, PerSAM-F. Freezing the entire SAM, we introduce a scale-aware fine-tuning to aggregate multi-scale masks, which only tunes 2 parameters within 10 seconds for improved performance. To demonstrate our efficacy, we construct a new dataset, PerSeg, for the evaluation of personalized object segmentation, and also test our methods on various one-shot image and video segmentation benchmarks. Besides, we propose to leverage PerSAM to improve DreamBooth for personalized text-to-image synthesis. By mitigating the disturbance of training-set backgrounds, our approach showcases better target appearance generation and higher fidelity to the input text prompt. Code is released at https://github.com/ZrrSkywalker/Personalize-SAM. Renrui Zhang, Zhengkai Jiang 0001, Shilin Yan, Junting Pan, Hao Dong 0003, Yu Qiao 0001, Peng Gao 0007, Hongsheng Li 0001 |
ICLR | 8 |
| 2024 | LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero-initialized AttentionabstractWith the rising tide of large language models (LLMs), there has been a growing interest in developing general-purpose instruction-following models, e.g., ChatGPT. To this end, we present LLaMA-Adapter, a lightweight adaption method for efficient instruction tuning of LLaMA. Using 52K self-instruct demonstrations, LLaMA-Adapter only introduces 1.2M learnable parameters upon the frozen LLaMA 7B model, and costs less than one hour for fine-tuning. Specifically, a zero-initialized attention mechanism is proposed. It adopts a learnable zero gating to adaptively inject the instructional cues into LLaMA within self-attention layers, contributing to a stable training process and superior final performance. In this way, LLaMA-Adapter can generate high-quality responses to diverse language instructions, comparable to Alpaca with fully fine-tuned 7B parameters. Besides language commands, by incorporating an image encoder, our approach can be simply extended to a multi-modal LLM for image-conditioned instruction following, which achieves superior multi-modal reasoning capacity on several popular benchmarks (MME, MMBench, LVLM-eHub). Furthermore, we also verify the proposed zero-initialized attention mechanism for fine-tuning other pre-trained models (ViT, RoBERTa, CLIP) on traditional vision and language tasks, demonstrating the effectiveness and generalizability of our approach. Code and models are released at https://github.com/OpenGVLab/LLaMA-Adapter. Renrui Zhang, Jiaming Han, Chris Liu, Aojun Zhou, Pan Lu, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
ICLR | 8 |
| 2024 | FreeBind: Free Lunch in Unified Multimodal Space via Knowledge FusionabstractUnified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained unified spaces. In this work, we propose FreeBind, an idea that treats multimodal representation spaces as basic units, and freely augments pre-trained unified space by integrating knowledge from extra expert spaces via “space bonds". Specifically, we introduce two kinds of basic space bonds: 1) Space Displacement Bond and 2) Space Combination Bond. Based on these basic bonds, we design Complex Sequential & Parallel Bonds to effectively integrate multiple spaces simultaneously. Benefiting from the modularization concept, we further propose a coarse-to-fine customized inference strategy to flexibly adjust the enhanced unified space for different purposes. Experimentally, we bind ImageBind with extra image-text and audio-text expert spaces, resulting in three main variants: ImageBind++, InternVL_IB, and InternVL_IB++. These resulting spaces outperform ImageBind on 5 audio-image-text downstream tasks across 9 datasets. Moreover, via customized inference, it even surpasses the advanced audio-text and image-text expert spaces. Our code and checkpoints are released at https://github.com/zehanwang01/FreeBind Zehan Wang 0001, Xize Cheng, Rongjie Huang 0001, Luping Liu, Zhenhui Ye, Haifeng Huang 0001, Yang Zhao 0022, Tao Jin 0004, Peng Gao 0007, Zhou Zhao 0001 |
ICML | 10 |
| 2024 | InstructSpeech: Following Speech Editing Instructions via Large Language ModelsabstractInstruction-guided speech editing aims to follow the user’s natural language instruction to manipulate the semantic and acoustic attributes of a speech. In this work, we construct triplet paired data (instruction, input speech, output speech) to alleviate data scarcity and train a multi-task large language model named InstructSpeech. To mitigate the challenges of accurately executing user’s instructions, we 1) introduce the learned task embeddings with a fine-tuned Flan-T5-XL to guide the generation process towards the correct generative task; 2) include an extensive and diverse set of speech editing and processing tasks to enhance model capabilities; 3) investigate chain-of-thought reasoning for free-form semantic content editing; and 4) propose a hierarchical adapter that effectively updates a small portion of parameters for generalization to new tasks. To assess instruction speech editing in greater depth, we introduce a benchmark evaluation with contrastive instruction-speech pre-training (CISP) to test the speech quality and instruction-speech alignment faithfulness. Experimental results demonstrate that InstructSpeech achieves state-of-the-art results in eleven tasks, for the first time unlocking the ability to edit speech’s acoustic and semantic attributes following a user’s instruction. Audio samples are available at https://InstructSpeech.github.io Rongjie Huang 0001, Ruofan Hu 0002, Zehan Wang 0001, Xize Cheng, Ziyue Jiang 0001, Zhenhui Ye, Dongchao Yang, Luping Liu, Peng Gao 0007, Zhou Zhao 0001 |
ICML | 10 |
| 2024 | SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language ModelsabstractWe propose SPHINX-X, an extensive Multi-modality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual encoders, bypassing fully-padded sub-images with skip tokens, and simplifying multi-stage training into a one-stage all-in-one paradigm. To fully unleash the potential of MLLMs, we assemble a comprehensive multi-domain and multi-modal dataset covering publicly available resources in language, vision, and vision-language tasks. We further enrich this collection with our curated OCR intensive and Set-of-Mark datasets, extending the diversity and generality. By training over different base LLMs including TinyLlama-1.1B, InternLM2-7B, LLaMA2-13B, and Mixtral-8$\times$7B, we obtain a spectrum of MLLMs that vary in parameter size and multilingual capabilities. Comprehensive benchmarking reveals a strong correlation between the multi-modal performance with the data and parameter scales. Code and models are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory. Renrui Zhang, Longtian Qiu, Siyuan Huang 0004, Weifeng Lin, Shitian Zhao, Shijie Geng, Kaipeng Zhang, Wenqi Shao, Conghui He, Junjun He, Hao Shao, Pan Lu, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
ICML | 19 |
| 2024 | SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language ModelsabstractLarge Language Models (LLMs) have become pivotal in advancing the field of artificial intelligence, yet their immense sizes pose significant challenges for both fine-tuning and deployment. Current post-training pruning methods, while reducing the sizes of LLMs, often fail to maintain their original performance. To address these challenges, this paper introduces SPP, a **S**parsity-**P**reserved **P**arameter-efficient fine-tuning method. Different from existing post-training pruning approaches that struggle with performance retention, SPP proposes to employ lightweight learnable column and row matrices to optimize sparse LLM weights, *keeping the structure and sparsity of pruned pre-trained models intact*. By element-wise multiplication and residual addition, SPP ensures the consistency of model sparsity pattern and ratio during both training and weight-merging processes. We demonstrate the effectiveness of SPP by applying it to the LLaMA and LLaMA-2 model families with recent post-training pruning methods. Our results show that SPP significantly enhances the performance of models with different sparsity patterns (i.e. unstructured and N:M sparsity), especially for those with high sparsity ratios (e.g. 75%), making it a promising solution for the efficient fine-tuning of sparse LLMs. Code will be made available at https://github.com/Lucky-Lance/SPP. Aojun Zhou, Renrui Zhang, Peng Gao 0007, Hongsheng Li 0001 |
ICML | 5 |
| 2024 | MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIabstractLarge Vision-Language Models (LVLMs) show significant strides in general-propose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, and reasoning. MMT-Bench comprises $31,325$ meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering $32$ core meta-tasks and $162$ subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving $20$ publicly available LVLMs such as the proprietary GeminiProVision model, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence. Kaining Ying, Fanqing Meng, Zhiqian Li, Hao Zhang 0117, Wenbo Zhang 0009, Yuqi Lin, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu 0035, Renrui Zhang, Haozhe Zhang 0002, Peng Gao 0007, Yali Wang 0001, Yu Qiao 0001, Ping Luo 0002, Kaipeng Zhang, Wenqi Shao |
ICML | 17 |
| 2024 | Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation SkillabstractZero-shot object navigation is a challenging task for home-assistance robots. This task emphasizes visual grounding, commonsense inference and locomotion abilities, where the first two are inherent in foundation models. But for the locomotion part, most works still depend on map-based planning approaches. The gap between RGB space and map space makes it difficult to directly transfer the knowledge from foundation models to navigation tasks. In this work, we propose a Pixel-guided Navigation skill (PixNav), which bridges the gap between the foundation models and the embodied navigation task. It is straightforward for recent foundation models to indicate an object by pixels, and with pixels as the goal specification, our method becomes a versatile navigation policy towards all different kinds of objects. Besides, our PixNav is a pure RGB-based policy that can reduce the cost of homeassistance robots. Experiments demonstrate the robustness of the PixNav which achieves 80+% success rate in the local path-planning task. To perform long-horizon object navigation, we design an LLM-based planner to utilize the commonsense knowledge between objects and rooms to select the best waypoint. Evaluations across both photorealistic indoor simulators and real-world environments validate the effectiveness of our proposed navigation strategy. More details are accessible via our project website https://sites.google.com/view/pixnav/. Wenzhe Cai, Siyuan Huang 0004, Guangran Cheng, Yuxing Long, Peng Gao 0007, Changyin Sun 0001, Hao Dong 0003 |
ICRA | 5 |
| 2024 | Collaborative Decision-Making Using Spatiotemporal Graphs in Connected AutonomyabstractCollaborative decision-making is an essential capability for multi-robot systems, such as connected vehicles, to collaboratively control autonomous vehicles in accident-prone scenarios. Under limited communication bandwidth, capturing comprehensive situational awareness by integrating connected agents’ observation is very challenging. In this paper, we propose a novel collaborative decision-making method that efficiently and effectively integrates collaborators’ representations to control the ego vehicle in accident-prone scenarios. Our approach formulates collaborative decision-making as a classification problem. We first represent sequences of raw observations as spatiotemporal graphs, which significantly reduce the package size to share among connected vehicles. Then we design a novel spatiotemporal graph neural network based on heterogeneous graph learning, which analyzes spatial and temporal connections of objects in a unified way for collaborative decision-making. We evaluate our approach using a high-fidelity simulator that considers realistic traffic, communication bandwidth, and vehicle sensing among connected autonomous vehicles. The experimental results show that our representation achieves over 100x reduction in the shared data size that meets the requirements of communication bandwidth for connected autonomous driving. In addition, our approach achieves over 30% improvements in driving safety. Peng Gao 0007, Ming C. Lin |
ICRA | 1 |
| 2024 | MTG: Mapless Trajectory Generator with Traversability Coverage for Outdoor NavigationabstractWe present a novel learning-based trajectory generation algorithm for outdoor robot navigation. Our goal is to compute collision-free paths that also satisfy the environment-specific traversability constraints. Our approach is designed for global planning using limited onboard robot perception in mapless environments while ensuring comprehensive coverage of all traversable directions. Our formulation uses a Conditional Variational Autoencoder (CVAE) generative model that is enhanced with traversability constraints and an optimization formulation used for the coverage. We highlight the benefits of our approach over state-of-the-art trajectory generation approaches and demonstrate its performance in challenging and large outdoor environments, including around buildings, across intersections, along trails, and off-road terrain, using a Clearpath Husky and a Boston Dynamics Spot robot. In practice, our approach results in a 6% improvement in coverage of traversable areas and an 89% reduction in trajectory portions residing in non-traversable regions. Our video is here: https://youtu.be/3eJ2soAzXnU Jing Liang 0006, Peng Gao 0007, Xuesu Xiao, Adarsh Jagan Sathyamoorthy, Mohamed Elnoor, Ming C. Lin, Dinesh Manocha |
ICRA | 2 |
| 2024 | ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language ModelsabstractWhile the integration of Multi-modal Large Language Models (MLLMs) with robotic systems has significantly improved robots’ ability to understand and execute natural language instructions, their performance in manipulation tasks remains limited due to a lack of robotics-specific knowledge. Conventional MLLMs are typically trained on generic image-text pairs, leaving them deficient in understanding affordances and physical concepts crucial for manipulation. To address this gap, we propose ManipVQA, a novel framework that infuses MLLMs with manipulation-centric knowledge through a Visual Question-Answering (VQA) format. This approach encompasses tool detection, affordance recognition, and a broader understanding of physical concepts. We curated a diverse dataset of images depicting interactive objects, to challenge robotic understanding in tool detection, affordance prediction, and physical concept comprehension. To effectively integrate this robotics-specific knowledge with the inherent vision-reasoning capabilities of MLLMs, we leverage a unified VQA format and devise a fine-tuning strategy. This strategy preserves the original vision-reasoning abilities while incorporating the newly acquired robotic insights. Empirical evaluations conducted in robotic simulators and across various vision task benchmarks demonstrate the robust performance of ManipVQA. The code and dataset are publicly available at https://github.com/SiyuanHuang95/ManipVQA. Siyuan Huang 0004, Iaroslav Ponomarenko, Zhengkai Jiang 0001, Xiaoqi Li 0020, Xiaobin Hu, Peng Gao 0007, Hongsheng Li 0001, Hao Dong 0003 |
IROS | 6 |
| 2024 | Phased Consistency ModelsabstractConsistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of Latent Consistency Models~(LCMs). We investigate the reasons behind these limitations and propose Phased Consistency Models (PCMs), which generalize the design space and address the identified limitations. Our evaluations demonstrate that PCMs outperform LCMs across 1--16 step generation settings. While PCMs are specifically designed for multi-step refinement, they achieve comparable 1-step generation results to previously state-of-the-art specifically designed 1-step methods. Furthermore, we show the methodology of PCMs is versatile and applicable to video generation, enabling us to train the state-of-the-art few-step text-to-video generator. Our code is available at https://github.com/G-U-N/Phased-Consistency-Model. Fu-Yun Wang, Alexander William Bergman, Dazhong Shen, Peng Gao 0007, Michael Lingelbach, Keqiang Sun, Weikang Bian, Guanglu Song, Yu Liu 0015, Xiaogang Wang 0001, Hongsheng Li 0001 |
NeurIPS | 5 |
| 2024 | Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiTabstractLumina-T2X is a nascent family of Flow-based Large Diffusion Transformers (Flag-DiT) that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters challenges including training instability, slow inference, and extrapolation artifacts. In this paper, we present Lumina-Next, an improved version of Lumina-T2X, showcasing stronger generation performance with increased training and inference efficiency. We begin with a comprehensive analysis of the Flag-DiT architecture and identify several suboptimal components, which we address by introducing the Next-DiT architecture with 3D RoPE and sandwich normalizations. To enable better resolution extrapolation, we thoroughly compare different context extrapolation methods applied to text-to-image generation with 3D RoPE, and propose Frequency- and Time-Aware Scaled RoPE tailored for diffusion transformers. Additionally, we introduce a sigmoid time discretization schedule for diffusion sampling, which achieves high-quality generation in 5-10 steps combined with higher-order ODE solvers. Thanks to these improvements, Lumina-Next not only improves the basic text-to-image generation but also demonstrates superior resolution extrapolation capabilities as well as multilingual generation using decoder-based LLMs as the text encoder, all in a zero-shot manner. To further validate Lumina-Next as a versatile generative framework, we instantiate it on diverse tasks including visual recognition, multi-views, audio, music, and point cloud generation, showcasing strong performance across these domains. By releasing all codes and model weights at https://github.com/Alpha-VLLM/Lumina-T2X, we aim to advance the development of next-generation generative AI capable of universal modeling. Le Zhuo, Ruoyi Du, Han Xiao 0010, Yangguang Li 0001, Rongjie Huang 0001, Wenze Liu, Fu-Yun Wang, Zhanyu Ma, Zehan Wang 0001, Kaipeng Zhang, Lirui Zhao, Si Liu 0001, Xiangyu Yue 0001, Wanli Ouyang, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
NeurIPS | 20 |
| 2024 | Efficient MAE towards Large-Scale Vision TransformersabstractMasked Autoencoder (MAE) has demonstrated superb pre-training efficiency for vision Transformer, thanks to its partial input paradigm and high mask ratio (0.75). However, MAE often suffers from severe performance drop under higher mask ratios, which hinders its potential toward larger-scale vision Transformers. In this work, we identify that the performance drop is largely attributed to the over-dominance of difficult reconstruction targets, as higher mask ratios lead to more sparse visible patches and fewer visual clues for reconstruction. To mitigate this issue, we design Efficient MAE that introduces a novel Difficulty-Flatten Loss and a decoder masking strategy, enabling a higher mask ratio for more efficient pre-training. The Difficulty-Flatten Loss provides balanced supervision on reconstruction targets of different difficulties, mitigating the performance drop under higher mask ratios effectively. Additionally, the decoder masking strategy discards the most difficult reconstruction targets, which further alleviates the optimization difficulty and accelerates the pre-training clearly. Our proposed Efficient MAE introduces 27% and 30% pre-training runtime accelerations for the ViT-Large and ViT-Huge models, provides valuable insights into MAE’s optimization, and paves the way for larger-scale vision Transformer pre-training. Code and pre-trained models will be released. Han Qiu 0008, Gongjie Zhang, Jiaxing Huang 0001, Peng Gao 0007, Shijian Lu |
WACV | 4 |
| 2024 | POS-BERT: Point cloud one-stage BERT pre-training
Kexue Fu 0001, Peng Gao 0007, Shaolei Liu, Linhao Qu, Longxiang Gao, Manning Wang |
Expert Syst. Appl. | 2 |
| 2024 | CLIP-Adapter: Better Vision-Language Models with Feature Adapters
Peng Gao 0007, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Hongsheng Li 0001, Yu Qiao 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Mimic before Reconstruct: Enhancing Masked Autoencoders with Feature Mimicking
Peng Gao 0007, Renrui Zhang, Rongyao Fang, Hongyang Li 0001, Hongsheng Li 0001, Yu Qiao 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | FeatAug-DETR: Enriching One-to-Many Matching for DETRs With Feature AugmentationabstractOne-to-one matching is a crucial design in DETR- like object detection frameworks. It enables the DETR to perform end-to-end detection. However, it also faces challenges of lacking positive sample supervision and slow convergence speed. Several recent works proposed the one-to-many matching mechanism to accelerate training and boost detection performance. We revisit these methods and model them in a unified format of augmenting the object queries. In this paper, we propose two methods that realize one-to-many matching from a different perspective of augmenting images or image features. The first method is One-to-many Matching via Data Augmentation (denoted asDataAug-DETR). It spatially transforms the images and includes multiple augmented versions of each image in the same training batch. Such a simple augmentation strategy already achieves one-to-many matching and surprisingly improves DETR's performance. The second method is One-to-many matching via Feature Augmentation (denoted asFeatAug-DETR). UnlikeDataAug-DETR, it augments the image features instead of the original images and includes multiple augmented features in the same batch to realize one-to-many matching.FeatAug-DETRsignificantly accelerates DETR training and boosts detection performance while keeping the inference speed unchanged. We conduct extensive experiments to evaluate the effectiveness of the proposed approach on DETR variants, including DAB-DETR, Deformable-DETR, and$\mathcal {H}$-Deformable-DETR. Without extra training data,FeatAug-DETRshortens the training convergence periods of Deformable-DETR [1] to 24 epochs and achieves 58.3 AP on COCOval2017set with Swin-L as the backbone. Rongyao Fang, Peng Gao 0007, Aojun Zhou, Yingjie Cai, Si Liu 0001, Jifeng Dai, Hongsheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Resilient Binary Neural NetworkabstractBinary neural networks (BNNs) have received ever-increasing popularity for their great capability of reducing storage burden as well as quickening inference time. However, there is a severe performance drop compared with {real-valued} networks, due to its intrinsic frequent weight oscillation during training. In this paper, we introduce a Resilient Binary Neural Network (ReBNN) to mitigate the frequent oscillation for better BNNs' training. We identify that the weight oscillation mainly stems from the non-parametric scaling factor. To address this issue, we propose to parameterize the scaling factor and introduce a weighted reconstruction loss to build an adaptive training objective. For the first time, we show that the weight oscillation is controlled by the balanced parameter attached to the reconstruction loss, which provides a theoretical foundation to parameterize it in back propagation. Based on this, we learn our ReBNN by calculating the balanced parameter based on its maximum magnitude, which can effectively mitigate the weight oscillation with a resilient training process. Extensive experiments are conducted upon various network models, such as ResNet and Faster-RCNN for computer vision, as well as BERT for natural language processing. The results demonstrate the overwhelming performance of our ReBNN over prior arts. For example, our ReBNN achieves 66.9% Top-1 accuracy with ResNet-18 backbone on the ImageNet dataset, surpassing existing state-of-the-arts by a significant margin. Our code is open-sourced at https://github.com/SteveTsui/ReBNN. Sheng Xu 0007, Yanjing Li, Teli Ma, Mingbao Lin, Hao Dong 0003, Baochang Zhang 0001, Peng Gao 0007, Jinhu Lü 0001 |
AAAI | 7 |
| 2023 | Q-DETR: An Efficient Low-Bit Quantized Detection TransformerabstractThe recent detection transformer (DETR) has advanced object detection, but its application on resource-constrained devices requires massive computation and memory resources. Quantization stands out as a solution by representing the network in low-bit parameters and operations. However, there is a significant performance drop when performing low-bit quantized DETR (Q-DETR) with existing quantization methods. We find that the bottle-necks of Q-DETR come from the query information distortion through our empirical analyses. This paper addresses this problem based on a distribution rectification distillation (DRD). We formulate our DRD as a bi-level optimization problem, which can be derived by generalizing the information bottleneck (IB) principle to the learning of Q-DETR. At the inner level, we conduct a distribution alignment for the queries to maximize the self-information entropy. At the upper level, we introduce a new foreground-aware query matching scheme to effectively transfer the teacher information to distillation-desired features to minimize the conditional information entropy. Extensive experimental results show that our method performs much better than prior arts. For example, the 4-bit Q-DETR can theoretically accelerate DETR with ResNet-50 backbone by 6.6× and achieve 39.4% AP, with only 2.6% performance gaps than its real-valued counterpart on the COCO dataset11Code: https://github.com/SteveTsui/Q-DETR. Sheng Xu 0007, Yanjing Li, Mingbao Lin, Peng Gao 0007, Guodong Guo, Jinhu Lü 0001, Baochang Zhang 0001 |
CVPR | 4 |
| 2023 | Stare at What You See: Masked Image Modeling without ReconstructionabstractMasked Autoencoders (MAE) have been prevailing paradigms for large-scale vision representation pretraining. By reconstructing masked image patches from a small portion of visible image regions, MAE forces the model to infer semantic correlation within an image. Recently, some approaches apply semantic-rich teacher models to extract image features as the reconstruction target, leading to better performance. However, unlike the low-level features such as pixel values, we argue the features extracted by powerful teacher models already encode rich semantic correlation across regions in an intact image. This raises one question: is reconstruction necessary in Masked Image Modeling (MIM) with a teacher model? In this paper, we propose an efficient MIM paradigm named MaskAlign. MaskAlign simply learns the consistency of visible patch features extracted by the student model and intact image features extracted by the teacher model. To further advance the performance and tackle the problem of input inconsistency between the student and teacher model, we propose a Dynamic Alignment (DA) module to apply learnable alignment. Our experimental results demonstrate that masked modeling does not lose effectiveness even without reconstruction on masked regions. Combined with Dynamic Alignment, MaskAlign can achieve state-of-the-art performance with much higher efficiency. Code and models will be available at https://github.com/OpenPerceptionX/maskalign. Hongwei Xue, Peng Gao 0007, Hongyang Li 0001, Yu Qiao 0001, Houqiang Li, Jiebo Luo 0001 |
CVPR | 2 |
| 2023 | Prompt, Generate, Then Cache: Cascade of Foundation Models Makes Strong Few-Shot LearnersabstractVisual recognition in low-data regimes requires deep neural networks to learn generalized representations from limited training samples. Recently, CLIP-based methods have shown promising few-shot performance benefited from the contrastive language-image pre-training. We then question, if the more diverse pre-training knowledge can be cascaded to further assist few-shot representation learning. In this paper, we propose CaFo, a Cascade of Foundation models that incorporates diverse prior knowledge of various pretraining paradigms for better few-shot learning. Our CaFo incorporates CLIP's language-contrastive knowledge, DINO's vision-contrastive knowledge, DALL-E's vision-generative knowledge, and GPT-3's language-generative knowledge. Specifically, CaFo works by ‘Prompt, Generate, then Cache’. Firstly, we leverage GPT-3 to produce textual inputs for prompting CLIP with rich downstream linguistic semantics. Then, we generate synthetic images via DALL-E to expand the few-shot training data without any manpower. At last, we introduce a learnable cache model to adaptively blend the predictions from CLIP and DINO. By such collaboration, CaFo can fully unleash the potential of different pre-training methods and unify them to perform state-of-the-art for few-shot classification. Code is available at https://github.com/ZrrSkywalker/CaFo. Renrui Zhang, Xiangfei Hu, Siyuan Huang 0004, Hanqiu Deng, Yu Qiao 0001, Peng Gao 0007, Hongsheng Li 0001 |
CVPR | 7 |
| 2023 | Learning 3D Representations from 2D Pre-Trained Models via Image-to-Point Masked AutoencodersabstractPre-training by numerous image data has become defacto for robust 2D representations. In contrast, due to the expensive data processing, a paucity of 3D datasets severely hinders the learning for high-quality 3D features. In this paper, we propose an alternative to obtain superior 3D representations from 2D pre-trained models via Image-to-Point Masked Autoencoders, named as I2P-MAE. By self-supervised pre-training, we leverage the well learned 2D knowledge to guide 3D masked autoencoding, which reconstructs the masked point tokens with an encoder-decoder architecture. Specifically, we first utilize off-the-shelf 2D models to extract the multi-view visual features of the input point cloud, and then conduct two types of image-to-point learning schemes. For one, we introduce a 2D-guided masking strategy that maintains semantically important point tokens to be visible. Compared to random masking, the network can better concentrate on significant 3D structures with key spatial cues. For another, we enforce these visible tokens to reconstruct multi-view 2D features after the decoder. This enables the network to effectively inherit high-level 2D semantics for discriminative 3D modeling. Aided by our image-to-point pre-training, the frozen I2P-MAE, without any fine-tuning, achieves 93.4% accuracy for linear SVM on ModelNet40, competitive to existing fully trained methods. By further fine-tuning on on ScanObjectNN's hardest split, I2P-MAE attains the state-of-the-art 90.11% accuracy, +3.68% to the second-best, demonstrating superior transferable capacity. Code is available at https://github.com/ZrrSkywalker/I2P-MAE. Renrui Zhang, Liuhui Wang, Yu Qiao 0001, Peng Gao 0007, Hongsheng Li 0001 |
CVPR | 4 |
| 2023 | Starting from Non-Parametric Networks for 3D Point Cloud AnalysisabstractWe present a Non-parametric Network for 3D point cloud analysis, Point-NN, which consists of purely non-learnable components: farthest point sampling (FPS), k-nearest neighbors (k-NN), and pooling operations, with trigonometric functions. Surprisingly, it performs well on various 3D tasks, requiring no parameters or training, and even surpasses existing fully trained models. Starting from this basic non-parametric model, we propose two extensions. First, Point-NN can serve as a base architectural framework to construct Parametric Networks by simply inserting linear layers on top. Given the superior non-parametric foundation, the derived Point-PN exhibits a high performance-efficiency tradeoff with only a few learnable parameters. Second, Point-NN can be regarded as a plug-and-play module for the already trained 3D models during inference. Point-NN captures the complementary geometric knowledge and enhances existing methods for different 3D benchmarks without retraining. We hope our work may cast a light on the community for understanding 3D point clouds with non-parametric methods. Code is available at https://github.com/ZrrSkywalker/Point-NN. Renrui Zhang, Liuhui Wang, Yali Wang 0001, Peng Gao 0007, Hongsheng Li 0001, Jianbo Shi |
CVPR | 4 |
| 2023 | MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionabstractMonocular 3D object detection has long been a challenging task in autonomous driving. Most existing methods follow conventional 2D detectors to first localize object centers, and then predict 3D attributes by neighboring features. However, only using local visual features is insufficient to understand the scene-level 3D spatial structures and ignores the long-range inter-object depth relations. In this paper, we introduce the first DETR framework for Monocular DEtection with a depth-guided TRansformer, named MonoDETR. We modify the vanilla transformer to be depth-aware and guide the whole detection process by contextual depth cues. Specifically, concurrent to the visual encoder that captures object appearances, we introduce to predict a foreground depth map, and specialize a depth encoder to extract non-local depth embeddings. Then, we formulate 3D object candidates as learnable queries and propose a depth-guided decoder to conduct object-scene depth interactions. In this way, each object query estimates its 3D attributes adaptively from the depth-guided regions on the image and is no longer constrained to local visual features. On KITTI benchmark with monocular images as input, MonoDETR achieves state-of-the-art performance and requires no extra dense depth annotations. Besides, our depth-guided modules can also be plug-and-play to enhance multi-view 3D object detectors on nuScenes dataset, demonstrating our superior generalization capacity. Code is available at https://github.com/ZrrSkywalker/MonoDETR. Renrui Zhang, Han Qiu 0010, Ziteng Cui, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
ICCV | 8 |
| 2023 | SparseMAE: Sparse Training Meets Masked AutoencodersabstractMasked Autoencoders (MAE) and its variants have proven to be effective for pretraining large-scale Vision Transformers (ViTs). However, small-scale models do not benefit from the pretraining mechanisms due to limited capacity. Sparse training is a method of transferring representations from large models to small ones by pruning unimportant parameters. However, naively combining MAE finetuning with sparse training make the network task-specific, resulting in the loss of task-agnostic knowledge, which is crucial for model generalization. In this paper, we aim to reduce model complexity from large vision transformers pretrained by MAE with assistant of sparse training. We summarize various sparse training methods to prune large vision transformers during MAE pretraining and finetuning stages, and discuss their shortcomings. To improve learning both task-agnostic and task-specific knowledge, we propose SparseMAE, a novel two-stage sparse training method that includes sparse pretraining and sparse finetuning. In sparse pretraining, we dynamically prune a small-scale sub-network from a ViT-Base. During finetuning, the sparse sub-network adaptively changes its topology connections under the task-agnostic knowledge of the full model. Extensive experimental results demonstrate the effectiveness of our method and its superiority on small-scale vision transformers. Code will be available at https://github.com/aojunzz/SparseMAE. Aojun Zhou, Yang Li 0116, Zipeng Qin, Junting Pan, Renrui Zhang, Rui Zhao 0001, Peng Gao 0007, Hongsheng Li 0001 |
ICCV | 8 |
| 2023 | PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningabstractLarge-scale pre-trained models have shown promising open-world performance for both vision and language tasks. However, their transferred capacity on 3D point clouds is still limited and only constrained to the classification task. In this paper, we first collaborate CLIP and GPT to be a unified 3D open-world learner, named as Point-CLIP V2, which fully unleashes their potential for zero-shot 3D classification, segmentation, and detection. To better align 3D data with the pre-trained language knowledge, Point-CLIP V2 contains two key designs. For the visual end, we prompt CLIP via a shape projection module to generate more realistic depth maps, narrowing the domain gap between projected point clouds with natural images. For the textual end, we prompt the GPT model to generate 3D-specific text as the input of CLIP’s textual encoder. Without any training in 3D domains, our approach significantly surpasses PointCLIP by +42.90%, +40.44%, and +28.75% accuracy on three datasets for zero-shot 3D classification. On top of that, V2 can be extended to few-shot 3D classification, zero-shot 3D part segmentation, and 3D object detection in a simple manner, demonstrating our generalization ability for unified 3D open-world learning. Code is available at https://github.com/yangyangyang127/PointCLIP_V2. Renrui Zhang, Bowei He, Ziyao Zeng, Zipeng Qin, Shanghang Zhang, Peng Gao 0007 |
ICCV | 8 |
| 2023 | Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior RefinementabstractThe popularity of Contrastive Language-Image Pretraining (CLIP) has propelled its application to diverse downstream vision tasks. To improve its capacity on downstream tasks, few-shot learning has become a widelya-dopted technique. However, existing methods either exhibit limited performance or suffer from excessive learnable parameters. In this paper, we propose APE, an Adaptive Prior rEfinement method for CLIP’s pre-trained knowledge, which achieves superior accuracy with high computational efficiency. Via a prior refinement module, we analyze the inter-class disparity in the downstream data and decouple the domain-specific knowledge from the CLIP-extracted cache model. On top of that, we introduce two model variants, a training-free APE and a training-required APE-T. We explore the trilateral affinities between the test image, prior cache model, and textual representations, and only enable a lightweight category-residual module to be trained. For the average accuracy over 11 benchmarks, both APE and APE-T attain state-of-the-art and respectively outperform the second-best by +1.59% and +1.99% under 16 shots with ×30 less learnable parameters. Code is available at https://github.com/yangyangyang127/APE. Renrui Zhang, Bowei He, Aojun Zhou, Dong Wang 0004, Bin Zhao 0001, Peng Gao 0007 |
ICCV | 7 |
| 2023 | Auxiliary Modality Learning with Generalized Curriculum DistillationabstractDriven by the need from real-world applications, Auxiliary Modality Learning (AML) offers the possibility to utilize more information from auxiliary data in training, while only requiring data from one or fewer modalities in test, to save the overall computational cost and reduce the amount of input data for inferencing. In this work, we formally define ``Auxiliary Modality Learning'' (AML), systematically classify types of auxiliary modality (in visual computing) and architectures for AML, and analyze their performance. We also analyze the conditions under which AML works well from the optimization and data distribution perspectives. To guide various choices to achieve optimal performance using AML, we propose a novel method to assist in choosing the best auxiliary modality and estimating an upper bound performance before executing AML. In addition, we propose a new AML method using generalized curriculum distillation to enable more effective curriculum learning. Our method achieves the best performance compared to other SOTA methods. Xijun Wang 0002, Peng Gao 0007, Ming C. Lin |
ICML | 3 |
| 2023 | Hybrid Transformer Network for Change Detection Under Self-Supervised PretrainingabstractThis paper presents a Siamese network architecture based on a multi-scale hybrid convolution-Transformer (CTUNet) for Change Detection (CD) in a pair of co-registered optical remote sensing images. Different form CD frameworks based on convolution neural networks (CNNs) and pure Transformer networks, this method combines a convolution-Transformer hybrid encoder with a multi-scale change information extraction decoder in a Siamese network architecture. It overcomes the inherent limitations of CNN and Transformer and effectively integrates the multi-scale information required for accurate CD. To learn better discriminative representations from various scales, we propose a masked auto-encoder scheme (CTMAE) to adapt to building targets with varying morphological scales, further unleashing the potential of CTUNet. Experiments on two CD datasets show that the proposed self-supervised pre-trained hybrid convolution-Transformer CTUNet architecture achieves better CD performance than previous methods. Yongjing Cui, Yin Zhuang, Shan Dong, Peng Gao 0007, He Chen 0004, Liang Chen 0004 |
IGARSS | 5 |
| 2023 | Visual, Spatial, Geometric-Preserved Place Recognition for Cross-View and Cross-Modal Collaborative PerceptionabstractPlace recognition plays an important role in multi-robot collaborative perception, such as aerial-ground search and rescue, in order to identify the same place they have visited. Recently, approaches based on semantics showed the promising performance to address cross-view and cross-modal challenges in place recognition, which can be further categorized as graph-based and geometric-based methods. However, both methods have shortcomings, including ignoring geometric cues and affecting by large non-overlapped regions between observations. In this paper, we introduce a novel approach that integrates semantic graph matching and distance fields (DF) matching for cross-view and cross-modal place recognition. Our method uses a graph representation to encode visual-spatial cues of semantics and uses a set of class-wise DFs to encode geometric cues of a scene. Then, we formulate place recognition as a two-step matching problem. We first perform semantic graph matching to identify the correspondence of semantic objects. Then, we estimate the overlapped regions based on the identified correspondences and further align these regions to compute their geometric-based DF similarity. Finally, we integrate graph-based similarity and geometry-based DF similarity to match places. We evaluate our approach over two public benchmark datasets, including KITTI and AirSim. Compared with the previous methods, our approach achieves around 10% improvement in ground-ground place recognition in KITTI and 35% improvement in aerial-ground place recognition in AirSim. Peng Gao 0007, Jing Liang 0006, Sanghyun Son 0003, Ming C. Lin |
IROS | 1 |
| 2023 | SUG: Single-dataset Unified Generalization for 3D Point Cloud ClassificationabstractAlthough Domain Generalization (DG) problem has been fast-growing in the 2D image tasks, its exploration on 3D point cloud data is still insufficient and challenged by more complex and uncertain cross-domain variances with uneven inter-class modality distribution. In this paper, different from previous 2D DG works, we focus on the 3D DG problem and propose a Single-dataset Unified Generalization (SUG) framework that only leverages a single source dataset to alleviate the unforeseen domain differences faced by a well-trained source model. Specifically, we first design a Multi-grained Sub-domain Alignment (MSA) method, which can constrain the learned representations to be domain-agnostic and discriminative, by performing a multi-grained feature alignment process between the splitted sub-domains from the single source dataset. Then, a Sample-level Domain-aware Attention (SDA) strategy is presented, which can selectively enhance easy-to-adapt samples from different sub-domains according to the sample-level inter-domain distance to avoid the negative transfer. Experiments demonstrate that our SUG can boost the generalization ability for unseen target domains, even outperforming the existing unsupervised domain adaptation methods that have to access extensive target domain data. Siyuan Huang 0004, Bo Zhang 0069, Botian Shi, Hongsheng Li 0001, Yikang Li 0002, Peng Gao 0007 |
ACM Multimedia | 6 |
| 2023 | UniFormer: Unifying Convolution and Self-Attention for Visual RecognitionabstractIt is a challenging task to learn discriminative representation from images and videos, due to large local redundancy and complex global dependency in these visual data. Convolution neural networks (CNNs) and vision transformers (ViTs) have been two dominant frameworks in the past few years. Though CNNs can efficiently decrease local redundancy by convolution within a small neighborhood, the limited receptive field makes it hard to capture global dependency. Alternatively, ViTs can effectively capture long-range dependency via self-attention, while blind similarity comparisons among all the tokens lead to high redundancy. To resolve these problems, we propose a novel Unified transFormer (UniFormer), which can seamlessly integrate the merits of convolution and self-attention in a concise transformer format. Different from the typical transformer blocks, the relation aggregators in our UniFormer block are equipped with local and global token affinity respectively in shallow and deep layers, allowing tackling both redundancy and dependency for efficient and effective representation learning. Finally, we flexibly stack our blocks into a new powerful backbone, and adopt it for various vision tasks from image to video domain, from classification to dense prediction. Without any extra training data, our UniFormer achieves 86.3 top-1 accuracy on ImageNet-1 K classification task. With only ImageNet-1 K pre-training, it can simply achieve state-of-the-art performance in a broad range of downstream tasks. It obtains 82.9/84.8 top-1 accuracy on Kinetics-400/600, 60.9/71.2 top-1 accuracy on Something-Something V1/V2 video classification tasks, 53.8 box AP and 46.4 mask AP on COCO object detection task, 50.8 mIoU on ADE20 K semantic segmentation task, and 77.4 AP on COCO pose estimation task. Moreover, we build an efficient UniFormer with a concise hourglass design of token shrinking and recovering, which achieves 2-4[Formula: see text] higher throughput than the recent lightweight models. Kunchang Li 0002, Yali Wang 0001, Junhao Zhang 0001, Peng Gao 0007, Guanglu Song, Yu Liu 0015, Hongsheng Li 0001, Yu Qiao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Hybrid token transformer for deep face recognition
Weicong Su, Yali Wang 0001, Kunchang Li 0002, Peng Gao 0007, Yu Qiao 0001 |
Pattern Recognit. | 4 |
| 2022 | You Only Need 90K Parameters to Adapt Light: a Light Weight Transformer for Image Enhancement and Exposure Correction
Ziteng Cui, Kunchang Li 0002, Lin Gu 0003, Shenghan Su, Peng Gao 0007, Zhengkai Jiang 0001, Yu Qiao 0001, Tatsuya Harada |
BMVC | 5 |
| 2022 | Unleashing the Potential of Vision-Language Models for Long-Tailed Visual Recognition
Teli Ma, Shijie Geng, Mengmeng Wang 0005, Sheng Xu 0007, Hongsheng Li 0001, Baochang Zhang 0001, Peng Gao 0007, Yu Qiao 0001 |
BMVC | 7 |
| 2022 | PointCLIP: Point Cloud Understanding by CLIPabstractRecently, zero-shot and few-shot learning via Contrastive Vision-Language Pre-training (CLIP) have shown inspirational performance on 2D visual recognition, which learns to match images with their corresponding texts in open-vocabulary settings. However, it remains under explored that whether CLIP, pre-trained by large-scale image-text pairs in 2D, can be generalized to 3D recognition. In this paper, we identify such a setting is feasible by proposing PointCLIP, which conducts alignment between CLIP-encoded point clouds and 3D category texts. Specifically, we encode a point cloud by projecting it onto multi-view depth maps and aggregate the view-wise zero-shot prediction in an end-to-end manner, which achieves efficient knowledge transfer from 2D to 3D. We further design an inter-view adapter to better extract the global feature and adaptively fuse the 3D few-shot knowledge into CLIP pre-trained in 2D. By just fine-tuning the adapter under few-shot settings, the performance of PointCLIP could be largely improved. In addition, we observe the knowledge complementary property between PointCLIP and classical 3D-supervised networks. Via simple ensemble during inference, PointCLIP contributes to favorable performance enhancement over state-of-the-art 3D networks. Therefore, PointCLIP is a promising alternative for effective 3D point cloud understanding under low data regime with marginal resource cost. We conduct thorough experiments on Model-NetlO, ModelNet40 and ScanObjectNN to demonstrate the effectiveness of PointCLIP. Code is available at https://github.com/ZrrSkywalker/PointCLIP. Renrui Zhang, Wei Zhang 0394, Kunchang Li 0002, Xupeng Miao, Bin Cui 0001, Yu Qiao 0001, Peng Gao 0007, Hongsheng Li 0001 |
CVPR | 8 |
| 2022 | Prototypical Contrast Adaptation for Domain Adaptive Semantic Segmentation
Zhengkai Jiang 0001, Yuxi Li 0009, Ceyuan Yang, Peng Gao 0007, Yabiao Wang, Ying Tai, Chengjie Wang 0001 |
ECCV (34) | 4 |
| 2022 | Frozen CLIP Models are Efficient Video Learners
Shijie Geng, Renrui Zhang, Peng Gao 0007, Gerard de Melo, Xiaogang Wang 0001, Jifeng Dai, Yu Qiao 0001, Hongsheng Li 0001 |
ECCV (35) | 4 |
| 2022 | Recurrent Bilinear Optimization for Binary Neural Networks
Sheng Xu 0007, Yanjing Li, Teli Ma, Baochang Zhang 0001, Peng Gao 0007, Yu Qiao 0001, Jinhu Lü 0001, Guodong Guo |
ECCV (24) | 6 |
| 2022 | IDa-Det: An Information Discrepancy-Aware Distillation for 1-Bit Detectors
Sheng Xu 0007, Yanjing Li, Bohan Zeng, Teli Ma, Baochang Zhang 0001, Xianbin Cao 0001, Peng Gao 0007, Jinhu Lü 0001 |
ECCV (11) | 7 |
| 2022 | Tip-Adapter: Training-Free Adaption of CLIP for Few-Shot Classification
Renrui Zhang, Wei Zhang 0394, Rongyao Fang, Peng Gao 0007, Kunchang Li 0002, Jifeng Dai, Yu Qiao 0001, Hongsheng Li 0001 |
ECCV (35) | 4 |
| 2022 | Audio-Visual Scene-Aware Dialog and Reasoning Using Audio-Visual Transformers with Joint Student-Teacher LearningabstractIn previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges (DSTC7, DSTC8). In these challenges, the best-performing systems relied heavily on human-generated descriptions of the video content, which were available in the datasets but would be unavailable in real-world applications. To promote further advancements for real-world applications, we proposed a third AVSD challenge, at DSTC10, with two modifications: 1) the human-created description is unavailable at inference time, and 2) systems must demonstrate temporal reasoning by finding evidence from the video to support each answer. This paper introduces the new task that includes temporal reasoning and our new extension of the AVSD dataset for DSTC10, for which we collected human-generated temporal reasoning data. We also introduce a baseline system built using an AV-transformer, which we released along with the new dataset. Finally, this paper introduces a new system that extends our baseline system with attentional multimodal fusion, joint student-teacher learning (JSTL), and model combination techniques, achieving state-of-the-art performances on the AVSD datasets for DSTC7, DSTC8, and DSTC10. We also propose two temporal reasoning methods for AVSD: one attention-based, and one based on a time-domain region proposal network. Ankit P. Shah, Shijie Geng, Peng Gao 0007, Anoop Cherian, Takaaki Hori, Tim K. Marks, Jonathan Le Roux, Chiori Hori |
ICASSP | 3 |
| 2022 | UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation Learning
Kunchang Li 0002, Yali Wang 0001, Peng Gao 0007, Guanglu Song, Yu Liu 0015, Hongsheng Li 0001, Yu Qiao 0001 |
ICLR | 3 |
| 2022 | Adaptive Local Context Embedding for Small Vehicle Detection from Aerial Optical Remote Sensing ImagesabstractSmall vehicle detection is one of the remaining challenging task because the ambiguous appearance is against complex background interference. Consequently, in order to improve the performance of small vehicle detection from aerial optical remote sensing images, a novel adaptive local context (ALC) embedding way is designed and further introduced into an anchor free detection manner which is called ALC-Net, and in ALC-Net, it can adaptively set up the effective local context feature to improve keypoint description of small vehicles and boost the detection performance without adding extra prior information. Finally, several experiments are carried out on two widely used datasets (e.g., UCAS-AOD [1] and VEDAI [2]) and the results indicate that the proposed ALC-Net can exhibit the competitive small vehicle detection performance than other detectors. Shanjunyu Liu, Yin Zhuang, Hao Dong 0003, Peng Gao 0007, Guanqun Wang, Tong Zhang 0028, Liang Chen 0004, He Chen 0004, LianLin Li |
IGARSS | 4 |
| 2022 | MCMAE: Masked Convolution Meets Masked AutoencodersabstractVision Transformers (ViT) become widely-adopted architectures for various vision tasks. Masked auto-encoding for feature pretraining and multi-scale hybrid convolution-transformer architectures can further unleash the potentials of ViT, leading to state-of-the-art performances on image classification, detection and semantic segmentation. In this paper, our MCMAE framework demonstrates that multi-scale hybrid convolution-transformer can learn more discriminative representations via the mask auto-encoding scheme. However, directly using the original masking strategy leads to the heavy computational cost and pretraining-finetuning discrepancy. To tackle the issue, we adopt the masked convolution to prevent information leakage in the convolution blocks. A simple block-wise masking strategy is proposed to ensure computational efficiency. We also propose to more directly supervise the multi-scale features of the encoder to boost multi-scale features. Based on our pretrained MCMAE models, MCMAE-Base improves ImageNet-1K finetuning accuracy by 1.4% compared with MAE-Base. On object detection, MCMAE-Base finetuned for only 25 epochs surpasses MAE-Base fined-tuned for 100 epochs by 2.9% box AP and 2.2% mask AP respectively. Code and pretrained models are available at \url{https://github.com/Alpha-VL/ConvMAE}. Peng Gao 0007, Teli Ma, Hongsheng Li 0001, Jifeng Dai, Yu Qiao 0001 |
NeurIPS | 1 |
| 2022 | Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerabstractThe large pre-trained vision transformers (ViTs) have demonstrated remarkable performance on various visual tasks, but suffer from expensive computational and memory cost problems when deployed on resource-constrained devices. Among the powerful compression approaches, quantization extremely reduces the computation and memory consumption by low-bit parameters and bit-wise operations. However, low-bit ViTs remain largely unexplored and usually suffer from a significant performance drop compared with the real-valued counterparts. In this work, through extensive empirical analysis, we first identify the bottleneck for severe performance drop comes from the information distortion of the low-bit quantized self-attention map. We then develop an information rectification module (IRM) and a distribution guided distillation (DGD) scheme for fully quantized vision transformers (Q-ViT) to effectively eliminate such distortion, leading to a fully quantized ViTs. We evaluate our methods on popular DeiT and Swin backbones. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, our Q-ViT can theoretically accelerates the ViT-S by 6.14x and achieves about 80.9% Top-1 accuracy, even surpassing the full-precision counterpart by 1.0% on ImageNet dataset. Our codes and models are attached on https://github.com/YanjingLi0202/Q-ViT Yanjing Li, Sheng Xu 0007, Baochang Zhang 0001, Xianbin Cao 0001, Peng Gao 0007, Guodong Guo |
NeurIPS | 5 |
| 2022 | Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingabstractMasked Autoencoders (MAE) have shown great potentials in self-supervised pre-training for language and 2D image transformers. However, it still remains an open question on how to exploit masked autoencoding for learning 3D representations of irregular point clouds. In this paper, we propose Point-M2AE, a strong Multi-scale MAE pre-training framework for hierarchical self-supervised learning of 3D point clouds. Unlike the standard transformer in MAE, we modify the encoder and decoder into pyramid architectures to progressively model spatial geometries and capture both fine-grained and high-level semantics of 3D shapes. For the encoder that downsamples point tokens by stages, we design a multi-scale masking strategy to generate consistent visible regions across scales, and adopt a local spatial self-attention mechanism during fine-tuning to focus on neighboring patterns. By multi-scale token propagation, the lightweight decoder gradually upsamples point tokens with complementary skip connections from the encoder, which further promotes the reconstruction from a global-to-local perspective. Extensive experiments demonstrate the state-of-the-art performance of Point-M2AE for 3D representation learning. With a frozen encoder after pre-training, Point-M2AE achieves 92.9% accuracy for linear SVM on ModelNet40, even surpassing some fully trained methods. By fine-tuning on downstream tasks, Point-M2AE achieves 86.43% accuracy on ScanObjectNN, +3.36% to the second-best, and largely benefits the few-shot classification, part segmentation and 3D object detection with the hierarchical pre-training scheme. Code is available at https://github.com/ZrrSkywalker/Point-M2AE. Renrui Zhang, Peng Gao 0007, Rongyao Fang, Bin Zhao 0001, Dong Wang 0004, Yu Qiao 0001, Hongsheng Li 0001 |
NeurIPS | 3 |
| 2021 | Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled TransformersabstractGiven an input video, its associated audio, and a brief caption, the audio-visual scene aware dialog (AVSD) task requires an agent to indulge in a question-answer dialog with a human about the audio-visual content. This task thus poses a challenging multi-modal representation learning and reasoning scenario, advancements into which could influence several human-machine interaction applications. To solve this task, we introduce a semantics-controlled multi-modal shuffled Transformer reasoning framework, consisting of a sequence of Transformer modules, each taking a modality as input and producing representations conditioned on the input question. Our proposed Transformer variant uses a shuffling scheme on their multi-head outputs, demonstrating better regularization. To encode fine-grained visual information, we present a novel dynamic scene graph representation learning pipeline that consists of an intra-frame reasoning layer producing spatio-semantic graph representations for every frame, and an inter-frame aggregation module capturing temporal cues. Our entire pipeline is trained end-to-end. We present experiments on the benchmark AVSD dataset, both on answer generation and selection tasks. Our results demonstrate state-of-the-art performances on all evaluation metrics. Shijie Geng, Peng Gao 0007, Moitreya Chatterjee, Chiori Hori, Jonathan Le Roux, Hongsheng Li 0001, Anoop Cherian |
AAAI | 2 |
| 2021 | End-to-End Object Detection with Adaptive Clustering Transformer
Minghang Zheng, Peng Gao 0007, Renrui Zhang, Kunchang Li 0002, Hongsheng Li 0001, Hao Dong 0003 |
BMVC | 2 |
| 2021 | Fast Convergence of DETR with Spatially Modulated Co-AttentionabstractThe recently proposed Detection Transformer (DETR) model successfully applies Transformer to objects detection and achieves comparable performance with two-stage object detection frameworks, such as Faster-RCNN. However, DETR suffers from its slow convergence. Training DETR [4] from scratch needs 500 epochs to achieve a high accuracy. To accelerate its convergence, we propose a simple yet effective scheme for improving the DETR framework, namely Spatially Modulated Co-Attention (SMCA) mechanism. The core idea of SMCA is to conduct location-aware co-attention in DETR by constraining co-attention responses to be high near initially estimated bounding box locations. Our proposed SMCA increases DETR’s convergence speed by replacing the original co-attention mechanism in the decoder while keeping other operations in DETR unchanged. Furthermore, by integrating multi-head and scale-selection attention designs into SMCA, our fully-fledged SMCA can achieve better performance compared to DETR with a dilated convolution-based backbone (45.6 mAP at 108 epochs vs. 43.3 mAP at 500 epochs). We perform extensive ablation studies on COCO dataset to validate SMCA. Code is released at https://github.com/gaopengcuhk/SMCA-DETR. Peng Gao 0007, Minghang Zheng, Xiaogang Wang 0001, Jifeng Dai, Hongsheng Li 0001 |
ICCV | 1 |
| 2021 | Dense Contrastive Visual-Linguistic PretrainingabstractInspired by the success of BERT, several multimodal representation learning approaches have been proposed that jointly represent image and text. These approaches achieve superior performance by capturing high-level semantic information from large-scale multimodal pretraining. In particular, LXMERT and UNITER adopt visual region feature regression and label classification as pretext tasks. However, they tend to suffer from the problems of noisy labels and sparse semantic annotations, based on the visual features having been pretrained on a crowdsourced dataset with limited and inconsistent semantic labeling. To overcome these issues, we propose unbiased Dense Contrastive Visual-Linguistic Pretraining (DCVLP), which replaces the region regression and classification with cross-modality region contrastive learning that requires no annotations. Two data augmentation strategies (Mask Perturbation and Intra-Inter-Adversarial Perturbation) are developed to improve the quality of negative samples used in contrastive learning. Overall, DCVLP allows cross-modality dense region contrastive learning in a self-supervised setting independent of any object annotations. We compare our method against prior visual-linguistic pretraining frameworks to validate the superiority of dense contrastive learning on multimodal representation learning. Lei Shi 0002, Kai Shuang, Shijie Geng, Peng Gao 0007, Zuohui Fu, Gerard de Melo, Yunpeng Chen, Sen Su |
ACM Multimedia | 4 |
| 2021 | Container: Context Aggregation NetworksabstractConvolutional neural networks (CNNs) are ubiquitous in computer vision, with a myriad of effective and efficient variations. Recently, Transformers -- originally introduced in natural language processing -- have been increasingly adopted in computer vision. While early adopters continued to employ CNN backbones, the latest networks are end-to-end CNN-free Transformer solutions. A recent surprising finding now shows that a simple MLP based solution without any traditional convolutional or Transformer components can produce effective visual representations. While CNNs, Transformers and MLP-Mixers may be considered as completely disparate architectures, we provide a unified view showing that they are in fact special cases of a more general method to aggregate spatial context in a neural network stack. We present the \model (CONText AggregatIon NEtwoRk), a general-purpose building block for multi-head context aggregation that can exploit long-range interactions \emph{a la} Transformers while still exploiting the inductive bias of the local convolution operation leading to faster convergence speeds, often seen in CNNs. Our \model architecture achieves 82.7 \% Top-1 accuracy on ImageNet using 22M parameters, +2.8 improvement compared with DeiT-Small, and can converge to 79.9 \% Top-1 accuracy in just 200 epochs. In contrast to Transformer-based methods that do not scale well to downstream tasks that rely on larger input image resolutions, our efficient network, named \modellight, can be employed in object detection and instance segmentation networks such as DETR, RetinaNet and Mask-RCNN to obtain an impressive detection mAP of 38.9, 43.8, 45.1 and mask mAP of 41.3, providing large improvements of 6.6, 7.3, 6.9 and 6.6 pts respectively, compared to a ResNet-50 backbone with a comparable compute and parameter size. Our method also achieves promising results on self-supervised learning compared to DeiT on the DINO framework. Code is released at https://github.com/allenai/container. Peng Gao 0007, Jiasen Lu, Hongsheng Li 0001, Roozbeh Mottaghi, Aniruddha Kembhavi |
NeurIPS | 1 |
| 2021 | Dual-stream Network for Visual RecognitionabstractTransformers with remarkable global representation capacities achieve competitive results for visual tasks, but fail to consider high-level local pattern information in input images. In this paper, we present a generic Dual-stream Network (DS-Net) to fully explore the representation capacity of local and global pattern features for image classification. Our DS-Net can simultaneously calculate fine-grained and integrated features and efficiently fuse them. Specifically, we propose an Intra-scale Propagation module to process two different resolutions in each block and an Inter-Scale Alignment module to perform information interaction across features at dual scales. Besides, we also design a Dual-stream FPN (DS-FPN) to further enhance contextual information for downstream dense predictions. Without bells and whistles, the proposed DS-Net outperforms DeiT-Small by 2.4\% in terms of top-1 accuracy on ImageNet-1k and achieves state-of-the-art performance over other Vision Transformers and ResNets. For object detection and instance segmentation, DS-Net-Small respectively outperforms ResNet-50 by 6.4\% and 5.5 \% in terms of mAP on MSCOCO 2017, and surpasses the previous state-of-the-art scheme, which significantly demonstrates its potential to be a general backbone in vision tasks. The code will be released soon. Mingyuan Mao, Peng Gao 0007, Renrui Zhang, Honghui Zheng, Teli Ma, Errui Ding, Baochang Zhang 0001, Shumin Han |
NeurIPS | 2 |
| 2020 | Learning Where to Focus for Efficient Video Object Detection
Zhengkai Jiang 0001, Yu Liu 0015, Ceyuan Yang, Jihao Liu, Peng Gao 0007, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
ECCV (16) | 5 |
| 2020 | Multi-Layer Content Interaction Through Quaternion Product for Visual Question AnsweringabstractMulti-modality fusion technologies have greatly improved the performance of neural network-based Video Description/Caption, Visual Question Answering (VQA) and Audio Visual Scene-aware Dialog (AVSD) over the recent years. Most previous approaches only explore the last layers of multiple layer feature fusion while omitting the importance of intermediate layers. To solve the issue for the intermediate layers, we propose an efficient Quaternion Block Network (QBN) to learn interaction not only for the last layer but also for all intermediate layers simultaneously. In our proposed QBN, we use the holistic text features to guide the update of visual features. In the meantime, Hamilton quaternion products can efficiently perform information flow from higher layers to lower layers for both visual and text modalities. The evaluation results show our QBN improve the performance on VQA 2.0, furthermore surpasses the approach using large scale BERT or visual BERT pre-trained models. Extensive ablation study has been carried out to examine the influence of each proposed module in this study. Lei Shi 0002, Shijie Geng, Kai Shuang, Chiori Hori, Songxiang Liu, Peng Gao 0007, Sen Su |
ICASSP | 6 |
| 2019 | Video Object Detection with Locally-Weighted Deformable NeighborsabstractDeep convolutional neural networks have achieved great success on various image recognition tasks. However, it is nontrivial to transfer the existing networks to video due to the fact that most of them are developed for static image. Frame-byframe processing is suboptimal because temporal information that is vital for video understanding is totally abandoned. Furthermore, frame-by-frame processing is slow and inefficient, which can hinder the practical usage. In this paper, we propose LWDN (Locally-Weighted Deformable Neighbors) for video object detection without utilizing time-consuming optical flow extraction networks. LWDN can latently align the high-level features between keyframes and keyframes or nonkeyframes. Inspired by (Zhu et al. 2017a) and (Hetang et al. 2017) who propose to aggregate features between keyframes and keyframes, we adopt brain-inspired memory mechanism to propagate and update the memory feature from keyframes to keyframes. We call this process Memory-Guided Propagation. With such a memory mechanism, the discriminative ability of features in keyframes and non-keyframes are both enhanced, which helps to improve the detection accuracy. Extensive experiments on VID dataset demonstrate that our method achieves superior performance in a speed and accuracy trade-off, i.e., 76.3% on the challenging VID dataset while maintaining 20fps in speed on Titan X GPU. Zhengkai Jiang 0001, Peng Gao 0007, Chaoxu Guo, Qian Zhang 0009, Shiming Xiang, Chunhong Pan |
AAAI | 2 |
| 2019 | Dynamic Fusion With Intra- and Inter-Modality Attention Flow for Visual Question AnsweringabstractLearning effective fusion of multi-modality features is at the heart of visual question answering. We propose a novel method of dynamically fuse multi-modal features with intra- and inter-modality information flow, which alternatively pass dynamic information between and across the visual and language modalities. It can robustly capture the high-level interactions between language and vision domains, thus significantly improves the performance of visual question answering. We also show that, the proposed dynamic intra modality attention flow conditioned on the other modality can dynamically modulate the intra-modality attention of the current modality, which is vital for multimodality feature fusion. Experimental evaluations on the VQA 2.0 dataset show that the proposed method achieves the state-of-the-art VQA performance. Extensive ablation studies are carried out for the comprehensive analysis of the proposed method. Peng Gao 0007, Zhengkai Jiang 0001, Haoxuan You, Pan Lu, Steven C. H. Hoi, Xiaogang Wang 0001, Hongsheng Li 0001 |
CVPR | 1 |
| 2019 | Multi-Modality Latent Interaction Network for Visual Question AnsweringabstractExploiting relationships between visual regions and question words have achieved great success in learning multi-modality features for Visual Question Answering (VQA). However, we argue that existing methods mostly model relations between individual visual regions and words, which are not enough to correctly answer the question. From humans' perspective, answering a visual question requires understanding the summarizations of visual and language information. In this paper, we proposed the Multi-modality Latent Interaction module (MLI) to tackle this problem. The proposed module learns the cross-modality relationships between latent visual and language summarizations, which summarize visual regions and question into a small number of latent representations to avoid modeling uninformative individual region-word relations. The cross-modality information between the latent summarizations are propagated to fuse valuable information from both modalities and are used to update the visual and word features. Such MLI modules can be stacked for several stages to model complex and latent relations between the two modalities and achieves highly competitive performance on public VQA benchmarks, VQA v2.0 and TDIUC . In addition, we show that the performance of our methods could be significantly improved by combining with pre-trained language model BERT. Peng Gao 0007, Haoxuan You, Xiaogang Wang 0001, Hongsheng Li 0001 |
ICCV | 1 |
| 2018 | Question-Guided Hybrid Convolution for Visual Question Answering
Peng Gao 0007, Hongsheng Li 0001, Shuang Li 0013, Pan Lu, Yikang Li 0002, Steven C. H. Hoi, Xiaogang Wang 0001 |
ECCV (1) | 1 |