VLDB 2026 Research / reviewers in the wild / expert
Li Yuan 0007
dblp:98/4583-7
· DBLP profile ↗
114ranked-venue papers
9as first author
105since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 90 · 6 first-author · 85 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 7 first-author · 55 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 360Explorer: Exploring 4D Controllable World in Panoramic VideosabstractWe present 360Explorer, a novel approach for generating 4D controllable panoramic videos conditioned on user-provided 3D instructions for exploring and manipulating dynamic worlds. Compared to existing perspective-based methods struggle to address spatial consistency during camera rotation in place, we introduce the panoramic view in controllable video generation models to inherently maintain the view recall consistency. By introducing dynamic point clouds as the 4D scene representations, 360Explorer unifies the modeling of camera transformations and object movements as incomplete renders to describe precise control instructions in 3D worlds. To tackle the data limitation in acquiring multi-viewpoint panoramic videos, we further propose a reverse warping strategy to construct the training dataset on easily accessible monocular panoramic videos. Extensive experiments demonstrate that 360Explorer achieves superior performance in creating 4D controllable panoramic videos with camera transformation and object movements aligned with diverse provided instructions. Xinhua Cheng, Haiyang Zhou, Wangbo Yu, Tanghui Jia, Bin Lin 0014, Yunyang Ge, Li Yuan 0007 |
AAAI | 8 |
| 2026 | Next Patch Prediction for AutoRegressive Visual GenerationabstractAutoregressive models, built based on the Next Token Prediction (NTP) paradigm, show great potential in developing a unified framework that integrates both language and vision tasks. Pioneering works introduce NTP to autoregressive visual generation tasks. In this work, we rethink the NTP for autoregressive image generation and extend it to a novel Next Patch Prediction (NPP) paradigm. Our key idea is to group and aggregate image tokens into patch tokens with higher information density. By using patch tokens as a more compact input sequence, the autoregressive model is trained to predict the next patch, significantly reducing computational costs. To further exploit the natural hierarchical structure of image data, we propose a multi-scale coarse-to-fine patch grouping strategy. With this strategy, the training process begins with a large patch size and ends with vanilla NTP where the patch size is 1x1, thus maintaining the original inference process without modifications. Extensive experiments across a diverse range of model sizes demonstrate that NPP could reduce the training cost to around 0.6 times while improving image generation quality by up to 1.0 FID score on the ImageNet 256x256 generation benchmark. Notably, our method retains the original autoregressive model architecture without introducing additional trainable parameters or specifically designing a custom image tokenizer, offering a flexible and plug-and-play solution for enhancing autoregressive visual generation. Yatian Pang, Peng Jin 0001, Bin Lin 0014, Chaoran Feng 0001, Zhenyu Tang 0004, Liuhan Chen, Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Li Yuan 0007 |
AAAI | 12 |
| 2026 | NeuralGS: Bridging Neural Fields and 3D Gaussian Splatting for Compact 3D Representationsabstract3D Gaussian Splatting (3DGS) achieves impressive quality and rendering speed, but with millions of 3D Gaussians and significant storage and transmission costs. In this paper, we aim to develop a simple yet effective method called NeuralGS that compresses the original 3DGS into a compact representation. Our observation is that neural fields like NeRF can represent complex 3D scenes with Multi-Layer Perceptron (MLP) neural networks using only a few megabytes. Thus, NeuralGS effectively adopts the neural field representation to encode the attributes of 3D Gaussians with MLPs, only requiring a small storage size even for a large-scale scene. To achieve this, we adopt a clustering strategy and fit the Gaussians within each cluster using different tiny MLPs, based on importance scores of Gaussians as fitting weights. We experiment on multiple datasets, achieving a 91$\times$ average model size reduction without harming the visual quality. Zhenyu Tang 0004, Chaoran Feng 0001, Xinhua Cheng, Wangbo Yu, Junwu Zhang, Yuan Liu 0025, Xiaoxiao Long, Wenping Wang 0001, Li Yuan 0007 |
AAAI | 9 |
| 2026 | Look-Back: Implicit Visual Re-focusing in MLLM ReasoningabstractMultimodal Large Language Models (MLLMs) have achieved remarkable progress in multimodal reasoning. However, they often excessively rely on textual information during the later stages of inference, neglecting the crucial integration of visual input. Current methods typically address this by explicitly injecting visual information to guide the reasoning process. In this work, through an analysis of MLLM attention patterns, we made an intriguing observation: with appropriate guidance, MLLMs can spontaneously re-focus their attention on visual inputs during the later stages of reasoning, even without explicit visual injection. This spontaneous shift in focus suggests that MLLMs are intrinsically capable of performing visual fusion reasoning. Building on this insight, we introduce Look-Back, an implicit approach designed to guide MLLMs to look back at visual information in a self-directed manner during reasoning. Look-Back empowers the model to autonomously determine when, where, and how to re-focus on visual inputs, eliminating the need for explicit model-structure constraints or additional input. We demonstrate that Look-Back significantly enhances the model's reasoning and perception capabilities, as evidenced by extensive empirical evaluations on multiple multimodal benchmarks. Yuwei Niu, Bin Lin 0014, Li Yuan 0007 |
AAAI | 6 |
| 2026 | AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety BasinabstractFine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely compromise safety measures. We observe that perturbations orthogonal to the alignment direction—defined by weight differences between aligned (safe) and unaligned models—rapidly compromise model safety. In contrast, updates along the alignment direction largely preserve it, revealing the parameter space as a "narrow safety basin". To address this, we propose AsFT (Anchoring Safety in Fine-Tuning) to maintain safety by explicitly constraining update directions during fine-tuning. By penalizing updates orthogonal to the alignment direction, AsFT effectively constrains the model within the "narrow safety basin," thus preserving its inherent safety. Extensive experiments on multiple datasets and models show that AsFT reduces harmful behaviors by up to 7.60%, improves task performance by 3.44%, and consistently outperforms existing methods across multiple tasks. Qihui Zhang, Yue Huang 0001, Xiaojun Jia, Kun-Peng Ning, Jia-Yu Yao, Jigang Wang, Hailiang Dai, Yibing Song, Li Yuan 0007 |
AAAI | 11 |
| 2026 | SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement LearningabstractPeidong Wang, Zhiming Ma, Xin Dai, YongKang Liu, Shi Feng, Xiaocui Yang, Wenxing Hu, Zhihao Wang, Mingjun Pan, Li Yuan, Daling Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Peidong Wang 0001, Zhiming Ma, Yongkang Liu 0002, Shi Feng 0001, Xiaocui Yang, Wenxing Hu, Mingjun Pan, Li Yuan 0007, Daling Wang |
ACL (1) | 10 |
| 2026 | Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-Based Large Language ModelsabstractVideo-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user inquiries. In pursuit of the ultimate goal of achieving artificial general intelligence, a truly intelligent Video-LLM model should not only see and understand the surroundings, but also possess human-level commonsense, and make well-informed decisions for users. To guide the development of such a model, the establishment of a robust and comprehensive evaluation system becomes crucial. To this end, this paper proposes Video-Bench, a new comprehensive benchmark along with a toolkit specifically designed for evaluating Video-LLMs. The benchmark comprises 10 meticulously crafted tasks, evaluating the capabilities of Video-LLMs across three distinct levels: video-exclusive understanding, prior knowledge-based question-answering, and comprehension and decision-making. In addition, we introduce an automatic toolkit tailored to process model outputs for various tasks, facilitating the calculation of metrics and conveniently generating final scores. We evaluate 9 representative Video-LLMs using Video-Bench. The findings reveal that current Video-LLMs still fall considerably short of achieving human-like comprehension and analysis of real-world video, and offer valuable insights for future research directions. The benchmark and toolkit are available at https://github.com/PKU-YuanGroup/Video-Bench. Munan Ning, Yujia Xie, Bin Lin 0014, Jiaxi Cui, Lu Yuan 0001, Dongdong Chen 0001, Li Yuan 0007 |
Comput. Vis. Media | 8 |
| 2026 | Next-Gen AIGC: a review of multimodal foundation models for text-to-media innovations
Jingru Fan, Jinfa Huang, Jinyuan Fu, Tao Mei 0001, Li Yuan 0007, Jiebo Luo 0001 |
Frontiers Comput. Sci. | 6 |
| 2026 | Sparse Orthogonal Parameters Tuning for Continual Learning
Kun-Peng Ning, Hai-Jian Ke, Jia-Yu Yao, Yonghong Tian 0001, Li Yuan 0007 |
Int. J. Comput. Vis. | 6 |
| 2026 | Evidence Conflict Sampling for Open-set Active Learning
Kun-Peng Ning, Hai-Jian Ke, Jia-Yu Yao, Yonghong Tian 0001, Li Yuan 0007 |
Int. J. Comput. Vis. | 6 |
| 2026 | NOFA++: Tuning-Free NeRF-Based One-Shot Facial Avatar Reconstructionabstract3D facial avatar reconstruction is a fundamental problem in computer vision and graphics with applications in digital humans, virtual reality, and telepresence. Recent neural radiance field (NeRF)-based methods have greatly improved fidelity, yet most remain subject-specific, requiring multi-view images with diverse expressions and extensive test-time finetuning, which limits their generalization to unseen identities. Achieving high-quality reconstruction from a single image is particularly challenging due to missing multi-view supervision and the need to balance efficiency with fidelity. To address these issues, we present NOFA++, a generalizable one-shot framework that reconstructs photorealistic 3D facial avatars from a single input. Our method leverages the generative prior of a pretrained 3D GAN in an encoder–generator pipeline to recover a canonical neural volume, and introduces a coarse-to-fine residual generation strategy to synthesize identity-specific details without per-subject optimization. We further design a deformation field conditioned on identity and expression parameters to model facial dynamics, enabling controllable reenactment from video or audio. Extensive experiments show that NOFA++ surpasses state-of-the-art baselines in both reconstruction fidelity and reenactment quality, while eliminating test-time finetuning and generalizing robustly across unseen subjects. Wangbo Yu, Chaoran Feng 0001, Li Yuan 0007, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Triple-View Knowledge Distillation for Semi-Supervised Semantic SegmentationabstractTo alleviate the expensive human labeling problem, semi-supervised semantic segmentation utilizes a few labeled images along with an abundance of unlabeled images to predict the pixel-level label maps with the same size. Previous methods often rely on co-training with two convolutional networks with the same architecture but different initialization, which fails to capture sufficiently diverse features. This limitation motivates us to employ tri-training and design a triple-view encoder to utilize encoders with different architectures to derive diverse features, while leveraging knowledge distillation to capture complementary semantics among these encoders. Moreover, existing approaches simply concatenate features from both encoder and decoder, and the simple concatenation requires a large memory cost. This inspires us to present a dual-frequency decoder that selects those important features by projecting the spatial-domain features into the frequency domain, where a dual-frequency channel attention mechanism is applied to evaluate the feature importance. Therefore, we propose a Triple-view Knowledge Distillation framework, termed TriKD, for semi-supervised semantic segmentation. It comprises the triple-view encoder and the dual-frequency decoder. Extensive experiments conducted on two benchmarks,i.e., Pascal VOC 2012 and Cityscapes, validate the superiority of our method, achieving a satisfying tradeoff between precision and inference speed. Our code is available at GitHub. Ping Li 0006, Li Yuan 0007, Xianghua Xu, Mingli Song |
IEEE Trans. Multim. | 3 |
| 2026 | MoE-LLaVA: Mixture of Experts for Large Vision-Language ModelsabstractRecently, remarkable progress has been made in scaling up Large Language Models (LLMs) through the use of the sparse Mixture-of-Expert (MoE) layers without significantly increasing computational cost. However, the transition from a pre-trained LLM to a sparse Large Vision-Language Model (LVLM) with MoE remains an open challenge. Directly fine-tuning an LLM to a sparse LVLM often leads to training collapse, characterized by (1) a large modality feature distribution gap and (2) expert load imbalance. This paper proposes a three-stage decoupled weight training process. In the first two stages, the model learns to adapt the LLM to an LVLM. In the third stage, the FFN weights from the second stage are used as lossless initialization for expert weights, effectively constructing a sparse model with a vast number of parameters while maintaining constant computational cost. Through extensive ablation experiments, we derive three empirical guidelines and propose a sparse LVLM termedMoE-LLaVA. MoE-LLaVA is a MoE-based sparse LVLM architecture, which uniquely activates only the top-$k$experts through routers during deployment, keeping the remaining experts inactive. Extensive experiments demonstrate that MoE-LLaVA outperforms LLaVA-1.5-7B with an average improvement of 4.6 across nine visual understanding benchmarks. Notably, with only 2.2B active parameters, our MoE-LLaVA shows comparable result with LLaVA-1.5-13B (87.0 vs. 85.9) on POPE benchmark. Our work establishes a baseline for sparse LVLMs and provides empirical guidelines for exploring the sparse LVLMs. Our code is available at:https://github.com/PKU-YuanGroup/MoE-LLaVA. Bin Lin 0014, Zhenyu Tang 0004, Jinfa Huang, Junwu Zhang, Yatian Pang, Peng Jin 0001, Munan Ning, Jiebo Luo 0001, Li Yuan 0007 |
IEEE Trans. Multim. | 10 |
| 2026 | HoloDreamer: Holistic 3D Panoramic Scene Generation From Text Descriptionsabstract3D scene generation is in high demand across various domains, including virtual reality, gaming, and the film industry. Owing to the powerful generative capabilities of text-to-image diffusion models that provide reliable priors, creating 3D scenes using only text prompts has become viable, thereby significantly advancing research in text-driven 3D scene generation. Prevailing methods typically employ the diffusion model to generate an initial local image, followed by iteratively outpainting the local image to gradually generate scenes. Nevertheless, these outpainting-based approaches are prone to producing globally inconsistent results with low completeness, restricting their broader applications. To tackle these problems, we introduce HoloDreamer, a framework that begins by generating a high-definition panorama to holistically initialize the full scene, and leverages 3D Gaussian Splatting (3D-GS) for rapid 3D scene reconstruction, thereby facilitating the creation of view-consistent and fully enclosed 3D scenes. Specifically, we propose Stylized Equirectangular Panorama Generation, a pipeline that combines multiple diffusion models to enable stylized and detailed equirectangular panorama generation from complex text prompts. Subsequently, Enhanced Two-Stage Panorama Reconstruction is introduced, conducting a two-stage optimization of 3D-GS to inpaint the missing region and enhance the integrity of the scene. Comprehensive experiments demonstrated that our method outperforms prior works in terms of overall visual consistency and harmony, as well as reconstruction quality and rendering robustness when generating fully enclosed scenes. Haiyang Zhou, Xinhua Cheng, Wangbo Yu, Yonghong Tian 0001, Li Yuan 0007 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger ScenesabstractCompared to frame-based methods, computational neuromorphic imaging using event cameras offers significant advantages, such as minimal motion blur, enhanced temporal resolution, and high dynamic range. The multi-view consistency of Neural Radiance Fields combined with the unique benefits of event cameras, has spurred recent research into reconstructing NeRF from data captured by moving event cameras. While showing impressive performance, existing methods rely on ideal conditions with the availability of uniform and high-quality event sequences and accurate camera poses, and mainly focus on the object level reconstruction, thus limiting their practical applications. In this work, we propose AE-NeRF to address the challenges of learning event-based NeRF from non-ideal conditions, including non-uniform event sequences, noisy poses, and various scales of scenes. Our method exploits the density of event streams and jointly learn a pose correction module with an event-based NeRF (e-NeRF) framework for robust 3D reconstruction from inaccurate camera poses. To generalize to larger scenes, we propose hierarchical event distillation with a proposal e-NeRF network and a vanilla e-NeRF network to resample and refine the reconstruction process. We further propose an event reconstruction loss and a temporal loss to improve the view consistency of the reconstructed scene. We established a comprehensive benchmark that includes large-scale scenes to simulate practical non-ideal conditions, incorporating both synthetic and challenging real-world event datasets. The experimental results show that our method achieves a new state-of-the-art in event-based 3D reconstruction. Chaoran Feng 0001, Wangbo Yu, Xinhua Cheng, Zhenyu Tang 0004, Junwu Zhang, Li Yuan 0007, Yonghong Tian 0001 |
AAAI | 6 |
| 2025 | Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction CycleabstractRecent 3D large reconstruction models typically employ a two-stage process, including first generate multi-view images by a multi-view diffusion model, and then utilize a feed-forward model to reconstruct images to 3D content. However, multi-view diffusion models often produce low-quality and inconsistent images, adversely affecting the quality of the final 3D reconstruction. To address this issue, we propose a unified 3D generation framework called Cycle3D, which cyclically utilizes a 2D diffusion-based generation module and a feed-forward 3D reconstruction module during the multi-step diffusion process. Concretely, 2D diffusion model is applied for generating high-quality texture, and the reconstruction model guarantees multi-view consistency. Moreover, 2D diffusion model can further control the generated content and inject reference-view information for unseen views, thereby enhancing the diversity and texture consistency of 3D generation during the denoising process. Extensive experiments demonstrate the superior ability of our method to create 3D content with high-quality and consistency compared with state-of-the-art baselines. Zhenyu Tang 0004, Junwu Zhang, Xinhua Cheng, Wangbo Yu, Chaoran Feng 0001, Yatian Pang, Bin Lin 0014, Li Yuan 0007 |
AAAI | 8 |
| 2025 | Is Parameter Collision Hindering Continual Learning in LLMs?abstractLarge Language Models (LLMs) often suffer from catastrophic forgetting when learning multiple tasks sequentially, making continual learning (CL) essential for their dynamic deployment. Existing state-of-the-art (SOTA) methods, such as O-LoRA, typically focus on constructing orthogonality tasks to decouple parameter interdependence from various domains.In this paper, we reveal that building non-collision parameters is a more critical factor in addressing CL challenges. Our theoretical and experimental analyses demonstrate that non-collision parameters provide better task orthogonality, which is a sufficient but unnecessary condition. Furthermore, knowledge from multiple domains will be preserved in non-collision parameter subspaces, making it more difficult to forget previously seen data. Leveraging this insight, we propose Non-collision Low-Rank Adaptation (N-LoRA), a simple yet effective approach leveraging low collision rates to enhance CL in LLMs. Experimental results on multiple CL benchmarks indicate that N-LoRA achieves superior performance (+2.9%), higher task orthogonality (×4.1times), and lower parameter collision (×58.1times) than SOTA methods. Kun-Peng Ning, Jia-Yu Yao, Yonghong Tian 0001, Yi-Bing Song, Li Yuan 0007 |
COLING | 7 |
| 2025 | RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene TexturingabstractIndoor scene texture synthesis has garnered significant interest due to its important potential applications in virtual reality, digital media and creative arts. Existing diffusion-model-based researches either rely on per-view inpainting techniques, which are plagued by severe cross-view inconsistencies and conspicuous seams, or adopt optimization-based approaches that involve substantial computational overhead. In this work, we present RoomPainter, a frame-work that seamlessly integrates efficiency and consistency to achieve high-fidelity texturing of indoor scenes. The core of RoomPainter features a zero-shot technique that effectively adapts a 2D diffusion model for 3D-consistent texture synthesis, along with a two-stage generation strategy that ensures both global and local consistency. Specifically, we introduce Attention-Guided Multi-View Integrated Sampling (MVIS) combined with a neighbor-integrated attention mechanism for zero-shot texture map generation. Using the MVIS, we firstly generate texture map for the entire room to ensure global consistency, then adopt its variant, namely Attention-Guided Multi-View Integrated Repaint Sampling (MVRS) to repaint individual instances within the room, thereby further enhancing local consistency and addressing the occlusion problem. Experiments demonstrate that RoomPainter achieves superior performance for indoor scene texture synthesis in visual quality, global consistency and generation efficiency. Zhipeng Huang 0001, Wangbo Yu, Xinhua Cheng, ChengShu Zhao, Yunyang Ge, Mingyi Guo, Li Yuan 0007, Yonghong Tian 0001 |
CVPR | 7 |
| 2025 | WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion ModelabstractVideo Variational Autoencoder (VAE) encodes videos into a low-dimensional latent space, becoming a key component of most Latent Video Diffusion Models (LVDMs) to reduce model training costs. However, as the resolution and duration of generated videos increase, the encoding cost of Video VAEs becomes a limiting bottleneck in training LVDMs. Moreover, the block-wise inference method adopted by most LVDMs can lead to discontinuities of latent space when processing long-duration videos. The key to addressing the computational bottleneck lies in decomposing videos into distinct components and efficiently encoding the critical information. Wavelet transform can decompose videos into multiple frequency-domain components and improve the efficiency significantly, we thus propose Wavelet Flow VAE (WF-VAE), an autoencoder that leverages multi-level wavelet transform to facilitate low-frequency energy flow into latent representation. Furthermore, we introduce a method called Causal Cache, which maintains the integrity of latent space during block-wise inference. Compared to state-of-the-art video VAEs, WF-VAE demonstrates superior performance in both PSNR and LPIPS metrics, achieving 2× higher throughput and 4× lower memory consumption while maintaining competitive reconstruction quality. Our code and models are available at https://github.com/PKU-YuanGroup/WF-VAE. Zongjian Li, Bin Lin 0014, Liuhan Chen, Xinhua Cheng, Shenghai Yuan 0002, Li Yuan 0007 |
CVPR | 7 |
| 2025 | Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter TuningabstractThree key challenges hinder the development of current deepfake video detection: (1) Temporal features can be complex and diverse: how can we identify general temporal artifacts to enhance model generalization? (2) Spatiotemporal models are proven to lean heavily on one type of forgery artifact and ignore the other (e.g., learning spatial only): how can we ensure balanced learning from both? (3) Videos are naturally resource-intensive: how can we tackle efficiency without compromising accuracy? This paper attempts to tackle the three challenges jointly. First, inspired by the notable generality of using image-level blending data for image forgery detection, we investigate whether and how video-level blending can be effective in video. We then perform a thorough analysis and identify a previously underexplored temporal forgery artifact: Facial Feature Drift (FFD), which commonly exists across different deep-fakes. To reproduce FFD, we then propose a novel Video-level Blending data (VB), which is implemented by blending the original image and its warped version frame-by-frame, serving as a hard negative sample to mine more general artifacts. Second, we carefully design a lightweight Spatiotemporal Adapter (StA) to equip a pretrained image model with the ability to capture both spatial and temporal features jointly and efficiently. StA is designed with two-stream 3D-Conv with varying kernel sizes, allowing it to process spatial and temporal features separately. Extensive experiments validate the effectiveness of our methods. Zhiyuan Yan 0002, Yandan Zhao, Shen Chen 0004, Mingyi Guo, Xinghe Fu, Taiping Yao, Shouhong Ding, Yunsheng Wu, Li Yuan 0007 |
CVPR | 9 |
| 2025 | Identity-Preserving Text-to-Video Generation by Frequency DecompositionabstractIdentity-Preserving text-to-video (IPT2V) generation aims to create high-fidelity videos with consistent human identity. It is an important task in video generation but remains an open problem for generative models. This paper pushes the technical frontier of IPT2V in two directions that have not been resolved in the literature: (1) A tuning-free pipeline without tedious case-by-case finetuning, and (2) A frequency-aware heuristic identity-preserving Diffusion Transformer (DiT)-based control scheme. To achieve these goals, we propose ConsisID, a tuning-free DiT-based controllable IPT2V model to keep human-identity consistent in the generated video. Inspired by prior findings in frequency analysis of vision/diffusion transformers, it employs identity-control signals base on frequency domain, since facial features can be decomposed into low-frequency global features (e.g., profile, proportions) and high-frequency intrinsic features (e.g., identity markers that remain unaffected by pose changes). Extensive experiments demonstrate that our frequency-aware heuristic scheme provides an optimal control solution for DiT-based models, making strides towards more effective IPT2V. Shenghai Yuan 0002, Jinfa Huang, Xianyi He, Yunyang Ge, Yujun Shi, Liuhan Chen, Jiebo Luo 0001, Li Yuan 0007 |
CVPR | 8 |
| 2025 | UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model EvaluationabstractMultimodal Large Language Models (MLLMs) have emerged to tackle the challenges of Visual Question Answering (VQA), sparking a new research focus on conducting objective evaluations of these models. Existing evaluation methods face limitations due to the significant human workload required to design Q&A pairs for visual images, which inherently restricts the scale and scope of evaluations. Although automated MLLM-as-judge approaches attempt to reduce the human workload through automatic evaluations, they often introduce biases. To address these problems, we propose an Unsupervised Peer review MLLM Evaluation framework. It utilizes only image data, allowing models to automatically generate questions and conduct peer review assessments of answers from other models, effectively alleviating the reliance on human workload. Additionally, we introduce the vision-language scoring system to mitigate the bias issues, which focuses on three aspects: (i) response correctness; (ii) visual understanding and reasoning; and (iii) image-text correlation. Experimental results demonstrate that UPME achieves a Pearson correlation of 0.944 with human evaluations on the MMstar dataset and 0.814 on the ScienceQA dataset, indicating that our framework closely aligns with human-designed benchmarks and inherent human preferences. Qihui Zhang, Munan Ning, Zheyuan Liu 0012, Yue Huang 0001, Yanbo Wang 0005, Jiayi Ye, Yibing Song, Li Yuan 0007 |
CVPR | 10 |
| 2025 | Spiking Transformer with Spatial-Temporal Spiking Self-AttentionabstractSpiking Neural Networks are celebrated for energy efficiency and biological plausibility. Building on Spiking Self-Attention (SSA), Spiking Transformers are extensively studied due to their exceptional performance. However, SSA focuses solely on spatial dimension at each time step, overlooking the crucial features across temporal dimension. To address this, we propose the Spatial-Temporal Spiking Self-Attention (STSSA), a spike-driven mechanism that leverages both spatial and temporal information with negligible additional computational overhead. Specifically, we extract the Representative Spiking Temporal Tokens (RSTT) and apply temporal window masking to the RSTT. These tokens are inserted between the Query and Key to integrate temporal features. Furthermore, we design a Multi-dimensional Learnable Scaling Factor (MLSF) to adapt to STSSA. Our results consistently demonstrate that STSSA outperforms SSA across extensive experiments on sequential, neuromorphic, and static datasets. Notably, STSSA achieves performance improvements of 5.7% and 2.9% over SSA on Sequential CIFAR-100 and CIFAR-10DVS, respectively. STSSA provides a powerful alternative within the family of Spiking Self-Attention mechanisms. Zhaokun Zhou, Li Yuan 0007, Yuesheng Zhu |
ICASSP | 4 |
| 2025 | DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D PosesabstractIn this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and user-friendly manner. Concretely, baseline methods relying on only 2D pose guidance lack the cues of 3D information, leading to suboptimal results, while methods using 3D representation as guidance achieve higher quality but involve a cumbersome and time-intensive process. To address these limitations, DreamDance enriches 3D geometry cues from 2D poses by introducing an efficient diffusion model, enabling high-quality human image animation with various guidance. Our key insight is that human images naturally exhibit multiple levels of correlation, progressing from coarse skeleton poses to fine-grained geometry cues, and further from these geometry cues to explicit appearance details. Capturing such correlations could enrich the guidance signals, facilitating intra-frame coherency and inter-frame consistency. Specifically, we construct the TikTok-Dance5K dataset, comprising 5K high-quality dance videos with detailed frame annotations, including human pose, depth, and normal maps. Next, we introduce a Mutually Aligned Geometry Diffusion Model to generate fine-grained depth and normal maps for enriched guidance. Finally, a Cross-domain Controller incorporates multi-level guidance to animate human images effectively with a video diffusion model. Extensive experiments demonstrate that our method achieves state-of-the-art performance in animating human images. Yatian Pang, Bin Bin, Mingzhe Zheng, Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Li Yuan 0007 |
ICCV | 8 |
| 2025 | LlaVA-CoT: Let Vision Language Models Reason Step-By-Step
Guowei Xu 0001, Peng Jin 0001, Ziang Wu, Yibing Song, Lichao Sun 0001, Li Yuan 0007 |
ICCV | 7 |
| 2025 | Evagaussians: Event Stream Assisted Gaussian Splatting from Blurry Images
Wangbo Yu, Chaoran Feng 0001, Jianing Li 0001, Jiye Tang, Jiashu Yang, Zhenyu Tang 0004, Meng Cao 0002, Xu Jia 0012, Li Yuan 0007, Yonghong Tian 0001 |
ICCV | 10 |
| 2025 | Epona: Autoregressive Diffusion World Model for Autonomous DrivingabstractDiffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is because conventional video diffusion models rely on global joint distribution modeling of fixed-length frame sequences rather than sequentially constructing localized distributions at each timestep. In this work, we propose Epona, an autoregressive diffusion world model that enables localized spatiotemporal distribution modeling through two key innovations: 1) Decoupled spatiotemporal factorization that separates temporal dynamics modeling from fine-grained future world generation, and 2) Modular trajectory and video prediction that seamlessly integrate motion planning with visual modeling in an end-to-end framework. Our architecture enables high-resolution, long-duration generation while introducing a novel chain-of-forward training strategy to address error accumulation in autoregressive loops. Experimental results demonstrate state-of-the-art performance with 7.4\% FVD improvement and minutes longer prediction duration compared to prior works. The learned world model further serves as a real-time motion planner, outperforming strong end-to-end planners on NAVSIM benchmarks. Code will be publicly available at \href{https://github.com/Kevin-thu/Epona/}{https://github.com/Kevin-thu/Epona/}. Kaiwen Zhang 0015, Zhenyu Tang 0004, Xiaotao Hu, Xingang Pan, Yuan Liu 0025, Li Yuan 0007, Qian Zhang 0001, Xiao-Xiao Long, Xun Cao, Wei Yin 0006 |
ICCV | 8 |
| 2025 | MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation ExpertsabstractIn this work, we aim to simultaneously enhance the effectiveness and efficiency of Mixture-of-Experts (MoE) methods. To achieve this, we propose MoE++, a general and heterogeneous MoE framework that integrates both Feed-Forward Network (FFN) and zero-computation experts. Specifically, we introduce three types of zero-computation experts: the zero expert, copy expert, and constant expert, which correspond to discard, skip, and replace operations, respectively. This design offers three key advantages: (i) **Low Computing Overhead**: Unlike the uniform mixing mechanism for all tokens within vanilla MoE, MoE++ allows each token to engage with a dynamic number of FFNs, be adjusted by constant vectors, or even skip the MoE layer entirely. (ii) **High Performance**: By enabling simple tokens to utilize fewer FFN experts, MoE++ allows more experts to focus on challenging tokens, thereby unlocking greater performance potential than vanilla MoE. (iii) **Deployment Friendly**: Given that zero-computation experts have negligible parameters, we can deploy all zero-computation experts on each GPU, eliminating the significant communication overhead and expert load imbalance associated with FFN experts distributed across different GPUs. Moreover, we leverage gating residuals, enabling each token to consider the pathway taken in the previous layer when selecting the appropriate experts. Extensive experimental results demonstrate that MoE++ achieves better performance while delivering 1.1$\sim$2.1$\times$ expert forward throughput compared to a vanilla MoE model of the same size, which lays a solid foundation for developing advanced and efficient MoE-related models. Peng Jin 0001, Li Yuan 0007, Shuicheng Yan |
ICLR | 3 |
| 2025 | PiCO: Peer Review in LLMs based on Consistency OptimizationabstractExisting large language models (LLMs) evaluation methods typically focus on testing the performance on some closed-environment and domain-specific benchmarks with human annotations. In this paper, we explore a novel unsupervised evaluation direction, utilizing peer-review mechanisms to measure LLMs automatically without any human feedback. In this setting, both open-source and closed-source LLMs lie in the same environment, capable of answering unlabeled questions and evaluating each other, where each LLM’s response score is jointly determined by other anonymous ones. During this process, we found that those answers that are more recognized by other ``reviewers'' (models) usually come from LLMs with stronger abilities, while these models can also evaluate others' answers more accurately. We formalize it as a consistency assumption, i.e., the ability and score of the model usually have consistency. We exploit this to optimize each model's confidence, thereby re-ranking the LLMs to be closer to human rankings. We perform experiments on multiple datasets with standard rank-based metrics, validating the effectiveness of the proposed approach. Kun-Peng Ning, Jia-Yu Yao, Zhen-Hui Liu, Yonghong Tian 0001, Yibing Song, Li Yuan 0007 |
ICLR | 8 |
| 2025 | OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion ModelabstractVariational Autoencoder (VAE), compressing videos into latent representations, is a crucial preceding component of Latent Video Diffusion Models (LVDMs). However, most LVDMs utilize 2D image VAE, which only compresses video spatially. This will lead to temporally redundant representations, reducing the efficiency of LVDMs. To eliminate this issue, we propose an omni-dimension compression VAE, named OD-VAE, which can temporally and spatially compress videos based on 3D-Causal-CNN architecture. To obtain a better trade-off between video reconstruction quality and compression speed, we further introduce and analyze four model variants of OD-VAE. In addition, a novel initialization method is designed to train our OD-VAE more efficiently, and a novel inference strategy is proposed to enable OD-VAE to handle videos of arbitrary length with limited GPU memory. Comprehensive experiments on video reconstruction and LVDM-based video generation demonstrate the effectiveness and efficiency of our proposed methods. The source code and models are available at here. Liuhan Chen, Zongjian Li, Bin Lin 0014, Qian Wang 0062, Shenghai Yuan 0002, Xinhua Cheng, Li Yuan 0007 |
ICME | 9 |
| 2025 | OpenDUN: To Discover Unknown Number of Visual CategoriesabstractOpen-Set methods have relaxed the underlying assumption made by most image recognition studies that all samples in the test and training datasets belong to the same classes by considering only a part of classes are known in training dataset. However, most of these approaches require a known or predefined number of novel classes, which is often not the case in real applications. In this study, we aim at a more difficult but practical scenario, where the number of novel classes is unknown. By merging the unlabeled samples into clusters instead of directly assigning categorical labels to them, the proposed end-to-end framework can simultaneously estimate the number of novel classes and learn the appropriate division of unlabeled samples. In addition, a cluster scattering strategy is introduced such that the erroneous merging can be alleviated. Comprehensive experiments on three benchmark datasets are conducted to demonstrate the superiority of the proposed method in both the estimation of the novel class number and the classification of unlabeled samples. Sik Chit Wu, Munan Ning, Dong Wei 0004, Yefeng Zheng 0001, Donghuan Lu, Li Yuan 0007 |
ICME | 6 |
| 2025 | MoH: Multi-Head Attention as Mixture-of-Head AttentionabstractIn this work, we upgrade the multi-head attention mechanism, the core of the Transformer model, to reduce computational costs while maintaining or surpassing the previous accuracy level. We show that multi-head attention can be expressed in the summation form. Drawing on the insight that not all attention heads hold equal significance, we propose Mixture-of-Head attention (MoH), a new architecture that treats attention heads as experts in the Mixture-of-Experts (MoE) mechanism. MoH has two significant advantages: First, MoH enables each token to select the appropriate attention heads, enhancing inference efficiency without compromising accuracy or increasing the number of parameters. Second, MoH replaces the standard summation in multi-head attention with a weighted summation, introducing flexibility to the attention mechanism and unlocking extra performance potential. Extensive experiments on ViT, DiT, and LLMs demonstrate that MoH outperforms multi-head attention by using only 50%$\sim$90% of the attention heads. Moreover, we demonstrate that pre-trained multi-head attention models, such as LLaMA3-8B, can be further continue-tuned into our MoH models. Notably, MoH-LLaMA3-8B achieves an average accuracy of 64.0% across 14 benchmarks, outperforming LLaMA3-8B by 2.4% by utilizing only 75% of the attention heads. We believe the proposed MoH is a promising alternative to multi-head attention and provides a strong foundation for developing advanced and efficient attention-based models. Peng Jin 0001, Li Yuan 0007, Shuicheng Yan |
ICML | 3 |
| 2025 | Orthogonal Subspace Decomposition for Generalizable AI-Generated Image DetectionabstractDetecting AI-generated images (AIGIs), such as natural images or face images, has become increasingly important yet challenging. In this paper, we start from a new perspective to excavate the reason behind the failure generalization in AIGI detection, named the asymmetry phenomenon, where a naively trained detector tends to favor overfitting to the limited and monotonous fake patterns, causing the feature space to become highly constrained and low-ranked, which is proved seriously limiting the expressivity and generalization. One potential remedy is incorporating the pre-trained knowledge within the vision foundation models (higher-ranked) to expand the feature space, alleviating the model's overfitting to fake. To this end, we employ Singular Value Decomposition (SVD) to decompose the original feature space into two orthogonal subspaces. By freezing the principal components and adapting only the remained components, we preserve the pre-trained knowledge while learning fake patterns. Compared to existing full-parameters and LoRA-based tuning methods, we explicitly ensure orthogonality, enabling the higher rank of the whole feature space, effectively minimizing overfitting and enhancing generalization. We finally identify a crucial insight: our method implicitly learns a vital prior that fakes are actually derived from the real, indicating a hierarchical relationship rather than independence. Modeling this prior, we believe, is essential for achieving superior generalization. Our codes are publicly available at https://github.com/YZY-stack/Effort-AIGI-Detection. Zhiyuan Yan 0002, Jiangming Wang, Peng Jin 0001, Ke-Yue Zhang, Chengchun Liu, Shen Chen 0004, Taiping Yao, Shouhong Ding, Baoyuan Wu, Li Yuan 0007 |
ICML | 10 |
| 2025 | E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event CamerasabstractNovel view synthesis and 4D reconstruction techniques predominantly rely on RGB cameras, thereby inheriting inherent limitations such as the dependence on adequate lighting, susceptibility to motion blur, and a limited dynamic range. Event cameras, offering advantages of low power, high temporal resolution and high dynamic range, have brought a new perspective to addressing the scene reconstruction challenges in high-speed motion and low-light scenes. To this end, we propose E-4DGS, the first event-driven dynamic Gaussian Splatting approach, for novel view synthesis from multi-view event streams with fast-moving cameras. Specifically, we introduce an event-based initialization scheme to ensure stable training and propose event-adaptive slicing splatting for time-aware reconstruction. Additionally, we employ intensity importance pruning to eliminate floating artifacts and enhance 3D consistency, while incorporating an adaptive contrast threshold for more precise optimization. We design a synthetic multi-view camera setup with six moving event cameras surrounding the object in a 360-degree configuration and provide a benchmark multi-view event stream dataset that captures challenging motion scenarios. Our approach outperforms both event-only and event-RGB fusion baselines and paves the way for the exploration of multi-view event-based reconstruction as a novel approach for rapid scene capture. Chaoran Feng 0001, Zhenyu Tang 0004, Wangbo Yu, Yatian Pang, Yian Zhao, Jianbin Zhao, Li Yuan 0007, Yonghong Tian 0001 |
ACM Multimedia | 7 |
| 2025 | HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
Haiyang Zhou, Wangbo Yu, Jiawen Guan, Xinhua Cheng, Yonghong Tian 0001, Li Yuan 0007 |
ACM Multimedia | 6 |
| 2025 | Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical OperationsabstractWhile large language models (LLMs) with Chain-of-Thought (CoT) reasoning excel in mathematics and coding, their potential for systematic reasoning in chemistry, a domain demanding rigorous structural analysis for real-world tasks like drug design and reaction engineering, remains untapped. Current benchmarks focus on simple knowledge retrieval, neglecting step-by-step reasoning required for complex tasks such as molecular optimization and reaction prediction. To address this, we introduce ChemCoTBench, a reasoning framework that bridges molecular structure understanding with arithmetic-inspired operations, including addition, deletion, and substitution, to formalize chemical problem-solving into transparent, step-by-step workflows. By treating molecular transformations as modular "chemical operations", the framework enables slow-thinking reasoning, mirroring the logic of mathematical proofs while grounding solutions in real-world chemical constraints. We evaluate models on two high-impact tasks: Molecular Property Optimization and Chemical Reaction Prediction. These tasks mirror real-world challenges while providing structured evaluability. We further provide ChemCoTDataset, a pioneering 22,000-instance chemical reasoning dataset with expert-annotated chains of thought to facilitate LLM fine-tuning. By providing annotated trainable datasets, a reasoning taxonomy, and baseline evaluations, our work bridges the gap between abstract reasoning methods and practical chemical discovery, establishing a foundation for advancing LLMs as tools for AI-driven scientific innovation. Hao Li 0073, He Cao, Bin Feng 0001, Daniel Shao, Robert Tang, Zhiyuan Yan 0002, Yonghong Tian 0001, Li Yuan 0007, Yu Li 0003 |
NeurIPS | 8 |
| 2025 | GS2E: Gaussian Splatting is an Effective Data Generator for Event Stream GenerationabstractWe introduce GS2E (Gaussian Splatting to Event Generation), a large-scale synthetic event dataset designed for high-fidelity event vision tasks, captured from real-world sparse multi-view RGB images. Existing event datasets are often synthesized from dense RGB videos, which typically suffer from limited viewpoint diversity and geometric inconsistency, or rely on expensive, hard-to-scale hardware setups. GS2E addresses these limitations by first reconstructing photorealistic static scenes using 3D Gaussian Splatting, followed by a novel, physically-informed event simulation pipeline. This pipeline integrates adaptive trajectory interpolation with physically-consistent event contrast threshold modeling. As a result, it generates temporally dense and geometrically consistent event streams under diverse motion and lighting conditions, while maintaining strong alignment with the underlying scene structure. Experimental results on event-based 3D reconstruction highlight GS2E’s superior generalization capabilities and its practical value as a benchmark for advancing event vision research. Chaoran Feng 0001, Zhenyu Tang 0004, Kaiyuan Deng, Wangbo Yu, Yonghong Tian 0001, Li Yuan 0007 |
NeurIPS | 7 |
| 2025 | CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-stepabstractCurrent text-to-image (T2I) generation models struggle to align spatial composition with the input text, especially in complex scenes.
Even layout-based approaches yield suboptimal spatial control, as their generation process is decoupled from layout planning, making it difficult to refine the layout during synthesis.
We present CoT-Diff, a framework that brings step-by-step CoT-style reasoning into T2I generation by tightly integrating Multimodal Large Language Model (MLLM)-driven 3D layout planning with the diffusion process.
CoT-Diff enables layout-aware reasoning inline within a single diffusion round: at each denoising step, the MLLM evaluates intermediate predictions, dynamically updates the 3D scene layout, and continuously guides the generation process.
The updated layout is converted into semantic conditions and depth maps, which are fused into the diffusion model via a condition-aware attention mechanism, enabling precise spatial control and semantic injection.
Experiments on 3D Scene benchmarks show that CoT-Diff significantly improves spatial alignment and compositional fidelity, and outperforms the state-of-the-art method by 34.7% in complex scene spatial accuracy, thereby validating the effectiveness of this entangled generation paradigm. Zheyuan Liu 0012, Munan Ning, Qihui Zhang, Yiwei Yang 0007, Yibing Song, Fan Wang 0019, Li Yuan 0007 |
NeurIPS | 11 |
| 2025 | ImgEdit: A Unified Image Editing Dataset and BenchmarkabstractRecent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quality data and insufficient benchmarks.To overcome these limitations, we introduce ImgEdit, a large-scale, high-quality image-editing dataset comprising one million carefully curated edit pairs, which contain both novel and complex single-turn edits, as well as challenging multi-turn tasks.To ensure the data quality, we employ a multi-stage pipeline that integrates a cutting-edge vision-language model, a detection model, a segmentation model, alongside task-specific in-painting procedures and strict post-processing. ImgEdit surpasses existing datasets in both task novelty and data quality.Using ImgEdit, we train ImgEdit-E1, an editing model using Vision Language Model to process the reference image and editing prompt, which outperforms existing open-source models on multiple tasks, highlighting the value of ImgEdit and model design.For comprehensive evaluation, we introduce ImgEdit-Bench, a benchmark designed to evaluate image editing performance in terms of instruction adherence, editing quality, and detail preservation.It includes a basic testsuite, a challenging single-turn suite, and a dedicated multi-turn suite. We evaluate both open-source and proprietary models, as well as ImgEdit-E1, providing deep analysis and actionable insights into the current behavior of image-editing models. Xianyi He, Zongjian Li, Bin Lin 0014, Shenghai Yuan 0002, Zhiyuan Yan 0002, Bohan Hou, Li Yuan 0007 |
NeurIPS | 8 |
| 2025 | OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video GenerationabstractSubject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we propose OpenS2V-Nexus, consisting of (i) OpenS2V‑Eval, a fine‑grained benchmark, and (ii) OpenS2V‑5M, a million‑scale dataset.In contrast to existing S2V benchmarks inherited from VBench that focus on global and coarse-grained assessment of generated videos, OpenS2V-Eval focuses on the model's ability to generate subject-consistent videos with natural subject appearance and identity fidelity. For these purposes, OpenS2V-Eval introduces 180 prompts from seven major categories of S2V, which incorporate both real and synthetic test data. Furthermore, to accurately align human preferences with S2V benchmarks, we propose three automatic metrics, NexusScore, NaturalScore and GmeScore, to separately quantify subject consistency, naturalness, and text relevance in generated videos. Building on this, we conduct a comprehensive evaluation of 18 representative S2V models, highlighting their strengths and weaknesses across different content. Moreover, we create the first open-source large-scale S2V generation dataset OpenS2V-5M, which consists of five million high-quality 720P subject-text-video triplets. Specifically, we ensure subject‐information diversity in our dataset by (1) segmenting subjects and building pairing information via cross‐video associations and (2) prompting GPT-4o on raw frames to synthesize multi-view representations. Through OpenS2V-Nexus, we deliver a robust infrastructure to accelerate future S2V generation research. Shenghai Yuan 0002, Xianyi He, Yufan Deng, Jinfa Huang, Bin Lin 0014, Chongyang Ma, Jiebo Luo 0001, Li Yuan 0007 |
NeurIPS | 9 |
| 2025 | Multi-objective Aligned Bidword Generation Model for E-commerce Search AdvertisingabstractThe retrieval system is a crucial module in e-commerce search advertising that matches user queries with ads. The diverse expressions of users often produce massive tail queries that cannot match merchant bidwords, leading to poor retrieval efficiency. Existing methods, such as query log mining and vector matching, fail to optimize relevance, authenticity, and ad revenue of the rewrite. Zhenhui Liu, Chunyuan Yuan, Ming Pang, Li Yuan 0007, Changping Peng, Zhangang Lin, Jingping Shao |
SIGIR | 5 |
| 2025 | TaxDiff: taxonomic-guided diffusion model for protein sequence generation
Zongying Lin, Hao Li 0073, Liuzhenghao Lv, Yu Wang 0027, Bin Lin 0014, Junwu Zhang, Calvin Yu-Chian Chen, Li Yuan 0007, Yonghong Tian 0001 |
Sci. China Inf. Sci. | 9 |
| 2025 | Learning Box Regression and Mask Segmentation Under Long-Tailed Distribution with Gradient Transfusing
Tao Wang 0053, Li Yuan 0007, Xinchao Wang, Jiashi Feng |
Int. J. Comput. Vis. | 2 |
| 2025 | Hierarchical Banzhaf Interaction for General Video-Language Representation LearningabstractMultimodal representation learning, with contrastive learning, plays an important role in the artificial intelligence domain. As an important subfield, video-language representation learning focuses on learning representations using global semantic interactions between pre-defined video-text pairs. However, to enhance and refine such coarse-grained global interactions, more detailed interactions are necessary for fine-grained multimodal learning. In this study, we introduce a creative approach that models video-text as game players using multivariate cooperative game theory to handle uncertainty during fine-grained semantic interactions with diverse granularity, flexible combination, and vague intensity. Specifically, we design the Hierarchical Banzhaf Interaction to simulate the finegrained correspondence between video clips and textual words from hierarchical perspectives. Furthermore, to mitigate the bias in calculations within Banzhaf Interaction, we propose reconstructing the representation through a fusion of single-modal and crossmodal components. This reconstructed representation ensures fine granularity comparable to that of the single-modal representation, while also preserving the adaptive encoding characteristics of cross-modal representation. Additionally, we extend our original structure into a flexible encoder-decoder framework, enabling the model to adapt to various downstream tasks. Extensive experiments on commonly used text-video retrieval, video-question answering, and video captioning benchmarks, with superior performance, validate the effectiveness and generalization of our method. The code is available at https://github.com/jpthu17/HBI. Peng Jin 0001, Hao Li 0073, Li Yuan 0007, Shuicheng Yan, Jie Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | MagicTime: Time-Lapse Video Generation Models as Metamorphic SimulatorsabstractRecent advances in text-to-video generation (T2V) have achieved remarkable success in synthesizing high-quality general videos from textual descriptions. A largely overlooked problem in T2V is that existing models have not adequately encoded physical knowledge of the real world, thus generated videos tend to have limited motion and poor variations. In this paper, we propose MagicTime, a metamorphic time-lapse video generation model, which learns real-world physics knowledge from time-lapse videos and implements metamorphic generation. First, we design a simple yet effective two-stage Magic Adaptive Strategy, encode more physical knowledge from metamorphic videos, and transform pre-trained T2V models to generate metamorphic videos. Second, we introduce a Dynamic Frames Extraction strategy to adapt to metamorphic time-lapse videos, which have a wider variation range and cover dramatic object metamorphic processes, thus embodying more physical knowledge than general videos. Finally, we introduce a Magic Text-Encoder to improve the understanding of metamorphic video prompts. Furthermore, we create a time-lapse video-text dataset called ChronoMagic, specifically curated to unlock the metamorphic video generation ability. Extensive experiments demonstrate the superiority and effectiveness of MagicTime for generating high-quality and dynamic metamorphic videos, suggesting time-lapse video generation is a promising path toward building metamorphic simulators of the physical world. Shenghai Yuan 0002, Jinfa Huang, Yujun Shi, Yongqi Xu, Rui-Jie Zhu 0003, Bin Lin 0014, Xinhua Cheng, Li Yuan 0007, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | Point Tree Transformer for Point Cloud RegistrationabstractPoint cloud registration is a fundamental task in the fields of computer vision and robotics. Recent advancements in transformer-based methods have demonstrated enhanced performance in this domain. However, the standard attention mechanisms employed in these approaches tend to incorporate numerous points of low relevance, and therefore struggle to focus their attention weights on sparse yet meaningful points. This inefficiency leads to limited local structure modeling capabilities and quadratic computational complexity. To overcome these limitations, we propose the Point Tree Transformer (PTT), a novel transformer-based approach for point cloud registration that efficiently extracts comprehensive local and global features while maintaining linear computational complexity. The PTT constructs hierarchical feature trees from point clouds in a coarse-to-dense manner, and introduces a novel Point Tree Attention (PTA) mechanism. This mechanism adheres to the tree structure to facilitate the progressive convergence of attended regions toward salient points. Specifically, each tree layer selectively identifies a subset of relevant points with the highest attention scores, and subsequent layers focus attention on areas of significant relevance, derived from the child points of the selected point set. The feature extraction process additionally incorporates coarse point features that capture high-level semantic information, thus facilitating local structure modeling and the progressive integration of multiscale information. Consequently, the PTA enables the model to focus on essential local structures and extract intricate local information while maintaining linear computational complexity. Extensive experiments conducted on the 3DMatch, ModelNet40, and KITTI datasets demonstrate that our method outperforms state-of-the-art methods in terms of performance. The code for our method is publicly available at https://github.com/CGuangyan-BIT/PTT. Meiling Wang 0002, Guangyan Chen, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Parallel Vertex Diffusion for Unified Visual GroundingabstractUnified visual grounding (UVG) capitalizes on a wealth of task-related knowledge across various grounding tasks via one-shot training, which curtails retraining costs and task-specific architecture design efforts. Vertex generation-based UVG methods achieve this versatility by unified modeling object box and contour prediction and provide a text-powered interface to vast related multi-modal tasks, e.g., visual question answering and captioning. However, these methods typically generate vertexes sequentially through autoregression, which is prone to be trapped in error accumulation and heavy computation, especially for high-dimension sequence generation in complex scenarios. In this paper, we develop Parallel Vertex Diffusion (PVD) based on the parallelizability of diffusion models to accurately and efficiently generate vertexes in a parallel and scalable manner. Since the coordinates fluctuate greatly, it typically encounters slow convergence when training diffusion models without geometry constraints. Therefore, we consummate our PVD by two critical components, i.e., center anchor mechanism and angle summation loss, which serve to normalize coordinates and adopt a differentiable geometry descriptor from the point-in-polygon problem of computational geometry to constrain the overall difference of prediction and label vertexes. These innovative designs empower our PVD to demonstrate its superiority with state-of-the-art performance across various grounding tasks. Zesen Cheng, Kehan Li 0002, Peng Jin 0001, Siheng Li, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
AAAI | 6 |
| 2024 | Towards Better Seach Query Classification with Distribution-Diverse Multi-Expert Knowledge Distillation in JD Ads SearchabstractIn the dynamic landscape of online advertising, decoding user intent remains a pivotal challenge, particularly in the context of query classification. Swift classification models, exemplified by FastText, cater to the demand for real-time responses but encounter limitations in handling intricate queries. Conversely, accuracy-centric models like BERT introduce challenges associated with increased latency. This paper undertakes a nuanced exploration, navigating the delicate balance between efficiency and accuracy. It unveils FastText's latent potential as an 'online dictionary' for historical queries while harnessing the semantic robustness of BERT for novel and complex scenarios. The proposed Distribution-Diverse Multi-Expert (DDME) framework employs multiple teacher models trained from diverse data distributions. Through meticulous data categorization and enrichment, it elevates the classification performance across the query spectrum. Empirical results within the JD ads search system validate the superiority of our proposed approaches. Kun-Peng Ning, Ming Pang, Xiwei Zhao, Changping Peng, Zhangang Lin, Jinghe Hu, Jingping Shao, Li Yuan 0007 |
CIKM | 10 |
| 2024 | Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video UnderstandingabstractLarge language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multi-modal conversations. However, existing methods encounter challenges in effectively handling both image and video understanding, particularly with limited visual tokens. In this work, we introduce Chat-UniVi, a Unified Vision-language model capable of comprehending and engaging in conver-sations involving images and videos through a unified visual representation. Specifically, we employ a set of dynamic visual tokens to uniformly represent images and videos. This representation framework empowers the model to ef-ficiently utilize a limited number of visual tokens to simul-taneously capture the spatial details necessary for images and the comprehensive temporal relationship required for videos. Moreover, we leverage a multi-scale representation, enabling the model to perceive both high-level seman-tic concepts and low-level visual details. Notably, Chat-UniVi is trained on a mixed dataset containing both images and videos, allowing direct application to tasks involving both mediums without requiring any modifications. Exten-sive experimental results demonstrate that Chat- UniVi con-sistently outperforms even existing methods exclusively de-signed for either images or videos. Code is available at https://github.com/PKu-Yuan Group/Chat-UniVi. Peng Jin 0001, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, Li Yuan 0007 |
CVPR | 5 |
| 2024 | SynSP: Synergy of Smoothness and Precision in Pose Sequences RefinementabstractPredicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily concentrated on striking a balance between the two objectives, i.e., smoothness and precision, while optimizing the predicted pose sequences. However, it has come to our attention that the tension between these two objectives can provide additional quality cues about the predicted pose sequences. These cues, in turn, are able to aid the network in optimizing lower-quality poses. To leverage this quality information, we propose a motion refinement network, termed SynSP, to achieve a Synergy of Smoothness and Precision in the sequence refinement tasks. Moreover, SynSP can also address multi-view poses of one person simultaneously, fixing inaccuracies in predicted poses through heightened attention to similar poses from other views, thereby amplifying the resultant quality cues and overall performance. Compared with previous methods, SynSP benefits from both pose quality and multi-view information with a much shorter input sequence length, achieving state-of-the-art results among four challenging datasets involving 2D, 3D, and SMPL pose representations in both single-view and multi-view scenes. Github code: https://github.com/InvertedForest/SynSP. Tao Wang 0011, Lei Jin 0003, Zheng Wang 0007, Jianshu Li, Liang Li 0003, Fang Zhao 0006, Yu Cheng 0009, Li Yuan 0007, Junliang Xing, Jian Zhao 0006 |
CVPR | 8 |
| 2024 | Regressor-Segmenter Mutual Prompt Learning for Crowd CountingabstractCrowd counting has achieved significant progress by training regressors to predict instance positions. In heavily crowded scenarios, however, regressors are challenged by uncontrollable annotation variance, which causes density map bias and context information inaccuracy. In this study, we propose mutual prompt learning (mPrompt), which leverages a regressor and a segmenter as guidance for each other, solving bias and inaccuracy caused by annotation variance while distinguishing foreground from back-ground. In specific, mPrompt leverages point annotations to tune the segmenter and predict pseudo head masks in a way of point prompt learning. It then uses the predicted segmentation masks, which serve as spatial constraint, to rectify biased point annotations as context prompt learning. mPrompt defines a way of mutual information maximization from prompt learning, mitigating the impact of annotation variance while improving model accuracy. Experiments show that mPrompt significantly reduces the Mean Aver-age Error (MAE), demonstrating the potential to be general framework for down-stream vision tasks. Code is available at https://github.com/csguomy/mPrompt. Mingyue Guo 0001, Li Yuan 0007, Zhaoyi Yan, Binghui Chen, Yaowei Wang 0001, Qixiang Ye |
CVPR | 2 |
| 2024 | GraCo: Granularity-Controllable Interactive SegmentationabstractInteractive Segmentation (IS) segments specific objects or parts in the image according to user input. Current IS pipelines fall into two categories: single-granularity out-put and multi-granularity output. The latter aims to allevi-ate the spatial ambiguity present in the former. However, the multi-granularity output pipeline suffers from limited interaction flexibility and produces redundant results. In this work, we introduce Granularity-Controllable Interactive Segmentation (GraCo), a novel approach that allows precise control of prediction granularity by introducing ad-ditional parameters to input. This enhances the customization of the interactive system and eliminates redundancy while resolving ambiguity. Nevertheless, the exorbitant cost of annotating multi-granularity masks and the lack of avail-able datasets with granularity annotations make it difficult for models to acquire the necessary guidance to control out-put granularity. To address this problem, we design an any-granularity mask generator that exploits the semantic property of the pre-trained IS model to automatically gen-erate abundant mask-granularity pairs without requiring additional manual annotation. Based on these pairs, we propose a granularity-controllable learning strategy that efficiently imparts the granularity controllability to the IS model. Extensive experiments on intricate scenarios at ob-ject and part levels demonstrate that our GraCo has signifi-cant advantages over previous methods. This highlights the potential of GraCo to be a flexible annotation tool, capable of adapting to diverse segmentation scenarios. The project page: https://zhao-yian.github.io/GraCo. Yian Zhao, Kehan Li 0002, Zesen Cheng, Pengchong Qiao, Xiawu Zheng, Rongrong Ji, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
CVPR | 8 |
| 2024 | Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
Peng Jin 0001, Hao Li 0073, Zesen Cheng, Kehan Li 0002, Runyi Yu 0002, Chang Liu 0047, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
ECCV (25) | 8 |
| 2024 | Learning Pseudo 3D Guidance for View-Consistent Texturing with 2D Diffusion
Kehan Li 0002, Yanbo Fan, Yang Wu 0001, Zhongqian Sun, Wei Yang 0019, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
ECCV (86) | 7 |
| 2024 | FreestyleRet: Retrieving Images from Style-Diversified Queries
Hao Li 0073, Yanhao Jia, Peng Jin 0001, Zesen Cheng, Kehan Li 0002, Jialu Sui, Chang Liu 0047, Li Yuan 0007 |
ECCV (23) | 8 |
| 2024 | HiFi-123: Towards High-Fidelity One Image to 3D Content Generation
Wangbo Yu, Li Yuan 0007, Yan-Pei Cao 0001, Xiangjun Gao, Xiaoyu Li 0002, Wenbo Hu 0002, Ying Shan, Yonghong Tian 0001 |
ECCV (73) | 2 |
| 2024 | Repaint123: Fast and High-Quality One Image to 3D Generation with Progressive Controllable Repainting
Junwu Zhang, Zhenyu Tang 0004, Yatian Pang, Xinhua Cheng, Peng Jin 0001, Yida Wei, Munan Ning, Li Yuan 0007 |
ECCV (25) | 9 |
| 2024 | Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionabstractLarge Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding.Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models.However, due to the lack of unified tokenization for images and videos, namely misalignment before projection, it becomes challenging for a Large Language Model (LLM) to learn multi-modal interactions from several poor projection layers.In this work, we unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM.As a result, we establish a simple but robust LVLM baseline, Video-LLaVA, which learns from a mixed dataset of images and videos, mutually enhancing each other.As a result, Video-LLaVA outperforms Video-ChatGPT by 5.8%, 9.9%, 18.6%, and 10.1% on MSRVTT, MSVD, TGIF, and ActivityNet, respectively.Additionally, our Video-LLaVA also achieves superior performances on a broad range of 9 image benchmarks.Notably, extensive experiments demonstrate that Video-LLaVA mutually benefits images and videos within a unified visual representation, outperforming models designed specifically for images or videos.We aim for this work to provide modest insights into the multi-modal inputs for the LLM. Bin Lin 0014, Jiaxi Cui, Munan Ning, Peng Jin 0001, Li Yuan 0007 |
EMNLP | 7 |
| 2024 | A Multi-modal Spiking Meta-learner with Brain-Inspired Task-Aware Modulation Scheme
Zhaokun Zhou, Kaiwei Che, Li Yuan 0007 |
ICANN (10) | 4 |
| 2024 | Temporal Contrastive Learning for Spiking Neural Networks
Haonan Qiu, Zeyin Song, Yanqi Chen, Munan Ning, Wei Fang 0006, Zhengyu Ma, Li Yuan 0007, Yonghong Tian 0001 |
ICANN (10) | 8 |
| 2024 | Optimal ANN-SNN Conversion with Group NeuronsabstractSpiking Neural Networks (SNNs) have emerged as a promising third generation of neural networks, offering unique characteristics such as binary outputs, high sparsity, and biological plausibility. However, the lack of effective learning algorithms remains a challenge for SNNs. For instance, while converting artificial neural networks (ANNs) to SNNs circumvents the need for direct training of SNNs, it encounters issues related to conversion errors and high inference time delays. In order to reduce or even eliminate conversion errors while decreasing inference time-steps, we have introduced a novel type of neuron called Group Neurons (GNs). One GN is composed of multiple Integrate-and-Fire (IF) neurons as members, and its neural dynamics are meticulously designed. Based on GNs, we have optimized the traditional ANN-SNN conversion framework. Specifically, we replace the IF neurons in the SNNs obtained by the traditional conversion framework with GNs. The resulting SNNs, which utilize GNs, are capable of achieving accuracy levels comparable to ANNs even within extremely short inference time-steps. The experiments on CIFAR10, CIFAR100, and ImageNet datasets demonstrate the superiority of the proposed methods in terms of both inference accuracy and latency. Code is available at https://github.com/Lyu6PosHao/ANN2SNN_GN. Liuzhenghao Lv, Wei Fang 0006, Li Yuan 0007, Yonghong Tian 0001 |
ICASSP | 3 |
| 2024 | Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsabstractRecent text-to-3D generation methods achieve impressive 3D content creation capacity thanks to the advances in image diffusion models and optimizing strategies. However, current methods struggle to generate correct 3D content for a complex prompt in semantics, i.e., a prompt describing multiple interacted objects binding with different attributes. In this work, we propose a general framework named Progressive3D, which decomposes the entire generation into a series of locally progressive editing steps to create precise 3D content for complex prompts, and we constrain the content change to only occur in regions determined by user-defined region prompts in each editing step. Furthermore, we propose an overlapped semantic component suppression technique to encourage the optimization process to focus more on the semantic differences between prompts. Extensive experiments demonstrate that the proposed Progressive3D framework generates precise 3D content for prompts with complex semantics through progressive editing steps and is general for various text-to-3D methods driven by different 3D representations. Xinhua Cheng, Tianyu Yang 0003, Yu Li 0003, Lei Zhang 0001, Jian Zhang 0018, Li Yuan 0007 |
ICLR | 7 |
| 2024 | LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentabstractThe video-language (VL) pretraining has achieved remarkable improvement in multiple downstream tasks. However, the current VL pretraining framework is hard to extend to multiple modalities (N modalities, N ≥ 3) beyond vision and language. We thus propose LanguageBind, taking the language as the bind across different modalities because the language modality is well-explored and contains rich semantics. Specifically, we freeze the language encoder acquired by VL pretraining and then train encoders for other modalities with contrastive learning. As a result, all modalities are mapped to a shared feature space, implementing multi-modal semantic alignment. While LanguageBind ensures that we can extend VL modalities to N modalities, we also need a high-quality dataset with alignment data pairs centered on language. We thus propose VIDAL-10M with 10 Million data with Video, Infrared, Depth, Audio and their corresponding Language. In our VIDAL-10M, all videos are from short video platforms with complete semantics rather than truncated segments from long videos, and all the video, depth, infrared, and audio modalities are aligned to their textual descriptions. LanguageBind has achieved superior performance on a wide range of 15 benchmarks covering video, audio, depth, and infrared. Moreover, multiple experiments have provided evidence for the effectiveness of LanguageBind in achieving indirect alignment and complementarity among diverse modalities. Bin Lin 0014, Munan Ning, Jiaxi Cui, Hongfa Wang, Yatian Pang, Junwu Zhang, Caiwan Zhang, Zhifeng Li 0001, Wei Liu 0005, Li Yuan 0007 |
ICLR | 14 |
| 2024 | Fast and Robust Point Cloud Registration with Tree-based TransformerabstractPoint cloud registration is essential in computer vision and robotics. Recently, transformer-based methods have achieved advanced point cloud registration performance. However, the standard attention mechanism utilized in these methods considers many low-relevance points, and it has difficulty focusing its attention weights on sparse and meaningful points, leading to limited local structure modeling capabilities and quadratic computational complexity. To address these limitations, we present the Tree-based Transformer (TrT), which is able to extract abundant local and global features with linear computational complexity. Specifically, the TrT builds coarse-to-dense feature trees, and a novel Tree-based Attention (TrA) is proposed to guide the progressive convergence of the attended regions toward meaningful points and to structurize point clouds following tree structures. In each layer, the top ${\mathcal{S}}$ key points with the highest attention scores are selected, such that in the next layer, attention is evaluated only within the specified high-relevance regions, corresponding to the child points of these selected ${\mathcal{S}}$ points. Additionally, coarse features containing high-level semantic information are incorporated into the child points to guide the feature extraction process, facilitating local structure modeling and multiscale information integration. Consequently, TrA enables the model to focus on critical local structures and extract rich local information with linear computational complexity. Experiments demonstrate that our method achieves state-of-the-art performance on 3DMatch and KITTI benchmarks. The code for our method is publicly available at https://github.com/CGuangyan-BIT/TrT. Guangyan Chen, Meiling Wang 0002, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
ICRA | 4 |
| 2024 | Prompt2Poster: Automatically Artistic Chinese Poster Creation from Prompt OnlyabstractAs a critical component in graphic design, artistic posters are widely applied in the advertising and entertainment industry, thus the automatic poster creation from user-provided prompts has become increasingly desired recently. Although existing Text2Image methods create impressive images aligned with given prompts, they fail to generate ideal artistic posters, especially with Chinese texts. To create desired artistic Chinese posters including an aligned background, reasonable layouts, and stylized graphical texts from given prompts only, we propose an automatic poster creation framework, named Prompt2Poster. Our framework utilizes the capacity of the powerful Large Language Model (LLM) to extract user intention from provided prompts and generate the aligned background. Although only taking a user prompt as the input, linguistic, visual, and geometrical information is fully utilized in the framework, bringing the ability to fit different distributions. To achieve the use of multi-modal information in the framework, two carefully designed modules, Controllable Layout Generator (CLG) and Graphical Text Generator (GTG) are proposed, leading to accurate and pleasurable visual results. Comprehensive experiments demonstrate that our Prompt2Poster achieves superior performance, especially in text quality and visual harmony. Yunyang Ge, Liuhan Chen, Haiyang Zhou, Qian Wang 0062, Xinhua Cheng, Li Yuan 0007 |
ACM Multimedia | 7 |
| 2024 | ShareGPT4Video: Improving Video Understanding and Generation with Better CaptionsabstractWe present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) via dense and precise captions. The series comprises: 1) ShareGPT4Video, 40K GPT4V annotated dense captions of videos with various lengths and sources, developed through carefully designed data filtering and annotating strategy. 2) ShareCaptioner-Video, an efficient and capable captioning model for arbitrary videos, with 4.8M high-quality aesthetic videos annotated by it. 3) ShareGPT4Video-8B, a simple yet superb LVLM that reached SOTA performance on three advancing video benchmarks. To achieve this, taking aside the non-scalable costly human annotators, we find using GPT4V to caption video with a naive multi-frame or frame-concatenation input strategy leads to less detailed and sometimes temporal-confused results. We argue the challenge of designing a high-quality video captioning strategy lies in three aspects: 1) Inter-frame precise temporal change understanding. 2) Intra-frame detailed content description. 3) Frame-number scalability for arbitrary-length videos. To this end, we meticulously designed a differential video captioning strategy, which is stable, scalable, and efficient for generating captions for videos with arbitrary resolution, aspect ratios, and length. Based on it, we construct ShareGPT4Video, which contains 40K high-quality videos spanning a wide range of categories, and the resulting captions encompass rich world knowledge, object attributes, camera movements, and crucially, detailed and precise temporal descriptions of events. Based on ShareGPT4Video, we further develop ShareCaptioner-Video, a superior captioner capable of efficiently generating high-quality captions for arbitrary videos. We annotated 4.8M aesthetically appealing videos by it and verified their effectiveness on a 10-second text2video generation task. For video understanding, we verified the effectiveness of ShareGPT4Video on several current LVLM architectures and presented our superb new LVLM ShareGPT4Video-8B. All the models, strategies, and annotations will be open-sourced and we hope this project can serve as a pivotal resource for advancing both the LVLMs and T2VMs community. Lin Chen 0016, Xilin Wei, Jinsong Li 0001, Xiaoyi Dong, Pan Zhang 0001, Yuhang Zang, Haodong Duan, Lin Bin, Zhenyu Tang 0004, Li Yuan 0007, Yu Qiao 0001, Dahua Lin, Feng Zhao 0004, Jiaqi Wang 0003 |
NeurIPS | 11 |
| 2024 | VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained ActionsabstractVisual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, current VIL methods naively employ VLMs to learn high-level plans from human videos, relying on pre-defined motion primitives for executing physical interactions, which remains a major bottleneck. In this work, we present VLMimic, a novel paradigm that harnesses VLMs to directly learn even fine-grained action levels, only given a limited number of human videos. Specifically, VLMimic first grounds object-centric movements from human videos, and learns skills using hierarchical constraint representations, facilitating the derivation of skills with fine-grained action levels from limited human videos. These skills are refined and updated through an iterative comparison strategy, enabling efficient adaptation to unseen environments. Our extensive experiments exhibit that our VLMimic, using only 5 human videos, yields significant improvements of over 27% and 21% in RLBench and real-world manipulation tasks, and surpasses baselines by more than 37% in long-horizon tasks. Code and videos are available on our anonymous homepage. Guangyan Chen, Meiling Wang 0002, Te Cui, Yao Mu 0001, Tianxing Zhou, Zicai Peng, Mengxiao Hu, Haizhou Li 0004, Li Yuan 0007, Yi Yang 0009, Yufeng Yue |
NeurIPS | 10 |
| 2024 | DF40: Toward Next-Generation Deepfake DetectionabstractWe propose a new comprehensive benchmark to revolutionize the current deepfake detection field to the next generation. Predominantly, existing works identify top-notch detection algorithms and models by adhering to the common practice: training detectors on one specific dataset (e.g., FF++) and testing them on other prevalent deepfake datasets. This protocol is often regarded as a "golden compass" for navigating SoTA detectors. But can these stand-out "winners" be truly applied to tackle the myriad of realistic and diverse deepfakes lurking in the real world? If not, what underlying factors contribute to this gap? In this work, we found the dataset (both train and test) can be the "primary culprit" due to the following: (1) forgery diversity: Deepfake techniques are commonly referred to as both face forgery (face-swapping and face-reenactment) and entire image synthesis (AIGC, especially face). Most existing datasets only contain partial types of them, with limited forgery methods implemented (e.g., 2 swapping and 2 reenactment methods in FF++); (2) forgery realism: The dominated training dataset, FF++, contains out-of-date forgery techniques from the past four years. "Honing skills" on these forgeries makes it difficult to guarantee effective detection generalization toward nowadays' SoTA deepfakes; (3) evaluation protocol: Most detection works perform evaluations on one type, e.g., face-swapping types only, which hinders the development of universal deepfake detectors.To address this dilemma, we construct a highly diverse and large-scale deepfake detection dataset called DF40, which comprises 40 distinct deepfake techniques (10 times larger than FF++). We then conduct comprehensive evaluations using 4 standard evaluation protocols and 8 representative detection methods, resulting in over 2,000 evaluations. Through these evaluations, we provide an extensive analysis from various perspectives, leading to 7 new insightful findings contributing to the field. We also open up 4 valuable yet previously underexplored research questions to inspire future works. We release our dataset, code, and pre-trained weights at https://github.com/YZY-stack/DF40. Zhiyuan Yan 0002, Taiping Yao, Shen Chen 0004, Yandan Zhao, Xinghe Fu, Donghao Luo 0001, Chengjie Wang 0001, Shouhong Ding, Yunsheng Wu, Li Yuan 0007 |
NeurIPS | 11 |
| 2024 | ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video GenerationabstractWe propose a novel text-to-video (T2V) generation benchmark, ChronoMagic-Bench, to evaluate the temporal and metamorphic knowledge skills in time-lapse video generation of the T2V models (e.g. Sora and Lumiere). Compared to existing benchmarks that focus on visual quality and text relevance of generated videos, ChronoMagic-Bench focuses on the models’ ability to generate time-lapse videos with significant metamorphic amplitude and temporal coherence. The benchmark probes T2V models for their physics, biology, and chemistry capabilities, in a free-form text control. For these purposes, ChronoMagic-Bench introduces 1,649 prompts and real-world videos as references, categorized into four major types of time-lapse videos: biological, human creation, meteorological, and physical phenomena, which are further divided into 75 subcategories. This categorization ensures a comprehensive evaluation of the models’ capacity to handle diverse and complex transformations. To accurately align human preference on the benchmark, we introduce two new automatic metrics, MTScore and CHScore, to evaluate the videos' metamorphic attributes and temporal coherence. MTScore measures the metamorphic amplitude, reflecting the degree of change over time, while CHScore assesses the temporal coherence, ensuring the generated videos maintain logical progression and continuity. Based on the ChronoMagic-Bench, we conduct comprehensive manual evaluations of eighteen representative T2V models, revealing their strengths and weaknesses across different categories of prompts, providing a thorough evaluation framework that addresses current gaps in video generation research. More encouragingly, we create a large-scale ChronoMagic-Pro dataset, containing 460k high-quality pairs of 720p time-lapse videos and detailed captions. Each caption ensures high physical content and large metamorphic amplitude, which have a far-reaching impact on the video generation community. The source data and code are publicly available on https://pku-yuangroup.github.io/ChronoMagic-Bench. Shenghai Yuan 0002, Jinfa Huang, Yongqi Xu, Yaoyang Liu, Shaofeng Zhang, Yujun Shi, Rui-Jie Zhu 0003, Xinhua Cheng, Jiebo Luo 0001, Li Yuan 0007 |
NeurIPS | 10 |
| 2024 | Spiking Transformer with Experts MixtureabstractSpiking Neural Networks (SNNs) provide a sparse spike-driven mechanism which is believed to be critical for energy-efficient deep learning.
Mixture-of-Experts (MoE), on the other side, aligns with the brain mechanism of distributed and sparse processing, resulting in an efficient way of enhancing model capacity and conditional computation.
In this work, we consider how to incorporate SNNs’ spike-driven and MoE’s conditional computation into a unified framework.
However, MoE uses softmax to get the dense conditional weights for each expert and TopK to hard-sparsify the network, which does not fit the properties of SNNs.
To address this issue, we reformulate MoE in SNNs and introduce the Spiking Experts Mixture Mechanism (SEMM) from the perspective of sparse spiking activation.
Both the experts and the router output spiking sequences, and their element-wise operation makes SEMM computation spike-driven and dynamic sparse-conditional.
By developing SEMM into Spiking Transformer, the Experts Mixture Spiking Attention (EMSA) and the Experts Mixture Spiking Perceptron (EMSP) are proposed, which performs routing allocation for head-wise and channel-wise spiking experts, respectively. Experiments show that SEMM realizes sparse conditional computation and obtains a stable improvement on neuromorphic and static datasets with approximate computational overhead based on the Spiking Transformer baselines. Zhaokun Zhou, Yijie Lu, Yanhao Jia, Kaiwei Che, Liwei Huang, Yuesheng Zhu, Guoqi Li 0002, Zhaofei Yu, Li Yuan 0007 |
NeurIPS | 11 |
| 2024 | QKFormer: Hierarchical Spiking Transformer using Q-K AttentionabstractSpiking Transformers, which integrate Spiking Neural Networks (SNNs) with Transformer architectures, have attracted significant attention due to their potential for low energy consumption and high performance. However, there remains a substantial gap in performance between SNNs and Artificial Neural Networks (ANNs). To narrow this gap, we have developed QKFormer, a direct training spiking transformer with the following features: i) _Linear complexity and high energy efficiency_, the novel spike-form Q-K attention module efficiently models the token or channel attention through binary vectors and enables the construction of larger models. ii) _Multi-scale spiking representation_, achieved by a hierarchical structure with the different numbers of tokens across blocks. iii) _Spiking Patch Embedding with Deformed Shortcut (SPEDS)_, enhances spiking information transmission and integration, thus improving overall performance. It is shown that QKFormer achieves significantly superior performance over existing state-of-the-art SNN models on various mainstream datasets. Notably, with comparable size to Spikformer (66.34 M, 74.81\%), QKFormer (64.96 M) achieves a groundbreaking top-1 accuracy of **85.65\%** on ImageNet-1k, substantially outperforming Spikformer by **10.84\%**. To our best knowledge, this is the first time that directly training SNNs have exceeded 85\% accuracy on ImageNet-1K. Chenlin Zhou, Han Zhang 0035, Zhaokun Zhou, Liutao Yu, Liwei Huang, Xiaopeng Fan 0001, Li Yuan 0007, Zhengyu Ma, Yonghong Tian 0001 |
NeurIPS | 7 |
| 2024 | Fully Transformer-Equipped Architecture for end-to-end Referring Video Object Segmentation
Ping Li 0006, Li Yuan 0007, Xianghua Xu |
Inf. Process. Manag. | 3 |
| 2024 | Self-architectural knowledge distillation for spiking neural networks
Haonan Qiu, Munan Ning, Zeyin Song, Wei Fang 0006, Yanqi Chen, Zhengyu Ma, Li Yuan 0007, Yonghong Tian 0001 |
Neural Networks | 8 |
| 2024 | Efficient Long-Short Temporal Attention network for unsupervised Video Object Segmentation
Ping Li 0006, Li Yuan 0007, Huaxin Xiao, Binbin Lin 0001, Xianghua Xu |
Pattern Recognit. | 3 |
| 2024 | Adversarial Attacks on Video Object Segmentation With Hard Region DiscoveryabstractVideo object segmentation has been applied to various computer vision tasks, such as video editing, autonomous driving, and human-robot interaction. However, the methods based on deep neural networks are vulnerable to adversarial examples, which are the inputs attacked by almost human-imperceptible perturbations, and the adversary (i.e., attacker) will fool the segmentation model to make incorrect pixel-level predictions. This will rise the security issues in highly-demanding tasks because small perturbations to the input video will result in potential attack risks. Though adversarial examples have been extensively used for classification, it is rarely studied in video object segmentation. Existing related methods in computer vision either require prior knowledge of categories or cannot be directly applied due to the special design for certain tasks, failing to consider the pixel-wise region attack. Hence, this work develops an object-agnostic adversary that has adversarial impacts on VOS by first-frame attacking via hard region discovery. Particularly, the gradients from the segmentation model are exploited to discover the easily confused region, in which it is difficult to identify the pixel-wise objects from the background in a frame. This provides a hardness map that helps to generate perturbations with a stronger adversarial power for attacking the first frame. Empirical studies on three benchmarks indicate that our attacker significantly degrades the performance of several state-of-the-art video object segmentation models. Ping Li 0006, Li Yuan 0007, Jian Zhao 0006, Xianghua Xu, Xiaoqin Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | An Organ-Aware Diagnosis Framework for Radiology Report GenerationabstractRadiology report generation (RRG) is crucial to save the valuable time of radiologists in drafting the report, therefore increasing their work efficiency. Compared to typical methods that directly transfer image captioning technologies to RRG, our approach incorporates organ-wise priors into the report generation. Specifically, in this paper, we propose Organ-aware Diagnosis (OaD) to generate diagnostic reports containing descriptions of each physiological organ. During training, we first develop a task distillation (TD) module to extract organ-level descriptions from reports. We then introduce an organ-aware report generation module that, for one thing, provides a specific description for each organ, and for another, simulates clinical situations to provide short descriptions for normal cases. Furthermore, we design an auto-balance mask loss to ensure balanced training for normal/abnormal descriptions and various organs simultaneously. Being intuitively reasonable and practically simple, our OaD outperforms SOTA alternatives by large margins on commonly used IU-Xray and MIMIC-CXR datasets, as evidenced by a 3.4% BLEU-1 improvement on MIMIC-CXR and 2.0% BLEU-2 improvement on IU-Xray. Pengchong Qiao, Lin Wang 0026, Munan Ning, Li Yuan 0007, Yefeng Zheng 0001, Jie Chen 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Full Transformer Framework for Robust Point Cloud Registration With Deep Information InteractionabstractPoint cloud registration is an essential technology in computer vision and robotics. Recently, transformer-based methods have achieved advanced performance in point cloud registration by utilizing the advantages of the transformer in order-invariance and modeling dependencies to aggregate information. However, they still suffer from indistinct feature extraction, sensitivity to noise, and outliers, owing to three major limitations: 1) the adoption of CNNs fails to model global relations due to their local receptive fields, resulting in extracted features susceptible to noise; 2) the shallow-wide architecture of transformers and the lack of positional information lead to indistinct feature extraction due to inefficient information interaction; and 3) the insufficient consideration of geometrical compatibility leads to the ambiguous identification of incorrect correspondences. To address the above-mentioned limitations, a novel full transformer network for point cloud registration is proposed, named the deep interaction transformer (DIT), which incorporates: 1) a point cloud structure extractor (PSE) to retrieve structural information and model global relations with the local feature integrator (LFI) and transformer encoders; 2) a deep-narrow point feature transformer (PFT) to facilitate deep information interaction across a pair of point clouds with positional information, such that transformers establish comprehensive associations and directly learn the relative position between points; and 3) a geometric matching-based correspondence confidence evaluation (GMCCE) method to measure spatial consistency and estimate correspondence confidence by the designed triangulated descriptor. Extensive experiments on the ModelNet40, ScanObjectNN, and 3DMatch datasets demonstrate that our method is capable of precisely aligning point clouds, consequently, achieving superior performance compared with state-of-the-art methods. The code is publicly available at https://github.com/CGuangyan-BIT/DIT. Guangyan Chen, Meiling Wang 0002, Qingxiang Zhang, Li Yuan 0007, Yufeng Yue |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Learnable Central Similarity Quantization for Efficient Image and Video RetrievalabstractData-dependent hashing methods aim to learn hash functions from the pairwise or triplet relationships among the data, which often lead to low efficiency and low collision rate by only capturing the local distribution of the data. To solve the limitation, we propose central similarity, in which the hash codes of similar data pairs are encouraged to approach a common center and those of dissimilar pairs to converge to different centers. As a new global similarity metric, central similarity can improve the efficiency and retrieval accuracy of hash learning. By introducing a new concept, hash centers, we principally formulate the computation of the proposed central similarity metric, in which the hash centers refer to a set of points scattered in the Hamming space with a sufficient mutual distance between each other. To construct well-separated hash centers, we provide two efficient methods: 1) leveraging the Hadamard matrix and Bernoulli distributions to generate data-independent hash centers and 2) learning data-dependent hash centers from data representations. Based on the proposed similarity metric and hash centers, we propose central similarity quantization (CSQ) that optimizes the central similarity between data points with respect to their hash centers instead of optimizing the local similarity to generate a high-quality deep hash function. We also further improve the CSQ with data-dependent hash centers, dubbed as CSQ with learnable center (CSQLC). The proposed CSQ and CSQLC are generic and applicable to image and video hashing scenarios. We conduct extensive experiments on large-scale image and video retrieval tasks, and the proposed CSQ yields noticeably boosted retrieval performance, i.e., 3%-20% in mean average precision (mAP) over the previous state-of-the-art methods, which also demonstrates that our methods can generate cohesive hash codes for similar data pairs and dispersed hash codes for dissimilar pairs. Li Yuan 0007, Tao Wang 0053, Xiaopeng Zhang 0008, Francis E. H. Tay, Zequn Jie, Yonghong Tian 0001, Wei Liu 0005, Jiashi Feng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | ACSeg: Adaptive Conceptualization for Unsupervised Semantic SegmentationabstractRecently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship among pixel-level representations typically contains rich class-aware information that semantically identical pixel embeddings in the representation space gather together to form sophisticated concepts. However, leveraging the learned models to ascertain semantically consistent pixel groups or regions in the image is non-trivial since over/ under-clustering overwhelms the conceptualization procedure under various semantic distributions of different images. In this work, we investigate the pixel-level semantic aggregation in self-supervised ViT pre-trained models as image Segmentation and propose the Adaptive Conceptualization approach for USS, termed ACSeg. Concretely, we explicitly encode concepts into learnable prototypes and design the Adaptive Concept Generator (ACG), which adaptively maps these prototypes to informative concepts for each image. Meanwhile, considering the scene complexity of different images, we propose the modularity loss to optimize ACG independent of the concept number based on estimating the intensity of pixel pairs belonging to the same concept. Finally, we turn the USS task into classifying the discovered concepts in an unsupervised manner. Extensive experiments with state-of-the-art results demonstrate the effectiveness of the proposed ACSeg. Kehan Li 0002, Zhennan Wang 0001, Zesen Cheng, Runyi Yu 0002, Yian Zhao, Guoli Song, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
CVPR | 8 |
| 2023 | Out-of-Candidate Rectification for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation is typically inspired by class activation maps, which serve as pseudo masks with class-discriminative regions highlighted. Although tremendous efforts have been made to recall precise and complete locations for each class, existing methods still commonly suffer from the unsolicited Out-of-Candidate (OC) error predictions that do not belong to the label candidates, which could be avoidable since the contradiction with image-level class tags is easy to be detected. In this paper, we develop a group ranking-based Out-of-f;Candidate Rectification (OCR) mechanism in a plug-and-play fashion. Firstly, we adaptively split the semantic categories into In-Candidate (IC) and OC groups for each OC pixel according to their prior annotation correlation and posterior prediction correlation. Then, we derive a differentiable rectification loss to force OC pixels to shift to the IC group. Incorporating OCR with seminal baselines (e.g., AffinityNet, SEAM, MCTformer), we can achieve remarkable performance gains on both Pascal VOC (+3.2%, +3.3%, +0.8% mIoU) and MS COCO (+1.0%, +1.3%, +0.5% mIoU) datasets with negligible extra training overhead, which jus-tifies the effectiveness and generality of OCR.††Ŋ github.com/sennnnn/Out-of-Candidate-Rectification Zesen Cheng, Pengchong Qiao, Kehan Li 0002, Siheng Li, Pengxu Wei, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
CVPR | 7 |
| 2023 | Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation LearningabstractContrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined video-text pairs. To clarify this coarse-grained global interaction and move a step further, we have to encounter challenging shell-breaking interactions for fine-grained cross-modal learning. In this paper, we creatively model video-text as game players with multivariate cooperative game theory to wisely handle the uncertainty during fine-grained semantic interaction with diverse granularity, flexible combination, and vague intensity. Concretely, we propose Hierarchical Banzhaf Interaction (HBI) to value possible correspondence between video frames and text words for sensitive and explainable cross-modal contrast. To efficiently realize the cooperative game of multiple video frames and multiple text words, the proposed method clusters the original video frames (text words) and computes the Banzhaf Interaction between the merged tokens. By stacking token merge modules, we achieve cooperative games at different semantic levels. Extensive experiments on commonly used text-video retrieval and video-question answering bench-marks with superior performances justify the efficacy of our HBI. More encouragingly, it can also serve as a visualization tool to promote the understanding of cross-modal interaction, which have a far-reaching impact on the community. Project page is available at https://jpthu17.github.io/HBI/. Peng Jin 0001, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu 0030, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
CVPR | 7 |
| 2023 | Learning with Fantasy: Semantic-Aware Virtual Contrastive Constraint for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) aims at learning to classify new classes continually from limited samples without forgetting the old classes. The mainstream framework tackling FSCIL is first to adopt the cross-entropy (CE) loss for training at the base session, then freeze the feature extractor to adapt to new classes. However, in this work, we find that the CE loss is not ideal for the base session training as it suffers poor class separation in terms of representations, which further degrades generalization to novel classes. One tempting method to mitigate this problem is to apply an additional naïve supervised contrastive learning (SCL) in the base session. Unfortunately, we find that although SCL can create a slightly better representation separation among different base classes, it still struggles to separate base classes and new classes. Inspired by the observations made, we propose Semantic-Aware Virtual Contrastive model (SAVC), a novel method that facilitates separation between new classes and base classes by introducing virtual classes to SCL. These virtual classes, which are generated via pre-defined transformations, not only act as placeholders for unseen classes in the representation space, but also provide diverse semantic information. By learning to recognize and contrast in the fantasy space fostered by virtual classes, our SAVC significantly boosts base class separation and novel class generalization, achieving new state-of-the-art performance on the three widely-used FSCIL benchmark datasets. Code is available at: https://github.com/zysong0113/SAVC. Zeyin Song, Yifan Zhao 0002, Yujun Shi, Peixi Peng, Li Yuan 0007, Yonghong Tian 0001 |
CVPR | 5 |
| 2023 | DiffusionRet: Generative Text-Video Retrieval with Diffusion ModelabstractExisting text-video retrieval solutions are, in essence, discriminant models focused on maximizing the conditional likelihood, i.e., p(candidates|query). While straightforward, this de facto paradigm overlooks the underlying data distribution p(query), which makes it challenging to identify out-of-distribution data. To address this limitation, we creatively tackle this task from a generative viewpoint and model the correlation between the text and the video as their joint probability p(candidates,query). This is accomplished through a diffusion-based text-video retrieval framework (Diffusion-Ret), which models the retrieval task as a process of gradually generating joint distribution from noise. During training, DiffusionRet is optimized from both the generation and discrimination perspectives, with the generator being optimized by generation loss and the feature extractor trained with contrastive loss. In this way, DiffusionRet cleverly leverages the strengths of both generative and discriminative methods. Extensive experiments on five commonly used text-video retrieval benchmarks, including MSRVTT, LSMDC, MSVD, ActivityNet Captions, and DiDeMo, with superior performances, justify the efficacy of our method. More encouragingly, without any modification, DiffusionRet even performs well in out-domain retrieval settings. We believe this work brings fundamental insights into the related fields. Code is available at https://github.com/jpthu17/DiffusionRet. Peng Jin 0001, Hao Li 0073, Zesen Cheng, Kehan Li 0002, Xiangyang Ji, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
ICCV | 7 |
| 2023 | Multi-granularity Interaction Simulation for Unsupervised Interactive SegmentationabstractInteractive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and medical image analysis. Typically, massive and expansive pixel-level annotations are spent to train deep models by object-oriented interactions with manually labeled object masks. In this work, we reveal that informative interactions can be made by simulation with semantic-consistent yet diverse region exploration in an unsupervised paradigm. Concretely, we introduce a Multi-granularity Interaction Simulation (MIS) approach to open up a promising direction for unsupervised interactive segmentation. Drawing on the high-quality dense features produced by recent self-supervised models, we propose to gradually merge patches or regions with similar features to form more extensive regions and thus, every merged region serves as a semantic-meaningful multi-granularity proposal. By randomly sampling these proposals and simulating possible interactions based on them, we provide meaningful interaction at multiple granularities to teach the model to understand interactions. Our MIS significantly outperforms non-deep learning unsupervised methods and is even comparable with some previous deep-supervised methods without any annotation. Kehan Li 0002, Yian Zhao, Zhennan Wang 0001, Zesen Cheng, Peng Jin 0001, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
ICCV | 7 |
| 2023 | Rethinking Point Cloud Registration as Masking and ReconstructionabstractPoint cloud registration is essential in computer vision and robotics. In this paper, a critical observation is made that the invisible parts of each point cloud can be directly utilized as inherent masks, and the aligned point cloud pair can be regarded as the reconstruction target. Motivated by this observation, we rethink the point cloud registration problem as a masking and reconstruction task. To this end, a generic and concise auxiliary training network, the Masked Reconstruction Auxiliary Network (MRA), is proposed. The MRA reconstructs the complete point cloud by separately using the encoded features of each point cloud obtained from the backbone, guiding the contextual features in the backbone to capture fine-grained geometric details and the overall structures of point cloud pairs. Unlike recently developed high-performing methods that incorporate specific encoding methods into transformer models, which sacrifice versatility and introduce significant computational complexity during the inference process, our MRA can be easily inserted into other methods to further improve registration accuracy. Additionally, the MRA is detached after training, thereby avoiding extra computational complexity during the inference process. Building upon the MRA, we present a novel transformer-based method, the Masked Reconstruction Transformer (MRT), which achieves both precise and efficient alignment using standard transformers. Extensive experiments conducted on the 3DMatch, ModelNet40, and KITTI datasets demonstrate the superior performance of our MRT over state-of-the-art methods. Codes are available at https://github.com/CGuangyan-BIT/MRA. Guangyan Chen, Meiling Wang 0002, Li Yuan 0007, Yi Yang 0009, Yufeng Yue |
ICCV | 3 |
| 2023 | Learning Sparse Neural Networks with Identity Layers
Mingjian Ni, Xiawu Zheng, Peixi Peng, Li Yuan 0007, Yonghong Tian 0001 |
ICIG (3) | 5 |
| 2023 | Spikformer: When Spiking Neural Network Meets Transformer
Zhaokun Zhou, Yuesheng Zhu, Yaowei Wang 0001, Shuicheng Yan, Yonghong Tian 0001, Li Yuan 0007 |
ICLR | 7 |
| 2023 | Deep Interactive Full Transformer Framework for Point Cloud RegistrationabstractPoint cloud registration is a crucial technology in the fields of robotics and computer vision. Despite the significant advances in point cloud registration enabled by Transformer-based methods, limitations persist due to indistinct feature extraction, noise sensitivity, and outlier handling. These limitations stem from three factors: (1) the inefficiency of convolutional neural networks (CNNs) to capture global relationships due to their local receptive fields, resulting in extracted features susceptible to noise; (2) the shallow-wide architecture of Transformers, coupled with a lack of positional information, leading to inefficient information interaction and indistinct feature extraction; and (3) the omission of geometrical compatibility leads to ambiguous identification of incorrect correspondences. To overcome these limitations, we propose the Deep Interactive Full Transformer (DIFT) network for point cloud registration, which consists of three key components: (1) a Point Cloud Structure Extractor (PSE) for modeling global relationships and retrieving structural information; (2) a Point Feature Transformer (PFT) for establishing comprehensive associations and directly learning the relative positions between points; and (3) a Geometric Matching-based Correspondence Confidence Evaluation (GMCCE) method for measuring spatial consistency and estimating correspondence confidence. Experimental results on ModelNet40 and 3DMatch datasets demonstrate the superior performance of our proposed method compared to existing state-of-the-art methods. The code for our method is publicly available at https://github.com/CGuangyan-BIT/DIFT. Guangyan Chen, Meiling Wang 0002, Qingxiang Zhang, Li Yuan 0007, Tong Liu 0009, Yufeng Yue |
ICRA | 4 |
| 2023 | Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set AlignmentabstractText-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they fail to leverage the heterogeneous concepts in data. In this paper, we propose the Disentangled Conceptualization and Set-to-set Alignment (DiCoSA) to simulate the conceptualizing and reasoning process of human beings. For disentangled conceptualization, we divide the coarse feature into multiple latent factors related to semantic concepts. For set-to-set alignment, where a set of visual concepts correspond to a set of textual concepts, we propose an adaptive pooling method to aggregate semantic concepts to address the partial matching. In particular, since we encode concepts independently in only a few dimensions, DiCoSA is superior at efficiency and granularity, ensuring fine-grained interactions using a similar computational complexity as coarse-grained alignment. Extensive experiments on five datasets, including MSR-VTT, LSMDC, MSVD, ActivityNet, and DiDeMo, demonstrate that our method outperforms the existing state-of-the-art methods. Peng Jin 0001, Hao Li 0073, Zesen Cheng, Jinfa Huang, Zhennan Wang 0001, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
IJCAI | 6 |
| 2023 | PointGPT: Auto-regressively Generative Pre-training from Point CloudsabstractLarge language models (LLMs) based on the generative pre-training transformer (GPT) have demonstrated remarkable effectiveness across a diverse range of downstream tasks. Inspired by the advancements of the GPT, we present PointGPT, a novel approach that extends the concept of GPT to point clouds, addressing the challenges associated with disorder properties, low information density, and task gaps. Specifically, a point cloud auto-regressive generation task is proposed to pre-train transformer models. Our method partitions the input point cloud into multiple point patches and arranges them in an ordered sequence based on their spatial proximity. Then, an extractor-generator based transformer decode, with a dual masking strategy, learns latent representations conditioned on the preceding point patches, aiming to predict the next one in an auto-regressive manner. To explore scalability and enhance performance, a larger pre-training dataset is collected. Additionally, a subsequent post-pre-training stage is introduced, incorporating a labeled hybrid dataset. Our scalable approach allows for learning high-capacity models that generalize well, achieving state-of-the-art performance on various downstream tasks. In particular, our approach achieves classification accuracies of 94.9% on the ModelNet40 dataset and 93.4% on the ScanObjectNN dataset, outperforming all other transformer models. Furthermore, our method also attains new state-of-the-art accuracies on all four few-shot learning benchmarks. Codes are available at https://github.com/CGuangyan-BIT/PointGPT. Guangyan Chen, Meiling Wang 0002, Yi Yang 0009, Li Yuan 0007, Yufeng Yue |
NeurIPS | 5 |
| 2023 | Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic GraphsabstractMost text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text representations may overemphasize the action names at the expense of other important properties and lack fine-grained details to guide the synthesis of subtly distinct motion. In this paper, we propose hierarchical semantic graphs for fine-grained control over motion generation. Specifically, we disentangle motion descriptions into hierarchical semantic graphs including three levels of motions, actions, and specifics. Such global-to-local structures facilitate a comprehensive understanding of motion description and fine-grained control of motion generation. Correspondingly, to leverage the coarse-to-fine topology of hierarchical semantic graphs, we decompose the text-to-motion diffusion process into three semantic levels, which correspond to capturing the overall motion, local actions, and action specifics. Extensive experiments on two benchmark human motion datasets, including HumanML3D and KIT, with superior performances, justify the efficacy of our method. More encouragingly, by modifying the edge weights of hierarchical semantic graphs, our method can continuously refine the generated motion, which may have a far-reaching impact on the community. Code and pre-trained weights are available at https://github.com/jpthu17/GraphMotion. Peng Jin 0001, Yang Wu 0001, Yanbo Fan, Zhongqian Sun, Wei Yang 0019, Li Yuan 0007 |
NeurIPS | 6 |
| 2023 | Spike-driven TransformerabstractSpiking Neural Networks (SNNs) provide an energy-efficient deep learning option due to their unique spike-based event-driven (i.e., spike-driven) paradigm. In this paper, we incorporate the spike-driven paradigm into Transformer by the proposed Spike-driven Transformer with four unique properties: (1) Event-driven, no calculation is triggered when the input of Transformer is zero; (2) Binary spike communication, all matrix multiplications associated with the spike matrix can be transformed into sparse additions; (3) Self-attention with linear complexity at both token and channel dimensions; (4) The operations between spike-form Query, Key, and Value are mask and addition. Together, there are only sparse addition operations in the Spike-driven Transformer. To this end, we design a novel Spike-Driven Self-Attention (SDSA), which exploits only mask and addition operations without any multiplication, and thus having up to $87.2\times$ lower computation energy than vanilla self-attention. Especially in SDSA, the matrix multiplication between Query, Key, and Value is designed as the mask operation. In addition, we rearrange all residual connections in the vanilla Transformer before the activation functions to ensure that all neurons transmit binary spike signals. It is shown that the Spike-driven Transformer can achieve 77.1\% top-1 accuracy on ImageNet-1K, which is the state-of-the-art result in the SNN field. Man Yao, Zhaokun Zhou, Li Yuan 0007, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002 |
NeurIPS | 4 |
| 2023 | Vision Permutator: A Permutable MLP-Like Architecture for Visual RecognitionabstractIn this paper, we present Vision Permutator, a conceptually simple and data efficient MLP-like architecture for visual recognition. By realizing the importance of the positional information carried by 2D feature representations, unlike recent MLP-like models that encode the spatial information along the flattened spatial dimensions, Vision Permutator separately encodes the feature representations along the height and width dimensions with linear projections. This allows Vision Permutator to capture long-range dependencies and meanwhile avoid the attention building process in transformers. The outputs are then aggregated in a mutually complementing manner to form expressive representations. We show that our Vision Permutators are formidable competitors to convolutional neural networks (CNNs) and vision transformers. Without the dependence on spatial convolutions or attention mechanisms, Vision Permutator achieves 81.5% top-1 accuracy on ImageNet without extra large-scale training data (e.g., ImageNet-22k) using only 25M learnable parameters, which is much better than most CNNs and vision transformers under the same model size constraint. When scaling up to 88M, it attains 83.2% top-1 accuracy, greatly improving the performance of recent state-of-the-art MLP-like networks for visual recognition. We hope this work could encourage research on rethinking the way of encoding spatial information and facilitate the development of MLP-like models. PyTorch/MindSpore/Jittor code is available at https://github.com/Andrew-Qibin/VisionPermutator. Qibin Hou, Zihang Jiang, Li Yuan 0007, Ming-Ming Cheng, Shuicheng Yan, Jiashi Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | MADAv2: Advanced Multi-Anchor Based Active Domain Adaptation SegmentationabstractUnsupervised domain adaption has been widely adopted in tasks with scarce annotated data. Unfortunately, mapping the target-domain distribution to the source-domain unconditionally may distort the essential structural information of the target-domain data, leading to inferior performance. To address this issue, we first propose to introduce active sample selection to assist domain adaptation regarding the semantic segmentation task. By innovatively adopting multiple anchors instead of a single centroid, both source and target domains can be better characterized as multimodal distributions, in which way more complementary and informative samples are selected from the target domain. With only a little workload to manually annotate these active samples, the distortion of the target-domain distribution can be effectively alleviated, achieving a large performance gain. In addition, a powerful semi-supervised domain adaptation strategy is proposed to alleviate the long-tail distribution problem and further improve the segmentation performance. Extensive experiments are conducted on public datasets, and the results demonstrate that the proposed approach outperforms state-of-the-art methods by large margins and achieves similar performance to the fully-supervised upperbound, i.e., 71.4% mIoU on GTA5 and 71.8% mIoU on SYNTHIA. The effectiveness of each component is also verified by thorough ablation studies. Munan Ning, Donghuan Lu, Yujia Xie, Dongdong Chen 0001, Dong Wei 0004, Yefeng Zheng 0001, Yonghong Tian 0001, Shuicheng Yan, Li Yuan 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2023 | VOLO: Vision Outlooker for Visual RecognitionabstractRecently, Vision Transformers (ViTs) have been broadly explored in visual recognition. With low efficiency in encoding fine-level features, the performance of ViTs is still inferior to the state-of-the-art CNNs when trained from scratch on a midsize dataset like ImageNet. Through experimental analysis, we find it is because of two reasons: 1) the simple tokenization of input images fails to model the important local structure such as edges and lines, leading to low training sample efficiency; 2) the redundant attention backbone design of ViTs leads to limited feature richness for fixed computation budgets and limited training samples. To overcome such limitations, we present a new simple and generic architecture, termed Vision Outlooker (VOLO), which implements a novel outlook attention operation that dynamically conduct the local feature aggregation mechanism in a sliding window manner across the input image. Unlike self-attention that focuses on modeling global dependencies of local features at a coarse level, our outlook attention targets at encoding finer-level features, which is critical for recognition but ignored by self-attention. Outlook attention breaks the bottleneck of self-attention whose computation cost scales quadratically with the input spatial dimension, and thus is much more memory efficient. Compared to our Tokens-To-Token Vision Transformer (T2T-ViT), VOLO can more efficiently encode fine-level features that are essential for high-performance visual recognition. Experiments show that with only 26.6 M learnable parameters, VOLO achieves 84.2% top-1 accuracy on ImageNet-1 K without using extra training data, 2.7% better than T2T-ViT with a comparable number of parameters. When the model size is scaled up to 296 M parameters, its performance can be further improved to 87.1%, setting a new record for ImageNet-1 K classification. In addition, we also take the proposed VOLO as pretrained models and report superior performance on downstream tasks, such as semantic segmentation. Code is available at https://github.com/sail-sg/volo. Li Yuan 0007, Qibin Hou, Zihang Jiang, Jiashi Feng, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Truncated attention-aware proposal networks with multi-scale dilation for temporal action detection
Ping Li 0006, Jiachen Cao, Li Yuan 0007, Qinghao Ye, Xianghua Xu |
Pattern Recognit. | 3 |
| 2022 | Improving Vision Transformers by Revisiting High-Frequency Components
Jiawang Bai, Li Yuan 0007, Shutao Xia, Shuicheng Yan, Zhifeng Li 0001, Wei Liu 0005 |
ECCV (24) | 2 |
| 2022 | Locality Guidance for Improving Vision Transformers on Tiny Datasets
Kehan Li 0002, Runyi Yu 0002, Zhennan Wang 0001, Li Yuan 0007, Guoli Song, Jie Chen 0001 |
ECCV (24) | 4 |
| 2022 | Masked Autoencoders for Point Cloud Self-supervised Learning
Yatian Pang, Wenxiao Wang 0001, Francis E. H. Tay, Wei Liu 0005, Yonghong Tian 0001, Li Yuan 0007 |
ECCV (2) | 6 |
| 2022 | DynaMixer: A Vision MLP Architecture with Dynamic MixingabstractRecently, MLP-like vision models have achieved promising performances on mainstream visual recognition tasks. In contrast with vision transformers and CNNs, the success of MLP-like models shows that simple information fusion operations among tokens and channels can yield a good representation power for deep recognition models. However, existing MLP-like models fuse tokens through static fusion operations, lacking adaptability to the contents of the tokens to be mixed. Thus, customary information fusion procedures are not effective enough. To this end, this paper presents an efficient MLP-like network architecture, dubbed DynaMixer, resorting to dynamic information fusion. Critically, we propose a procedure, on which the DynaMixer model relies, to dynamically generate mixing matrices by leveraging the contents of all the tokens to be mixed. To reduce the time complexity and improve the robustness, a dimensionality reduction technique and a multi-segment fusion mechanism are adopted. Our proposed DynaMixer model (97M parameters) achieves 84.3% top-1 accuracy on the ImageNet-1K dataset without extra training data, performing favorably against the state-of-the-art vision MLP models. When the number of parameters is reduced to 26M, it still achieves 82.7% top-1 accuracy, surpassing the existing MLP-like models with a similar capacity. The code is available at \url{https://github.com/ziyuwwang/DynaMixer}. Li Yuan 0007, Yibing Song, Wei Liu 0005 |
ICML | 4 |
| 2021 | Continual Learning via Bit-Level Information PreservingabstractContinual learning tackles the setting of learning different tasks sequentially. Despite the lots of previous solutions, most of them still suffer significant forgetting or expensive memory cost. In this work, targeted at these problems, we first study the continual learning process through the lens of information theory and observe that forgetting of a model stems from the loss of information gain on its parameters from the previous tasks when learning a new task. From this viewpoint, we then propose a novel continual learning approach called Bit-Level Information Preserving (BLIP) that preserves the information gain on model parameters through updating the parameters at the bit level, which can be conveniently implemented with parameter quantization. More specifically, BLIP first trains a neural network with weight quantization on the new incoming task and then estimates information gain on each parameter provided by the task data to determine the bits to be frozen to prevent forgetting. We conduct extensive experiments ranging from classification tasks to reinforcement learning tasks, and the results show that our method produces better or on par results comparing to previous state-of-the-arts. Indeed, BLIP achieves close to zero forgetting while only requiring constant memory overheads throughout continual learning1. Yujun Shi, Li Yuan 0007, Yunpeng Chen, Jiashi Feng |
CVPR | 2 |
| 2021 | Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetabstractTransformers, which are popular for language modeling, have been explored for solving vision tasks recently, e.g., the Vision Transformer (ViT) for image classification. The ViT model splits each image into a sequence of tokens with fixed length and then applies multiple Transformer layers to model their global relation for classification. However, ViT achieves inferior performance to CNNs when trained from scratch on a midsize dataset like ImageNet. We find it is because: 1) the simple tokenization of input images fails to model the important local structure such as edges and lines among neighboring pixels, leading to low training sample efficiency; 2) the redundant attention backbone design of ViT leads to limited feature richness for fixed computation budgets and limited training samples. To overcome such limitations, we propose a new Tokens-To-Token Vision Transformer (T2T-VTT), which incorporates 1) a layer-wise Tokens-to-Token (T2T) transformation to progressively structurize the image to tokens by recursively aggregating neighboring Tokens into one Token (Tokens-to-Token), such that local structure represented by surrounding tokens can be modeled and tokens length can be reduced; 2) an efficient backbone with a deep-narrow structure for vision transformer motivated by CNN architecture design after empirical study. Notably, T2T-ViT reduces the parameter count and MACs of vanilla ViT by half, while achieving more than 3.0% improvement when trained from scratch on ImageNet. It also outperforms ResNets and achieves comparable performance with MobileNets by directly training on ImageNet. For example, T2T-ViT with comparable size to ResNet50 (21.5M parameters) can achieve 83.3% top1 accuracy in image resolution 384x384 on ImageNet.1 Li Yuan 0007, Yunpeng Chen, Tao Wang 0053, Weihao Yu 0001, Yujun Shi, Zihang Jiang, Francis E. H. Tay, Jiashi Feng, Shuicheng Yan |
ICCV | 1 |
| 2021 | PnP-DETR: Towards Efficient Visual Analysis with TransformersabstractRecently, DETR [3] pioneered the solution of vision tasks with transformers, it directly translates the image feature map into the object detection result. Though effective, translating the full feature map can be costly due to redundant computation on some area like the background. In this work, we encapsulate the idea of reducing spatial redundancy into a novel poll and pool (PnP) sampling module, with which we build an end-to-end PnP-DETR architecture that adaptively allocates its computation spatially to be more efficient. Concretely, the PnP module abstracts the image feature map into fine foreground object feature vectors and a small number of coarse background contextual feature vectors. The transformer models information interaction within the fine-coarse feature space and translates the features into the detection result. Moreover, the PnP-augmented model can instantly achieve various desired trade-offs between performance and computation with a single model by varying the sampled feature length, without requiring to train multiple models as existing methods. Thus it offers greater flexibility for deployment in diverse scenarios with varying computation constraint. We further validate the generalizability of the PnP module on panoptic segmentation and the recent transformer-based image recognition model ViT [7] and show consistent efficiency gain. We believe our method makes a step for efficient visual analysis with transformers, wherein spatial redundancy is commonly observed. Code and models will be available. Tao Wang 0053, Li Yuan 0007, Yunpeng Chen, Jiashi Feng, Shuicheng Yan |
ICCV | 2 |
| 2021 | All Tokens Matter: Token Labeling for Training Better Vision TransformersabstractIn this paper, we present token labeling---a new training objective for training high-performance vision transformers (ViTs). Different from the standard training objective of ViTs that computes the classification loss on an additional trainable class token, our proposed one takes advantage of all the image patch tokens to compute the training loss in a dense manner. Specifically, token labeling reformulates the image classification problem into multiple token-level recognition problems and assigns each patch token with an individual location-specific supervision generated by a machine annotator. Experiments show that token labeling can clearly and consistently improve the performance of various ViT models across a wide spectrum. For a vision transformer with 26M learnable parameters serving as an example, with token labeling, the model can achieve 84.4% Top-1 accuracy on ImageNet. The result can be further increased to 86.4% by slightly scaling the model size up to 150M, delivering the minimal-sized model among previous models (250M+) reaching 86%. We also show that token labeling can clearly improve the generalization capability of the pretrained models on downstream tasks with dense prediction, such as semantic segmentation. Our code and model are publiclyavailable at https://github.com/zihangJiang/TokenLabeling. Zihang Jiang, Qibin Hou, Li Yuan 0007, Daquan Zhou, Yujun Shi, Xiaojie Jin 0004, Anran Wang 0001, Jiashi Feng |
NeurIPS | 3 |
| 2021 | Exploring global diverse attention via pairwise temporal relation for video summarization
Ping Li 0006, Qinghao Ye, Li Yuan 0007, Xianghua Xu, Ling Shao 0001 |
Pattern Recognit. | 4 |
| 2020 | Revisiting Knowledge Distillation via Label Smoothing RegularizationabstractKnowledge Distillation (KD) aims to distill the knowledge of a cumbersome teacher model into a lightweight student model. Its success is generally attributed to the privileged information on similarities among categories provided by the teacher model, and in this sense, only strong teacher models are deployed to teach weaker students in practice. In this work, we challenge this common belief by following experimental observations: 1) beyond the acknowledgment that the teacher can improve the student, the student can also enhance the teacher significantly by reversing the KD procedure; 2) a poorly-trained teacher with much lower accuracy than the student can still improve the latter significantly. To explain these observations, we provide a theoretical analysis of the relationships between KD and label smoothing regularization. We prove that 1) KD is a type of learned label smoothing regularization and 2) label smoothing regularization provides a virtual teacher model for KD. From these results, we argue that the success of KD is not fully due to the similarity information between categories from teachers, but also to the regularization of soft targets, which is equally or even more important. Based on these analyses, we further propose a novel Teacher-free Knowledge Distillation (Tf-KD) framework, where a student model learns from itself or manually-designed regularization distribution. The Tf-KD achieves comparable performance with normal KD from a superior teacher, which is well applied when a stronger teacher model is unavailable. Meanwhile, Tf-KD is generic and can be directly deployed for training deep neural networks. Without any extra computation cost, Tf-KD achieves up to 0.65\% improvement on ImageNet over well-established baseline models, which is superior to label smoothing regularization. Li Yuan 0007, Francis E. H. Tay, Guilin Li 0001, Tao Wang 0053, Jiashi Feng |
CVPR | 1 |
| 2020 | Central Similarity Quantization for Efficient Image and Video RetrievalabstractExisting data-dependent hashing methods usually learn hash functions from pairwise or triplet data relationships, which only capture the data similarity locally, and often suffer from low learning efficiency and low collision rate. In this work, we propose a new \emph{global} similarity metric, termed as \emph{central similarity}, with which the hash codes of similar data pairs are encouraged to approach a common center and those for dissimilar pairs to converge to different centers, to improve hash learning efficiency and retrieval accuracy. We principally formulate the computation of the proposed central similarity metric by introducing a new concept, i.e., \emph{hash center} that refers to a set of data points scattered in the Hamming space with a sufficient mutual distance between each other. We then provide an efficient method to construct well separated hash centers by leveraging the Hadamard matrix and Bernoulli distributions. Finally, we propose the Central Similarity Quantization (CSQ) that optimizes the central similarity between data points w.r.t.\ their hash centers instead of optimizing the local similarity. CSQ is generic and applicable to both image and video hashing scenarios. Extensive experiments on large-scale image and video retrieval tasks demonstrate that CSQ can generate cohesive hash codes for similar data pairs and dispersed hash codes for dissimilar pairs, achieving a noticeable boost in retrieval performance, i.e. 3\%-20\% in mAP over the previous state-of-the-arts. Li Yuan 0007, Tao Wang 0053, Xiaopeng Zhang 0008, Francis E. H. Tay, Zequn Jie, Wei Liu 0005, Jiashi Feng |
CVPR | 1 |
| 2020 | A Simple Baseline for Pose Tracking in Videos of Crowed ScenesabstractThis paper presents our solution to ACM MM challenge: Large-scale Human-centric Video Analysis in Complex Events[13]; specifically, here we focus on Track3: Crowd Pose Tracking in Complex Events. Remarkable progress has been made in multi-pose training in recent years. However, how to track the human pose in crowded and complex environments has not been well addressed. We formulate the problem as several subproblems to be solved. First, we use a multi-object tracking method to assign human ID to each bounding box generated by the detection model. After that, a pose is generated to each bounding box with ID. At last, optical flow is used to take advantage of the temporal information in the videos and generate the final pose tracking result. Li Yuan 0007, Shuning Chang, Ziyuan Huang 0003, Xuecheng Nie, Francis E. H. Tay, Jiashi Feng, Shuicheng Yan |
ACM Multimedia | 1 |
| 2020 | Towards Accurate Human Pose Estimation in Videos of Crowded ScenesabstractVideo-based human pose estimation in crowed scenes is a challenging problem due to occlusion, motion blur, scale variation and viewpoint change, etc. Prior approaches always fail to deal with this problem because of (1) lacking of usage of temporal information; (2) lacking of training data in crowded scenes. In this paper, we focus on improving human pose estimation in videos of crowded scenes from the perspectives of exploiting temporal context and collecting new data. In particular, we first follow the top-down strategy to detect persons and perform single-person pose estimation for each frame. Then, we refine the frame-based pose estimation with temporal contexts deriving from the optical-flow. Specifically, for one frame, we forward the historical poses from the previous frames and backward the future poses from the subsequent frames to current frame, leading to stable and accurate human pose estimation in videos. In addition, we mine new data of similar scenes to HIE dataset from the Internet for improving the diversity of training set. In this way, our model achieves best performance on 7 out of 13 videos and 56.33 average wAP on test dataset of HIE challenge. Shuning Chang, Li Yuan 0007, Xuecheng Nie, Ziyuan Huang 0003, Yunpeng Chen, Jiashi Feng, Shuicheng Yan |
ACM Multimedia | 2 |
| 2020 | Toward Accurate Person-level Action Recognition in Videos of Crowed ScenesabstractDetecting and recognizing human action in videos with crowed scenes is a challenging problem due to the complex environment and diversity events. Prior works always fail to deal with this problem in two aspects: (1) lacking utilizing information of the scenes; (2) lacking training data in the crowd and complex scenes. In this paper, we focus on improving spatio-temporal action recognition by fully-utilizing the information of scenes and collecting new data. A top-down strategy is used to overcome the limitations. Specifically, we adopt a strong human detector to detect the spatial location of each frame. We then apply action recognition models to learn the spatio-temporal information from video frames on both the HIE dataset and new data with diverse scenes from the internet, which can improve the generalization ability of our model. Besides, the scenes information is extracted by the semantic segmentation model to assistant the process. As a result, our method achieved an average 26.05 wf\_mAP (ranking 1st place in the ACM MM grand challenge 2020: Human in Events). Li Yuan 0007, Shuning Chang, Ziyuan Huang 0003, Xuecheng Nie, Tao Wang 0053, Jiashi Feng, Shuicheng Yan |
ACM Multimedia | 1 |
| 2020 | Unsupervised Video Summarization With Cycle-Consistent Adversarial LSTM NetworksabstractVideo summarization is an important technique to browse, manage and retrieve a large amount of videos efficiently. The main objective of video summarization is to minimize the information loss when selecting a subset of video frames from the original video, hence the summary video can faithfully represent the overall story of the original video. Recently developed unsupervised video summarization approaches are free of requiring tedious annotation on important frames to train a video summarization model and thus are practically attractive. However, their performance is still limited due to the difficulty of minimizing information loss between the summary and original videos. In this paper, we address unsupervised video summarization by developing a novel Cycle-consistent Adversarial LSTM architecture to effectively reduce the information loss in the summary video. The proposed model, named Cycle-SUM, consists of a frame selector and a cycle-consistent learning based evaluator. The selector is a bi-directional LSTM network to capture the long-range relationship between video frames. To overcome the difficulty of specifying a suitable information preserving metric between original video and summary video, the evaluator is introduced to “supervise” selector to improve the video summarization quality. Specifically, the evaluator is composed of two generative adversarial networks (GANs), in which the forward GAN component is learned to reconstruct the original video from summary video, while the backward GAN learns to invert the process. We establish the relation between mutual information maximization and such cycle learning procedure and further introduce cycle-consistent loss to regularize the summarization. Extensive experiments on three video summarization benchmark datasets demonstrate a state-of-the-art performance, and show the superiority of the Cycle-SUM model compared with other unsupervised approaches. Li Yuan 0007, Francis E. H. Tay, Ping Li 0006, Jiashi Feng |
IEEE Trans. Multim. | 1 |
| 2019 | Cycle-SUM: Cycle-Consistent Adversarial LSTM Networks for Unsupervised Video SummarizationabstractIn this paper, we present a novel unsupervised video summarization model that requires no manual annotation. The proposed model termed Cycle-SUM adopts a new cycleconsistent adversarial LSTM architecture that can effectively maximize the information preserving and compactness of the summary video. It consists of a frame selector and a cycle-consistent learning based evaluator. The selector is a bi-direction LSTM network that learns video representations that embed the long-range relationships among video frames. The evaluator defines a learnable information preserving metric between original video and summary video and “supervises” the selector to identify the most informative frames to form the summary video. In particular, the evaluator is composed of two generative adversarial networks (GANs), in which the forward GAN is learned to reconstruct original video from summary video while the backward GAN learns to invert the processing. The consistency between the output of such cycle learning is adopted as the information preserving metric for video summarization. We demonstrate the close relation between mutual information maximization and such cycle learning procedure. Experiments on two video summarization benchmark datasets validate the state-of-theart performance and superiority of the Cycle-SUM model over previous baselines. Li Yuan 0007, Francis E. H. Tay, Ping Li 0006, Jiashi Feng |
AAAI | 1 |
| 2019 | Distilling Object Detectors With Fine-Grained Feature ImitationabstractState-of-the-art CNN based recognition models are often computationally prohibitive to deploy on low-end devices. A promising high level approach tackling this limitation is knowledge distillation, which let small student model mimic cumbersome teacher model's output to get improved generalization. However, related methods mainly focus on simple task of classification while do not consider complex tasks like object detection. We show applying the vanilla knowledge distillation to detection model gets minor gain. To address the challenge of distilling knowledge in detection model, we propose a fine-grained feature imitation method exploiting the cross-location discrepancy of feature response. Our intuition is that detectors care more about local near object regions. Thus the discrepancy of feature response on the near object anchor locations reveals important information of how teacher model tends to generalize. We design a novel mechanism to estimate those locations and let student model imitate the teacher on them to get enhanced performance. We first validate the idea on a developed lightweight toy detector which carries simplest notion of current state-of-the-art anchor based detection models on challenging KITTI dataset, our method generates up to 15% boost of mAP for the student model compared to the non-imitated counterpart. We then extensively evaluate the method with Faster R-CNN model under various scenarios with common object detection benchmark of Pascal VOC and COCO, imitation alleviates up to 74% performance drop of student model compared to teacher. Codes released at https://github.com/twangnh/Distilling-Object-Detectors. Tao Wang 0053, Li Yuan 0007, Xiaopeng Zhang 0008, Jiashi Feng |
CVPR | 2 |
| 2019 | Few-Shot Adaptive Faster R-CNNabstractTo mitigate the detection performance drop caused by domain shift, we aim to develop a novel few-shot adaptation approach that requires only a few target domain images with limited bounding box annotations. To this end, we first observe several significant challenges. First, the target domain data is highly insufficient, making most existing domain adaptation methods ineffective. Second, object detection involves simultaneous localization and classification, further complicating the model adaptation process. Third, the model suffers from over-adaptation (similar to overfitting when training with a few data example) and instability risk that may lead to degraded detection performance in the target domain. To address these challenges, we first introduce a pairing mechanism over source and target features to alleviate the issue of insufficient target domain samples. We then propose a bi-level module to adapt the source trained detector to the target domain: 1) the split pooling based image level adaptation module uniformly extracts and aligns paired local patch features over locations, with different scale and aspect ratio; 2) the instance level adaptation module semantically aligns paired object features while avoids inter-class confusion. Meanwhile, a source model feature regularization (SMFR) is applied to stabilize the adaptation process of the two modules. Combining these contributions gives a novel few-shot adaptive Faster-RCNN framework, termed FAFRCNN, which effectively adapts to target domain with a few labeled samples. Experiments with multiple datasets show that our model achieves new state-of-the-art performance under both the interested few-shot domain adaptation(FDA) and unsupervised domain adaptation(UDA) setting. Tao Wang 0053, Xiaopeng Zhang 0008, Li Yuan 0007, Jiashi Feng |
CVPR | 3 |