EDBT 2026 Demo / reviewers in the wild / expert
Kai Chen 0026
dblp:181/2839-26
· DBLP profile ↗
110ranked-venue papers
8as first author
102since 2021 · last 2026
0000-0002-6820-2325ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 95 · 3 first-author · 87 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 5 first-author · 44 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Powering Verifiable Learning via Automated Evolutionary Data SynthesisabstractReliable verifiable data has become a key driver of capability gains in modern language models, enabling stable reinforcement learning with verifiable rewards and effective distillation that transfers competence across math, coding, and agentic tasks.Yet constructing generalizable synthetic verifiable data remains difficult due to hallucination-prone generation, and weak or trivial verification artifacts that fail to separate strong from weak solutions.Existing approaches often rely on task-specific heuristics or post-hoc filters that do not transfer across domains and lack a principled, universal evaluator of verifiability.In this work, we introduce an evolutionary, task-agnostic, strategy-guided, executably-checkable data synthesis framework that, from minimal seed supervision, jointly synthesizes problems, diverse candidate solutions, and verification artifacts, and iteratively discovers strategies via a consistencybased evaluator that enforces agreement between human-annotated and strategy-induced checks.This pipeline upgrades filtering into principled synthesis: it reliably assembles coherent, verifiable training instances and generalizes without domain-specific rules.Our experiments demonstrate the effectiveness of the proposed approach under both RLVR and model distillation training paradigms.The results show that training with our synthesized data yields significant improvements on both the LiveCodeBench and AgentBench-OS tasks, highlighting the robust generalization of our framework 1 . He Du, Bowen Li 0002, Aijun Yang, Siyang He, Qipeng Guo, Kai Chen 0026, Dacheng Tao |
ACL (1) | 6 |
| 2026 | Timely Machine: Awareness of Time Makes Test-Time Scaling AgenticabstractYichuan Ma, Linyang Li, Yongkang Chen, Peiji Li, Xiaozhe Li, Qipeng Guo, Dahua Lin, Kai Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yichuan Ma, Linyang Li, Peiji Li, Xiaozhe Li, Qipeng Guo, Dahua Lin, Kai Chen 0026 |
ACL (1) | 8 |
| 2026 | Efficient Personalized Reranking with Semi-autoregressive Generation and Online Knowledge Distillation
Kai Chen 0026, Wei Guo 0006, Weiwen Liu, Yong Liu 0020, Enhong Chen |
DASFAA (1) | 1 |
| 2026 | MDTace: Agentic Context Engineering for Multi-disciplinary Team Medical Consultation
Kai Chen 0026, Yang Gao 0001 |
ICIC (8) | 1 |
| 2026 | MedMentor: Teacher-Student Collaboration with Decoupled Experience for Clinical Diagnosis
Kai Chen 0026, Yang Gao 0001 |
ICIC (8) | 1 |
| 2026 | FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Yiming Zhang 0031, Yicheng Gu, Yanhong Zeng, Zhening Xing, Yuancheng Wang, Zhizheng Wu 0001, Kai Chen 0026 |
Int. J. Comput. Vis. | 8 |
| 2026 | EnrichGAN: Exploiting enriched discriminator representations for training GANs under limited data
Wenhao Mu, Kai Chen 0026, Lizhuang Ma, Nan Wang 0027, Qingchao Jiang, Bingcang Huang |
Neurocomputing | 2 |
| 2026 | StyleShot: A Snapshot on Any StyleabstractImage Style Transfer aims to replicate the style of a reference image based on the content from a text description or another image. With the significant advancements in image generation through diffusion models, recent studies have attempted to either fine-tuning embeddings to learn the single style or utilizing the pre-trained CLIP image encoder to extract style representations. However, style-tuning requires substantial computational resources and the pre-trained CLIP image encoder is trained for semantic understanding rather than for style representation. To address these challenges, we introduce a style-aware encoder and a well-organized style dataset called StyleGallery to learn a good style representation that is crucial and sufficient for generalized style transfer without test-time tuning. With dedicated design for style learning, this style-aware encoder is trained to extract expressive style representation from multi-level patches with decoupling training strategy, and StyleGallery enables the generalization ability. Moreover, we employ a content extraction and content-fusion encoder to enhance image-driven style transfer. We highlight that, our approach, named StyleShot, is simple yet effective in mimicking various desired styles, i.e., 3D, flat, abstract or even fine-grained styles, without test-time tuning. Rigorous experiments validate that, StyleShot achieves superior performance across a wide range of styles compared to existing state-of-the-art text- and image-driven methods. Junyao Gao 0002, Yanan Sun 0005, Yinhao Tang, Yanhong Zeng, Ding Qi, Kai Chen 0026, Cairong Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous DrivingabstractRecent advancements in vision foundation models (VFMs) have revolutionized visual perception in 2D, yet their potential for 3D scene understanding, particularly in autonomous driving applications, remains underexplored. In this paper, we introduce LargeAD, a versatile and scalable framework designed for large-scale 3D pretraining across diverse real-world driving datasets. Our framework leverages VFMs to extract semantically rich superpixels from 2D images, which are aligned with LiDAR point clouds to generate high-quality contrastive samples. This alignment facilitates cross-modal representation learning, enhancing the semantic consistency between 2D and 3D data. We introduce several key innovations: (i) VFM-driven superpixel generation for detailed semantic representation, (ii) a VFM-assisted contrastive learning strategy to align multimodal features, (iii) superpoint temporal consistency to maintain stable representations across time, and (iv) multi-source data pretraining to generalize across various LiDAR configurations. Our approach achieves substantial gains over state-of-the-art methods in linear probing and fine-tuning for LiDAR-based segmentation and object detection. Extensive experiments on 11 large-scale multi-sensor datasets highlight our superior performance, demonstrating adaptability, efficiency, and robustness in real-world autonomous driving scenarios. Lingdong Kong, Xiang Xu 0009, Youquan Liu, Jun Cen, Runnan Chen, Liang Pan, Kai Chen 0026, Ziwei Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | Enhanced Spatiotemporal Consistency for Image-to-LiDAR Data PretrainingabstractLiDAR representation learning has emerged as a promising approach to reducing reliance on costly and labor-intensive human annotations. While existing methods primarily focus on spatial alignment between LiDAR and camera sensors, they often overlook the temporal dynamics critical for capturing motion and scene continuity in driving scenarios. To address this limitation, we propose SuperFlow++, a novel framework that integrates spatiotemporal cues in both pretraining and downstream tasks using consecutive LiDAR-camera pairs. SuperFlow++ introduces four key components: (1) a view consistency alignment module to unify semantic information across camera views, (2) a dense-to-sparse consistency regularization mechanism to enhance feature robustness across varying point cloud densities, (3) a flow-based contrastive learning approach that models temporal relationships for improved scene understanding, and (4) a temporal voting strategy that propagates semantic information across LiDAR scans to improve prediction consistency. Extensive evaluations on 11 heterogeneous LiDAR datasets demonstrate that SuperFlow++ outperforms state-of-the-art methods across diverse tasks and driving conditions. Furthermore, by scaling both 2D and 3D backbones during pretraining, we uncover emergent properties that provide deeper insights into developing scalable 3D foundation models. With strong generalizability and computational efficiency, SuperFlow++ establishes a new benchmark for data-efficient LiDAR-based perception in autonomous driving. Xiang Xu 0009, Lingdong Kong, Hui Shuai, Liang Pan, Kai Chen 0026, Ziwei Liu 0002, Qingshan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Double Buffer Vaccination: Bolstering Immunity Against Catastrophic Forgetting in Continual Medical Image Segmentation
Kai Chen 0026, Hewei Wang 0001, Pinzhuo Tian, Jing Huo, Taihang Zhen, Yang Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | MG-LLaVA: Toward Multi-Granularity Visual Instruction TuningabstractMulti-modal large language models (MLLMs) have made significant strides in various visual understanding tasks. However, the majority of these models are constrained to process low-resolution images, which limits their effectiveness in perception tasks that necessitate detailed visual information. In our study, we present MG-LLaVA, an innovative MLLM that enhances the model’s visual processing capabilities by incorporating a multi-granularity vision flow, which includes low-resolution, high-resolution, and object-centric features. We propose the integration of an additional high-resolution visual encoder to capture fine-grained details, which are then fused with base visual features through a Conv-Gate fusion network. To further refine the model’s object recognition abilities, we incorporate object-level features derived from bounding boxes identified by offline detectors. Being trained solely on publicly available multimodal data through instruction tuning, MG-LLaVA demonstrates exceptional perception skills. We instantiate MG-LLaVA with a wide variety of language encoders, ranging from 3.8B to 34B, to evaluate the model’s performance comprehensively. Extensive evaluations across multiple benchmarks demonstrate that MG-LLaVA outperforms existing MLLMs of comparable parameter sizes, showcasing its remarkable efficacy. Xiangtai Li, Haodong Duan, Haian Huang, Kai Chen 0026, Hua Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best PracticesabstractZhi Chen, Qiguang Chen, Libo Qin, Qipeng Guo, Haijun Lv, Yicheng Zou, Hang Yan, Kai Chen, Dahua Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhi Chen 0006, Qiguang Chen, Libo Qin 0001, Qipeng Guo, Haijun Lv, Yicheng Zou, Hang Yan 0001, Kai Chen 0026, Dahua Lin |
ACL (1) | 8 |
| 2025 | Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling LawabstractScaling law builds the relationship between training computation and validation loss, enabling researchers to effectively predict the loss trending of models across different levels of computation. However, a gap still remains between validation loss and the model’s downstream capabilities, making it untrivial to apply scaling law to direct performance prediction for downstream tasks. The loss typically represents a cumulative penalty for predicted tokens, which are implicitly considered to have equal importance. Nevertheless, our studies have shown evidence that when considering different training data distributions, we cannot directly model the relationship between downstream capability and computation or token loss. To bridge the gap between validation loss and downstream task capabilities, in this work, we introduce Capability Salience Vector, which decomposes the overall loss and assigns different importance weights to tokens to assess a specific meta-capability, aligning the validation loss with downstream task performance in terms of the model’s capabilities. Experiments on various popular benchmarks demonstrate that our proposed Capability Salience Vector could significantly improve the predictability of language model performance on downstream tasks. Qiming Ge, Shuhao Xing, Songyang Gao, Yunhua Zhou, Yicheng Zou, Songyang Zhang 0001, Zhi Chen 0006, Hang Yan 0001, Qi Zhang 0001, Qipeng Guo, Kai Chen 0026 |
ACL (1) | 11 |
| 2025 | CritiQ: Mining Data Quality Criteria from Human PreferencesabstractHonglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun, Kai Chen, Xipeng Qiu, Tao Gui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Honglin Guo, Kai Lv 0001, Qipeng Guo, Tianyi Liang 0002, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun 0031, Kai Chen 0026, Xipeng Qiu, Tao Gui |
ACL (1) | 9 |
| 2025 | Scaling up the State Size of RNN LLMs for Long-Context ScenariosabstractThe Transformer architecture has become the standard LLM architecture due to its powerful self-attention mechanism. However, it suffers from quadratic computational complexity and linear memory complexity. RNN-based LLMs have been proposed as alternatives. Yet, RNN models struggle in long-context scenarios, making it challenging to replace self-attention with RNNs. We identify the state size as a critical bottleneck, which is significantly smaller than that of Transformers with a basic context length of 2k. However, simply increasing the state size significantly raises the number of parameters and lowers training efficiency. In this paper, we propose an efficient scaling method to scale the state size of RNN models to match the 2k context length of Transformers, with small parameters overhead. Experimental results demonstrate that scaling the state size significantly enhances long-context understanding. Retrieval performance scales almost linearly with state size, with a 454M model featuring an expanded state achieving performance comparable to a 1.47B model on FDA, a recall-intensive task. These findings highlight state scaling as a promising approach for advancing RNN-based LLMs. Jianfei Gao 0003, Kai Chen 0026 |
ACL (1) | 3 |
| 2025 | Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and RefinementabstractMaosongcao Maosongcao, Taolin Zhang, Mo Li, Chuyu Zhang, Yunxin Liu, Conghui He, Haodong Duan, Songyang Zhang, Kai Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Maosongcao, Taolin Zhang 0003, Mo Li 0012, Chuyu Zhang, Conghui He, Haodong Duan, Songyang Zhang 0001, Kai Chen 0026 |
ACL (1) | 9 |
| 2025 | Redundancy Principles for MLLMs BenchmarksabstractZicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, Guangtao Zhai. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xinyu Fang, Chunyi Li 0001, Xiaohong Liu 0001, Xiongkuo Min, Haodong Duan, Kai Chen 0026, Guangtao Zhai |
ACL (1) | 8 |
| 2025 | OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human PreferenceabstractXiangyu Zhao, Shengyuan Ding, Zicheng Zhang, Haian Huang, Maosongcao Maosongcao, Jiaqi Wang, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang, Haodong Duan, Kai Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shengyuan Ding, Haian Huang, Maosongcao, Jiaqi Wang 0003, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang 0001, Haodong Duan, Kai Chen 0026 |
ACL (1) | 13 |
| 2025 | Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case StudyabstractRecent advancements in large language models (LLMs) have significantly enhanced their coding capabilities. However, existing benchmarks predominantly focused on simplified or isolated aspects of coding, such as single-file code generation or repository issue debugging, falling short of measuring the full spectrum of challenges raised by real-world programming activities. In this case study, we explore the performance of LLMs across the entire software development lifecycle with DevEval, encompassing stages including software design, environment setup, implementation, acceptance testing, and unit testing. DevEval features four programming languages, multiple domains, high-quality data collection, and carefully designed and verified metrics for each task. Empirical studies show that current LLMs, including GPT-4, fail to solve the challenges presented within DevEval. Our findings offer actionable insights for the future development of LLMs toward real-world programming applications. Bowen Li 0002, Ziwei Tang, John Yang 0002, Jinyang Li 0003, Shunyu Yao 0006, Chen Qian 0006, Binyuan Hui, Qicheng Zhang, Zhiyin Yu, He Du, Dahua Lin, Chao Peng 0002, Kai Chen 0026 |
COLING | 16 |
| 2025 | Auto Cherry-Picker: Learning from High-quality Generative Data Driven by LanguageabstractDiffusion models can generate realistic and diverse images, potentially facilitating data availability for data-intensive perception tasks. However, leveraging these models to boost performance on downstream tasks with synthetic data poses several challenges, including aligning with real data distribution, scaling synthetic sample volumes, and ensuring their quality. To bridge these gaps, we present Auto Cherry-Picker (ACP), a novel framework that generates high-quality cross-modality training samples at scale to augment perception and multi-modal training. ACP first uses LLMs to sample descriptions and layouts based on object combinations from real data priors, eliminating the need for ground truth image captions or annotations. Next, we use an off-the-shelf controllable diffusion model to generate multiple images. Then, the generated data are refined using a comprehensively designed metric, Composite Layout and Image Score (CLIS), to ensure quality. Our customized synthetic high-quality samples boost performance in various scenarios, especially in addressing challenges associated with long-tailed distribution and imbalanced datasets. Experiment results on downstream tasks demonstrate that ACP can significantly improve the performance of existing models. In addition, we find a positive correlation between CLIS and performance gains in downstream tasks. This finding shows the potential for evaluation metrics as the role for various visual perception and MLLM tasks. Xiangtai Li, Yanhong Zeng, Jianzong Wu, Kai Chen 0026 |
CVPR | 7 |
| 2025 | CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome RewardabstractShudong Liu, Hongwei Liu, Junnan Liu, Linchen Xiao, Songyang Gao, Chengqi Lyu, Yuzhe Gu, Wenwei Zhang, Derek F. Wong, Songyang Zhang, Kai Chen. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Shudong Liu 0007, Linchen Xiao, Songyang Gao, Chengqi Lyu, Yuzhe Gu, Derek F. Wong, Songyang Zhang 0001, Kai Chen 0026 |
EMNLP | 11 |
| 2025 | UnitCoder: Scalable Code Synthesis from Pre-training CorporaabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities in various tasks, yet code generation remains a major challenge.Despite the abundant sources of code data, constructing high-quality training datasets at scale poses a significant challenge.Pre-training code data typically suffers from inconsistent data quality issues.Conversely, instruction-based methods which use a high-quality subset as seed samples suffer from limited task diversity.In this paper, we introduce UnitCoder, which directly supervises pre-training data quality through automatically generated unit tests, while ensuring the correctness via an iterative fix and refine flow.Code synthesized by Unit-Coder benefits from both the diversity of pretraining corpora and the high quality ensured by unit test supervision.Our experiments demonstrate that models fine-tuned on our synthetic dataset exhibit consistent performance improvements.Our work presents a scalable approach that leverages model-generated unit tests to guide the synthesis of high-quality code data from pre-training corpora, demonstrating the potential for producing diverse and high-quality post-training data at scale.All code and data will be released 1 . Yichuan Ma, Yunfan Shao, Peiji Li, Demin Song, Qipeng Guo, Linyang Li, Xipeng Qiu, Kai Chen 0026 |
EMNLP | 8 |
| 2025 | A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language ModelsabstractWe propose a multi-agent approach (SeM-Agents) based on large language models for medical consultations. This framework incorporates various doctor roles and auxiliary roles, with agents communicating through natural language. Using a residual structure, the system conducts multi-round medical consultations based on the patient’s treatment background and symptoms. In the final summary and output stage of the consultation, it utilizes two experience databases—the Correct Consultation Experience Database and the Chain of Thought (CoT) Experience Database—which evolve with accumulated experience during consultations. This evolution drives the framework’s self-improvement, significantly enhancing the rationality and accuracy of the consultations. To ensure that the conclusions are safe, reliable, and aligned with human values, the final decisions undergo a safety review before being provided to the patient. This framework achieved accuracy rates of 89.2% and 83.1% on the MedQA and PubMedQA datasets, respectively. Kai Chen 0026, Jing Huo, Pinzhuo Tian, Yang Gao 0001 |
ICASSP | 1 |
| 2025 | Exposure-Limited Image Enhancement with Generative Diffusion PriorabstractMany consumer cameras are equipped with 8-bit image sensors, which often struggle to capture scenes with a High Dynamic Range (HDR). This limitation can result in overexposed or underexposed regions, a loss of fine details due to low bit-depth compression, skewed color distributions, and noticeable noise in dark areas. Traditional Standard Dynamic Range (SDR) image enhancement methods typically focus on color mapping by expanding the color range and adjusting brightness. However, they often fail to restore details in dynamic range extremes, i.e. regions where pixel values approach the minimum or maximum limits. We define “exposure-limited image enhancement” as the process of enhancing images with large missing areas due to exposure issues within the SDR space, which differs from existing “mis-exposed image enhancement” methods primarily aimed at correcting color distributions. To enhance these exposure-limited images and overcome the limitations of current models, we propose a novel two-stage approach. In the first stage, we remap color and brightness to a suitable range while preserving existing details. In the second stage, we use a diffusion prior to generate content in severely overexposed or underexposed regions, which are otherwise lost during capture. Notably, this generative refinement module can also serve as a plug-and-play component alongside existing enhancement methods. Extensive experiments demonstrate that our method significantly improves image quality and detail, outperforming state-of-the-art techniques in dynamic range extremes. The project page is at https://Sagiri0208.github.io. Baiang Li, Sizhuo Ma, Yanhong Zeng, Xiaogang Xu 0002, Youqing Fang, Zhao Zhang 0001, Jian Wang 0100, Kai Chen 0026 |
ICCP | 8 |
| 2025 | Creation-Mmbench: Assessing Context-Aware Creative Intelligence in Mllms
Xinyu Fang, Kai Lan, Lixin Ma, Shengyuan Ding, Yingji Liang, Farong Wen, Guofeng Zhang 0001, Haodong Duan, Kai Chen 0026, Dahua Lin |
ICCV | 12 |
| 2025 | Information Density Principle for MLLM BenchmarksabstractWith the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and whether the test results meet their requirements. Therefore, we propose a critical principle of Information Density, which examines how much insight a benchmark can provide for the development of MLLMs. We characterize it from four key dimensions: (1) Fallacy, (2) Difficulty, (3) Redundancy, (4) Diversity. Through a comprehensive analysis of more than 10,000 samples, we measured the information density of 19 MLLM benchmarks. Experiments show that using the latest benchmarks in testing can provide more insight compared to previous ones, but there is still room for improvement in their information density. We hope this principle can promote the development and application of future MLLM benchmarks. Project page: https://github.com/lcysyzxdxc/bench4bench Chunyi Li 0001, Xiaozhe Li, Yuan Tian 0017, Ziheng Jia, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Haodong Duan, Kai Chen 0026, Guangtao Zhai |
ICCV | 10 |
| 2025 | MotionShot: Adaptive Motion Transfer Across Arbitrary Objects for Text-to-Video GenerationabstractExisting text-to-video methods struggle to transfer motion smoothly from a reference object to a target object with significant differences in appearance or structure between them. To address this challenge, we introduce MotionShot, a training-free framework capable of parsing reference-target correspondences in a fine-grained manner, thereby achieving high-fidelity motion transfer while preserving coherence in appearance. To be specific, MotionShot first performs semantic feature matching to ensure high-level alignments between the reference and target objects. It then further establishes low-level morphological alignments through reference-to-target shape retargeting. By encoding motion with temporal attention, our MotionShot can coherently transfer motion across objects, even in the presence of significant appearance and structure disparities, demonstrated by extensive experiments. The project page is available at: https://motionshot.github.io/. Yanan Sun 0005, Zhening Xing, Junyao Gao 0002, Kai Chen 0026, Wenjie Pei |
ICCV | 5 |
| 2025 | FaceShot: Bring Any Character into LifeabstractIn this paper, we present ***FaceShot***, a novel training-free portrait animation framework designed to bring any character into life from any driven video without fine-tuning or retraining.
We achieve this by offering precise and robust reposed landmark sequences from an appearance-guided landmark matching module and a coordinate-based landmark retargeting module.
Together, these components harness the robust semantic correspondences of latent diffusion models to produce facial motion sequence across a wide range of character types.
After that, we input the landmark sequences into a pre-trained landmark-driven animation model to generate animated video.
With this powerful generalization capability, FaceShot can significantly extend the application of portrait animation by breaking the limitation of realistic portrait landmark detection for any stylized character and driven video.
Also, FaceShot is compatible with any landmark-driven animation model, significantly improving overall performance.
Extensive experiments on our newly constructed character benchmark CharacBench confirm that FaceShot consistently surpasses state-of-the-art (SOTA) approaches across any character domain.
More results are available at our project website https://faceshot2024.github.io/faceshot/. Junyao Gao 0002, Yanan Sun 0005, Fei Shen 0004, Xin Jiang 0010, Zhening Xing, Kai Chen 0026, Cairong Zhao |
ICLR | 6 |
| 2025 | MindSearch: Mimicking Human Minds Elicits Deep AI SearcherabstractInformation seeking and integration is a complex cognitive task that consumes enormous time and effort. Inspired by the remarkable progress of Large Language Models, recent works attempt to solve this task by combining LLMs and search engines. However, these methods still obtain unsatisfying performance due to three challenges: (1) complex requests often cannot be accurately and completely retrieved by the search engine once (2) corresponding information to be integrated is spread over multiple web pages along with massive noise, and (3) a large number of web pages with long contents may quickly exceed the maximum context length of LLMs. Inspired by the cognitive process when humans solve these problems, we introduce MindSearch to mimic the human minds in web information seeking and integration, which can be instantiated by a simple yet effective LLM-based multi-agent framework. The WebPlanner models the human mind of multi-step information seeking as a dynamic graph construction process: it decomposes the user query into atomic sub-questions as nodes in the graph and progressively extends the graph based on the search result from WebSearcher. Tasked with each sub-question, WebSearcher performs hierarchical information retrieval with search engines and collects valuable information for WebPlanner. The multi-agent design of MindSearch enables the whole framework to seek and integrate information parallelly from larger-scale (e.g., more than 300) web pages in 3 minutes, which is worth 3 hours of human effort. MindSearch demonstrates significant improvement in the response quality in terms of depth and breadth, on both close-set and open-set QA problems. Besides, responses from MindSearch based on InternLM2.5-7B are preferable by humans to ChatGPT-Web and Perplexity.ai applications, which implies that MindSearch can already deliver a competitive solution to the proprietary AI search engine. Kuikun Liu, Qiuchen Wang, Jiangning Liu, Kai Chen 0026, Feng Zhao 0004 |
ICLR | 6 |
| 2025 | Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMsabstractLarge language models (LLMs) exhibit hallucinations (i.e., unfaithful or nonsensical information) when serving as AI assistants in various domains. Since hallucinations always come with truthful content in the LLM responses, previous factuality alignment methods that conduct response-level preference learning inevitably introduced noises during training. Therefore, this paper proposes a fine-grained factuality alignment method based on Direct Preference Optimization (DPO), called Mask-DPO. Incorporating sentence-level factuality as mask signals, Mask-DPO only learns from factually correct sentences in the preferred samples and prevents the penalty on factual contents in the not preferred samples, which resolves the ambiguity in the preference learning. Extensive experimental results demonstrate that Mask-DPO can significantly improve the factuality of LLMs responses to questions from both in-domain and out-of-domain datasets, although these questions and their corresponding topics are unseen during training. Only trained on the ANAH train set, the score of Llama3.1-8B-Instruct on the ANAH test set is improved from 49.19% to 77.53%, even surpassing the score of Llama3.1-70B-Instruct (53.44%), while its FactScore on the out-of-domain Biography dataset is also improved from 30.29% to 39.39%. We further study the generalization property of Mask-DPO using different training sample scaling strategies and find that scaling the number of topics in the dataset is more effective than the number of questions. We provide a hypothesis of what factual alignment is doing with LLMs, on the implication of this phenomenon, and conduct proof-of-concept experiments to verify it. We hope the method and the findings pave the way for future research on scaling factuality alignment. Yuzhe Gu, Chengqi Lyu, Dahua Lin, Kai Chen 0026 |
ICLR | 5 |
| 2025 | RMP-SAM: Towards Real-Time Multi-Purpose Segment AnythingabstractRecent segmentation methods, which adopt large-scale data training and transformer architecture, aim to create one foundation model that can perform multiple tasks.
However, most of these methods rely on heavy encoder and decoder frameworks, hindering their performance in real-time scenarios.
To explore real-time segmentation, recent advancements primarily focus on semantic segmentation within specific environments, such as autonomous driving. However, they often overlook the generalization ability of these models across diverse scenarios.
Therefore, to fill this gap, this work explores a novel real-time segmentation setting called real-time multi-purpose segmentation.
It contains three fundamental sub-tasks: interactive segmentation, panoptic segmentation, and video instance segmentation.
Unlike previous methods, which use a specific design for each task, we aim to use only a single end-to-end model to accomplish all these tasks in real-time.
To meet real-time requirements and balance multi-task learning, we present a novel dynamic convolution-based method, Real-Time Multi-Purpose SAM (RMP-SAM).
It contains an efficient encoder and an efficient decoupled adapter to perform prompt-driven decoding.
Moreover, we further explore different training strategies and one new adapter design to boost co-training performance further.
We benchmark several strong baselines by extending existing works to support our multi-purpose segmentation.
Extensive experiments demonstrate that RMP-SAM is effective and generalizes well on proposed benchmarks and other specific semantic tasks.
Our implementation of RMP-SAM achieves the optimal balance between accuracy and speed for these tasks. The code is released at
\url{https://github.com/xushilin1/RAP-SAM} Shilin Xu 0001, Haobo Yuan, Lu Qi 0001, Jingbo Wang 0001, Kai Chen 0026, Yunhai Tong, Bernard Ghanem, Xiangtai Li, Ming-Hsuan Yang 0001 |
ICLR | 8 |
| 2025 | Spherical Scissor-Like Reconfigurable Palm Design in Robotic Hands: Insights from Human Hand FunctionalityabstractThe human palm demonstrates spatial reconfigurability during the gripping process and forms a spherical grasping envelope. Based on these observations, this study designs a reconfigurable spherical palm that incorporates a spatial scissor mechanism, which only requires a single actuator to reshape the palm into a range of spherical forms. We conduct a kinematic analysis and modelling of the structure, abstracting three key parameters and analysing their influence on the motion characteristics of the palm. Through multi-objective optimisation, a set of dimensional parameters is derived to balance workspace, human-like motion, and mechanical performance. The performance of the reconfigurability and the grasping capability of the proposed palm is compared to a planar folding palm by superquadrics, and the results show that the spherical design and the reconfigurable characteristics provide larger grasping arrangement and stronger grasping capability of the palm on most of the testing surfaces. Kai Chen 0026, Chang Liu 0030, Guoniu Zhu, Qiujie Lu, Zhongxue Gan 0001 |
IROS | 3 |
| 2025 | Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language ModelsabstractLarge Language Models (LLMs) have shown strong abilities in general language tasks, yet adapting them to specific domains remains a challenge.
Current method like Domain Adaptive Pretraining (DAPT) requires costly full-parameter training and suffers from catastrophic forgetting.
Meanwhile, Retrieval-Augmented Generation (RAG) introduces substantial inference latency due to expensive nearest-neighbor searches and longer context.
This paper introduces \textit{Memory Decoder}, a plug-and-play pretrained memory that enables efficient domain adaptation without changing the original model's parameters.
Memory Decoder employs a small transformer decoder that learns to imitate the behavior of an external non-parametric retriever.
Once trained, Memory Decoder can be seamlessly integrated with any pretrained language model that shares the same tokenizer, requiring no model-specific modifications.
Experimental results demonstrate that Memory Decoder enables effective adaptation of various Qwen and Llama models to three distinct specialized domains: biomedicine, finance, and law, reducing perplexity by an average of 6.17 points.
Overall, Memory Decoder introduces a novel paradigm centered on a specially pretrained memory component designed for domain-specific adaptation. This memory architecture can be integrated in a plug-and-play manner, consistently enhancing performance across multiple models within the target domain. Jiaqi Cao 0002, Rubin Wei, Qipeng Guo, Kai Chen 0026, Bowen Zhou 0002, Zhouhan Lin |
NeurIPS | 5 |
| 2025 | Pre-Trained Policy Discriminators are General Reward ModelsabstractWe offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance.
For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines.
POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks.
Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99.
The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models. Shihan Dou, Shichun Liu, Yuming Yang 0001, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Haijun Lv, Demin Song, Songyang Gao, Chengqi Lyu, Enyu Zhou, Honglin Guo, Zhiheng Xi, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Kai Chen 0026 |
NeurIPS | 22 |
| 2025 | Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of GoabstractLarge language models (LLMs) have demonstrated exceptional performance in reasoning tasks such as mathematics and coding, matching or surpassing human capabilities. However, these impressive reasoning abilities face significant challenges in specialized domains. Taking Go as an example, although AlphaGo has established the high performance ceiling of AI systems in Go, mainstream LLMs still struggle to reach even beginner-level proficiency, let alone perform natural language reasoning. This performance gap between general-purpose LLMs and domain experts is significantly limiting the application of LLMs on a wider range of domain-specific tasks. In this work, we aim to bridge the divide between LLMs' general reasoning capabilities and expert knowledge in domain-specific tasks. We perform mixed fine-tuning with structured Go expertise and general long Chain-of-Thought (CoT) reasoning data as a cold start, followed by reinforcement learning to integrate expert knowledge in Go with general reasoning capabilities. Through this methodology, we present LoGos, a powerful LLM that not only maintains outstanding general reasoning abilities, but also conducts Go gameplay in natural language, demonstrating effective strategic reasoning and accurate next-move prediction. LoGos achieves performance comparable to human professional players, substantially surpassing all existing LLMs. Through this work, we aim to contribute insights on applying general LLM reasoning capabilities to specialized domains. We will release the first large-scale Go dataset for LLM training, the first LLM Go evaluation benchmark, and the first general LLM that reaches human expert-level performance in Go. Yichuan Ma, Linyang Li, Peiji Li, Jiasheng Ye, Qipeng Guo, Dahua Lin, Kai Chen 0026 |
NeurIPS | 8 |
| 2025 | Rethinking Verification for LLM Code Generation: From Generation to TestingabstractLarge language models (LLMs) have recently achieved notable success in code‑generation benchmarks such as HumanEval and LiveCodeBench. However, a detailed examination reveals that these evaluation suites often comprise only a limited number of homogeneous test cases, resulting in subtle faults going undetected. This not only artificially inflates measured performance but also compromises accurate reward estimation in reinforcement learning frameworks utilizing verifiable rewards (RLVR). To address these critical shortcomings, we systematically investigate the test-case generation (TCG) task by proposing multi-dimensional metrics designed to rigorously quantify test-suite thoroughness. Furthermore, we introduce a human-LLM collaborative method (SAGA), leveraging human programming expertise with LLM reasoning capability, aimed at significantly enhancing both the coverage and the quality of generated test cases. In addition, we develop a TCGBench to facilitate the study of the TCG task. Experiments show that SAGA achieves a detection rate of 90.62\% and a verifier accuracy of 32.58\% on TCGBench. The Verifier Accuracy (Verifier Acc) of the code generation evaluation benchmark synthesized by SAGA is 10.78\% higher than that of LiveCodeBench-v6. These results demonstrate the effectiveness of our proposed method. We hope this work contributes to building a scalable foundation for reliable LLM code evaluation, further advancing RLVR in code generation, and paving the way for automated adversarial test synthesis and adaptive benchmark integration. Zihan Ma 0010, Taolin Zhang 0003, Maosongcao, Minnan Luo, Songyang Zhang 0001, Kai Chen 0026 |
NeurIPS | 8 |
| 2025 | Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking ReasoningabstractEnhancing large vision-language models (LVLMs) with visual slow-thinking reasoning is crucial for solving complex multimodal tasks. However, since LVLMs are mainly trained with vision-language alignment, it is difficult to adopt on-policy reinforcement learning (RL) to develop the slow thinking ability because the rollout space is restricted by its initial abilities. Off-policy RL offers a way to go beyond the current policy, but directly distilling trajectories from external models may cause visual hallucinations due to mismatched visual perception abilities across models. To address these issues, this paper proposes **SOPHIA**, a simple and scalable **S**emi-**O**ff-**P**olicy RL for vision-language slow-t**HI**nking re**A**soning. SOPHIA builds a semi-off-policy behavior model by combining on-policy visual understanding from a trainable LVLM with off-policy slow-thinking reasoning from a language model, assigns outcome-based rewards to reasoning, and propagates visual rewards backward. Then LVLM learns slow-thinking reasoning ability from the obtained reasoning trajectories using propagated rewards via off-policy RL algorithms. Extensive experiments with InternVL2.5 and InternVL3.0 with 8B and 38B sizes show the effectiveness of SOPHIA. Notably, SOPHIA improves InternVL3.0-38B by 8.50\% in average, reaching state-of-the-art performance among open-source LVLMs on multiple multimodal reasoning benchmarks, and even outperforms some closed-source models (e.g., GPT-4.1) on the challenging MathVision and OlympiadBench, achieving 49.08\% and 49.95\% pass@1 accuracy, respectively. Analysis shows SOPHIA outperforms supervised fine-tuning and direct on-policy RL methods, offering a better policy initialization for further on-policy training. Haiteng Zhao, Yuzhe Gu, Songyang Gao, Kuikun Liu, Haian Huang, Jianfei Gao 0003, Dahua Lin, Kai Chen 0026 |
NeurIPS | 10 |
| 2025 | Calib3D: Calibrating Model Preferences for Reliable 3D Scene UnderstandingabstractSafety-critical 3D scene understanding tasks necessitate not only accurate but also confident predictions from 3D perception models. This study introduces Calib3D, a pioneering effort to benchmark and scrutinize the reliability of 3D scene understanding models from an uncertainty estimation viewpoint. We comprehensively evaluate 28 state of the art models across 10 diverse 3D datasets, uncovering insightful phenomena that cope with both the aleatoric and epistemic uncertainties in 3D scene understanding. We discover that despite achieving impressive levels of accuracy, existing models frequently fail to provide reliable uncertainty estimates - a pitfall that critically undermines their applicability in safety-sensitive contexts. Through extensive analysis of key factors such as network capacity, LiDAR representations, rasterization resolutions, and 3D data augmentation techniques, we correlate these aspects directly with the model calibration efficacy. Furthermore, we introduce DeptS, a novel depth-aware scaling approach aimed at enhancing 3D model calibration. Extensive experiments across a wide range of configurations validate the superiority of our method. We hope this work could serve as a cornerstone for fostering reliable 3D scene understanding. Code and benchmark toolkit are publicly available11https://github.com/ldkong1205/Calib3D. Lingdong Kong, Xiang Xu 0009, Jun Cen, Liang Pan, Kai Chen 0026, Ziwei Liu 0002 |
WACV | 6 |
| 2025 | Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous DrivingabstractEfficient data utilization is crucial for advancing 3D scene understanding in autonomous driving, where reliance on heavily human-annotated LiDAR point clouds challenges fully supervised methods. Addressing this, our study extends into semi-supervised learning for LiDAR semantic segmentation, leveraging the intrinsic spatial priors of driving scenes and multi-sensor complements to augment the efficacy of unlabeled datasets. We introduce LaserMix++, an evolved framework that integrates laser beam manipulations from disparate LiDAR scans and incorporates LiDAR-camera correspondences to further assist data-efficient learning. Our framework is tailored to enhance 3D scene consistency regularization by incorporating multi-modality, including 1) multi-modal LaserMix operation for fine-grained cross-sensor interactions; 2) camera-to-LiDAR feature distillation that enhances LiDAR feature learning; and 3) language-driven knowledge guidance generating auxiliary supervisions using open-vocabulary models. The versatility of LaserMix++ enables applications across LiDAR representations, establishing it as a universally applicable solution. Our framework is rigorously validated through theoretical analysis and extensive experiments on popular driving perception datasets. Results demonstrate that LaserMix++ markedly outperforms fully supervised alternatives, achieving comparable accuracy with five times fewer annotations and significantly improving the supervised-only baselines. This substantial advancement underscores the potential of semi-supervised approaches in reducing the reliance on extensive labeled data in LiDAR-based 3D scene understanding systems. Lingdong Kong, Xiang Xu 0009, Jiawei Ren 0001, Liang Pan, Kai Chen 0026, Wei Tsang Ooi, Ziwei Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Benchmarking and Improving Bird's Eye View Perception Robustness in Autonomous DrivingabstractRecent advancements in bird's eye view (BEV) representations have shown remarkable promise for in-vehicle 3D perception. However, while these methods have achieved impressive results on standard benchmarks, their robustness in varied conditions remains insufficiently assessed. In this study, we present RoboBEV, an extensive benchmark suite designed to evaluate the resilience of BEV algorithms. This suite incorporates a diverse set of camera corruption types, each examined over three severity levels. Our benchmarks also consider the impact of complete sensor failures that occur when using multi-modal models. Through RoboBEV, we assess 33 state-of-the-art BEV-based perception models spanning tasks like detection, map segmentation, depth estimation, and occupancy prediction. Our analyses reveal a noticeable correlation between the model's performance on in-distribution datasets and its resilience to out-of-distribution challenges. Our experimental results also underline the efficacy of strategies like pre-training and depth-free BEV transformations in enhancing robustness against out-of-distribution data. Furthermore, we observe that leveraging extensive temporal information significantly improves the model's robustness. Based on our observations, we design an effective robustness enhancement strategy based on the CLIP model. The insights from this study pave the way for the development of future BEV models that seamlessly combine accuracy with real-world robustness. Shaoyuan Xie, Lingdong Kong, Jiawei Ren 0001, Liang Pan, Kai Chen 0026, Ziwei Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Language-Aware Vision Transformer for Referring SegmentationabstractReferring segmentation is a fundamental vision-language task that aims to segment out an object from an image or video in accordance with a natural language description. One of the key challenges behind this task is leveraging the referring expression for highlighting relevant positions in the image or video frames. A paradigm for tackling this problem in both the image and the video domains is to leverage a powerful vision-language ("cross-modal") decoder to fuse features independently extracted from a vision encoder and a language encoder. Recent methods have made remarkable advances in this paradigm by exploiting Transformers as cross-modal decoders, concurrent to the Transformer's overwhelming success in many other vision-language tasks. Adopting a different approach in this work, we show that significantly better cross-modal alignments can be achieved through the early fusion of linguistic and visual features in intermediate layers of a vision Transformer encoder network. Based on the idea of conducting cross-modal feature fusion in the visual feature encoding stage, we propose a unified framework named Language-Aware Vision Transformer (LAVT), which leverages the well-proven correlation modeling power of a Transformer encoder for excavating helpful multi-modal context. This way, accurate segmentation results can be harvested with a light-weight mask predictor. One of the key components in the proposed system is a dense attention mechanism for collecting pixel-specific linguistic cues. When dealing with video inputs, we present the video LAVT framework and design a 3D version of this component by introducing multi-scale convolutional operators arranged in a parallel fashion, which can exploit spatio-temporal dependencies at different granularity levels. We further introduce unified LAVT as a unified framework that could handle both image and video inputs with enhanced segmentation capability on unified referring segmentation task. Our methods surpass previous state-of-the-art methods on seven benchmarks for referring image segmentation and referring video segmentation. The code to reproduce our experiments is available at LAVT-RS. Zhao Yang 0002, Jiaqi Wang 0003, Xubing Ye, Yansong Tang, Kai Chen 0026, Hengshuang Zhao, Philip Torr 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by StepabstractZehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu, Jiangning Liu, Miao Zheng, Jingming Zhuo, Songyang Zhang, Dahua Lin, Kai Chen, Feng Zhao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Weihua Du, Kuikun Liu, Jiangning Liu, Miao Zheng, Jingming Zhuo, Songyang Zhang 0001, Dahua Lin, Kai Chen 0026, Feng Zhao 0004 |
ACL (1) | 10 |
| 2024 | ANAH: Analytical Annotation of Hallucinations in Large Language ModelsabstractReducing the 'hallucination' problem of Large Language Models (LLMs) is crucial for their wide applications.A comprehensive and finegrained measurement of the hallucination is the first key step for the governance of this issue but is under-explored in the community.Thus, we present ANAH, a bilingual dataset that offers ANalytical Annotation of Hallucinations in LLMs within Generative Question Answering.Each answer sentence in our dataset undergoes rigorous annotation, involving the retrieval of a reference fragment, the judgment of the hallucination type, and the correction of hallucinated content.ANAH consists of ∼12k sentence-level annotations for ∼4.3kLLM responses covering over 700 topics, constructed by a human-in-the-loop pipeline.Thanks to the fine granularity of the hallucination annotations, we can quantitatively confirm that the hallucinations of LLMs progressively accumulate in the answer and use ANAH to train and evaluate hallucination annotators.We conduct extensive experiments on studying generative and discriminative annotators and show that, although current open-source LLMs have difficulties in fine-grained hallucination annotation, the generative annotator trained with ANAH can surpass all open-source LLMs and GPT-3.5, obtain performance competitive with GPT-4, and exhibits better generalization ability on unseen questions. 1 Ziwei Ji 0001, Yuzhe Gu, Chengqi Lyu, Dahua Lin, Kai Chen 0026 |
ACL (1) | 6 |
| 2024 | OMG-Seg: Is One Model Good Enough for all Segmentation?abstractIn this work, we address various segmentation tasks, each traditionally tackled by distinct or partially unified models. We propose OMG-Seg, One Model that is Good enough to efficiently and effectively handle all the segmentation tasks, including image semantic, instance, and panoptic segmentation, as well as their video counterparts, open vocabulary settings, prompt-driven, interactive segmentation like SAM, and video object segmentation. To our knowledge, this is the first model to handle all these tasks in one model and achieve satisfactory performance. We show that OMG-Seg, a transformer-based encoder-decoder architecture with task-specific queries and outputs, can support over ten distinct segmentation tasks and yet significantly reduce computational and parameter overhead across various tasks and datasets. We rigorously evaluate the inter-task influences and correlations during co-training. Code and models are available at https://github.com/lxtGH/OMG-Seg. Xiangtai Li, Haobo Yuan, Wei Li 0319, Henghui Ding, Size Wu, Kai Chen 0026, Chen Change Loy |
CVPR | 8 |
| 2024 | From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsabstractScene graph generation (SGG) aims to parse a visual scene into an intermediate graph representation for downstream reasoning tasks. Despite recent advancements, existing methods struggle to generate scene graphs with novel visual relation concepts. To address this challenge, we introduce a new open-vocabulary SGG framework based on sequence generation. Our framework leverages vision-language pre-trained models (VLM) by incorporating an image-to-graph generation paradigm. Specifically, we generate scene graph sequences via image-to-text generation with VLM and then construct scene graphs from these sequences. By doing so, we harness the strong capabilities of VLM for open-vocabulary SGG and seamlessly integrate explicit relational modeling for enhancing the VL tasks. Experimental results demonstrate that our design not only achieves superior performance with an open vocabulary but also enhances downstream vision-language task performance through explicit relation modeling knowledge. Rongjie Li, Songyang Zhang 0001, Dahua Lin, Kai Chen 0026, Xuming He 0001 |
CVPR | 4 |
| 2024 | RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose EstimationabstractReal-time multi-person pose estimation presents signif-icant challenges in balancing speed and precision. While two-stage top-down methods slow down as the number of people in the image increases, existing one-stage meth-ods often fail to simultaneously deliver high accuracy and real-time performance. This paper introduces RTMO, a one-stage pose estimation framework that seamlessly inte-grates coordinate classification by representing keypoints using dual I-D heatmaps within the YOLO architecture, achieving accuracy comparable to top-down methods while maintaining high speed. We propose a dynamic coordi-nate classifier and a tailored loss function for heatmap learning, specifically designed to address the incompati-bilities between coordinate classification and dense pre-diction models. RTMO outperforms state-of-the-art one-stage pose estimators, achieving 1.1% higher AP on COCO while operating about 9 times faster with the same back-bone. Our largest model, RTMO-1, attains 74.8% AP on COCO va12017 and 141 FPS on a single V100 GPU, demonstrating its efficiency and accuracy. The code and models are available at https://github.com/open-mmlab/mmpose/tree/main/projects/rtmo. Xiangtai Li, Kai Chen 0026, Wenming Yang |
CVPR | 5 |
| 2024 | Make-It-Vivid: Dressing Your Animatable Biped Cartoon Characters from TextabstractCreating and animating 3D biped cartoon characters is crucial and valuable in various applications. Compared with geometry, the diverse texture design plays an important role in making 3D biped cartoon characters vivid and charming. Therefore, we focus on automatic texture design for cartoon characters based on input instructions. This is challenging for domain-specific requirements and a lack of high-quality data. To address this challenge, we propose Make-It-Vivid, the first attempt to enable high-quality texture generation from text in UV space. We prepare a detailed text-texture paired data for 3D characters by using vision-question-answering agents. Then we customize a pretrained text-to-image model to generate texture map with template structure while preserving the natural 2D image knowledge. Furthermore, to enhance fine-grained details, we propose a novel adversarial learning scheme to shorten the domain gap between original dataset and realistic texture domain. Extensive experiments show that our approach outperforms current texture generation methods, resulting in efficient character texturing and faithful generation with prompts. Besides, we showcase various applications such as out of domain generation and texture stylization. We also provide an efficient generation system for automatic text-guided textured character generation and animation. Junshu Tang, Yanhong Zeng, Xuheng Wang, Bo Dai 0002, Kai Chen 0026, Lizhuang Ma |
CVPR | 6 |
| 2024 | EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AIabstractIn the realm of computer vision and robotics, embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-person observations and contextualize them into language for interaction. However, traditional research focuses more on scene-level input and output setups from a global view. To address the gap, we introduce EmbodiedScan, a multi-modal, ego-centric 3D perception dataset and benchmark for holistic 3D scene understanding. It encompasses over 5k scans encapsulating 1M ego-centric RGB-D views, 1M language prompts, 160k 3D-oriented boxes spanning over 760 categories, some of which partially align with LVIS, and dense semantic occupancy with 80 common categories. Building upon this database, we introduce a baseline framework named Embodied Perceptron. It is capable of processing an arbitrary number of multi-modal inputs and demonstrates remarkable 3D perception capabilities, both within the two series of benchmarks we set up, i.e., fundamental 3D perception tasks and language-grounded tasks, and in the wild. Xiaohan Mao, Chenming Zhu, Runsen Xu, Ruiyuan Lyu, Peisen Li, Xiao Chen 0016, Kai Chen 0026, Tianfan Xue, Xihui Liu, Cewu Lu, Dahua Lin, Jiangmiao Pang |
CVPR | 9 |
| 2024 | Towards Language-Driven Video Inpainting via Multimodal Large Language ModelsabstractWe introduce a new task - language-driven video inpainting, which uses natural language instructions to guide the inpainting process. This approach overcomes the limitations of traditional video inpainting methods that depend on manually labeled binary masks, a process often tedious and labor-intensive. We present the Remove Objects from Videos by Instructions (ROVI) dataset, containing 5,650 videos and 9,091 inpainting results, to support training and evaluation for this task. We also propose a novel diffusion-based language-driven video inpainting framework, the first end-to-end baseline for this task, integrating Multimodal Large Language Models to understand and execute complex language-based inpaintingrequests effectively. Our comprehensive results showcase the dataset's versatility and the model's effectiveness in various language-instructed inpainting scenarios. We have made datasets, code, and models publicly available at https://github.com/jianzongwu/Language-Driven-Video-Inpainting. Jianzong Wu, Xiangtai Li, Chenyang Si, Shangchen Zhou, Jiangning Zhang, Kai Chen 0026, Yunhai Tong, Ziwei Liu 0002, Chen Change Loy |
CVPR | 8 |
| 2024 | PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image ModelsabstractRecent advancements in personalized text-to-image (T2I) models have revolutionized content creation, empowering non-experts to generate stunning images with unique styles. While promising, animating these personalized images with realistic motions poses significant challenges in preserving distinct styles, high-fidelity details, and achieving motion controllability by text. In this paper, we present PIA, a Personalized Image Animator that excels in aligning with condition images, achieving motion controllability by text, and the compatibility with various personalized T2I models without specific tuning. To achieve these goals, PIA builds upon a base T2I model with well-trained temporal alignment layers, allowing for the seamless transformation of any personalized T2I model into an image animation model. A key component of PIA is the introduction of the condition module, which takes as inputs the conditionframe and inter-frame affinity. This module leverages the affinity hint to transfer appearance information from the condition frame to individual frames in the latent space. This design mitigates the challenges of appearance-related frame alignment within PIA and allows for a stronger focus on aligning with motion-related guidance. To address the lack of a benchmark for this field, we introduce AnimateBench, a comprehensive benchmark comprising diverse personalized T2I models, curated images, and motion-related prompts. We show extensive evaluations and applications on AnimateBench to verify the superiority of PIA. Yiming Zhang 0031, Zhening Xing, Yanhong Zeng, Youqing Fang, Kai Chen 0026 |
CVPR | 5 |
| 2024 | MMBench: Is Your Multi-modal Model an All-Around Player?
Yuan Liu 0025, Haodong Duan, Yuanhan Zhang, Bo Li 0080, Songyang Zhang 0001, Wangbo Zhao, Yike Yuan, Jiaqi Wang 0003, Conghui He, Ziwei Liu 0002, Kai Chen 0026, Dahua Lin |
ECCV (6) | 11 |
| 2024 | AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation
Yanan Sun 0005, Yinhao Tang, Wenjie Pei, Kai Chen 0026 |
ECCV (11) | 5 |
| 2024 | 4D Contrastive Superflows are Dense 3D Representation Learners
Xiang Xu 0009, Lingdong Kong, Hui Shuai, Liang Pan, Kai Chen 0026, Ziwei Liu 0002, Qingshan Liu 0001 |
ECCV (1) | 6 |
| 2024 | Open-Vocabulary SAM: Segment and Recognize Twenty-Thousand Classes Interactively
Haobo Yuan, Xiangtai Li, Chong Zhou, Kai Chen 0026, Chen Change Loy |
ECCV (43) | 5 |
| 2024 | ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities
Chenming Zhu, Kai Chen 0026, Xihui Liu |
ECCV (8) | 4 |
| 2024 | A Task Is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting
Junhao Zhuang, Yanhong Zeng, Wenran Liu, Chun Yuan 0003, Kai Chen 0026 |
ECCV (58) | 5 |
| 2024 | LawBench: Benchmarking Legal Knowledge of Large Language ModelsabstractZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Alan Huang, Songyang Zhang, Kai Chen, Zhixin Yin, Zongwen Shen, Jidong Ge, Vincent Ng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhiwei Fei, Xiaoyu Shen 0001, Fengzhe Zhou, Zhuo Han, Alan Huang, Songyang Zhang 0001, Kai Chen 0026, Zhixin Yin, Zongwen Shen, Jidong Ge, Vincent Ng 0001 |
EMNLP | 8 |
| 2024 | Can AI Assistants Know What They Don't Know?abstractAI assistants powered by Large Language Models (LLMs) have demonstrated impressive performance in various tasks. However, LLMs still make factual errors in knowledge-intensive tasks such as open-domain question answering. These untruthful responses from AI assistants can pose significant risks in practical applications. Therefore, in this paper, we ask the question Can AI assistants know what they don’t know and express this awareness through natural language? To investigate this, we construct a model-specific "I don’t know" (Idk) dataset. This dataset includes Supervised Fine-tuning data and preference data, categorizing questions based on whether the assistant knows or does not know the answers. Then, we align the assistant with its corresponding Idk dataset using different alignment methods, including Supervised Fine-tuning and preference optimization. Experimental results show that, after alignment with the Idk dataset, the assistant is more capable of declining to answer questions outside its knowledge scope. The assistant aligned with the Idk dataset shows significantly higher truthfulness than the original assistant. Qinyuan Cheng, Tianxiang Sun, Zhangyue Yin, Linyang Li, Zhengfu He, Kai Chen 0026, Xipeng Qiu |
ICML | 9 |
| 2024 | Differentiable Model Scaling using Differentiable TopkabstractOver the past few years, as large language models have ushered in an era of intelligence emergence, there has been an intensified focus on scaling networks. Although Neural Architecture Search (NAS) methods have been proposed to automate this process, they suffer from low search efficiency. This study introduces Differentiable Model Scaling (DMS), increasing the efficiency for searching optimal width and depth in networks. DMS can model both width and depth in a direct and fully differentiable way, making it easy to optimize. We have evaluated our DMS across diverse tasks, ranging from vision tasks to NLP tasks and various network architectures, including CNNs and Transformers. Results consistently indicate that our DMS can find improved structures and outperforms state-of-the-art NAS methods. Specifically, for image classification on ImageNet, our DMS improves the top-1 accuracy of EfficientNet-B0 and Deit-Tiny by 1.4% and 0.6%, respectively, and outperforms the state-of-the-art zero-shot NAS method, ZiCo, by 1.3% while requiring only 0.4 GPU days for searching. For object detection on COCO, DMS improves the mAP of Yolo-v8-n by 2.0%. For language modeling, our pruned Llama-7B outperforms the prior method with lower perplexity and higher zero-shot classification accuracy. Our code is available at https://github.com/LKJacky/Differentiable-Model-Scaling. Ruohui Wang, Jianfei Gao 0003, Kai Chen 0026 |
ICML | 4 |
| 2024 | VLMEvalKit: An Open-Source ToolKit for Evaluating Large Multi-Modality Models
Haodong Duan, Junming Yang 0001, Yuxuan Qiao, Xinyu Fang, Lin Chen 0026, Yuan Liu 0025, Xiaoyi Dong, Yuhang Zang, Pan Zhang 0001, Jiaqi Wang 0003, Dahua Lin, Kai Chen 0026 |
ACM Multimedia | 12 |
| 2024 | Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarksabstractChonghua Wang, Haodong Duan, Songyang Zhang, Dahua Lin, Kai Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chonghua Wang, Haodong Duan, Songyang Zhang 0001, Dahua Lin, Kai Chen 0026 |
NAACL-HLT | 5 |
| 2024 | HumanVid: Demystifying Training Data for Camera-controllable Human Image AnimationabstractHuman image animation involves generating videos from a character photo, allowing user control and unlocking the potential for video and movie production. While recent approaches yield impressive results using high-quality training data, the inaccessibility of these datasets hampers fair and transparent benchmarking. Moreover, these approaches prioritize 2D human motion and overlook the significance of camera motions in videos, leading to limited control and unstable video generation. To demystify the training data, we present HumanVid, the first large-scale high-quality dataset tailored for human image animation, which combines crafted real-world and synthetic data. For the real-world data, we compile a vast collection of real-world videos from the internet. We developed and applied careful filtering rules to ensure video quality, resulting in a curated collection of 20K high-resolution (1080P) human-centric videos. Human and camera motion annotation is accomplished using a 2D pose estimator and a SLAM-based method. To expand our synthetic dataset, we collected 10K 3D avatar assets and leveraged existing assets of body shapes, skin textures and clothings. Notably, we introduce a rule-based camera trajectory generation method, enabling the synthetic pipeline to incorporate diverse and precise camera motion annotation, which can rarely be found in real-world data. To verify the effectiveness of HumanVid, we establish a baseline model named CamAnimate, short for Camera-controllable Human Animation, that considers both human and camera motions as conditions. Through extensive experimentation, we demonstrate that such simple baseline training on our HumanVid achieves state-of-the-art performance in controlling both human pose and camera motions, setting a new benchmark. Demo, data and code could be found in the project website: https://humanvid.github.io/. Zhenzhi Wang 0001, Yixuan Li 0002, Yanhong Zeng, Youqing Fang, Yuwei Guo 0002, Wenran Liu, Jing Tan 0002, Kai Chen 0026, Tianfan Xue, Bo Dai 0002, Dahua Lin |
NeurIPS | 8 |
| 2024 | CriticEval: Evaluating Large-scale Language Model as CriticabstractCritique ability, i.e., the capability of Large Language Models (LLMs) to identify and rectify flaws in responses, is crucial for their applications in self-improvement and scalable oversight. While numerous studies have been proposed to evaluate critique ability of LLMs, their comprehensiveness and reliability are still limited. To overcome this problem, we introduce CriticEval, a novel benchmark designed to comprehensively and reliably evaluate critique ability of LLMs. Specifically, to ensure the comprehensiveness, CriticEval evaluates critique ability from four dimensions across nine diverse task scenarios. It evaluates both scalar-valued and textual critiques, targeting responses of varying quality. To ensure the reliability, a large number of critiques are annotated to serve as references, enabling GPT-4 to evaluate textual critiques reliably. Extensive evaluations of open-source and closed-source LLMs first validate the reliability of evaluation in CriticEval. Then, experimental results demonstrate the promising potential of open-source LLMs, the effectiveness of critique datasets and several intriguing relationships between the critique ability and some critical factors, including task types, response qualities and critique dimensions. Tian Lan 0003, Heyan Huang, Dahua Lin, Kai Chen 0026, Xianling Mao |
NeurIPS | 6 |
| 2024 | InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HDabstractThe Large Vision-Language Model (LVLM) field has seen significant advancements, yet its progression has been hindered by challenges in comprehending fine-grained visual content due to limited resolution. Recent efforts have aimed to enhance the high-resolution understanding capabilities of LVLMs, yet they remain capped at approximately 1500 $\times$ 1500 pixels and constrained to a relatively narrow resolution range. This paper represents InternLM-XComposer2-4KHD, a groundbreaking exploration into elevating LVLM resolution capabilities up to 4K HD (3840 × 1600) and beyond. Concurrently, considering the ultra-high resolution may not be necessary in all scenarios, it supports a wide range of diverse resolutions from 336 pixels to 4K standard, significantly broadening its scope of applicability. Specifically, this research advances the patch division paradigm by introducing a novel extension: dynamic resolution with automatic patch configuration. It maintains the training image aspect ratios while automatically varying patch counts and configuring layouts based on a pre-trained Vision Transformer (ViT) (336 $\times$ 336), leading to dynamic training resolution from 336 pixels to 4K standard. Our research demonstrates that scaling training resolution up to 4K HD leads to consistent performance enhancements without hitting the ceiling of potential improvements. InternLM-XComposer2-4KHD shows superb capability that matches or even surpasses GPT-4V and Gemini Pro in 10 of the 16 benchmarks. Xiaoyi Dong, Pan Zhang 0001, Yuhang Zang, Yuhang Cao, Bin Wang 0065, Linke Ouyang, Songyang Zhang 0001, Haodong Duan, Hang Yan 0001, Yang Gao 0042, Zhe Chen 0017, Xinyue Zhang 0005, Wei Li 0320, Wenhai Wang, Kai Chen 0026, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao 0001, Dahua Lin, Jiaqi Wang 0003 |
NeurIPS | 18 |
| 2024 | MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video UnderstandingabstractThe advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative metrics, often fail to encompass the full spectrum of video content and inadequately assess models' temporal comprehension. To address these limitations, we introduce MMBench-Video, a quantitative benchmark designed to rigorously evaluate LVLMs' proficiency in video understanding. MMBench-Video incorporates lengthy videos from YouTube and employs free-form questions, mirroring practical use cases. The benchmark is meticulously crafted to probe the models' temporal reasoning skills, with all questions human-annotated according to a carefully constructed ability taxonomy.We employ GPT-4 for automated assessment, demonstrating superior accuracy and robustness over earlier LLM-based evaluations. Utilizing MMBench-Video, we have conducted comprehensive evaluations that include both proprietary and open-source LVLMs for images and videos. MMBench-Video stands as a valuable resource for the research community, facilitating improved evaluation of LVLMs and catalyzing progress in the field of video understanding. Xinyu Fang, Kangrui Mao, Haodong Duan, Dahua Lin, Kai Chen 0026 |
NeurIPS | 7 |
| 2024 | ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language ModelsabstractLarge language models (LLMs) exhibit hallucinations in long-form question-answering tasks across various domains and wide applications. Current hallucination detection and mitigation datasets are limited in domain and size, which struggle to scale due to prohibitive labor costs and insufficient reliability of existing hallucination annotators. To facilitate the scalable oversight of LLM hallucinations, this paper introduces an iterative self-training framework that simultaneously and progressively scales up the annotation dataset and improves the accuracy of the annotator. Based on the Expectation Maximization algorithm, in each iteration, the framework first applies an automatic hallucination annotation pipeline for a scaled dataset and then trains a more accurate annotator on the dataset. This new annotator is adopted in the annotation pipeline for the next iteration. Extensive experimental results demonstrate that the finally obtained hallucination annotator with only 7B parameters surpasses GPT-4 and obtains new state-of-the-art hallucination detection results on HaluEval and HalluQA by zero-shot inference. Such an annotator can not only evaluate the hallucination levels of various LLMs on the large-scale dataset but also help to mitigate the hallucination of LLMs generations, with the Natural Language Inference metric increasing from 25% to 37% on HaluEval. Yuzhe Gu, Ziwei Ji 0001, Chengqi Lyu, Dahua Lin, Kai Chen 0026 |
NeurIPS | 6 |
| 2024 | Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained OptimizationabstractRecent research indicates that large language models (LLMs) are susceptible to jailbreaking attacks that can generate harmful content. This paper introduces a novel token-level attack method, Adaptive Dense-to-Sparse Constrained Optimization (ADC), which has been shown to successfully jailbreak multiple open-source LLMs. Drawing inspiration from the difficulties of discrete token optimization, our method relaxes the discrete jailbreak optimization into a continuous optimization process while gradually increasing the sparsity of the optimizing vectors. This technique effectively bridges the gap between discrete and continuous space optimization. Experimental results demonstrate that our method is more effective and efficient than state-of-the-art token-level methods. On Harmbench, our approach achieves the highest attack success rate on seven out of eight LLMs compared to the latest jailbreak methods. \textcolor{red}{Trigger Warning: This paper contains model behavior that can be offensive in nature.} Weichen Yu, Tianjun Yao, Wenhe Liu, Lijun Yu, Kai Chen 0026, Matt Fredrikson |
NeurIPS | 9 |
| 2024 | Prism: A Framework for Decoupling and Assessing the Capabilities of VLMsabstractVision Language Models (VLMs) demonstrate remarkable proficiency in addressing a wide array of visual questions, which requires strong perception and reasoning faculties. Assessing these two competencies independently is crucial for model refinement, despite the inherent difficulty due to the intertwined nature of seeing and reasoning in existing VLMs. To tackle this issue, we present Prism, an innovative framework designed to disentangle the perception and reasoning processes involved in visual question solving. Prism comprises two distinct stages: a perception stage that utilizes a VLM to extract and articulate visual information in textual form, and a reasoning stage that formulates responses based on the extracted visual information using a Large Language Model (LLM). This modular design enables the systematic comparison and assessment of both proprietary and open-source VLM for their perception and reasoning strengths. Our analytical framework provides several valuable insights, underscoring Prism's potential as a cost-effective solution for vision-language tasks.
By combining a streamlined VLM focused on perception with a powerful LLM tailored for reasoning, Prism achieves superior results in general vision-language tasks while substantially cutting down on training and operational expenses. Quantitative evaluations show that Prism, when configured with a vanilla 2B LLaVA and freely accessible GPT-3.5, delivers performance on par with VLMs $10 \times$ larger on the rigorous multimodal benchmark MMStar. Yuxuan Qiao, Haodong Duan, Xinyu Fang, Junming Yang 0001, Lin Chen 0026, Songyang Zhang 0001, Jiaqi Wang 0003, Dahua Lin, Kai Chen 0026 |
NeurIPS | 9 |
| 2024 | AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source DataabstractOpen-source Large Language Models (LLMs) and their specialized variants, particularly Code LLMs, have recently delivered impressive performance. However, previous Code LLMs are typically fine-tuned on single-source data with limited quality and diversity, which may insufficiently elicit the potential of pre-trained Code LLMs. In this paper, we present AlchemistCoder, a series of Code LLMs with enhanced code generation and generalization capabilities fine-tuned on multi-source data. To achieve this, we pioneer to unveil inherent conflicts among the various styles and qualities in multi-source code corpora and introduce data-specific prompts with hindsight relabeling, termed AlchemistPrompts, to harmonize different data sources and instruction-response pairs. Additionally, we propose incorporating the data construction process into the fine-tuning data as code comprehension tasks, including instruction evolution, data filtering, and code review. Extensive experiments demonstrate that AlchemistCoder holds a clear lead among all models of the same size (6.7B/7B) and rivals or even surpasses larger models (15B/33B/70B), showcasing the efficacy of our method in refining instruction-following capabilities and advancing the boundaries of code intelligence. Source code and models are available at https://github.com/InternLM/AlchemistCoder. Zifan Song, Yudong Wang 0002, Kuikun Liu, Chengqi Lyu, Demin Song, Qipeng Guo, Hang Yan 0001, Dahua Lin, Kai Chen 0026, Cairong Zhao |
NeurIPS | 10 |
| 2024 | GTA: A Benchmark for General Tool AgentsabstractIn developing general-purpose agents, significant focus has been placed on integrating large language models (LLMs) with various tools. This poses a challenge to the tool-use capabilities of LLMs. However, there are evident gaps between existing tool evaluations and real-world scenarios. Current evaluations often use AI-generated queries, single-step tasks, dummy tools, and text-only inputs, which fail to reveal the agents' real-world problem-solving abilities effectively. To address this, we propose GTA, a benchmark for General Tool Agents, featuring three main aspects: (i) Real user queries: human-written queries with simple real-world objectives but implicit tool-use, requiring the LLM to reason the suitable tools and plan the solution steps. (ii) Real deployed tools: an evaluation platform equipped with tools across perception, operation, logic, and creativity categories to evaluate the agents' actual task execution performance. (iii) Real multimodal inputs: authentic image files, such as spatial scenes, web page screenshots, tables, code snippets, and printed/handwritten materials, used as the query contexts to align with real-world scenarios closely. We designed 229 real-world tasks and executable tool chains to evaluate mainstream LLMs. Our findings show that real-world user queries are challenging for existing LLMs, with GPT-4 completing less than 50\% of the tasks and most LLMs achieving below 25\%. This evaluation reveals the bottlenecks in the tool-use capabilities of current LLMs in real-world scenarios, which is beneficial for the advancement of general-purpose tool agents. Dataset and code are available at https://github.com/open-compass/GTA. Jize Wang, Zerun Ma, Songyang Zhang 0001, Cailian Chen, Kai Chen 0026, Xinyi Le |
NeurIPS | 6 |
| 2024 | MotionBooth: Motion-Aware Customized Text-to-Video GenerationabstractIn this work, we present MotionBooth, an innovative framework designed for animating customized subjects with precise control over both object and camera movements. By leveraging a few images of a specific object, we efficiently fine-tune a text-to-video model to capture the object's shape and attributes accurately. Our approach presents subject region loss and video preservation loss to enhance the subject's learning performance, along with a subject token cross-attention loss to integrate the customized subject with motion control signals. Additionally, we propose training-free techniques for managing subject and camera motions during inference. In particular, we utilize cross-attention map manipulation to govern subject motion and introduce a novel latent shift module for camera movement control as well. MotionBooth excels in preserving the appearance of subjects while simultaneously controlling the motions in generated videos. Extensive quantitative and qualitative evaluations demonstrate the superiority and effectiveness of our method. Models and codes will be made publicly available. Jianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang, Qianyu Zhou 0001, Yunhai Tong, Kai Chen 0026 |
NeurIPS | 8 |
| 2024 | Lean Workbook: A large-scale Lean problem set formalized from natural language math problemsabstractLarge language models have demonstrated impressive capabilities across various natural language processing tasks, especially in solving mathematical problems. However, large language models are not good at math theorem proving using formal languages like Lean. A significant challenge in this area is the scarcity of training data available in these formal languages. To address this issue, we propose a novel pipeline that iteratively generates and filters synthetic data to translate natural language mathematical problems into Lean 4 statements, and vice versa. Our results indicate that the synthetic data pipeline can provide useful training data and improve the performance of LLMs in translating and understanding complex mathematical problems and proofs. Our final dataset contains about 57K formal-informal question pairs along with searched proof from the math contest forum and 21 new IMO questions. We open-source our code at \url{https://github.com/InternLM/InternLM-Math} and our data at \url{https://huggingface.co/datasets/InternLM/Lean-Workbook}. Huaiyuan Ying, Zijian Wu 0002, Yihan Geng, Dahua Lin, Kai Chen 0026 |
NeurIPS | 6 |
| 2024 | Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech GenerationabstractRecent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/. Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu 0008, Jiaqi Li 0030, Peiyang Shi, Yuancheng Wang, Kai Chen 0026, Pengyuan Zhang, Zhizheng Wu 0001 |
SLT | 12 |
| 2024 | Amphion: an Open-Source Audio, Music, and Speech Generation ToolkitabstractAmphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion. Xueyao Zhang, Liumeng Xue, Yicheng Gu, Yuancheng Wang, Jiaqi Li 0030, Haorui He, Chaoren Wang, Songting Liu, Junan Zhang, Zihao Fang, Haopeng Chen, Tze Ying Tang, Lexiao Zou, Mingxuan Wang, Kai Chen 0026, Haizhou Li 0001, Zhizheng Wu 0001 |
SLT | 17 |
| 2024 | A Novel Contrastive Learning Model for Aerial ImagesabstractIn the field of remote sensing, tasks related to images often require a large amount of manually annotated data for model training, thus requiring a large amount of human resources with professional knowledge. To alleviate this problem, some self-supervised methods have been proposed for learning image representations for downstream tasks. As a specific form, contrastive learning can be effectively used for unlabeled image data. Using this method, we achieve the ability to learn informative feature representations from extensive unlabeled data, thereby establishing a robust initialization model for subsequent tasks. But during the model learning process, the abundance and quality of negative samples significantly impact the model’s overall performance. Consequently, we improve the existing contrastive learning architecture of the conventional two-pathway network by incorporating an additional branch equipped with a hybrid encoder. This modification facilitates the fusion of both global and local features. Furthermore, we introduce other optimization techniques, including image reconstruction and indicator vector transformation. Notably, our approach emphasizes maintaining diversity among negative samples stored in the historical feature storage queue. Our model provides competitive results on aerial image dataset (AID) classification. Besides, our model outperforms other self-supervised baselines under different proportions of labeled data fine-tuning. Taihang Zhen, Kai Chen 0026, Yang Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Transformer-Based Visual Segmentation: A SurveyabstractVisual segmentation seeks to partition images, video frames, or point clouds into multiple segments or groups. This technique has numerous real-world applications, such as autonomous driving, image editing, robot sensing, and medical analysis. Over the past decade, deep learning-based methods have made remarkable strides in this area. Recently, transformers, a type of neural network based on self-attention originally designed for natural language processing, have considerably surpassed previous convolutional or recurrent approaches in various vision processing tasks. Specifically, vision transformers offer robust, unified, and even simpler solutions for various segmentation tasks. This survey provides a thorough overview of transformer-based visual segmentation, summarizing recent advancements. We first review the background, encompassing problem definitions, datasets, and prior convolutional methods. Next, we summarize a meta-architecture that unifies all recent transformer-based approaches. Based on this meta-architecture, we examine various method designs, including modifications to the meta-architecture and associated applications. We also present several specific subfields, including 3D point cloud segmentation, foundation model tuning, domain-aware segmentation, efficient segmentation, and medical segmentation. Additionally, we compile and re-evaluate the reviewed methods on several well-established datasets. Finally, we identify open challenges in this field and propose directions for future research. Xiangtai Li, Henghui Ding, Haobo Yuan, Jiangmiao Pang, Kai Chen 0026, Ziwei Liu 0002, Chen Change Loy |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Semantics-Aware Dynamic Localization and Refinement for Referring Image SegmentationabstractReferring image segmentation segments an image from a language expression. With the aim of producing high-quality masks, existing methods often adopt iterative learning approaches that rely on RNNs or stacked attention layers to refine vision-language features. Despite their complexity, RNN-based methods are subject to specific encoder choices, while attention-based methods offer limited gains. In this work, we introduce a simple yet effective alternative for progressively learning discriminative multi-modal features. The core idea of our approach is to leverage a continuously updated query as the representation of the target object and at each iteration, strengthen multi-modal features strongly correlated to the query while weakening less related ones. As the query is initialized by language features and successively updated by object features, our algorithm gradually shifts from being localization-centric to segmentation-centric. This strategy enables the incremental recovery of missing object parts and/or removal of extraneous parts through iteration. Compared to its counterparts, our method is more versatile—it can be plugged into prior arts straightforwardly and consistently bring improvements. Experimental results on the challenging datasets of RefCOCO, RefCOCO+, and G-Ref demonstrate its advantage with respect to the state-of-the-art methods. Zhao Yang 0002, Jiaqi Wang 0003, Yansong Tang, Kai Chen 0026, Hengshuang Zhao, Philip Torr 0001 |
AAAI | 4 |
| 2023 | RIFormer: Keep Your Vision Backbone Effective But Removing Token MixerabstractThis paper studies how to keep a vision backbone effective while removing token mixers in its basic building blocks. Token mixers, as self-attention for vision transformers (ViTs), are intended to perform information communication between different spatial tokens but suffer from considerable computational cost and latency. However, directly removing them will lead to an incomplete model structure prior, and thus brings a significant accuracy drop. To this end, we first develop an RepIdentityFormer base on the re-parameterizing idea, to study the token mixer free model architecture. And we then explore the improved learning paradigm to break the limitation of simple token mixer free backbone, and summarize the empirical practice into 5 guidelines. Equipped with the proposed optimization strategy, we are able to build an extremely simple vision backbone with encouraging performance, while enjoying the high efficiency during inference. Extensive experiments and ablative analysis also demonstrate that the inductive bias of network architecture, can be incorporated into simple network structure with appropriate optimization strategy. We hope this work can serve as a starting point for the exploration of optimization-driven efficient network design. Jiahao Wang 0005, Songyang Zhang 0001, Yong Liu 0033, Taiqiang Wu, Yujiu Yang 0001, Xihui Liu, Kai Chen 0026, Ping Luo 0002, Dahua Lin |
CVPR | 7 |
| 2023 | Dense Distinct Query for End-to-End Object DetectionabstractOne-to-one label assignment in object detection has successfully obviated the need for non-maximum suppression (NMS) as postprocessing and makes the pipeline end-to-end. However, it triggers a new dilemma as the widely used sparse queries cannot guarantee a high recall, while dense queries inevitably bring more similar queries and encounter optimization difficulties. As both sparse and dense queries are problematic, then what are the expected queries in end-to-end object detection? This paper shows that the solution should be Dense Distinct Queries (DDQ). Concretely, we first lay dense queries like traditional detectors and then select distinct ones for one-to-one assignments. DDQ blends the advantages of traditional and recent end-to-end detectors and significantly improves the performance of various detectors including FCN, R-CNN, and DETRs. Most impressively, DDQ-DETR achieves 52.1 AP on MS-COCO dataset within 12 epochs using a ResNet-50 backbone, outperforming all existing detectors in the same setting. DDQ also shares the benefit of end-to-end detectors in crowded scenes and achieves 93.8 AP on Crowd-Human. We hope DDQ can inspire researchers to consider the complementarity between traditional methods and end-to-end detectors. The source code can be found at https://github.com/jshilong/DDQ. Xinjiang Wang, Jiaqi Wang 0003, Jiangmiao Pang, Chengqi Lyu, Ping Luo 0002, Kai Chen 0026 |
CVPR | 8 |
| 2023 | Robo3D: Towards Robust and Reliable 3D Perception against CorruptionsabstractThe robustness of 3D perception systems under natural corruptions from environments and sensors is pivotal for safety-critical applications. Existing large-scale 3D perception datasets often contain data that are meticulously cleaned. Such configurations, however, cannot reflect the reliability of perception models during the deployment stage. In this work, we present Robo3D, the first comprehensive benchmark heading toward probing the robustness of 3D detectors and segmentors under out-of-distribution scenarios against natural corruptions that occur in real-world environments. Specifically, we consider eight corruption types stemming from severe weather conditions, external disturbances, and internal sensor failure. We uncover that, although promising results have been progressively achieved on standard benchmarks, state-of-the-art 3D perception models are at risk of being vulnerable to corruptions. We draw key observations on the use of data representations, augmentation schemes, and training strategies, that could severely affect the model's performance. To pursue better robustness, we propose a density-insensitive training framework along with a simple flexible voxelization strategy to enhance the model resiliency. We hope our benchmark and approach could inspire future research in designing more robust and reliable 3D perception models. Our robustness benchmark suite is publicly available1. Lingdong Kong, Youquan Liu, Xin Li 0110, Runnan Chen, Jiawei Ren 0001, Liang Pan, Kai Chen 0026, Ziwei Liu 0002 |
ICCV | 8 |
| 2023 | Improving Pixel-based MIM by Reducing Wasted Modeling CapabilityabstractThere has been significant progress in Masked Image Modeling (MIM). Existing MIM methods can be broadly categorized into two groups based on the reconstruction target: pixel-based and tokenizer-based approaches. The former offers a simpler pipeline and lower computational cost, but it is known to be biased toward high-frequency details. In this paper, we provide a set of empirical studies to confirm this limitation of pixel-based MIM and propose a new method that explicitly utilizes low-level features from shallow layers to aid pixel reconstruction. By incorporating this design into our base method, MAE, we reduce the wasted modeling capability of pixel-based MIM, improving its convergence and achieving non-trivial improvements across various downstream tasks. To the best of our knowledge, we are the first to systematically investigate multilevel feature fusion for isotropic architectures like the standard Vision Transformer (ViT). Notably, when applied to a smaller model (e.g., ViT-S), our method yields significant performance gains, such as 1.2% on fine-tuning, 2.8% on linear probing, and 2.6% on semantic segmentation. Code and models are available in MMPretrain1. Yuan Liu 0025, Songyang Zhang 0001, Zhaohui Yu, Kai Chen 0026, Dahua Lin |
ICCV | 5 |
| 2023 | TG-VQA: Ternary Game of Video Question AnsweringabstractVideo question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions, i.e., annotations or priors, current contrastive learning-based VideoQA methods remains challenging to perform fine-grained visual-linguistic alignments. In this work, we innovatively resort to game theory, which can simulate complicated relationships among multiple players with specific interaction strategies, e.g., video, question, and answer as ternary players, to achieve fine-grained alignment for VideoQA task. Specifically, we carefully design a VideoQA-specific interaction strategy to tailor the characteristics of VideoQA, which can mathematically generate the fine-grained visual-linguistic alignment label without label-intensive efforts. Our TG-VQA outperforms existing state-of-the-art by a large margin (more than 5%) on long-term and short-term VideoQA datasets, verifying its effectiveness and generalization ability. Thanks to the guidance of game-theoretic interaction, our model impressively convergences well on limited data (10^4 videos), surpassing most of those pre-trained on large-scale data (10^7 videos). Hao Li 0073, Peng Jin 0001, Zesen Cheng, Songyang Zhang 0001, Kai Chen 0026, Zhennan Wang 0001, Chang Liu 0030, Jie Chen 0001 |
IJCAI | 5 |
| 2023 | Segment Any Point Cloud Sequences by Distilling Vision Foundation ModelsabstractRecent advancements in vision foundation models (VFMs) have opened up new possibilities for versatile and efficient visual perception. In this work, we introduce Seal, a novel framework that harnesses VFMs for segmenting diverse automotive point cloud sequences. Seal exhibits three appealing properties: i) Scalability: VFMs are directly distilled into point clouds, obviating the need for annotations in either 2D or 3D during pretraining. ii) Consistency: Spatial and temporal relationships are enforced at both the camera-to-LiDAR and point-to-segment regularization stages, facilitating cross-modal representation learning. iii) Generalizability: Seal enables knowledge transfer in an off-the-shelf manner to downstream tasks involving diverse point clouds, including those from real/synthetic, low/high-resolution, large/small-scale, and clean/corrupted datasets. Extensive experiments conducted on eleven different point cloud datasets showcase the effectiveness and superiority of Seal. Notably, Seal achieves a remarkable 45.0% mIoU on nuScenes after linear probing, surpassing random initialization by 36.9% mIoU and outperforming prior arts by 6.1% mIoU. Moreover, Seal demonstrates significant performance gains over existing methods across 20 different few-shot fine-tuning tasks on all eleven tested point cloud datasets. The code is available at this link. Youquan Liu, Lingdong Kong, Jun Cen, Runnan Chen, Liang Pan, Kai Chen 0026, Ziwei Liu 0002 |
NeurIPS | 7 |
| 2022 | LAVT: Language-Aware Vision Transformer for Referring Image SegmentationabstractReferring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring expression for highlighting relevant positions in the image. A paradigm for tackling this problem is to leverage a powerful vision-language (“cross-madal”) decoder to fuse features independently extracted from a vision encoder and a language encoder. Recent methods have made remarkable advancements in this paradigm by exploiting Transformers as cross-modal decoders, concurrent to the Transformer's overwhelming success in many other vision-language tasks. Adopting a different approach in this work, we show that significantly better cross-modal alignments can be achieved through the early fusion of linguistic and visual features in intermediate layers of a vision Transformer encoder network. By conducting cross-modal feature fusion in the visual feature encoding stage, we can leverage the well-proven correlation modeling power of a Transformer encoder for excavating helpful multi-modal context. This way, accurate segmentation results are readily harvested with a light-weight mask predictor. Without bells and whistles, our method surpasses the previous state-of-the-art methods on Ref CoCo, RefCOCO+, and G-Ref by large margins. Zhao Yang 0002, Jiaqi Wang 0003, Yansong Tang, Kai Chen 0026, Hengshuang Zhao, Philip Torr 0001 |
CVPR | 4 |
| 2022 | Revisiting Skeleton-based Action RecognitionabstractHuman skeleton, as a compact representation of human action, has received increasing attention in recent years. Many skeleton-based action recognition methods adopt GCNs to extract features on top of human skeletons. Despite the positive results shown in these attempts, GCN-based methods are subject to limitations in robustness, interoperability, and scalability. In this work, we propose PoseConv3D, a new approach to skeleton-based action recognition. PoseConv3D relies on a 3D heatmap volume instead of a graph sequence as the base representation of human skeletons. Compared to GCN-based methods, PoseConv3D is more effective in learning spatiotemporal features, more robust against pose estimation noises, and generalizes better in cross-dataset settings. Also, PoseConv3D can handle multiple-person scenarios without additional computation costs. The hierarchical features can be easily integrated with other modalities at early fusion stages, providing a great design space to boost the performance. PoseConv3D achieves the state-of-the-art on five of six standard skeleton-based action recognition benchmarks. Once fused with other modalities, it achieves the state-of-the-art on all eight multi-modality action recognition benchmarks. Code has been made available at: https://github.com/kennymckormick/pyskl. Haodong Duan, Yue Zhao 0006, Kai Chen 0026, Dahua Lin, Bo Dai 0002 |
CVPR | 3 |
| 2022 | TransRank: Self-supervised Video Representation Learning via Ranking-based Transformation RecognitionabstractRecognizing transformation types applied to a video clip (RecogTrans) is a long-established paradigm for selfsupervised video representation learning, which achieves much inferior performance compared to instance discrimination approaches (InstDisc) in recent works. However, based on a thorough comparison of representative Recog-Trans and InstDisc methods, we observe the great potential of RecogTrans on both semantic-related and temporalrelated downstream tasks. Based on hard-label classification, existing RecogTrans approaches suffer from noisy supervision signals in pre-training. To mitigate this problem, we developed TransRank, a unified framework for recognizing Transformations in a Ranking formulation. TransRank provides accurate supervision signals by recognizing transformations relatively, consistently outperforming the classification-based formulation. Meanwhile, the unified framework can be instantiated with an arbitrary set of temporal or spatial transformations, demonstrating good generality. With a ranking-based formulation and several empirical practices, we achieve competitive performance on video retrieval and action recognition. Under the same setting, TransRank surpasses the previous state-of-the-art method [28] by 6.4% on UCF101 and 8.3% on HMDB51 for action recognition (Topl Acc); improves video retrieval on UCF101 by 20.4% (R@1). The promising results validate that RecogTrans is still a worth exploring paradigm for video self-supervised learning. Codes will be released at https://github.com/kennymckormick/TransRank. Haodong Duan, Nanxuan Zhao, Kai Chen 0026, Dahua Lin |
CVPR | 3 |
| 2022 | Video K-Net: A Simple, Strong, and Unified Baseline for Video SegmentationabstractThis paper presents Video K-Net, a simple, strong, and unified framework for fully end-to-end video panoptic seg-mentation. The method is built upon K-Net, a method that unifies image segmentation via a group of learnable ker-nels. We observe that these learnable kernels from K-Net, which encode object appearances and contexts, can naturally associate identical instances across video frames. Motivated by this observation, Video K-Net learns to simultaneously segment and track “things” and “stuff” in a video with simple kernel-based appearance modeling and cross-temporal kernel interaction. Despite the simplicity, it achieves state-of-the-art video panoptic segmentation results on Citscapes-VPS and KITTI-STEP without bells and whistles. In particular on KITTI-STEP, the simple method can boost almost 12% relative improvements over previous methods. We also validate its generalization on video semantic segmentation, where we boost various baselines by 2% on the VSPW dataset. Moreover, we extend K-Net into clip-level video framework for video instance segmentation where we obtain 40.5% for ResNet50 backbone and 51.5% mAP for Swin-base on YouTube-2019 validation set. We hope this simple yet effective method can serve as a new flexible baseline in video segmentation.11Both code and models are released at here. Xiangtai Li, Jiangmiao Pang, Kai Chen 0026, Yunhai Tong, Chen Change Loy |
CVPR | 4 |
| 2022 | OCSampler: Compressing Videos to One Clip with Single-step SamplingabstractVideos incorporate rich semantics as well as redundant information. Seeking a compact yet effective video representation, e.g., sample informative frames from the entire video, is critical to efficient video recognition. There have been works that formulate frame sampling as a sequential decision task by selecting frames one by one according to their importance. In this paper, we present a more efficient framework named OCSampler, which explores such a representation with one short clip. OCSampler designs a new paradigm of learning instance-specific video condensation policies to select frames only in a single step. Rather than picking up frames sequentially like previous methods, we simply process a whole sequence at once. Accordingly, these policies are derived from a light-weighted skim network together with a simple yet effective policy network. Moreover, we extend the proposed method with a frame number budget, enabling the framework to produce correct predictions in high confidence with as few frames as possible. Experiments on various benchmarks demonstrate the effectiveness of OCSampler over previous methods in terms of accuracy and efficiency. Specifically, it achieves 76.9% mAP and 21.7 GFLOPs on ActivityNet with an impressive throughput: 123.9 Video/s on a single TITAN Xp GPU. Jintao Lin, Haodong Duan, Kai Chen 0026, Dahua Lin, Limin Wang 0002 |
CVPR | 3 |
| 2022 | Group R-CNN for Weakly Semi-supervised Object Detection with PointsabstractWe study the problem of weakly semi-supervised object detection with points (WSSOD-P), where the training data is combined by a small set of fully annotated images with bounding boxes and a large set of weakly-labeled images with only a single point annotated for each instance. The core of this task is to train a point-to-box regressor on well-labeled images that can be used to predict credible bounding boxes for each point annotation. We challenge the prior belief that existing CNN-based detectors are not compatible with this task. Based on the classic R-CNN architecture, we propose an effective point-to-box regressor: Group R-CNN. Group R-CNN first uses instance-level proposal grouping to generate a group of proposals for each point annotation and thus can obtain a high recall rate. To better distinguish different instances and improve precision, we propose instance-level proposal assignment to replace the vanilla assignment strategy adopted in original R-CNN methods. As naive instance-level assignment brings converging difficulty, we propose instance aware representation learning which consists of instance aware feature enhancement and instance-aware parameter generation to overcome this issue. Comprehensive experiments on the MS-COCO benchmark demonstrate the effectiveness of our method. Specifically, Group R-CNN significantly outperforms the prior method Point DETR by 3.9 mAP with 5% well-labeled images, which is the most challenging scenario. The source code can be found at https://github.com/jshilong/GroupRCNN. Zhuoran Yu, Liyang Liu, Xinjiang Wang, Aojun Zhou, Kai Chen 0026 |
CVPR | 6 |
| 2022 | Dense Siamese Network for Dense Unsupervised Learning
Jiangmiao Pang, Kai Chen 0026, Chen Change Loy |
ECCV (30) | 3 |
| 2022 | PYSKL: Towards Good Practices for Skeleton Action RecognitionabstractWe present PYSKL: an open-source toolbox for skeleton-based action recognition based on PyTorch. The toolbox supports a wide variety of skeleton action recognition algorithms, including approaches based on GCN and CNN. In contrast to existing open-source skeleton action recognition projects that include only one or two algorithms, PYSKL implements six different algorithms under a unified framework with both the latest and original good practices to ease the comparison of efficacy and efficiency. We also provide an original GCN-based skeleton action recognition model named ST-GCN++, which achieves competitive recognition performance without any complicated attention schemes, serving as a strong baseline. Meanwhile, PYSKL supports the training and testing of nine skeleton-based action recognition benchmarks and achieves state-of-the-art recognition performance on eight of them. To facilitate future research on skeleton action recognition, we also provide a large number of trained models and detailed benchmark results to give some insights. PYSKL is released at https://github.com/kennymckormick/pyskl and is actively maintained. Haodong Duan, Jiaqi Wang 0003, Kai Chen 0026, Dahua Lin |
ACM Multimedia | 3 |
| 2022 | MMRotate: A Rotated Object Detection Benchmark using PyTorchabstractWe present an open-source toolbox, named MMRotate, which provides a coherent algorithm framework of training, inferring, and evaluation for the popular rotated object detection algorithm based on deep learning. MMRotate implements 18 state-of-the-art algorithms and supports the three most frequently used angle definition methods. To facilitate future research and industrial applications of rotated object detection-related problems, we also provide a large number of trained models and detailed benchmarks to give insights into the performance of rotated object detection. MMRotate is publicly released at https://github.com/open-mmlab/mmrotate. Yue Zhou 0005, Xue Yang 0005, Gefan Zhang, Yanyi Liu, Liping Hou, Xue Jiang 0001, Xingzhao Liu, Junchi Yan, Chengqi Lyu, Kai Chen 0026 |
ACM Multimedia | 12 |
| 2022 | Deliberated Domain Bridging for Domain Adaptive Semantic SegmentationabstractIn unsupervised domain adaptation (UDA), directly adapting from the source to the target domain usually suffers significant discrepancies and leads to insufficient alignment. Thus, many UDA works attempt to vanish the domain gap gradually and softly via various intermediate spaces, dubbed domain bridging (DB). However, for dense prediction tasks such as domain adaptive semantic segmentation (DASS), existing solutions have mostly relied on rough style transfer and how to elegantly bridge domains is still under-explored. In this work, we resort to data mixing to establish a deliberated domain bridging (DDB) for DASS, through which the joint distributions of source and target domains are aligned and interacted with each in the intermediate space. At the heart of DDB lies a dual-path domain bridging step for generating two intermediate domains using the coarse-wise and the fine-wise data mixing techniques, alongside a cross-path knowledge distillation step for taking two complementary models trained on generated intermediate samples as ‘teachers’ to develop a superior ‘student’ in a multi-teacher distillation manner. These two optimization steps work in an alternating way and reinforce each other to give rise to DDB with strong adaptation power. Extensive experiments on adaptive segmentation tasks with different settings demonstrate that our DDB significantly outperforms state-of-the-art methods. Lin Chen 0026, Zhixiang Wei, Xin Jin 0014, Huaian Chen, Miao Zheng, Kai Chen 0026, Yi Jin 0002 |
NeurIPS | 6 |
| 2022 | CARAFE++: Unified Content-Aware ReAssembly of FEaturesabstractFeature reassembly, i.e. feature downsampling and upsampling, is a key operation in a number of modern convolutional network architectures, e.g., residual networks and feature pyramids. Its design is critical for dense prediction tasks such as object detection and semantic/instance segmentation. In this work, we propose unified Content-Aware ReAssembly of FEatures (CARAFE++), a universal, lightweight, and highly effective operator to fulfill this goal. CARAFE++ has several appealing properties: (1) Unlike conventional methods such as pooling and interpolation that only exploit sub-pixel neighborhood, CARAFE++ aggregates contextual information within a large receptive field. (2) Instead of using a fixed kernel for all samples (e.g. convolution and deconvolution), CARAFE++ generates adaptive kernels on-the-fly to enable instance-specific content-aware handling. (3) CARAFE++ introduces little computational overhead and can be readily integrated into modern network architectures. We conduct comprehensive evaluations on standard benchmarks in object detection, instance/semantic segmentation, and image inpainting. CARAFE++ shows consistent and substantial gains on mainstream methods across all the tasks with negligible computational overhead. It shows great potential to serve as a strong building block for modern deep networks. Jiaqi Wang 0003, Kai Chen 0026, Rui Xu 0014, Ziwei Liu 0002, Chen Change Loy, Dahua Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Temporal ROI Align for Video Object RecognitionabstractVideo object detection is challenging in the presence of appearance deterioration in certain video frames. Therefore, it is a natural choice to aggregate temporal information from other frames of the same video into the current frame. However, ROI Align, as one of the most core procedures of video detectors, still remains extracting features from a single-frame feature map for proposals, making the extracted ROI features lack temporal information from videos. In this work, considering the features of the same object instance are highly similar among frames in a video, a novel Temporal ROI Align operator is proposed to extract features from other frames feature maps for current frame proposals by utilizing feature similarity. The proposed Temporal ROI Align operator can extract temporal information from the entire video for proposals. We integrate it into single-frame video detectors and other state-of-the-art video detectors, and conduct quantitative experiments to demonstrate that the proposed Temporal ROI Align operator can consistently and significantly boost the performance. Besides, the proposed Temporal ROI Align can also be applied into video instance segmentation. Kai Chen 0026, Xinjiang Wang, Qi Chu 0001, Feng Zhu 0006, Dahua Lin, Nenghai Yu, Huamin Feng |
AAAI | 2 |
| 2021 | Seesaw Loss for Long-Tailed Instance SegmentationabstractInstance segmentation has witnessed a remarkable progress on class-balanced benchmarks. However, they fail to perform as accurately in real-world scenarios, where the category distribution of objects naturally comes with a long tail. Instances of head classes dominate a long-tailed dataset and they serve as negative samples of tail categories. The overwhelming gradients of negative samples on tail classes lead to a biased learning process for classifiers. Consequently, objects of tail categories are more likely to be misclassified as backgrounds or head categories. To tackle this problem, we propose Seesaw Loss to dynamically re-balance gradients of positive and negative samples for each category, with two complementary factors, i.e., mitigation factor and compensation factor. The mitigation factor reduces punishments to tail categories w.r.t. the ratio of cumulative training instances between different categories. Meanwhile, the compensation factor increases the penalty of misclassified instances to avoid false positives of tail categories. We conduct extensive experiments on Seesaw Loss with mainstream frameworks and different data sampling strategies. With a simple end-to-end training pipeline, Seesaw Loss obtains significant gains over Cross-Entropy Loss, and achieves state-of-the-art performance on LVIS dataset without bells and whistles. Code is available at https://github.com/open-mmlab/mmdetection. Jiaqi Wang 0003, Yuhang Zang, Yuhang Cao, Jiangmiao Pang, Kai Chen 0026, Ziwei Liu 0002, Chen Change Loy, Dahua Lin |
CVPR | 7 |
| 2021 | Positional Encoding As Spatial Inductive Bias in GANsabstractSinGAN shows impressive capability in learning internal patch distribution despite its limited effective receptive field. We are interested in knowing how such a translationinvariant convolutional generator could capture the global structure with just a spatially i.i.d. input. In this work, taking SinGAN and StyleGAN2 as examples, we show that such capability, to a large extent, is brought by the implicit positional encoding when using zero padding in the generators. Such positional encoding is indispensable for generating images with high fidelity. The same phenomenon is observed in other generative architectures such as DCGAN and PGGAN. We further show that zero padding leads to an unbalanced spatial bias with a vague relation between locations. To offer a better spatial inductive bias, we investigate alternative positional encodings and analyze their effects. Based on a more flexible positional encoding explicitly, we propose a new multi-scale training strategy and demonstrate its effectiveness in the state-of-the-art unconditional generator StyleGAN2. Besides, the explicit spatial inductive bias substantially improves SinGAN for more versatile image manipulation.1 Rui Xu 0014, Xintao Wang 0002, Kai Chen 0026, Bolei Zhou, Chen Change Loy |
CVPR | 3 |
| 2021 | MMOCR: A Comprehensive Toolbox for Text Detection, Recognition and UnderstandingabstractWe present MMOCR---an open-source toolbox which provides a comprehensive pipeline for text detection and recognition, as well as their downstream tasks such as named entity recognition and key information extraction. MMOCR implements 14 state-of-the-art algorithms, which is significantly more than all the existing open-source OCR projects we are aware of to date. To facilitate future research and industrial applications of text recognition-related problems, we also provide a large number of trained models and detailed benchmarks to give insights into the performance of text detection, recognition and understanding. MMOCR is publicly released at https://github.com/open-mmlab/mmocr. Zhanghui Kuang, Zhizhong Li 0002, Xiaoyu Yue, Tsui Hin Lin, Jianyong Chen, Huaqiang Wei, Yiqin Zhu, Kai Chen 0026, Wayne Zhang 0001, Dahua Lin |
ACM Multimedia | 11 |
| 2021 | Few-Shot Object Detection via Association and DIscriminationabstractObject detection has achieved substantial progress in the last decade. However, detecting novel classes with only few samples remains challenging, since deep learning under low data regime usually leads to a degraded feature space. Existing works employ a holistic fine-tuning paradigm to tackle this problem, where the model is first pre-trained on all base classes with abundant samples, and then it is used to carve the novel class feature space. Nonetheless, this paradigm is still imperfect. Durning fine-tuning, a novel class may implicitly leverage the knowledge of multiple base classes to construct its feature space, which induces a scattered feature space, hence violating the inter-class separability. To overcome these obstacles, we propose a two-step fine-tuning framework, Few-shot object detection via Association and DIscrimination (FADI), which builds up a discriminative feature space for each novel class with two integral steps. 1) In the association step, in contrast to implicitly leveraging multiple base classes, we construct a compact novel class feature space via explicitly imitating a specific base class feature space. Specifically, we associate each novel class with a base class according to their semantic similarity. After that, the feature space of a novel class can readily imitate the well-trained feature space of the associated base class. 2) In the discrimination step, to ensure the separability between the novel classes and associated base classes, we disentangle the classification branches for base and novel classes. To further enlarge the inter-class separability between all classes, a set-specialized margin loss is imposed. Extensive experiments on standard Pascal VOC and MS-COCO datasets demonstrate that FADI achieves new state-of-the-art performance, significantly improving the baseline in any shot/split by +18.7. Notably, the advantage of FADI is most announced on extremely few-shot scenarios (e.g. 1- and 3- shot). Yuhang Cao, Jiaqi Wang 0003, Kai Chen 0026, Ziwei Liu 0002, Dahua Lin |
NeurIPS | 5 |
| 2021 | K-Net: Towards Unified Image SegmentationabstractSemantic, instance, and panoptic segmentations have been addressed using different and specialized frameworks despite their underlying connections. This paper presents a unified, simple, and effective framework for these essentially similar tasks. The framework, named K-Net, segments both instances and semantic categories consistently by a group of learnable kernels, where each kernel is responsible for generating a mask for either a potential instance or a stuff class. To remedy the difficulties of distinguishing various instances, we propose a kernel update strategy that enables each kernel dynamic and conditional on its meaningful group in the input image. K-Net can be trained in an end-to-end manner with bipartite matching, and its training and inference are naturally NMS-free and box-free. Without bells and whistles, K-Net surpasses all previous published state-of-the-art single-model results of panoptic segmentation on MS COCO test-dev split and semantic segmentation on ADE20K val split with 55.2% PQ and 54.3% mIoU, respectively. Its instance segmentation performance is also on par with Cascade Mask R-CNN on MS COCO with 60%-90% faster inference speeds. Code and models will be released at https://github.com/ZwwWayne/K-Net/. Jiangmiao Pang, Kai Chen 0026, Chen Change Loy |
NeurIPS | 3 |
| 2021 | Towards Balanced Learning for Instance Recognition
Jiangmiao Pang, Kai Chen 0026, Qi Li 0018, Zhi-hai Xu, Huajun Feng, Jianping Shi, Wanli Ouyang, Dahua Lin |
Int. J. Comput. Vis. | 2 |
| 2020 | Prime Sample Attention in Object DetectionabstractIt is a common paradigm in object detection frameworks to treat all samples equally and target at maximizing the performance on average. In this work, we revisit this paradigm through a careful study on how different samples contribute to the overall performance measured in terms of mAP. Our study suggests that the samples in each mini-batch are neither independent nor equally important, and therefore a better classifier on average does not necessarily result in higher mAP. Motivated by this study, we propose the notion of Prime Samples, those that play a key role in driving the detection performance. We further develop a simple yet effective sampling and learning strategy called PrIme Sample Attention (PISA) that directs the focus of the training process towards such samples. Our experiments demonstrate that it is often more effective to focus on prime samples than hard samples when training a detector. Particularly, on the MSCOCO dataset, PISA outperforms the random sampling baseline and hard mining schemes, \eg~OHEM and Focal Loss, consistently by around 2\% on both single-stage and two-stage detectors, even with a strong backbone ResNeXt-101. Code is available at: \url{https://github.com/open-mmlab/mmdetection}. Yuhang Cao, Kai Chen 0026, Chen Change Loy, Dahua Lin |
CVPR | 2 |
| 2020 | Side-Aware Boundary Localization for More Precise Object Detection
Jiaqi Wang 0003, Yuhang Cao, Kai Chen 0026, Jiangmiao Pang, Jianping Shi, Chen Change Loy, Dahua Lin |
ECCV (4) | 4 |
| 2019 | Hybrid Task Cascade for Instance SegmentationabstractCascade is a classic yet powerful architecture that has boosted performance on various tasks. However, how to introduce cascade to instance segmentation remains an open question. A simple combination of Cascade R-CNN and Mask R-CNN only brings limited gain. In exploring a more effective approach, we find that the key to a successful instance segmentation cascade is to fully leverage the reciprocal relationship between detection and segmentation. In this work, we propose a new framework, Hybrid Task Cascade (HTC), which differs in two important aspects: (1) instead of performing cascaded refinement on these two tasks separately, it interweaves them for a joint multi-stage processing; (2) it adopts a fully convolutional branch to provide spatial context, which can help distinguishing hard foreground from cluttered background. Overall, this framework can learn more discriminative features progressively while integrating complementary features together in each stage. Without bells and whistles, a single HTC obtains 38.4% and 1.5% improvement over a strong Cascade Mask R-CNN baseline on MSCOCO dataset. Moreover, our overall system achieves 48.6 mask AP on the test-challenge split, ranking 1st in the COCO 2018 Challenge Object Detection Task. Code is available at https://github.com/open-mmlab/mmdetection. Kai Chen 0026, Jiangmiao Pang, Jiaqi Wang 0003, Shuyang Sun, Wansen Feng, Ziwei Liu 0002, Jianping Shi, Wanli Ouyang, Chen Change Loy, Dahua Lin |
CVPR | 1 |
| 2019 | Libra R-CNN: Towards Balanced Learning for Object DetectionabstractCompared with model architectures, the training process, which is also crucial to the success of detectors, has received relatively less attention in object detection. In this work, we carefully revisit the standard training practice of detectors, and find that the detection performance is often limited by the imbalance during the training process, which generally consists in three levels - sample level, feature level, and objective level. To mitigate the adverse effects caused thereby, we propose Libra R-CNN, a simple but effective framework towards balanced learning for object detection. It integrates three novel components: IoU-balanced sampling, balanced feature pyramid, and balanced L1 loss, respectively for reducing the imbalance at sample, feature, and objective level. Benefitted from the overall balanced design, Libra R-CNN significantly improves the detection performance. Without bells and whistles, it achieves 2.5 points and 2.0 points higher Average Precision (AP) than FPN Faster R-CNN and RetinaNet respectively on MSCOCO. Jiangmiao Pang, Kai Chen 0026, Jianping Shi, Huajun Feng, Wanli Ouyang, Dahua Lin |
CVPR | 2 |
| 2019 | Region Proposal by Guided AnchoringabstractRegion anchors are the cornerstone of modern object detection techniques. State-of-the-art detectors mostly rely on a dense anchoring scheme, where anchors are sampled uniformly over the spatial domain with a predefined set of scales and aspect ratios. In this paper, we revisit this foundational stage. Our study shows that it can be done much more effectively and efficiently. Specifically, we present an alternative scheme, named Guided Anchoring, which leverages semantic features to guide the anchoring. The proposed method jointly predicts the locations where the center of objects of interest are likely to exist as well as the scales and aspect ratios at different locations. On top of predicted anchor shapes, we mitigate the feature inconsistency with a feature adaption module. We also study the use of high-quality proposals to improve detection performance. The anchoring scheme can be seamlessly integrated into proposal methods and detectors. With Guided Anchoring, we achieve 9.1% higher recall on MS COCO with 90% fewer anchors than the RPN baseline. We also adopt Guided Anchoring in Fast R-CNN, Faster R-CNN and RetinaNet, respectively improving the detection mAP by 2.2%, 2.7% and 1.2%. Code is available at https://github.com/open-mmlab/mmdetection. Jiaqi Wang 0003, Kai Chen 0026, Shuo Yang 0003, Chen Change Loy, Dahua Lin |
CVPR | 2 |
| 2019 | CARAFE: Content-Aware ReAssembly of FEaturesabstractFeature upsampling is a key operation in a number of modern convolutional network architectures, e.g. feature pyramids. Its design is critical for dense prediction tasks such as object detection and semantic/instance segmentation. In this work, we propose Content-Aware ReAssembly of FEatures (CARAFE), a universal, lightweight and highly effective operator to fulfill this goal. CARAFE has several appealing properties: (1) Large field of view. Unlike previous works (e.g. bilinear interpolation) that only exploit subpixel neighborhood, CARAFE can aggregate contextual information within a large receptive field. (2) Content-aware handling. Instead of using a fixed kernel for all samples (e.g. deconvolution), CARAFE enables instance-specific content-aware handling, which generates adaptive kernels on-the-fly. (3) Lightweight and fast to compute. CARAFE introduces little computational overhead and can be readily integrated into modern network architectures. We conduct comprehensive evaluations on standard benchmarks in object detection, instance/semantic segmentation and inpainting. CARAFE shows consistent and substantial gains across all the tasks (1.2% AP, 1.3% AP, 1.8% mIoU, 1.1dB respectively) with negligible computational overhead. It has great potential to serve as a strong building block for future research. Code and models are available at https://github.com/open-mmlab/mmdetection. Jiaqi Wang 0003, Kai Chen 0026, Rui Xu 0014, Ziwei Liu 0002, Chen Change Loy, Dahua Lin |
ICCV | 2 |
| 2018 | Optimizing Video Object Detection via a Scale-Time LatticeabstractHigh-performance object detection relies on expensive convolutional networks to compute features, often leading to significant challenges in applications, e.g. those that require detecting objects from video streams in real time. The key to this problem is to trade accuracy for efficiency in an effective way, i.e. reducing the computing cost while maintaining competitive performance. To seek a good balance, previous efforts usually focus on optimizing the model architectures. This paper explores an alternative approach, that is, to reallocate the computation over a scale-time space. The basic idea is to perform expensive detection sparsely and propagate the results across both scales and time with substantially cheaper networks, by exploiting the strong correlations among them. Specifically, we present a unified framework that integrates detection, temporal propagation, and across-scale refinement on a Scale-Time Lattice. On this framework, one can explore various strategies to balance performance and cost. Taking advantage of this flexibility, we further develop an adaptive scheme with the detector invoked on demand and thus obtain improved tradeoff. On ImageNet VID dataset, the proposed method can achieve a competitive mAP 79.6% at 20 fps, or 79.0% at 62 fps as a performance/speed tradeoff.1 Kai Chen 0026, Jiaqi Wang 0003, Shuo Yang 0003, Xingcheng Zhang, Yuanjun Xiong, Chen Change Loy, Dahua Lin |
CVPR | 1 |
| 2017 | Discover and Learn New Objects from DocumentariesabstractDespite the remarkable progress in recent years, detecting objects in a new context remains a challenging task. Detectors learned from a public dataset can only work with a fixed list of categories, while training from scratch usually requires a large amount of training data with detailed annotations. This work aims to explore a novel approach – learning object detectors from documentary films in a weakly supervised manner. This is inspired by the observation that documentaries often provide dedicated exposition of certain object categories, where visual presentations are aligned with subtitles. We believe that object detectors can be learned from such a rich source of information. Towards this goal, we develop a joint probabilistic framework, where individual pieces of information, including video frames and subtitles, are brought together via both visual and linguistic links. On top of this formulation, we further derive a weakly supervised learning algorithm, where object model learning and training set mining are unified in an optimization procedure. Experimental results on a real world dataset demonstrate that this is an effective approach to learning new object detectors. Kai Chen 0026, Chen Change Loy, Dahua Lin |
CVPR | 1 |