VLDB 2026 Research / reviewers in the wild / expert
Xiangxiang Chu
dblp:207/8002
· DBLP profile ↗
49ranked-venue papers
13as first author
44since 2021 · last 2026
0000-0003-2548-0605ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 12 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 10 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical RevisitingabstractReinforcement learning (RL) has demonstrated considerable potential for enhancing reasoning in large language models (LLMs). However, existing methods suffer from Gradient Starvation and Policy Degradation when training directly on samples with mixed difficulty. To mitigate this, prior approaches leverage Chain-of-Thought (CoT) data, but the construction of high-quality CoT annotations remains labor-intensive. Alternatively, curriculum learning strategies have been explored but frequently encounter challenges, such as difficulty mismatch, reliance on manual curriculum design, and catastrophic forgetting. To address these issues, we propose AdaCuRL, a Adaptive Curriculum Reinforcement Learning framework that integrates coarse-to-fine difficulty estimation with adaptive curriculum scheduling. This approach dynamically aligns data difficulty with model capability and incorporates a data revisitation mechanism to mitigate catastrophic forgetting. Furthermore, AdaCuRL employs adaptive reference and sparse KL strategies to prevent Policy Degradation. Extensive experiments across diverse reasoning benchmarks demonstrate that AdaCuRL consistently achieves significant performance improvements on both LLMs and MLLMs. Renda Li, Hailang Huang, Fei Wei, Xiangxiang Chu |
AAAI | 6 |
| 2026 | Omni-Effects: Unified and Spatially-Controllable Visual Effects GenerationabstractVisual effects (VFX) are essential visual enhancements fundamental to modern cinematic production. Although video generation models offer cost-efficient solutions for VFX production, current methods are constrained by per-effect LoRA training, which limits generation to single effects. This fundamental limitation impedes applications that require spatially controllable composite effects, i.e., the concurrent generation of multiple effects at designated locations. However, integrating diverse effects into a unified framework faces major challenges: interference from effect variations and spatial uncontrollability during multi-VFX joint training. To tackle these challenges, we propose Omni-Effects, a first unified framework capable of generating prompt-guided effects and spatially controllable composite effects. The core of our framework comprises two key innovations: (1) LoRA-based Mixture of Experts (LoRA-MoE), which employs a group of expert LoRAs, integrating diverse effects within a unified model while effectively mitigating cross-task interference. (2) Spatial-Aware Prompt (SAP) incorporates spatial mask information into the text token, enabling precise spatial control. Furthermore, we introduce an Independent-Information Flow (IIF) module integrated within the SAP, isolating the control signals corresponding to individual effects to prevent any unwanted blending. To facilitate this research, we construct a comprehensive VFX dataset Omni-VFX via a novel data collection pipeline combining image editing and First-Last Frame-to-Video (FLF2V) synthesis, and introduce a dedicated VFX evaluation framework for validating model performance. Extensive experiments demonstrate that Omni-Effects achieves precise spatial control and diverse effect generation, enabling users to specify both the category and location of desired effects. Fangyuan Mao, Aiming Hao, Dongxia Liu, Xiaokun Feng, Jiashu Zhu, Meiqi Wu, Chubin Chen, Jiahong Wu 0005, Xiangxiang Chu |
AAAI | 10 |
| 2026 | ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency ConstraintsabstractVideo generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training distributions. Existing methods typically apply test-time scaling for improving video quality, but their fixed search spaces and static reward designs limit adaptability to imaginative scenarios. To fill this gap, we propose ImagerySearch, a dynamic test-time scaling law strategy inspired by imagery that adaptively adjusts the inference search space and reward guided by prompts, effectively enhancing generation quality in imaginative scenarios. Furthermore, we introduce LDT-Bench, the first benchmark targeting long-distance semantic prompts, designed to evaluate the creativity of video generation models. It comprises 2,839 challenging concept pairs from diverse recognition datasets and incorporates an automatic evaluation protocol to assess creative capacity. Extensive experiments on LDT-Bench demonstrate that our approach consistently outperforms general generation models and test-time scaling approaches. Additionally, ImagerySearch achieves strong performance on VBench, confirming its effectiveness in improving video generation quality under diverse conditions. Meiqi Wu, Jiashu Zhu, Xiaokun Feng, Chubin Chen, Bingze Song, Fangyuan Mao, Jiahong Wu 0005, Xiangxiang Chu, Kaiqi Huang |
AAAI | 9 |
| 2026 | SCALAR: Scale-wise Controllable Visual Autoregressive LearningabstractControllable image synthesis, which enables fine-grained control over generated outputs, has emerged as a key focus in visual generative modeling. However, controllable generation remains challenging for Visual Autoregressive (VAR) models due to their hierarchical, next-scale prediction style. Existing VAR-based methods often suffer from inefficient control encoding and disruptive injection mechanisms that compromise both fidelity and efficiency. In this work, we present SCALAR, a controllable generation method based on VAR, incorporating a Scale-wise Conditional Decoding mechanism. SCALAR leverages a pretrained image encoder to extract semantic control signal encodings, which are projected into scale-specific representations and injected into the corresponding layers of the VAR backbone. This design provides persistent and structurally aligned guidance throughout the generation process. Building on SCALAR, we develop SCALAR-Uni, a unified extension that aligns multiple control modalities into a shared latent space, supporting flexible multi-conditional guidance in a single model. Extensive experiments show that SCALAR achieves superior generation quality and control precision across various tasks. Ryan Xu, Dongyang Jin, Yancheng Bai, Rui Lan, Xu Duan, Xiangxiang Chu |
AAAI | 7 |
| 2026 | Visually-Guided Policy Optimization for Multimodal ReasoningabstractZengbin Wang, Feng Xiong, Liang Lin, Xuecai Hu, Yong Wang, Yanlin Wang, Man Zhang, Xiangxiang Chu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zengbin Wang, Liang Lin 0004, Xuecai Hu, Xiangxiang Chu |
ACL (1) | 8 |
| 2026 | CoEvolve: Training LLM Agents via Agent-Data Mutual EvolutionabstractReinforcement learning for LLM agents is typically conducted on a static data distribution, which fails to adapt to the agent's evolving behavior and leads to poor coverage of complex environment interactions.To address these challenges, we propose CoEvolve, an agentdata mutual evolution framework that enables LLM agents to improve through closed-loop, interaction-driven training.Specifically, CoEvolve extracts feedback signals such as forgetting and uncertainty from rollout trajectories to identify failure-prone interaction patterns, and utilizes them to guide LLM-based task synthesis.The synthesized tasks are validated through environment interaction and utilized to update the data distribution, enabling joint adaptation of the agent and its data.Extensive experiments on AppWorld and BFCL across Qwen2.5-7B,Qwen3-4B, and Qwen3-30B-A3B demonstrate consistent and significant improvements over strong base models, yielding absolute gains of 19.43%, 15.58%, and 18.14%, respectively. Shidong Yang, Ziyu Ma, Tongwen Huang, Yiming Hu, Xiangxiang Chu |
ACL (1) | 6 |
| 2026 | L2Dir: Integrating L_2-Norm and Directional Alignment for Unsupervised Contrastive Representation Learning in Multimodal RetrievalabstractTianyu Zong, Rui Dai, Hongzhu Yi, Yuanxiang Wang, Zhenghao Zhang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Yueyang Ding, Xiangxiang Chu, Kaikui Liu, Jungang Xu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianyu Zong, Hongzhu Yi, Yuanxiang Wang, Yujia Yang, Bingkang Shi, Yueyang Ding, Xiangxiang Chu, Kaikui Liu, Jungang Xu |
ACL (1) | 10 |
| 2026 | Breaking Block Boundaries: Anchor-based History-stable Decoding for Diffusion Large Language ModelsabstractShun Zou, Yong Wang, Zehui Chen, Lin Chen, Chongyang Tao, Feng Zhao, Xiangxiang Chu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shun Zou, Chongyang Tao, Feng Zhao 0004, Xiangxiang Chu |
ACL (1) | 7 |
| 2026 | Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task DifficultyabstractVisual instruction tuning is a key training stage of large multimodal models. However, when learning multiple visual tasks simultaneously, this approach often results in suboptimal and imbalanced overall performance due to latent knowledge conflicts across tasks. To mitigate this issue, we propose a novel Adaptive Task Balancing approach tailored for visual instruction tuning (VisATB). Specifically, we measure two critical dimensions for visual task balancing based on validation performance: (1) Inter-Task Contribution, the mechanism where learning one task enhances the performance on others owing to shared knowledge across tasks, and (2) Intra-Task Difficulty, which denotes the inherent learning difficulty of a single task. Furthermore, we propose prioritizing three categories of tasks with greater weight: those that offer substantial contributions to others, those that receive minimal contributions from others, and those that present high learning difficulties. Among these three task weighting strategies, the first and third focus on improving overall performance, and the second targets the mitigation of performance imbalance. Extensive experiments on three benchmarks demonstrate that our VisATB approach consistently achieves superior and more balanced overall performance in visual instruction tuning. The data, code, and models are available at https://github.com/YanqiDai/VisATB. Yanqi Dai, Zebin You, Dong Jing, Xiangxiang Chu, Zhiwu Lu 0001 |
WWW | 5 |
| 2026 | FastPillars: A Deployment-Friendly Pillar-Based 3D DetectorabstractThe deployment of 3D detectors strikes one of the major challenges in real-world self-driving scenarios. Existing BEV-based (i.e., Bird Eye View) detectors favor sparse convolutions (known as SPConv) to speed up training and inference, which puts a hard barrier for deployment, especially for on-device applications. In this paper, in order to tackle the challenge of efficient 3D object detection from an industry perspective, we devise a deployment-friendly pillar-based 3D detector, termed FastPillars. Specifically, aiming to compensate the geometric information loss of pillar encoding. First, we design a novel lightweight Max-and-Attention Pillar Encoding (MAPE) module specially for enhancing small objects. Second, we propose a simple yet effective backbone design for pillar-based 3D detection, enhancing pillar representations. We construct FastPillars based on these designs, achieving high performance and low latency without SPConv. Extensive experiments on two large-scale datasets demonstrate the effectiveness and efficiency of FastPillars for on-device 3D detection regarding both performance and speed. Specifically, FastPillars delivers real-time state-of-the-art accuracy on Waymo Open Dataset with 1.8 × speed up and 3.8 mAPH/L2 improvement over CenterPoint (SPConv-based). Code will be opened soon in: https://github.com/StiphyJay/FastPillars. Sifan Zhou, Xinyu Zhang 0015, Xiangxiang Chu, Bo Zhang 0046, Xiaobo Lu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Towards Efficient Foundation Model for Zero-shot Amodal SegmentationabstractAiming to predict the complete shape of partially occluded objects, amodal segmentation is an important capacity towards visual intelligence. In order to promote the practicability, zero-shot foundation model competent for the open world gains growing attention in this field. Nevertheless, prior models exhibit deficiencies in efficiency and stability. To address this problem, utilizing the implicit prior knowledge, we propose the first SAM-based amodal segmentation foundation model, SAMBA. Methodologically, a novel framework with multilevel facilitation is designed to better adapt the task characteristics and unleash the potential capabilities of SAM. In the modality level, a separation-to-fusion structure is employed that jointly learns modal and amodal segmentation to enhance mutual coordination. In the instance level, to ease the complexity of amodal feature extraction, we introduce a principal focusing mechanism to indicate objects of interest. In the pixel level, mixture-of-experts is incorporated with a specialized distribution loss, by which distinct occlusion rates correspond to different experts to improve the accuracy. Experiments are conducted on several eminent datasets, and the results show that the performance of SAMBA is superior to existing zero-shot and even supervised approaches. Furthermore, our proposed model has notable advantages in terms of speed and size. Zhaochen Liu, Limeng Qiao, Xiangxiang Chu, Lin Ma 0002, Tingting Jiang 0001 |
CVPR | 3 |
| 2025 | POSITION BIAS MITIGATES POSITION BIAS: Mitigate Position Bias Through Inter-Position Knowledge DistillationabstractPositional bias (PB), manifesting as nonuniform sensitivity across different contextual locations, significantly impairs long-context comprehension and processing capabilities.Previous studies have addressed PB either by modifying the underlying architectures or by employing extensive contextual awareness training.However, the former approach fails to effectively eliminate the substantial performance disparities, while the latter imposes significant data and computational overhead.To address PB effectively, we introduce Pos2Distill, a position to position knowledge distillation framework.Pos2Distill transfers the superior capabilities from advantageous positions to less favorable ones, thereby reducing the huge performance gaps.The conceptual principle is to leverage the inherent, position-induced disparity to counteract the PB itself.We identify distinct manifestations of PB under Retrieval and Reasoning paradigms, thereby designing two specialized instantiations: Pos2Distill-R 1 and Pos2Distill-R 2 respectively, both grounded in this core principle.By employing our approach, we achieve enhanced uniformity and significant performance gains across all contextual positions in long-context retrieval and reasoning tasks.Crucially, both specialized systems exhibit strong cross-task generalization mutually, while achieving superior performance on their respective tasks. Linjing Li, Xiangxiang Chu, Daniel Dajun Zeng |
EMNLP | 5 |
| 2025 | HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget ReallocationabstractSelf-taught reasoners (STaRs) enhance the mathematical reasoning abilities of large language models (LLMs) by leveraging selfgenerated responses for self-training.Recent studies have incorporated reward models to guide response selection or decoding, aiming to obtain higher-quality data.However, they typically allocate a uniform sampling budget across all problems, overlooking the varying utility of problems at different difficulty levels.In this work, we conduct an empirical study and find that problems near the boundary of the LLM's reasoning capability offer significantly greater learning utility than both easy and overly difficult ones.To identify and exploit such problems, we propose HS-STAR, a Hierarchical Sampling framework for Self-Taught Reasoners.Given a fixed sampling budget, HS-STAR first performs lightweight presampling with a reward-guided difficulty estimation strategy to efficiently identify boundarylevel problems.Subsequently, it dynamically reallocates the remaining budget toward these high-utility problems during a re-sampling phase, maximizing the generation of valuable training data.Extensive experiments across multiple reasoning benchmarks and backbone LLMs demonstrate that HS-STAR significantly outperforms other baselines without requiring additional sampling budget.78 81 84 87 GSM8K 80.1 81.0 79.5 81.1 80.2 79.9 82.8 80. Hongling Xu, Runxi Cheng, Xiangxiang Chu |
EMNLP | 6 |
| 2025 | Lenna: Language Enhanced Reasoning Detection AssistantabstractWith the fast-paced development of multimodal large language models (MLLMs), we can now converse with AI systems in natural languages to understand images. However, the reasoning power and world knowledge embedded in the large language models have been much less investigated and exploited for image perception tasks. In this paper, we propose Lenna, a Language enhanced reasoning detection assistant, which utilizes the robust multimodal feature representation of MLLMs, while preserving location information for detection. This is achieved by incorporating an additionaltoken in the MLLM vocabulary that is free of explicit semantic context but serves as a prompt for the detector to identify the corresponding position. To evaluate the reasoning capability of Lenna, we construct a ReasonDet dataset to measure its performance on reasoning-based detection. Remarkably, Lenna demonstrates outstanding performance on ReasonDet and comes with significantly low training costs. It also incurs minimal transferring overhead when extended to other tasks. Fei Wei, Xinyu Zhang 0015, Bo Zhang 0046, Xiangxiang Chu |
ICASSP | 5 |
| 2025 | USP: Unified Self-Supervised Pretraining for Image Generation and UnderstandingabstractRecent studies have highlighted the interplay between diffusion models and representation learning. Intermediate representations from diffusion models can be leveraged for downstream visual tasks, while self-supervised vision models can enhance the convergence and generation quality of diffusion models. However, transferring pretrained weights from vision models to diffusion models is challenging due to input mismatches and the use of latent spaces. To address these challenges, we propose Unified Self-supervised Pretraining (USP), a framework that initializes diffusion models via masked latent modeling in a Variational Autoencoder (VAE) latent space. USP achieves comparable performance in understanding tasks while significantly improving the convergence speed and generation quality of diffusion models. Our code will be publicly available at https://github.com/AMAP-ML/USP. Xiangxiang Chu, Renda Li |
ICCV | 1 |
| 2025 | LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior SamplingabstractUnified image restoration is a significantly challenging task in low-level vision. Existing methods either make tailored designs for specific tasks, limiting their generalizability across various types of degradation, or rely on training with paired datasets, thereby suffering from closed-set constraints. To address these issues, we propose a novel, dataset-free, and unified approach through recurrent posterior sampling utilizing a pretrained latent diffusion model. Our method incorporates the multimodal understanding model to provide sematic priors for the generative model under a task-blind condition. Furthermore, it utilizes a lightweight module to align the degraded input with the generated preference of the diffusion model, and employs recurrent refinement for posterior sampling. Extensive experiments demonstrate that our method outperforms state-of-the-art methods, validating its effectiveness and robustness. Our code and data are available at https://github.com/AMAP-ML/LD-RPS. Huaqiu Li, Tongwen Huang, Hailang Huang, Haoqian Wang, Xiangxiang Chu |
ICCV | 6 |
| 2025 | VMBench: A Benchmark for Perception-Aligned Video Motion Generation
Xinran Ling, Meiqi Wu, Xiaokun Feng, Cundian Yang, Aiming Hao, Jiashu Zhu, Jiahong Wu 0005, Xiangxiang Chu |
ICCV | 10 |
| 2025 | UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation EnhancementabstractZero-shot domain adaptation (ZSDA) presents substantial challenges due to the lack of images in the target domain. Previous approaches leverage Vision-Language Models (VLMs) to tackle this challenge, exploiting their zero-shot learning capabilities. However, these methods primarily address domain distribution shifts and overlook the misalignment between the detection task and VLMs, which rely on manually crafted prompts. To overcome these limitations, we propose the unified prompt and representation enhancement (UPRE) framework, which jointly optimizes both textual prompts and visual representations. Specifically, our approach introduces a multi-view domain prompt that combines linguistic domain priors with detection-specific knowledge, and a visual representation enhancement module that produces domain style variations. Furthermore, we introduce multi-level enhancement strategies, including relative domain distance and positive-negative separation, which align multi-modal representations at the image level and capture diverse visual representations at the instance level, respectively. Extensive experiments conducted on nine benchmark datasets demonstrate the superior performance of our framework in ZSDA detection scenarios. Code is available at https://github.com/AMAP-ML/UPRE. Xiao Zhang 0050, Fei Wei, Wenda Zhao 0003, Feiyi Li, Xiangxiang Chu |
ICCV | 6 |
| 2025 | Contrastive Instruction Fine-Tuning Large Multimodal Model for Hateful Meme ClassificationabstractDetecting hateful memes requires a model that possesses extensive background knowledge and robust reasoning abilities, especially when the memes contain ambiguous descriptions. Previous research has used large language models (LLMs) and large multimodal models (LMMs) to interpret and categorize these memes. However, distinguishing subtly different hateful and non-hateful memes is still challenging. In recognition of this, our study introduces a unique contrastive instruction fine-tuning approach, InstructMemeCL. This method improves an LMM's ability to discern between memes that have similar visual or textual elements by intensifying its focus on semantic subtleties that separate hateful from non-hateful content. We evaluated our model using AUROC and accuracy metrics on three publicly available hateful meme datasets. The results indicate that our improved LMM more accurately identifies hateful and non-hateful memes, demonstrating superior performance compared to conventional LLMs and LMMs used in similar tasks. Ming Shan Hee, Xiangxiang Chu, Roy Ka-Wei Lee, Zengchang Qin |
ICWSM | 4 |
| 2025 | FingER: Content Aware Fine-grained Evaluation with Reasoning for AI-Generated VideosabstractRecent advances in video generation have posed great challenges in the assessment of AI-generated content, particularly with the emergence of increasingly sophisticated models. The various inconsistencies and defects observed in such videos are inherently complex, making overall scoring notoriously difficult. In this paper, we emphasize the critical importance of integrating fine-grained reasoning into video evaluation. We propose FingER, a novel entity-level reasoning evaluation framework that first automatically generates Fine-grained Entity-level questions, and then answers those questions by a Reasoning model with scores, which can be subsequently weighted summed to an overall score for different applications. Specifically, we leverage LLMs to derive entity-level questions across five distinct perspectives, which (i) often focus on some specific entities of the content, thereby making answering or scoring much easier for MLLMs, and (ii) are more interpretable. Then we construct a FingER dataset, consisting of approximately 3.3k videos and corresponding 60k fine-grained QA annotations, each with detailed reasons. Based on that, we further investigate various training protocols to best incentivize the reasoning capability of MLLMs for correct answer prediction. Extensive experiments demonstrate that a reasoning model trained using GRPO with a cold-start strategy achieves the best performance. Notably, our model surpasses existing methods by a relative margin of 11.8% on GenAI-Bench and 5.5% on MonetBench with only 3.3k training videos, which is at most one-tenth of the training samples utilized by other methods. Our codes and datasets have been released. Rui Chen 0016, Lei Sun 0009, Jing Tang 0006, Xiangxiang Chu |
ACM Multimedia | 5 |
| 2024 | Make RepVGG Greater Again: A Quantization-Aware ApproachabstractThe tradeoff between performance and inference speed is critical for practical applications. Architecture reparameterization obtains better tradeoffs and it is becoming an increasingly popular ingredient in modern convolutional neural networks. Nonetheless, its quantization performance is usually too poor to deploy (e.g. more than 20% top-1 accuracy drop on ImageNet) when INT8 inference is desired. In this paper, we dive into the underlying mechanism of this failure, where the original design inevitably enlarges quantization error. We propose a simple, robust, and effective remedy to have a quantization-friendly structure that also enjoys reparameterization benefits. Our method greatly bridges the gap between INT8 and FP32 accuracy for RepVGG. Without bells and whistles, the top-1 accuracy drop on ImageNet is reduced within 2% by standard post-training quantization. Extensive experiments on detection and semantic segmentation tasks verify its generalization. Xiangxiang Chu, Liang Li 0003, Bo Zhang 0046 |
AAAI | 1 |
| 2024 | Norm Tweaking: High-Performance Low-Bit Quantization of Large Language ModelsabstractAs the size of large language models (LLMs) continues to grow, model compression without sacrificing accuracy has become a crucial challenge for deployment. While some quantization methods, such as GPTQ, have made progress in achieving acceptable 4-bit weight-only quantization, attempts at lower-bit quantization often result in severe performance degradation. In this paper, we introduce a technique called norm tweaking, which can be used as a plugin in current PTQ methods to achieve high precision while being cost-efficient. Our approach is inspired by the observation that rectifying the quantized activation distribution to match its float counterpart can readily restore accuracy for LLMs. To achieve this, we carefully design a tweaking strategy that includes calibration data generation and channel-wise distance constraint to update the weights of normalization layers for better generalization. We conduct extensive experiments on various datasets using several open-sourced LLMs. Our method demonstrates significant improvements in both weight-only quantization and joint quantization of weights and activations, surpassing existing PTQ methods. On GLM-130B and OPT-66B, our method even achieves the same level of accuracy at 2-bit quantization as their float ones. Our simple and effective approach makes it more practical for real-world applications. Liang Li 0003, Qingyuan Li 0001, Bo Zhang 0046, Xiangxiang Chu |
AAAI | 4 |
| 2024 | SCTNet: Single-Branch CNN with Transformer Semantic Information for Real-Time SegmentationabstractRecent real-time semantic segmentation methods usually adopt an additional semantic branch to pursue rich long-range context. However, the additional branch incurs undesirable computational overhead and slows inference speed. To eliminate this dilemma, we propose SCTNet, a single branch CNN with transformer semantic information for real-time segmentation. SCTNet enjoys the rich semantic representations of an inference-free semantic branch while retaining the high efficiency of lightweight single branch CNN. SCTNet utilizes a transformer as the training-only semantic branch considering its superb ability to extract long-range context. With the help of the proposed transformer-like CNN block CFBlock and the semantic information alignment module, SCTNet could capture the rich semantic information from the transformer branch in training. During the inference, only the single branch CNN needs to be deployed. We conduct extensive experiments on Cityscapes, ADE20K, and COCO-Stuff-10K, and the results show that our method achieves the new state-of-the-art performance. The code and model is available at https://github.com/xzz777/SCTNet. Zhengze Xu, Dongyue Wu, Changqian Yu, Xiangxiang Chu, Nong Sang, Changxin Gao |
AAAI | 4 |
| 2024 | PeLK: Parameter-Efficient Large Kernel ConvNets with Peripheral ConvolutionabstractRecently, some large kernel convnets strike back with appealing performance and efficiency. However, given the square complexity of convolution, scaling up kernels can bring about an enormous amount of parameters and the proliferated parameters can induce severe optimization problem. Due to these issues, current CNNs compromise to scale up to 51 × 51 in the form of stripe convolution (i.e., 51 ×5 + 5 ×51) and start to saturate as the kernel size continues growing. In this paper, we delve into addressing these vital issues and explore whether we can continue scaling up kernels for more performance gains. Inspired by human vision, we propose a human-like peripheral convolution that efficiently reduces over 90% parameter count of dense grid convolution through parameter sharing, and manage to scale up kernel size to extremely large. Our peripheral convolution behaves highly similar to human, reducing the complexity of convolution from O(K2) to O(logK) without backfiring performance. Built on this, we propose Parameter-efficient Large Kernel Network (PeLK). Our PeLK outperforms modern vision Transformers and ConvNet architectures like Swin, ConvNeXt, RepLKNet and SLaK on various vision tasks including ImageNet classification, semantic segmentation on ADE20K and object detection on MS COCO. For the first time, we successfully scale up the kernel size of CNNs to an unprecedented 101 × 101 and demonstrate consistent improvements. Xiangxiang Chu, Xin Zhao 0012, Kaiqi Huang |
CVPR | 2 |
| 2024 | VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
Xiangxiang Chu, Jianlin Su, Bo Zhang 0046, Chunhua Shen |
ECCV (66) | 1 |
| 2024 | Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition
Xiangxiang Chu |
ECCV (71) | 4 |
| 2024 | LiDAR-PTQ: Post-Training Quantization for Point Cloud 3D Object DetectionabstractDue to highly constrained computing power and memory, deploying 3D lidar-based detectors on edge devices equipped in autonomous vehicles and robots poses a crucial challenge. Being a convenient and straightforward model compression approach, Post-Training Quantization (PTQ) has been widely adopted in 2D vision tasks. However, applying it directly to 3D lidar-based tasks inevitably leads to performance degradation. As a remedy, we propose an effective PTQ method called LiDAR-PTQ, which is particularly curated for 3D lidar detection (both SPConv-based and SPConv-free). Our LiDAR-PTQ features three main components, (1) a sparsity-based calibration method to determine the initialization of quantization parameters, (2) an adaptive rounding-to-nearest operation to minimize the layerwise reconstruction error, (3) a Task-guided Global Positive Loss (TGPL) to reduce the disparity between the final predictions before and after quantization. Extensive experiments demonstrate that our LiDAR-PTQ can achieve state-of-the-art quantization performance when applied to CenterPoint (both Pillar-based and Voxel-based). To our knowledge, for the very first time in lidar-based 3D detection tasks, the PTQ INT8 model's accuracy is almost the same as the FP32 model while enjoying 3X inference speedup. Moreover, our LiDAR-PTQ is cost-effective being 6X faster than the quantization-aware training method. The code will be released. Sifan Zhou, Liang Li 0003, Xinyu Zhang 0015, Bo Zhang 0046, Shipeng Bai, Xiaobo Lu, Xiangxiang Chu |
ICLR | 9 |
| 2024 | Revealing the Dark Secrets of Extremely Large Kernel ConvNets on RobustnessabstractRobustness is a vital aspect to consider when deploying deep learning models into the wild. Numerous studies have been dedicated to the study of the robustness of vision transformers (ViTs), which have dominated as the mainstream backbone choice for vision tasks since the dawn of 2020s. Recently, some large kernel convnets make a comeback with impressive performance and efficiency. However, it still remains unclear whether large kernel networks are robust and the attribution of their robustness. In this paper, we first conduct a comprehensive evaluation of large kernel convnets’ robustness and their differences from typical small kernel counterparts and ViTs on six diverse robustness benchmark datasets. Then to analyze the underlying factors behind their strong robustness, we design experiments from both quantitative and qualitative perspectives to reveal large kernel convnets’ intriguing properties that are completely different from typical convnets. Our experiments demonstrate for the first time that pure CNNs can achieve exceptional robustness comparable or even superior to that of ViTs. Our analysis on occlusion invariance, kernel attention patterns and frequency characteristics provide novel insights into the source of robustness. Code available at: https://github.com/Lauch1ng/LKRobust. Xiaokun Feng, Xiangxiang Chu, Kaiqi Huang |
ICML | 4 |
| 2024 | AODet: Aerial Object Detection Using Transformers for Foreground RegionsabstractAerial object detection is an important task and has received significant attention in recent years. Aerial images typically depict small and sparse instances against a simple background. Nevertheless, the simple background can only provide limited information. Based on the observation, we present a new transformer-based framework for aerial object detection. In contrast to previous methods that address sparsity through multi-stage pipelines involving Region-of-Interest (RoI) techniques or Sparse Convolutions, our method, referred as AODet, enjoy two significant advantages: 1) AODet is a simple yet accurate object detector which is specialized for aerial object detection. AODet identifies the background regions earlier and then only operates on the regions which most likely include the foreground objects, thereby significantly reducing the redundant computations. The utilization of transformer exploits more context information between foreground regions, helping to retain high-quality detection results. 2) Instead of involving the sparse operations like Sparse Convolutions or Clustering algorithms/ROI operations, AODet employs transformer to detect objects from foreground proposals. Our approach is simpler and can be easily implemented with simple tensor manipulations. Extensive experiments have conducted on VisDrone and DOTA. AODet achieves 40.9 AP on Visdrone and 79.6 mAP DOTA, demonstrating the effectiveness of AODet. Xiaoming Wang 0010, Hao Chen 0041, Xiangxiang Chu, Peng Wang 0015 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | AeDet: Azimuth-Invariant Multi-View 3D Object DetectionabstractRecent LSS-based multi-view 3D object detection has made tremendous progress, by processing the features in Brid-Eye-View (BEV) via the convolutional detector. However, the typical convolution ignores the radial symmetry of the BEV features and increases the difficulty of the detector optimization. To preserve the inherent property of the BEV features and ease the optimization, we propose an azimuth-equivariant convolution (AeConv) and an azimuth-equivariant anchor. The sampling grid of AeConv is always in the radial direction, thus it can learn azimuth-invariant BEV features. The proposed anchor enables the detection head to learn predicting azimuth-irrelevant targets. In addition, we introduce a camera-decoupled virtual depth to unify the depth prediction for the images with different camera intrinsic parameters. The resultant detector is dubbed Azimith-equivariant Detector (AeDet). Extensive experiments are conducted on nuScenes, and AeDet achieves a 62.0% NDS, surpassing the recent multi-view 3D object detectors such as PETRv2 and BEVDepth by a large margin. Project page: https://fcjian.github.io/aedet. Chengjian Feng, Zequn Jie, Xiangxiang Chu, Lin Ma 0002 |
CVPR | 4 |
| 2023 | MixPath: A Unified Approach for One-shot Neural Architecture SearchabstractBlending multiple convolutional kernels is proved advantageous in neural architecture design. However, current two-stage neural architecture search methods are mainly limited to single-path search spaces. How to efficiently search models of multi-path structures remains a difficult problem. In this paper, we are motivated to train a one-shot multi-path supernet to accurately evaluate the candidate architectures. Specifically, we discover that in the studied search spaces, feature vectors summed from multiple paths are nearly multiples of those from a single path. Such disparity perturbs the supernet training and its ranking ability. Therefore, we propose a novel mechanism called Shadow Batch Normalization (SBN) to regularize the disparate feature statistics. Extensive experiments prove that SBNs are capable of stabilizing the optimization and improving ranking performance. We call our unified multi-path one-shot approach as MixPath, which generates a series of models that achieve state-of-the-art results on ImageNet. Xiangxiang Chu, Shun Lu 0001, Xudong Li 0003, Bo Zhang 0046 |
ICCV | 1 |
| 2023 | ROME: Robustifying Memory-Efficient NAS via Topology Disentanglement and Gradient AccumulationabstractAlbeit being a prevalent architecture searching approach, differentiable architecture search (DARTS) is largely hindered by its substantial memory cost since the entire supernet resides in the memory. This is where the single-path DARTS comes in, which only chooses a single-path submodel at each step. While being memory-friendly, it also comes with low computational costs. Nonetheless, we discover a critical issue of single-path DARTS that has not been primarily noticed. Namely, it also suffers from severe performance collapse since too many parameter-free operations like skip connections are derived, just like DARTS does. In this paper, we propose a new algorithm called RObustifying Memory-Efficient NAS (ROME) to give a cure. First, we disentangle the topology search from the operation search to make searching and evaluation consistent. We then adopt Gumbel-Top2 reparameterization and gradient accumulation to robustify the unwieldy bi-level optimization. We verify ROME extensively across 15 benchmarks to demonstrate its effectiveness and robustness. Xiaoxing Wang, Xiangxiang Chu, Yuda Fan, Zhexi Zhang, Bo Zhang 0046, Xiaokang Yang 0001, Junchi Yan |
ICCV | 2 |
| 2023 | Conditional Positional Encodings for Vision Transformers
Xiangxiang Chu, Zhi Tian, Bo Zhang 0046, Chunhua Shen |
ICLR | 1 |
| 2022 | A Unified Mixture-View Framework for Unsupervised Representation Learning
Xiangxiang Chu, Xiaohang Zhan, Bo Zhang 0046 |
BMVC | 1 |
| 2022 | EAPruning: Evolutionary Pruning for Vision Transformers and CNNs
Qingyuan Li 0001, Bo Zhang 0046, Xiangxiang Chu |
BMVC | 3 |
| 2022 | Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationabstractText-based video segmentation aims to segment the target object in a video based on a describing sentence. Incorporating motion information from optical flow maps with appearance and linguistic modalities is crucial yet has been largely ignored by previous work. In this paper, we design a method to fuse and align appearance, motion, and linguistic features to achieve accurate segmentation. Specifically, we propose a multi-modal video transformer, which can fuse and aggregate multi-modal and temporal features between frames. Furthermore, we design a language-guided feature fusion module to progressively fuse appearance and motion features in each feature level with guidance from linguistic features. Finally, a multi-modal alignment loss is proposed to alleviate the semantic gap between features from different modalities. Extensive experiments on A2D Sentences and J-HMDB Sentences verify the performance and the generalization ability of our method compared to the state-of-the-art methods. Wangbo Zhao, Kai Wang 0036, Xiangxiang Chu, Fuzhao Xue, Xinchao Wang, Yang You 0001 |
CVPR | 3 |
| 2022 | PromptDet: Towards Open-Vocabulary Detection Using Uncurated Images
Chengjian Feng, Zequn Jie, Xiangxiang Chu, Haibing Ren, Xiaolin Wei, Weidi Xie, Lin Ma 0002 |
ECCV (9) | 4 |
| 2022 | SegViT: Semantic Segmentation with Plain Vision TransformersabstractWe explore the capability of plain Vision Transformers (ViTs) for semantic segmentation and propose the SegViT. Previous ViT-based segmentation networks usually learn a pixel-level representation from the output of the ViT. Differently, we make use of the fundamental component—attention mechanism, to generate masks for semantic segmentation. Specifically, we propose the Attention-to-Mask (ATM) module, in which the similarity maps between a set of learnable class tokens and the spatial feature maps are transferred to the segmentation masks. Experiments show that our proposed SegViT using the ATM module outperforms its counterparts using the plain ViT backbone on the ADE20K dataset and achieves new state-of-the-art performance on COCO-Stuff-10K and PASCAL-Context datasets. Furthermore, to reduce the computational cost of the ViT backbone, we propose query-based down-sampling (QD) and query-based up-sampling (QU) to build a Shrunk structure. With our Shrunk structure, the model can save up to 40% computations while maintaining competitive performance. Bowen Zhang 0009, Zhi Tian, Quan Tang 0001, Xiangxiang Chu, Xiaolin Wei, Chunhua Shen, Yifan Liu 0001 |
NeurIPS | 4 |
| 2022 | Fully Convolutional One-Stage 3D Object Detection on LiDAR Range ImagesabstractWe present a simple yet effective fully convolutional one-stage 3D object detector for LiDAR point clouds of autonomous driving scenes, termed FCOS-LiDAR. Unlike the dominant methods that use the bird-eye view (BEV), our proposed detector detects objects from the range view (RV, a.k.a. range image) of the LiDAR points. Due to the range view's compactness and compatibility with the LiDAR sensors' sampling process on self-driving cars, the range view-based object detector can be realized by solely exploiting the vanilla 2D convolutions, departing from the BEV-based methods which often involve complicated voxelization operations and sparse convolutions. For the first time, we show that an RV-based 3D detector with standard 2D convolutions alone can achieve comparable performance to state-of-the-art BEV-based detectors while being significantly faster and simpler. More importantly, almost all previous range view-based detectors only focus on single-frame point clouds since it is challenging to fuse multi-frame point clouds into a single range view. In this work, we tackle this challenging issue with a novel range view projection mechanism, and for the first time demonstrate the benefits of fusing multi-frame point clouds for a range-view based detector. Extensive experiments on nuScenes show the superiority of our proposed method and we believe that our work can be strong evidence that an RV-based 3D detector can compare favourably with the current mainstream BEV-based detectors. Code will be made publicly available. Zhi Tian, Xiangxiang Chu, Xiaolin Wei, Chunhua Shen |
NeurIPS | 2 |
| 2021 | Noisy Differentiable Architecture Search
Xiangxiang Chu, Bo Zhang 0047 |
BMVC | 1 |
| 2021 | AutoKWS: Keyword Spotting with Differentiable Architecture SearchabstractSmart audio devices are gated by an always-on lightweight keyword spotting program to reduce power consumption. It is however challenging to design models that have both high accuracy and low latency for accurate and fast responsiveness. Many efforts have been made to develop end-to-end neural networks, in which depthwise separable convolutions, temporal convolutions, and LSTMs are adopted as building units. Nonetheless, these networks designed with human expertise may not achieve an optimal trade-off in an expansive search space. In this paper, we propose to leverage recent advances in differentiable neural architecture search to discover more efficient networks. Our searched model attains 97.2% top-1 accuracy on Google Speech Command Dataset v1 with only nearly 100K parameters. Bo Zhang 0046, Qingyuan Li 0001, Weiji Zhuang, Xiangxiang Chu |
ICASSP | 5 |
| 2021 | FairNAS: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture SearchabstractOne of the most critical problems in weight-sharing neural architecture search is the evaluation of candidate models within a predefined search space. In practice, a one-shot supernet is trained to serve as an evaluator. A faithful ranking certainly leads to more accurate searching results. However, current methods are prone to making misjudgments. In this paper, we prove that their biased evaluation is due to inherent unfairness in the supernet training. In view of this, we propose two levels of constraints: expectation fairness and strict fairness. Particularly, strict fairness ensures equal optimization opportunities for all choice blocks throughout the training, which neither overestimates nor underestimates their capacity. We demonstrate that this is crucial for improving the confidence of models’ ranking. Incorporating the one-shot supernet trained under the proposed fairness constraints with a multi-objective evolutionary search algorithm, we obtain various state-of-the-art models, e.g., FairNAS-A attains 77.5% top-1 validation accuracy on ImageNet. Xiangxiang Chu, Bo Zhang 0046, Ruijun Xu |
ICCV | 1 |
| 2021 | DARTS-: Robustly Stepping out of Performance Collapse Without Indicators
Xiangxiang Chu, Xiaoxing Wang, Bo Zhang 0046, Shun Lu 0001, Xiaolin Wei, Junchi Yan |
ICLR | 1 |
| 2021 | Twins: Revisiting the Design of Spatial Attention in Vision TransformersabstractVery recently, a variety of vision transformer architectures for dense prediction tasks have been proposed and they show that the design of spatial attention is critical to their success in these tasks. In this work, we revisit the design of the spatial attention and demonstrate that a carefully devised yet simple spatial attention mechanism performs favorably against the state-of-the-art schemes. As a result, we propose two vision transformer architectures, namely, Twins- PCPVT and Twins-SVT. Our proposed architectures are highly efficient and easy to implement, only involving matrix multiplications that are highly optimized in modern deep learning frameworks. More importantly, the proposed architectures achieve excellent performance on a wide range of visual tasks including image-level classification as well as dense detection and segmentation. The simplicity and strong performance suggest that our proposed architectures may serve as stronger backbones for many vision tasks. Xiangxiang Chu, Zhi Tian, Bo Zhang 0046, Haibing Ren, Xiaolin Wei, Huaxia Xia, Chunhua Shen |
NeurIPS | 1 |
| 2020 | Accurate and Efficient Single Image Super-Resolution with Matrix Channel Attention Network
Xiangxiang Chu, Bo Zhang 0046 |
ACCV (2) | 2 |
| 2020 | Fair DARTS: Eliminating Unfair Advantages in Differentiable Architecture Search
Xiangxiang Chu, Tianbao Zhou, Bo Zhang 0046 |
ECCV (15) | 1 |
| 2020 | MoGA: Searching Beyond Mobilenetv3abstractThe evolution of MobileNets has laid a solid foundation for neural network applications on mobile end. With the latest MobileNetV3, neural architecture search again claimed its supremacy in network design. Unfortunately, till today all mobile methods mainly focus on CPU latencies instead of GPU, the latter, however, is much preferred in practice because it has faster speed, lower overhead and less interference. Bearing the target hardware in mind, we propose the first Mobile GPU-Aware (MoGA) neural architecture search in order to be precisely tailored for real-world applications. Further, the ultimate objective to devise a mobile network lies in achieving better performance by maximizing the utilization of bounded resources. Urging higher capability while restraining time consumption is not reconcilable. We alleviate this tension by weighted evolution techniques. Moreover, we encourage increasing the number of parameters for higher representational power. With 200× fewer GPU days than MnasNet, we obtain a series of models that outperform MobileNetV3 under the similar latency constraints1. Xiangxiang Chu, Bo Zhang 0046, Ruijun Xu |
ICASSP | 1 |
| 2020 | Fast, Accurate and Lightweight Super-Resolution with Neural Architecture SearchabstractDeep convolutional neural networks demonstrate impressive results in the super-resolution domain. A series of studies concentrate on improving peak signal noise ratio (PSNR) by using much deeper layers, which are not friendly to constrained resources. Pursuing a trade-off between the restoration capacity and the simplicity of models is still non-trivial. Recent contributions are struggling to manually maximize this balance, while our work achieves the same goal automatically with neural architecture search. Specifically, we handle super-resolution with a multi-objective approach. We also propose an elastic search tactic at both micro and macro level, based on a hybrid controller that profits from evolutionary computation and reinforcement learning. Quantitative experiments help us to draw a conclusion that our generated models dominate most of the state-of-the-art methods with respect to the individual FLOPS. Xiangxiang Chu, Bo Zhang 0046, Ruijun Xu, Qingyuan Li 0001 |
ICPR | 1 |
| 2020 | Neural Architecture Search on Acoustic Scene ClassificationabstractConvolutional neural networks are widely adopted in Acoustic Scene Classification (ASC) tasks, but they generally carry a heavy computational burden. In this work, we propose a lightweight yet high-performing baseline network inspired by MobileNetV2, which replaces square convolutional kernels with unidirectional ones to extract features alternately in temporal and frequency dimensions. Furthermore, we explore a dynamic architecture space built on the basis of the proposed baseline with the recent Neural Architecture Search (NAS) paradigm, which first trains a supernet that incorporates all candidate networks and then applies a well-known evolutionary algorithm NSGA-II to discover more efficient networks with higher accuracy and lower computational cost. Experimental results demonstrate that our searched network is competent in ASC tasks, which achieves 90.3% F1-score on the DCASE2018 task 5 evaluation set, marking a new state-of-the-art performance while saving 25% of FLOPs compared to our baseline network. Chuming Liang, Bo Zhang 0046, Xiangxiang Chu |
INTERSPEECH | 6 |