Lin Zhang 0055

dblp:37/1629-55 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0006-5293-9169ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 30% Deep learning architectures and training · 16% 3D vision · 15%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.622025
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models · CVPR 2025
Lightweight Model Pre-Training via Language Guided Knowledge Distillation · IEEE Trans. Multim. 2024
Machine learning › Deep learning architectures and training
feedforward neural network
0.912025
PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding · NeurIPS 2025
Computer vision › Vision and language
image captioning
0.912025
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning · ICCV 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context understanding
0.912025
PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › parameter compression
mixture-of-experts compression
0.912025
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models · CVPR 2025
Computer vision › Video understanding and tracking › motion analysis
motion understanding
0.912025
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding · NeurIPS 2025
Performance modeling and evaluation
benchmarking
0.912025
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding · NeurIPS 2025
Computer vision › 3D vision
3d object detection
0.812024
3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection · NeurIPS 2024
Computer vision › 3D vision › 3d object detection
indoor 3d object detection
0.812024
3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.812024
Lightweight Model Pre-Training via Language Guided Knowledge Distillation · IEEE Trans. Multim. 2024
Machine learning › Deep learning architectures and training
state space model
0.812024
3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection · NeurIPS 2024
Machine learning › Deep learning architectures and training
mixture of experts
0.312025
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models · CVPR 2025
Machine learning › Reinforcement learning
reward design
0.312025
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning · ICCV 2025
Computer vision › 3D vision
point cloud processing
0.212024
3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-supervised distillation
0.212024
Lightweight Model Pre-Training via Language Guided Knowledge Distillation · IEEE Trans. Multim. 2024

Methods — techniques the papers use, named apart from their topics

LLM-assisted evaluation · 1.7sparsification · 0.9scene graph parsing · 0.9reinforcement learning · 0.9quantization · 0.9persistent activity mechanism · 0.9multiple-choice question answering · 0.9low-rank approximation · 0.9expert clustering · 0.9direct preference optimization · 0.9decomposition · 0.9
YearPublicationVenuePosition
2026 Adapter-X: A general parameter-efficient fine-tuning framework for 2D and 3D vision
Peng Ye 0006, Lin Zhang 0055, Bizhe Bai, Tao Chen 0003
Neurocomputing3
2025 DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
abstract
Upcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into MoE layers. However, these models still suffer from significant parameter inefficiency due to the introduction of multiple experts. In this work, we propose a novel DeRS (Decompose, Replace, and Synthesis) paradigm to overcome this shortcoming, which is motivated by our observations about the unique redundancy mechanisms of upcycled MoE experts. Specifically, DeRS decomposes the experts into one expert-shared base weight and multiple expert-specific delta weights, and subsequently represents these delta weights in lightweight forms. Our proposed DeRS paradigm can be applied to enhance parameter efficiency in two different scenarios, including: 1) DeRS Compression for inference stage, using sparsification or quantization to compress vanilla upcycled MoE models; and 2) DeRS Up-cycling for training stage, employing lightweight sparse or low-rank matrixes to efficiently upcycle dense models into MoE models. Extensive experiments across three different tasks show that the proposed methods can achieve extreme parameter efficiency while maintaining the performance for both training and compression of upcycled MoE models.
Yongqi Huang, Peng Ye 0006, Chenyu Huang 0001, Jianjian Cao, Lin Zhang 0055, Baopu Li, Gang Yu 0002, Tao Chen 0003
CVPR5
2025 SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
abstract
We propose SC-Captioner, a reinforcement learning framework that enables the self-correcting capability of image caption models. Our crucial technique lies in the design of the reward function to incentivize accurate caption corrections. Specifically, the predicted and reference captions are decomposed into object, attribute, and relation sets using scene-graph parsing algorithms. We calculate the set difference between sets of initial and self-corrected captions to identify added and removed elements. These elements are matched against the reference sets to calculate correctness bonuses for accurate refinements and mistake punishments for wrong additions and removals, thereby forming the final reward. For image caption quality assessment, we propose a set of metrics refined from CAPTURE that alleviate its incomplete precision evaluation and inefficient relation matching problems. Furthermore, we collect a fine-grained annotated image caption dataset, RefinedCaps, consisting of 6.5K diverse images from COCO dataset. Experiments show that applying SC-Captioner on large visual-language models can generate better image captions across various scenarios, significantly outperforming the direct preference optimization training strategy.
Lin Zhang 0055, Xianfang Zeng, Kangcong Li, Gang Yu 0002, Tao Chen 0003
ICCV1
2025 PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding
abstract
While Large Language Models (LLMs) demonstrate strong performance across domains, their long-context capabilities are limited by transient neural activations causing information decay and unstructured feed-forward network (FFN) weights leading to semantic fragmentation. Inspired by the brain’s working memory and cortical modularity, we propose PaceLLM, featuring two innovations: (1) a Persistent Activity (PA) Mechanism that mimics prefrontal cortex (PFC) neurons’ persistent firing by introducing an activation-level memory bank to dynamically retrieve, reuse, and update critical FFN states, addressing contextual decay; and (2) Cortical Expert (CE) Clustering that emulates task-adaptive neural specialization to reorganize FFN weights into semantic modules, establishing cross-token dependencies and mitigating fragmentation. Extensive evaluations show that PaceLLM achieves 6% improvement on LongBench’s Multi-document QA and 12.5–17.5% performance gains on $\infty$-Bench tasks, while extending measurable context length to 200K tokens in Needle-In-A-Haystack (NIAH) tests. This work pioneers brain-inspired LLM optimization and is complementary to other works. Besides, it can be generalized to any model and enhance their long-context performance and interpretability without structural overhauls.
Kangcong Li, Peng Ye 0006, Chongjun Tu, Lin Zhang 0055, Chunfeng Song, Qihao Zheng, Tao Chen 0003
NeurIPS4
2025 FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
abstract
Multimodal Large Language Models (MLLMs) have shown impressive video content understanding capabilities but struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding ability of existing MLLMs, we introduce FAVOR-Bench, which comprises 1,776 videos from both ego-centric and third-person perspectives and enables assessment through both close-ended and open-ended tasks. For close-ended evaluation, we carefully design 8,184 multiple-choice question-answer pairs spanning six distinct sub-tasks. For open-ended evaluation, we employ the GPT-assisted evaluation and develop a novel cost-efficient LLM-free assessment method, where the latter can enhance benchmarking interpretability and accessibility. Comprehensive experiments with21 state-of-the-art MLLMs reveal significant limitations in their ability to comprehend and describe detailed temporal dynamics in video motions. To alleviate this limitation, we further build FAVOR-Train, a dataset of 17,152 videos with fine-grained motion annotations. Finetuning Qwen2.5-VL on FAVOR-Train yields consistent improvements on motion-related tasks across TVBench, MotionBenchand our FAVOR-Bench. Our assessment results demonstrate that the proposed FAVOR-Bench and FAVOR-Train provide valuable tools for the community to develop more powerful video understanding models.
Chongjun Tu, Lin Zhang 0055, Pengtao Chen, Peng Ye 0006, Xianfang Zeng, Gang Yu 0002, Tao Chen 0003
NeurIPS2
2025 Learnable Bi-directional Data Augmentation for few-shot cross-domain point cloud classification
Lin Zhang 0055, Jiakang Yuan, Tao Chen 0003
Neurocomputing2
2024 3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection
abstract
Transformer-based architectures have been proven successful in detecting 3D objects from point clouds. However, the quadratic complexity of the attention mechanism struggles to encode rich information as point cloud resolution increases. Recently, state space models (SSM) such as Mamba have gained great attention due to their linear complexity and long sequence modeling ability for language understanding. To exploit the potential of Mamba on 3D scene-level perception, for the first time, we propose 3DET-Mamba, which is a novel SSM-based model designed for indoor 3d object detection. Specifically, we divide the point cloud into different patches and use a lightweight yet effective Inner Mamba to capture local geometric information. To observe the scene from a global perspective, we introduce a novel Dual Mamba module that models the point cloud in terms of spatial distribution and continuity. Additionally, we design a Query-aware Mamba module that decodes context features into object sets under the guidance of learnable queries. Extensive experiments demonstrate that 3DET-Mamba surpasses previous 3DETR on indoor 3D detection benchmarks such as ScanNet, improving AP25/AP50 from 65.0\%/47.0\% to 70.4\%/54.4\%, respectively.
Mingsheng Li, Jiakang Yuan, Sijin Chen, Lin Zhang 0055, Anyu Zhu, Tao Chen 0003
NeurIPS4
2024 Push-and-Pull: A General Training Framework With Differential Augmentor for Domain Generalized Point Cloud Classification
abstract
As a fundamental task of 3D perception, point cloud recognition has shown significant progress in recent years. However, existing methods still face challenges when dealing with geometry differences, resulting in performance degradation when a distribution gap exists between the training and testing data, also known as domain generalization. In this work, we focus on this problem and propose a general training framework, named Push-and-Pull, aimed at effectively improving the generalization ability of models on unseen target domains. Specifically, our framework first introduces a learnable 3D data augmentor to generate new training point clouds, which helps to reduce the domain bias and enrich the source training set. Also, an adversarial training strategy is proposed topushthe augmented samples away from the original ones in the latent space and meanwhile keep the geometric structure. Second, based on the original and augmented samples, a dual-level consistency regularization strategy on logits and feature spaces is designed topullthe deviated representations back to their original space as close as possible, and promote discriminative and domain-agnostic representations. These two steps are iteratively optimized to enhance the overall performance. Extensive experiments on the PointDA-10 and Sim2Real benchmarks consistently demonstrate the effectiveness of our proposed framework.
Xinzhu Ma, Lin Zhang 0055, Bo Zhang 0069, Tao Chen 0003
IEEE Trans. Circuits Syst. Video Technol.3
2024 Few-Shot Cross-Domain Object Detection With Instance-Level Prototype-Based Meta-Learning
abstract
In typical unsupervised domain adaptive object detection, it is assumed that extensive unlabeled training data from the target domain can be easily obtained. However, in some access-constrained scenarios, massive target data cannot be guaranteed, but acquiring only a few target samples and annotating them may costs less. Therefore, inspired by the meta-learning success in few-shot tasks, we propose an Instance-level Prototype learning Network (IPNet) for solving the domain adaptive object detection under the supervised few-shot scenario in this work. To compensate for the target domain data deficiency, we fuse cropped instances from labeled images in both domains to learn a representative prototype for each class, by enforcing features of the same class’s instances but from different domains to be as close as possible. These prototypes are further employed to discriminate various features’ salience in an image, and separate foreground and background regions for respective domain alignment. Extensive experiments are conducted on several cross-domain scenarios, and their results show the consistent accuracy gains of the IPNet over state-of-the-art methods, e.g., 10.4% mAP increase on Cityscapes-to-FoggyCityscapes setting and 3.0% mAP increase on Sim10k-to-Cityscapes setting.
Lin Zhang 0055, Bo Zhang 0069, Botian Shi, Jiayuan Fan 0001, Tao Chen 0003
IEEE Trans. Circuits Syst. Video Technol.1
2024 Lightweight Model Pre-Training via Language Guided Knowledge Distillation
abstract
This paper studies the problem of pre-training for small models, which is essential for many mobile devices. Current state-of-the-art methods on this problem transfer the representational knowledge of a large network (as a Teacher) into a smaller model (as a Student) using self-supervised distillation, improving the performance of the small model on downstream tasks. However, existing approaches are insufficient in extracting the crucial knowledge that is useful for discerning categories in downstream tasks during the distillation process. In this paper, for the first time, we introduce language guidance to the distillation process and propose a new method named Language-Guided Distillation (LGD) system, which uses category names of the target downstream task to help refine the knowledge transferred between the teacher and student. To this end, we utilize a pre-trained text encoder to extract semantic embeddings from language and construct a textual semantic space called Textual Semantics Bank (TSB). Furthermore, we design a Language-Guided Knowledge Aggregation (LGKA) module to construct the visual semantic space, also named Visual Semantics Bank (VSB). The task-related knowledge is transferred by driving a student encoder to mimic the similarity score distribution inferred by a teacher over TSB and VSB. Compared with other small models obtained by either ImageNet pre-training or self-supervised distillation, experiment results show that the distilled lightweight model using the proposed LGD method presents state-of-the-art performance and is validated on various downstream tasks, including classification, detection, and segmentation.
Mingsheng Li, Lin Zhang 0055, Mingzhen Zhu, Gang Yu 0002, Jiayuan Fan 0001, Tao Chen 0003
IEEE Trans. Multim.2