Taolin Zhang 0003

dblp:270/2482-3 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
0009-0006-2441-2861ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 32% Vision and language · 19% Transfer learning and domain adaptation · 15%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 33% Program synthesis and code generation · 33% Software testing · 33%
Computer graphics and multimedia
2 papers
Image and video processing · 67% Multimedia analysis and retrieval · 33%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.912025
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement · ACL (1) 2025
Machine learning › Generative modeling › autoregressive model
autoregressive image generation
0.912025
FastVAR: Linear Visual Autoregressive Modeling Via Cached Token Pruning · ICCV 2025
Machine learning › Efficient and distributed learning
inference acceleration
0.912025
FastVAR: Linear Visual Autoregressive Modeling Via Cached Token Pruning · ICCV 2025
Machine learning › Efficient and distributed learning › model compression
token pruning
0.912025
FastVAR: Linear Visual Autoregressive Modeling Via Cached Token Pruning · ICCV 2025
Compilers and program optimization
code generation
0.912025
Rethinking Verification for LLM Code Generation: From Generation to Testing · NeurIPS 2025
Program synthesis and code generation
code generation evaluation
0.912025
Rethinking Verification for LLM Code Generation: From Generation to Testing · NeurIPS 2025
Software testing
test generation
0.912025
Rethinking Verification for LLM Code Generation: From Generation to Testing · NeurIPS 2025
Computer vision › 3D vision
3d scene understanding
0.812024
Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding · AAAI 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding · AAAI 2024
Machine learning › Efficient and distributed learning › memory-efficient training
memory-efficient fine-tuning
0.812024
Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach · ECCV (45) 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach · ECCV (45) 2024
Machine learning › Transfer learning and domain adaptation
test-time adaptation
0.812024
BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping · NeurIPS 2024
Computer vision › Vision and language › vision-language model
vision-language model adaptation
0.812024
BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping · NeurIPS 2024
Computer vision › Vision and language
vision-language pretraining
0.812024
Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding · AAAI 2024
Machine learning › Transfer learning and domain adaptation › model adaptation
vision transformer adaptation
0.812024
Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach · ECCV (45) 2024
Information retrieval
ranking
0.812024
Multimodal Label Relevance Ranking via Reinforcement Learning · ECCV (66) 2024
Information retrieval
retrieval augmentation
0.812024
ReFIR: Grounding Large Restoration Models with Retrieval Augmentation · NeurIPS 2024
Image and video processing › image restoration › deep image restoration
diffusion-based image restoration
0.812024
ReFIR: Grounding Large Restoration Models with Retrieval Augmentation · NeurIPS 2024
Image and video processing
image restoration
0.812024
ReFIR: Grounding Large Restoration Models with Retrieval Augmentation · NeurIPS 2024
Natural language and speech › Language models and text generation
instruction tuning
0.312025
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement · ACL (1) 2025
Computer vision › Vision and language › visual grounding
object grounding
0.212024
Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding · AAAI 2024
Computer vision › Vision and language
visual grounding
0.212024
Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding · AAAI 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.5nearest neighbor retrieval · 1.5diffusion model · 1.5cross-image injection · 1.5reinforcement learning with verifiable rewards · 0.9large language model · 0.9knowledge-driven data synthesis · 0.9flashattention · 0.9data refinement · 0.9cached token pruning · 0.9vision-language pretraining · 0.8regional bootstrapping · 0.8key-value memory · 0.8contrastive learning · 0.8
YearPublicationVenuePosition
2025 Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
abstract
Maosongcao Maosongcao, Taolin Zhang, Mo Li, Chuyu Zhang, Yunxin Liu, Conghui He, Haodong Duan, Songyang Zhang, Kai Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Maosongcao, Taolin Zhang 0003, Mo Li 0012, Chuyu Zhang, Conghui He, Haodong Duan, Songyang Zhang 0001, Kai Chen 0026
ACL (1)2
2025 FastVAR: Linear Visual Autoregressive Modeling Via Cached Token Pruning
abstract
Visual Autoregressive (VAR) modeling has gained popularity for its shift towards next-scale prediction. However, existing VAR paradigms process the entire token map at each scale step, leading to the complexity and runtime scaling dramatically with image resolution. To address this challenge, we propose FastVAR, a post-training acceleration method for efficient resolution scaling with VARs. Our key finding is that the majority of latency arises from the large-scale step where most tokens have already converged. Leveraging this observation, we develop the cached token pruning strategy that only forwards pivotal tokens for scale-specific modeling while using cached tokens from previous scale steps to restore the pruned slots. This significantly reduces the number of forwarded tokens and improves the efficiency at larger resolutions. Experiments show the proposed FastVAR can further speedup FlashAttention-accelerated VAR by 2.7$\times$ with negligible performance drop of <1%. We further extend FastVAR to zero-shot generation of higher resolution images. In particular, FastVAR can generate one 2K image with 15GB memory footprints in 1.5s on a single NVIDIA 3090 GPU. Code is available at https://github.com/csguoh/FastVAR.
Hang Guo 0002, Yawei Li 0001, Taolin Zhang 0003, Jiangshan Wang, Tao Dai 0001, Shutao Xia, Luca Benini
ICCV3
2025 Rethinking Verification for LLM Code Generation: From Generation to Testing
abstract
Large language models (LLMs) have recently achieved notable success in code‑generation benchmarks such as HumanEval and LiveCodeBench. However, a detailed examination reveals that these evaluation suites often comprise only a limited number of homogeneous test cases, resulting in subtle faults going undetected. This not only artificially inflates measured performance but also compromises accurate reward estimation in reinforcement learning frameworks utilizing verifiable rewards (RLVR). To address these critical shortcomings, we systematically investigate the test-case generation (TCG) task by proposing multi-dimensional metrics designed to rigorously quantify test-suite thoroughness. Furthermore, we introduce a human-LLM collaborative method (SAGA), leveraging human programming expertise with LLM reasoning capability, aimed at significantly enhancing both the coverage and the quality of generated test cases. In addition, we develop a TCGBench to facilitate the study of the TCG task. Experiments show that SAGA achieves a detection rate of 90.62\% and a verifier accuracy of 32.58\% on TCGBench. The Verifier Accuracy (Verifier Acc) of the code generation evaluation benchmark synthesized by SAGA is 10.78\% higher than that of LiveCodeBench-v6. These results demonstrate the effectiveness of our proposed method. We hope this work contributes to building a scalable foundation for reliable LLM code evaluation, further advancing RLVR in code generation, and paving the way for automated adversarial test synthesis and adaptive benchmark integration.
Zihan Ma 0010, Taolin Zhang 0003, Maosongcao, Minnan Luo, Songyang Zhang 0001, Kai Chen 0026
NeurIPS2
2024 Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding
abstract
In recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to point cloud data, existing works mainly focus on building task-specific models, and fail to extract universal 3D vision-language embedding that generalize well. We carefully investigate three common tasks in semantic 3D scene understanding, and derive key insights into the development of a pre-training model. Motivated by these observations, we propose a vision-language pre-training framework 3DVLP (3D vision-language pre-training with object contrastive learning), which transfers flexibly on 3D vision-language downstream tasks. 3DVLP takes visual grounding as the proxy task and introduces Object-level IoU-guided Detection (OID) loss to obtain high-quality proposals in the scene. Moreover, we design Object-level Cross-Contrastive alignment (OCC) task and Object-level Self-Contrastive learning (OSC) task to align the objects with descriptions and distinguish different objects in the scene, respectively. Extensive experiments verify the excellent performance of 3DVLP on three 3D vision-language tasks, reflecting its superiority in semantic 3D scene understanding. Code is available at https://github.com/iridescentttt/3DVLP.
Taolin Zhang 0003, Sunan He, Tao Dai 0001, Zhi Wang 0001, Bin Chen 0011, Shutao Xia
AAAI1
2024 Multimodal Label Relevance Ranking via Reinforcement Learning
Taian Guo, Taolin Zhang 0003, Haoqian Wu, Hanjun Li 0002, Ruizhi Qiao, Xing Sun 0001
ECCV (66)2
2024 Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach
Taolin Zhang 0003, Jiawang Bai, Zhihe Lu, Dongze Lian, Genping Wang, Xinchao Wang, Shutao Xia
ECCV (45)1
2024 BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping
abstract
Adaptation of pretrained vision-language models such as CLIP to various downstream tasks have raised great interest in recent researches. Previous works have proposed a variety of test-time adaptation (TTA) methods to achieve strong generalization without any knowledge of the target domain. However, existing training-required TTA approaches like TPT necessitate entropy minimization that involves large computational overhead, while training-free methods like TDA overlook the potential for information mining from the test samples themselves. In this paper, we break down the design of existing popular training-required and training-free TTA methods and bridge the gap between them within our framework. Specifically, we maintain a light-weight key-value memory for feature retrieval from instance-agnostic historical samples and instance-aware boosting samples. The historical samples are filtered from the testing data stream and serve to extract useful information from the target distribution, while the boosting samples are drawn from regional bootstrapping and capture the knowledge of the test sample itself. We theoretically justify the rationality behind our method and empirically verify its effectiveness on both the out-of-distribution and the cross-domain datasets, showcasing its applicability in real-world situations.
Taolin Zhang 0003, Jinpeng Wang 0002, Hang Guo 0002, Tao Dai 0001, Bin Chen 0011, Shutao Xia
NeurIPS1
2024 ReFIR: Grounding Large Restoration Models with Retrieval Augmentation
abstract
Recent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from the hallucination dilemma, i.e., producing incorrect contents or textures when dealing with severe degradations, due to their heavy reliance on limited internal knowledge. In this paper, we propose an orthogonal solution called the Retrieval-augmented Framework for Image Restoration (ReFIR), which incorporates retrieved images as external knowledge to extend the knowledge boundary of existing LRMs in generating details faithful to the original scene. Specifically, we first introduce the nearest neighbor lookup to retrieve content-relevant high-quality images as reference, after which we propose the cross-image injection to modify existing LRMs to utilize high-quality textures from retrieved images. Thanks to the additional external knowledge, our ReFIR can well handle the hallucination challenge and facilitate faithfully results. Extensive experiments demonstrate that ReFIR can achieve not only high-fidelity but also realistic restoration results. Importantly, our ReFIR requires no training and is adaptable to various LRMs.
Hang Guo 0002, Tao Dai 0001, Zhihao Ouyang, Taolin Zhang 0003, Yaohua Zha, Bin Chen 0011, Shutao Xia
NeurIPS4
2024 FedEgo: Privacy-preserving Personalized Federated Graph Learning with Ego-graphs
abstract
As special information carriers containing both structure and feature information, graphs are widely used in graph mining, e.g., Graph Neural Networks (GNNs). However, graph data are stored separately in multiple distributed parties in some practical scenarios, which may not be directly shared due to conflicts of interest. Hence, federated graph neural networks are proposed to address such data silo issues while preserving each party’s privacy (or client). Nevertheless, different graph data distributions of various parties, which is known as the statistical heterogeneity, may degrade the performance of naive federated learning algorithms like FedAvg. In this article, we propose FedEgo, a federated graph learning framework based on ego-graphs to tackle the challenges above, in which each client will train their local models while also contributing to the training of a global model. FedEgo applies GraphSAGE over ego-graphs to make full use of the structure information and utilizes Mixup for privacy concerns. To deal with the statistical heterogeneity, we integrate personalization into learning and propose an adaptive mixing coefficient strategy that enables clients to achieve their optimal personalization. Extensive experimental results and in-depth analysis demonstrate the effectiveness of FedEgo.
Taolin Zhang 0003, Chengyuan Mai, Yaomin Chang, Chuan Chen 0001, Zibin Zheng
ACM Trans. Knowl. Discov. Data1