Chongjie Si

dblp:324/2201 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
10since 2021 · last 2026
0009-0001-3017-7574ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 36% Segmentation and scene understanding · 18% Learning paradigms · 11%
Databases, data mining, and information retrieval
1 paper
Data mining · 91% Machine learning and data management · 9%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
3.642026
Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning · ICLR 2025
Unleashing the Power of Task-Specific Directions in Parameter Efficient Fine-tuning · ICLR 2025
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
1.922026
Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Unleashing the Power of Task-Specific Directions in Parameter Efficient Fine-tuning · ICLR 2025
Computer vision › Segmentation and scene understanding
semantic segmentation
1.622025
OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information · NeurIPS 2025
Tendency-Driven Mutual Exclusivity for Weakly Supervised Incremental Semantic Segmentation · ECCV (35) 2024
Machine learning › Learning paradigms › weakly supervised learning
partial label learning
1.422024
Partial Label Learning with a Partner · AAAI 2024
Complementary Classifier Induced Partial Label Learning · KDD 2023
Machine learning › Efficient and distributed learning
model compression
1.012026
Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Vision and language › vision-language model
CLIP
0.912025
OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
foundation model adaptation
0.912025
Generalized Tensor-Based Parameter-Efficient Fine-Tuning via Lie Group Transformations · ICCV 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.912025
OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation › model adaptation
task adaptation
0.912025
Unleashing the Power of Task-Specific Directions in Parameter Efficient Fine-tuning · ICLR 2025
Computer vision › Vision and language
vision-language model
0.912025
OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information · NeurIPS 2025
Natural language and speech › Information extraction and text analysis
disambiguation
0.812024
Partial Label Learning with a Partner · AAAI 2024
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.812024
Partial Label Learning with a Partner · AAAI 2024
Computer vision › Segmentation and scene understanding
mutual exclusivity
0.812024
Tendency-Driven Mutual Exclusivity for Weakly Supervised Incremental Semantic Segmentation · ECCV (35) 2024
Data mining › predictive modeling
classification
0.812024
Multi-Label Classification With High-Rank and High-Order Label Correlations · IEEE Trans. Knowl. Data Eng. 2024
Data mining › predictive modeling › classification › multi-label classification
label correlation exploitation
0.812024
Multi-Label Classification With High-Rank and High-Order Label Correlations · IEEE Trans. Knowl. Data Eng. 2024
Data mining › predictive modeling › classification
multi-label classification
0.812024
Multi-Label Classification With High-Rank and High-Order Label Correlations · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Learning paradigms
weakly supervised learning
0.712023
Complementary Classifier Induced Partial Label Learning · KDD 2023
Machine learning › Trustworthy machine learning › interpretability
attention map
0.312025
OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.312025
OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

LoRA · 2.7initialization strategy · 1.0training-free refinement · 0.9tensor decomposition · 0.9semantic attention alignment · 0.9low-rank adaptation · 0.9lie group transformations · 0.9fine-tuning · 0.9exponential map · 0.9context-aware attention injection · 0.9matrix factorization · 0.8local geometric structure · 0.8
YearPublicationVenuePosition
2026 Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-Tuning
abstract
Large language models demonstrate impressive performance on downstream tasks, yet requiring extensive resource consumption when fully fine-tuning all parameters. To mitigate this, Parameter Efficient Fine-Tuning (PEFT) strategies, such as LoRA, have been developed. In this paper, we delve into the concept of task-specific directions (TSDs)-critical for transitioning large models from pretrained states to task-specific enhancements in PEFT. We propose a framework to clearly define these directions and explore their properties, and practical utilization challenges. We then introduce a novel approach, LoRA-Dash, which aims to maximize the impact of TSDs during the fine-tuning process, thereby enhancing model performance on targeted tasks. Additionally, based on our exploration of TSD, we focus on an important issue in PEFT: the initialization of LoRA. While some works have pointed out the significance of initialization for LoRA's performance and proposed various strategies, these methods are often empirical and not task-specific. To address this issue, we propose LoRA-Init. Starting from TSD, we identify the directions that require the most adjustment during fine-tuning for downstream tasks. By initializing the matrices in LoRA with these directions, LoRA-Init significantly enhances LoRA's performance. Moreover, we can combine LoRA-Dash and LoRA-Init to create the final version of LoRA based on TSDs, which we refer to as LoRA-TSD. Extensive experiments have conclusively demonstrated the effectiveness of these methods, and in-depth analyses further reveal the underlying mechanisms of these methods.
Chongjie Si, Zhiyi Shi, Shifan Zhang, Xiaokang Yang 0001, Hanspeter Pfister, Wei Shen 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Generalized Tensor-Based Parameter-Efficient Fine-Tuning via Lie Group Transformations
abstract
Adapting pre-trained foundation models for diverse downstream tasks is a core practice in artificial intelligence. However, the wide range of tasks and high computational costs make full fine-tuning impractical. To overcome this, parameter-efficient fine-tuning (PEFT) methods like LoRA have emerged and are becoming a growing research focus. Despite the success of these methods, they are primarily designed for linear layers, focusing on two-dimensional matrices while largely ignoring higher-dimensional parameter spaces like convolutional kernels. Moreover, directly applying these methods to higher-dimensional parameter spaces often disrupts their structural relationships. Given the rapid advancements in matrix-based PEFT methods, rather than designing a specialized strategy, we propose a generalization that extends matrix-based PEFT methods to higher-dimensional parameter spaces without compromising their structural properties. Specifically, we treat parameters as elements of a Lie group, with updates modeled as perturbations in the corresponding Lie algebra. These perturbations are mapped back to the Lie group through the exponential map, ensuring smooth, consistent updates that preserve the inherent structure of the parameter space. Extensive experiments on computer vision and natural language processing validate the effectiveness and versatility of our approach, demonstrating clear improvements over existing methods.
Chongjie Si, Zhiyi Shi, Yichen Xiao, Xiaokang Yang 0001, Wei Shen 0002
ICCV1
2025 Unleashing the Power of Task-Specific Directions in Parameter Efficient Fine-tuning
abstract
Large language models demonstrate impressive performance on downstream tasks, yet requiring extensive resource consumption when fully fine-tuning all parameters. To mitigate this, Parameter Efficient Fine-Tuning (PEFT) strategies, such as LoRA, have been developed. In this paper, we delve into the concept of task-specific directions (TSDs)—critical for transitioning large models from pretrained states to task-specific enhancements in PEFT. We propose a framework to clearly define these directions and explore their properties, and practical utilization challenges. We then introduce a novel approach, LoRA-Dash, which aims to maximize the impact of TSDs during the fine-tuning process, thereby enhancing model performance on targeted tasks. Extensive experiments have conclusively demonstrated the effectiveness of LoRA-Dash, and in-depth analyses further reveal the underlying mechanisms of LoRA-Dash.
Chongjie Si, Zhiyi Shi, Shifan Zhang, Xiaokang Yang 0001, Hanspeter Pfister, Wei Shen 0002
ICLR1
2025 Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning
abstract
Adapting pre-trained foundation models for various downstream tasks has been prevalent in artificial intelligence. Due to the vast number of tasks and high costs, adjusting all parameters becomes unfeasible. To mitigate this, several fine-tuning techniques have been developed to update the pre-trained model weights in a more resource-efficient manner, such as through low-rank adjustments. Yet, almost all of these methods focus on linear weights, neglecting the intricacies of parameter spaces in higher dimensions like 4D. Alternatively, some methods can be adapted for high-dimensional parameter space by compressing changes in the original space into two dimensions and then employing low-rank matrix adaptations. However, these approaches destructs the structural integrity of the involved high-dimensional spaces. To tackle the diversity of dimensional spaces across different foundation models and provide a more precise representation of the changes within these spaces, this paper introduces a generalized parameter-efficient fine-tuning framework, designed for various dimensional parameter space. Specifically, our method asserts that changes in each dimensional parameter space are based on a low-rank core space which maintains the consistent topological structure with the original space. It then models the changes through this core space alongside corresponding weights to reconstruct alterations in the original space. It effectively preserves the structural integrity of the change of original N-dimensional parameter space, meanwhile models it via low-rank tensor adaptation. Extensive experiments on computer vision, natural language processing and multi-modal tasks validate the effectiveness of our method.
Chongjie Si, Xue Yang 0005, Zhengqin Xu, Qingyun Li, Jifeng Dai, Yu Qiao 0001, Xiaokang Yang 0001, Wei Shen 0002
ICLR1
2025 Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
abstract
This paper presents a pioneering exploration of reinforcement learning (RL) via group relative policy optimization for unified multimodal large language models (ULMs), aimed at simultaneously reinforcing generation and understanding capabilities. Through systematic pilot studies, we uncover the significant potential of ULMs to enable the synergistic co-evolution of dual capabilities within a shared policy optimization framework. Building on this insight, we introduce \textbf{CoRL}, a \textbf{Co}-\textbf{R}einforcement \textbf{L}earning framework comprising a unified RL stage for joint optimization and a refined RL stage for task-specific enhancement. With the proposed CoRL, our resulting model, \textbf{ULM-R1}, achieves average improvements of 7\% on three text-to-image generation datasets and 23\% on nine multimodal understanding benchmarks. These results demonstrate the effectiveness of CoRL and highlight the substantial benefits of reinforcement learning in facilitating cross-task synergy and optimization for ULMs. Code is available at \url{https://github.com/mm-vl/ULM-R1}.
Jingjing Jiang, Chongjie Si, Jun Luo 0001, Hanwang Zhang
NeurIPS2
2025 OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information
abstract
Open-vocabulary semantic segmentation assigns every pixel a label drawn from an open-ended, text-defined space. Vision–language models such as CLIP excel at zero-shot recognition, yet their image-level pre-training hinders dense prediction. Current approaches either fine-tune CLIP—at high computational cost—or adopt training-free attention refinements that favor local smoothness while overlooking global semantics. In this paper, we present OPMapper, a lightweight, plug-and-play module that injects both local compactness and global connectivity into attention maps of CLIP. It combines Context-aware Attention Injection, which embeds spatial and semantic correlations, and Semantic Attention Alignment, which iteratively aligns the enriched weights with textual prompts. By jointly modeling token dependencies and leveraging textual guidance, OPMapper enhances visual understanding. OPMapper is highly flexible and can be seamlessly integrated into both training-based and training-free paradigms with minimal computational overhead. Extensive experiments demonstrate its effectiveness, yielding significant improvements across 8 open-vocabulary segmentation benchmarks.
Chongjie Si, Xue Yang 0005, Yuzhi Zhao, Wenhai Wang, Xiaokang Yang 0001, Wei Shen 0002
NeurIPS2
2024 Partial Label Learning with a Partner
abstract
In partial label learning (PLL), each instance is associated with a set of candidate labels among which only one is ground-truth. The majority of the existing works focuses on constructing robust classifiers to estimate the labeling confidence of candidate labels in order to identify the correct one. However, these methods usually struggle to rectify mislabeled samples. To help existing PLL methods identify and rectify mislabeled samples, in this paper, we introduce a novel partner classifier and propose a novel ``mutual supervision'' paradigm. Specifically, we instantiate the partner classifier predicated on the implicit fact that non-candidate labels of a sample should not be assigned to it, which is inherently accurate and has not been fully investigated in PLL. Furthermore, a novel collaborative term is formulated to link the base classifier and the partner one. During each stage of mutual supervision, both classifiers will blur each other's predictions through a blurring mechanism to prevent overconfidence in a specific label. Extensive experiments demonstrate that the performance and disambiguation ability of several well-established stand-alone and deep-learning based PLL approaches can be significantly improved by coupling with this learning paradigm.
Chongjie Si, Zekun Jiang, Yan Wang 0033, Xiaokang Yang 0001, Wei Shen 0002
AAAI1
2024 Tendency-Driven Mutual Exclusivity for Weakly Supervised Incremental Semantic Segmentation
Chongjie Si, Xiaokang Yang 0001, Wei Shen 0002
ECCV (35)1
2024 Multi-Label Classification With High-Rank and High-Order Label Correlations
abstract
Exploiting label correlations is important to multi-label classification. Previous methods capture the high-order label correlations mainly by transforming the label matrix to a latent label space with low-rank matrix factorization. However, the label matrix is generally a full-rank or approximate full-rank matrix, making the low-rank factorization inappropriate. Besides, in the latent space, the label correlations will become implicit. To this end, we propose a simple yet effective method to depict the high-order label correlations explicitly, and at the same time maintain the high-rank of the label matrix. Moreover, we estimate the label correlations and infer model parameters simultaneously via the local geometric structure of the input to achieve mutual enhancement. Comparative studies over twelve benchmark data sets validate the effectiveness of the proposed algorithm in multi-label classification. The exploited high-order label correlations are consistent with common sense empirically.Our code is publicly available athttps://github.com/Chongjie-Si/HOMI.
Chongjie Si, Yuheng Jia, Ran Wang 0001, Min-Ling Zhang, Yang-He Feng, Chongxiao Qu
IEEE Trans. Knowl. Data Eng.1
2023 Complementary Classifier Induced Partial Label Learning
abstract
In partial label learning (PLL), each training sample is associated with a set of candidate labels, among which only one is valid. The core of PLL is to disambiguate the candidate labels to get the ground-truth one. In disambiguation, the existing works usually do not fully investigate the effectiveness of the non-candidate label set (a.k.a. complementary labels), which accurately indicates a set of labels that do not belong to a sample. In this paper, we use the non-candidate labels to induce a complementary classifier, which naturally forms an adversarial relationship against the traditional PLL classifier, to eliminate the false-positive labels in the candidate label set. Besides, we assume the feature space and the label space share the same local topological structure captured by a dynamic graph, and use it to assist disambiguation. Extensive experimental results validate the superiority of the proposed approach against state-of-the-art PLL methods on 4 controlled UCI data sets and 6 real-world data sets and reveal the usefulness of complementary learning in PLL. The code has been released in the link https://github.com/Chongjie-Si/PL-CL
Yuheng Jia, Chongjie Si, Min-Ling Zhang
KDD2