Shuhao Chen

dblp:43/2127 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
0009-0002-0410-5961ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Graph learning · 21% Efficient and distributed learning · 21% Language models and text generation · 20%
Software engineering, system software, and programming languages
1 paper
Program analysis · 100%

Topics — the 14 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model merging
adapter merging
0.912025
Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation · ICML 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
SPE Attention: Making Attention Equivariant to Semantic-Preserving Permutation for Code Processing · EMNLP 2025
Natural language and speech › Language models and text generation
code analysis
0.912025
SPE Attention: Making Attention Equivariant to Semantic-Preserving Permutation for Code Processing · EMNLP 2025
Natural language and speech › Language models and text generation › text summarization › domain-specific summarization
code summarization
0.912025
SPE Attention: Making Attention Equivariant to Semantic-Preserving Permutation for Code Processing · EMNLP 2025
Machine learning › Graph learning
link prediction
0.912025
Open Your Eyes: Vision Enhances Message Passing Neural Networks in Link Prediction · ICML 2025
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.912025
Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation · ICML 2025
Machine learning › Graph learning › graph neural network
message passing
0.912025
Open Your Eyes: Vision Enhances Message Passing Neural Networks in Link Prediction · ICML 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation · ICML 2025
Program analysis
error detection
0.912025
SPE Attention: Making Attention Equivariant to Semantic-Preserving Permutation for Code Processing · EMNLP 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models · NeurIPS 2024
Machine learning › Learning theory
empirical risk minimization
0.812024
Rethinking Guidance Information to Utilize Unlabeled Samples: A Label Encoding Perspective · ICML 2024
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › semantic embedding
label embedding
0.812024
Rethinking Guidance Information to Utilize Unlabeled Samples: A Label Encoding Perspective · ICML 2024
Machine learning › Learning theory
model selection
0.812024
RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models · NeurIPS 2024
Machine learning › Learning paradigms
semi-supervised learning
0.812024
Rethinking Guidance Information to Utilize Unlabeled Samples: A Label Encoding Perspective · ICML 2024

Methods — techniques the papers use, named apart from their topics

symmetry mask · 1.7directed layered graph · 1.7visual structural awareness · 0.9stochastic adapter deactivation · 0.9progressive training · 0.9message passing · 0.9graph neural network · 0.9cooperative game theory · 0.9entropy minimization · 0.8dual contrastive learning · 0.8
YearPublicationVenuePosition
2025 Enhancing Local Search for MaxSAT with Deep Differentiation Clause Weighting
abstract
Partial Maximum Satisfiability (PMS) and Weighted Partial Maximum Satisfiability (WPMS) generalize Maximum Satisfiability (MaxSAT), with broad real-world applications. Recent advances in Stochastic Local Search (SLS) algorithms for solving (W)PMS have mainly focused on designing clause weighting schemes. However, existing methods often fail to adequately distinguish between PMS and WPMS, typically employing uniform update strategies for clause weights and overlooking critical structural differences between the two problem types. In this work, we present a novel clause weighting scheme that, for the first time, updates the clause weights of PMS and WPMS instances according to distinct conditions. This scheme also introduces a new initialization method, which better accommodates the unique characteristics of both instance types. Furthermore, we propose a decimation method that prioritizes satisfying unit and hard clauses, effectively complementing our proposed clause weighting scheme. Building on these methods, we develop a new SLS solver for (W)PMS named DeepDist. Experimental results on benchmarks from the anytime tracks of recent MaxSAT Evaluations show that DeepDist outperforms state-of-the-art SLS solvers. Notably, a hybrid solver combining DeepDist with TT-Open-WBO-Inc surpasses the performance of the MaxSAT Evaluation 2024 winners, SPB-MaxSAT-c-Band and SPB-MaxSAT-c-FPS, highlighting the effectiveness of our approach. The code is available at https://github.com/jmhmaxsat/DeepDist
Menghua Jiang 0001, Haokai Gao, Shuhao Chen
ECAI3
2025 SPE Attention: Making Attention Equivariant to Semantic-Preserving Permutation for Code Processing
abstract
Codes serve as the fundamental language for human to communicate with machines, and various Transformer-based models are trained to process codes in recent advancements.A unique symmetry of code is its semanticpreserving permutation, which allows certain lines to be rearranged without altering the overall meaning.To capture such symmetry, we propose a novel attention mechanism that incorporates semantic-preserving permutation equivariance, called the SPE attention.By leveraging the symmetry relationships within code, we introduce a directed layered graph to represent the code structure, and this graph is then summarized into a symmetry mask.The SPE attention integrates those symmetry masks, granting semantic-preserving permutations equivariance to the model.Experiments on various code related tasks, including code summarization and error detection, demonstrate the effectiveness of the proposed SPE attention.
Chengyu Jiao, Shuhao Chen
EMNLP2
2025 Open Your Eyes: Vision Enhances Message Passing Neural Networks in Link Prediction
abstract
Message-passing graph neural networks (MPNNs) and structural features (SFs) are cornerstones for the link prediction task. However, as a common and intuitive mode of understanding, the potential of visual perception has been overlooked in the MPNN community. For the first time, we equip MPNNs with vision structural awareness by proposing an effective framework called Graph Vision Network (GVN), along with a more efficient variant (E-GVN). Extensive empirical results demonstrate that with the proposed frameworks, GVN consistently benefits from the vision enhancement across seven link prediction datasets, including challenging large-scale graphs. Such improvements are compatible with existing state-of-the-art (SOTA) methods and GVNs achieve new SOTA results, thereby underscoring a promising novel direction for link prediction.
Yanbin Wei, Xuehao Wang, Zhan Zhuang, Yang Chen 0031, Shuhao Chen, Yulong Zhang 0005, James T. Kwok, Yu Zhang 0006
ICML5
2025 Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
abstract
Low-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal minima near their initialization. This hampers model generalization and limits downstream operators such as adapter merging and pruning. Here, we propose CoTo, a progressive training strategy that gradually increases adapters’ activation probability over the course of fine-tuning. By stochastically deactivating adapters, CoTo encourages more balanced optimization and broader exploration of the loss landscape. We provide a theoretical analysis showing that CoTo promotes layer-wise dropout stability and linear mode connectivity, and we adopt a cooperative-game approach to quantify each adapter’s marginal contribution. Extensive experiments demonstrate that CoTo consistently boosts single-task performance, enhances multi-task merging accuracy, improves pruning robustness, and reduces training overhead, all while remaining compatible with diverse LoRA variants. Code is available at https://github.com/zwebzone/coto.
Zhan Zhuang, Xiequn Wang, Yulong Zhang 0005, Qiushi Huang, Shuhao Chen, Xuehao Wang, Yanbin Wei, Yuhe Nie, Kede Ma, Yu Zhang 0006, Ying Wei 0001
ICML6
2025 SNRWLS: Improve (W)PMS Solver with Weighting Strategies Related to Number of Soft Clauses
Shuhao Chen, Menghua Jiang 0001
TASE1
2025 Domain-guided conditional diffusion model for unsupervised domain adaptation
Yulong Zhang 0005, Shuhao Chen, Weisen Jiang, Yu Zhang 0006, Jiangang Lu, James T. Kwok
Neural Networks2
2024 Rethinking Guidance Information to Utilize Unlabeled Samples: A Label Encoding Perspective
abstract
Empirical Risk Minimization (ERM) is fragile in scenarios with insufficient labeled samples. A vanilla extension of ERM to unlabeled samples is Entropy Minimization (EntMin), which employs the soft-labels of unlabeled samples to guide their learning. However, EntMin emphasizes prediction discriminability while neglecting prediction diversity. To alleviate this issue, in this paper, we rethink the guidance information to utilize unlabeled samples. By analyzing the learning objective of ERM, we find that the guidance information for labeled samples in a specific category is the corresponding label encoding. Inspired by this finding, we propose a Label-Encoding Risk Minimization (LERM). It first estimates the label encodings through prediction means of unlabeled samples and then aligns them with their corresponding ground-truth label encodings. As a result, the LERM ensures both prediction discriminability and diversity, and it can be integrated into existing methods as a plugin. Theoretically, we analyze the relationships between LERM and ERM as well as EntMin. Empirically, we verify the superiority of the LERM under several label insufficient scenarios. The codes are available at https://github.com/zhangyl660/LERM.
Yulong Zhang 0005, Yuan Yao 0016, Shuhao Chen, Pengrong Jin, Yu Zhang 0006, Jiangang Lu
ICML3
2024 RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models
abstract
Recent works show that assembling multiple off-the-shelf large language models (LLMs) can harness their complementary abilities. To achieve this, routing is a promising method, which learns a router to select the most suitable LLM for each query. However, existing routing models are ineffective when multiple LLMs perform well for a query. To address this problem, in this paper, we propose a method called query-based Router by Dual Contrastive learning (RouterDC). The RouterDC model, which consists of an encoder and LLM embeddings, is trained by two proposed contrastive losses (sample-LLM and sample-sample losses). Experimental results show that RouterDC is effective in assembling LLMs and largely outperforms individual top-performing LLMs as well as existing routing methods on both in-distribution (+2.76\%) and out-of-distribution (+1.90\%) tasks. The source code is available at https://github.com/shuhao02/RouterDC.
Shuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok, Yu Zhang 0006
NeurIPS1