VLDB 2026 Research / reviewers in the wild / expert
Yuhua Li 0003
dblp:79/5796-3
· DBLP profile ↗
81ranked-venue papers
3as first author
50since 2021 · last 2026
0000-0002-1846-4941ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 2 first-author · 37 since 2021Databases, data management, data science and information retrieval · 20 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 12 since 2021Systems, architecture and hardware · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoupling Template Bias in CLIP: Harnessing Empty Prompts for Enhanced Few-Shot LearningabstractThe Contrastive Language-Image Pre-Training (CLIP) model excels in few-shot learning by aligning visual and textual representations. Our study shows that template-sample similarity (TSS), defined as the resemblance between a text template and an image sample, introduces bias. This bias leads the model to rely on template proximity rather than true sample-to-category alignment, reducing both accuracy and robustness in classification. We present a framework that uses empty prompts, textual inputs that convey the idea of “emptiness” without category information. These prompts capture unbiased template features and offset TSS bias. The framework employs two stages. During pre-training, empty prompts reveal and reduce template-induced bias within the CLIP encoder. During few-shot fine-tuning, a bias calibration loss enforces correct alignment between images and their categories, ensuring the model focuses on relevant visual cues. Experiments across multiple benchmarks demonstrate that our template correction method significantly reduces performance fluctuations caused by TSS, yielding higher classification accuracy and stronger robustness. Zhenyu Zhang 0035, Yixiong Zou, Zhimeng Huang, Yuhua Li 0003 |
AAAI | 5 |
| 2026 | UMPIRE: Unveiling LLM-generated Posts via Redundant ExpressionsabstractXiaoquan Yi, Haixing Wu, Haozhao Wang, Yichen Li, Yuhua Li, Rui Zhang, Ruixuan Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoquan Yi, Haixing Wu, Haozhao Wang, Yichen Li 0006, Yuhua Li 0003, Rui Zhang 0003, Ruixuan Li 0001 |
ACL (1) | 5 |
| 2026 | Language-Guided Game-Theoretic Fairness in Web-Enabled Energy NetworksabstractWeb platforms are reshaping resource allocation in distributed energy networks globally, from off-grid communities to lunar bases. Algorithmic decision-makers face the fundamental challenge of fairly distributing scarce resources among heterogeneous stakeholders. Traditional approaches assume complete rationality with perfect information and unlimited computation, yet distributed networks only permit local observation, requiring fairness to emerge from individual strategic interactions. Centralized optimization fails due to exponential complexity, rule-based methods cannot adapt to disruptions, and existing platforms translate economic inequality into energy access inequality. Recognizing the unattainability of complete rationality necessitates bounded rationality: pursuing provably convergent satisficing solutions under incomplete information and limited computation, translating natural language ethics into computable constraints, and designing incentives so self-interested behavior satisfies fairness at equilibrium. We propose a unified semantic-game-distributed framework. Large language models map ambiguous ethical principles into game-theoretic parameters through semantic parameterization, with contrastive learning ensuring semantic consistency and temporal stability. A two-layer Stackelberg game implements incentive design: the platform signals through differentiated pricing while nodes optimize locally, enabling fairness to emerge from equilibrium. Distributed asynchronous iteration achieves global convergence through local communication, with cognitive models adaptively adjusting step sizes and differential perturbation preserving privacy. Theoretical analysis establishes equilibrium existence and convergence guarantees, while extreme scenarios validate robustness under information scarcity and high uncertainty. Yuhua Li 0003, Yuntao Zou, Qianqi Zhang, Ruixuan Li 0001, Zeling Xu, Wei Wang 0395 |
WWW | 2 |
| 2026 | Community-strength-based fairness in GNN explanations for heterogeneous link prediction
Yanhong Wen, Yuhua Li 0003, Yixiong Zou, Kai Shu, Quan Fu, Ruixuan Li 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Alleviating noise memorization for adversarially robust few-shot learning
Yiman Hu, Yixiong Zou, Xiaosen Wang, Yuhua Li 0003, Kun He 0001, Ruixuan Li 0001 |
Neural Networks | 4 |
| 2026 | Rethinking Graph Contrastive Learning for Heterophilic Graphs: An Effective Method for Heterophilic GCL Methods With Regularization and Stabilization Techniques Enhanced High-Pass FilterabstractGraph contrastive learning (GCL) is a powerful self-supervised learning approach. However, existing GCL methods are designed for homophilic graphs, using low-pass filters that struggle to capture high-frequency components in heterophilic graphs. We proposeGraphContrastiveLearning withRegularization and stabilization techniques enhanced high-passFilter (GCLRF).REgularization andStabilization techniques enhancedHigh-pass filter (RESH) can serve as a mutually promoting plug-in, significantly improving the performance of various homophilic GCL training strategies on heterophilic graphs. We also investigate four component orderings in RESH and identify the optimal fusion mechanism, demonstrating its critical impact on performance. Experiments show GCLRF achieves state-of-the-art (SOTA) performance across six benchmark datasets in node classification and clustering. Notably, on the Cornell dataset, GCLRF outperformers classification accuracy by 6.76% and achieves a 23.64%relative improvement in clustering normalized mutual information (NMI). Yuhua Li 0003, Yixiong Zou, Keke Huang, Rui Zhang 0003, Ruixuan Li 0001, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Reconstruction Target Matters in Masked Image Modeling for Cross-Domain Few-Shot LearningabstractCross-Domain Few-Shot Learning (CDFSL) requires the model to transfer knowledge from the data-abundant source domain to data-scarce target domains for fast adaptation, where the large domain gap makes CDFSL a challenging problem. Masked Autoencoder (MAE) excels in effectively using unlabeled data and learning image’s global structures, enhancing model generalization and robustness. However, in the CDFSL task with significant domain shifts, we find MAE even shows lower performance than the baseline supervised models. In this paper, we first delve into this phenomenon for an interpretation. We find that MAE tends to focus on low-level domain information during reconstructing pixels while changing the reconstruction target to token features could mitigate this problem. However, not all features are beneficial, as we then find reconstructing high-level features can hardly improve the model’s transferability, indicating a trade-off between filtering domain information and preserving the image’s global structure. In all, the reconstruction target matters for the CDFSL task. Based on the above findings and interpretations, we further propose Domain-Agnostic Masked Image Modeling (DAMIM) for the CDFSL task. DAMIM includes an Aggregated Feature Reconstruction module to automatically aggregate features for reconstruction, with balanced learning of domain-agnostic information and images’ global structure, and a Lightweight Decoder module to further benefit the encoder’s generalizability. Experiments on four CDFSL datasets demonstrate that our method achieves state-of-the-art performance. Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
AAAI | 3 |
| 2025 | The Devil is in Low-Level Features for Cross-Domain Few-Shot SegmentationabstractCross-Domain Few-Shot Segmentation (CDFSS) is proposed to transfer the pixel-level segmentation capabilities learned from large-scale source-domain datasets to downstream target-domain datasets, with only a few annotated images per class. In this paper, we focus on a well-observed but under-explored phenomenon in CDFSS: for target domains, particularly those distant from the source domain, segmentation performance peaks at the very early epochs, and declines sharply as the source-domain training proceeds. We delve into this phenomenon for an interpretation: low-level features are vulnerable to domain shifts, leading to sharper loss landscapes during the source-domain training, which is the devil of CDFSS. Based on this phenomenon and interpretation, we further propose a method that includes two plug-and-play modules: one to flatten the loss landscapes for low-level features during source-domain training as a novel sharpness-aware minimization method, and the other to directly supplement target-domain information to the model during target-domain testing by low-level-based calibration. Extensive experiments on four target datasets validate our rationale and demonstrate that our method surpasses the state-of-the-art method in CDFSS signifcantly by 3.71% and 5.34% average MIoU in 1-shot and 5-shot scenarios, respectively. Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
CVPR | 3 |
| 2025 | Revisiting Pool-Based Prompt Learning for Few-Shot Class-Incremental Learning
Yongwei Jiang, Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
ICCV | 3 |
| 2025 | Beyond Zero Initialization: Investigating the Impact of Non-Zero Initialization on LoRA Fine-Tuning DynamicsabstractLow-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method. In standard LoRA layers, one of the matrices, $A$ or $B$, is initialized to zero, ensuring that fine-tuning starts from the pretrained model. However, there is no theoretical support for this practice. In this paper, we investigate the impact of non-zero initialization on LoRA’s fine-tuning dynamics from an infinite-width perspective. Our analysis reveals that, compared to zero initialization, simultaneously initializing $A$ and $B$ to non-zero values improves LoRA’s robustness to suboptimal learning rates, particularly smaller ones. Further analysis indicates that although the non-zero initialization of $AB$ introduces random noise into the pretrained weight, it generally does not affect fine-tuning performance. In other words, fine-tuning does not need to strictly start from the pretrained model. The validity of our findings is confirmed through extensive experiments across various models and datasets. The code is available at https://github.com/Leopold1423/non_zero_lora-icml25. Shiwei Li 0002, Xiandi Luo, Xing Tang 0007, Haozhao Wang, Weihong Luo, Yuhua Li 0003, Xiuqiang He 0001, Ruixuan Li 0001 |
ICML | 7 |
| 2025 | The Panaceas for Improving Low-Rank Decomposition in Communication-Efficient Federated LearningabstractTo improve the training efficiency of federated learning (FL), previous research has employed low-rank decomposition techniques to reduce communication overhead.
In this paper, we seek to enhance the performance of these low-rank decomposition methods. Specifically, we focus on three key issues related to decomposition in FL: what to decompose, how to decompose, and how to aggregate. Subsequently, we introduce three novel techniques: Model Update Decomposition (MUD), Block-wise Kronecker Decomposition (BKD), and Aggregation-Aware Decomposition (AAD), each targeting a specific issue. These techniques are complementary and can be applied simultaneously to achieve optimal performance. Additionally, we provide a rigorous theoretical analysis to ensure the convergence of the proposed MUD. Extensive experimental results show that our approach achieves faster convergence and superior accuracy compared to relevant baseline methods. The code is available at https://github.com/Leopold1423/fedmud-icml25. Shiwei Li 0002, Xiandi Luo, Haozhao Wang, Xing Tang 0007, Weihong Luo, Yuhua Li 0003, Xiuqiang He 0001, Ruixuan Li 0001 |
ICML | 7 |
| 2025 | Adapter Naturally Serves as Decoupler for Cross-Domain Few-Shot Semantic SegmentationabstractCross-domain few-shot segmentation (CD-FSS) is proposed to first pre-train the model on a source-domain dataset with sufficient samples, and then transfer the model to target-domain datasets where only a few training samples are available for efficient finetuning. There are majorly two challenges in this task: (1) the domain gap and (2) finetuning with scarce data. To solve these challenges, we revisit the adapter-based methods, and discover an intriguing insight not explored in previous works: the adapter not only helps the fine-tuning of downstream tasks but also naturally serves as a domain information decoupler. Then, we delve into this finding for an interpretation, and we find the model's inherent structure could lead to a natural decoupling of domain information. Building upon this insight, we propose the Domain Feature Navigator (DFN), which is a structure-based decoupler instead of loss-based ones like current works, to capture domain-specific information, thereby directing the model's attention towards domain-agnostic knowledge. Moreover, to prevent the potential excessive overfitting of DFN during the source-domain training, we further design the SAM-SVN method to constrain DFN from learning sample-specific knowledge. On target domains, we freeze the model and fine-tune the DFN to learn knowledge specific to target domains. Extensive experiments demonstrate that our method surpasses the state-of-the-art method in CD-FSS significantly by 2.69% and 4.68% average MIoU in 1-shot and 5-shot scenarios, respectively. Jintao Tong, Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
ICML | 5 |
| 2025 | Self-Disentanglement and Re-Composition for Cross-Domain Few-Shot SegmentationabstractCross-Domain Few-Shot Segmentation (CD-FSS) aims to transfer knowledge from a large-scale source-domain dataset to unseen target-domain datasets with limited annotated samples. Current methods typically compare the distance between training and testing samples for mask prediction. However, a problem of feature entanglement exists in this well-adopted method, which binds multiple patterns together and harms the transferability. However, we find an entanglement problem exists in this widely adopted method, which tends to bind source-domain patterns together and make each of them hard to transfer. In this paper, we aim to address this problem for the CD-FSS task. We first find a natural decomposition of the ViT structure, based on which we delve into the entanglement problem for an interpretation. We find the decomposed ViT components are crossly compared between images in distance calculation, where the rational comparisons are entangled with those meaningless ones by their equal importance, leading to the entanglement problem. Based on this interpretation, we further propose to address the entanglement problem by learning to weigh for all comparisons of ViT components, which learn disentangled features and re-compose them for the CD-FSS task, benefiting both the generalization and finetuning. Experiments show that our model outperforms the state-of-the-art CD-FSS method by 1.92% and 1.88% in average accuracy under 1-shot and 5-shot settings, respectively. Jintao Tong, Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
ICML | 4 |
| 2025 | Revisiting Continuity of Image Tokens for Cross-domain Few-shot LearningabstractVision Transformer (ViT) has achieved remarkable success due to its large-scale pretraining on general domains, but it still faces challenges when applying it to downstream distant domains that have only scarce training data, which gives rise to the Cross-Domain Few-Shot Learning (CDFSL) task. Inspired by Self-Attention's insensitivity to token orders, we find an interesting phenomenon neglected in current works: disrupting the continuity of image tokens (i.e., making pixels not smoothly transited across patches) in ViT leads to a noticeable performance decline in the general (source) domain but only a marginal decrease in downstream target domains. This questions the role of image tokens' continuity in ViT's generalization under large domain gaps. In this paper, we delve into this phenomenon for an interpretation. We find continuity aids ViT in learning larger spatial patterns, which are harder to transfer than smaller ones, enlarging domain distances. Meanwhile, it implies that only smaller patterns within each patch could be transferred under extreme domain gaps. Based on this interpretation, we further propose a simple yet effective method for CDFSL that better disrupts the continuity of image tokens, encouraging the model to rely less on large patterns and more on smaller ones. Extensive experiments show the effectiveness of our method in reducing domain gaps and outperforming state-of-the-art works. Codes and models are available at https://github.com/shuaiyi308/ReCIT. Shuai Yi, Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
ICML | 3 |
| 2025 | Random Registers for Cross-Domain Few-Shot LearningabstractCross-domain few-shot learning (CDFSL) aims to transfer knowledge from a data-sufficient source domain to data-scarce target domains. Although Vision Transformer (ViT) has shown superior capability in many vision tasks, its transferability against huge domain gaps in CDFSL is still under-explored. In this paper, we find an intriguing phenomenon: during the source-domain training, prompt tuning, as a common way to train ViT, could be harmful for the generalization of ViT in target domains, but setting them to random noises (i.e., random registers) could consistently improve target-domain performance. We then delve into this phenomenon for an interpretation. We find that learnable prompts capture domain information during the training on the source dataset, which views irrelevant visual patterns as vital cues for recognition. This can be viewed as a kind of overfitting and increases the sharpness of the loss landscapes. In contrast, random registers are essentially a novel way of perturbing attention for the sharpness-aware minimization, which helps the model find a flattened minimum in loss landscapes, increasing the transferability. Based on this phenomenon and interpretation, we further propose a simple but effective approach for CDFSL to enhance the perturbation on attention maps by adding random registers on the semantic regions of image tokens, improving the effectiveness and efficiency of random registers. Extensive experiments on four benchmarks validate our rationale and state-of-the-art performance. Codes and models are available at https://github.com/shuaiyi308/REAP. Shuai Yi, Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
ICML | 3 |
| 2025 | Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank AdaptationabstractLow-rank adaptation (LoRA) is a parameter-efficient fine-tuning (PEFT) method widely used in large language models (LLMs).
LoRA essentially describes the projection of an input space into a low-dimensional output space, with the dimensionality determined by the LoRA rank.
In standard LoRA, all input tokens share the same weights and undergo an identical input-output projection.
This limits LoRA's ability to capture token-specific information due to the inherent semantic differences among tokens.
To address this limitation, we propose **Token-wise Projected Low-Rank Adaptation (TopLoRA)**, which dynamically adjusts LoRA weights according to the input token, thereby learning token-wise input-output projections in an end-to-end manner.
Formally, the weights of TopLoRA can be expressed as $B\Sigma_X A$, where $A$ and $B$ are low-rank matrices (as in standard LoRA), and $\Sigma_X$ is a diagonal matrix generated from each input token $X$.
Notably, TopLoRA does not increase the rank of LoRA weights but achieves more granular adaptation by learning token-wise LoRA weights (i.e., token-wise input-output projections).
Extensive experiments across multiple models and datasets demonstrate that TopLoRA consistently outperforms LoRA and its variants.
The code is available at https://github.com/Leopold1423/toplora-neurips25. Shiwei Li 0002, Xiandi Luo, Haozhao Wang, Xing Tang 0007, Ziqiang Cui, Dugang Liu, Yuhua Li 0003, Xiuqiang He 0001, Ruixuan Li 0001 |
NeurIPS | 7 |
| 2025 | FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language ModelsabstractLarge vision-language models (LVLMs) excel at multimodal understanding but suffer from high computational costs due to redundant vision tokens. Existing pruning methods typically rely on single-layer attention scores to rank and prune redundant visual tokens to solve this inefficiency. However, as the interaction between tokens and layers is complicated, this raises a basic question: Is such a simple single-layer criterion sufficient to identify redundancy? To answer this question, we rethink the emergence of redundant visual tokens from a fundamental perspective: information flow, which models the interaction between tokens and layers by capturing how information moves between tokens across layers. We find (1) the CLS token acts as an information relay, which can simplify the complicated flow analysis; (2) the redundancy emerges progressively and dynamically via layer-wise attention concentration; and (3) relying solely on attention scores from single layers can lead to contradictory redundancy identification. Based on this, we propose FlowCut, an information-flow-aware pruning framework, mitigating the insufficiency of the current criterion for identifying redundant tokens and better aligning with the model's inherent behaviors. Extensive experiments show FlowCut achieves superior results, outperforming SoTA by 1.6% on LLaVA-1.5-7B with 88.9% token reduction, and by 4.3% on LLaVA-NeXT-7B with 94.4% reduction, delivering 3.2$\times$ speed-up in the prefilling stage. Our code is available at https://github.com/TungChintao/FlowCut. Jintao Tong, Wenwei Jin, Pengda Qin, Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
NeurIPS | 7 |
| 2025 | ChatbotID: Identifying Chatbots with Granger Causality TestabstractWith the increasing sophistication of Large Language Models (LLMs), it is crucial to develop reliable methods to accurately identify whether an interlocutor in real-time dialogue is human or chatbot. However, existing detection methods are primarily designed for analyzing full documents, not the unique dynamics and characteristics of dialogue. These approaches frequently overlook the nuances of interaction that are essential in conversational contexts. This work identifies two key patterns in dialogues: (1) Human-Human (H-H) interactions exhibit significant bidirectional sentiment influence, while (2) Human-Chatbot (H-C) interactions display a clear asymmetric pattern. We propose an innovative approach named ChatbotID, which
applies the Granger Causality Test (GCT) to extract a novel set of interactional features that capture the evolving, predictive relationships between conversational attributes. By synergistically fusing these GCT-based interactional features with contextual embeddings, and optimizing the model through a meticulous loss function. Experimental results across multiple datasets and detection models demonstrate the effectiveness of our framework, with significant improvements in accuracy for distinguishing between H-H and H-C dialogues. Xiaoquan Yi, Haozhao Wang, Yining Qi, Wenchao Xu 0001, Rui Zhang 0003, Yuhua Li 0003, Ruixuan Li 0001 |
NeurIPS | 6 |
| 2025 | Multi-modal Robustness Fake News Detection with Cross-Modal and Propagation Network Contrastive Learning
Yuhua Li 0003, Yujing Zhang 0001, Kai Shu, Ruixuan Li 0001, Philip S. Yu |
Knowl. Based Syst. | 4 |
| 2025 | Fair path explanations for heterogeneous link prediction from a community perspective
Yanhong Wen, Yuhua Li 0003, Yixiong Zou, Quan Fu, Ruixuan Li 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Community-influencing path explanation for link prediction in heterogeneous graph neural network
Yanhong Wen, Yuhua Li 0003, Yixiong Zou, Kai Shu, Jinxian Ye, Quan Fu, Ruixuan Li 0001 |
Neural Networks | 2 |
| 2025 | A Survey on Self-Supervised Graph Foundation Models: Knowledge-Based PerspectiveabstractThe field of graph foundation models (GFMs) has seen a dramatic rise in interest in recent years. Their powerful generalization ability is believed to be endowed by self-supervised pre-training and downstream tuning techniques. There is a wide variety of knowledge patterns embedded in the graph data, such as node properties and clusters, which are crucial for learning generalized representations for GFMs. We present a comprehensive survey of self-supervised GFMs from a novel knowledge-based perspective. Our main contribution is a knowledge-based taxonomy that categorizes self-supervised graph models by the specific graph knowledge utilized: microscopic (nodes, links, etc.), mesoscopic (context, clusters, etc.), and macroscopic (global structure, manifolds, etc.). It covers a total of 9 knowledge categories and 300 references for self-supervised pre-training as well as various downstream tuning strategies. Such a knowledge-based taxonomy allows us to more clearly re-examine potential GFM architectures, including large language models (LLMs), as well as provide deeper insights for constructing future GFMs. Yixin Su 0001, Yuhua Li 0003, Yixiong Zou, Ruixuan Li 0001, Rui Zhang 0003 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Flatten Long-Range Loss Landscapes for Cross-Domain Few-Shot LearningabstractCross-domain few-shot learning (CDFSL) aims to acquire knowledge from limited training data in the target domain by leveraging prior knowledge transferred from source domains with abundant training samples. CDFSL faces challenges in transferring knowledge across dissimilar domains and fine-tuning models with limited training data. To address these challenges, we initially extend the analysis of loss landscapes from the parameter space to the representation space, which allows us to simultaneously interpret the transferring and fine-tuning difficulties of CDFSL models. We observe that sharp minima in the loss landscapes of the representation space result in representations that are hard to transfer and fine-tune. More-over, existing flatness-based methods have limited generalization ability due to their short-range flatness. To enhance the transferability and facilitate fine-tuning, we introduce a simple yet effective approach to achieve long-range flat-tening of the minima in the loss landscape. This approach considers representations that are differently normalized as minima in the loss landscape and flattens the high-loss region in the middle by randomly sampling interpolated representations. We implement this method as a new normalization layer that replaces the original one in both CNNs and ViTs. This layer is simple and lightweight, introducing only a minimal number of additional parameters. Experimental results on 8 datasets demonstrate that our approach outperforms state-of-the-art methods in terms of average accuracy. Moreover, our method achieves performance improvements of up to 9% compared to the current best approaches on individual datasets. Our code will be released. Yixiong Zou, Yiman Hu, Yuhua Li 0003, Ruixuan Li 0001 |
CVPR | 4 |
| 2024 | Adversarial Attack for Explanation Robustness of Rationalization ModelsabstractRationalization models, which select a subset of input text as rationale—crucial for humans to understand and trust predictions—have recently emerged as a prominent research area in eXplainable Artificial Intelligence (XAI). However, most of previous studies mainly focus on improving the quality of the rationale, ignoring its robustness to malicious attack. Specifically, whether the rationalization models can still generate high-quality rationale under the adversarial attack remains unknown. To explore this, this paper proposes UAT2E, which aims to undermine the explainability of rationalization models without altering their predictions, thereby eliciting distrust in these models from human users. UAT2E employs the gradient-based search on triggers and then inserts them into the original input to conduct both the non-target and target attack. Experimental results on five datasets reveal the vulnerability of rationalization models in terms of explanation, where they tend to select more meaningless tokens under attacks. Based on this, we make a series of recommendations for improving rationalization models in terms of explanation. Yuankai Zhang 0002, Lingxiao Kong, Haozhao Wang, Ruixuan Li 0001, Jun Wang 0018, Yuhua Li 0003, Wei Liu 0144 |
ECAI | 6 |
| 2024 | FedTA: Unsupervised Federated Prototype Learning with Temperature AdaptationabstractFederated Learning (FL) has emerged as a foundational paradigm that enables collaborative training of deep neural networks across distributed clients while ensuring data privacy. However, most of the existing research primarily concentrates on supervised federated learning tailored for specific downstream tasks. In this paper, we concentrate on unsupervised federated learning, where clients collaborate to identify data patterns, structures, or representations. Contrastive learning methods have proven highly effective for unsupervised learning in server-based environments. However, a straightforward adaptation of the contrastive learning approach for this setting falls short. Upon analyzing client behaviors during training, we identified an unstable training process stemming from NonIID issues, which can result in diminished performance. To address this problem, we propose a novel Temperature-Adaption based unsupervised Federated prototype learning approach, termed FedTA. This method incorporates two straightforward yet potent solutions: 1. Prototype Enhancement and 2. Temperature-Adaption. Experimental results show that our proposed approach surpasses the state-of-the-art, achieving a performance accuracy improvement of up to 5.3%. Juan Zhao 0010, Xiaoquan Yi, Ruixuan Li 0001, Yuhua Li 0003, Haozhao Wang, Yichen Li 0006, Zhiying Deng |
HPCC | 4 |
| 2024 | FedBAT: Communication-Efficient Federated Learning via Learnable BinarizationabstractFederated learning is a promising distributed machine learning paradigm that can effectively exploit large-scale data without exposing users’ privacy. However, it may incur significant communication overhead, thereby potentially impairing the training efficiency. To address this challenge, numerous studies suggest binarizing the model updates. Nonetheless, traditional methods usually binarize model updates in a post-training manner, resulting in significant approximation errors and consequent degradation in model accuracy. To this end, we propose Federated Binarization-Aware Training (FedBAT), a novel framework that directly learns binary model updates during the local training process, thus inherently reducing the approximation errors. FedBAT incorporates an innovative binarization operator, along with meticulously designed derivatives to facilitate efficient learning. In addition, we establish theoretical guarantees regarding the convergence of FedBAT. Extensive experiments are conducted on four popular datasets. The results show that FedBAT significantly accelerates the convergence and exceeds the accuracy of baselines by up to 9%, even surpassing that of FedAvg in some cases. Shiwei Li 0002, Wenchao Xu 0001, Haozhao Wang, Xing Tang 0007, Yining Qi, Weihong Luo, Yuhua Li 0003, Xiuqiang He 0001, Ruixuan Li 0001 |
ICML | 8 |
| 2024 | Compositional Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) is proposed to continually learn from novel classes with only a few samples after the (pre-)training on base classes with sufficient data. However, this remains a challenge. In contrast, humans can easily recognize novel classes with a few samples. Cognitive science demonstrates that an important component of such human capability is compositional learning. This involves identifying visual primitives from learned knowledge and then composing new concepts using these transferred primitives, making incremental learning both effective and interpretable. To imitate human compositional learning, we propose a cognitive-inspired method for the FSCIL task. We define and build a compositional model based on set similarities, and then equip it with a primitive composition module and a primitive reuse module. In the primitive composition module, we propose to utilize the Centered Kernel Alignment (CKA) similarity to approximate the similarity between primitive sets, allowing the training and evaluation based on primitive compositions. In the primitive reuse module, we enhance primitive reusability by classifying inputs based on primitives replaced with the closest primitives from other classes. Experiments on three datasets validate our method, showing it outperforms current state-of-the-art methods with improved interpretability. Our code is available at https://github.com/Zoilsen/Comp-FSCIL. Yixiong Zou, Shanghang Zhang, Haichen Zhou, Yuhua Li 0003, Ruixuan Li 0001 |
ICML | 4 |
| 2024 | Delve into Base-Novel Confusion: Redundancy Exploration for Few-Shot Class-Incremental Learning
Haichen Zhou, Yixiong Zou, Ruixuan Li 0001, Yuhua Li 0003, Kui Xiao |
IJCAI | 4 |
| 2024 | Masked Random Noise for Communication-Efficient Federated LearningabstractFederated learning is a promising distributed training paradigm that effectively safeguards data privacy. However, it may involve significant communication costs, which hinders training efficiency. In this paper, we aim to enhance communication efficiency from a new perspective. Specifically, we request the distributed clients to find optimal model updates relative to global model parameters within predefined random noise. For this purpose, we propose Federated Masked Random Noise (FedMRN), a novel framework that enables clients to learn a 1-bit mask for each model parameter and apply masked random noise (i.e., the Hadamard product of random noise and masks) to represent model updates. To make FedMRN feasible, we propose an advanced mask training strategy, called progressive stochastic masking (PSM). After local training, each client only need to transmit local masks and a random seed to the server. Additionally, we provide theoretical guarantees for the convergence of FedMRN under both strongly convex and non-convex assumptions. Extensive experiments are conducted on four popular datasets. The results show that FedMRN exhibits superior convergence speed and test accuracy compared to relevant baselines, while attaining a similar level of accuracy as FedAvg. Shiwei Li 0002, Yingyi Cheng, Haozhao Wang, Xing Tang 0007, Weihong Luo, Yuhua Li 0003, Dugang Liu, Xiuqiang He 0001, Ruixuan Li 0001 |
ACM Multimedia | 7 |
| 2024 | Learning Unknowns from Unknowns: Diversified Negative Prototypes Generator for Few-shot Open-Set RecognitionabstractFew-shot open-set recognition (FSOR) is a challenging task that requires a model to recognize known classes and identify unknown classes with limited labeled data. Existing approaches, particularly Negative-Prototype-Based methods, generate negative prototypes based solely on known class data. However, as the unknown space is infinite while the known space is limited, these methods suffer from limited representation capability. To address this limitation, we propose a novel approach, termed Diversified Negative Prototypes Generator (DNPG), which adopts the principle of "learning unknowns from unknowns." Our method leverages the unknown space information learned from base classes to generate more representative negative prototypes for novel classes. During the pre-training phase, we learn the unknown space representation of the base classes. This representation, along with inter-class relationships, is then utilized in the meta-learning process to construct negative prototypes for novel classes. To prevent prototype collapse and ensure adaptability to varying data compositions, we introduce the Swap Alignment (SA) module. Our DNPG model, by learning from the unknown space, generates negative prototypes that cover a broader unknown space, thereby achieving state-of-the-art performance on three standard FSOR datasets. The repository of this project is available at https://github.com/iCGY96/DNPG. Zhenyu Zhang 0035, Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
ACM Multimedia | 4 |
| 2024 | MICM: Rethinking Unsupervised Pretraining for Enhanced Few-shot LearningabstractHumans exhibit a remarkable ability to learn quickly from a limited number of labeled samples, a capability that starkly contrasts with that of current machine learning systems. Unsupervised Few-Shot Learning (U-FSL) seeks to bridge this divide by reducing reliance on annotated datasets during initial training phases. In this work, we first quantitatively assess the impacts of Masked Image Modeling (MIM) and Contrastive Learning (CL) on few-shot learning tasks. Our findings highlight the respective limitations of MIM and CL in terms of discriminative and generalization abilities, which contribute to their underperformance in U-FSL contexts. To address these trade-offs between generalization and discriminability in unsupervised pretraining, we introduce a novel paradigm named Masked Image Contrastive Modeling (MICM). MICM creatively combines the targeted object learning strength of CL with the generalized visual feature learning capability of MIM, significantly enhancing its efficacy in downstream few-shot learning inference. Extensive experimental analyses confirm the advantages of MICM, demonstrating significant improvements in both generalization and discrimination capabilities for few-shot learning. Our comprehensive quantitative evaluations further substantiate the superiority of MICM, showing that our two-stage U-FSL framework based on MICM markedly outperforms existing leading baselines. The repository of this project is available at https://github.com/iCGY96/MICM. Zhenyu Zhang 0035, Yixiong Zou, Zhimeng Huang, Yuhua Li 0003, Ruixuan Li 0001 |
ACM Multimedia | 5 |
| 2024 | Generate Universal Adversarial Perturbations for Few-Shot LearningabstractDeep networks are known to be vulnerable to adversarial examples which are deliberately designed to mislead the trained model by introducing imperceptible perturbations to input samples. Compared to traditional perturbations crafted specifically for each data point, Universal Adversarial Perturbations (UAPs) are input-agnostic and shown to be more practical in the real world. However, UAPs are typically generated in a close-set scenario that shares the same classification task during the training and testing phases. This paper demonstrates the ineffectiveness of traditional UAPs in open-set scenarios like Few-Shot Learning (FSL). Through analysis, we identify two primary challenges that hinder the attacking process: the task shift and the semantic shift. To enhance the transferability of UAPs in FSL, we propose a unifying attacking framework addressing these two shifts. The task shift is addressed by aligning proxy tasks to the downstream tasks, while the semantic shift is handled by leveraging the generalizability of pre-trained encoders.The proposed Few-Shot Attacking FrameWork, denoted as FSAFW, can effectively generate UAPs across various FSL training paradigms and different downstream tasks. Our approach not only sets a new standard for state-of-the-art works but also significantly enhances attack performance, exceeding the baseline method by over 16\%. Yiman Hu, Yixiong Zou, Ruixuan Li 0001, Yuhua Li 0003 |
NeurIPS | 4 |
| 2024 | Lightweight Frequency Masker for Cross-Domain Few-Shot Semantic SegmentationabstractCross-domain few-shot segmentation (CD-FSS) is proposed to first pre-train the model on a large-scale source-domain dataset, and then transfer the model to data-scarce target-domain datasets for pixel-level segmentation. The significant domain gap between the source and target datasets leads to a sharp decline in the performance of existing few-shot segmentation (FSS) methods in cross-domain scenarios. In this work, we discover an intriguing phenomenon: simply filtering different frequency components for target domains can lead to a significant performance improvement, sometimes even as high as 14% mIoU. Then, we delve into this phenomenon for an interpretation, and find such improvements stem from the reduced inter-channel correlation in feature maps, which benefits CD-FSS with enhanced robustness against domain gaps and larger activated regions for segmentation. Based on this, we propose a lightweight frequency masker, which further reduces channel correlations by an Amplitude-Phase Masker (APM) module and an Adaptive Channel Phase Attention (ACPA) module. Notably, APM introduces only 0.01% additional parameters but improves the average performance by over 10%, and ACPA imports only 2.5% parameters but further improves the performance by over 1.5%, which significantly surpasses the state-of-the-art CD-FSS methods. Jintao Tong, Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
NeurIPS | 3 |
| 2024 | Attention Temperature Matters in ViT-Based Cross-Domain Few-Shot LearningabstractCross-domain few-shot learning (CDFSL) is proposed to transfer knowledge from large-scale source-domain datasets to downstream target-domain datasets with only a few training samples. However, Vision Transformer (ViT), as a strong backbone network to achieve many top performances, is still under-explored in the CDFSL task in its transferability against large domain gaps. In this paper, we find an interesting phenomenon of ViT in the CDFSL task: by simply multiplying a temperature (even as small as 0) to the attention in ViT blocks, the target-domain performance consistently increases, even though the attention map is downgraded to a uniform map. In this paper, we delve into this phenomenon for an interpretation. Through experiments, we interpret this phenomenon as a remedy for the ineffective target-domain attention caused by the query-key attention mechanism under large domain gaps. Based on it, we further propose a simple but effective method for the CDFSL task to boost ViT's transferability by resisting the learning of query-key parameters and encouraging that of non-query-key ones. Experiments on four CDFSL datasets validate the rationale of our interpretation and method, showing we can consistently outperform state-of-the-art methods. Our codes are available at https://github.com/Zoilsen/Attn_Temp_CDFSL. Yixiong Zou, Yuhua Li 0003, Ruixuan Li 0001 |
NeurIPS | 3 |
| 2024 | A Closer Look at the CLS Token for Cross-Domain Few-Shot LearningabstractVision Transformer (ViT) has shown great power in learning from large-scale datasets. However, collecting sufficient data for expert knowledge is always difficult. To handle this problem, Cross-Domain Few-Shot Learning (CDFSL) has been proposed to transfer the source-domain knowledge learned from sufficient data to target domains where only scarce data is available. In this paper, we find an intriguing phenomenon neglected by previous works for the CDFSL task based on ViT: leaving the CLS token to random initialization, instead of loading source-domain trained parameters, could consistently improve target-domain performance. We then delve into this phenomenon for an interpretation. We find **the CLS token naturally absorbs domain information** due to the inherent structure of the ViT, which is represented as the low-frequency component in the Fourier frequency space of images. Based on this phenomenon and interpretation, we further propose a method for the CDFSL task to decouple the domain information in the CLS token during the source-domain training, and adapt the CLS token on the target domain for efficient few-shot learning. Extensive experiments on four benchmarks validate our rationale and state-of-the-art performance. Our codes are available at https://github.com/Zoilsen/CLS_Token_CDFSL. Yixiong Zou, Shuai Yi, Yuhua Li 0003, Ruixuan Li 0001 |
NeurIPS | 3 |
| 2024 | Masked Graph Autoencoder with Non-discrete BandwidthsabstractMasked graph autoencoders have emerged as a powerful graph self-supervised learning method that has yet to be fully explored. In this paper, we unveil that the existing discrete edge masking and binary link reconstruction strategies are insufficient to learn topologically informative representations, from the perspective of message propagation on graph neural networks. These limitations include blocking message flows, vulnerability to over-smoothness, and suboptimal neighborhood discriminability. Inspired by these understandings, we explore non-discrete edge masks, which are sampled from a continuous and dispersive probability distribution instead of the discrete Bernoulli distribution. These masks restrict the amount of output messages for each edge, referred to as "bandwidths". We propose a novel, informative, and effective topological masked graph autoencoder using bandwidth masking and a layer-wise bandwidth prediction objective. We demonstrate its powerful graph topological learning ability both theoretically and empirically. Our proposed framework outperforms representative baselines in both self-supervised link prediction (improving the discrete edge reconstructors by at most 20%) and node classification on numerous datasets, solely with a structure-learning pretext. Our implementation is available at https://github.com/Newiz430/Bandana. Yuhua Li 0003, Yixiong Zou, Jiliang Tang, Ruixuan Li 0001 |
WWW | 2 |
| 2024 | DCMSL: Dual influenced community strength-boosted multi-scale graph contrastive learning
Yuhua Li 0003, Philip S. Yu, Yixiong Zou, Ruixuan Li 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Spectral Decomposition and Transformation for Cross-domain Few-shot Learning
Yixiong Zou, Ruixuan Li 0001, Yuhua Li 0003 |
Neural Networks | 4 |
| 2023 | XFed: Improving Explainability in Federated Learning by Intersection Over Union Ratio Extended Client SelectionabstractFederated Learning (FL) allows massive clients to collaboratively train a global model without revealing their private data. Because of the participants’ not independently and identically distributed (non-IID) statistical characteristics, it will cause divergence among the client’s Deep Neural Network model weights and require more communication rounds before training can be converged. Moreover, models trained from non-IID data may also extract biased features and the rationale behind the model is still not fully analyzed and exploited. In this paper, we propose eXplainable-Fed (XFed) which is a novel client selection mechanism that takes both accuracy and explainability into account. Specifically, XFed selects participants in each round based on a small test set’s accuracy via cross-entropy loss and interpretability via XAI-accuracy. XAI-accuracy is calculated by Intersection over Union Ratio between the heat map and the truth mask to evaluate the overall rationale of accuracy. The results of our experiments show that our method has comparable accuracy to state-of-the-art methods specially designed for accuracy while increasing explainability by 14%-35% in terms of rationality. Juan Zhao 0010, Yuankai Zhang 0002, Ruixuan Li 0001, Yuhua Li 0003, Haozhao Wang, Xiaoquan Yi, Zhiying Deng |
ECAI | 4 |
| 2023 | CSGCL: Community-Strength-Enhanced Graph Contrastive LearningabstractGraph Contrastive Learning (GCL) is an effective way to learn generalized graph representations in a self-supervised manner, and has grown rapidly in recent years. However, the underlying community semantics has not been well explored by most previous GCL methods. Research that attempts to leverage communities in GCL regards them as having the same influence on the graph, leading to extra representation errors. To tackle this issue, we define ''community strength'' to measure the difference of influence among communities. Under this premise, we propose a Community-Strength-enhanced Graph Contrastive Learning (CSGCL) framework to preserve community strength throughout the learning process. Firstly, we present two novel graph augmentation methods, Communal Attribute Voting (CAV) and Communal Edge Dropping (CED), where the perturbations of node attributes and edges are guided by community strength. Secondly, we propose a dynamic ''Team-up'' contrastive learning scheme, where community strength is used to progressively fine-tune the contrastive objective. We report extensive experiment results on three downstream tasks: node classification, node clustering, and link prediction. CSGCL achieves state-of-the-art performance compared with other GCL methods, validating that community strength brings effectiveness and generality to graph representations. Our code is available at https://github.com/HanChen-HUST/CSGCL. Yuhua Li 0003, Yixiong Zou, Ruixuan Li 0001, Rui Zhang 0003 |
IJCAI | 3 |
| 2023 | SSTP: Social and Spatial-Temporal Aware Next Point-of-Interest RecommendationabstractAbstract The expansion of available information in location-based social networks (LBSNs) has led to information overload, making it urgent to discover users’ next point-of-interest (POI). Some existing works only consider certain modal information in LBSNs and do not transform them into high-dimensional structures, which hinders the alleviation of the data sparsity problem. Moreover, many approaches rely solely on social relationships, making it difficult to recommend POIs to new users without association information. To tackle these challenges, we propose a social- and spatial–temporal-aware next point-of-Interest (SSTP) recommendation model. SSTP uses two feature encoders based on self-attention mechanism and gate recurrent unit to model users’ check-in enhancement sequence hierarchically. We also design a random neighborhood sampling approach to mine user social relationships, thus alleviating the user cold start problem. Finally, we propose a geographical-aware graph attention network to learn the sensitivity of users to distance. Extensive experiments on two real-world datasets show that SSTP outperforms state-of-the-art models, improving Hit@k by 2.26–6.55 $$\%$$ % and MAP@k by 3.49–6.55 $$\%$$ % . Moreover, SSTP has better performance on sparse data, with an average improvement of 6.09 $$\%$$ % on the Hit@k. The code can be downloaded at https://github.com/Rih0/sstp . Junzhuang Wu, Yujing Zhang 0001, Yuhua Li 0003, Yixiong Zou, Ruixuan Li 0001, Zhenyu Zhang 0035 |
Data Sci. Eng. | 3 |
| 2022 | Deep Neural Factorization Machine for Recommender System
Zhenlong Zhu, Changzheng Liu, Yuhua Li 0003, Ruixuan Li 0001 |
KSEM (2) | 4 |
| 2022 | Margin-Based Few-Shot Class-Incremental Learning with Class-Level Overfitting MitigationabstractFew-shot class-incremental learning (FSCIL) is designed to incrementally recognize novel classes with only few training samples after the (pre-)training on base classes with sufficient samples, which focuses on both base-class performance and novel-class generalization. A well known modification to the base-class training is to apply a margin to the base-class classification. However, a dilemma exists that we can hardly achieve both good base-class performance and novel-class generalization simultaneously by applying the margin during the base-class training, which is still under explored. In this paper, we study the cause of such dilemma for FSCIL. We first interpret this dilemma as a class-level overfitting (CO) problem from the aspect of pattern learning, and then find its cause lies in the easily-satisfied constraint of learning margin-based patterns. Based on the analysis, we propose a novel margin-based FSCIL method to mitigate the CO problem by providing the pattern learning process with extra constraint from the margin-based patterns themselves. Extensive experiments on CIFAR100, Caltech-USCD Birds-200-2011 (CUB200), and miniImageNet demonstrate that the proposed method effectively mitigates the CO problem and achieves state-of-the-art performance. Yixiong Zou, Shanghang Zhang, Yuhua Li 0003, Ruixuan Li 0001 |
NeurIPS | 3 |
| 2022 | Community detection using multitopology and attributes in social networksabstractSummary Community detection is a fundamental research problem in social networks. However, most existing research focuses on homogeneous networks while ignoring the multitopology and attributes in social media. In this article, we propose community detection algorithms based on community kernels to detect high‐quality communities in heterogeneous social networks. It is noticed that the social community has multiple topology structures, as nodes or users in social media networks have multiple attributions. For example, users can be friends and coworkers in a research group simultaneously. Hence, we propose a multilayer and attribute combined measure (MACM), a novel measurement based on the multilayer structure and common neighboring attributes, which includes the similarity measure between nodes and the importance measure for individual node in multilayer networks. Two improved community kernel detection algorithms based on MACM are subsequently proposed. They are the MA‐Greedy, which is based on the greedy algorithm, and the MA‐WeBA, which is a weighted balanced algorithm. The multilayer structure and attributes are comprehensively considered when calculating the similarity and importance of nodes in these strategies. Extensive experimental results on two public data sets demonstrate that the multilayer structure and attribute information can be used to enhance the precision of community detection. Changzheng Liu, Fengling Huang, Ruixuan Li 0001, Qi Yang 0009, Yuhua Li 0003, Shui Yu 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | Android malware obfuscation variants detection method based on multi-granularity opcode features
Junwei Tang, Ruixuan Li 0001, Xiwu Gu, Yuhua Li 0003 |
Future Gener. Comput. Syst. | 5 |
| 2022 | Gradient Scheduling With Global Momentum for Asynchronous Federated Learning in Edge EnvironmentabstractFederated Learning has attracted widespread attention in recent years because it allows massive edge nodes to collaboratively train machine learning models without sharing their private data sets. However, these edge nodes are usually heterogeneous in computational capability and statistically different in data distribution, i.e., non-independent and identically distributed (IID), leading to significant performance degradation. Although status quo asynchronous training methods can solve the heterogeneity issue, they cannot prevent the non-IID problem from reducing the convergence rate. In this article, we propose a novel paradigm that schedules the gradient with partially averaged gradients and applies the global momentum (GSGM) for asynchronous training over non-IID data sets in an edge environment. Our key idea is to apply global momentum and partial average on the biased gradients calculated on edge nodes after scheduling, to make the training process stable. Empirical results demonstrate that GSGM can well adapt to different degrees of non-IID data and bring 20% performance gains in terms of training stability for popular optimization algorithms with enhanced accuracy over Fashion-Mnist and CIFAR-10 data sets. Haozhao Wang, Ruixuan Li 0001, Pan Zhou 0001, Yuhua Li 0003, Wenchao Xu 0001, Song Guo 0001 |
IEEE Internet Things J. | 5 |
| 2022 | Trajectory planning for UAV navigation in dynamic environments with matrix alignment Dijkstra
Yuhua Li 0003, Ruixuan Li 0001, Kejing Chu |
Soft Comput. | 2 |
| 2021 | Chinese Administrative Penalty Event Extraction for Due Diligence in Financial Markets
Jun Wang 0018, Ruixuan Li 0001, Yuhua Li 0003 |
WISA | 5 |
| 2021 | Word Sense Disambiguation based on Sequence Topic Model using sense dependencyabstractWord sense disambiguation is a challenging task that aims to distinguish the correct sense for a target word. Unlike typical methods which only use the context information, we present the basic idea that incorporates the global sense distribution and contextual sense dependency to determine the correct sense. In this paper, we leverage the generative process of the sequence topic model to model the sense distribution from the whole global corpus directly. Since contextual sense dependency is more likely to happen between successive words, we hypothesize that the sense assignment of the target word depends on the sense of the previous word. Hence, the sense dependency could be modeled by the probabilistic topic model with hidden Markov chain assumption. Furthermore, the information in the sense inventory of WordNet is used as prior knowledge for obtaining the non-uniform sense distribution over words. It is notable that the prior knowledge contributes to parameter learning and inference during the Gibbs Sampling procedure, hence we call the proposed method a knowledge-based unsupervised approach. We evaluate the proposed method on Senseval-2, Senseval-3, SemEval-2007, SemEval-2013, and SemEval-2015 English All-Word WSD datasets and the experimental results show that although the performance of the proposed method is restricted by the sparse problem, it still achieves comparable performance compared with the state-of-the-art knowledge-based approaches. Qi Yang 0009, Ruixuan Li 0001, Yuhua Li 0003, Xiwu Gu |
IJCNN | 3 |
| 2021 | Multi-topical authority sensitive influence maximization with authority based graph pruning and three-stage heuristic optimization
Yuhua Li 0003, Ruixuan Li 0001, Xiaoqing Xiong, Xiwu Gu, Tianan Liang, Mingli Xu, Yumeng Yuan |
Appl. Intell. | 1 |
| 2020 | Fine-Grained Text Sentiment Transfer via Dependency Parsing
Lulu Xiao, Xiaoye Qu, Ruixuan Li 0001, Jun Wang 0018, Pan Zhou 0001, Yuhua Li 0003 |
ECAI | 6 |
| 2020 | A Multi-Task Learning Approach to Improve Sentiment Analysis with Explicit RecommendationabstractWhen expressing sentiment towards products, customers often explicitly indicate their recommendation status. Nevertheless, most existing literature focuses on sentiment analysis but neglects the rich correlation information that may be brought by explicit recommendation classification. We argue that the two tasks are correlated and hence, the knowledge in explicit recommendation classification can also be beneficial to sentiment analysis. Consequently, in this paper, a novel bidirectional encoder representations from transformers (BERT)-enhanced multi-task learning (BeMTL) approach is proposed to improve sentiment analysis with explicit recommendation classification. Specifically, the proposed MTL approach takes contextualized word embeddings produced by the pre-trained BERT-based embedding layer. Then, it learns the sentence contextual features shared between both tasks with a convolutional multi-head attention neural network. To fully exploit the correlation information between sentiment analysis and explicit recommendation classification tasks, a novel inter-task matching layer (IML) is designed to match their representations. In nutshell, our study reveals the potential of multi-task learning models on such types of problems, and experimental results on two Amazon datasets show that our approach outperforms the state-of-the-art baseline approaches for sentiment analysis. Olivier Habimana, Yuhua Li 0003, Ruixuan Li 0001, Xiwu Gu, Yuqi Peng |
IJCNN | 2 |
| 2020 | Sentiment analysis using deep learning approaches: an overview
Olivier Habimana, Yuhua Li 0003, Ruixuan Li 0001, Xiwu Gu, Ge Yu 0001 |
Sci. China Inf. Sci. | 2 |
| 2020 | Adversarial joint domain adaptation of asymmetric feature mapping based on least squares distance
Yumeng Yuan, Yuhua Li 0003, Zhenlong Zhu, Ruixuan Li 0001, Xiwu Gu |
Pattern Recognit. Lett. | 2 |
| 2020 | AMNN: Attention-Based Multimodal Neural Network Model for Hashtag RecommendationabstractIn the real-world social networks, hashtags are widely applied for understanding the content of an individual microblog. However, users do not always take the initiative in attaching hashtags when posting a microblog so that much effort has been invested for automatically hashtag recommendation. As a new trend, users no longer only post texts but prefer to share with multimodal data, such as images. To deal with these situations, we propose an attention-based multimodal neural network model (AMNN) to learn the representations of multimodal microblogs and recommend relevant hashtags. In this article, we convert the hashtag recommendation task into a sequence generation problem. Then, we propose a hybrid neural network approach to extract the features of both texts and images and incorporate them into the sequence-to-sequence model for hashtag recommendation. Experimental results on the data set collected on Instagram and two public data sets demonstrate that the proposed method outperforms state-of-the-art methods. Our model achieves the best performance in three different metrics: precision, recall, and accuracy. The source code of this article can be obtained from “https://github.com/w5688414/AMNN.” Qi Yang 0009, Gaosheng Wu, Yuhua Li 0003, Ruixuan Li 0001, Xiwu Gu, Huicai Deng, Junzhuang Wu |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2020 | HeteroYARN: A Heterogeneous FPGA-Accelerated Architecture Based on YARNabstractIn recent years, the heterogeneous distributed platform integrating with FPGAs to accelerate computation tasks has been widely studied to deal with the deluge of data. However, most of current works suffer from poor universality and low resource utilization that run specific algorithms with the highly customized structure. Moreover, there are still many challenges, such as data curation, task scheduling, and resource management, which further limit the scalability of a CPU-FPGA distributed platform. In this paper, we present HeteroYARN, an FPGA-accelerated heterogeneous architecture based on YARN platform, which provides resource management and programming support for computing-intensive applications using FPGAs. In particular, the HeteroYARN abstracts FPGA accelerators as general resources and provides programming APIs to utilize those accelerators easily. Our HeteroYARN simplifies the request and usage of FPGA resources to enhance the efficiency of the heterogeneous framework while maintaining previous workflow unchanged. Experimental results using two representative algorithms, K-means and Naive Bayes classifier, which are accelerated by FPGAs, demonstrate the usability of the HeteroYARN framework and show performance speedup improvement by 7.5x (K-means) and 2.3x (Naive Bayes) respectively compared to conventional CPU-only applications provided by Mahout. Ruixuan Li 0001, Qi Yang 0009, Yuhua Li 0003, Xiwu Gu, Weijun Xiao, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Imbalance Rectification in Deep Logistic Regression for Multi-Label Image Classification Using Random Noise SamplesabstractLogistic regression (LR) is the most commonly used loss function in multi-label image classification. However, it suffers from class imbalance problem caused by the huge difference in quantity between positive and negative samples as well as between different classes. First, we find that feeding randomly generated noise samples into an LR classifier is an effective way to detect class imbalances, and further define an informative imbalance metric named inference tendency based on noise sample analysis. Second, we design an efficient moving average based method for calculating inference tendency, which can be easily done during training with negligible overhead. Third, two novel rectification methods called extremum shift (ES) and tendency constraint (TC) are designed to offset or constrain inference tendency in the loss function, and mitigate class imbalances significantly. Finally, comparative experiments with Resnet on Microsoft COCO, NUS-WIDE and DeepFashion demonstrate the effectiveness of inference tendency and the superiority of our approach over the baseline LR and several state-of-the-art alternatives. Wenjin Yan, Ruixuan Li 0001, Jun Wang 0018, Yuhua Li 0003, Pan Zhou 0001, Xiwu Gu |
CIKM | 4 |
| 2019 | ST-RNet: A Time-aware Point-of-interest Recommendation Method based on Neural NetworkabstractPoint-of-interest (POI) recommendation is one of the most important services in the rapid growing location-based social networks (LBSNs). Good POI recommendation can help people explore the locations they haven't visited but are interested in, and help merchants find their target users. Time-aware POI recommendation aims to recommend unvisited POIs for a given user at a specified time in a day. However, previous methods, such as user-based collaborative filtering, lack the mining of the features of POIs and the learning of abstract spatio-temporal interactions. In this paper, we propose a novel time-aware POI recommendation method named ST-RNet (Spatio-Temporal Recommender Network) to address these shortages. ST-RNet works in the following fashion. Firstly, we analyze the crucial features in LBSNs to alleviate data sparsity problem and further measure the similarities between POIs. For subsequent network training, we then construct the embedding matrices with same dimension for users and POIs by POI-based Collaborative Filtering (PCF). Furthermore, the positive and negative check-in records are fed into a novel recommender neural network (RNet) to learn the embedding matrix of times and the abstract interactions between users, POIs and times. Finally, ST-RNet recommends the unvisited POIs most likely to be visited to a given user at a given time. The experimental results on Foursquare real-world dataset show that ST-RNet is effective on time-aware POI recommendation task and is capable of analyzing the hidden patterns behind spatio-temporal interactions. Yuhua Li 0003, Ruixuan Li 0001, Zhenlong Zhu, Xiwu Gu, Olivier Habimana |
IJCNN | 2 |
| 2019 | Personalizing Session-based Recommendation with Dual Attentive Neural NetworkabstractSession-based recommendation which aims to recommend the next item in an anonymous session for users, becomes one of the most popular tasks in recommendation area. Traditional methods such as matrix factorization and item-to-item perform very poorly because they only take into account the last click of the session and ignore the information of the whole click sequence. On the other hand, Recurrent Neural Network (RNN) based methods have performed excellently in session-based recommendation. However, they only consider the user's sequential behavior in the current session or only use cross-session information to track user's interests over time, whereas the user preference is not emphasized in cross-session. Therefore, in this work, we design a novel neural network framework for personalized session-based recommendation, named Dual Attentive Neural Network (DANN). DANN considers user's main purpose of current session and user's personalized preference of cross-session. Specifically, in DANN we exploit a user-level attention mechanism to model user's personalized preference and capture user's main purpose in the current session via a session-level attention mechanism. The experimental results on two real-world datasets show that our DANN model outperforms other baseline models. Furthermore, we find that DANN achieves improvement when modeling user personalized preferences, which shows the advantage of modeling user's preference and user's purpose simultaneously. Tianan Liang, Yuhua Li 0003, Ruixuan Li 0001, Xiwu Gu, Olivier Habimana |
IJCNN | 2 |
| 2019 | TDP: Personalized Taxi Demand Prediction Based on Heterogeneous Graph EmbeddingabstractPredicting users' irregular trips in a short term period is one of the crucial tasks in the intelligent transportation system. With the prediction, the taxi requesting services, such as Didi Chuxing in China, can manage the transportation resources to offer better services. There are several different transportation scenes, such as commuting scene and entertainment scene. The origin and the destination of entertainment scene are more unsure than that of commuting scene, so both origin and destination should be predicted. Moreover, users' trips on Didi platform is only a part of their real life, so these transportation data are only few weak samples. To address these challenges, in this paper, we propose Taxi Demand Prediction (TDP) model in challenging entertainment scene based on heterogeneous graph embedding and deep neural predicting network. TDP aims to predict next possible trip edges that have not appeared in historical data for each user in entertainment scene. Experimental results on the real-world dataset show that TDP achieves significant improvements over the state-of-the-art methods. Zhenlong Zhu, Ruixuan Li 0001, Minghui Shan, Yuhua Li 0003, Jixing Xu, Xiwu Gu |
SIGIR | 4 |
| 2019 | EADP: An extended adaptive density peaks clustering for overlapping community detection in social networks
Mingli Xu, Yuhua Li 0003, Ruixuan Li 0001, Fuhao Zou, Xiwu Gu |
Neurocomputing | 2 |
| 2018 | Hashtag Recommendation with Attention-Based Neural Image Hashtagging Network
Gaosheng Wu, Yuhua Li 0003, Wenjin Yan, Ruixuan Li 0001, Xiwu Gu, Qi Yang 0009 |
ICONIP (2) | 2 |
| 2018 | Self-inhibition Residual Convolutional Networks for Chinese Sentence Classification
Mengting Xiong, Ruixuan Li 0001, Yuhua Li 0003, Qi Yang 0009 |
ICONIP (1) | 3 |
| 2018 | Topic-Bigram Enhanced Word Embedding Model
Qi Yang 0009, Ruixuan Li 0001, Yuhua Li 0003, Qilei Liu |
ICONIP (3) | 3 |
| 2018 | Stock Price Prediction Using Time Convolution Long Short-Term Memory Network
Xukuan Zhan, Yuhua Li 0003, Ruixuan Li 0001, Xiwu Gu, Olivier Habimana, Haozhao Wang |
KSEM (1) | 2 |
| 2018 | Distant Domain Adaptation for Text Classification
Zhenlong Zhu, Yuhua Li 0003, Ruixuan Li 0001, Xiwu Gu |
KSEM (1) | 2 |
| 2018 | Topical Authority-Sensitive Influence Maximization
Xiaoqing Xiong, Ruixuan Li 0001, Yuhua Li 0003, Xiwu Gu, Tianan Liang |
WISE (1) | 3 |
| 2018 | FRFB: Top-k Followee Recommendation by exploring the Following Behaviors in social networksabstractSummary As social networks such as micro‐blogging sites rapidly grow, deciding whom to follow (followee recommendation) becomes a significantly important problem. Most existing works exclusively rely on two traditional factors: the proximity between two users in the network topology or the similarity of the user‐generated contents in the social network, disregarding the effect of users' following behaviors. The challenge of how to effectively combine these two factors remains largely open. Moreover, most research studies simply sort the scores to find top‐k users, which is time‐consuming, especially for large‐scale networks. In this paper, we propose the idea that “predict users' following behaviors by following behaviors themselves.” We consider a user's following to others as a normal process of dynamic and coherent behavior, and we model the potential propagation of the users' following behaviors. Furthermore, based on our previous research on top‐k selection problem, we propose an effective top‐k followee recommendation algorithm, called FRFB. FRFB has low complexity and high scalability and, moreover, good adaptability to real‐life dynamic social networks. We conduct extensive experiments, with two real social network data sets (Wiki and Twitter), which show that FRFB outperforms the well‐known topology‐based followee recommendation algorithms. Zhengyuan Xue, Ruixuan Li 0001, Yuhua Li 0003, Xiwu Gu, Weijun Xiao |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Role mining using answer set programming
Ruixuan Li 0001, Xiwu Gu, Yuhua Li 0003, Kunmei Wen |
Future Gener. Comput. Syst. | 4 |
| 2015 | Subtopic-Level Sentiment Analysis of EmergenciesabstractWith the rapid development of microblog, millions of Internet users share their opinions on different aspects of daily life. By analyzing and monitoring sentiment information extracting from tweets related to an important event, we are able to gain insights into variation trends of users’ sentiment. In this paper, we focus on extracting public sentiment of microblog emergencies. A subtopic-level opinion mining method is proposed based on two-phase optimization. Different subtopics of emergencies are extracted based on retweets. Opinion tweets are classified to different subtopics. The sentiment score of opinion holders is calculated. The above results are optimized based on users and endorsement interactions between users. Experimental results validate the effectiveness of the proposed method. Kunmei Wen, Zhijiang Liu, Ruixuan Li 0001, Yuhua Li 0003, Xiwu Gu, Jie Zan |
KSEM | 5 |
| 2015 | Role mining based on cardinality constraintsabstractSummary Role mining was recently proposed to automatically find roles among user‐permission assignments using data mining technologies. However, the current studies about role mining mainly focus on how to find roles, without considering the constraints that are essentially required in role‐based access control systems. In this paper, we present a role mining algorithm with constraints, especially for the cardinality constraints. We illustrate it is essential for role mining to take cardinality constraints into account, and introduce the concepts of the cardinality constraints of roles and permissions. We further propose a role mining algorithm to generate roles based on these two kinds of cardinality constraints. The algorithm uses graph theory to model the role mining problem and maps the relation of two roles to the relation of graph elements. We set an optimization goal for role mining and employ graph optimization theory to find roles that satisfy the aforementioned cardinality constraints. We carry out the experiments to evaluate our approach. The experimental results demonstrate the rationality and effectiveness of the proposed algorithm. Copyright © 2015 John Wiley & Sons, Ltd. Ruixuan Li 0001, Huaqing Li 0001, Xiwu Gu, Yuhua Li 0003, Xiaopu Ma |
Concurr. Comput. Pract. Exp. | 4 |
| 2014 | LIMTopic: A Framework of Incorporating Link Based Importance into Topic ModelingabstractTopic modeling has become a widely used tool for document management. However, there are few topic models distinguishing the importance of documents on different topics. In this paper, we propose a framework LIMTopic to incorporate link based importance into topic modeling. To instantiate the framework, RankTopic and HITSTopic are proposed by incorporating topical pagerank and topical HITS into topic modeling respectively. Specifically, ranking methods are first used to compute the topical importance of documents. Then, a generalized relation is built between link importance and topic modeling. We empirically show that LIMTopic converges after a small number of iterations in most experimental settings. The necessity of incorporating link importance into topic modeling is justified based on KL-Divergences between topic distributions converted from topical link importance and those computed by basic topic models. To investigate the document network summarization performance of topic models, we propose a novel measure called log-likelihood of ranking-integrated document-word matrix. Extensive experimental results show that LIMTopic performs better than baseline models in generalization performance, document clustering and classification, topic interpretability and document network summarization performance. Moreover, RankTopic has comparable performance with relational topic model (RTM) and HITSTopic performs much better than baseline models in document clustering and classification. Dongsheng Duan, Yuhua Li 0003, Ruixuan Li 0001, Rui Zhang 0003, Xiwu Gu, Kunmei Wen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Keyword-Matched Data Skyline in Peer-to-Peer Systems
Khaled M. Banafaa, Ruixuan Li 0001, Kunmei Wen, Xiwu Gu, Yuhua Li 0003 |
DASFAA (1) | 5 |
| 2013 | MEI: Mutual Enhanced Infinite Community-Topic Model for Analyzing Text-Augmented Social NetworksabstractThe community and the topic can help summarize the text-augmented social networks. Existing works mixed up the community and the topic by regarding them as the same. However, there is an inherent difference between the community and the topic such that considering them as the same is not so flexible. We propose a mutual enhanced infinite (MEI) community–topic model to detect communities and topics simultaneously in text-augmented social networks. The community and the topic are correlated via a community–topic distribution. The mutual enhancement effect between the community and the topic is validated by introducing two novel measures perplexity with community (perplexityc) and mean of the rank (MRK) with topic (MRKt). To determine the numbers of communities and topics automatically, the Dirichlet Process Mixture (DPM) model and the Hierarchical Dirichlet Process mixture (HDP) model are used to model the community and the topic, respectively. We further introduce parameters to model the weight of the community and the topic responsible for community inference. Experiments on the co-author network built from a subset of DBLP data show that MEI outperforms the baseline models in terms of the generalization performance. Parameter study shows that MEI is averagely improved by 3.7 and 15.5% in perplexityc and MRKt, respectively, by setting a low weight for the topic. We also experimentally validate that MEI can determine the appropriate numbers of communities and topics. Dongsheng Duan, Yuhua Li 0003, Ruixuan Li 0001, Zhengding Lu, Aiming Wen |
Comput. J. | 2 |
| 2012 | ℓ1-Graph Based Community Detection in Online Social Networks
Ruixuan Li 0001, Yuhua Li 0003, Xiwu Gu, Kunmei Wen, Zhiyong Xu 0003 |
APWeb | 3 |
| 2012 | RankTopic: Ranking Based Topic ModelingabstractTopic modeling has become a widely used tool for document management due to its superior performance. However, there are few topic models distinguishing the importance of documents on different topics. In this paper, we investigate how to utilize the importance of documents to improve topic modeling and propose to incorporate link based ranking into topic modeling. Specifically, topical pagerank is used to compute the topic level ranking of documents, which indicates the importance of documents on different topics. By retreating the topical ranking of a document as the probability of the document involved in corresponding topic, a generalized relation is built between ranking and topic modeling. Based on the relation, a ranking based topic model Rank Topic is proposed. With Rank Topic, a mutual enhancement framework is established between ranking and topic modeling. Extensive experiments on paper citation data and Twitter data are conducted to compare the performance of Rank Topic with that of some state-of-the-art topic models. Experimental results show that Rank Topic performs much better than some baseline models and is comparable with the state-of-the-art link combined relational topic model (RTM) in generalization performance, document clustering and classification by setting a proper balancing parameter. It is also demonstrated in both quantitative and qualitative ways that topics detected by Rank Topic are more interpretable than those detected by some baseline models and still competitive with RTM. Dongsheng Duan, Yuhua Li 0003, Ruixuan Li 0001, Rui Zhang 0003, Aiming Wen |
ICDM | 2 |
| 2011 | MEI: Mutual Enhanced Infinite Generative Model for Simultaneous Community and Topic Detection
Dongsheng Duan, Yuhua Li 0003, Ruixuan Li 0001, Zhengding Lu, Aiming Wen |
Discovery Science | 2 |
| 2011 | Incorporating User Feedback into Name Disambiguation of Scientific Cooperation Network
Yuhua Li 0003, Aiming Wen, Quan Lin, Ruixuan Li 0001, Zhengding Lu |
WAIM | 1 |
| 2011 | Type-2 fuzzy description logic
Ruixuan Li 0001, Kunmei Wen, Xiwu Gu, Yuhua Li 0003, Xiaolin Sun 0001, Bing Li 0010 |
Frontiers Comput. Sci. China | 4 |
| 2010 | TGP: Mining Top-K Frequent Closed Graph Pattern without Minimum Support
Yuhua Li 0003, Quan Lin, Ruixuan Li 0001, Dongsheng Duan |
ADMA (1) | 1 |
| 2010 | Semantic Grounding of Hybridization for Tag Recommendation
Yanan Jin, Ruixuan Li 0001, Yi Cai 0001, Qing Li 0001, Ali Daud, Yuhua Li 0003 |
WAIM | 6 |