Hong Yu 0005

dblp:55/6749-5 · DBLP profile ↗
← Back
49ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0003-4807-1812ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 6 since 2021Security and privacy · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 RUQuant: Towards Refining Uniform Quantization for Large Language Models
abstract
The increasing size and complexity of large language models (LLMs) have raised significant challenges in deployment efficiency, particularly under resource constraints. Post-training quantization (PTQ) has emerged as a practical solution by compressing models without requiring retraining. While existing methods focus on uniform quantization schemes for both weights and activations, they often suffer from substantial accuracy degradation due to the non-uniform nature of activation distributions. In this work, we revisit the activation quantization problem from a theoretical perspective grounded in the Lloyd-Max optimality conditions. We identify the core issue as the non-uniform distribution of activations within the quantization interval, which causes the optimal quantization point under the Lloyd-Max criterion to shift away from the midpoint of the interval. To address this issue, we propose a two-stage orthogonal transformation method, RUQuant. In the first stage, activations are divided into blocks. Each block is mapped to uniformly sampled target vectors using composite orthogonal matrices, which are constructed from Householder reflections and Givens rotations. In the second stage, a global Householder reflection is fine-tuned to further minimize quantization error using Transformer output discrepancies. Empirical results show that our method achieves near-optimal quantization performance without requiring model fine-tuning: RUQuant achieves 99.8% of full-precision accuracy with W6A6 and 97% with W4A4 quantization for a 13B LLM, within approximately one minute. A fine-tuned variant yields even higher accuracy, demonstrating the effectiveness and scalability of our approach.
Han Liu 0008, Changya Li, Feng Zhang 0027, Xiaotong Zhang 0003, Wei Wang 0077, Hong Yu 0005
KDD (1)7
2026 SEP-Attack: A Simple and Effective Paradigm for Transfer-Based Textual Adversarial Attack
Han Liu 0008, Zhi Xu 0008, Xiaotong Zhang 0003, Feng Zhang 0027, Xiaoming Xu 0003, Wei Wang 0077, Fenglong Ma, Hong Yu 0005
WWW8
2025 Multi-Label Few-Shot Image Classification via Pairwise Feature Augmentation and Flexible Prompt Learning
abstract
Multi-label few-shot image classification is a crucial and challenging task due to limited annotated data and elusive category specificity. However, research on this topic is still in the rudimentary stage and few methods are available. Existing methods either leverage data augmentation to alleviate data scarcity or utilize label features as auxiliary knowledge to eliminate the negative effect caused by irrelevant categories, but they ignore the utilization of image region features for data augmentation, and overlook to learn appropriate text feature to better match the image features of specific categories. Moreover, these methods only focus on one side and do not effectively tackle the above two issues simultaneously. In this paper, we introduce a novel prototype-based multi-label few-shot learning framework that seamlessly integrates pairwise feature augmentation and flexible prompt learning. Specifically, by pairwise feature augmentation, we leverage the region features of images in the support set to generate more image features and construct image prototypes, thus alleviating the issue of data scarcity. By flexible prompt learning, we adaptively acquire class-specific prompts to build text prototypes that highly match the image features of specific classes, thereby mitigating the impact of irrelevant classes. Finally, with adaptive learnable parameters, we merge image and text prototypes to obtain the final prototypes, achieving a more powerful classifier for multi-label few-shot image classification. Extensive experimental results demonstrate that our proposed method can push the performance to a higher level.
Han Liu 0008, Xiaotong Zhang 0003, Feng Zhang 0027, Wei Wang 0077, Fenglong Ma, Hong Yu 0005
AAAI7
2025 AdaDHP: Fine-Grained Fine-Tuning via Dual Hadamard Product and Adaptive Parameter Selection
abstract
With the continuously expanding parameters, efficiently adapting large language models to downstream tasks is crucial in resource-limited conditions.Many parameter-efficient finetuning methods have emerged to address this challenge.However, they lack flexibility, like LoRA requires manually selecting trainable parameters and rank size, (IA) 3 can only scale the activations along columns, yielding inferior results due to less precise fine-tuning.To address these issues, we propose a novel method named AdaDHP with fewer parameters and finer granularity, which can adaptively select important parameters for each task.Specifically, we introduce two trainable vectors for each parameter and fine-tune the parameters through Hadamard product along both rows and columns.This significantly reduces the number of trainable parameters, with our parameter count capped at the lower limit of LoRA.Moreover, we design an adaptive parameter selection strategy to select important parameters for downstream tasks dynamically.This allows our method to flexibly remove unimportant parameters for downstream tasks.Finally, we demonstrate the superiority of our method on the T5-base model across 17 NLU tasks and on complex mathematical tasks with the Llama series models.
Han Liu 0008, Changya Li, Xiaotong Zhang 0003, Feng Zhang 0027, Fenglong Ma, Wei Wang 0077, Hong Yu 0005
ACL (1)7
2025 Non-Autoregressive Image Captioning with Multi-Label Classification and Self-Critical Sequence Training
abstract
Most current image captioning models rely on the autoregressive approach, which unfortunately results in significant inference delays that hinder their practical use. In contrast, non-autoregressive methods show promising potential for increasing inference speeds. However, there is often a performance gap between the non-autoregressive and autoregressive image captioning models due to issues like word repetition and semantic inconsistencies in the generated captions. Autoregressive models benefit from self-critical sequence training, which helps produce more coherent and fluid captions. While non-autoregressive models are difficult to benefit from as they predict words independently. In this paper, we introduce a two-stage training strategy designed to harness self-critical sequence training for enhancing the non-autoregressive image captioning model. Our approach initially treats the image captioning task as a multi-label classification problem, which allows for the stable production of multiple candidate captions. In the second stage, we employ these candidate captions to compute sequence-level evaluation metric scores that serve as reward scores for self-critical sequence training. Extensive experiments demonstrate the effectiveness of our proposed method and show that our model achieves a new state-of-the-art performance in inference accuracy and speed.
Yuanqiu Liu, Hong Yu 0005, Xiaotong Zhang 0003, Han Liu 0008
ICASSP2
2025 SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
abstract
Large language models (LLMs) have shown remarkable performance in various domains, but they are constrained by massive computational and storage costs. Quantization, an effective technique for compressing models to fit resource-limited devices while preserving generative quality, encompasses two primary methods: quantization aware training (QAT) and post-training quantization (PTQ). QAT involves additional retraining or fine-tuning, thus inevitably resulting in high training cost and making it unsuitable for LLMs. Consequently, PTQ has become the research hotspot in recent quantization methods. However, existing PTQ methods usually rely on various complex computation procedures and suffer from considerable performance degradation under low-bit quantization settings. To alleviate the above issues, we propose a simple and effective post-training quantization paradigm for LLMs, named SEPTQ. Specifically, SEPTQ first calculates the importance score for each element in the weight matrix and determines the quantization locations in a static global manner. Then it utilizes the mask matrix which represents the important locations to quantize and update the associated weights column-by-column until the appropriate quantized weight matrix is obtained. Compared with previous methods, SEPTQ simplifies the post-training quantization procedure into only two steps, and considers the effectiveness and efficiency simultaneously. Experimental results on various datasets across a suite of models ranging from millions to billions in different quantization bit-levels demonstrate that SEPTQ significantly outperforms other strong baselines, especially in low-bit quantization scenarios.
Han Liu 0008, Xiaotong Zhang 0003, Changya Li, Feng Zhang 0027, Wei Wang 0077, Fenglong Ma, Hong Yu 0005
KDD (1)8
2025 HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained Models
abstract
Black-box adversarial attack on vision-language pre-trained models is a practical and challenging task, as text and image perturbations need to be considered simultaneously, and only the predicted results are accessible. Research on this problem is in its infancy, and only a handful of methods are available. Nevertheless, existing methods either rely on a complex iterative cross-search strategy, which inevitably consumes numerous queries, or only consider reducing the similarity of positive image-text pairs but ignore that of negative ones, which will also be implicitly diminished, thus inevitably affecting the attack performance. To alleviate the above issues, we propose a simple yet effective framework to generate high-quality adversarial examples on vision-language pre-trained models, named HQA-VLAttack, which consists of text and image attack stages. For text perturbation generation, it leverages the counter-fitting word vector to generate the substitute word set, thus guaranteeing the semantic consistency between the substitute word and the original word. For image perturbation generation, it first initializes the image adversarial example via the layer-importance guided strategy, and then utilizes contrastive learning to optimize the image adversarial perturbation, which ensures that the similarity of positive image-text pairs is decreased while that of negative image-text pairs is increased. In this way, the optimized adversarial images and texts are more likely to retrieve negative examples, thereby enhancing the attack success rate. Experimental results on three benchmark datasets demonstrate that HQA-VLAttack significantly outperforms strong baselines in terms of attack success rate.
Han Liu 0008, Zhi Xu 0008, Xiaotong Zhang 0003, Xiaoming Xu 0003, Fenglong Ma, Yuanman Li, Hong Yu 0005
NeurIPS8
2025 Transformer-Based Nonautoregressive Image Captioning via Guided Keyword Generation and Learnable Positional Encoding for IoT Devices
abstract
The emergence of the Intelligent Internet of Things (IIoT) has brought data processing closer to data sources, especially for real-time processing of surveillance video and image analysis. Image captioning plays a crucial role in understanding images. However, the Transformer architecture, which has become prevalent in recent applications, has been observed to increase the computational resources required for image captioning models. Conventionally, most existing methods use the autoregressive paradigm, which reduces their computational efficiency on edge devices and results in significant inference delays. In this paper, we use non-autoregressive paradigms to improve its inference speed and model efficiency. Nevertheless, the lack of effective inputs results in a performance gap between non-autoregressive and autoregressive models. To bridge this gap, we propose the learnable positional encoding and keyword guided non-autoregressive image captioning. Firstly, a diffusion model guided by image features is employed to generate keywords that accurately reflect the image content, thereby infusing a substantial amount of semantic information into the non-autoregressive decoder. Secondly, positional encoding is utilized to guide the decoder in generating appropriate words at the correct positions within the caption. Extensive experiments on widely used benchmarks demonstrate that our model achieves state-of-the-art performance in non-autoregressive image captioning. Furthermore, our model maintains a competitive inference speed.
Yuanqiu Liu, Hong Yu 0005, Xiaotong Zhang 0003, Han Liu 0008
IEEE Internet Things J.2
2024 Depression Detection via Capsule Networks with Contrastive Learning
abstract
Depression detection is a challenging and crucial task in psychological illness diagnosis. Utilizing online user posts to predict whether a user suffers from depression seems an effective and promising direction. However, existing methods suffer from either poor interpretability brought by the black-box models or underwhelming performance caused by the completely separate two-stage model structure. To alleviate these limitations, we propose a novel capsule network integrated with contrastive learning for depression detection (DeCapsNet). The highlights of DeCapsNet can be summarized as follows. First, it extracts symptom capsules from user posts by leveraging meticulously designed symptom descriptions, and then distills them into class-indicative depression capsules. The overall workflow is in an explicit hierarchical reasoning manner and can be well interpreted by the Patient Health Questionnaire-9 (PHQ9), which is one of the most widely adopted questionnaires for depression diagnosis. Second, it integrates with contrastive learning, which can facilitate the embeddings from the same class to be pulled closer, while simultaneously pushing the embeddings from different classes apart. In addition, by adopting the end-to-end training strategy, it does not necessitate additional data annotation, and mitigates the potential adverse effects from the upstream task to the downstream task. Extensive experiments on three widely-used datasets show that in both within-dataset and cross-dataset scenarios our proposed method outperforms other strong baselines significantly.
Han Liu 0008, Changya Li, Xiaotong Zhang 0003, Feng Zhang 0027, Wei Wang 0077, Fenglong Ma, Hongyang Chen 0001, Hong Yu 0005, Xianchao Zhang 0001
AAAI8
2024 Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification Reframing
abstract
Few-shot and zero-shot text classification aim to recognize samples from novel classes with limited labeled samples or no labeled samples at all. While prevailing methods have shown promising performance via transferring knowledge from seen classes to unseen classes, they are still limited by (1) Inherent dissimilarities among classes make the transformation of features learned from seen classes to unseen classes both difficult and inefficient. (2) Rare labeled novel samples usually cannot provide enough supervision signals to enable the model to adjust from the source distribution to the target distribution, especially for complicated scenarios. To alleviate the above issues, we propose a simple and effective strategy for few-shot and zero-shot text classification. We aim to liberate the model from the confines of seen classes, thereby enabling it to predict unseen categories without the necessity of training on seen classes. Specifically, for mining more related unseen category knowledge, we utilize a large pre-trained language model to generate pseudo novel samples, and select the most representative ones as category anchors. After that, we convert the multi-class classification task into a binary classification task and use the similarities of query-anchor pairs for prediction to fully leverage the limited supervision signals. Extensive experiments on six widely used public datasets show that our proposed method can outperform other strong baselines significantly in few-shot and zero-shot tasks, even without using any seen class samples.
Han Liu 0008, Siyang Zhao, Xiaotong Zhang 0003, Feng Zhang 0027, Wei Wang 0077, Fenglong Ma, Hongyang Chen 0001, Hong Yu 0005, Xianchao Zhang 0001
AAAI8
2024 Alignment-Enhanced Network for Temporal Language Grounding in Videos
Hong Yu 0005, Yuanqiu Liu, Han Liu 0008
ICANN (3)1
2024 Boosting Node Injection Attack with Graph Local Sparsity
abstract
Graph neural networks have achieved tremendous success in various tasks over the past decade. However, recent studies have shown their vulnerabilities to well-designed adversarial attacks, in which even tiny perturbations can lead to the misclassification of the model. In this paper, we investigate the global evasion injection attack, aiming to degrade model performance on test nodes by injecting additional nodes into the graph. We propose the Graph Local Sparsity Attack(GLSA), where we introduce the concept of graph local sparsity, defining a node’s vulnerability by incorporating its neighborhood. Our method pre-computes scores to select vulnerable nodes using the sparsity and the gradual consistency condition. Subsequently, an iterative selection strategy is employed to inject the nodes. Finally, scores are dynamically updated, nodes are injected and features are finely generated to effectively degrade model performance. Experiments on several large-scale datasets demonstrate the effectiveness of our proposed method compared with the state-of-the-art approaches.
Wenxin Liang, Bingkai Liu, Han Liu 0008, Hong Yu 0005
ICME4
2024 Dual Branch Non-Autoregressive Image Captioning
Yuanqiu Liu, Hong Yu 0005, Han Liu 0008
ICPR (18)2
2024 Correlation-Guided Semantic Consistency Network for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) has raised more attention in night-time surveillance applications due to the struggle to capture valid appearance information under poor illumination conditions via visible cameras. Existing works usually separate the modality-specific and modality-irrelevant information in visible and infrared features, or project features of two modalities into a unified embedding feature space directly, which aims to eliminate huge modality discrepancies. However, these methods neglect the intra-modality and inter-modality correlations. We argue that the correlations can implicitly guide the network to discover the modality-irrelevant information, thus more beneficial for eliminating huge modality discrepancies and preserving individual differences. To this end, we propose a novel framework, termed as correlation-guided semantic consistency network (CSC-Net), to explore and exploit the intra-modality and inter-modality correlations. Specifically, CSC-Net consists of a cross-modality semantic alignment (CSA) module, a cross-granularity discrepancy awareness (CDA) module, and a probability consistency constraint (PCC) module. CSA mines the inter-modality correlation by calculating the semantic similarity between modalities to explore modality-irrelevant features, and then transfers the learned features to the backbone network to face the input of only single modality images. To preserve the individual differences, CDA sufficiently utilizes the intra-modality correlation via exploring the multi-granularity discriminative information. Finally, PCC constrains the network at the probability level, cooperating with the CSA which constrains at the feature level, to further alleviate the modality discrepancy. Extensive experiments on two public VI-ReID datasets SYSU-MM01 and RegDB have verified the effectiveness of our approach.
Qijie Peng, Shijie Wang 0003, Hong Yu 0005, Zhihui Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 SSPAttack: A Simple and Sweet Paradigm for Black-Box Hard-Label Textual Adversarial Attack
abstract
Hard-label textual adversarial attack is a challenging task, as only the predicted label information is available, and the text space is discrete and non-differentiable. Relevant research work is still in fancy and just a handful of methods are proposed. However, existing methods suffer from either the high complexity of genetic algorithms or inaccurate gradient estimation, thus are arduous to obtain adversarial examples with high semantic similarity and low perturbation rate under the tight-budget scenario. In this paper, we propose a simple and sweet paradigm for hard-label textual adversarial attack, named SSPAttack. Specifically, SSPAttack first utilizes initialization to generate an adversarial example, and removes unnecessary replacement words to reduce the number of changed words. Then it determines the replacement order and searches for an anchor synonym, thus avoiding going through all the synonyms. Finally, it pushes substitution words towards original words until an appropriate adversarial example is obtained. The core idea of SSPAttack is just swapping words whose mechanism is simple. Experimental results on eight benchmark datasets and two real-world APIs have shown that the performance of SSPAttack is sweet in terms of similarity, perturbation rate and query efficiency.
Han Liu 0008, Zhi Xu 0008, Xiaotong Zhang 0003, Xiaoming Xu 0003, Feng Zhang 0027, Fenglong Ma, Hongyang Chen 0001, Hong Yu 0005, Xianchao Zhang 0001
AAAI8
2023 Boosting Few-Shot Text Classification via Distribution Estimation
abstract
Distribution estimation has been demonstrated as one of the most effective approaches in dealing with few-shot image classification, as the low-level patterns and underlying representations can be easily transferred across different tasks in computer vision domain. However, directly applying this approach to few-shot text classification is challenging, since leveraging the statistics of known classes with sufficient samples to calibrate the distributions of novel classes may cause negative effects due to serious category difference in text domain. To alleviate this issue, we propose two simple yet effective strategies to estimate the distributions of the novel classes by utilizing unlabeled query samples, thus avoiding the potential negative transfer issue. Specifically, we first assume a class or sample follows the Gaussian distribution, and use the original support set and the nearest few query samples to estimate the corresponding mean and covariance. Then, we augment the labeled samples by sampling from the estimated distribution, which can provide sufficient supervision for training the classification model. Extensive experiments on eight few-shot text classification datasets show that the proposed method outperforms state-of-the-art baselines significantly.
Han Liu 0008, Feng Zhang 0027, Xiaotong Zhang 0003, Siyang Zhao, Fenglong Ma, Xiao-Ming Wu 0003, Hongyang Chen 0001, Hong Yu 0005, Xianchao Zhang 0001
AAAI8
2023 Boosting Meta-Learning Cold-Start Recommendation with Graph Neural Network
abstract
Meta-learning methods have shown to be effective in dealing with cold-start recommendation. However, most previous methods rely on an ideal assumption that there exists a similar data distribution between source and target tasks, which are unsuitable for the scenario that only extremely limited number of new user or item interactions are available. In this paper, we propose to boost meta-learning cold-start recommendation with graph neural network (MeGNN). First, it utilizes the global neighborhood translation learning to obtain consistent potential interactions for all new user and item nodes, which can refine their representations. Second, it employs the local neighborhood translation learning to predict specific potential interactions for each node, thus guaranteeing the personalized requirement. In experiments, we combine MeGNN with two representative meta-learning models MeLU and TaNP. Extensive results on two widely-used datasets show the superiority of MeGNN in four different scenarios.
Han Liu 0008, Hongxiang Lin, Xiaotong Zhang 0003, Fenglong Ma, Hongyang Chen 0001, Lei Wang 0005, Hong Yu 0005, Xianchao Zhang 0001
CIKM7
2023 Boosting Visual Question Answering Through Geometric Perception and Region Features
abstract
Visual question answering (VQA) is a crucial yet challenging task in multimodal understanding. To correctly answer questions about an image, VQA models are required to comprehend the fine-grained semantics of both the image and the question. Recent advances have shown that both grid and region features contribute to improving the VQA performance, while grid features surprisingly outperform region features. However, grid features will inevitably induce visual semantic noise due to fine granularity. Besides, the ignorance of geometric relationship makes VQA models difficult to understand the object relative positions in the image and answer questions accurately. In this paper, we propose a visual enhancement network for VQA that leverages region features and position information to enhance grid features, thus generating richer visual grid semantics. First, the grid enhancement multi-head guided-attention module utilizes regions around the grid to provide visual context, forming rich visual grid semantics and effectively compensating for the fine granularity of the grid. Second, a novel geometric perception multi-head self-attention is introduced to process two types of features, incorporating geometric relations such as relative direction between objects while exploring internal semantic interactions. Extensive experiments demonstrate that the proposed method can obtain competitive results over other strong baselines.
Hong Yu 0005, Zhiyue Wang, Yuanqiu Liu, Han Liu 0008
ECAI1
2023 End-to-End Non-Autoregressive Image Captioning
abstract
Most of the existing image captioning models use the autoregressive approach to generate captions, which leads to high latency in the inference process. Non-autoregressive decoding generates words in parallel, which greatly improves the model inference speed. However, non-autoregressive decoding usually leads to performance loss due to the loss of word input. In this paper, we propose a semantic retrieval module that uses image features to retrieve semantic information as input of the non-autoregressive decoder, narrowing the performance gap between the non-autoregressive and the autoregressive model. Furthermore, we adopt Swin-Transformer instead of Faster R-CNN to extract image features, thus building an end-to-end image caption model. Experiments conducted on the MSCOCO dataset show that our model achieves new state-of-the-art performances of 122.6% CIDEr score on the ’Karpathy’ offline test split with 37× inference speedup.
Hong Yu 0005, Yuanqiu Liu, Baokun Qi, Zhao-Long Hu, Han Liu 0008
ICASSP1
2023 Boosting Decision-Based Black-Box Adversarial Attack with Gradient Priors
abstract
Decision-based methods have shown to be effective in black-box adversarial attacks, as they can obtain satisfactory performance and only require to access the final model prediction. Gradient estimation is a critical step in black-box adversarial attacks, as it will directly affect the query efficiency. Recent works have attempted to utilize gradient priors to facilitate score-based methods to obtain better results. However, these gradient priors still suffer from the edge gradient discrepancy issue and the successive iteration gradient direction issue, thus are difficult to simply extend to decision-based methods. In this paper, we propose a novel Decision-based Black-box Attack framework with Gradient Priors (DBA-GP), which seamlessly integrates the data-dependent gradient prior and time-dependent prior into the gradient estimation procedure. First, by leveraging the joint bilateral filter to deal with each random perturbation, DBA-GP can guarantee that the generated perturbations in edge locations are hardly smoothed, i.e., alleviating the edge gradient discrepancy, thus remaining the characteristics of the original image as much as possible. Second, by utilizing a new gradient updating strategy to automatically adjust the successive iteration gradient direction, DBA-GP can accelerate the convergence speed, thus improving the query efficiency. Extensive experiments have demonstrated that the proposed method outperforms other strong baselines significantly.
Han Liu 0008, Xingshuo Huang, Xiaotong Zhang 0003, Qimai Li, Fenglong Ma, Wei Wang 0077, Hongyang Chen 0001, Hong Yu 0005, Xianchao Zhang 0001
IJCAI8
2023 HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on Text
abstract
Black-box hard-label adversarial attack on text is a practical and challenging task, as the text data space is inherently discrete and non-differentiable, and only the predicted label is accessible. Research on this problem is still in the embryonic stage and only a few methods are available. Nevertheless, existing methods rely on the complex heuristic algorithm or unreliable gradient estimation strategy, which probably fall into the local optimum and inevitably consume numerous queries, thus are difficult to craft satisfactory adversarial examples with high semantic similarity and low perturbation rate in a limited query budget. To alleviate above issues, we propose a simple yet effective framework to generate high quality textual adversarial examples under the black-box hard-label attack scenarios, named HQA-Attack. Specifically, after initializing an adversarial example randomly, HQA-attack first constantly substitutes original words back as many as possible, thus shrinking the perturbation rate. Then it leverages the synonym set of the remaining changed words to further optimize the adversarial example with the direction which can improve the semantic similarity and satisfy the adversarial condition simultaneously. In addition, during the optimizing procedure, it searches a transition synonym word for each changed word, thus avoiding traversing the whole synonym set and reducing the query number to some extent. Extensive experimental results on five text classification datasets, three natural language inference datasets and two real-world APIs have shown that the proposed HQA-Attack method outperforms other strong baselines significantly.
Han Liu 0008, Zhi Xu 0008, Xiaotong Zhang 0003, Feng Zhang 0027, Fenglong Ma, Hongyang Chen 0001, Hong Yu 0005, Xianchao Zhang 0001
NeurIPS7
2022 Label-enhanced Prototypical Network with Contrastive Learning for Multi-label Few-shot Aspect Category Detection
abstract
Multi-label aspect category detection allows a given review sentence to contain multiple aspect categories, which is shown to be more practical in sentiment analysis and attracting increasing attention. As annotating large amounts of data is time-consuming and labor-intensive, data scarcity occurs frequently in real-world scenarios, which motivates multi-label few-shot aspect category detection. However, research on this problem is still in infancy and few methods are available. In this paper, we propose a novel label-enhanced prototypical network (LPN) for multi-label few-shot aspect category detection. The highlights of LPN can be summarized as follows. First, it leverages label description as auxiliary knowledge to learn more discriminative prototypes, which can retain aspect-relevant information while eliminating the harmful effect caused by irrelevant aspects. Second, it integrates with contrastive learning, which encourages that the sentences with the same aspect label are pulled together in embedding space while simultaneously pushing apart the sentences with different aspect labels. In addition, it introduces an adaptive multi-label inference module to predict the aspect count in the sentence, which is simple yet effective. Extensive experimental results on three datasets demonstrate that our proposed model LPN can consistently achieve state-of-the-art performance.
Han Liu 0008, Feng Zhang 0027, Xiaotong Zhang 0003, Siyang Zhao, Junjie Sun, Hong Yu 0005, Xianchao Zhang 0001
KDD6
2022 A Simple Meta-learning Paradigm for Zero-shot Intent Classification with Mixture Attention Mechanism
abstract
Zero-shot intent classification is a vital and challenging task in dialogue systems, which aims to deal with numerous fast-emerging unacquainted intents without annotated training data. To obtain more satisfactory performance, the crucial points lie in two aspects: extracting better utterance features and strengthening the model generalization ability. In this paper, we propose a simple yet effective meta-learning paradigm for zero-shot intent classification. To learn better semantic representations for utterances, we introduce a new mixture attention mechanism, which encodes the pertinent word occurrence patterns by leveraging the distributional signature attention and multi-layer perceptron attention simultaneously. To strengthen the transfer ability of the model from seen classes to unseen classes, we reformulate zero-shot intent classification with a meta-learning strategy, which trains the model by simulating multiple zero-shot classification tasks on seen categories, and promotes the model generalization ability with a meta-adapting procedure on mimic unseen categories. Extensive experiments on two real-world dialogue datasets in different languages show that our model outperforms other strong baselines on both standard and generalized zero-shot intent classification tasks.
Han Liu 0008, Siyang Zhao, Xiaotong Zhang 0003, Feng Zhang 0027, Junjie Sun, Hong Yu 0005, Xianchao Zhang 0001
SIGIR6
2021 Self-Guided Deep Multi-View Subspace Clustering Network
abstract
To cluster the data with complex structures, Deep Subspace Clustering Network (DSCN) extracts the subspace relations among non-linear latent features. However, the performance improvement has encountered bottlenecks due to the lack of supervision. Meantime, in multi-view settings, most of DSCN-based methods underestimate the significance of view-fusion, which always adopt simple tactics. To address these issues, we propose a self-supervised model for simultaneous subspace clustering, consensus construction and self-guided learning, named as Self-Guided Deep Multi-view Subspace Clustering Network (SG-DMSC). We utilize DSCN to learn a complex subspace representation for each single-view. Considering their different importance, we design the view-fusion layer to establish the agreement. We construct a novel loss term, the spectral supervisor, so that the consensus can be more clustering-friendly by the self-guidance of pseudo labels. Theoretical support is provided to reflect the validity of this self-guided strategy. An alternate iterative optimization algorithm is presented to handle SG-DMSC. Experiments on real-world datasets confirm its efficacy compared with others.
Beilei Cui, Hong Yu 0005, Linlin Zong
ICME2
2021 A Dual-Questioning Attention Network for Emotion-Cause Pair Extraction with Context Awareness
abstract
Emotion-cause pair extraction (ECPE), an emerging task in sentiment analysis, aims at extracting pairs of emotions and their corresponding causes in documents. This is a more challenging problem than emotion cause extraction (ECE), since it requires no emotion signals which are demonstrated as an important role in the ECE task. Existing work follows a two-stage pipeline which identifies emotions and causes at the first step and pairs them at the second step. However, error propagation across steps and pair combining without contextual information limits the effectiveness. Therefore, we propose a Dual-Questioning Attention Network to alleviate these limitations. Specifically, we question candidate emotions and causes to the context independently through attention networks for a contextual and semantical answer. Also, we explore how weighted loss functions in controlling error propagation between steps. Empirical results show that our method performs better than baselines in terms of multiple evaluation metrics. The source code can be obtained at https://github.com/QixuanSun/DQAN.
Qixuan Sun, Yaqi Yin, Hong Yu 0005
IJCNN3
2021 Incomplete multi-view clustering with partially mapped instances and clusters
Linlin Zong, Faqiang Miao, Xianchao Zhang 0001, Xinyue Liu 0002, Hong Yu 0005
Knowl. Based Syst.5
2020 Multi-view clustering via clusterwise weights learning
Qianli Zhao, Linlin Zong, Xianchao Zhang 0001, Xinyue Liu 0002, Hong Yu 0005
Knowl. Based Syst.5
2020 Multi-view clustering on data with partial instances and clusters
Linlin Zong, Xianchao Zhang 0001, Xinyue Liu 0002, Hong Yu 0005
Neural Networks4
2019 Self-Weighted Multi-View Clustering with Deep Matrix Factorization
abstract
Due to the efficiency of exploring multiple views of the real-word data, Multi-View Clustering (MVC) has attracted extensive attention from the scholars and researches based on it have made significant progress. However, multi-view data with numerous complementary information is vulnerable to various factors (such as noise). So it is an important and challenging task to discover the intrinsic characteristics hidden deeply in the data. In this paper, we present a novel MVC algorithm based on deep matrix factorization, named Self-Weighted Multi-view Clustering with Deep Matrix Factorization (SMDMF). By performing the deep decomposition structure, SMDMF can eliminate interference and reveal semantic information of the multi-view data. To properly integrate the complementary information among views, it assigns an automatic weight for each view without introducing supernumerary parameters. We also analyze the convergence of the algorithm and discuss the hierarchical parameters. The experimental results on four datasets show our algorithm is superior to other comparisons in all aspects.
Beilei Cui, Hong Yu 0005, Siwen Li
ACML2
2019 One Shot Learning with Margin
Xianchao Zhang 0001, Jinlong Nie, Linlin Zong, Hong Yu 0005, Wenxin Liang
PAKDD (2)4
2018 Weighted Multi-View Spectral Clustering Based on Spectral Perturbation
abstract
Considering the diversity of the views, assigning the multiviews with different weights is important to multi-view clustering. Several multi-view clustering algorithms have been proposed to assign different weights to the views. However, the existing weighting schemes do not simultaneously consider the characteristic of multi-view clustering and the characteristic of related single-view clustering. In this paper, based on the spectral perturbation theory of spectral clustering, we propose a weighted multi-view spectral clustering algorithm which employs the spectral perturbation to model the weights of the views. The proposed weighting scheme follows the two basic principles: 1) the clustering results on each view should be close to the consensus clustering result, and 2) views with similar clustering results should be assigned similar weights. According to spectral perturbation theory, the largest canonical angle is used to measure the difference between spectral clustering results. In this way, the weighting scheme can be formulated into a standard quadratic programming problem. Experimental results demonstrate the superiority of the proposed algorithm.
Linlin Zong, Xianchao Zhang 0001, Xinyue Liu 0002, Hong Yu 0005
AAAI4
2018 Co-regularized Multi-view Subspace Clustering
abstract
For many clustering applications, Multi-view data sets are very common. Multi-view clustering aims to exploit information across views instead of individual views, which is promising to improve clustering performance. Note that a high-dimensional data set usually distributes on certain low-dimensional subspace. Thus, many multi-view subspace clustering algorithms have been developed. However, existing multi-view subspace clustering methods rarely perform clustering on the subspace representation of each view simultaneously as well as keep the indicator consistency among the representations, i.e., the same data point in different views should be assigned to the same cluster. In this paper, we propose a novel multi-view subspace clustering method. In our method, we use the indicator matrix to ensure that we perform clustering on the subspace representation of each view simultaneously. And at the same time, a co-regularized term is added to guarantee the consistency of the indicator matrices. Experiments on several real-world multi-view datasets demonstrate the effectiveness and superiority of our proposed method.
Hong Yu 0005, Yahong Lian
ACML1
2018 Web Items Recommendation Based on Multi-View Clustering
abstract
Nowadays, using recommendation system to provide users with personalized recommendation service is significantly meaningful. However, traditional collaborative filtering methods may suffer from the cold start problem, while another common recommendation model called content-based recommendation may not have the ability to dig out the potential semantic features of web items sufficiently. In this paper, we propose a novel multi-content clustering collaborative filtering model (MCCCF) for recommendation system. The proposed model can apply multi-view clustering to the mining of the similarity and relevance of web items so that they can be used to improve the classic collaborative filtering. Consequently, the data sparsity problem can be solved. We propose to use multi-view clustering to analyse web items or users from different views such as user ratings and user comments so that it discovers deeper similarity and relevance. At the same time, features from multiple views can be used to complement the user views or item views where the features are deficient, which declines the problem of data sparsity drastically. In this way, we can analyse users' preference by their historical interaction features and supplementary behaviour features to give corresponding recommendation. Above all, the weak spots of the traditional model can be filled in and its performance can be improved. Extensive experiments on real world datasets show that our method outperforms the baselines remarkably.
Hong Yu 0005, Yahong Lian
COMPSAC (1)1
2017 View-Weighted Multi-view K-means Clustering
Hong Yu 0005, Yahong Lian
ICANN (2)1
2017 Self-Paced Learning Based Multi-view Spectral Clustering
abstract
Multi-view data are prevalent in both machine learning and artificial intelligence. A panoply of multi-view clustering algorithms have been proposed to deal with multiview data. However, most of them just blindly concatenate all the views in spite of characteristic of different views. Self-paced learning is a kind of learning scheme which comes from human learning. It progresses from easy example to complex example during learning process. In analogy with these intuitions, we can learn the easiness of multiple views. Therefore, in this paper, we first present a new Self-Paced Learning Regularizer which is a kind of mixture weighting scheme to allocate different weight to the different view by considering views' complexity. To recap the effectiveness of our self-paced learning regularizer, we propose a novel self-paced learning based multi-view spectral clustering algorithm(SPLMVC), which can define complexity across views and then automatically assign weight to each view. Extensive experiments on real-world multi-view datasets reveal its strength by comparison with other state-of-art methods.
Hong Yu 0005, Yahong Lian, Linlin Zong, Linlin Tian
ICTAI1
2017 Multi-view clustering via multi-manifold regularized non-negative matrix factorization
Linlin Zong, Xianchao Zhang 0001, Hong Yu 0005, Qianli Zhao
Neural Networks4
2016 Recommending Features of Mobile Applications for Developer
Hong Yu 0005, Yahong Lian, Shuotao Yang, Linlin Tian, Xiaowei Zhao 0003
ADMA1
2016 Constraint Based Subspace Clustering for High Dimensional Uncertain Data
Xianchao Zhang 0001, Hong Yu 0005
PAKDD (2)3
2016 Local linear neighbor reconstruction for multi-view data
Linlin Zong, Xianchao Zhang 0001, Hong Yu 0005, Qianli Zhao, Feng Ding 0004
Pattern Recognit. Lett.3
2015 Constrained NMF-Based Multi-View Clustering on Unmapped Data
abstract
Existing multi-view clustering algorithms require thatthe data is completely or partially mapped betweeneach pair of views. However, this requirement couldnot be satisfied in most practical settings. In this paper,we tackle the problem of multi-view clustering for unmappeddata in the framework of NMF based clustering.With the help of inter-view constraints, we definethe disagreement between each pair of views by the factthat the indicator vectors of two instances from two differentviews should be similar if they belong to the samecluster and dissimilar otherwise. The overall objectiveof our algorithm is to minimize the loss function of NMFin each view as well as the disagreement betweeneach pair of views. Experimental results show that, witha small number of constraints, the proposed algorithmgets good performance on unmapped data, and outperformsexisting algorithms on partially mapped data andcompletely mapped data.
Xianchao Zhang 0001, Linlin Zong, Xinyue Liu 0002, Hong Yu 0005
AAAI4
2015 Mobile Application Recommendations Based on Complex Information
Shuotao Yang, Hong Yu 0005, Xiaochen Lai
IEA/AIE2
2014 iDBMM: A Novel Algorithm to Model Dynamic Behavior in Large Evolving Graphs
abstract
In the dynamic social network, how to use data mining tools to find the hidden dynamic knowledge in the social network has become the focus of the study. It can be applied to a wide range of areas with good practical value and application significance. We propose a novel algorithm called iDBMM based on the improvement of DBMM algorithm. At first, iDBMM algorithm classifies the training set to obtain the basic characteristics of each role. Then it scores the test set relative to each role and distribute the role of the highest score to the corresponding node. Finally, the transition model is obtained by the statistical method. Experimental results show that new method determines the distribution of the roles of the nodes effectively to make up for the shortcoming of non-negative matrix factorization and improve the prediction accuracy.
Xiujuan Xu, Wei Wang 0077, Yu Liu 0035, Hong Yu 0005, Xiaowei Zhao 0003
DASC4
2014 The Role of Probability of Learning and Reconnecting in the Evolution of Cooperation
abstract
Cooperative phenomenon is widely researched within the fields of computational genomics, artificial intelligence, machine learning and data mining technologies. A key idea behind complex system constructed by intelligent agents is to establish cooperation between different agents so as to solve problems more effectively and efficiently than a single agent can do. Understanding the evolutionary mechanisms that promote and maintain cooperative behavior is recognized as a major theoretical problem where the intricacy increases with the complexity of the participating individuals. We presents current research on the effect of the "probability of learning" (Pl) and the "probability of reconnecting" (Pr) in NIPD game in evolution experiment based on complex network system, static and dynamic. We show that the "probability of learning" (Pl) is not the decisive factor on static network but it plays an important role on dynamic network where the "probability of reconnecting" (Pr) is not zero.
Xiaowei Zhao 0003, Hong Yu 0005, Zhenzhen Xu, Tinlin Tian, Xiujuan Xu
DASC2
2014 Multi-view Clustering via Multi-manifold Regularized Nonnegative Matrix Factorization
abstract
Multi-view clustering integrates complementary information from multiple views to gain better clustering performance rather than relying on a single view. NMF based multi-view clustering algorithms have shown their competitiveness among different multi-view clustering algorithms. However, NMF fails to preserve the locally geometrical structure of the data space. In this paper, we propose a multi-manifold regularized nonnegative matrix factorization framework (MMNMF) which can preserve the locally geometrical structure of the manifolds for multi-view clustering. MMNMF regards that the intrinsic manifold of the dataset is embedded in a convex hull of all the views' manifolds, and incorporates such an intrinsic manifold and an intrinsic (consistent) coefficient matrix with a multi-manifold regularizer to preserve the locally geometrical structure of the multi-view data space. We use linear combination to construct the intrinsic manifold, and propose two strategies to find the intrinsic coefficient matrix, which lead to two instances of the framework. Experimental results show that the proposed algorithms outperform existing NMF based algorithms for multi-view clustering.
Xianchao Zhang 0001, Linlin Zong, Xinyue Liu 0002, Hong Yu 0005
ICDM5
2014 Minimum Similarity Sampling Scheme for Nyström Based Spectral Clustering on Large Scale High-Dimensional Data
Zhicheng Zeng, Ming Zhu 0001, Hong Yu 0005, Honglian Ma
IEA/AIE (2)3
2012 An Extended ISOMAP by Enhancing Similarity for Clustering
Hong Yu 0005, Xianchao Zhang 0001, Yuansheng Yang, Xiaowei Zhao 0003
IEA/AIE1
2012 Agents' Cooperation Based on Long-Term Reciprocal Altruism
Xiaowei Zhao 0003, Haoxiang Xia, Hong Yu 0005, Linlin Tian
IEA/AIE3
2011 Local density adaptive similarity measurement for spectral clustering
Xianchao Zhang 0001, Hong Yu 0005
Pattern Recognit. Lett.3
2007 A Clustering Algorithm Based on Mechanics
Xianchao Zhang 0001, He Jiang 0001, Xinyue Liu 0002, Hong Yu 0005
PAKDD4