Yi Zhu 0006

dblp:67/4972-6 · DBLP profile ↗
← Back
60ranked-venue papers
15as first author
52since 2021 · last 2026
0000-0003-3045-2588ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 10 first-author · 33 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021
YearPublicationVenuePosition
2026 An Early Malicious Users Prediction Benchmark for Chinese Esports via Bullet Chats
Xiang Xing, Yi Zhu 0006, Chaowei Zhang 0001, Jipeng Qiang
ICIC (22)3
2026 Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
abstract
The widespread proliferation of online content has intensified concerns about clickbait, deceptive or exaggerated headlines designed to attract attention. While Large Language Models (LLMs) offer a promising avenue for addressing this issue, their effectiveness is often hindered by Sycophancy, a tendency to produce reasoning that matches users' beliefs over truthful ones, which deviates from instruction-following principles. Rather than treating sycophancy as a flaw to be eliminated, this work proposes a novel approach that initially harnesses this behavior to generate contrastive reasoning from opposing perspectives. Specifically, we design a Self-renewal Opposing-stance Reasoning Generation (SORG) framework that prompts LLMs to produce high-quality ''agree'' and ''disagree'' reasoning pairs for a given news title without requiring ground-truth labels. To utilize the generated reasoning, we develop a local Opposing Reasoning-based Clickbait Detection (ORCD) model that integrates three BERT encoders to represent the title and its associated reasoning. The model leverages contrastive learning, guided by soft labels derived from LLM-generated credibility scores, to enhance detection robustness. Experimental evaluations on three benchmark datasets demonstrate that our method consistently outperforms LLM prompting, fine-tuned smaller language models, and state-of-the-art clickbait detection baselines. Our code is available in https://github.com/126541/ORCD.
Chaowei Zhang 0001, Xiansheng Luo, Zewei Zhang, Yi Zhu 0006, Jipeng Qiang, Longwei Wang
WWW4
2026 Analyzing bullet chats for recommendation intent identification: Dataset and method
Yi Zhu 0006, Qinqin Han, Yun-Hao Yuan 0001, Chaowei Zhang 0001, Jipeng Qiang, Xindong Wu 0001
Artif. Intell.1
2026 LLM4CGDS: Large language model-based agents for Chinese graded document simplification
abstract
Graded reading tailors text difficulty to learners’ proficiency by producing multiple versions of the same content—an approach long embraced in language education but still dependent on labor-intensive, expert-driven adaptation. In this paper, we introduce the task of C hinese G raded D ocument S implification (CGDS) for non-native learners, which seeks to automate the creation of multi-level reading materials in accordance with established proficiency standards. Guided by the three stages of the Hanyu Shuiping Kaoshi (HSK) 3.0 framework (Levels 1–3 for Advanced, Levels 4–6 for Intermediate, and Levels 7–9 for Beginner learners), we propose Large Language Model for Chinese Graded Document Simplification (LLM4CGDS), a rule-guided, large language model (LLM)-based framework that integrates HSK-level readability constraints and external knowledge retrieval to control document-level simplification without requiring supervised fine-tuning. To foster further research, we construct two complementary datasets: J ourney to the W est D ocument S implification (JWDS) and M ulti- D omain D ocument S implification (MDDS) that covering diverse genres and difficulty levels. Experimental evaluation on two datasets demonstrates that LLM4CGDS substantially outperforms direct prompting of state-of-the-art LLMs in both readability control and meaning preservation.
Dengzhao Fang, Jipeng Qiang, Wenjie Hou, Yi Zhu 0006, Jingtong Gao, Zhaoxiang Zhang 0001
Eng. Appl. Artif. Intell.4
2026 SubAttack: A word-level adversarial textual attack method via antonym substitution
abstract
Over the past few years, various word-level textual attack approaches have been proposed to reveal the vulnerability in existing deep neural networks and even large language models (LLMs) for Natural Language Processing (NLP). The textual attack aims to fool existing models into making erroneous predictions by altering the text without affecting the user’s understanding. However, current methods either struggle to construct semantically preserved adversarial texts and altered the semantics of the original text, or fail to consider the semantic perturbation constraints and are prone to invalid adversarial examples. In this paper, we propose an efficient and effective framework SubAttack to address these issues. SubAttack is a word-level adversarial textual attack method via antonym substitution, which replaces semantic indicator keywords to generate high-quality adversarial samples with considering both semantically preservation and semantic perturbation. Specifically, the process first involves tokenizing the text and performing part-of-speech tagging Identifying the semantic indicator keywords. Then, the antonym ranking is designed to decide the substitutions of candidate words to fit the context. Finally, while retaining the original text, the ranked antonyms are integrated into the text and the instructions are added for both semantically preservation and semantic perturbation. Extensive experiments reveal that state-of-the-art (SOTA) LLMs (e.g. Llama and QWen) are still vulnerable to our SubAttack. Further experiments show that the adversarial examples crafted by SubAttack usually have higher quality, exhibit better fluency and barely affect human performance and can bring more robustness improvement to victim models by adversarial training.
Chenqi Hua, Yi Zhu 0006, Chaowei Zhang 0001, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Eng. Appl. Artif. Intell.3
2026 Personalized recommendation with clustering via prompt-tuning
abstract
The personalized recommendation aims to address the information overload problem, which can find interesting items for users from massive amounts of information. The research paradigm of personalized recommendation evolved from deep neural networks to pre-trained language models (PLMs) like BERT and, more recently, into large language models (LLMs). However, it is always very difficult to find the target item among a massive number of data or information, which is not only time-consuming but also often has low accuracy. In this paper, we propose a Personalized Recommendation method with Clustering via Prompt-tuning (PRCP), a candidate item set is developed and a prompt-tuning model with a designed verbalizer is constructed for recommendation. Specifically, the target users are first selected by the similarity calculation, and items are then clustered by the preferences of similar users to form a candidate item set. Then the prompt-tuning model is introduced to predict the masked label for candidate items, and three different strategies are designed to expand the label word space for verbalizer optimization. Extensive experiments conducted on three datasets validated the effectiveness of the proposed method compared to other state-of-the-art baselines including LLMs.
Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Intell. Data Anal.3
2026 Dualmark: A novel dual watermarking approach for large language models
Zihao Qiang, Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Chaowei Zhang 0001, Yan Liu 0038, Wei Li 0121
Inf. Process. Manag.4
2026 A smaller model can be better: Domain adaptation for LLM-generated text detection via soft prompt-tuning
Yi Zhu 0006, Pei-Pei Li 0001
Inf. Process. Manag.2
2026 Turning hallucinations into knowledge: Towards identifying clickbait using LLM-generated fallacies
Chaowei Zhang 0001, Zhicong Wang, Zewei Zhang, Yi Zhu 0006, Jipeng Qiang, Yuchao Huang
Inf. Process. Manag.4
2025 Collaborative Document Simplification Using Multi-Agent Systems
abstract
Research on text simplification has been ongoing for many years. However, the task of document simplification (DS) remains a significant challenge due to the need to consider complex factors such as technical terminology, metaphors, and overall coherence. In this work, we introduce a novel multi-agent framework for document simplification (AgentSimp) based on large language models (LLMs). This framework emulates the collaborative process of a human expert team through the roles played by multiple agents, addressing the intricate demands of document simplification. We explore two communication strategies among agents (pipeline-style and synchronous) and two document reconstruction strategies (Direct and Iterative ). According to both automatic evaluation metrics and human evaluation results, the documents simplified by AgentSimp are deemed to be more thoroughly simplified and more coherent on a variety of articles across different types and styles.
Dengzhao Fang, Jipeng Qiang, Xiaoye Ouyang, Yi Zhu 0006, Yun-Hao Yuan 0001, Yun Li 0010
COLING4
2025 Post-Hoc Watermarking for Robust Detection in Text Generated by Large Language Models
abstract
Research on text simplification has been ongoing for many years, yet document simplification remains a significant challenge due to the need to address complex factors such as technical terminology, metaphors, and overall coherence. In this work, we introduce a novel multi-agent framework AgentSimp for document simplification, based on large language models. This framework simulates the collaborative efforts of a team of human experts through the roles played by multiple agents, effectively meeting the intricate demands of document simplification. We investigate two communication strategies among agents (pipeline-style and synchronous) and two document reconstruction strategies (Direct and Iterative). According to both automatic evaluation metrics and human evaluation results, AgentSimp produces simplified documents that are more thoroughly simplified and more coherent across various articles and styles.
Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Xiaoye Ouyang
COLING3
2025 Learning Simultaneous Facial Canonical Correlation Representation for Face Hallucination
abstract
The low resolution (LR) problem is rather challenging in face analysis. Most existing face hallucination methods assume that LR face images have only one resolution, but multiple resolutions may be available from different sources. To solve this issue, we propose a novel simultaneous facial canonical correlation representation learning method for face hallucination, which seeks latent correlation subspaces for multi-resolution views. Our method jointly solves multiple linear transformations by optimizing a correlation summation criterion of all pairs of resolutions. The neighborhood reconstruction is used to infer the HR facial canonical correlation representation of LR face inputs. Extensive experimental results show the superiority of our proposed method in terms of quantitative and qualitative evaluations.
Yun-Hao Yuan 0001, Jin Li 0028, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001, Yun Li 0010
ICASSP4
2025 A domain adaptation method to Defend Chinese textual adversarial attacks via prompt-tuning
Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Eng. Appl. Artif. Intell.1
2025 Soft Prompt-tuning with Self-Resource Verbalizer for short text streams
Yi Zhu 0006, Ye Wang 0022, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001
Eng. Appl. Artif. Intell.1
2025 Simplified multi-view graph neural network for multilingual knowledge graph completion
Bingbing Dong, Chenyang Bu, Yi Zhu 0006, Shengwei Ji, Xindong Wu 0001
Frontiers Comput. Sci.3
2025 Robust and semantic-faithful post-hoc watermarking of text generated by black-box language models
Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Xiaocheng Hu, Xiaoye Ouyang
Frontiers Comput. Sci.3
2025 Domain adaptation for textual adversarial defense via prompt-tuning
Yi Zhu 0006, Chenqi Hua, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Neurocomputing2
2025 Multi-modal soft prompt-tuning for Chinese Clickbait Detection
Ye Wang 0022, Yi Zhu 0006, Yun Li 0010, Liting Wei, Yun-Hao Yuan 0001, Jipeng Qiang
Neurocomputing2
2025 Soft prompt-tuning for unsupervised domain adaptation via self-supervision
Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Neurocomputing1
2024 Incomplete Multi-Kernel k-Means Clustering With Fractional-Order Embedding
abstract
Multiple kernel clustering (MKC) has received increasing attention in the community of machine learning, which takes advantage of multiple pre-specified kernels to perform clustering tasks. Traditional MKC algorithms cannot effectively deal with the incomplete views where some samples are missing. Thus, incomplete MKC (IMKC) has been developed to solve this problem and obtained promising results. Nevertheless, the samples may be noisy or limited in real-world applications, which will result in the performance deterioration of existing IMKC algorithms. To address this issue, in this paper we propose a simple yet effective clustering method for incomplete data, termed fractional-order embedding incomplete multi-kernel k-means clustering (FE-MKKM-IK). Specifically, FE-MKKM-IK introduces the idea of fractional-order embedding to reconstruct the kernel matrix computed by the samples. On this basis, a new incomplete multiple kernel k-means clustering is developed. Performance evaluation is conducted on four widely used datasets, which shows that FE-MKKM-IK is effective to cluster the incomplete data.
Deheng Xu, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Yi Zhu 0006
IEEE Big Data5
2024 Learning Spectral Canonical ℱ-Correlation Representation for Face Super-Resolution
abstract
Face super-resolution (FSR) is a powerful technique for restoring high-resolution face images from the captured low-resolution ones with the assistance of prior information. Existing FSR methods based on explicit or implicit covariance matrices are difficult to reveal complex nonlinear relationships between features, as conventional covariance computation is essentially a linear operation process. Besides, the limited number of training samples and noise disturbance lead to the deviation of sample covariance matrices. To solve these issues, we propose a novel FSR method via using spectral canonical ℱ-correlation representation. The proposed method first defines intra-resolution and inter-resolution covariation matrices by considering the nonlinear relationship between different features, and then uses the fractional order idea to rebuild covariation matrices. The qualitative and quantitative results have validated the superiority of the proposed method.
Yun-Hao Yuan 0001, Mingzhi Hao, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001
ICASSP5
2024 Face Super-Resolution Using Covariation-Guided Orthonormalized Partial Least Squares
Mingzhi Hao, Yun-Hao Yuan 0001, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010, Runmei Zhang
ICONIP (8)4
2024 Short text classification with Soft Knowledgeable Prompt-tuning
Yi Zhu 0006, Ye Wang 0022, Jianyuan Mu, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001, Xindong Wu 0001
Expert Syst. Appl.1
2024 Representation learning: serial-autoencoder for personalized recommendation
Yi Zhu 0006, Yishuai Geng, Yun Li 0010, Jipeng Qiang, Xindong Wu 0001
Frontiers Comput. Sci.1
2024 Prompt-Learning for Short Text Classification
abstract
In the short text, the extremely short length, feature sparsity, and high ambiguity pose huge challenges to classification tasks. Recently, as an effective method for tuning Pre-trained Language Models for specific downstream tasks, prompt-learning has attracted a vast amount of attention and research. The main intuition behind the prompt-learning is to insert the template into the input and convert the tasks into equivalent cloze-style tasks. However, most prompt-learning methods only consider the class name and monotonous strategy for knowledge incorporating in cloze-style prediction, which will inevitably incur omissions and bias in short text classification tasks. In this paper, we propose a short text classification method with prompt-learning. Specifically, the top$M$concepts related to the entity in the short text are retrieved from the open Knowledge Graph like Probase, these concepts are first selected by the distance with class labels, which takes both the short text itself and the class name into consideration during expanding label word space. Then, we conducted four additional strategies for the integration of the expanded concepts, and the union of these concepts are adopted finally in the verbalizer of prompt-learning. Experimental results show that the obvious improvement is obtained compared with other state-of-the-art methods on five well-known datasets.
Yi Zhu 0006, Ye Wang 0022, Jipeng Qiang, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.1
2024 Iterative Soft Prompt-Tuning for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation aims to facilitate learning tasks in unlabeled target domain with knowledge in the related source domain, which has achieved awesome performance with the pre-trained language models (PLMs). Recently, inspired by GPT, the prompt-tuning model has been widely explored in stimulating rich knowledge in PLMs for language understanding. However, existing prompt-tuning methods still directly applied the model that was learned in the source domain into the target domain to minimize the discrepancy between different domains, e.g., the prompts or the template are trained separately to learn embeddings for transferring to the target domain, which is actually the intuition of end-to-end deep-based approach. In this paper, we propose an Iterative Soft Prompt-Tuning method (ItSPT) for better unsupervised domain adaptation. On the one hand, the prompt-tuning model learned in the source domain is converted into an iterative model to find the true label information in the target domain, the domain adaptation method is then regarded as a few-shot learning task. On the other hand, instead of hand-crafted templates, ItSPT adopts soft prompts for both considering the automatic template generation and classification performance. Experiments on both English and Chinese datasets demonstrate that our method surpasses the performance of SOTA methods.
Yi Zhu 0006, Jipeng Qiang, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.1
2023 ParaLS: Lexical Substitution via Pretrained Paraphraser
abstract
Lexical substitution (LS) aims at finding appropriate substitutes for a target word in a sentence.Recently, LS methods based on pretrained language models have made remarkable progress, generating potential substitutes for a target word through analysis of its contextual surroundings.However, these methods tend to overlook the preservation of the sentence's meaning when generating the substitutes.This study explores how to generate the substitute candidates from a paraphraser, as the generated paraphrases from a paraphraser contain variations in word choice and preserve the sentence's meaning.Since we cannot directly generate the substitutes via commonly used decoding strategies, we propose two simple decoding strategies that focus on the variations of the target word during decoding.Experimental results show that our methods outperform state-of-theart LS methods based on pre-trained language models on three benchmarks.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006
ACL (1)5
2023 Multilingual Lexical Simplification via Paraphrase Generation
abstract
Lexical simplification (LS) methods based on pretrained language models have made remarkable progress, generating potential substitutes for a complex word through analysis of its contextual surroundings. However, these methods require separate pretrained models for different languages and disregard the preservation of sentence meaning. In this paper, we propose a novel multilingual LS method via paraphrase generation, as paraphrases provide diversity in word selection while preserving the sentence’s meaning. We regard paraphrasing as a zero-shot translation task within multilingual neural machine translation that supports hundreds of languages. After feeding the input sentence into the encoder of paraphrase modeling, we generate the substitutes based on a novel decoding strategy that concentrates solely on the lexical variations of the complex word. Experimental results demonstrate that our approach surpasses BERT-based methods and zero-shot GPT3-based method significantly on English, Spanish, and Portuguese.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006, Kaixun Hua
ECAI5
2023 Chinese Lexical Substitution: Dataset and Method
abstract
Existing lexical substitution (LS) benchmarks were collected by asking human annotators to think of substitutes from memory, resulting in benchmarks with limited coverage and relatively small scales.To overcome this problem, we propose a novel annotation method to construct an LS dataset based on human and machine collaboration.Based on our annotation method, we construct the first Chinese LS dataset CHNLS which consists of 33,695 instances and 144,708 substitutes, covering three text genres (News, Novel, and Wikipedia).Specifically, we first combine four unsupervised LS methods as an ensemble method to generate the candidate substitutes, and then let human annotators judge these candidates or add new ones.This collaborative process combines the diversity of machine-generated substitutes with the expertise of human annotators.Experimental results that the ensemble method outperforms other LS methods.To our best knowledge, this is the first study for the Chinese LS task.
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xiaocheng Hu, Xiaoye Ouyang
EMNLP5
2023 Learning Supervised Covariation Projection Through General Covariance
abstract
Canonical correlation analysis (CCA) is a classical yet powerful tool for learning two-view feature representation in various fields. But, most CCA approaches are based on the conventional covariance measure, which makes them difficult to uncover the complicatedly nonlinear relationship between distinct features. In this paper, we address the preceding problem and propose two novel CCA approaches in a supervised manner by using a general covariance metric. The proposed approaches not only consider the label information of training data, but also the nonlinear relationship between different features rather than samples, which leads to greater flexibility in many practical applications. A series of experimental results on five benchmark datasets demonstrate the effectiveness of our proposed methods in terms of classification accuracy.
Xiangze Bao, Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006
ICASSP5
2023 Many Is Better Than One: Multiple Covariation Learning for Latent Multiview Representation
Yun-Hao Yuan 0001, Pengwei Qian, Jin Li 0028, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010
ICONIP (9)5
2023 Natural language watermarking via paraphraser-based lexical substitution
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001
Artif. Intell.4
2023 Joint user profiling with hierarchical attention networks
Yi Zhu 0006, Xindong Wu 0001
Frontiers Comput. Sci.2
2023 Lexical simplification via single-word generation
Jipeng Qiang, Yang Li 0186, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006
Frontiers Comput. Sci.5
2023 Unsupervised statistical text simplification using pre-trained language modeling for initialization
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006, Xindong Wu 0001
Frontiers Comput. Sci.5
2023 Representation learning via an integrated autoencoder for unsupervised domain adaptation
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Yun-Hao Yuan 0001, Yun Li 0010
Frontiers Comput. Sci.1
2023 A hybrid classification method via keywords screening and attention mechanisms in extreme short text
abstract
Short text classification has provoked a vast amount of attention and research in recent decades. However, most existing methods only focus on the short texts that contain dozens of words like Twitter and Microblog, while pay far less attention to the extreme short texts like news headline and search snippets. Meanwhile, contemporary short text classification methods that extend the features via external knowledge sources always introduce lots of useless concepts, which may be detrimental to classification performance. Moreover, unlike traditional short text classification methods, the classification results of extreme short texts are often determined by a few even one or two keywords. To address these problems, we propose a novel hybrid classification method via Keywords Screening and Attention Mechanisms in extreme short text, called KSAM. More specifically, firstly, the attention-based BiLSTM is introduced in our method to enhance the role of keywords. Secondly, we screen the keywords in the extreme short text for obtaining the true class label, and the concepts concerning the keywords are retrieved from external open knowledge sources like DBpedia. Thirdly, the attention mechanisms are introduced to acquire the weight of these retrieved concepts. Finally, conceptual information is utilized to assist the classification of the extreme short text. Extensive experiments have demonstrated the effectiveness of our method compared to other state-of-the-art methods.
Xinke Zhou, Yi Zhu 0006, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001, Xingdong Wu, Runmei Zhang
Intell. Data Anal.2
2023 Chinese Idiom Paraphrasing
abstract
Abstract Idioms are a kind of idiomatic expression in Chinese, most of which consist of four Chinese characters. Due to the properties of non-compositionality and metaphorical meaning, Chinese idioms are hard to be understood by children and non-native speakers. This study proposes a novel task, denoted as Chinese Idiom Paraphrasing (CIP). CIP aims to rephrase idiom-containing sentences to non-idiomatic ones under the premise of preserving the original sentence’s meaning. Since the sentences without idioms are more easily handled by Chinese NLP systems, CIP can be used to pre-process Chinese datasets, thereby facilitating and improving the performance of Chinese NLP tasks, e.g., machine translation systems, Chinese idiom cloze, and Chinese idiom embeddings. In this study, we can treat the CIP task as a special paraphrase generation task. To circumvent difficulties in acquiring annotations, we first establish a large-scale CIP dataset based on human and machine collaboration, which consists of 115,529 sentence pairs. In addition to three sequence-to-sequence methods as the baselines, we further propose a novel infill-based approach based on text infilling. The results show that the proposed method has better performance than the baselines based on the established CIP dataset.
Jipeng Qiang, Yang Li 0186, Chaowei Zhang 0001, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001
Trans. Assoc. Comput. Linguistics5
2023 Knowledge Graph for China's Genealogy11.A shorter version of this paper won the Best Paper Award at IEEE ICKG 2020 (the 11th IEEE International Conference on Knowledge Graph, ickg 2020.bigke.org)
abstract
Genealogical knowledge graphs depict the relationships of family networks and the development of family histories. They can help researchers to analyze and understand genealogical data, search for genealogical descendant paths, and explore the origins of a family more easily. However, the heterogenous, autonomous, complex, and evolving natures of genealogical data bring challenges to the development of contemporary genealogical knowledge graph models. Applying existing methods to genealogical data may be improper because general knowledge graph models lack in-depth domain knowledge. In this paper, we propose a genealogical knowledge graph model named Huapu-KG that combines HAO intelligence (human intelligence + artificial intelligence + organizational intelligence) to implement the construction and applications of genealogical knowledge graphs. Furthermore, challenges in constructing genealogical knowledge graphs are demonstrated, and experiments conducted on real-world genealogical datasets verify the feasibility and effectiveness of our proposed model.
Xindong Wu 0001, Tingting Jiang 0004, Yi Zhu 0006, Chenyang Bu
IEEE Trans. Knowl. Data Eng.3
2022 Learning Canonical F-Correlation Projection for Compact Multiview Representation
abstract
Canonical correlation analysis (CCA) matters in multi-view representation learning. But, CCA and its most variants are essentially based on explicit or implicit covariance matrices. It means that they have no ability to model the nonlinear relationship among features due to intrinsic linearity of covariance. In this paper, we address the preceding problem and propose a novel canonical F-correlation framework by exploring and exploiting the nonlinear relationship between different features. The framework projects each feature rather than observation into a certain new space by an arbitrary nonlinear mapping, thus resulting in more flexibility in real applications. With this frame-work as a tool, we propose a correlative covariation projection (CCP) method by using an explicit nonlinear mapping. Moreover, we further propose a multiset version of CCP dubbed MCCP for learning compact representation of more than two views. The proposed MCCP is solved by an iterative method, and we prove the convergence of this iteration. A series of experimental results on six benchmark datasets demonstrate the effectiveness of our proposed CCP and MCCP methods.
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001, Jianping Gou
CVPR5
2022 Hypernode: Entity Fusion for Data Traceability and Link Prediction
abstract
In the era of big data, fragmented knowledge, multisource heterogeneity, and different representation forms of the same entities in various data sources have posed considerable challenges to entity fusion. How to effectively integrate multisource knowledge for the same entities has provoked vast amounts of attention and research from multiple disciplines. Most existing methods for entity fusion can be categorized into two classes: one is to establish an association between the same entities, and the other is to delete duplicate entities after knowledge fusion and create a new fusion entity. However, in these two classes of methods, the former does not achieve true knowledge fusion and semantic interoperability, while the latter may cause irreversible loss of original information. In this paper, we propose a novel entity fusion scheme: Hypernode. Hypernode fuses the same entity in different data sources into a new entity while retaining the original data. We verify the effectiveness of Hypernode on multiple models of link prediction experiments. Several practical application cases illustrate the applicability of Hypernode in data traceability, open domain knowledge fusion, and multi-modal knowledge graph fusion.
Bingbing Dong, Zan Zhang 0002, Yi Zhu 0006, Chenyang Bu, Xindong Wu 0001
ICDM4
2022 Personalized recommendation with knowledge graph via dual-autoencoder
Yang Yang 0002, Yi Zhu 0006, Yun Li 0010
Appl. Intell.2
2022 Combining embedding-based and symbol-based methods for entity alignment
Tingting Jiang 0004, Chenyang Bu, Yi Zhu 0006, Xindong Wu 0001
Pattern Recognit.3
2022 Representation learning with deep sparse auto-encoder for multi-task learning
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001
Pattern Recognit.1
2021 Fractional Multi-view Hashing with Semantic Correlation Maximization
Ruijie Gao, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Yi Zhu 0006
ICONIP (5)5
2021 Multi-view Fractional Deep Canonical Correlation Analysis for Subspace Clustering
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001
ICONIP (2)5
2021 Domain Adaptation with Stacked Convolutional Sparse Autoencoder
Yi Zhu 0006, Xinke Zhou, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001
ICONIP (5)1
2021 Representation learning with collaborative autoencoder for personalized recommendation
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Yun-Hao Yuan 0001, Yun Li 0010
Expert Syst. Appl.1
2021 Microbloggers' interest inference using a subgraph stream
abstract
Inferring user interest over large-scale microblogs have attracted much attention in recent years. However, the emergence of the massive data, dynamic change of information and persistence of microblogs pose challenges to interest inference. Most of the existing approaches rarely take into account the combination of these microbloggers’ characteristics within the model, which may incur information loss with nontrivial magnitude in real-time extraction of user interest and massive social data processing. To address these problems, in this paper, we propose a novel User-Networked Interest Topic Extraction in the form of Subgraph Stream (UNITE_SS) for microbloggers’ interest inference. To be specific, we develop several strategies for the construction of subgraph stream to select the better strategy for user interest inference. Moreover, the information of microblogs in each subgraph is utilized to obtain a real-time and effective interest for microbloggers. The experimental evaluation on a large dataset from Sina Weibo, one of the most popular microblogs in China, demonstrates that the proposed approach outperforms the state-of-the-art baselines in terms of precision, mean reciprocal rank (MRR) as well as runtime from the effectiveness and efficiency perspectives.
Hao Wang 0008, Lei Li 0002, Yi Zhu 0006, Chengxiang Hu
Intell. Data Anal.4
2021 Hybrid collaborative recommendation of co-embedded item attributes and graph features
Bingbing Dong, Yi Zhu 0006, Lei Li 0002, Xindong Wu 0001
Neurocomputing2
2021 LSBert: Lexical Simplification Based on BERT
abstract
Lexical simplification (LS) aims at replacing complex words with simpler alternatives. LS commonly consists of three main steps: complex word identification, substitute generation, and substitute ranking. Existing LS methods focus on the contextual information of the complex word in the last step (substitute ranking). However, they miss out the following two facts: (1) The word complexity of a polysemous word is very closely related to its context; (2) The step of substitute generation regardless of the context will inevitably produce a large number of spurious candidates. Therefore, we propose a novel LS system LSBert based on pretrained language model BERT to address the aforementioned issues, which is capable of making use of the wider context when both identifying the words in need of simplification and generating substitute candidates for the complex words. Specifically, LSBert consists of a network for complex word identification by fine-tuning BERT and a network for substitute generation based on BERT. Experimental results show that LSBert performs well in both complex word identification and substitute generation, achieving state-of-the-art results in three benchmarks. To facilitate reproducibility, the code of the LSBert system is available at https://github.com/qiang2100/BERT-LS.
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Yang Shi 0003, Xindong Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Stacked Convolutional Sparse Auto-Encoders for Representation Learning
abstract
Deep learning seeks to achieve excellent performance for representation learning in image datasets. However, supervised deep learning models such as convolutional neural networks require a large number of labeled image data, which is intractable in applications, while unsupervised deep learning models like stacked denoising auto-encoder cannot employ label information. Meanwhile, the redundancy of image data incurs performance degradation on representation learning for aforementioned models. To address these problems, we propose a semi-supervised deep learning framework called stacked convolutional sparse auto-encoder, which can learn robust and sparse representations from image data with fewer labeled data records. More specifically, the framework is constructed by stacking layers. In each layer, higher layer feature representations are generated by features of lower layers in a convolutional way with kernels learned by a sparse auto-encoder. Meanwhile, to solve the data redundance problem, the algorithm of Reconstruction Independent Component Analysis is designed to train on patches for sphering the input data. The label information is encoded using a Softmax Regression model for semi-supervised learning. With this framework, higher level representations are learned by layers mapping from image data. It can boost the performance of the base subsequent classifiers such as support vector machines. Extensive experiments demonstrate the superior classification performance of our framework compared to several state-of-the-art representation learning methods.
Yi Zhu 0006, Lei Li 0002, Xindong Wu 0001
ACM Trans. Knowl. Discov. Data1
2020 Lexical Simplification with Pretrained Encoders
abstract
Lexical simplification (LS) aims to replace complex words in a given sentence with their simpler alternatives of equivalent meaning. Recently unsupervised lexical simplification approaches only rely on the complex word itself regardless of the given sentence to generate candidate substitutions, which will inevitably produce a large number of spurious candidates. We present a simple LS approach that makes use of the Bidirectional Encoder Representations from Transformers (BERT) which can consider both the given sentence and the complex word during generating candidate substitutions for the complex word. Specifically, we mask the complex word of the original sentence for feeding into the BERT to predict the masked token. The predicted results will be used as candidate substitutions. Despite being entirely unsupervised, experimental results show that our approach obtains obvious improvement compared with these baselines leveraging linguistic databases and parallel corpus, outperforming the state-of-the-art by more than 12 Accuracy points on three well-known benchmarks.
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001
AAAI3
2020 Regularized Multiset Neighborhood Correlation Analysis for Semi-paired Multiview Learning
Yun-Hao Yuan 0001, Zhaoqi Wu, Yun Li 0010, Jipeng Qiang, Jianping Gou, Yi Zhu 0006
ICONIP (2)6
2020 Semi-supervised representation learning via dual autoencoders for domain adaptation
Shuai Yang 0003, Hao Wang 0008, Yuhong Zhang 0002, Pei-Pei Li 0001, Yi Zhu 0006, Xuegang Hu
Knowl. Based Syst.5
2020 Wasserstein GAN based on Autoencoder with back-translation for cross-lingual embedding mappings
Yuhong Zhang 0002, Yuling Li 0001, Yi Zhu 0006, Xuegang Hu
Pattern Recognit. Lett.3
2019 Two-Stage Entity Alignment: Combining Hybrid Knowledge Graph Embedding with Similarity-Based Relation Alignment
Tingting Jiang 0004, Chenyang Bu, Yi Zhu 0006, Xindong Wu 0001
PRICAI (1)3
2019 Representation learning via serial autoencoders for domain adaptation
Shuai Yang 0003, Yuhong Zhang 0002, Yi Zhu 0006, Pei-Pei Li 0001, Xuegang Hu
Neurocomputing3
2019 Transfer learning with deep manifold regularized auto-encoders
Yi Zhu 0006, Xindong Wu 0001, Pei-Pei Li 0001, Yuhong Zhang 0002, Xuegang Hu
Neurocomputing1
2018 Transfer learning with stacked reconstruction independent component analysis
Yi Zhu 0006, Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001
Knowl. Based Syst.1