Jipeng Qiang

dblp:138/2494 · also Ji-Peng Qiang · DBLP profile ↗
← Back
78ranked-venue papers
19as first author
56since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 13 first-author · 33 since 2021Databases, data management, data science and information retrieval · 13 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Early Malicious Users Prediction Benchmark for Chinese Esports via Bullet Chats
Xiang Xing, Yi Zhu 0006, Chaowei Zhang 0001, Jipeng Qiang
ICIC (22)5
2026 ProEchoMem: Enhancing Long Video Understanding via Multi-Trace Probe-Echo Memory
abstract
Large vision-language models (LVLMs) have shown significant progress in video understanding, but they struggle to scale to long videos due to limited context windows. Existing methods reduce input dimensionality via frame sampling and feature compression, yet discard details and incur high computational cost for post-training. In contrast, retrieval-augmented generation (RAG) that indexes long videos for query retrieval and memory-based methods that maintain evolving long-term stores, offer a lighter and deployment-friendly solution. Nevertheless, they rely on shallow retrieval that selects only top-ranked segments and fails to integrate information across multiple relevant video episodes. Inspired by Multiple-Trace Theory in cognitive psychology, we revisit long video understanding from a probe-echo perspective, in which human episodic memories are activated and integrated in parallel. Building on this insight, we propose ProEchoMem, a cognitive-inspired framework that simulates the probe-echo mechanism: (1) Incremental Episodic Memory Construction builds structured knowledge graphs from video streams; (2) Probe-Driven Memory Activation generates probe signals from user queries to activate all stored traces simultaneously; (3) Memory Echo Synthesis integrates activated traces into a coherent and structured memory echo. Experiments on LongerVideos, LVBench, and cross-domain settings demonstrate the effectiveness of ProEchoMem, with multi-trace probing achieving up to 14.2% higher relevance and ablation studies validating the contribution of each module. The code is available at https://github.com/Applied-Machine-Learning-Lab/SIGIR26_ProEchoMem
Derong Xu, Yanxin Chen, Pengyue Jia, Chao Zhang 0096, Maolin Wang 0001, Yiqi Wang 0001, Jipeng Qiang, Xuetao Wei, Hongzhi Yin, Tong Xu 0001, Xiangyu Zhao 0001
SIGIR8
2026 Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
abstract
The widespread proliferation of online content has intensified concerns about clickbait, deceptive or exaggerated headlines designed to attract attention. While Large Language Models (LLMs) offer a promising avenue for addressing this issue, their effectiveness is often hindered by Sycophancy, a tendency to produce reasoning that matches users' beliefs over truthful ones, which deviates from instruction-following principles. Rather than treating sycophancy as a flaw to be eliminated, this work proposes a novel approach that initially harnesses this behavior to generate contrastive reasoning from opposing perspectives. Specifically, we design a Self-renewal Opposing-stance Reasoning Generation (SORG) framework that prompts LLMs to produce high-quality ''agree'' and ''disagree'' reasoning pairs for a given news title without requiring ground-truth labels. To utilize the generated reasoning, we develop a local Opposing Reasoning-based Clickbait Detection (ORCD) model that integrates three BERT encoders to represent the title and its associated reasoning. The model leverages contrastive learning, guided by soft labels derived from LLM-generated credibility scores, to enhance detection robustness. Experimental evaluations on three benchmark datasets demonstrate that our method consistently outperforms LLM prompting, fine-tuned smaller language models, and state-of-the-art clickbait detection baselines. Our code is available in https://github.com/126541/ORCD.
Chaowei Zhang 0001, Xiansheng Luo, Zewei Zhang, Yi Zhu 0006, Jipeng Qiang, Longwei Wang
WWW5
2026 Analyzing bullet chats for recommendation intent identification: Dataset and method
Yi Zhu 0006, Qinqin Han, Yun-Hao Yuan 0001, Chaowei Zhang 0001, Jipeng Qiang, Xindong Wu 0001
Artif. Intell.5
2026 LLM4CGDS: Large language model-based agents for Chinese graded document simplification
abstract
Graded reading tailors text difficulty to learners’ proficiency by producing multiple versions of the same content—an approach long embraced in language education but still dependent on labor-intensive, expert-driven adaptation. In this paper, we introduce the task of C hinese G raded D ocument S implification (CGDS) for non-native learners, which seeks to automate the creation of multi-level reading materials in accordance with established proficiency standards. Guided by the three stages of the Hanyu Shuiping Kaoshi (HSK) 3.0 framework (Levels 1–3 for Advanced, Levels 4–6 for Intermediate, and Levels 7–9 for Beginner learners), we propose Large Language Model for Chinese Graded Document Simplification (LLM4CGDS), a rule-guided, large language model (LLM)-based framework that integrates HSK-level readability constraints and external knowledge retrieval to control document-level simplification without requiring supervised fine-tuning. To foster further research, we construct two complementary datasets: J ourney to the W est D ocument S implification (JWDS) and M ulti- D omain D ocument S implification (MDDS) that covering diverse genres and difficulty levels. Experimental evaluation on two datasets demonstrates that LLM4CGDS substantially outperforms direct prompting of state-of-the-art LLMs in both readability control and meaning preservation.
Dengzhao Fang, Jipeng Qiang, Wenjie Hou, Yi Zhu 0006, Jingtong Gao, Zhaoxiang Zhang 0001
Eng. Appl. Artif. Intell.2
2026 SubAttack: A word-level adversarial textual attack method via antonym substitution
abstract
Over the past few years, various word-level textual attack approaches have been proposed to reveal the vulnerability in existing deep neural networks and even large language models (LLMs) for Natural Language Processing (NLP). The textual attack aims to fool existing models into making erroneous predictions by altering the text without affecting the user’s understanding. However, current methods either struggle to construct semantically preserved adversarial texts and altered the semantics of the original text, or fail to consider the semantic perturbation constraints and are prone to invalid adversarial examples. In this paper, we propose an efficient and effective framework SubAttack to address these issues. SubAttack is a word-level adversarial textual attack method via antonym substitution, which replaces semantic indicator keywords to generate high-quality adversarial samples with considering both semantically preservation and semantic perturbation. Specifically, the process first involves tokenizing the text and performing part-of-speech tagging Identifying the semantic indicator keywords. Then, the antonym ranking is designed to decide the substitutions of candidate words to fit the context. Finally, while retaining the original text, the ranked antonyms are integrated into the text and the instructions are added for both semantically preservation and semantic perturbation. Extensive experiments reveal that state-of-the-art (SOTA) LLMs (e.g. Llama and QWen) are still vulnerable to our SubAttack. Further experiments show that the adversarial examples crafted by SubAttack usually have higher quality, exhibit better fluency and barely affect human performance and can bring more robustness improvement to victim models by adversarial training.
Chenqi Hua, Yi Zhu 0006, Chaowei Zhang 0001, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Eng. Appl. Artif. Intell.7
2026 Personalized recommendation with clustering via prompt-tuning
abstract
The personalized recommendation aims to address the information overload problem, which can find interesting items for users from massive amounts of information. The research paradigm of personalized recommendation evolved from deep neural networks to pre-trained language models (PLMs) like BERT and, more recently, into large language models (LLMs). However, it is always very difficult to find the target item among a massive number of data or information, which is not only time-consuming but also often has low accuracy. In this paper, we propose a Personalized Recommendation method with Clustering via Prompt-tuning (PRCP), a candidate item set is developed and a prompt-tuning model with a designed verbalizer is constructed for recommendation. Specifically, the target users are first selected by the similarity calculation, and items are then clustered by the preferences of similar users to form a candidate item set. Then the prompt-tuning model is introduced to predict the masked label for candidate items, and three different strategies are designed to expand the label word space for verbalizer optimization. Extensive experiments conducted on three datasets validated the effectiveness of the proposed method compared to other state-of-the-art baselines including LLMs.
Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Intell. Data Anal.6
2026 Dualmark: A novel dual watermarking approach for large language models
Zihao Qiang, Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Chaowei Zhang 0001, Yan Liu 0038, Wei Li 0121
Inf. Process. Manag.3
2026 Turning hallucinations into knowledge: Towards identifying clickbait using LLM-generated fallacies
Chaowei Zhang 0001, Zhicong Wang, Zewei Zhang, Yi Zhu 0006, Jipeng Qiang, Yuchao Huang
Inf. Process. Manag.5
2026 CellPredX, a computational framework for cross-data type, cross-sample, and cross-protocol cell type annotation through domain adaptation and deep metric learning
abstract
Accurate cell type annotation is fundamental to single-cell analysis, yet remains challenging across heterogeneous datasets and modalities. In particular, transferring labels between scRNA-seq and scATAC-seq data poses unique difficulties due to discrepancies in sequencing protocols and feature spaces. Existing methods typically handle only a subset of these challenges, often requiring scenario-specific adjustments and offering limited interpretability. Here, we present CellPredX, a structurally unified but adaptively parameterized, semi-supervised cross-modality framework for label transfer across scRNA-seq, scATAC-seq, and cross-protocol datasets. While maintaining a unified model architecture and optimization strategy, CellPredX allows adaptive tuning of loss-weight hyperparameters to account for the varying degree of similarity or discrepancy between different reference-query dataset pairs. CellPredX integrates domain adaptation and deep metric learning to align heterogeneous embeddings, and introduces a sparse center loss with an attention mechanism to enhance discriminative representations while suppressing noise. Moreover, an integrated interpreter module based on gradient attribution enables biological interpretability by identifying key markers and feature dimensions driving model predictions. Through extensive benchmarking across scRNA to scATAC, scATAC to scATAC, and scRNA to scRNA transfers, CellPredX consistently outperforms state-of-the-art annotation methods in both accuracy and robustness. The interpreter module further reveals biologically meaningful marker patterns that are consistent with known cell hierarchies. Together, these results demonstrate that CellPredX provides an interpretable and scalable solution for cross-modality cell type annotation in single-cell multi-omic integration.
Yan Liu 0038, Long-Chen Shen, Jipeng Qiang
PLoS Comput. Biol.6
2025 Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News Detection
abstract
The questionable responses caused by knowledge hallucination may lead to LLMs' unstable ability in decision-making. However, it has never been investigated whether the LLMs' hallucination is possibly usable for generating negative reasoning to assist fake news detection. In this paper, we propose a novel supervised self-reinforced reasoning rectification approach - SR^3 that not only yields common reasonable reasoning for news but also forces LLMs to generate the wrong understandings of news via LLMs reflection for semantic consistency learning. Upon that, we construct a negative reasoning-based news learning model called - NRFE, which leverages positive or negative news-reasoning pairs for learning the semantic consistency between them. To avoid the impact of label-implicated reasoning, we deploy a student model - NRFE-D that only takes news content as input to inspect the performance of our method by distilling the knowledge from NRFE. The experimental results verified on three popular fake news datasets demonstrate the superiority of our method compared with three kinds of baselines including prompting-based LLMs, fine-tuning-based PLMs, and other representative fake news detection methods.
Chaowei Zhang 0001, Zongling Feng, Zewei Zhang, Jipeng Qiang, Guandong Xu, Yun Li 0010
AAAI4
2025 Collaborative Document Simplification Using Multi-Agent Systems
abstract
Research on text simplification has been ongoing for many years. However, the task of document simplification (DS) remains a significant challenge due to the need to consider complex factors such as technical terminology, metaphors, and overall coherence. In this work, we introduce a novel multi-agent framework for document simplification (AgentSimp) based on large language models (LLMs). This framework emulates the collaborative process of a human expert team through the roles played by multiple agents, addressing the intricate demands of document simplification. We explore two communication strategies among agents (pipeline-style and synchronous) and two document reconstruction strategies (Direct and Iterative ). According to both automatic evaluation metrics and human evaluation results, the documents simplified by AgentSimp are deemed to be more thoroughly simplified and more coherent on a variety of articles across different types and styles.
Dengzhao Fang, Jipeng Qiang, Xiaoye Ouyang, Yi Zhu 0006, Yun-Hao Yuan 0001, Yun Li 0010
COLING2
2025 Post-Hoc Watermarking for Robust Detection in Text Generated by Large Language Models
abstract
Research on text simplification has been ongoing for many years, yet document simplification remains a significant challenge due to the need to address complex factors such as technical terminology, metaphors, and overall coherence. In this work, we introduce a novel multi-agent framework AgentSimp for document simplification, based on large language models. This framework simulates the collaborative efforts of a team of human experts through the roles played by multiple agents, effectively meeting the intricate demands of document simplification. We investigate two communication strategies among agents (pipeline-style and synchronous) and two document reconstruction strategies (Direct and Iterative). According to both automatic evaluation metrics and human evaluation results, AgentSimp produces simplified documents that are more thoroughly simplified and more coherent across various articles and styles.
Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Xiaoye Ouyang
COLING2
2025 Learning Simultaneous Facial Canonical Correlation Representation for Face Hallucination
abstract
The low resolution (LR) problem is rather challenging in face analysis. Most existing face hallucination methods assume that LR face images have only one resolution, but multiple resolutions may be available from different sources. To solve this issue, we propose a novel simultaneous facial canonical correlation representation learning method for face hallucination, which seeks latent correlation subspaces for multi-resolution views. Our method jointly solves multiple linear transformations by optimizing a correlation summation criterion of all pairs of resolutions. The neighborhood reconstruction is used to infer the HR facial canonical correlation representation of LR face inputs. Extensive experimental results show the superiority of our proposed method in terms of quantitative and qualitative evaluations.
Yun-Hao Yuan 0001, Jin Li 0028, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001, Yun Li 0010
ICASSP3
2025 Simplify and Translate: A Unified Framework for Accessible Machine Translation
Jipeng Qiang
NLPCC (3)2
2025 A domain adaptation method to Defend Chinese textual adversarial attacks via prompt-tuning
Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Eng. Appl. Artif. Intell.5
2025 Soft Prompt-tuning with Self-Resource Verbalizer for short text streams
Yi Zhu 0006, Ye Wang 0022, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001
Eng. Appl. Artif. Intell.4
2025 Robust and semantic-faithful post-hoc watermarking of text generated by black-box language models
Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Xiaocheng Hu, Xiaoye Ouyang
Frontiers Comput. Sci.2
2025 Domain adaptation for textual adversarial defense via prompt-tuning
Yi Zhu 0006, Chenqi Hua, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Neurocomputing6
2025 Multi-modal soft prompt-tuning for Chinese Clickbait Detection
Ye Wang 0022, Yi Zhu 0006, Yun Li 0010, Liting Wei, Yun-Hao Yuan 0001, Jipeng Qiang
Neurocomputing6
2025 Soft prompt-tuning for unsupervised domain adaptation via self-supervision
Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Neurocomputing5
2024 Incomplete Multi-Kernel k-Means Clustering With Fractional-Order Embedding
abstract
Multiple kernel clustering (MKC) has received increasing attention in the community of machine learning, which takes advantage of multiple pre-specified kernels to perform clustering tasks. Traditional MKC algorithms cannot effectively deal with the incomplete views where some samples are missing. Thus, incomplete MKC (IMKC) has been developed to solve this problem and obtained promising results. Nevertheless, the samples may be noisy or limited in real-world applications, which will result in the performance deterioration of existing IMKC algorithms. To address this issue, in this paper we propose a simple yet effective clustering method for incomplete data, termed fractional-order embedding incomplete multi-kernel k-means clustering (FE-MKKM-IK). Specifically, FE-MKKM-IK introduces the idea of fractional-order embedding to reconstruct the kernel matrix computed by the samples. On this basis, a new incomplete multiple kernel k-means clustering is developed. Performance evaluation is conducted on four widely used datasets, which shows that FE-MKKM-IK is effective to cluster the incomplete data.
Deheng Xu, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Yi Zhu 0006
IEEE Big Data4
2024 Learning Spectral Canonical ℱ-Correlation Representation for Face Super-Resolution
abstract
Face super-resolution (FSR) is a powerful technique for restoring high-resolution face images from the captured low-resolution ones with the assistance of prior information. Existing FSR methods based on explicit or implicit covariance matrices are difficult to reveal complex nonlinear relationships between features, as conventional covariance computation is essentially a linear operation process. Besides, the limited number of training samples and noise disturbance lead to the deviation of sample covariance matrices. To solve these issues, we propose a novel FSR method via using spectral canonical ℱ-correlation representation. The proposed method first defines intra-resolution and inter-resolution covariation matrices by considering the nonlinear relationship between different features, and then uses the fractional order idea to rebuild covariation matrices. The qualitative and quantitative results have validated the superiority of the proposed method.
Yun-Hao Yuan 0001, Mingzhi Hao, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001
ICASSP4
2024 Face Super-Resolution Using Covariation-Guided Orthonormalized Partial Least Squares
Mingzhi Hao, Yun-Hao Yuan 0001, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010, Runmei Zhang
ICONIP (8)3
2024 Short text classification with Soft Knowledgeable Prompt-tuning
Yi Zhu 0006, Ye Wang 0022, Jianyuan Mu, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001, Xindong Wu 0001
Expert Syst. Appl.5
2024 Representation learning: serial-autoencoder for personalized recommendation
Yi Zhu 0006, Yishuai Geng, Yun Li 0010, Jipeng Qiang, Xindong Wu 0001
Frontiers Comput. Sci.4
2024 Prompt-Learning for Short Text Classification
abstract
In the short text, the extremely short length, feature sparsity, and high ambiguity pose huge challenges to classification tasks. Recently, as an effective method for tuning Pre-trained Language Models for specific downstream tasks, prompt-learning has attracted a vast amount of attention and research. The main intuition behind the prompt-learning is to insert the template into the input and convert the tasks into equivalent cloze-style tasks. However, most prompt-learning methods only consider the class name and monotonous strategy for knowledge incorporating in cloze-style prediction, which will inevitably incur omissions and bias in short text classification tasks. In this paper, we propose a short text classification method with prompt-learning. Specifically, the top$M$concepts related to the entity in the short text are retrieved from the open Knowledge Graph like Probase, these concepts are first selected by the distance with class labels, which takes both the short text itself and the class name into consideration during expanding label word space. Then, we conducted four additional strategies for the integration of the expanded concepts, and the union of these concepts are adopted finally in the verbalizer of prompt-learning. Experimental results show that the obvious improvement is obtained compared with other state-of-the-art methods on five well-known datasets.
Yi Zhu 0006, Ye Wang 0022, Jipeng Qiang, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.3
2024 Iterative Soft Prompt-Tuning for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation aims to facilitate learning tasks in unlabeled target domain with knowledge in the related source domain, which has achieved awesome performance with the pre-trained language models (PLMs). Recently, inspired by GPT, the prompt-tuning model has been widely explored in stimulating rich knowledge in PLMs for language understanding. However, existing prompt-tuning methods still directly applied the model that was learned in the source domain into the target domain to minimize the discrepancy between different domains, e.g., the prompts or the template are trained separately to learn embeddings for transferring to the target domain, which is actually the intuition of end-to-end deep-based approach. In this paper, we propose an Iterative Soft Prompt-Tuning method (ItSPT) for better unsupervised domain adaptation. On the one hand, the prompt-tuning model learned in the source domain is converted into an iterative model to find the true label information in the target domain, the domain adaptation method is then regarded as a few-shot learning task. On the other hand, instead of hand-crafted templates, ItSPT adopts soft prompts for both considering the automatic template generation and classification performance. Experiments on both English and Chinese datasets demonstrate that our method surpasses the performance of SOTA methods.
Yi Zhu 0006, Jipeng Qiang, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.3
2023 ParaLS: Lexical Substitution via Pretrained Paraphraser
abstract
Lexical substitution (LS) aims at finding appropriate substitutes for a target word in a sentence.Recently, LS methods based on pretrained language models have made remarkable progress, generating potential substitutes for a target word through analysis of its contextual surroundings.However, these methods tend to overlook the preservation of the sentence's meaning when generating the substitutes.This study explores how to generate the substitute candidates from a paraphraser, as the generated paraphrases from a paraphraser contain variations in word choice and preserve the sentence's meaning.Since we cannot directly generate the substitutes via commonly used decoding strategies, we propose two simple decoding strategies that focus on the variations of the target word during decoding.Experimental results show that our methods outperform state-of-theart LS methods based on pre-trained language models on three benchmarks.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006
ACL (1)1
2023 Multilingual Lexical Simplification via Paraphrase Generation
abstract
Lexical simplification (LS) methods based on pretrained language models have made remarkable progress, generating potential substitutes for a complex word through analysis of its contextual surroundings. However, these methods require separate pretrained models for different languages and disregard the preservation of sentence meaning. In this paper, we propose a novel multilingual LS method via paraphrase generation, as paraphrases provide diversity in word selection while preserving the sentence’s meaning. We regard paraphrasing as a zero-shot translation task within multilingual neural machine translation that supports hundreds of languages. After feeding the input sentence into the encoder of paraphrase modeling, we generate the substitutes based on a novel decoding strategy that concentrates solely on the lexical variations of the complex word. Experimental results demonstrate that our approach surpasses BERT-based methods and zero-shot GPT3-based method significantly on English, Spanish, and Portuguese.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006, Kaixun Hua
ECAI2
2023 Chinese Lexical Substitution: Dataset and Method
abstract
Existing lexical substitution (LS) benchmarks were collected by asking human annotators to think of substitutes from memory, resulting in benchmarks with limited coverage and relatively small scales.To overcome this problem, we propose a novel annotation method to construct an LS dataset based on human and machine collaboration.Based on our annotation method, we construct the first Chinese LS dataset CHNLS which consists of 33,695 instances and 144,708 substitutes, covering three text genres (News, Novel, and Wikipedia).Specifically, we first combine four unsupervised LS methods as an ensemble method to generate the candidate substitutes, and then let human annotators judge these candidates or add new ones.This collaborative process combines the diversity of machine-generated substitutes with the expertise of human annotators.Experimental results that the ensemble method outperforms other LS methods.To our best knowledge, this is the first study for the Chinese LS task.
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xiaocheng Hu, Xiaoye Ouyang
EMNLP1
2023 Learning Supervised Covariation Projection Through General Covariance
abstract
Canonical correlation analysis (CCA) is a classical yet powerful tool for learning two-view feature representation in various fields. But, most CCA approaches are based on the conventional covariance measure, which makes them difficult to uncover the complicatedly nonlinear relationship between distinct features. In this paper, we address the preceding problem and propose two novel CCA approaches in a supervised manner by using a general covariance metric. The proposed approaches not only consider the label information of training data, but also the nonlinear relationship between different features rather than samples, which leads to greater flexibility in many practical applications. A series of experimental results on five benchmark datasets demonstrate the effectiveness of our proposed methods in terms of classification accuracy.
Xiangze Bao, Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006
ICASSP4
2023 Many Is Better Than One: Multiple Covariation Learning for Latent Multiview Representation
Yun-Hao Yuan 0001, Pengwei Qian, Jin Li 0028, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010
ICONIP (9)4
2023 Natural language watermarking via paraphraser-based lexical substitution
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001
Artif. Intell.1
2023 Lexical simplification via single-word generation
Jipeng Qiang, Yang Li 0186, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006
Frontiers Comput. Sci.1
2023 Unsupervised statistical text simplification using pre-trained language modeling for initialization
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006, Xindong Wu 0001
Frontiers Comput. Sci.1
2023 Safeguarding text generation API's intellectual property through meaning-preserving lexical watermarks
Yun Li 0010, Xiaoye Ouyang, Xiaocheng Hu, Jipeng Qiang
Frontiers Comput. Sci.5
2023 Representation learning via an integrated autoencoder for unsupervised domain adaptation
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Yun-Hao Yuan 0001, Yun Li 0010
Frontiers Comput. Sci.3
2023 A hybrid classification method via keywords screening and attention mechanisms in extreme short text
abstract
Short text classification has provoked a vast amount of attention and research in recent decades. However, most existing methods only focus on the short texts that contain dozens of words like Twitter and Microblog, while pay far less attention to the extreme short texts like news headline and search snippets. Meanwhile, contemporary short text classification methods that extend the features via external knowledge sources always introduce lots of useless concepts, which may be detrimental to classification performance. Moreover, unlike traditional short text classification methods, the classification results of extreme short texts are often determined by a few even one or two keywords. To address these problems, we propose a novel hybrid classification method via Keywords Screening and Attention Mechanisms in extreme short text, called KSAM. More specifically, firstly, the attention-based BiLSTM is introduced in our method to enhance the role of keywords. Secondly, we screen the keywords in the extreme short text for obtaining the true class label, and the concepts concerning the keywords are retrieved from external open knowledge sources like DBpedia. Thirdly, the attention mechanisms are introduced to acquire the weight of these retrieved concepts. Finally, conceptual information is utilized to assist the classification of the extreme short text. Extensive experiments have demonstrated the effectiveness of our method compared to other state-of-the-art methods.
Xinke Zhou, Yi Zhu 0006, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001, Xingdong Wu, Runmei Zhang
Intell. Data Anal.4
2023 Fuzzy clustering analysis for the loan audit short texts
Zhidong Liu, Jipeng Qiang, Zhuangyi Zhang
Knowl. Inf. Syst.3
2023 Chinese Idiom Paraphrasing
abstract
Abstract Idioms are a kind of idiomatic expression in Chinese, most of which consist of four Chinese characters. Due to the properties of non-compositionality and metaphorical meaning, Chinese idioms are hard to be understood by children and non-native speakers. This study proposes a novel task, denoted as Chinese Idiom Paraphrasing (CIP). CIP aims to rephrase idiom-containing sentences to non-idiomatic ones under the premise of preserving the original sentence’s meaning. Since the sentences without idioms are more easily handled by Chinese NLP systems, CIP can be used to pre-process Chinese datasets, thereby facilitating and improving the performance of Chinese NLP tasks, e.g., machine translation systems, Chinese idiom cloze, and Chinese idiom embeddings. In this study, we can treat the CIP task as a special paraphrase generation task. To circumvent difficulties in acquiring annotations, we first establish a large-scale CIP dataset based on human and machine collaboration, which consists of 115,529 sentence pairs. In addition to three sequence-to-sequence methods as the baselines, we further propose a novel infill-based approach based on text infilling. The results show that the proposed method has better performance than the baselines based on the established CIP dataset.
Jipeng Qiang, Yang Li 0186, Chaowei Zhang 0001, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001
Trans. Assoc. Comput. Linguistics1
2022 Learning Canonical F-Correlation Projection for Compact Multiview Representation
abstract
Canonical correlation analysis (CCA) matters in multi-view representation learning. But, CCA and its most variants are essentially based on explicit or implicit covariance matrices. It means that they have no ability to model the nonlinear relationship among features due to intrinsic linearity of covariance. In this paper, we address the preceding problem and propose a novel canonical F-correlation framework by exploring and exploiting the nonlinear relationship between different features. The framework projects each feature rather than observation into a certain new space by an arbitrary nonlinear mapping, thus resulting in more flexibility in real applications. With this frame-work as a tool, we propose a correlative covariation projection (CCP) method by using an explicit nonlinear mapping. Moreover, we further propose a multiset version of CCP dubbed MCCP for learning compact representation of more than two views. The proposed MCCP is solved by an iterative method, and we prove the convergence of this iteration. A series of experimental results on six benchmark datasets demonstrate the effectiveness of our proposed CCP and MCCP methods.
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001, Jianping Gou
CVPR4
2022 Dynamic clustering for short text stream based on Dirichlet process
Wanyin Xu, Yun Li 0010, Jipeng Qiang
Appl. Intell.3
2022 Representation learning with deep sparse auto-encoder for multi-task learning
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Xuegang Hu, Yuhong Zhang 0002, Pei-Pei Li 0001
Pattern Recognit.3
2022 Short Text Topic Modeling Techniques, Applications, and Performance: A Survey
abstract
Analyzing short texts infers discriminative and coherent latent topics that is a critical and fundamental task since many real-world applications require semantic understanding of short texts. Traditional long text topic modeling algorithms (e.g., PLSA and LDA) based on word co-occurrences cannot solve this problem very well since only very limited word co-occurrence information is available in short texts. Therefore, short text topic modeling has already attracted much attention from the machine learning research community in recent years, which aims at overcoming the problem of sparseness in short texts. In this survey, we conduct a comprehensive review of various short text topic modeling techniques proposed in the literature. We present three categories of methods based on Dirichlet multinomial mixture, global word co-occurrences, and self-aggregation, with example of representative approaches in each category and analysis of their performance on various tasks. We develop the first comprehensive open-source library, called STTM, for use in Java that integrates all surveyed algorithms within a unified interface, benchmark datasets, to facilitate the expansion of new methods in this research field. Finally, we evaluate these state-of-the-art methods on many real-world datasets and compare their performance against one another and versus long text topic modeling algorithm.
Jipeng Qiang, Zhenyu Qian 0006, Yun Li 0010, Yun-Hao Yuan 0001, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.1
2022 Novel Discrete-Time Recurrent Neural Networks Handling Discrete-Form Time-Variant Multi-Augmented Sylvester Matrix Problems and Manipulator Application
abstract
In this article, the discrete-form time-variant multi-augmented Sylvester matrix problems, including discrete-form time-variant multi-augmented Sylvester matrix equation (MASME) and discrete-form time-variant multi-augmented Sylvester matrix inequality (MASMI), are formulated first. In order to solve the above-mentioned problems, in continuous time-variant environment, aided with the Kronecker product and vectorization techniques, the multi-augmented Sylvester matrix problems are transformed into simple linear matrix problems, which can be solved by using the proposed discrete-time recurrent neural network (RNN) models. Second, the theoretical analyses and comparisons on the computational performance of the recently developed discretization formulas are presented. Based on these theoretical results, a five-instant discretization formula with superior property is leveraged to establish the corresponding discrete-time RNN (DTRNN) models for solving the discrete-form time-variant MASME and discrete-form time-variant MASMI, respectively. Note that these DTRNN models are zero stable, consistent, and convergent with satisfied precision. Furthermore, illustrative numerical experiments are given to substantiate the excellent performance of the proposed DTRNN models for solving discrete-form time-variant multi-augmented Sylvester matrix problems. In addition, an application of robot manipulator further extends the theoretical research and physical realizability of RNN methods.
Yang Shi 0003, Long Jin 0001, Shuai Li 0002, Jian Li 0018, Jipeng Qiang, Dimitrios Gerontitis
IEEE Trans. Neural Networks Learn. Syst.5
2021 Fractional Multi-view Hashing with Semantic Correlation Maximization
Ruijie Gao, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Yi Zhu 0006
ICONIP (5)4
2021 Multi-view Fractional Deep Canonical Correlation Analysis for Subspace Clustering
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001
ICONIP (2)4
2021 Domain Adaptation with Stacked Convolutional Sparse Autoencoder
Yi Zhu 0006, Xinke Zhou, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001
ICONIP (5)4
2021 Composite nonlinear multiset canonical correlation analysis for multiview feature learning and recognition
abstract
Summary In this paper, we propose a composite nonlinear multiset canonical correlation projections (CNMCPs) framework where orthogonal constraints are imposed in each set. This makes CNMCP capable of learning uncorrelated low‐dimensional features with minimum redundancy in Hilbert space. With the CNMCP framework, we further present a particular algorithm called multikernel multiset canonical correlations or mKMCC, which introduces different weights into multiple nonlinear functions in all views. An alternating iterative optimization is designed for computational solution. Numerous experimental results on practical datasets have demonstrated the effectiveness and robustness of mKMCC, in contrast with existing kernel correlation learning approaches.
Yun-Hao Yuan 0001, Xiaobo Shen 0001, Yun Li 0010, Bin Li 0006, Jianping Gou, Jipeng Qiang, Xinfeng Zhang 0003, Quan-Sen Sun
Concurr. Comput. Pract. Exp.6
2021 Representation learning with collaborative autoencoder for personalized recommendation
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Yun-Hao Yuan 0001, Yun Li 0010
Expert Syst. Appl.3
2021 OPLS-SR: A novel face super-resolution learning method using orthonormalized coherent features
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jipeng Qiang, Bin Li 0006, Wankou Yang, Furong Peng
Inf. Sci.4
2021 Learning Unsupervised and Supervised Representations via General Covariance
abstract
Component analysis (CA) is a powerful technique for learning discriminative representations in various computer vision tasks. Typical CA methods are essentially based on the covariance matrix of training data. But, the covariance matrix has obvious disadvantages such as failing to model complex relationship among features and singularity in small sample size cases. In this letter, we propose a general covariance measure to achieve better data representations. The proposed covariance is characterized by a nonlinear mapping determined by domain-specific applications, thus leading to more advantages, flexibility, and applicability in practice. With general covariance, we further present two novel CA methods for learning compact representations and discuss their differences from conventional methods. A series of experimental results on nine benchmark data sets demonstrate the effectiveness of the proposed methods in terms of accuracy.
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jianping Gou, Jipeng Qiang
IEEE Signal Process. Lett.5
2021 Chinese Lexical Simplification
abstract
Lexical simplification has attracted much attention in many languages, which is the process of replacing complex words in a given sentence with simpler alternatives of equivalent meaning. Although the richness of vocabulary in Chinese makes the text very difficult to read for children and non-native speakers, there is no research work for the Chinese lexical simplification (CLS) task. To circumvent difficulties in acquiring annotations, we manually create the first benchmark dataset for CLS, which can be used for evaluating the lexical simplification systems automatically. To acquire a more thorough comparison, we present five different types of methods as baselines to generate substitute candidates for the complex word that includes synonym-based approach, word embedding-based approach, BERT-based approach, sememe-based approach, and a hybrid approach. Finally, we design the experimental evaluation of these baselines and discuss their advantages and disadvantages. To our best knowledge, this is the first study for CLS task.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Xindong Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 LSBert: Lexical Simplification Based on BERT
abstract
Lexical simplification (LS) aims at replacing complex words with simpler alternatives. LS commonly consists of three main steps: complex word identification, substitute generation, and substitute ranking. Existing LS methods focus on the contextual information of the complex word in the last step (substitute ranking). However, they miss out the following two facts: (1) The word complexity of a polysemous word is very closely related to its context; (2) The step of substitute generation regardless of the context will inevitably produce a large number of spurious candidates. Therefore, we propose a novel LS system LSBert based on pretrained language model BERT to address the aforementioned issues, which is capable of making use of the wider context when both identifying the words in need of simplification and generating substitute candidates for the complex words. Specifically, LSBert consists of a network for complex word identification by fine-tuning BERT and a network for substitute generation based on BERT. Experimental results show that LSBert performs well in both complex word identification and substitute generation, achieving state-of-the-art results in three benchmarks. To facilitate reproducibility, the code of the LSBert system is available at https://github.com/qiang2100/BERT-LS.
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Yang Shi 0003, Xindong Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Unsupervised Statistical Text Simplification
abstract
Most recent approaches for Text Simplification (TS) have drawn on insights from machine translation to learn simplification rewrites from the monolingual parallel corpus of complex and simple sentences, yet their effectiveness strongly relies on large amounts of parallel sentences. However, there has been a serious problem haunting TS for decades, that is, the availability of parallel TS corpora is scarce or not fit for the learning task. In this paper, we will focus on one especially useful and challenging problem of unsupervised TS without a single parallel sentence. To the best of our knowledge, we present the first unsupervised text simplification system based on phrase-based machine translation system, which leverages a careful initialization of phrase tables and language models. On the widely used WikiLarge and WikiSmall benchmarks, our system respectively obtains 39.08 and 25.12 SARI points, even outperforms some supervised baselines.
Jipeng Qiang, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.1
2020 Lexical Simplification with Pretrained Encoders
abstract
Lexical simplification (LS) aims to replace complex words in a given sentence with their simpler alternatives of equivalent meaning. Recently unsupervised lexical simplification approaches only rely on the complex word itself regardless of the given sentence to generate candidate substitutions, which will inevitably produce a large number of spurious candidates. We present a simple LS approach that makes use of the Bidirectional Encoder Representations from Transformers (BERT) which can consider both the given sentence and the complex word during generating candidate substitutions for the complex word. Specifically, we mask the complex word of the original sentence for feeding into the BERT to predict the masked token. The predicted results will be used as candidate substitutions. Despite being entirely unsupervised, experimental results show that our approach obtains obvious improvement compared with these baselines leveraging linguistic databases and parallel corpus, outperforming the state-of-the-art by more than 12 Accuracy points on three well-known benchmarks.
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001
AAAI1
2020 Learning Fractional Orthogonal Latent Consistent Features for Face Hallucination and Recognition
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jipeng Qiang, Bin Li 0006
ICASSP4
2020 Regularized Multiset Neighborhood Correlation Analysis for Semi-paired Multiview Learning
Yun-Hao Yuan 0001, Zhaoqi Wu, Yun Li 0010, Jipeng Qiang, Jianping Gou, Yi Zhu 0006
ICONIP (2)4
2019 Learning Super-Resolution Coherent Facial Features Using Nonlinear Multiset PLS for Low-Resolution Face Recognition
abstract
Face hallucination (FH) is an effective technique for super-resolving low-resolution (LR) face images. In real-world applications, a face image usually has multiple distinct low resolutions. Most existing FH methods can not effectively deal with multiple LR views simultaneously. To solve this issue, we present a multi-set partial least squares (MPLS) approach and its kernel extension for jointly learning the nonlinear consistency of multi-resolution facial features. With nonlinear MPLS, we present a novel simultaneous super-resolution coherent facial feature method for the face images with multiple LRs, which has capacity of jointly learning the nonlinear relationships between multiple facial resolutions. Experimental results demonstrate the effectiveness and robustness of our proposed FH method.
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jianping Gou, Jipeng Qiang, Quan-Sen Sun
ICIP5
2019 Learning Simultaneous Face Super-Resolution Using Multiset Partial Least Squares
abstract
Face super-resolution (FSR) is an effective way to solve low-resolution (LR) problems in face analysis. But, most FSR methods only consider that LR face images have a single resolution, which is usually not consistent with practical situations due to the existence of multiple resolutions. To date, simultaneously learning the mappings from multiple LRs to high resolution (HR) has not been given proper attention. To solve this issue, we first propose a multi-set partial least squares (MPLS) approach to jointly deal with multi-set random variables via a recursive optimization. With MPLS, we then present a novel FSR method called MPLS-FH to simultaneously learn multiple resolution-specific mappings for various LR views from the same source. Concretely, MPLS-FH first divides multi-resolution face images into many patches. Then, it jointly learns the latent coherent features of principal-component embeddings of multi-resolution patches. Last, it super-resolves the input LR face by cross-resolution neighborhood search. Experimental results demonstrate the effectiveness of the proposed method in terms of quantitative and qualitative evaluations.
Yun-Hao Yuan 0001, Jin Li 0028, Jianping Gou, Yun Li 0010, Jipeng Qiang, Bin Li 0006
ICME5
2019 D2PLS: A Novel Bilinear Method for Facial Feature Fusion
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Bin Li 0006, Jianping Gou
ICONIP (4)4
2019 Fuzzy Bilinear Latent Canonical Correlation Projection for Feature Learning
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Jianping Gou, Guangwei Gao, Bin Li 0006
ICONIP (1)4
2019 A practical algorithm for solving the sparseness problem of short text clustering
abstract
Dirichlet Multinomial Mixture (DMM) models have been successful in clustering short texts. However, the word co-occurrence information that can be captured by these models is limited to the short text corpus itself. If two words have strong relatedness but rarely co-occurring in short texts, these models can not fully capture the semantic relatedness between the two words. In this paper, we propose a novel model by incorporating word-word correlation into DMM, called WDMM. By constructing a sparse graph using word-word relationship, our model expands each short text using their neighboring words in each text that can help to solve the problem of sparseness in short texts. Therefore, the cluster label of each text is not only influenced by its words, but decided by their similar words in this corpus. Experimental results on real-world datasets demonstrated the substantial superiority of our WDMM model over the state-of-the-art methods.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Wei Liu 0010, Xindong Wu 0001
Intell. Data Anal.1
2019 Heterogeneous-Length Text Topic Modeling for Reader-Aware Multi-Document Summarization
abstract
More and more user comments like Tweets are available, which often contain user concerns. In order to meet the demands of users, a good summary generating from multiple documents should consider reader interests as reflected in reader comments. In this article, we focus on how to generate a summary from multi-document documents by considering reader comments, named as reader-aware multi-document summarization (RA-MDS). We present an innovative topic-based method for RA-MDA, which exploits latent topics to obtain the most salient and lessen redundancy summary from multiple documents. Since finding latent topics for RA-MDS is a crucial step, we also present a Heterogeneous-length Text Topic Modeling (HTTM) to extract topics from the corpus that includes both news reports and user comments, denoted as heterogeneous-length texts. In this case, the latent topics extract by HTTM cover not only important aspects of the event, but also aspects that attract reader interests. Comparisons on summary benchmark datasets also confirm that the proposed RA-MDS method is effective in improving the quality of extracted summaries. In addition, experimental results demonstrate that the proposed topic modeling method outperforms existing topic modeling algorithms.
Jipeng Qiang, Ping Chen 0001, Wei Ding 0003, Tong Wang 0007, Fei Xie 0002, Xindong Wu 0001
ACM Trans. Knowl. Discov. Data1
2018 Text Simplification with Self-Attention-Based Pointer-Generator Networks
Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001
ICONIP (5)3
2018 Supervised Two-Dimensional CCA for Multiview Data Representation
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Wenyan Bao
ICONIP (5)4
2018 Learning Parallel Canonical Correlations for Scale-Adaptive Low Resolution Face Recognition
abstract
Low resolution is one of the main obstacles in the application of face recognition. Although many methods have been proposed to improve the problem, they assume that low-resolution (LR) face images have a uniform scale. In real scenarios, this prerequisite is very harsh. In this paper, we propose a scale-adaptive LR face recognition approach based on two-dimensional multi-set canonical correlation analysis (2DM-CCA), where face image matrix does not need to be previously transformed into a vector. In the proposed method, training sets with different resolutions are treated as different views, and then projected in parallel into a latent coherent space where the consistency of multi-view face data is maximally enhanced. When a new LR face image with an arbitrary scale is input, we first transform it by using the left and right projection matrices of an appropriate training view, and then reconstruct its high resolution facial feature by neighborhood reconstruction. Experimental results show that our proposed method is more effective and efficient than several existing methods.
Yun-Hao Yuan 0001, Zhao Zhang 0018, Yun Li 0010, Jipeng Qiang, Bin Li 0006, Xiaobo Shen 0001
ICPR4
2018 Snapshot ensembles of non-negative matrix factorization for stability of topic modeling
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Wei Liu 0010
Appl. Intell.1
2018 Short text clustering based on Pitman-Yor process mixture model
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Xindong Wu 0001
Appl. Intell.1
2017 Supervised Deep Canonical Correlation Analysis for Multiview Feature Learning
Yan Liu 0038, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Min Ruan, Zhao Zhang 0018
ICONIP (6)4
2017 Face Hallucination and Recognition Using Kernel Canonical Correlation Analysis
Zhao Zhang 0018, Yun-Hao Yuan 0001, Yun Li 0010, Bin Li 0006, Jipeng Qiang
ICONIP (6)5
2017 Identifying the Number of Clusters in Short Text Using Bayesian Nonparametric Model
abstract
Before inferring the real number of clusters in short text clustering, Dirichlet Multinomial Mixture (DMM) model makes assumption that there are at most Kmax clusters. In some cases, it is difficult to choose a proper Kmax beforehand. In the paper, we propose a novel model based on Pitman-Yor Process to capture the power-law phenomenon of the cluster distribution. Specifically, each text chooses one of the active clusters or a new cluster with probabilities derived from the Pitman-Yor Process Mixture model (PYPM). Different from DMM model, our model does not require Kmax as input. Discriminative words and nondiscriminative words are identified automatically to help enhance text clustering. Parameters are estimated efficiently by collapsed Gibbs sampling. The experiments on real-world datasets validate the effectiveness of the proposed model in comparison with other state-of-theart models.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Tong Wang 0007
ICTAI1
2017 Topic Modeling over Short Texts by Incorporating Word Embeddings
Jipeng Qiang, Ping Chen 0001, Tong Wang 0007, Xindong Wu 0001
PAKDD (2)1
2016 Text Simplification Using Neural Machine Translation
abstract
Text simplification (TS) is the technique of reducing the lexical, syntactical complexity of text. Existing automatic TS systems can simplify text only by lexical simplification or by manually defined rules. Neural Machine Translation (NMT) is a recently proposed approach for Machine Translation (MT) that is receiving a lot of research interest. In this paper, we regard original English and simplified English as two languages, and apply a NMT model–Recurrent Neural Network (RNN) encoder-decoder on TS to make the neural network to learn text simplification rules by itself. Then we discuss challenges and strategies about how to apply a NMT model to the task of text simplification.
Tong Wang 0007, Ping Chen 0001, John Rochford, Jipeng Qiang
AAAI4
2016 Topic Discovery from Heterogeneous Texts
abstract
Recently many topic models such as Latent Dirich-let Allocation (LDA) have made important progress towards generating high-level knowledge from a large corpus. They assume that a text consists of a mixture of topics, which is usually the case for regular articles but may not hold for a short text that usually contains only one topic. In practice, a corpus may include both short texts and long texts, in this case neither methods developed for only long texts nor methods for only short texts can generate satisfying results. In this paper, we present an innovative method to discover latent topics from a heterogeneous corpus including both long and short texts. A new topic model based on collapsed Gibbs sampling algorithm is developed for modeling such heterogeneous texts. The experiments on real-world datasets validate the effectiveness of the proposed model in comparison with other state-of-the-art models.
Jipeng Qiang, Ping Chen 0001, Wei Ding 0003, Tong Wang 0007, Fei Xie 0002, Xindong Wu 0001
ICTAI1
2016 Multi-document summarization using closed patterns
Jipeng Qiang, Ping Chen 0001, Wei Ding 0003, Fei Xie 0002, Xindong Wu 0001
Knowl. Based Syst.1
2014 Pattern Matching with Flexible Wildcards
Xindong Wu 0001, Jipeng Qiang, Fei Xie 0002
J. Comput. Sci. Technol.2