Yun Li 0010

dblp:87/6284-10 · DBLP profile ↗
← Back
97ranked-venue papers
4as first author
62since 2021 · last 2026
0000-0002-7628-0358ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 13 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Security and privacy · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 MedVerse: Efficient and Reliable Medical Reasoning via DAG-Structured Parallel Execution
abstract
Jianwen Chen, Xinyu Yang, Peng Xia, Arian Azarang, Yueh Z Lee, Gang Li, Hongtu Zhu, Yun Li, Beidi Chen, Huaxiu Yao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinyu Yang 0002, Peng Xia 0005, Arian Azarang, Yueh Z. Lee, Gang Li 0001, Hongtu Zhu, Yun Li 0010, Beidi Chen, Huaxiu Yao
ACL (1)8
2026 Secure Lookup Tables: Faster, Leaner, and More General
Chongrong Li, Yun Li 0010, Zhanpeng Guo, Yuncong Hu, Cheng Hong 0001
SP3
2026 SubAttack: A word-level adversarial textual attack method via antonym substitution
abstract
Over the past few years, various word-level textual attack approaches have been proposed to reveal the vulnerability in existing deep neural networks and even large language models (LLMs) for Natural Language Processing (NLP). The textual attack aims to fool existing models into making erroneous predictions by altering the text without affecting the user’s understanding. However, current methods either struggle to construct semantically preserved adversarial texts and altered the semantics of the original text, or fail to consider the semantic perturbation constraints and are prone to invalid adversarial examples. In this paper, we propose an efficient and effective framework SubAttack to address these issues. SubAttack is a word-level adversarial textual attack method via antonym substitution, which replaces semantic indicator keywords to generate high-quality adversarial samples with considering both semantically preservation and semantic perturbation. Specifically, the process first involves tokenizing the text and performing part-of-speech tagging Identifying the semantic indicator keywords. Then, the antonym ranking is designed to decide the substitutions of candidate words to fit the context. Finally, while retaining the original text, the ranked antonyms are integrated into the text and the instructions are added for both semantically preservation and semantic perturbation. Extensive experiments reveal that state-of-the-art (SOTA) LLMs (e.g. Llama and QWen) are still vulnerable to our SubAttack. Further experiments show that the adversarial examples crafted by SubAttack usually have higher quality, exhibit better fluency and barely affect human performance and can bring more robustness improvement to victim models by adversarial training.
Chenqi Hua, Yi Zhu 0006, Chaowei Zhang 0001, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Eng. Appl. Artif. Intell.5
2026 Personalized recommendation with clustering via prompt-tuning
abstract
The personalized recommendation aims to address the information overload problem, which can find interesting items for users from massive amounts of information. The research paradigm of personalized recommendation evolved from deep neural networks to pre-trained language models (PLMs) like BERT and, more recently, into large language models (LLMs). However, it is always very difficult to find the target item among a massive number of data or information, which is not only time-consuming but also often has low accuracy. In this paper, we propose a Personalized Recommendation method with Clustering via Prompt-tuning (PRCP), a candidate item set is developed and a prompt-tuning model with a designed verbalizer is constructed for recommendation. Specifically, the target users are first selected by the similarity calculation, and items are then clustered by the preferences of similar users to form a candidate item set. Then the prompt-tuning model is introduced to predict the masked label for candidate items, and three different strategies are designed to expand the label word space for verbalizer optimization. Extensive experiments conducted on three datasets validated the effectiveness of the proposed method compared to other state-of-the-art baselines including LLMs.
Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Intell. Data Anal.4
2025 Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News Detection
abstract
The questionable responses caused by knowledge hallucination may lead to LLMs' unstable ability in decision-making. However, it has never been investigated whether the LLMs' hallucination is possibly usable for generating negative reasoning to assist fake news detection. In this paper, we propose a novel supervised self-reinforced reasoning rectification approach - SR^3 that not only yields common reasonable reasoning for news but also forces LLMs to generate the wrong understandings of news via LLMs reflection for semantic consistency learning. Upon that, we construct a negative reasoning-based news learning model called - NRFE, which leverages positive or negative news-reasoning pairs for learning the semantic consistency between them. To avoid the impact of label-implicated reasoning, we deploy a student model - NRFE-D that only takes news content as input to inspect the performance of our method by distilling the knowledge from NRFE. The experimental results verified on three popular fake news datasets demonstrate the superiority of our method compared with three kinds of baselines including prompting-based LLMs, fine-tuning-based PLMs, and other representative fake news detection methods.
Chaowei Zhang 0001, Zongling Feng, Zewei Zhang, Jipeng Qiang, Guandong Xu, Yun Li 0010
AAAI6
2025 Collaborative Document Simplification Using Multi-Agent Systems
abstract
Research on text simplification has been ongoing for many years. However, the task of document simplification (DS) remains a significant challenge due to the need to consider complex factors such as technical terminology, metaphors, and overall coherence. In this work, we introduce a novel multi-agent framework for document simplification (AgentSimp) based on large language models (LLMs). This framework emulates the collaborative process of a human expert team through the roles played by multiple agents, addressing the intricate demands of document simplification. We explore two communication strategies among agents (pipeline-style and synchronous) and two document reconstruction strategies (Direct and Iterative ). According to both automatic evaluation metrics and human evaluation results, the documents simplified by AgentSimp are deemed to be more thoroughly simplified and more coherent on a variety of articles across different types and styles.
Dengzhao Fang, Jipeng Qiang, Xiaoye Ouyang, Yi Zhu 0006, Yun-Hao Yuan 0001, Yun Li 0010
COLING6
2025 Post-Hoc Watermarking for Robust Detection in Text Generated by Large Language Models
abstract
Research on text simplification has been ongoing for many years, yet document simplification remains a significant challenge due to the need to address complex factors such as technical terminology, metaphors, and overall coherence. In this work, we introduce a novel multi-agent framework AgentSimp for document simplification, based on large language models. This framework simulates the collaborative efforts of a team of human experts through the roles played by multiple agents, effectively meeting the intricate demands of document simplification. We investigate two communication strategies among agents (pipeline-style and synchronous) and two document reconstruction strategies (Direct and Iterative). According to both automatic evaluation metrics and human evaluation results, AgentSimp produces simplified documents that are more thoroughly simplified and more coherent across various articles and styles.
Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Xiaoye Ouyang
COLING4
2025 Learning Simultaneous Facial Canonical Correlation Representation for Face Hallucination
abstract
The low resolution (LR) problem is rather challenging in face analysis. Most existing face hallucination methods assume that LR face images have only one resolution, but multiple resolutions may be available from different sources. To solve this issue, we propose a novel simultaneous facial canonical correlation representation learning method for face hallucination, which seeks latent correlation subspaces for multi-resolution views. Our method jointly solves multiple linear transformations by optimizing a correlation summation criterion of all pairs of resolutions. The neighborhood reconstruction is used to infer the HR facial canonical correlation representation of LR face inputs. Extensive experimental results show the superiority of our proposed method in terms of quantitative and qualitative evaluations.
Yun-Hao Yuan 0001, Jin Li 0028, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001, Yun Li 0010
ICASSP6
2025 MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
abstract
The advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) encounter factuality challenges due to modality misalignment, where the models prioritize textual knowledge over visual input, leading to hallucinations that contradict information in medical images. Previous attempts to enhance modality alignment in Med-LVLMs through preference optimization have inadequately addressed clinical relevance in preference data, making these samples easily distinguishable and reducing alignment effectiveness. In response, we propose MMedPO, a novel multimodal medical preference optimization approach that considers the clinical relevance of preference samples to enhance Med-LVLM alignment. MMedPO curates multimodal preference data by introducing two types of dispreference: (1) plausible hallucinations injected through target Med-LVLMs or GPT-4o to produce medically inaccurate responses, and (2) lesion region neglect achieved through local lesion-noising, disrupting visual understanding of critical areas. We then calculate clinical relevance for each sample based on scores from multiple Med-LLMs and visual tools, enabling effective alignment. Our experiments demonstrate that MMedPO significantly enhances factual accuracy in Med-LVLMs, achieving substantial improvements over existing preference optimization methods by 14.2% and 51.7% on the Med-VQA and report generation tasks, respectively. Our code are available in https://github.com/aiming-lab/MMedPO}{https://github.com/aiming-lab/MMedPO.
Kangyu Zhu, Peng Xia 0005, Yun Li 0010, Hongtu Zhu, Sheng Wang 0014, Huaxiu Yao
ICML3
2025 HyperPianist: Pianist with Linear-Time Prover and Logarithmic Communication Cost
abstract
Recent years have seen great improvements in zero-knowledge proofs (ZKPs). Among them, zero-knowledge SNARKs are notable for their compact and efficiently-verifiable proofs, but suffer from high prover costs. Wu et al. (Usenix Security 2018) proposed to distribute the proving task across multiple machines, and achieved significant improvements in proving time. However, existing distributed ZKP systems still have quasi-linear prover cost, and may incur a communication cost that is linear in circuit size. In this paper, we introduce HyperPianist. Inspired by the state-of-the-art distributed ZKP system Pianist (Liu et al., S&P 2024) and the multivariate proof system HyperPlonk (Chen et al., EUROCRYPT 2023), we design a distributed multivariate polynomial interactive oracle proof (PIOP) system with a linear-time prover cost and logarithmic communication cost. Unlike Pianist, HyperPianist incurs no extra overhead in prover time or communication when applied to general (non-data-parallel) circuits. To instantiate the PIOP system, we adapt two additively-homomorphic multivariate polynomial commitment schemes, multivariate KZG (Papamanthou et al., TCC 2013) and Dory (Lee et al., TCC 2021), into the distributed setting, and get HyperPianistKand HyperPianistDrespectively. Both systems have linear prover complexity and logarithmic communication cost; furthermore, HyperPianistDrequires no trusted setup. We also propose HyperPianist+, incorporating an optimized lookup argument based on Lasso (Setty et al., EUROCRYPT 2024) with lower prover cost. Experiments demonstrate HyperPianistKand HyperPianistDachieve speedups of 63.1x and 40.2x over HyperPlonk with 32 distributed machines. Compared to Pianist, HyperPianistKcan be 2.9x and 4.6x as fast and HyperPianistDcan be 2.4x and 3.8x as fast, on vanilla gates and custom gates respectively. With layered circuits, HyperPianistKis up to 5.9x as fast on custom gates, and HyperPianistDachieves a 4.7x speedup.
Chongrong Li, Yun Li 0010, Cheng Hong 0001, Wenjie Qu 0001, Jiaheng Zhang
SP3
2025 ZHE: Efficient Zero-Knowledge Proofs for HE Evaluations
abstract
Homomorphic Encryption (HE) allows computations on encrypted data without decryption. It can be used where the users' information are to be processed by an untrustful server, and has been a popular choice in privacy-preserving applications. However, in order to obtain meaningful results, we have to assume an honest-but-curious server, i.e., it will faithfully follow what was asked to do. If the server is malicious, there is no guarantee that the computed result is correct. The notion of verifiable HE (vHE) is introduced to detect malicious server's behaviors, but current vHE schemes are either more than four orders of magnitude slower than the underlying HE operations (Atapoor et. al, CIC 2024) or fast but incompatible with server-side private inputs (Chatel et. al, CCS 2024). In this work, we propose a vHE framework ZHE: efficient Zero-Knowledge Proofs (ZKPs) that prove the correct execution of HE evaluations while protecting the server's private inputs. More precisely, we first design two new highly-efficient ZKPs for modulo operations and (Inverse) Number Theoretic Transforms (NTTs), two of the basic operations of HE evaluations. Then we build a customized ZKP for HE evaluations, which is scalable, enjoys a fast prover time and has a non-interactive online phase. Our ZKP is applicable to all Ring-LWE based HE schemes, such as BGV and CKKS. Finally, we implement our protocols for both BGV and CKKS and conduct extensive experiments on various HE workloads. Compared to the state-of-the-art works, both of our prover time and verifier time are improved; especially, our prover cost is only roughly 27–36× more expensive than the underlying HE operations, this is two to three orders of magnitude cheaper than state-of-the-arts.
Zhelei Zhou, Yun Li 0010, Zhaomin Yang, Bingsheng Zhang, Cheng Hong 0001, Tao Wei 0002
SP2
2025 A comprehensive comparison on clustering methods for multi-slice spatially resolved transcriptomics data analysis
abstract
Spatial transcriptomics (ST) data, by providing spatial information, enable simultaneous analysis of gene expression distributions and their spatial patterns within tissue. Clustering or spatial domain detection represents an essential methodology for ST data, facilitating the exploration of spatial organizations with shared gene expression or histological characteristics. Traditionally, clustering algorithms for ST have focused on individual tissue sections. However, the emergence of numerous contiguous tissue sections derived from the same or similar tissue specimens within or across individuals has led to the development of multi-slice clustering methods. In this study, we assess seven single-slice and four multi-slice clustering methods on two simulated datasets and four real datasets. Additionally, we investigate the effectiveness of preprocessing techniques, including spatial coordinate alignment (e.g. PASTE) and gene expression batch effect removal (e.g. Harmony), on clustering performance. Our study provides a comprehensive comparison of clustering methods for multi-slice ST data, serving as a practical guide for method selection in various scenarios.
Caiwei Xiong, Muqing Zhou, Wenrong Wu, Xihao Li, Huaxiu Yao, Jiawen Chen 0002, Yun Li 0010
Briefings Bioinform.9
2025 A domain adaptation method to Defend Chinese textual adversarial attacks via prompt-tuning
Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Eng. Appl. Artif. Intell.3
2025 Soft Prompt-tuning with Self-Resource Verbalizer for short text streams
Yi Zhu 0006, Ye Wang 0022, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001
Eng. Appl. Artif. Intell.3
2025 Robust and semantic-faithful post-hoc watermarking of text generated by black-box language models
Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Xiaocheng Hu, Xiaoye Ouyang
Frontiers Comput. Sci.4
2025 Domain adaptation for textual adversarial defense via prompt-tuning
Yi Zhu 0006, Chenqi Hua, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Neurocomputing4
2025 Multi-modal soft prompt-tuning for Chinese Clickbait Detection
Ye Wang 0022, Yi Zhu 0006, Yun Li 0010, Liting Wei, Yun-Hao Yuan 0001, Jipeng Qiang
Neurocomputing3
2025 Soft prompt-tuning for unsupervised domain adaptation via self-supervision
Yi Zhu 0006, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang
Neurocomputing3
2024 Incomplete Multi-Kernel k-Means Clustering With Fractional-Order Embedding
abstract
Multiple kernel clustering (MKC) has received increasing attention in the community of machine learning, which takes advantage of multiple pre-specified kernels to perform clustering tasks. Traditional MKC algorithms cannot effectively deal with the incomplete views where some samples are missing. Thus, incomplete MKC (IMKC) has been developed to solve this problem and obtained promising results. Nevertheless, the samples may be noisy or limited in real-world applications, which will result in the performance deterioration of existing IMKC algorithms. To address this issue, in this paper we propose a simple yet effective clustering method for incomplete data, termed fractional-order embedding incomplete multi-kernel k-means clustering (FE-MKKM-IK). Specifically, FE-MKKM-IK introduces the idea of fractional-order embedding to reconstruct the kernel matrix computed by the samples. On this basis, a new incomplete multiple kernel k-means clustering is developed. Performance evaluation is conducted on four widely used datasets, which shows that FE-MKKM-IK is effective to cluster the incomplete data.
Deheng Xu, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Yi Zhu 0006
IEEE Big Data2
2024 Sublinear Distributed Product Checks on Replicated Secret-Shared Data over Z2k Without Ring Extensions
abstract
Multiple works have designed or used maliciously secure honest majority MPC protocols over Z2k using replicated secret sharing (e.g. Koti et al. USENIX'21). A recent trend in the design of such MPC protocols is to first execute a semi-honest protocol, and then use a check that verifies the correctness of the computation requiring only sublinear amount of communication in terms of the circuit size. The so-called Galois ring extensions are needed in order to execute such checks over Z2k, but these rings incur incredibly high computation overheads, which completely undermine any potential benefits the ring Z2k had to begin with.
Yun Li 0010, Daniel Escudero 0001, Yufei Duan, Cheng Hong 0001, Chao Zhang 0008, Yifan Song 0001
CCS1
2024 RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models
abstract
The recent emergence of Medical Large Vision Language Models (Med-LVLMs) has enhanced medical diagnosis.However, current Med-LVLMs frequently encounter factual issues, often generating responses that do not align with established medical facts.Retrieval-Augmented Generation (RAG), which utilizes external knowledge, can improve the factual accuracy of these models but introduces two major challenges.First, limited retrieved contexts might not cover all necessary information, while excessive retrieval can introduce irrelevant and inaccurate references, interfering with the model's generation.Second, in cases where the model originally responds correctly, applying RAG can lead to an over-reliance on retrieved contexts, resulting in incorrect answers.To address these issues, we propose RULE, which consists of two components.First, we introduce a provably effective strategy for controlling factuality risk through the calibrated selection of the number of retrieved contexts.Second, based on samples where over-reliance on retrieved contexts led to errors, we curate a preference dataset to fine-tune the model, balancing its dependence on inherent knowledge and retrieved contexts for generation.We demonstrate the effectiveness of RULE on medical VQA and report generation tasks across three datasets, achieving an average improvement of 47.4% in factual accuracy.We publicly release our benchmark and code in https: //github.com/richard-peng-xia/RULE.
Peng Xia 0005, Kangyu Zhu, Haoran Li 0011, Hongtu Zhu, Yun Li 0010, Gang Li 0001, Linjun Zhang, Huaxiu Yao
EMNLP5
2024 Learning Spectral Canonical ℱ-Correlation Representation for Face Super-Resolution
abstract
Face super-resolution (FSR) is a powerful technique for restoring high-resolution face images from the captured low-resolution ones with the assistance of prior information. Existing FSR methods based on explicit or implicit covariance matrices are difficult to reveal complex nonlinear relationships between features, as conventional covariance computation is essentially a linear operation process. Besides, the limited number of training samples and noise disturbance lead to the deviation of sample covariance matrices. To solve these issues, we propose a novel FSR method via using spectral canonical ℱ-correlation representation. The proposed method first defines intra-resolution and inter-resolution covariation matrices by considering the nonlinear relationship between different features, and then uses the fractional order idea to rebuild covariation matrices. The qualitative and quantitative results have validated the superiority of the proposed method.
Yun-Hao Yuan 0001, Mingzhi Hao, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001
ICASSP3
2024 Face Super-Resolution Using Covariation-Guided Orthonormalized Partial Least Squares
Mingzhi Hao, Yun-Hao Yuan 0001, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010, Runmei Zhang
ICONIP (8)5
2024 STimage-1K4M: A histopathology image-gene expression dataset for spatial transcriptomics
abstract
Recent advances in multi-modal algorithms have driven and been driven by the increasing availability of large image-text datasets, leading to significant strides in various fields, including computational pathology. However, in most existing medical image-text datasets, the text typically provides high-level summaries that may not sufficiently describe sub-tile regions within a large pathology image. For example, an image might cover an extensive tissue area containing cancerous and healthy regions, but the accompanying text might only specify that this image is a cancer slide, lacking the nuanced details needed for in-depth analysis. In this study, we introduce STimage-1K4M, a novel dataset designed to bridge this gap by providing genomic features for sub-tile images. STimage-1K4M contains 1,149 images derived from spatial transcriptomics data, which captures gene expression information at the level of individual spatial spots within a pathology image. Specifically, each image in the dataset is broken down into smaller sub-image tiles, with each tile paired with $15,000-30,000$ dimensional gene expressions. With $4,293,195$ pairs of sub-tile images and gene expressions, STimage-1K4M offers unprecedented granularity, paving the way for a wide range of advanced research in multi-modal data analysis an innovative applications in computational pathology, and beyond.
Jiawen Chen 0002, Muqing Zhou, Wenrong Wu, Yun Li 0010, Didong Li
NeurIPS5
2024 CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
abstract
Artificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized healthcare. However, the trustworthiness of Med-LVLMs remains unverified, posing significant risks for future model deployment. In this paper, we introduce CARES and aim to comprehensively evaluate the Trustworthiness of Med-LVLMs across the medical domain. We assess the trustworthiness of Med-LVLMs across five dimensions, including trustfulness, fairness, safety, privacy, and robustness. CARES comprises about 41K question-answer pairs in both closed and open-ended formats, covering 16 medical image modalities and 27 anatomical regions. Our analysis reveals that the models consistently exhibit concerns regarding trustworthiness, often displaying factual inaccuracies and failing to maintain fairness across different demographic groups. Furthermore, they are vulnerable to attacks and demonstrate a lack of privacy awareness. We publicly release our benchmark and code in https://github.com/richard-peng-xia/CARES.
Peng Xia 0005, Juanxi Tian, Yangrui Gong, Ruibo Hou, Zhenbang Wu, Zhiyuan Fan, Yiyang Zhou, Kangyu Zhu, Zhaoyang Wang 0004, Xiao Wang 0044, Xuchao Zhang, Chetan Bansal, Marc Niethammer, Junzhou Huang, Hongtu Zhu, Yun Li 0010, Jimeng Sun 0001, ZongYuan Ge, Gang Li 0001, James Zou 0001, Huaxiu Yao
NeurIPS19
2024 FPLS-DC: functional partial least squares through distance covariance for imaging genetics
abstract
MOTIVATION: Imaging genetics integrates imaging and genetic techniques to examine how genetic variations influence the function and structure of organs like the brain or heart, providing insights into their impact on behavior and disease phenotypes. The use of organ-wide imaging endophenotypes has increasingly been used to identify potential genes associated with complex disorders. However, analyzing organ-wide imaging data alongside genetic data presents two significant challenges: high dimensionality and complex relationships. To address these challenges, we propose a novel, nonlinear inference framework designed to partially mitigate these issues. RESULTS: We propose a functional partial least squares through distance covariance (FPLS-DC) framework for efficient genome wide analyses of imaging phenotypes. It consists of two components. The first component utilizes the FPLS-derived base functions to reduce image dimensionality while screening genetic markers. The second component maximizes the distance correlation between genetic markers and projected imaging data, which is a linear combination of the FPLS-basis functions, using simulated annealing algorithm. In addition, we proposed an iterative FPLS-DC method based on FPLS-DC framework, which effectively overcomes the influence of inter-gene correlation on inference analysis. We efficiently approximate the null distribution of test statistics using a gamma approximation. Compared to existing methods, FPLS-DC offers computational and statistical efficiency for handling large-scale imaging genetics. In real-world applications, our method successfully detected genetic variants associated with the hippocampus, demonstrating its value as a statistical toolbox for imaging genetic studies. AVAILABILITY AND IMPLEMENTATION: The FPLS-DC method we propose opens up new research avenues and offers valuable insights for analyzing functional and high-dimensional data. In addition, it serves as a useful tool for scientific analysis in practical applications within the field of imaging genetics research. The R package FPLS-DC is available in Github: https://github.com/BIG-S2/FPLSDC.
Wenliang Pan, Yue Shan, Tengfei Li 0001, Yun Li 0010, Hongtu Zhu
Bioinform.6
2024 Short text classification with Soft Knowledgeable Prompt-tuning
Yi Zhu 0006, Ye Wang 0022, Jianyuan Mu, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001, Xindong Wu 0001
Expert Syst. Appl.4
2024 Representation learning: serial-autoencoder for personalized recommendation
Yi Zhu 0006, Yishuai Geng, Yun Li 0010, Jipeng Qiang, Xindong Wu 0001
Frontiers Comput. Sci.3
2023 ParaLS: Lexical Substitution via Pretrained Paraphraser
abstract
Lexical substitution (LS) aims at finding appropriate substitutes for a target word in a sentence.Recently, LS methods based on pretrained language models have made remarkable progress, generating potential substitutes for a target word through analysis of its contextual surroundings.However, these methods tend to overlook the preservation of the sentence's meaning when generating the substitutes.This study explores how to generate the substitute candidates from a paraphraser, as the generated paraphrases from a paraphraser contain variations in word choice and preserve the sentence's meaning.Since we cannot directly generate the substitutes via commonly used decoding strategies, we propose two simple decoding strategies that focus on the variations of the target word during decoding.Experimental results show that our methods outperform state-of-theart LS methods based on pre-trained language models on three benchmarks.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006
ACL (1)3
2023 Multilingual Lexical Simplification via Paraphrase Generation
abstract
Lexical simplification (LS) methods based on pretrained language models have made remarkable progress, generating potential substitutes for a complex word through analysis of its contextual surroundings. However, these methods require separate pretrained models for different languages and disregard the preservation of sentence meaning. In this paper, we propose a novel multilingual LS method via paraphrase generation, as paraphrases provide diversity in word selection while preserving the sentence’s meaning. We regard paraphrasing as a zero-shot translation task within multilingual neural machine translation that supports hundreds of languages. After feeding the input sentence into the encoder of paraphrase modeling, we generate the substitutes based on a novel decoding strategy that concentrates solely on the lexical variations of the complex word. Experimental results demonstrate that our approach surpasses BERT-based methods and zero-shot GPT3-based method significantly on English, Spanish, and Portuguese.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006, Kaixun Hua
ECAI3
2023 Chinese Lexical Substitution: Dataset and Method
abstract
Existing lexical substitution (LS) benchmarks were collected by asking human annotators to think of substitutes from memory, resulting in benchmarks with limited coverage and relatively small scales.To overcome this problem, we propose a novel annotation method to construct an LS dataset based on human and machine collaboration.Based on our annotation method, we construct the first Chinese LS dataset CHNLS which consists of 33,695 instances and 144,708 substitutes, covering three text genres (News, Novel, and Wikipedia).Specifically, we first combine four unsupervised LS methods as an ensemble method to generate the candidate substitutes, and then let human annotators judge these candidates or add new ones.This collaborative process combines the diversity of machine-generated substitutes with the expertise of human annotators.Experimental results that the ensemble method outperforms other LS methods.To our best knowledge, this is the first study for the Chinese LS task.
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xiaocheng Hu, Xiaoye Ouyang
EMNLP4
2023 Learning Supervised Covariation Projection Through General Covariance
abstract
Canonical correlation analysis (CCA) is a classical yet powerful tool for learning two-view feature representation in various fields. But, most CCA approaches are based on the conventional covariance measure, which makes them difficult to uncover the complicatedly nonlinear relationship between distinct features. In this paper, we address the preceding problem and propose two novel CCA approaches in a supervised manner by using a general covariance metric. The proposed approaches not only consider the label information of training data, but also the nonlinear relationship between different features rather than samples, which leads to greater flexibility in many practical applications. A series of experimental results on five benchmark datasets demonstrate the effectiveness of our proposed methods in terms of classification accuracy.
Xiangze Bao, Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006
ICASSP3
2023 Many Is Better Than One: Multiple Covariation Learning for Latent Multiview Representation
Yun-Hao Yuan 0001, Pengwei Qian, Jin Li 0028, Jipeng Qiang, Yi Zhu 0006, Yun Li 0010
ICONIP (9)6
2023 On the Identifiability and Interpretability of Gaussian Process Models
abstract
In this paper, we critically examine the prevalent practice of using additive mixtures of Mat\'ern kernels in single-output Gaussian process (GP) models and explore the properties of multiplicative mixtures of Mat\'ern kernels for multi-output GP models. For the single-output case, we derive a series of theoretical results showing that the smoothness of a mixture of Mat\'ern kernels is determined by the least smooth component and that a GP with such a kernel is effectively equivalent to the least smooth kernel component. Furthermore, we demonstrate that none of the mixing weights or parameters within individual kernel components are identifiable. We then turn our attention to multi-output GP models and analyze the identifiability of the covariance matrix $A$ in the multiplicative kernel $K(x,y) = AK_0(x,y)$, where $K_0$ is a standard single output kernel such as Mat\'ern. We show that $A$ is identifiable up to a multiplicative constant, suggesting that multiplicative mixtures are well suited for multi-output tasks. Our findings are supported by extensive simulations and real applications for both single- and multi-output settings. This work provides insight into kernel selection and interpretation for GP models, emphasizing the importance of choosing appropriate kernel structures for different tasks.
Jiawen Chen 0002, Wancen Mu, Yun Li 0010, Didong Li
NeurIPS3
2023 Efficient 3PC for Binary Circuits with Application to Maliciously-Secure DNN Inference
Yun Li 0010, Yufei Duan, Cheng Hong 0001, Chao Zhang 0008, Yifan Song 0001
USENIX Security Symposium1
2023 Natural language watermarking via paraphraser-based lexical substitution
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001
Artif. Intell.3
2023 Lexical simplification via single-word generation
Jipeng Qiang, Yang Li 0186, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006
Frontiers Comput. Sci.3
2023 Unsupervised statistical text simplification using pre-trained language modeling for initialization
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Yi Zhu 0006, Xindong Wu 0001
Frontiers Comput. Sci.3
2023 Safeguarding text generation API's intellectual property through meaning-preserving lexical watermarks
Yun Li 0010, Xiaoye Ouyang, Xiaocheng Hu, Jipeng Qiang
Frontiers Comput. Sci.2
2023 Representation learning via an integrated autoencoder for unsupervised domain adaptation
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Yun-Hao Yuan 0001, Yun Li 0010
Frontiers Comput. Sci.5
2023 A hybrid classification method via keywords screening and attention mechanisms in extreme short text
abstract
Short text classification has provoked a vast amount of attention and research in recent decades. However, most existing methods only focus on the short texts that contain dozens of words like Twitter and Microblog, while pay far less attention to the extreme short texts like news headline and search snippets. Meanwhile, contemporary short text classification methods that extend the features via external knowledge sources always introduce lots of useless concepts, which may be detrimental to classification performance. Moreover, unlike traditional short text classification methods, the classification results of extreme short texts are often determined by a few even one or two keywords. To address these problems, we propose a novel hybrid classification method via Keywords Screening and Attention Mechanisms in extreme short text, called KSAM. More specifically, firstly, the attention-based BiLSTM is introduced in our method to enhance the role of keywords. Secondly, we screen the keywords in the extreme short text for obtaining the true class label, and the concepts concerning the keywords are retrieved from external open knowledge sources like DBpedia. Thirdly, the attention mechanisms are introduced to acquire the weight of these retrieved concepts. Finally, conceptual information is utilized to assist the classification of the extreme short text. Extensive experiments have demonstrated the effectiveness of our method compared to other state-of-the-art methods.
Xinke Zhou, Yi Zhu 0006, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001, Xingdong Wu, Runmei Zhang
Intell. Data Anal.3
2023 Chinese Idiom Paraphrasing
abstract
Abstract Idioms are a kind of idiomatic expression in Chinese, most of which consist of four Chinese characters. Due to the properties of non-compositionality and metaphorical meaning, Chinese idioms are hard to be understood by children and non-native speakers. This study proposes a novel task, denoted as Chinese Idiom Paraphrasing (CIP). CIP aims to rephrase idiom-containing sentences to non-idiomatic ones under the premise of preserving the original sentence’s meaning. Since the sentences without idioms are more easily handled by Chinese NLP systems, CIP can be used to pre-process Chinese datasets, thereby facilitating and improving the performance of Chinese NLP tasks, e.g., machine translation systems, Chinese idiom cloze, and Chinese idiom embeddings. In this study, we can treat the CIP task as a special paraphrase generation task. To circumvent difficulties in acquiring annotations, we first establish a large-scale CIP dataset based on human and machine collaboration, which consists of 115,529 sentence pairs. In addition to three sequence-to-sequence methods as the baselines, we further propose a novel infill-based approach based on text infilling. The results show that the proposed method has better performance than the baselines based on the established CIP dataset.
Jipeng Qiang, Yang Li 0186, Chaowei Zhang 0001, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001
Trans. Assoc. Comput. Linguistics4
2022 Learning Canonical F-Correlation Projection for Compact Multiview Representation
abstract
Canonical correlation analysis (CCA) matters in multi-view representation learning. But, CCA and its most variants are essentially based on explicit or implicit covariance matrices. It means that they have no ability to model the nonlinear relationship among features due to intrinsic linearity of covariance. In this paper, we address the preceding problem and propose a novel canonical F-correlation framework by exploring and exploiting the nonlinear relationship between different features. The framework projects each feature rather than observation into a certain new space by an arbitrary nonlinear mapping, thus resulting in more flexibility in real applications. With this frame-work as a tool, we propose a correlative covariation projection (CCP) method by using an explicit nonlinear mapping. Moreover, we further propose a multiset version of CCP dubbed MCCP for learning compact representation of more than two views. The proposed MCCP is solved by an iterative method, and we prove the convergence of this iteration. A series of experimental results on six benchmark datasets demonstrate the effectiveness of our proposed CCP and MCCP methods.
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001, Jianping Gou
CVPR3
2022 RT-FEND: Spark-Based Real Time FakE News Detection
abstract
Fake news is a rampant societal and organizational problem with various social media outlets further aggravating its spread. There is a pressing demand to assist people to identify misinformation from massive amount of news data in a timely manner. Detecting Fake news in a timely manner is critical for mitigating its impact. In this research, we propose a novel approach for detecting fake news in real time, RT-FEND (Real Time- FakE News Detection), which relies on distributed computing paradigm. The proposed methodology utilizes event and topic extraction techniques along with a topic- merging mechanism to process real time news data and reduce the number of topics for managing the curse of dimensionality. We report the findings from several experiments to compare RT-FEND with other systems to benchmark in different system settings. RT-FEND approach is more performance-improved and time-efficient in detecting fake news when compared to other fake news detection baselines.
Chaowei Zhang 0001, Ashish Gupta 0004, Hui Sun 0002, Yun Li 0010, Xiao Qin 0001
NAS4
2022 Dynamic clustering for short text stream based on Dirichlet process
Wanyin Xu, Yun Li 0010, Jipeng Qiang
Appl. Intell.2
2022 Personalized recommendation with knowledge graph via dual-autoencoder
Yang Yang 0002, Yi Zhu 0006, Yun Li 0010
Appl. Intell.3
2022 A comprehensive comparison on cell-type composition inference for spatial transcriptomics data
abstract
Spatial transcriptomics (ST) technologies allow researchers to examine transcriptional profiles along with maintained positional information. Such spatially resolved transcriptional characterization of intact tissue samples provides an integrated view of gene expression in its natural spatial and functional context. However, high-throughput sequencing-based ST technologies cannot yet reach single cell resolution. Thus, similar to bulk RNA-seq data, gene expression data at ST spot-level reflect transcriptional profiles of multiple cells and entail the inference of cell-type composition within each ST spot for valid and powerful subsequent analyses. Realizing the critical importance of cell-type decomposition, multiple groups have developed ST deconvolution methods. The aim of this work is to review state-of-the-art methods for ST deconvolution, comparing their strengths and weaknesses. In particular, we construct ST spots from single-cell level ST data to assess the performance of 10 methods, with either ideal reference or non-ideal reference. Furthermore, we examine the performance of these methods on spot- and bead-level ST data by comparing estimated cell-type proportions to carefully matched single-cell ST data. In comparing the performance on various tissues and technological platforms, we concluded that RCTD and stereoscope achieve more robust and accurate inferences.
Jiawen Chen 0002, Weifang Liu, Tianyou Luo, Zhentao Yu, Minzhi Jiang, Gaorav P. Gupta, Paola Giusti, Hongtu Zhu, Yun Li 0010
Briefings Bioinform.11
2022 CrisprVi: a software for visualizing and analyzing CRISPR sequences of prokaryotes
abstract
BACKGROUND: Clustered regularly interspaced short palindromic repeats (CRISPR) and their spacers are important components of prokaryotic CRISPR-Cas systems. In order to analyze the CRISPR loci of multiple genomes more intuitively and comparatively, here we propose a visualization analysis tool named CrisprVi. RESULTS: CrisprVi is a Python package consisting of a graphic user interface (GUI) for visualization, a module for commands parsing and data transmission, local SQLite and BLAST databases for data storage and a functions layer for data processing. CrisprVi can not only visually present information of CRISPR direct repeats (DRs) and spacers, such as their orders on the genome, IDs, start and end coordinates, but also provide interactive operation for users to display, label and align the CRISPR sequences, which help researchers investigate the locations, orders and components of the CRISPR sequences in a global view. In comparison to other CRISPR visualization tools such as CRISPRviz and CRISPRStudio, CrisprVi not only improves the interactivity and effects of the visualization, but also provides basic statistics of the CRISPR sequences, and the consensus sequences of DRs/spacers across the input strains can be inspected from a clustering heatmap based on the BLAST results of the CRISPR sequences hitting against the genomes. CONCLUSIONS: CrisprVi is a convenient tool for visualizing and analyzing the CRISPR sequences and it would be helpful for users to inspect novel CRISPR-Cas systems of prokaryotes.
Jinbiao Wang, Fu Yan, Gongming Wang, Yun Li 0010, Jinlin Huang
BMC Bioinform.5
2022 Weakly-Supervised Semantic Segmentation Network With Iterative dCRF
abstract
This Autonomous driving methods driven by big data are becoming more and more perfect, but the cost of existing data labeling is too high, so how to reduce or even not label data has attracted more and more attention. Semantic segmentation networks supervised by image-level annotations are all trained using pseudo-labels. Most methods use image classification networks to generate class activation maps (CAMs) and start with CAMs to diffuse features to other parts of the target to obtain pseudo-labels. However, due to its weak supervision information, it is difficult for the existing methods to obtain better results. Therefore, we propose a weakly-supervised semantic segmentation network with iterative dCRF based on graph convolution. Specifically, we use ResNet to generate CAMs and node features and then use graph convolution for feature propagation and merge the low-level and high-level semantic information of the image. Then execute dCRF in an iterative manner, and finally obtain refined pseudo-labels. On the PASCAL VOC 2012 data set, our model achieves an mIoU of 63.5%, which is 0.3% higher than the graph convolutional network method.
Yujie Li 0001, Yun Li 0010
IEEE Trans. Intell. Transp. Syst.3
2022 Short Text Topic Modeling Techniques, Applications, and Performance: A Survey
abstract
Analyzing short texts infers discriminative and coherent latent topics that is a critical and fundamental task since many real-world applications require semantic understanding of short texts. Traditional long text topic modeling algorithms (e.g., PLSA and LDA) based on word co-occurrences cannot solve this problem very well since only very limited word co-occurrence information is available in short texts. Therefore, short text topic modeling has already attracted much attention from the machine learning research community in recent years, which aims at overcoming the problem of sparseness in short texts. In this survey, we conduct a comprehensive review of various short text topic modeling techniques proposed in the literature. We present three categories of methods based on Dirichlet multinomial mixture, global word co-occurrences, and self-aggregation, with example of representative approaches in each category and analysis of their performance on various tasks. We develop the first comprehensive open-source library, called STTM, for use in Java that integrates all surveyed algorithms within a unified interface, benchmark datasets, to facilitate the expansion of new methods in this research field. Finally, we evaluate these state-of-the-art methods on many real-world datasets and compare their performance against one another and versus long text topic modeling algorithm.
Jipeng Qiang, Zhenyu Qian 0006, Yun Li 0010, Yun-Hao Yuan 0001, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.3
2022 Multi-feature fusion point cloud completion network
Xiu Chen, Yujie Li 0001, Yun Li 0010
World Wide Web3
2021 ZKCPlus: Optimized Fair-exchange Protocol Supporting Practical and Flexible Data Exchange
abstract
Devising a fair-exchange protocol for digital goods has been an appealing line of research in the past decades. The Zero-Knowledge Contingent Payment (ZKCP) protocol first achieves fair exchange in a trustless manner with the aid of the Bitcoin network and zero-knowledge proofs. However, it incurs setup issues and substantial proving overhead, and has difficulties handling complicated validation of large-scale data. In this paper, we propose an improved solution ZKCPlus for practical and flexible fair exchange. ZKCPlus incorporates a new commit-and-prove non-interactive zero-knowledge (CP-NIZK) argument of knowledge under standard discrete logarithmic assumption, which is prover-efficient for data-parallel computations. With this argument we avoid the setup issues of ZKCP and reduce seller's proving overhead, more importantly enable the protocol to handle complicated data validations. We have implemented a prototype of ZKCPlus and built several applications atop it. We rework a ZKCP's classic application of trading sudoku solutions, and ZKCPlus achieves 21-67 times improvement in seller efficiency than ZKCP, with only milliseconds of setup time and 1 MB public parameters. In particular, our CP-NIZK argument shows an order of magnitude higher proving efficiency than the zkSNARK adopted by ZKCP. We also built a realistic application of trading trained CNN models. For a 3-layer CNN containing 8,620 parameters, it takes less than 1 second to prove and verify an inference computation, and also about 1 second to deliver the parameters, which is very promising for practical use.
Yun Li 0010, Cun Ye, Yuguang Hu, Ivring Morpheus, Chao Zhang 0008, Yupeng Zhang 0001, Haodi Wang
CCS1
2021 Fractional Multi-view Hashing with Semantic Correlation Maximization
Ruijie Gao, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Yi Zhu 0006
ICONIP (5)2
2021 Multi-view Fractional Deep Canonical Correlation Analysis for Subspace Clustering
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Yi Zhu 0006, Xiaobo Shen 0001
ICONIP (2)3
2021 Domain Adaptation with Stacked Convolutional Sparse Autoencoder
Yi Zhu 0006, Xinke Zhou, Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001
ICONIP (5)3
2021 RAProducer: efficiently diagnose and reproduce data race bugs for binaries via trace analysis
abstract
A growing number of bugs have been reported by vulnerability discovery solutions. Among them, some bugs are hard to diagnose or reproduce, including data race bugs caused by thread interleavings. Few solutions are able to well address this issue, due to the huge space of interleavings to explore. What’s worse, in security analysis scenarios, analysts usually have no access to the source code of target programs and have troubles in comprehending them.
Ming Yuan 0003, Yeseop Lee, Chao Zhang 0008, Yun Li 0010, Yan Cai 0001, Bodong Zhao
ISSTA4
2021 Composite nonlinear multiset canonical correlation analysis for multiview feature learning and recognition
abstract
Summary In this paper, we propose a composite nonlinear multiset canonical correlation projections (CNMCPs) framework where orthogonal constraints are imposed in each set. This makes CNMCP capable of learning uncorrelated low‐dimensional features with minimum redundancy in Hilbert space. With the CNMCP framework, we further present a particular algorithm called multikernel multiset canonical correlations or mKMCC, which introduces different weights into multiple nonlinear functions in all views. An alternating iterative optimization is designed for computational solution. Numerous experimental results on practical datasets have demonstrated the effectiveness and robustness of mKMCC, in contrast with existing kernel correlation learning approaches.
Yun-Hao Yuan 0001, Xiaobo Shen 0001, Yun Li 0010, Bin Li 0006, Jianping Gou, Jipeng Qiang, Xinfeng Zhang 0003, Quan-Sen Sun
Concurr. Comput. Pract. Exp.3
2021 Representation learning with collaborative autoencoder for personalized recommendation
Yi Zhu 0006, Xindong Wu 0001, Jipeng Qiang, Yun-Hao Yuan 0001, Yun Li 0010
Expert Syst. Appl.5
2021 OPLS-SR: A novel face super-resolution learning method using orthonormalized coherent features
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jipeng Qiang, Bin Li 0006, Wankou Yang, Furong Peng
Inf. Sci.3
2021 Learning Unsupervised and Supervised Representations via General Covariance
abstract
Component analysis (CA) is a powerful technique for learning discriminative representations in various computer vision tasks. Typical CA methods are essentially based on the covariance matrix of training data. But, the covariance matrix has obvious disadvantages such as failing to model complex relationship among features and singularity in small sample size cases. In this letter, we propose a general covariance measure to achieve better data representations. The proposed covariance is characterized by a nonlinear mapping determined by domain-specific applications, thus leading to more advantages, flexibility, and applicability in practice. With general covariance, we further present two novel CA methods for learning compact representations and discuss their differences from conventional methods. A series of experimental results on nine benchmark data sets demonstrate the effectiveness of the proposed methods in terms of accuracy.
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jianping Gou, Jipeng Qiang
IEEE Signal Process. Lett.3
2021 Chinese Lexical Simplification
abstract
Lexical simplification has attracted much attention in many languages, which is the process of replacing complex words in a given sentence with simpler alternatives of equivalent meaning. Although the richness of vocabulary in Chinese makes the text very difficult to read for children and non-native speakers, there is no research work for the Chinese lexical simplification (CLS) task. To circumvent difficulties in acquiring annotations, we manually create the first benchmark dataset for CLS, which can be used for evaluating the lexical simplification systems automatically. To acquire a more thorough comparison, we present five different types of methods as baselines to generate substitute candidates for the complex word that includes synonym-based approach, word embedding-based approach, BERT-based approach, sememe-based approach, and a hybrid approach. Finally, we design the experimental evaluation of these baselines and discuss their advantages and disadvantages. To our best knowledge, this is the first study for CLS task.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Xindong Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 LSBert: Lexical Simplification Based on BERT
abstract
Lexical simplification (LS) aims at replacing complex words with simpler alternatives. LS commonly consists of three main steps: complex word identification, substitute generation, and substitute ranking. Existing LS methods focus on the contextual information of the complex word in the last step (substitute ranking). However, they miss out the following two facts: (1) The word complexity of a polysemous word is very closely related to its context; (2) The step of substitute generation regardless of the context will inevitably produce a large number of spurious candidates. Therefore, we propose a novel LS system LSBert based on pretrained language model BERT to address the aforementioned issues, which is capable of making use of the wider context when both identifying the words in need of simplification and generating substitute candidates for the complex words. Specifically, LSBert consists of a network for complex word identification by fine-tuning BERT and a network for substitute generation based on BERT. Experimental results show that LSBert performs well in both complex word identification and substitute generation, achieving state-of-the-art results in three benchmarks. To facilitate reproducibility, the code of the LSBert system is available at https://github.com/qiang2100/BERT-LS.
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Yang Shi 0003, Xindong Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Lexical Simplification with Pretrained Encoders
abstract
Lexical simplification (LS) aims to replace complex words in a given sentence with their simpler alternatives of equivalent meaning. Recently unsupervised lexical simplification approaches only rely on the complex word itself regardless of the given sentence to generate candidate substitutions, which will inevitably produce a large number of spurious candidates. We present a simple LS approach that makes use of the Bidirectional Encoder Representations from Transformers (BERT) which can consider both the given sentence and the complex word during generating candidate substitutions for the complex word. Specifically, we mask the complex word of the original sentence for feeding into the BERT to predict the masked token. The predicted results will be used as candidate substitutions. Despite being entirely unsupervised, experimental results show that our approach obtains obvious improvement compared with these baselines leveraging linguistic databases and parallel corpus, outperforming the state-of-the-art by more than 12 Accuracy points on three well-known benchmarks.
Jipeng Qiang, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001
AAAI2
2020 Learning Fractional Orthogonal Latent Consistent Features for Face Hallucination and Recognition
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jipeng Qiang, Bin Li 0006
ICASSP3
2020 Regularized Multiset Neighborhood Correlation Analysis for Semi-paired Multiview Learning
Yun-Hao Yuan 0001, Zhaoqi Wu, Yun Li 0010, Jipeng Qiang, Jianping Gou, Yi Zhu 0006
ICONIP (2)3
2020 Maximum likelihood-based influence maximization in social networks
Wei Liu 0010, Yun Li 0010
Appl. Intell.2
2019 Learning Super-Resolution Coherent Facial Features Using Nonlinear Multiset PLS for Low-Resolution Face Recognition
abstract
Face hallucination (FH) is an effective technique for super-resolving low-resolution (LR) face images. In real-world applications, a face image usually has multiple distinct low resolutions. Most existing FH methods can not effectively deal with multiple LR views simultaneously. To solve this issue, we present a multi-set partial least squares (MPLS) approach and its kernel extension for jointly learning the nonlinear consistency of multi-resolution facial features. With nonlinear MPLS, we present a novel simultaneous super-resolution coherent facial feature method for the face images with multiple LRs, which has capacity of jointly learning the nonlinear relationships between multiple facial resolutions. Experimental results demonstrate the effectiveness and robustness of our proposed FH method.
Yun-Hao Yuan 0001, Jin Li 0028, Yun Li 0010, Jianping Gou, Jipeng Qiang, Quan-Sen Sun
ICIP3
2019 Learning Simultaneous Face Super-Resolution Using Multiset Partial Least Squares
abstract
Face super-resolution (FSR) is an effective way to solve low-resolution (LR) problems in face analysis. But, most FSR methods only consider that LR face images have a single resolution, which is usually not consistent with practical situations due to the existence of multiple resolutions. To date, simultaneously learning the mappings from multiple LRs to high resolution (HR) has not been given proper attention. To solve this issue, we first propose a multi-set partial least squares (MPLS) approach to jointly deal with multi-set random variables via a recursive optimization. With MPLS, we then present a novel FSR method called MPLS-FH to simultaneously learn multiple resolution-specific mappings for various LR views from the same source. Concretely, MPLS-FH first divides multi-resolution face images into many patches. Then, it jointly learns the latent coherent features of principal-component embeddings of multi-resolution patches. Last, it super-resolves the input LR face by cross-resolution neighborhood search. Experimental results demonstrate the effectiveness of the proposed method in terms of quantitative and qualitative evaluations.
Yun-Hao Yuan 0001, Jin Li 0028, Jianping Gou, Yun Li 0010, Jipeng Qiang, Bin Li 0006
ICME4
2019 D2PLS: A Novel Bilinear Method for Facial Feature Fusion
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Bin Li 0006, Jianping Gou
ICONIP (4)3
2019 Fuzzy Bilinear Latent Canonical Correlation Projection for Feature Learning
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Jianping Gou, Guangwei Gao, Bin Li 0006
ICONIP (1)3
2019 A practical algorithm for solving the sparseness problem of short text clustering
abstract
Dirichlet Multinomial Mixture (DMM) models have been successful in clustering short texts. However, the word co-occurrence information that can be captured by these models is limited to the short text corpus itself. If two words have strong relatedness but rarely co-occurring in short texts, these models can not fully capture the semantic relatedness between the two words. In this paper, we propose a novel model by incorporating word-word correlation into DMM, called WDMM. By constructing a sparse graph using word-word relationship, our model expands each short text using their neighboring words in each text that can help to solve the problem of sparseness in short texts. Therefore, the cluster label of each text is not only influenced by its words, but decided by their similar words in this corpus. Experimental results on real-world datasets demonstrated the substantial superiority of our WDMM model over the state-of-the-art methods.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Wei Liu 0010, Xindong Wu 0001
Intell. Data Anal.2
2018 Low Resolution Face Recognition and Reconstruction Via Deep Canonical Correlation Analysis
abstract
Low-resolution (LR) face identification is always a challenge in computer vision. In this paper, we propose a new LR face recognition and reconstruction method using deep canonical correlation analysis (DCCA). Unlike linear CCA-based methods, our proposed method can learn flexible nonlinear representations by passing LR and high-resolution (HR) image principal component features through multiple stacked layers of nonlinear transformation. As the nonlinear transformation in deep neural networks is implicit, we apply radial basis function based neural network to learn an explicit mapping between principal components and correlational features. In addition, we also design two residual compensation methods for identification and vision enhancement, respectively. The proposed approach is compared with existing LR face recognition and reconstruction algorithms. A number of experimental results on benchmark datasets have demonstrated the effectiveness and robustness of our method.
Zhao Zhang 0018, Yun-Hao Yuan 0001, Xiaobo Shen 0001, Yun Li 0010
ICASSP4
2018 A Complete Canonical Correlation Analysis for Multiview Learning
abstract
Canonical correlation analysis (CCA) is an effective feature learning method, which has wide applications in pattern recognition and computer vision. However, CCA considers the correlation only between the one-to-one aligned samples in two views, ignoring the correlation between all the samples sharing the same label. In this paper, we propose a deep complete canonical correlation analysis (Deep Complete-CCA), which learns the relationships between all pairwise correspondences of sample points in the same classes. Unlike CCA, our method can learn discriminant representations that maximize the correlation between the two views while segregating the different classes on the learned space. We test Deep Complete-CCA on handwriting recognition and speech based emotion recognition using two popular MNIST and RAVDESS datasets. Experimental results show that our proposed method can obtain better performances than several related algorithms.
Yan Liu 0038, Yun Li 0010, Yun-Hao Yuan 0001
ICIP2
2018 Text Simplification with Self-Attention-Based Pointer-Generator Networks
Yun Li 0010, Jipeng Qiang, Yun-Hao Yuan 0001
ICONIP (5)2
2018 Supervised Two-Dimensional CCA for Multiview Data Representation
Yun-Hao Yuan 0001, Yun Li 0010, Jipeng Qiang, Wenyan Bao
ICONIP (5)3
2018 Learning Parallel Canonical Correlations for Scale-Adaptive Low Resolution Face Recognition
abstract
Low resolution is one of the main obstacles in the application of face recognition. Although many methods have been proposed to improve the problem, they assume that low-resolution (LR) face images have a uniform scale. In real scenarios, this prerequisite is very harsh. In this paper, we propose a scale-adaptive LR face recognition approach based on two-dimensional multi-set canonical correlation analysis (2DM-CCA), where face image matrix does not need to be previously transformed into a vector. In the proposed method, training sets with different resolutions are treated as different views, and then projected in parallel into a latent coherent space where the consistency of multi-view face data is maximally enhanced. When a new LR face image with an arbitrary scale is input, we first transform it by using the left and right projection matrices of an appropriate training view, and then reconstruct its high resolution facial feature by neighborhood reconstruction. Experimental results show that our proposed method is more effective and efficient than several existing methods.
Yun-Hao Yuan 0001, Zhao Zhang 0018, Yun Li 0010, Jipeng Qiang, Bin Li 0006, Xiaobo Shen 0001
ICPR3
2018 Snapshot ensembles of non-negative matrix factorization for stability of topic modeling
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Wei Liu 0010
Appl. Intell.2
2018 Short text clustering based on Pitman-Yor process mixture model
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Xindong Wu 0001
Appl. Intell.2
2018 Active contour model-based segmentation algorithm for medical robots recognition
Yujie Li 0001, Yun Li 0010, Hyoungseop Kim, Seiichi Serikawa
Multim. Tools Appl.2
2017 Fractional discriminative multiview correlation projection for face feature fusion
abstract
Multiple view data with different feature representations have widely arisen in various practical applications. Due to the information diversity, fusing multiview features is very valuable for classification purpose. In this paper, we propose a new multifeature fusion method called fractional-order discriminative multiview correlation projection (FDMCP), which is based on fractional-order scatter matrices with class label information of the samples. FDMCP first defines supervised covariance matrices in each view. It then constructs fractional supervised scatter matrices. Experimental results on three benchmark face image datasets show that our proposed FDMCP approach outperforms generalized multiview linear discriminant analysis.
Yun-Hao Yuan 0001, Yun Li 0010, Bin Li 0006, Hongkun Ji, Xiaobo Shen 0001
FUSION3
2017 Supervised Deep Canonical Correlation Analysis for Multiview Feature Learning
Yan Liu 0038, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang, Min Ruan, Zhao Zhang 0018
ICONIP (6)2
2017 Face Hallucination and Recognition Using Kernel Canonical Correlation Analysis
Zhao Zhang 0018, Yun-Hao Yuan 0001, Yun Li 0010, Bin Li 0006, Jipeng Qiang
ICONIP (6)3
2017 Identifying the Number of Clusters in Short Text Using Bayesian Nonparametric Model
abstract
Before inferring the real number of clusters in short text clustering, Dirichlet Multinomial Mixture (DMM) model makes assumption that there are at most Kmax clusters. In some cases, it is difficult to choose a proper Kmax beforehand. In the paper, we propose a novel model based on Pitman-Yor Process to capture the power-law phenomenon of the cluster distribution. Specifically, each text chooses one of the active clusters or a new cluster with probabilities derived from the Pitman-Yor Process Mixture model (PYPM). Different from DMM model, our model does not require Kmax as input. Discriminative words and nondiscriminative words are identified automatically to help enhance text clustering. Parameters are estimated efficiently by collapsed Gibbs sampling. The experiments on real-world datasets validate the effectiveness of the proposed model in comparison with other state-of-theart models.
Jipeng Qiang, Yun Li 0010, Yun-Hao Yuan 0001, Tong Wang 0007
ICTAI2
2017 A new closed frequent itemset mining algorithm based on GPU and improved vertical structure
abstract
Summary Vertical data structure is very important for closed frequent itemset mining. All closed frequent itemsets can be found by simply using the operations of AND/OR. However, it consumes a large amount of storage space, especially in the case of large‐size dataset. This paper proposes an algorithm for mining closed frequent itemsets based on a new vertical data structure. The proposed data structure is helpful to save storage space by using a multi‐layer index. At the same time, numerous CPU and graphics processing unit can be employed in parallel to achieve high‐efficiency computing. Especially when dealing with large datasets, the proposed algorithm can obtain a high‐speed computing with the help of graphics processing unit. The improved vertical structure reduces the storage space of the data. The experimental results show that our proposed algorithm requires much less computation time than other related methods. Copyright © 2016 John Wiley & Sons, Ltd.
Yun Li 0010, Yun-Hao Yuan 0001, Ling Chen 0005
Concurr. Comput. Pract. Exp.1
2017 Wound intensity correction and segmentation with convolutional neural networks
abstract
Summary Wound area changes over multiple weeks are highly predictive of the wound healing process. A big data eHealth system would be very helpful in evaluating these changes. We usually analyze images of the wound bed for diagnosing injury. Unfortunately, accurate measurements of wound region changes from images are difficult. Many factors affect the quality of images, such as intensity inhomogeneity and color distortion. To this end, we propose a fast level set model‐based method for intensity inhomogeneity correction and a spectral properties‐based color correction method to overcome these obstacles. State‐of‐the‐art level set methods can segment objects well. However, such methods are time‐consuming and inefficient. In contrast to conventional approaches, the proposed model integrates a new signed energy force function that can detect contours at weak or blurred edges efficiently. It ensures the smoothness of the level set function and reduces the computational complexity of re‐initialization. To increase the speed of the algorithm further, we also include an additive operator‐splitting algorithm in our fast level set model. In addition, we consider using a camera, lighting, and spectral properties to recover the actual color. Numerical synthetic and real‐world images demonstrate the advantages of the proposed method over state‐of‐the‐art methods. Experimental results also show that the proposed model is at least twice as fast as methods used widely. Copyright © 2016 John Wiley & Sons, Ltd.
Huimin Lu 0001, Bin Li 0006, Junwu Zhu, Yujie Li 0001, Yun Li 0010, Xing Xu 0001, Li He 0001, Xin Li 0034, Jianru Li, Seiichi Serikawa
Concurr. Comput. Pract. Exp.5
2017 Link prediction in multi-relational networks based on relational similarity
Caiyan Dai, Ling Chen 0005, Bin Li 0006, Yun Li 0010
Inf. Sci.4
2017 Projection-based link prediction in a bipartite network
Man Gao, Ling Chen 0005, Bin Li 0006, Yun Li 0010, Wei Liu 0010, Yongcheng Xu
Inf. Sci.4
2017 Laplacian multiset canonical correlations for multiview feature extraction and image recognition
Yun-Hao Yuan 0001, Yun Li 0010, Xiaobo Shen 0001, Quan-Sen Sun, Jinlong Yang 0002
Multim. Tools Appl.2
2016 Underwater image descattering and quality assessment
abstract
Vision-based underwater navigation and object detection requires robust computer vision algorithms to operate in turbid water. Many conventional methods aimed at improving visibility in low turbid water. In this paper, we propose a novel contrast enhancement to enhance high turbid underwater images using descattering and color correction. The proposed enhancement method removes the scatter and preserves colors. In addition, as a rule to compare the performance of different image enhancement algorithms, a more comprehensive image quality assessment index Qu is proposed. The index combines the benefits of SSIM index and color distance index. Experimental results show that the proposed approach statistically outperforms state-of-the-art general purpose underwater image contrast enhancement algorithms. The experiment also demonstrated that the proposed method performs well for image classification.
Huimin Lu 0001, Yujie Li 0001, Xing Xu 0001, Li He 0001, Yun Li 0010, Donald G. Dansereau, Seiichi Serikawa
ICIP5
2016 A Duplication Task Scheduling Algorithm in Cloud Environments
Min Ruan, Yun Li 0010, Yinjuan Zhang
IDEAL2
2016 Semi-discriminative Multiview Canonical Correlation Analysis for Recognition
Yun-Hao Yuan 0001, Yun Li 0010, Hongkun Ji, Chong-Guang Ren, Xiaobo Shen 0001, Quan-Sen Sun
IDEAL2
2016 Fractional-Order Multiview Discriminant Analysis
Yun-Hao Yuan 0001, Yun Li 0010, Xiaobo Shen 0001, Chong-Guang Ren, Chao-Fei Li
IDEAL2
2016 Learning multi-kernel multi-view canonical correlations for image recognition
abstract
canonical correlations (M 2 CCs) framework for subspace learning. In the proposed framework, the input data of each original view are mapped into multiple higher dimensional feature spaces by multiple nonlinear mappings determined by different kernels. This makes M 2 CC can discover multiple kinds of useful information of each original view in the feature spaces. With the framework, we further provide a specific multi-view feature learning method based on direct summation kernel strategy and its regularized version. The experimental results in visual recognition tasks demonstrate the effectiveness and robustness of the proposed method.
Yun-Hao Yuan 0001, Yun Li 0010, Xiaobo Shen 0001, Guoqing Zhang 0002, Quan-Sen Sun
Comput. Vis. Media2
2016 Sampling-based algorithm for link prediction in temporal networks
Nahla Mohamed Ahmed Ibrahim, Ling Chen 0005, Bin Li 0006, Yun Li 0010, Wei Liu 0010
Inf. Sci.5
2015 MSR4SM: Using topic models to effectively mining software repositories for software maintenance tasks
Xiaobing Sun 0001, Bixin Li, Hareton K. N. Leung, Bin Li 0006, Yun Li 0010
Inf. Softw. Technol.5
2014 Automatic generation of package diagram to understand Java packages
abstract
Program comprehension is a prerequisite in most software maintenance and evolution tasks. Given an unfamiliar system, it is difficult for practitioners to determine which software artifacts are relevant to the current task. Generally, there are a variety of packages in a Java software system. These packages often have different intents and different relationships between each other. Different information of packages and the relationships between different stereotypes packages form a signature of the system. This paper proposes a novel approach to automatically generate the description of the packages and its diagram to show relationships between the packages. The generated description and diagram can allow developers to more easily understand the main intent and structure of the system.
Xiaobing Sun 0001, Yun Li 0010, Xiangyue Liu 0002
ICIS3
2014 Supporting program comprehension with program summarization
abstract
A large amount of software maintenance effort is spent on program comprehension. How to accurately and quickly get the functional features in a program becomes a hot issue in program comprehension. Some studies in this area are focused on extracting the topics by analyzing linguistic information in the source code based on the textual mining techniques. However, the extracted topics are usually composed of some standalone words and difficult to understand. In this paper, we attempt to solve this problem based on a novel program summarization technique. First, we propose to use latent semantic indexing and clustering to group source artifacts with similar vocabulary to analyze the composition of each package in the program. Then, some topics composed of a vector of independent words can be extracted based on latent semantic indexing. Finally, we employ Minipar, a nature language parser, to help generate the summaries. The summaries can effectively organize the words from the topics in the form of the predefined sentence based on some rules. With such form of summaries, developers can understand what the features the program has and their corresponding source artifacts.
Xiaobing Sun 0001, Xiangyue Liu 0002, Yun Li 0010
ICIS4