Xutan Peng

dblp:228/6859 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2024
0000-0001-5787-9982ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 L-APPLE: Language-agnostic Prototype Prefix Learning for Cross-lingual Event Detection
abstract
Cross-lingual event detection (CLED) is a challenging information extraction task in which a model is trained in one language and evaluated in another. Most recent methods attack CLED by aligning source and target language representations based on fine-tuning multilingual pre-trained language models. However, they need to modify all the model parameters and store a complete copy for each source-target language pair, which is resource-intensive and requires significant memory. In contrast, prefix-tuning is a more lightweight alternative, but it relies solely on the labeled source language data during training, limiting its performance. To address the above problems, we propose a novel framework for CLED with Language-agnostic Prototypical Prefix-Learning (L-APPLE), which can integrate language-agnostic event information with prefix-tuning. In detail, inspired by vanilla prompt methods, L-APPLE divides the prefix into two parts: one optimized as continuous word embeddings while the other generated with cross-lingual aligned event prototypes. Meanwhile, we employ language alignment with contrastive learning to acquire cross-lingual aligned event prototypes, and finally, parameters are optimized using both task and alignment loss. The evaluation of public CLED benchmarks demonstrates that L-APPLE achieves significant improvements in CLED with only less than 0.1% of the parameters optimized compared to previous fine-tuning methods.
Ziqin Zhu, Xutan Peng, Qian Li 0033, Cheng Ji 0001, Qingyun Sun, Jianxin Li 0002
CIKM2
2024 Selective Run-Length Encoding
abstract
Background. Run-Length Encoding (RLE) is one of the most fundamental tools in data compression. The basic idea is to represent each element only once, followed by its count within the run. For example,can be stored as(namely an encoded variable ) together with(namely a run-control variable ). Yet, the compression power of RLE drops significantly if there lacks consecutive elements in the sequence. In extreme cases, the output of the encoder may require more space than the input (aka size inflation ).
Xutan Peng, Dejia Peng, Jiafa Zhu
DCC1
2024 Triplet-aware graph neural networks for factorized multi-modal knowledge graph entity alignment
Qian Li 0033, Jianxin Li 0002, Jia Wu 0001, Xutan Peng, Cheng Ji 0001, Hao Peng 0001, Philip S. Yu
Neural Networks4
2023 On the Vulnerabilities of Text-to-SQL Models
abstract
Although it has been demonstrated that Natural Language Processing (NLP) algorithms are vulnerable to deliberate attacks, the question of whether such weaknesses can lead to software security threats is under-explored. To bridge this gap, we conducted vulnerability tests on Text-to-SQL systems that are commonly used to create natural language interfaces to databases. We showed that the Text-to-SQL modules within six commercial applications can be manipulated to produce malicious code, potentially leading to data breaches and Denial of Service attacks.1This is the first demonstration that NLP models can be exploited as attack vectors in the wild. In addition, experiments using four open-source language models verified that straightforward backdoor attacks on Text-to-SQL systems achieve a 100% success rate without affecting their performance. The aim of this work is to draw the community’s attention to potential software security issues associated with NLP algorithms and encourage exploration of methods to mitigate against them.
Xutan Peng, Jingfeng Yang 0001, Mark Stevenson 0001
ISSRE1
2022 Generating Disentangled Arguments with Prompts: A Simple Event Extraction Framework That Works
abstract
Event Extraction bridges the gap between text and event signals. Based on the assumption of trigger-argument dependency, existing approaches have achieved state-of-the-art performance with expert-designed templates or complicated decoding constraints. In this paper, for the first time we introduce the prompt-based learning strategy to the domain of Event Extraction, which empowers the automatic exploitation of label semantics on both input and output sides. To validate the effectiveness of the proposed generative method, we conduct extensive experiments with 11 diverse baselines. Empirical results show that, in terms of F1 score on Argument Extraction, our simple architecture is stronger than any other generative counterpart and even competitive with algorithms that require template engineering. Regarding the measure of recall, it sets new overall records for both Argument and Trigger Extractions. We hereby recommend this framework to the community, with the code publicly available at https://github.com/RingBDStack/GDAP.
Jinghui Si, Xutan Peng, Chen Li 0046, Jianxin Li 0002
ICASSP2
2022 Cross-knowledge-graph entity alignment via relation prediction
Hongren Huang, Chen Li 0046, Xutan Peng, Lifang He 0001, Hao Peng 0001, Jianxin Li 0002
Knowl. Based Syst.3
2021 Graph-based Semi-Supervised Learning by Strengthening Local Label Consistency
abstract
Graph-based algorithms have drawn much attention thanks to their impressive success in semi-supervised setups. For better model performance, previous studies have learned to transform the topology of the input graph. However, these works only focus on optimizing the original nodes and edges, leaving the direction of augmenting existing data insufficiently explored. In this paper, we propose a novel heuristic pre-processing technique, namelyLocal Label Consistency Strengthening (ŁLCS), which automatically expands new nodes and edges to refine the label consistency within a dense subgraph. Our framework can effectively benefit downstream models by substantially enlarging the original training set with high-quality generated labeled data and refining the original graph topology. To justify the generality and practicality of ŁLCS, we couple it with the popular graph convolution network and graph attention network to perform extensive evaluations on three standard datasets. In all setups tested, our method boosts the average accuracy by a large margin of 4.7% and consistently outperforms the state-of-the-art.
Chen Li 0046, Xutan Peng, Hao Peng 0001, Jia Wu 0001, Philip S. Yu, Jianxin Li 0002, Lichao Sun 0001
CIKM2
2021 Summarising Historical Text in Modern Languages
abstract
We introduce the task of historical text summarisation, where documents in historical forms of a language are summarised in the corresponding modern language.This is a fundamentally important routine to historians and digital humanities researchers but has never been automated.We compile a high-quality gold-standard text summarisation dataset, which consists of historical German and Chinese news from hundreds of years ago summarised in modern German or Chinese.Based on cross-lingual transfer learning techniques, we propose a summarisation model that can be trained even with no cross-lingual (historical to modern) parallel data, and further benchmark it against state-of-the-art algorithms.We report automatic and human evaluations that distinguish the historic to modern language summarisation task from standard cross-lingual summarisation (i.e., modern to modern language), highlight the distinctness and value of our dataset, and demonstrate that our transfer learning approach outperforms standard cross-lingual benchmarks on this task.DE №34 Story Jhre Königl.Majest.befinden sich noch vnweit Thorn / ... / dahero zur Erledigung Hoffnung gemacht werden will.(Their Royal Majesties are still not far from Torn, ... , therefore completion of the hope is desired.)Summary Der Krieg zwischen Polen und Schweden dauert an.Von einem Friedensvertrag ist noch nicht der Rede.(The war between Poland and Sweden continues.There is still no talk on the peace treaty.)ZH №7 Story 有脚夫小民,三四千名集众围绕马监丞衙门,...,冒火突入,捧出敕印。 (Three to four thousand porters gathered around Majiancheng Yamen (a government office), ..., rushed into fire and salvaged the authority's seal.)Summary 小本生意免税条约未能落实,小商贩被严重剥削,以致百姓聚众闹事并火烧衙门,造成多人伤亡。王炀 抢救出公章。 (The tax-exemption act for small businesses was not well implemented and small traders were terribly exploited, leading to riot and arson attack on Yamen with many casualties.Yang Wang salvaged the authority's seal.)
Xutan Peng, Chenghua Lin 0002, Advaith Siddharthan
EACL1
2021 TextGTL: Graph-based Transductive Learning for Semi-supervised Text Classification via Structure-Sensitive Interpolation
abstract
Compared with traditional sequential learning models, graph-based neural networks exhibit excellent properties when encoding text, such as the capacity of capturing global and local information simultaneously. Especially in the semi-supervised scenario, propagating information along the edge can effectively alleviate the sparsity of labeled data. In this paper, beyond the existing architecture of heterogeneous word-document graphs, for the first time, we investigate how to construct lightweight non-heterogeneous graphs based on different linguistic information to better serve free text representation learning. Then, a novel semi-supervised framework for text classification that refines graph topology under theoretical guidance and shares information across different text graphs, namely Text-oriented Graph-based Transductive Learning (TextGTL), is proposed. TextGTL also performs attribute space interpolation based on dense substructure in graphs to predict low-entropy labels with high-quality feature nodes for data augmentation. To verify the effectiveness of TextGTL, we conduct extensive experiments on various benchmark datasets, observing significant performance gains over conventional heterogeneous graphs. In addition, we also design ablation studies to dive deep into the validity of components in TextTGL.
Chen Li 0046, Xutan Peng, Hao Peng 0001, Jianxin Li 0002
IJCAI2
2021 Highly Efficient Knowledge Graph Embedding Learning with Orthogonal Procrustes Analysis
abstract
Knowledge Graph Embeddings (KGEs) have been intensively explored in recent years due to their promise for a wide range of applications. However, existing studies focus on improving the final model performance without acknowledging the computational cost of the proposed approaches, in terms of execution time and environmental impact. This paper proposes a simple yet effective KGE framework which can reduce the training time and carbon footprint by orders of magnitudes compared with state-of-the-art approaches, while producing competitive performance. We highlight three technical innovations: full batch learning via relational matrices, closed-form Orthogonal Procrustes Analysis for KGEs, and non-negative-sampling training. In addition, as the first KGE method whose entity embeddings also store full relation information, our trained models encode rich semantics and are highly interpretable. Comprehensive experiments and ablation studies involving 13 strong baselines and two standard datasets verify the effectiveness and efficiency of our algorithm.
Xutan Peng, Guanyi Chen, Chenghua Lin 0002, Mark Stevenson 0001
NAACL-HLT1
2021 Cross-Lingual Word Embedding Refinement by $\ell_1$ Norm Optimisation
abstract
This is a repository copy of Cross-lingual word embedding refinement by ℓ1 norm optimisation.
Xutan Peng, Chenghua Lin 0002, Mark Stevenson 0001
NAACL-HLT1
2021 Learning graph attention-aware knowledge graph embedding
Chen Li 0046, Xutan Peng, Yuhang Niu, Shanghang Zhang, Hao Peng 0001, Chuan Zhou 0001, Jianxin Li 0002
Neurocomputing2
2020 Modeling relation paths for knowledge base completion via joint adversarial training
Chen Li 0046, Xutan Peng, Shanghang Zhang, Hao Peng 0001, Philip S. Yu, Linfeng Du
Knowl. Based Syst.2