EDBT 2026 Demo / reviewers in the wild / expert
Pengyu Xu
dblp:301/7227
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Information extraction and text analysis · 61% Representation and self-supervised learning · 18% Deep learning architectures and training · 10% | |
| Software engineering, system software, and programming languages
1 paper |
Debugging and program repair · 87% Program synthesis and code generation · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › text classification
multi-label text classification |
1.3 | 2 | 2023 | Triple Alliance Prototype Orthotist Network for Long-Tailed Multi-Label Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2023 Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text Classification · AAAI 2023 |
Natural language and speech › Information extraction and text analysis
text classification |
1.3 | 2 | 2023 | Triple Alliance Prototype Orthotist Network for Long-Tailed Multi-Label Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2023 Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text Classification · AAAI 2023 |
Computational science and engineering › scientific machine learning
neural operator |
1.0 | 1 | 2026 | Learning Neural Operators from Partial Observations via Latent Autoregressive Modeling · AAAI 2026 |
Computational science and engineering › scientific machine learning
physics-informed machine learning |
1.0 | 1 | 2026 | Learning Neural Operators from Partial Observations via Latent Autoregressive Modeling · AAAI 2026 |
Debugging and program repair
automated program repair |
0.9 | 1 | 2025 | InstructRepair: Instruct Large Language Models With Rich Bug Information for Automated Program Repair · IEEE Trans. Inf. Forensics Secur. 2025 |
Debugging and program repair › automated program repair
LLM-based program repair |
0.9 | 1 | 2025 | InstructRepair: Instruct Large Language Models With Rich Bug Information for Automated Program Repair · IEEE Trans. Inf. Forensics Secur. 2025 |
Machine learning › Deep learning architectures and training
data augmentation |
0.7 | 1 | 2023 | Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text Classification · AAAI 2023 |
Machine learning › Representation and self-supervised learning › feature transformation
feature augmentation |
0.7 | 1 | 2023 | Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text Classification · AAAI 2023 |
Natural language and speech › Information extraction and text analysis
keyphrase extraction |
0.7 | 1 | 2023 | Mitigating Over-Generation for Unsupervised Keyphrase Extraction with Heterogeneous Centrality Detection · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis › keyphrase extraction
unsupervised keyphrase extraction |
0.7 | 1 | 2023 | Mitigating Over-Generation for Unsupervised Keyphrase Extraction with Heterogeneous Centrality Detection · EMNLP 2023 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.5 | 1 | 2021 | Interpretable Deep Generative Recommendation Models · J. Mach. Learn. Res. 2021 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | Interpretable Deep Generative Recommendation Models · J. Mach. Learn. Res. 2021 |
Program synthesis and code generation
code generation with language models |
0.3 | 1 | 2025 | InstructRepair: Instruct Large Language Models With Rich Bug Information for Automated Program Repair · IEEE Trans. Inf. Forensics Secur. 2025 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.2 | 1 | 2023 | Triple Alliance Prototype Orthotist Network for Long-Tailed Multi-Label Text Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Methods — techniques the papers use, named apart from their topics
heterogeneous centrality detection · 1.3adaptive boundary-aware regularization · 1.3variational inference · 1.0physics-aware latent propagator · 1.0mask-to-predict training · 1.0autoregressive generation · 1.0large language model instruction tuning · 0.9second-order statistics transfer · 0.7prototype network · 0.7meta-learning · 0.7loss reweighting · 0.7attention mechanism · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Neural Operators from Partial Observations via Latent Autoregressive ModelingabstractReal-world scientific applications frequently encounter incomplete observational data due to sensor limitations, geographic constraints, or measurement costs. Although neural operators significantly advanced PDE solving in terms of computational efficiency and accuracy, their underlying assumption of fully-observed spatial inputs severely restricts applicability in real-world application. We introduce the first systematic framework for learning neural operators from partial observation. We identify and formalize two fundamental obstacles: (i) the supervision gap in unobserved regions that prevents effective learning of physical correlations, and (ii) the dynamic spatial mismatch between incomplete inputs and complete solution fields. Specifically, our proposed LANO (Latent Autoregressive Neural Operator) introduces two novel components designed explicitly to address the core difficulties of partial observations: (i) a mask-to-predict training strategy that creates artificial supervision by strategically masking observed regions, and (ii) a Physics-Aware Latent Propagator that reconstructs solutions through boundary-first autoregressive generation in latent space. Additionally, we develop POBench-PDE, a dedicated and comprehensive benchmark designed specifically for evaluating neural operators under partial observation conditions across three PDE-governed tasks. LANO achieves state-of-the-art performance with relative error reductions ranging from eighteen to sixty-nine percent across all benchmarks under patch-wise missingness with missing rates below fifty percent, including real-world climate prediction. Our approach effectively addresses practical scenarios with missing rates of up to seventy-five percent, to some extent bridging the existing gap between idealized research settings and the complexities of real-world scientific computing. Jingren Hou, Pengyu Xu, Chang Gao 0007, Huafeng Liu 0001, Liping Jing |
AAAI | 3 |
| 2025 | SSRepVM-UNet: a lightweight hybrid model for medical image segmentation based on channel parallelism
Yijing Guo, Fuhang Li, Kunhua Li, Pengyu Xu |
Appl. Intell. | 5 |
| 2025 | Pedestrian detection based on vision-language semantics with global adaptive adjustment
Yijing Guo, Fuhang Li, Pengyu Xu, Kunhua Li |
Pattern Recognit. Lett. | 4 |
| 2025 | InstructRepair: Instruct Large Language Models With Rich Bug Information for Automated Program Repair
Anmin Fu, Pengyu Xu, Jichunyang Li, Boyu Kuang, Yansong Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Taming Prompt-Based Data Augmentation for Long-Tailed Extreme Multi-Label Text ClassificationabstractIn extreme multi-label text classification (XMC), labels usually follow a long-tailed distribution, where most labels only contain a small number of documents and limit the performance of XMC. Data augmentation (DA) is a simple but effective strategy to solve such low-resource problems. In this paper, we propose a prompt-based DA method called XDA, which is specifically designed for XMC. First, we employ a soft prompt during the fine-tuning process of the T5 model for label-conditional DA, thereby enabling T5 to augment samples while preserving label-compatibility. Subsequently, XDA performs sample filtering on the augmented samples through the diversity of text and the consistency of labels, which enhances the quality of the DA. In contrast to traditional sample-level DA, we propose a pair-level DA method by masking the augmented sample-label pairs of head-labels during training, effectively mitigating the long-tailed problem. Comprehensive experiments on benchmark datasets have shown that the proposed XDA outperforms the state-of-the-art counterparts. Pengyu Xu, Sijin Lu, Liping Jing, Jian Yu 0001 |
ICASSP | 1 |
| 2024 | Global Optimizing Prestack Seismic Inversion Approach Using an Accurate Hessian Matrix Based on Exact Zoeppritz EquationsabstractTo increase the accuracy and vertical resolution of seismic inversion for exploratory purposes, a new method was developed for P-wave velocity, S-wave velocity and density inversion using prestack seismic data based on the mayfly optimization algorithm (MA), exact Zoeppritz equations and Bayesian framework. A new form of an accurate Hessian matrix was successfully derived. We innovatively used the MA nonlinear AVA inversion based on the accurate Hessian matrix (MANAI-Hessian) method for prestack seismic inversion and highlighted two main challenges for the first time. The popular and recent whale optimization algorithm (WOA) and a conventional Levenberg–Marquardt (LM) method were introduced to demonstrate the existence of these two challenges. Comprehensive partial derivative tests were well designed to verify the existence of the second-order partial derivatives of the P-wave reflection coefficients. A 3D special wedge model was introduced to test the accuracy and vertical resolution of the new method. Next, we applied the proposed method to the field data of deep carbonate rock from a study area in China. Compared with the conventional LM method and the accurate Jacobian matrix-based nonlinear AVA inversion method, which provides foundational approaches to address the two main challenges, the proposed approach shows superior performance in terms of accuracy and vertical resolution. Pengyu Xu, Huailai Zhou, Xingye Liu, Yuyong Yang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text ClassificationabstractMulti-label text classification (MLTC) involves tagging a document with its most relevant subset of labels from a label set. In real applications, labels usually follow a long-tailed distribution, where most labels (called as tail-label) only contain a small number of documents and limit the performance of MLTC. To facilitate this low-resource problem, researchers introduced a simple but effective strategy, data augmentation (DA). However, most existing DA approaches struggle in multi-label settings. The main reason is that the augmented documents for one label may inevitably influence the other co-occurring labels and further exaggerate the long-tailed problem. To mitigate this issue, we propose a new pair-level augmentation framework for MLTC, called Label-Specific Feature Augmentation (LSFA), which merely augments positive feature-label pairs for the tail-labels. LSFA contains two main parts. The first is for label-specific document representation learning in the high-level latent space, the second is for augmenting tail-label features in latent space by transferring the documents second-order statistics (intra-class semantic variations) from head labels to tail labels. At last, we design a new loss function for adjusting classifiers based on augmented datasets. The whole learning procedure can be effectively trained. Comprehensive experiments on benchmark datasets have shown that the proposed LSFA outperforms the state-of-the-art counterparts. Pengyu Xu, Sijin Lu, Liping Jing, Jian Yu 0001 |
AAAI | 1 |
| 2023 | Mitigating Over-Generation for Unsupervised Keyphrase Extraction with Heterogeneous Centrality DetectionabstractOver-generation errors occur when a keyphrase extraction model correctly determines a candidate keyphrase as a keyphrase because it contains a word that frequently appears in the document but at the same time erroneously outputs other candidates as keyphrases because they contain the same word.To mitigate this issue, we propose a new heterogeneous centrality detection approach (CentralityRank), which extracts keyphrases by simultaneously identifying both implicit and explicit centrality within a heterogeneous graph as the importance score of each candidate.More specifically, Centrali-tyRank detects centrality by taking full advantage of the content within the input document to construct graphs that encompass semantic nodes of varying granularity levels, not limited to just phrases.These additional nodes act as intermediaries between candidate keyphrases, enhancing inter-phrase relevance.Furthermore, we introduce a novel adaptive boundary-aware regularization that can leverage the position information of candidate keyphrases, thus influencing the importance of candidate keyphrases.Extensive experimental results demonstrate the superiority of CentralityRank over recent stateof-the-art unsupervised keyphrase extraction baselines on three benchmark datasets. Pengyu Xu, Huafeng Liu 0001, Liping Jing |
EMNLP | 2 |
| 2023 | Bootstrap Decomposition Enhanced Orthogonal-Basis Projection Model for Long-term Time Series ForecastingabstractLong-term time series forecasting facilitates early decision-making in a variety of fields, including weather forecasting and disease control. For the majority of data, the only way to predict the future is by discovering the hidden regularities of historical data. However, these regularities are difficult to learn. In the real world, time series patterns are complex, with a mixture of various trends and seasonality, and these sub-patterns are distinct. Thus, it is challenging for a single model to learn the regularities underlying these distinct sub-patterns. Second, these regularities are frequently obscured by noise, which makes it easy to overfit during the learning process. To address these problems, we propose a Bootstrap Decomposition Enhanced Orthogonal-Basis Projection model, or BDEOP: it first decomposes the data into several distinct sub-patterns with a bootstrap decomposition module; subsequently, it uses different modules to capture the hidden regularities and make the predictions based on these sub-patterns; simultaneously, the effect of noise is eliminated by randomly dropping the frequency components in the Fourier space. In terms of accuracy, our experimental investigation of four benchmark datasets from multiple domains demonstrates that the proposed model outperforms state-of-the-art models. Code is available at https://github.com/Nicholas0917/BDEOP-main. Pengyu Xu, Liping Jing |
IJCNN | 2 |
| 2023 | Textual tag recommendation with multi-tag topical attention
Pengyu Xu, Mingxuan Xia, Huafeng Liu 0001, Liping Jing, Jian Yu 0001 |
Neurocomputing | 1 |
| 2023 | Triple Alliance Prototype Orthotist Network for Long-Tailed Multi-Label Text ClassificationabstractText classification, is one of the key tasks for representing the semantic information of documents, multi-label text classification (MLTC) is an important branch of it. MLTC aims to tag the most relevant labels for the given document. Compared to the standard multi-class case where each document has only one label, it is considerably more difficulty to annotate new coming documents for multi-label text classification. Furthermore, it also suffers from the challenge of highly skewed long-tailed label distribution. i.e., a few labels are associated with a large number of documents (a.k.a. head labels), while a large fraction of labels are associated with a small number of documents (a.k.a. tail labels). Due to the relative infrequency of tail labels, this leads to an imbalance that biases towards predicting more head labels. As challenging as this task is, it is an essential task to tackle since it represents many real-world cases, such as text retrieval of news. To address the challenge, we propose a Triple Alliance Prototype Orthotist Network (TAPON) to build a generic meta-mapping from few-shot prototypes to many-shot classifier parameters, which aims to promote the generalizability of tail classifiers. To be specific, TAPON is a two-stage method. At the first stage focusing on head labels, TAPON obtains the meta-knowledge between many-shot classifier parameters and few-shot prototype of head labels. Head label classifiers are trained by many-shot documents. Meanwhile, the triple alliance prototype is obtained by adopting an Attentive Prototype with the aid of few-shot documents, label semantic information and label correlation. Additionally, a Prototype Orthotist module is especially designed to capture the meta-knowledge between the many-shot classifier and few-shot prototype. At the second stage of transferring, TAPON aims to transfer the generic meta-mapping from head labels to tail labels. It first uses Attentive Prototype to obtain triple alliance prototype for tail labels, and then uses the meta-knowledge obtained from the first stage to get many-shot classifiers for tail labels. By conducting extensive experiments on four benchmark datasets, we show that the proposed TAPON significantly outperforms other state-of-the-art methods for long-tailed multi-label text classification. Pengyu Xu, Huafeng Liu 0001, Liping Jing, Xiangliang Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Semantic guide for semi-supervised few-shot multi-label node classification
Pengyu Xu, Liping Jing, Uchenna Akujuobi, Xiangliang Zhang 0001 |
Inf. Sci. | 2 |
| 2022 | Bayesian Additive Matrix Approximation for Social RecommendationabstractSocial relations between users have been proven to be a good type of auxiliary information to improve the recommendation performance. However, it is a challenging issue to sufficiently exploit the social relations and correctly determine the user preference from both social and rating information. In this article, we propose a unified Bayesian Additive Matrix Approximation model (BAMA), which takes advantage of rating preference and social network to provide high-quality recommendation. The basic idea of BAMA is to extract social influence from social networks, integrate them to Bayesian additive co-clustering for effectively determining the user clusters and item clusters, and provide an accurate rating prediction. In addition, an efficient algorithm with collapsed Gibbs Sampling is designed to inference the proposed model. A series of experiments were conducted on six real-world social datasets. The results demonstrate the superiority of the proposed BAMA by comparing with the state-of-the-art methods from three views, all users, cold-start users, and users with few social relations. With the aid of social information, furthermore, BAMA has ability to provide the explainable recommendation. Huafeng Liu 0001, Liping Jing, Jingxuan Wen, Pengyu Xu, Jian Yu 0001, Michael Kwok-Po Ng |
ACM Trans. Knowl. Discov. Data | 4 |
| 2021 | Interpretable Deep Generative Recommendation ModelsabstractUser preference modeling in recommendation system aims to improve customer experience through discovering users’ intrinsic preference based on prior user behavior data. This is a challenging issue because user preferences usually have complicated structure, such as inter-user preference similarity and intra-user preference diversity. Among them, inter-user similarity indicates different users may share similar preference, while intra-user diversity indicates one user may have several preferences. In literatures, deep generative models have been successfully applied in recommendation systems due to its flexibility on statistical distributions and strong ability for non-linear representation learning. However, they suffer from the simple generative process when handling complex user preferences. Meanwhile, the latent representations learned by deep generative models are usually entangled, and may range from observed-level ones that dominate the complex correlations between users, to latent-level ones that characterize a user’s preference, which makes the deep model hard to explain and unfriendly for recommendation. Thus, in this paper, we propose an Interpretable Deep Generative Recommendation Model (InDGRM) to characterize inter-user preference similarity and intra-user preference diversity, which will simultaneously disentangle the learned representation from observed-level and latent-level. In InDGRM, the observed-level disentanglement on users is achieved by modeling the user-cluster structure (i.e., inter-user preference similarity) in a rich multimodal space, so that users with similar preferences are assigned into the same cluster. The observed-level disentanglement on items is achieved by modeling the intra-user preference diversity in a prototype learning strategy, where different user intentions are captured by item groups (one group refers to one intention). To promote disentangled latent representations, InDGRM adopts structure and sparsity-inducing penalty and integrates them into the generative procedure, which has ability to enforce each latent factor focus on a limited subset of items (e.g., one item group) and benefit latent-level disentanglement. Meanwhile, it can be efficiently inferred by minimizing its penalized upper bound with the aid of local variational optimization technique. Theoretically, we analyze the generalization error bound of InDGRM to guarantee its performance. A series of experimental results on four widely-used benchmark datasets demonstrates the superiority of InDGRM on recommendation performance and interpretability. Huafeng Liu 0001, Liping Jing, Jingxuan Wen, Pengyu Xu, Jiaqi Wang 0006, Jian Yu 0001, Michael Kwok-Po Ng |
J. Mach. Learn. Res. | 4 |