VLDB 2026 Research / reviewers in the wild / expert
Pengyuan Li 0001
dblp:131/2826-1
· DBLP profile ↗
19ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-8205-7611ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Context-Aware Visual Multi-turn Conversation Generation from Wikipedia and Wikidata
Basel Shbita, Pengyuan Li 0001, Anna Lisa Gentile |
ESWC (2) | 2 |
| 2024 | Towards Collecting Royalties for Copyrighted Data for Generative ModelsabstractAddressing issues of copyrighted data in the context of generative models has become an important issue for content creators, publishers, organizations training generative models, and those who deploy generative models for particular applications. Copyright holders want to ensure that they are fairly compensated for their work and users of training data and models do not want to expose themselves to litigation. However, traditional models of bulk-licensing data fit only poorly the context of model training. In this paper, we want to discuss why a traditional data license is not always a good fit, how data is used in the life-cycle of generative models and which impact data has on model output. This can be used as a foundation for a pay-per-(model)use compensation based how data contributes to a model’s output. Having a way to compensate copyright holders in this way reduces risk for model trainers, avoids large investments upfront, and encourage a lively data ecosystem in which the creation and distribution of original work is encouraged and fairly compensated. Heiko Ludwig, Yi Zhou 0015, Syed Zawad, Yuya Jeremy Ong, Pengyuan Li 0001, Eric Butler, Eelaaf Zahid |
ICWS | 5 |
| 2023 | Understanding Customer Requirements - An Enterprise Knowledge Graph Approach
Basel Shbita, Anna Lisa Gentile, Pengyuan Li 0001, Chad DeLuca |
ESWC | 3 |
| 2023 | Long-Form Information Retrieval for Enterprise MatchmakingabstractUnderstanding customer requirements is a key success factor for both business-to-consumer (B2C) and business-to-business (B2B) enterprises. In a B2C context, most requirements are directly related to products and therefore expressed in keyword-based queries. In comparison, B2B requirements contain more information about customer needs and as such the queries are often in a longer form. Such long-form queries pose significant challenges to the information retrieval task in B2B context. In this work, we address the long-form information retrieval challenges by proposing a combination of (i) traditional retrieval methods, to leverage the lexical match from the query, and (ii) state-of-the-art sentence transformers, to capture the rich context in the long queries. We compare our method against traditional TF-IDF and BM25 models on an internal dataset of 12,368 pairs of long-form requirements and products sold. The evaluation shows promising results and provides directions for future work. Pengyuan Li 0001, Anna Lisa Gentile, Chad DeLuca, Daniel Tan 0002, Sandeep Gopisetty |
SIGIR | 1 |
| 2022 | Toward ECG-based analysis of hypertrophic cardiomyopathy: a novel ECG segmentation method for handling abnormalitiesabstractOBJECTIVE: Abnormalities in impulse propagation and cardiac repolarization are frequent in hypertrophic cardiomyopathy (HCM), leading to abnormalities in 12-lead electrocardiograms (ECGs). Computational ECG analysis can identify electrophysiological and structural remodeling and predict arrhythmias. This requires accurate ECG segmentation. It is unknown whether current segmentation methods developed using datasets containing annotations for mostly normal heartbeats perform well in HCM. Here, we present a segmentation method to effectively identify ECG waves across 12-lead HCM ECGs. METHODS: We develop (1) a web-based tool that permits manual annotations of P, P', QRS, R', S', T, T', U, J, epsilon waves, QRS complex slurring, and atrial fibrillation by 3 experts and (2) an easy-to-implement segmentation method that effectively identifies ECG waves in normal and abnormal heartbeats. Our method was tested on 131 12-lead HCM ECGs and 2 public ECG sets to evaluate its performance in non-HCM ECGs. RESULTS: Over the HCM dataset, our method obtained a sensitivity of 99.2% and 98.1% and a positive predictive value of 92% and 95.3% when detecting QRS complex and T-offset, respectively, significantly outperforming a state-of-the-art segmentation method previously employed for HCM analysis. Over public ECG sets, it significantly outperformed 3 state-of-the-art methods when detecting P-onset and peak, T-offset, and QRS-onset and peak regarding the positive predictive value and segmentation error. It performed at a level similar to other methods in other tasks. CONCLUSION: Our method accurately identified ECG waves in the HCM dataset, outperforming a state-of-the-art method, and demonstrated similar good performance as other methods in normal/non-HCM ECG sets. Kasra Nezamabadi, Jacob Mayfield, Pengyuan Li 0001, Gabriela V. Greenland, Sebastian Rodriguez 0001, Bahadir Simsek, Parvin Mousavi, Hagit Shatkay, M. Roselle Abraham |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | Skin lesion image classification method based on extension theory and deep learning
Xiaofei Bian, Haiwei Pan, Kejia Zhang 0001, Pengyuan Li 0001, Chunling Chen |
Multim. Tools Appl. | 4 |
| 2021 | ANIMO: Annotation of Biomed Image ModalitiesabstractFigures within biomedical articles present essential evidence of the relevance of a publication in a curation workflow. In particular, visual cues of the image modality or experimental methods can help expert curators identify relevant papers from an increasing number of publications. Automating the identification of these content-bearing images can thus be helpful in computer-assisted curation. However, the paucity of labeled datasets and the specialized training required to label such images hinder the development of such tools. To address this problem, we present the design of ANIMO, a labeling system that integrates extraction and segmentation tools to ease the annotation burden. We first introduce two taxonomies of image modalities and experimental methods, derived in collaboration with curators. On the back-end of the system, we process batches of documents and create a labeling task per document. At the front-end, expert curators can access these tasks through a web interface and access the article of interest. We describe the evaluation of this system by a group of biocurators, and the human factor lessons learned from this interdisciplinary experience. Juan Trelles Trabucco, Pengyuan Li 0001, Cecilia N. Arighi, Daniela Raciti, Hagit Shatkay, G. Elisabeta Marai |
BIBM | 2 |
| 2021 | Corrigendum to: Utilizing image and caption information for biomedical document classificationabstractBioinformatics (2021), Volume 37(Suppl1), i468–i476, doi:10.1093/bioinformatics/btab331 The error is thus only in mis-typing the formulae themselves; not in the actual calculations of the precision and the recall used throughout the paper. As such, the other parts of the manuscript – specifically the experimental results reported, all stand as they appear in the original publication, and are not impacted by this correction. Pengyuan Li 0001, Xiangying Jiang, Juan Trelles Trabucco, Daniela Raciti, Cynthia L. Smith, Martin Ringwald, G. Elisabeta Marai, Cecilia N. Arighi, Hagit Shatkay |
Bioinform. | 1 |
| 2021 | Utilizing image and caption information for biomedical document classificationabstractMOTIVATION: Biomedical research findings are typically disseminated through publications. To simplify access to domain-specific knowledge while supporting the research community, several biomedical databases devote significant effort to manual curation of the literature-a labor intensive process. The first step toward biocuration requires identifying articles relevant to the specific area on which the database focuses. Thus, automatically identifying publications relevant to a specific topic within a large volume of publications is an important task toward expediting the biocuration process and, in turn, biomedical research. Current methods focus on textual contents, typically extracted from the title-and-abstract. Notably, images and captions are often used in publications to convey pivotal evidence about processes, experiments and results. RESULTS: We present a new document classification scheme, using both image and caption information, in addition to titles-and-abstracts. To use the image information, we introduce a new image representation, namely Figure-word, based on class labels of subfigures. We use word embeddings for representing captions and titles-and-abstracts. To utilize all three types of information, we introduce two information integration methods. The first combines Figure-words and textual features obtained from captions and titles-and-abstracts into a single larger vector for document representation; the second employs a meta-classification scheme. Our experiments and results demonstrate the usefulness of the newly proposed Figure-words for representing images. Moreover, the results showcase the value of Figure-words, captions and titles-and-abstracts in providing complementary information for document classification; these three sources of information when combined, lead to an overall improved classification performance. AVAILABILITY AND IMPLEMENTATION: Source code and the list of PMIDs of the publications in our datasets are available upon request. Pengyuan Li 0001, Xiangying Jiang, Juan Trelles Trabucco, Daniela Raciti, Cynthia L. Smith, Martin Ringwald, G. Elisabeta Marai, Cecilia N. Arighi, Hagit Shatkay |
Bioinform. | 1 |
| 2020 | Modality-Classification of Microscopy Images Using Shallow Variants of Deep NetworksabstractMicroscopy images are pervasive in biomedical research publications, where images obtained through various microscopy modalities (light, fluorescence, scanning, transmission) are often used to describe and summarize experiments and contributions. Hence, there is growing interest in automatically identifying these microscopy images' modality and utilizing this knowledge in automated search tools. However, identifying microscopy images poses challenges due to a lack of extensive collections of labeled images. We describe and evaluate two alternative approaches to microscopy image classification. In the first approach, we progressively fine-tuned layers of ResNet models. The second approach uses shallow variants of ResNet networks, where we leverage the outputs from previous convolutional blocks. We compare these results against a Support Vector Machine (SVM)-based baseline. Our results show that fine-tuning specific layers yields better results than fine-tuning the whole model. Furthermore, shallower variants produce competitive results when compared to the entire fine-tuned model. Juan Trelles Trabucco, Pengyuan Li 0001, Cecilia N. Arighi, Hagit Shatkay, G. Elisabeta Marai |
BIBM | 2 |
| 2019 | Figure and caption extraction from biomedical documentsabstractMOTIVATION: Figures and captions convey essential information in biomedical documents. As such, there is a growing interest in mining published biomedical figures and in utilizing their respective captions as a source of knowledge. Notably, an essential step underlying such mining is the extraction of figures and captions from publications. While several PDF parsing tools that extract information from such documents are publicly available, they attempt to identify images by analyzing the PDF encoding and structure and the complex graphical objects embedded within. As such, they often incorrectly identify figures and captions in scientific publications, whose structure is often non-trivial. The extraction of figures, captions and figure-caption pairs from biomedical publications is thus neither well-studied nor yet well-addressed. RESULTS: We introduce a new and effective system for figure and caption extraction, PDFigCapX. Unlike existing methods, we first separate between text and graphical contents, and then utilize layout information to effectively detect and extract figures and captions. We generate files containing the figures and their associated captions and provide those as output to the end-user.We test our system both over a public dataset of computer science documents previously used by others, and over two newly collected sets of publications focusing on the biomedical domain. Our experiments and results comparing PDFigCapX to other state-of-the-art systems show a significant improvement in performance, and demonstrate the effectiveness and robustness of our approach. AVAILABILITY AND IMPLEMENTATION: Our system is publicly available for use at: https://www.eecis.udel.edu/~compbio/PDFigCapX. The two new datasets are available at: https://www.eecis.udel.edu/~compbio/PDFigCapX/Downloads. Pengyuan Li 0001, Xiangying Jiang, Hagit Shatkay |
Bioinform. | 1 |
| 2018 | Extracting Figures and Captions from Scientific PublicationsabstractFigures and captions convey essential information in scientific publications. As such, there is a growing interest in mining published figures and in utilizing their respective captions as a source of knowledge. There is also much interest in image captioning systems that can automatically generate captions for images, whose training requires large datasets of image-caption pairs. Notably, the first fundamental step of obtaining figures and captions from publications is neither well-studied nor yet well-addressed. In this paper, we introduce a new and effective system for figure and caption extraction, PDFigCapX. Unlike current methods that extract figures by handling raw encoded contents of PDF documents, we separate text from graphical contents and utilize layout information to detect and disambiguate figures and captions. Files containing the figures and their associated captions are then produced as output to the end-user. We test PDFigCapX on both a previously used generic dataset and on two new sets of publications within the biomedical domain. Our experiments and results show a significant improvement in performance compared to the state-of-the-art, and demonstrate the effectiveness of our approach. Our system will be available for use at: https://www.eecis.udel.edu/~compbio/PDFigCapX. Pengyuan Li 0001, Xiangying Jiang, Hagit Shatkay |
CIKM | 1 |
| 2018 | Compound image segmentation of published biomedical figuresabstractMotivation: Images convey essential information in biomedical publications. As such, there is a growing interest within the bio-curation and the bio-databases communities, to store images within publications as evidence for biomedical processes and for experimental results. However, many of the images in biomedical publications are compound images consisting of multiple panels, where each individual panel potentially conveys a different type of information. Segmenting such images into constituent panels is an essential first step toward utilizing images. Results: In this article, we develop a new compound image segmentation system, FigSplit, which is based on Connected Component Analysis. To overcome shortcomings typically manifested by existing methods, we develop a quality assessment step for evaluating and modifying segmentations. Two methods are proposed to re-segment the images if the initial segmentation is inaccurate. Experimental results show the effectiveness of our method compared with other methods. Availability and implementation: The system is publicly available for use at: https://www.eecis.udel.edu/~compbio/FigSplit. The code is available upon request. Contact: [email protected]. Supplementary information: Supplementary data are available online at Bioinformatics. Pengyuan Li 0001, Xiangying Jiang, Chandra Kambhamettu, Hagit Shatkay |
Bioinform. | 1 |
| 2017 | Identifying articles relevant to drug-drug interaction: Addressing class imbalanceabstractInteractions between drugs (also known as drug-drug interactions or DDIs), which may cause adverse affects, are of much concern; predicting, anticipating and avoiding them is key for improving patient safety and treatment outcome. Knowledge of DDIs is important for physicians to avoid adverse effects when prescribing two drugs simultaneously. DDIs are often published in the biomedical literature; however, gathering information about DDIs is time consuming given the shear volume of publications. Automatic text classification can speed up access to documents related to DDIs. However, the biomedical literature contains a relatively small number of publications relevant to DDIs, compared to the vast amount of irrelevant publications. This imbalance can lead to incorrect classification. While methods addressing class imbalance have been introduced to correctly identify items in the minority (relevant) class to improve recall, they often misclassify items in the majority (irrelevant) class, which leads to low precision. To reduce the number of irrelevant documents misclassified as relevant (false positive), we develop a two-stage cascade classifier. In each step, we separate publication abstracts that are DDI-relevant from those that are either drug-irrelevant or drug-relevant but DDI-irrelevant. We compare our classifier with other popular learning methods that aim to handle imbalance, applying the methods to a well-curated corpus consisting of DDI-relevant and DDI-irrelevant PubMed abstracts. Our method achieves higher precision and F1 measure than other methods while maintaining similar recall. Moumita Bhattacharya, Heng-Yi Wu, Pengyuan Li 0001, Lang Li 0001, Hagit Shatkay |
BIBM | 4 |
| 2017 | A medical image retrieval method based on texture block coding tree
Haiwei Pan, Pengyuan Li 0001, Xiaoqin Xie, Zhiqiang Zhang 0010 |
Signal Process. Image Commun. | 3 |
| 2015 | Finding Frequent Approximate Subgraphs in medical image databaseabstractMedical images are one of the most important tools in doctors' diagnostic decision-making. It has been a research hotspot in medical big data that how to effectively represent medical images and find essential patterns hidden in them to assist doctors to achieve a better diagnosis. Several graph models have been developed to represent medical images. However, the unique structures of domain-specific images are not considered well to lose some essential information. Thus, aiming at brain CT images, we first construct a graph about the Topological Relations between Ventricles and Lesions (TRVL) and present the graph modeling process. Then we propose a method named Frequent Approximate Subgraph Mining based on Graph Edit Distance (FASMGED). This method uses an error-tolerant graph matching strategy that is accordant with ubiquitous noise in practice. Experimental results show that the graph modeling process is computationally scalable and FASMGED can find more significant patterns than current algorithms. Linlin Gao, Haiwei Pan, Qilong Han, Xiaoqin Xie, Zhiqiang Zhang 0010, Xiao Zhai, Pengyuan Li 0001 |
BIBM | 7 |
| 2014 | Brain CT Image Similarity Retrieval Method Based on Uncertain Location GraphabstractA number of brain computed tomography (CT) images stored in hospitals that contain valuable information should be shared to support computer-aided diagnosis systems. Finding the similar brain CT images from the brain CT image database can effectively help doctors diagnose based on the earlier cases. However, the similarity retrieval for brain CT images requires much higher accuracy than the general images. In this paper, a new model of uncertain location graph (ULG) is presented for brain CT image modeling and similarity retrieval. According to the characteristics of brain CT image, we propose a novel method to model brain CT image to ULG based on brain CT image texture. Then, a scheme for ULG similarity retrieval is introduced. Furthermore, an effective index structure is applied to reduce the searching time. Experimental results reveal that our method functions well on brain CT images similarity retrieval with higher accuracy and efficiency. Haiwei Pan, Pengyuan Li 0001, Qing Li 0001, Qilong Han, Xiaoning Feng, Linlin Gao |
IEEE J. Biomed. Health Informatics | 2 |
| 2013 | A Novel Model for Medical Image Similarity Retrieval
Pengyuan Li 0001, Haiwei Pan, Qilong Han, Xiaoqin Xie, Zhiqiang Zhang 0010 |
WAIM | 1 |
| 2012 | Medical Image Retrieval Method Based on Relevance Feedback
Haiwei Pan, Qilong Han, Jingzi Gu, Pengyuan Li 0001 |
ADMA | 5 |