VLDB 2026 Research / reviewers in the wild / expert
Pengfei Yin
dblp:97/8996
· DBLP profile ↗
23ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Computer networks · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unleashing the Potential of Multiple WiFi APs in Real-World Localization SystemabstractWiFi indoor localization plays an important role in many real-world applications and has gained widespread attentions from both academia and industry over the past decade. While existing works have already achieved remarkable performance under various practical scenarios, they solely utilize multiple WiFi Access Points (APs) to achieve a higher accuracy and do not deeply explore the intrinsic relationships among them. To further unleash the potential of multiple APs and reach the limits of WiFi indoor localization, in this paper, we propose Qidi, a novel multi-APs collaboration based localization system. To the best of our knowledge, we are the first to explicitly reveal three underlying relationships among multiple spatially distributed APs, i.e., consistency, continuity, and non-uniformity. By combining these three basic principles with the unique characteristics of specific tasks, a series of long-lasting practical challenges, such as automatic phase offset calibration, bilateral angle ambiguity, elevation angle estimation with Uniform Linear Array (ULA) and Non-Line-of-Sight (NLoS), can be efficiently resolved. Extensive experiments in various complex environments are provided to demonstrate that Qidi can achieve 2.4°, 3.2° median errors of joint azimuth and elevation angle estimation, and 0.4mlocalization median error even in 20m×20mexhibition hall. Moreover, a one-month longitudinal evaluation conducted on a real-world deployed WiFi ISAC system further validates the effectiveness and robustness of the proposed localization system. Guanzhong Wang, Xuecheng Xie, Pengfei Yin, Ruiyuan Song, Dongheng Zhang, Yan Chen 0007 |
IEEE Internet Things J. | 5 |
| 2026 | CAFE: Towards Practical WiFi Localization via Continuous Angle Focusing EffectabstractWiFi-based indoor localization serves as a critical foundation for numerous real-world applications and has attracted widespread attentions over the past decade. Recent advancements have demonstrated the feasibility of achieving decimeter-level accuracy by leveraging Angle of Arrival (AoA) information. However, existing commercial WiFi Access Points (APs) suffer from phase offset across different antennas, which significantly degrade the performance of AoA-based methods. Previous works either relied on labor-intensive manual calibration or involved inaccurate and non-robust automatic calibration, which hinders their widespread use in large-scale deployments. Moreover, to expand the signal coverage and enhance communication performance, the inter-antenna spacing in existing commercial APs typically exceeds the standard half-wavelength. The resulting angle ambiguity problem can mislead target detection results, which has not been well resolved in existing works. To address the above two practical challenges, in this paper, we propose CAFE, a practical WiFi indoor localization system based on theContinuousAngleFocusingEffect. The key insight lies on the fact that the angle information of multiple APs originates from the same client, and thus exhibits highly convergent properties in both the temporal and spatial dimensions. By further exploring the binary nature of phase offset and the periodicity of grating lobes, our approach can efficiently resolve the above two practical challenges. Extensive experiments are provided to demonstrate the effectiveness of the proposed CAFE system, which outperforms state-of-the-art methods by$22.1\%$in median localization error for simple scenarios and by$37.1\%$for complex multipath scenarios. Pengfei Yin, Guanzhong Wang, Dongheng Zhang, Yan Chen 0007 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | ToP: a Structured Pathway for Document-level Information Extraction with Large Language ModelsabstractIn Natural Language Processing (NLP), Document-level Information Extraction (DocIE) poses a significant challenge, requiring analyzing and synthesizing information across extensive contexts. This complexity extends beyond the scale of data, delving into the nuanced interplay of context, semantics, and text elements. Recent advancements in Large Language Models (LLMs) have opened new prospects for DocIE. However, their application is limited by the models’ struggles with extended contexts and complex narrative structures, a gap highlighted when comparing the efficacy of direct LLM prompting against Chain-of-Thought (CoT) prompting. Addressing these challenges, we introduce the Tree of Plans (Top) framework, a novel approach that mirrors human decision-making by engaging in slow, deliberate reasoning. Top decomposes the DocIE task into three structured stages: (1) Task-specific Seed Plan Decomposition, (2) LLM-driven Plan Sampling and Voting, and (3) Instruction-prompted Plan Execution. Each stage leverages zero-shot prompting of an LLM, ensuring coherent progression and comprehensive document-level analysis. We evaluate Top on four challenging DocIE datasets, which cover document-level relation extraction and document-level event extraction tasks, both requiring complex reasoning over the long context. Pengfei Yin, Yubing Ren, Fang Fang 0009, Boxiang Hu |
IJCNN | 1 |
| 2025 | Investigation of Tea Yield Forecasting in Xiangxi Prefecture Utilizing Random Forest and Gradient Boosting Techniques
Xinyi Tao, Pengfei Yin, Junping Shi, Fanghui Mo |
ISNN | 4 |
| 2025 | Unsupervised Vast-Receptive-Field attention for blind Super-Resolution
Pengfei Yin, Weihua Ou, Kaibing Zhang |
Expert Syst. Appl. | 1 |
| 2025 | Measuring and visualizing healthcare process variabilityabstractIMPORTANCE: Understanding factors that contribute to clinical variability in patient care is critical, as unwarranted variability can lead to increased adverse events and prolonged hospital stays. Determining when this variability becomes excessive can be a step in optimizing patient outcomes and healthcare efficiency. OBJECTIVE: Explore the association between clinical variation and clinical outcomes. This study aims to identify the point in time when the relationship between clinical variation and length of stay (LOS) becomes significant. METHODS: This cohort study uses MIMIC-IV, a dataset collecting electronic health records of the Beth Israel Deaconess Medical Center in the United States. We focused on adult patients who underwent elective coronary bypass surgery, generating 847 patient observations. Demographic factors such as age, race, insurance type, and the Charlson Comorbidity Index (CCI) were recorded. We performed a variability analysis where patients' clinical processes are represented as sequences of events. The data was segmented based on the initial day of recorded activity to establish observation windows. Using a regression analysis, we identified the temporal window where variability's impact on LOS becomes independently significant. RESULT: Regression analysis revealed that patients in the top 20 % of the variability distance group experienced an 81 % increase in LOS (95 % CI: 1.72 to 1.91, p < 0.001). Insurance types, such as Medicare and Other, were associated with 18 % (95 % CI: 0.73 to 0.92, p < 0.001) and 21 % (95 % CI: 0.71 to 0.88, p < 0.001) decreases in LOS, respectively. Neither age nor race significantly affected LOS, but a higher CCI was associated with a 3.3 % increase in LOS (95 % CI: 1.02 to 1.05, p < 0.001). These findings indicate that higher variability and CCI significantly influence LOS, with insurance type also playing a crucial role. CONCLUSION: In the studied cohort, patient journeys with greater variability were associated with longer LOS with a dose-response relationship: the higher the variability, the longer LOS. This study presents a standardized way to measure and visualize variability in clinical processes and measure its impact on patient-relevant outcomes. Pengfei Yin, Abel Armas-Cervantes, Daniel Capurro |
J. Biomed. Informatics | 1 |
| 2024 | AutoCali: Enhancing AoA-based Indoor Localization through Automatic Phase CalibrationabstractRecent advancements in WiFi indoor localization have demonstrated the potential for achieving decimeter-level accuracy based on Angle of Arrival (AoA). However, existing commercial WiFi Access Points (APs) suffer from phase offset across different antennas, which significantly degrade the performance of AoA-based methods in practical deployment. Previous work either relied on labor-intensive manual calibration or involved inaccurate and non-robust automatic calibration. In this paper, we propose AutoCali, an accurate and robust automatic phase offset calibration system. The key insight is to utilize the binary nature of phase offsets and the property that triangulation exhibits higher convergence when the correct combination of phase offsets is employed. Extensive experiments demonstrate that AutoCali outperforms state-of-the-art methods by 22.1% in median localization error for simple scenarios and by 37.1% for complex multipath scenarios. Pengfei Yin, Dongheng Zhang, Guanzhong Wang, Yang Hu 0006, Yan Chen 0007 |
ICASSP | 1 |
| 2024 | Joint computation offloading and resource allocation for end-edge collaboration in internet of vehicles via multi-agent reinforcement learning
Cong Wang 0009, Ying Yuan 0001, Sancheng Peng, Guorui Li, Pengfei Yin |
Neural Networks | 6 |
| 2024 | Multiple WiFi Access Points Co-Localization Through Joint AoA EstimationabstractIndoor localization is a fundamental task to many real-world applications, which however remains unresolved, especially with commodity WiFi Access Points (APs). In this paper, we tackle this problem and propose an accurate, robust, and real-time indoor localization system that can be directly deployed on commodity WiFi infrastructure. Specifically, the proposed system makes three key contributions: 1) we introduce a non-parametric metric to measure the accuracy of Angle of Arrival (AoA) estimation; 2) we are the first to explicitly consider the relationship among the AoAs of different APs and propose a multiple APs co-localization algorithm to exploit such a relationship to improve the localization performance; 3) we propose several strategies to reduce the computational complexity of our system to achieve real-time localization. Extensive experiments are conducted to evaluate the performance of the proposed system under various situations, which demonstrate that the proposed system can achieve a 4 degrees median error of AoA estimation and 30 cm localization median error, outperforming the state-of-the-art systems. Dongheng Zhang, Ruiyuan Song, Pengfei Yin, Yan Chen 0007 |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Parsing Clinical Trial Eligibility Criteria for Cohort Query by a Multi-Input Multi-Output Sequence Labeling ModelabstractTo enable electronic screening of eligible patients for clinical trials, free-text clinical trial eligibility criteria should be translated to a computable format. Natural language processing (NLP) techniques have the potential to automate this process. In this study, we explored a supervised multi-input multi-output (MIMO) sequence labelling model to parse eligibility criteria into combinations of fact and condition tuples. Our experiments on a small manually annotated training dataset showed that that the performance of the MIMO framework with a BERT-based encoder using all the input sequences achieved an overall lenient-level AUROC of 0.61. Although the performance is suboptimal, representing eligibility criteria into logical and semantically clear tuples can potentially make subsequent translation of these tuples into database queries more reliable. Shubo Tian, Pengfei Yin, Hansi Zhang, Arslan Erdengasileng, Jiang Bian 0001, Zhe He 0001 |
BIBM | 2 |
| 2023 | ScSmOP: a universal computational pipeline for single-cell single-molecule multiomics data analysisabstractSingle-cell multiomics techniques have been widely applied to detect the key signature of cells. These methods have achieved a single-molecule resolution and can even reveal spatial localization. These emerging methods provide insights elucidating the features of genomic, epigenomic and transcriptomic heterogeneity in individual cells. However, they have given rise to new computational challenges in data processing. Here, we describe Single-cell Single-molecule multiple Omics Pipeline (ScSmOP), a universal pipeline for barcode-indexed single-cell single-molecule multiomics data analysis. Essentially, the C language is utilized in ScSmOP to set up spaced-seed hash table-based algorithms for barcode identification according to ligation-based barcoding data and synthesis-based barcoding data, followed by data mapping and deconvolution. We demonstrate high reproducibility of data processing between ScSmOP and published pipelines in comprehensive analyses of single-cell omics data (scRNA-seq, scATAC-seq, scARC-seq), single-molecule chromatin interaction data (ChIA-Drop, SPRITE, RD-SPRITE), single-cell single-molecule chromatin interaction data (scSPRITE) and spatial transcriptomic data from various cell types and species. Additionally, ScSmOP shows more rapid performance and is a versatile, efficient, easy-to-use and robust pipeline for single-cell single-molecule multiomics data analysis. Kai Jing, Yewen Xu, Yang Yang 0145, Pengfei Yin, Duo Ning, Guangyu Huang, Yuqing Deng, Gengzhan Chen, Guoliang Li 0002, Simon Zhongyuan Tian, Meizhen Zheng |
Briefings Bioinform. | 4 |
| 2023 | Corrigendum to "Unsupervised Simple Siamese Representation Learning for Blind Super-Resolution" [Eng. Appl. Artif. Intell. 114 (2022) 105092]
Pengfei Yin, Hua Huo, Kaibing Zhang |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | MCIBox: a toolkit for single-molecule multi-way chromatin interaction visualization and micro-domains identificationabstractThe emerging ligation-free three-dimensional (3D) genome mapping technologies can identify multiplex chromatin interactions with single-molecule precision. These technologies not only offer new insight into high-dimensional chromatin organization and gene regulation, but also introduce new challenges in data visualization and analysis. To overcome these challenges, we developed MCIBox, a toolkit for multi-way chromatin interaction (MCI) analysis, including a visualization tool and a platform for identifying micro-domains with clustered single-molecule chromatin complexes. MCIBox is based on various clustering algorithms integrated with dimensionality reduction methods that can display multiplex chromatin interactions at single-molecule level, allowing users to explore chromatin extrusion patterns and super-enhancers regulation modes in transcription, and to identify single-molecule chromatin complexes that are clustered into micro-domains. Furthermore, MCIBox incorporates a two-dimensional kernel density estimation algorithm to identify micro-domains boundaries automatically. These micro-domains were stratified with distinctive signatures of transcription activity and contained different cell-cycle-associated genes. Taken together, MCIBox represents an invaluable tool for the study of multiple chromatin interactions and inaugurates a previously unappreciated view of 3D genome structure. Simon Zhongyuan Tian, Guoliang Li 0002, Duo Ning, Kai Jing, Yewen Xu, Yang Yang 0145, Melissa Jane Fullwood, Pengfei Yin, Guangyu Huang, Dariusz Plewczynski, Jixian Zhai, Ziwei Dai, Meizhen Zheng |
Briefings Bioinform. | 8 |
| 2021 | Flexible Non-Autoregressive Extractive Summarization with Threshold: How to Extract a Non-Fixed Number of Summary SentencesabstractSentence-level extractive summarization is a fundamental yet challenging task, and recent powerful approaches prefer to pick sentences sorted by the predicted probabilities until the length limit is reached, a.k.a. ``Top-K Strategy''. This length limit is fixed based on the validation set, resulting in the lack of flexibility. In this work, we propose a more flexible and accurate non-autoregressive method for single document extractive summarization, extracting a non-fixed number of summary sentences without the sorting step. We call our approach ThresSum as it picks sentences simultaneously and individually from the source document when the predicted probabilities exceed a threshold. During training, the model enhances sentence representation through iterative refinement and the intermediate latent variables receive some weak supervision with soft labels, which are generated progressively by adjusting the temperature with a knowledge distillation algorithm. Specifically, the temperature is initialized with high value and drops along with the iteration until a temperature of 1. Experimental results on CNN/DM and NYT datasets have demonstrated the effectiveness of ThresSum, which significantly outperforms BERTSUMEXT with a substantial improvement of 0.74 ROUGE-1 score on CNN/DM. Our source code will be available on Github. Ruipeng Jia, Yanan Cao 0001, Haichao Shi, Fang Fang 0009, Pengfei Yin, Shi Wang 0002 |
AAAI | 5 |
| 2021 | A Survey of Interface Representations in Visual Programming Language Environments for Children's Physical Computing KitsabstractPhysical computing toolkits for children expose young minds to the concepts of computing and electronics within a target activity. To this end, these kits usually make use of a custom Visual Programming Language (or VPL) environment that extends past the functionality of simply programming, often also incorporating representations of electronics aspects in the interface. These representations of the electronics function as a scaffold to help the child focus on programming, instead of having to handle both the programming and details of the electronics at the same time. This paper presents a review of existing physical computing toolkits, looking at the What, How, and Where of electronics representations in their VPL interfaces. We then discuss potential research directions for the design of VPL interfaces for physical computing toolkits for children. Sarah Anne Brown, Sharon Lynn Chu Yew Yee, Pengfei Yin |
IDC | 3 |
| 2021 | Data and Model Biases in Social Media Analyses: A Case Study of COVID-19 Tweets
Pengfei Yin, Yongqiu Li, Xing He 0003, Jingcheng Du, Cui Tao, Yi Guo 0005, Mattia Prosperi, Pierangelo Veltri, Xi Yang 0015, Yonghui Wu 0001, Jiang Bian 0001 |
AMIA | 2 |
| 2021 | No news is an island: Joint heterogeneous graph network for news classificationabstractNews classification is important for people to organize web information. There are many connections between daily news, for example, they may be different reports about the same event or the same person. However, previous works classify news mainly based on single news, they ignore relationships between multiple news. This paper innovatively utilizes relationships between multiple news, such as their relevance in time, place and people, to classify news. To take full advantage of these relationships and integrate various information of multiple news, we propose News Classification Graph (NCG), a heterogeneous graph with different types of nodes and edges. Furthermore, we propose Joint Heterogeneous graph Network (JHN) to properly embed the NCG. It fully utilizes the information of heterogeneous nodes and heterogeneous edges in NCG. Extensive experiments carried on four news datasets demonstrate the effectiveness of our work in solving the news classification problem. Zhezhou Kang, Yi Liu 0067, Guanqun Bi, Fang Fang 0009, Pengfei Yin |
ISCC | 5 |
| 2020 | Enhancing Textual Representation for Abstractive Summarization: Leveraging Masked DecoderabstractFor existing models of abstractive summarization, the paradigm of autoregressive decoder inherently prefers relying on former tokens and the prediction error will propagate subsequently. To effectively eliminate the errors, we need a way to remodeling dependency during text generation. In this paper, we introduce MDSumma (as shorthand for Masked Decoder for Summarization), which masks partial tokens in decoder, aiming to alleviate the over-reliance on the antecedent. Moreover, with further facilitating the flexibility and diversity of textual representation, we employ a variational autoencoder model, sampling continuous latent variables from the probability distribution to explicitly model underlying semantics of the target summaries. Our architecture gives good balance between encoder contextual representation and decoder prediction, sidestepping the gap between training and inference. Experimental results on three benchmark datasets validate the effectiveness that our proposed method significantly outperforms the existing state-of-the-art approaches both on ROUGE and diversity scores. Ruipeng Jia, Yannan Cao, Fang Fang 0009, Jinpeng Li 0003, Yanbing Liu 0007, Pengfei Yin |
IJCNN | 6 |
| 2020 | Enhancing Pre-trained Language Representation for Multi-Task Learning of Scientific SummarizationabstractThis paper aims to extract summarization and keywords from scientific articles simultaneously, while abstract extraction (AE) and key extraction (KE) are considered as auxiliary tasks to each other. For the data scarcity in scientific AE and KE tasks, we propose a multi-task learning framework which uses huge unlabeled data to learn scientific language representation (pre-training) and uses smaller annotated data to transfer the learned representation to AE and KE (fine-tuning). Although the pre-trained language model performs well in universal natural language tasks, its capacity still has a margin of improvement for specific tasks. Inspired by this intuition, we use another two tasks keyword masking and key sentence prediction before the fine-tuning phase to enhance the language representation for AE and KE. This language representation enhancing stage uses the same labeled data but different optimization objectives with the fine-tuning phase. In order to evaluate our model, we develop and release a high-quality annotated corpus for scientific papers with keywords and abstract. We conduct comparative experiments on this dataset, and experimental results show that our multi-task learning framework achieves the state-of-the-art performance, proving the effectiveness of the language model enhancing mechanism. Ruipeng Jia, Yannan Cao, Fang Fang 0009, Jinpeng Li 0003, Yanbing Liu 0007, Pengfei Yin |
IJCNN | 6 |
| 2020 | HIN: Hierarchical Inference Network for Document-Level Relation Extraction
Hengzhu Tang, Yanan Cao 0001, Zhenyu Zhang 0006, Jiangxia Cao, Fang Fang 0009, Shi Wang 0002, Pengfei Yin |
PAKDD (1) | 7 |
| 2010 | Spectral data analysis of ground objects in Chao Lake basinabstractBased on a great deal of spectral data for different kinds of ground objects which were denoised by wavelet transform method, the spectral characteristics and changing rules of water, paddy field, wheat, cole and vegetable greenhouse in Chao Lake basin were analyzed with ASD portable spectrum analyzer by field investigation and plot survey. The spectral data of typical ground objects in Chao Lake basin were also processed by using mathematics technology such as de-noising technology of derivative spectrum technology and normalization processing technology. It was shows that wavelet transform had advantage in de-noising because it could remove the noise from signal as well as preserve the detail information; derivative spectrum and normalization processing technology had better practicability in suppressing background effect and emphasizing the signal of objects. Qiu Yin, Li Li 0017, Zhenghua Chen, Yuhuan Ren, Weizhen Hou, Pengfei Yin |
IGARSS | 8 |
| 2010 | An atmospheric correction algorithm for hyperspectral imagery of lake water by Chinese satellite HJ-1AabstractThis paper demonstrates the Ruddick's algorithm to utilize atmospheric correction with the hyper-spectral imagery over Chinese turbid lake water obtained by China first hyperspectral imager (HSI) onboard HJ-1A. The paper studies on the sensor characteristics and analyzes the optical properties of turbid lake water in Taihu Lake. Based on consideration about real circumstances, the paper recalibrates parameterαof value taken as 1.43. Results indicate that the recalibrated parameter could enhance the algorithm performance and improve the accuracy through comparison with the in situ measurements. Xingfa Gu, Qiu Yin, Li Li 0017, Zhenghua Chen, Yuhuan Ren, Weizhen Hou, Pengfei Yin |
IGARSS | 9 |
| 2010 | Ocean surface winds measurement using reflected GNSS signalsabstractSignals of the Global Navigation Satellite System (GNSS) can be used for purposes rather than navigation, positioning and timing. Direct GNSS signals have already been used to obtain the earth's atmospheric profile. Now it has been recognized that the reflected GNSS signals can also be used to obtain ocean surface wind speed and direction. Based on the aircraft flight experiments taken in China, this paper presents a thorough and extended analysis of this surface wind retrieval algorithm. In order to verify the accuracy of the retrieval algorithm, the retrieval results have been compared with QuickSCAT satellite's ocean surface winds sounding results. The retrieval accuracies of wind vectors are 2m/s wind speed and 2 deg wind direction. Pengfei Yin, Qiu Yin, Yiqiang Zhang, Xingfeng Chen, Shangguo Ning |
IGARSS | 1 |