VLDB 2026 Research / reviewers in the wild / expert
Peize Li
dblp:254/1325
· DBLP profile ↗
17ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolving Beyond Snapshots: Harmonizing Structure and Sequence via Entity State Tuning for Temporal Knowledge Graph ForecastingabstractSiyuan Li, Yunjia Wu, Yiyong Xiao, Pingyang Huang, Peize Li, Ruitong Liu, Yan Wen, Te Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yunjia Wu, Yiyong Xiao, Pingyang Huang, Peize Li, Te Sun 0001 |
ACL (1) | 5 |
| 2026 | LISA-Net: Lateral Inhibition and Structure Alignment Network
Peize Li, Qilan Lin |
Pattern Recognit. | 2 |
| 2025 | Multimodal Hypothetical Summary for Retrieval-based Multi-image Question AnsweringabstractRetrieval-based multi-image question answering (QA) task involves retrieving multiple question-related images and synthesizing these images to generate an answer. Conventional "retrieve-then-answer" pipelines often suffer from cascading errors because the training objective of QA fails to optimize the retrieval stage. To address this issue, we propose a novel method to effectively introduce and reference retrieved information into the QA. Given the image set to be retrieved, we employ a multimodal large language model (visual perspective) and a large language model (textual perspective) to obtain multimodal hypothetical summary in question-form and description-form. By combining visual and textual perspectives, MHyS captures image content more specifically and replaces real images in retrieval, which eliminates the modality gap by transforming into text-to-text retrieval and helps improve retrieval. To more advantageously introduce retrieval with QA, we employ contrastive learning to align queries (questions) with MHyS. Moreover, we propose a coarse-to-fine strategy for calculating both sentence-level and word-level similarity scores, to further enhance retrieval and filter out irrelevant details. Our approach achieves a 3.7% absolute improvement over state-of-the-art methods on RETVQA and a 14.5% improvement over CLIP. Comprehensive experiments and detailed ablation studies demonstrate the superiority of our method. Peize Li, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Yan Wang 0028 |
AAAI | 1 |
| 2025 | KA-CDRE: Knowledge-Augmented Cross-Document Relation Extraction
Peize Li, Jingzi Gu, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005 |
ADMA (4) | 2 |
| 2025 | Research on the Simultaneous Measurement Method of Atmospheric Temperature and Pressure Profiles Based on the Lidar Echo Energy and Characteristics of the Scattering SpectrumabstractLidar plays an important role in the detection of atmospheric environmental parameters due to its property of high detection accuracy, good real-time performance, and high spatiotemporal resolution. Measuring the spectral properties of molecular scattered light enables to determine the profiles of atmospheric parameters such as wind speed, temperature, and pressure with high vertical resolution (~100 m) by using sophisticated multiparameter inversion algorithms. So far, atmospheric measurements of Rayleigh–Brillouin (RB) spectrum were used to only determine temperature, using pressure from respective models or from measurements such as radiosondes. In this article, we propose a new simultaneous inversion method of atmospheric temperature and pressure based on lidar echo energy and the information contained in the RB spectrum. The accuracy and uncertainty of the proposed method are analyzed based on numerical simulations and atmospheric measurements and compared to state-of-the-art retrieval methods. It is shown that the accuracy of retrieved temperature can be improved from 2.0 to 1.1 K and one of the pressures from 1.2% to 0.8%. Furthermore, the potential of additionally retrieving wind speed and the aerosol extinction coefficient based on spectral characteristics and intensity of the scattered light is investigated and discussed. This lays a foundation for the future development and application lidar instruments capable of multiparameter retrieval. Yangrui Xu, Yujun Gu, Peize Li, Yanpeng Zhao, Kun Liang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Object Attribute Matters in Visual Question AnsweringabstractVisual question answering is a multimodal task that requires the joint comprehension of visual and textual information. However, integrating visual and textual semantics solely through attention layers is insufficient to comprehensively understand and align information from both modalities. Intuitively, object attributes can naturally serve as a bridge to unify them, which has been overlooked in previous research. In this paper, we propose a novel VQA approach from the perspective of utilizing object attribute, aiming to achieve better object-level visual-language alignment and multimodal scene understanding. Specifically, we design an attribute fusion module and a contrastive knowledge distillation module. The attribute fusion module constructs a multimodal graph neural network to fuse attributes and visual features through message passing. The enhanced object-level visual features contribute to solving fine-grained problem like counting-question. The better object-level visual-language alignment aids in understanding multimodal scenes, thereby improving the model's robustness. Furthermore, to augment scene understanding and the out-of-distribution performance, the contrastive knowledge distillation module introduces a series of implicit knowledge. We distill knowledge into attributes through contrastive loss, which further strengthens the representation learning of attribute features and facilitates visual-linguistic alignment. Intensive experiments on six datasets, COCO-QA, VQAv2, VQA-CPv2, VQA-CPv1, VQAvs and TDIUC, show the superiority of the proposed method. Peize Li, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Yan Wang 0028 |
AAAI | 1 |
| 2024 | E/I Balanced Adaptive Sequential Neural Posterior Estimation for Inferring the Connection Weights in Mouse V1 ModelabstractEffectively utilizing biological firing rate data to estimate the numerous connection weights in the mouse primary visual cortex (V1) model from the Allen Institute is a challenging task. The existing iterative grid-search algorithm cannot enable the mouse V1 model to better fit the biological firing rate data. To tackle this issue, we propose an excitation-inhibition balanced adaptive sequential neural posterior estimation (E/I balanced ASNPE) approach to accurately infer the connection weights of the mouse V1 model, allowing the neurons’ firing rates to converge to the given biological data. This method fully leverages the structural information of the mouse V1 model, reducing the dimensionality of the weight parameters to be optimized. Initially, sampling is performed in the prior distribution based on the proposed non-dominated sorting adaptive genetic algorithm (NSAGA). This algorithm optimizes the sorting, crossover and mutation processes based on the fitness scores of the current samples and updates the proposal distribution based on these samples, increasing the likelihood of identifying high posterior probability regions in the prior distribution. To avoid bad simulations, we also explore the E/I balance in each layer of the mouse V1 model, adding biological constraints during weight inference with Automatic Posterior Transformation (APT). Experimental results confirm that the proposed E/I balanced ASNPE method significantly outperforms the baseline in all five firing rate fitness scores in the mouse V1 model. This study is pioneering in applying non-dominated sorting genetic algorithms combined with sequential neural posterior estimation to optimize connection weights in large-scale complex biological models. Luntian Mou, Peize Li, Lei Ma 0008, Tiejun Huang 0001 |
BIBM | 2 |
| 2024 | Single- Frame Background Reconstruction Based on Image Semantic PropagationabstractImage background recognition and reconstruction is an important research content in the field of computer vision. Instead of paying attention to all the targets in the image, this paper only focuses on the static background in the scene. Therefore, an adaptive semantic propagation method is proposed to reconstruct the complete background of static images. It executes single-connected region division and edge detection on a single frame. Experimental results demonstrate the robustness and adaptability of the method, achieving convincing performance under both simple and complex road conditions. Yang Zhang 0032, Peize Li |
ICARCV | 2 |
| 2024 | Multimodal Indoor Localization Using Crowdsourced Radio MapsabstractIndoor Positioning Systems (IPS) traditionally rely on odometry and building infrastructures like WiFi, often supplemented by building floor plans for increased accuracy. However, the limitation of floor plans in terms of availability and timeliness of updates challenges their wide applicability. In contrast, the proliferation of smartphones and WiFi-enabled robots has made crowdsourced radio maps – databases pairing locations with their corresponding Received Signal Strengths (RSS) – increasingly accessible. These radio maps not only provide WiFi fingerprint-location pairs but encode movement regularities akin to the constraints imposed by floor plans. This work investigates the possibility of leveraging these radio maps as a substitute for floor plans in multimodal IPS. We introduce a new framework to address the challenges of radio map inaccuracies and sparse coverage. Our proposed system integrates an uncertainty-aware neural network model for WiFi localization and a bespoken Bayesian fusion technique for optimal fusion. Extensive evaluations on multiple real-world sites indicate a significant performance enhancement, with results showing ∼ 25% improvement over the best baseline. Zhaoguang Yi, Xiangyu Wen 0001, Qiyue Xia, Peize Li, Francisco Zampella, Firas Alsehly, Xiaoxuan Lu 0001 |
ICRA | 4 |
| 2024 | Iterative Fine-Grained Genetic Algorithm for Inferring Connection Weights in Large-Scale Biophysical Mouse V1 Model
Peize Li, Tiejun Huang 0001 |
PRICAI (4) | 3 |
| 2024 | Cross-modality Multiple Relations Learning for Knowledge-based Visual Question AnsweringabstractKnowledge-based visual question answering not only needs to answer the questions based on images but also incorporates external knowledge to study reasoning in the joint space of vision and language. To bridge the gap between visual content and semantic cues, it is important to capture the question-related and semantics-rich vision-language connections. Most existing solutions model simple intra-modality relation or represent cross-modality relation using a single vector, which makes it difficult to effectively model complex connections between visual features and question features. Thus, we propose a cross-modality multiple relations learning model, aiming to better enrich cross-modality representations and construct advanced multi-modality knowledge triplets. First, we design a simple yet effective method to generate multiple relations that represent the rich cross-modality relations. The various cross-modality relations link the textual question to the related visual objects. These multi-modality triplets efficiently align the visual objects and corresponding textual answers. Second, to encourage multiple relations to better align with different semantic relations, we further formulate a novel global-local loss. The global loss enables the visual objects and corresponding textual answers close to each other through cross-modality relations in the vision-language space, and the local loss better preserves semantic diversity among multiple relations. Experimental results on the Outside Knowledge VQA and Knowledge-Routed Visual Question Reasoning datasets demonstrate that our model outperforms the state-of-the-art methods. Yan Wang 0028, Peize Li, Qingyi Si, Hanwen Zhang 0010, Wenyu Zang, Zheng Lin 0001, Peng Fu 0008 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Robust Human Detection under Visual Degradation via Thermal and mmWave Radar Fusion
Kaiwen Cai, Qiyue Xia, Peize Li, Xiaoxuan Lu 0001, John A. Stankovic |
EWSN | 3 |
| 2023 | Feature-based Visual Odometry for Bronchoscopy: A Dataset and BenchmarkabstractBronchoscopy is a medical procedure that involves the insertion of a flexible tube with a camera into the airways to survey, diagnose and treat lung diseases. Due to the complex branching anatomical structure of the bronchial tree and the similarity of the inner surfaces of the segmental airways, navigation systems are now being routinely used to guide the operator during procedures to access the lung periphery. Current navigation systems rely on sensor-integrated bronchoscopes to track the position of the bronchoscope in real-time. This approach has limitations, including increased cost and limited use in non-specialized settings. To address this issue, researchers have proposed visual odometry algorithms to track the bronchoscope camera without the need for external sensors. However, due to the lack of publicly available datasets, limited progress is made. To this end, we have developed a database of bronchoscopy videos in a phantom lung model and ex-vivo human lungs. The dataset contains 34 video sequences with over 23,000 frames with odometry ground truth data collected using electromagnetic tracking sensors. With our dataset, we empower the robotics and machine learning community to advance the field. We share our insights on challenges in endoscopic visual odometry. Furthermore, we provide benchmark results for this dataset. State-of-the-art feature extraction algorithms including SIFT, ORB, Superpoint, Shi- Tomasi, and LoFTR are tested on this dataset. The benchmark results demonstrate that the LoFTR algorithm outperforms other approaches, but still has significant errors in the presence of rapid movements and occlusions. Jianning Deng, Peize Li, Kevin Dhaliwal, Xiaoxuan Lu 0001, Mohsen Khadem |
IROS | 2 |
| 2023 | Democratizing Pathological Image Segmentation with Lay Annotators via Molecular-Empowered Learning
Ruining Deng, Peize Li, Jiacheng Wang 0007, Lucas W. Remedios, Saydolimkhon Agzamkhodjaev, Zuhayr Asad, Quan Liu 0002, Can Cui 0006, Yaohong Wang, Yucheng Tang, Haichun Yang, Yuankai Huo |
MICCAI (6) | 3 |
| 2022 | OdomBeyondVision: An Indoor Multi-modal Multi-platform Odometry Dataset Beyond the Visible SpectrumabstractThis paper presents a multimodal indoor odometry dataset, OdomBeyondVision, featuring multiple sensors across the different spectrum and collected with different mobile platforms. Not only does OdomBeyondVision contain the traditional navigation sensors, sensors such as IMUs, mechanical LiDAR, RGBD camera, it also includes several emerging sensors such as the single-chip mmWave radar, LWIR thermal camera and solid-state LiDAR. With the above sensors on UAV, UGV and handheld platforms, we respectively recorded the multimodal odometry data and their movement trajectories in various indoor scenes and different illumination conditions. We release the exemplar radar, radar-inertial and thermal-inertial odometry implementations to demonstrate their results for future works to compare against and improve upon. The full dataset including toolkit and documentation is publicly available at: https://github.com/MAPS-Lab/OdomBeyondVision. Peize Li, Kaiwen Cai, Muhamad Risqi Utama Saputra, Zhuangzhuang Dai, Xiaoxuan Lu 0001 |
IROS | 1 |
| 2021 | Motion Tracklet Oriented 6-DoF Inertial Tracking Using Commodity SmartphonesabstractMotion tracklets are the basic fragments of the track followed by a moving object and constitute various everyday motion behavior. An accurate estimation of motion tracklets in 3-D space can enable a wide range of applications, ranging from human computer interaction to medical rehabilitation. This paper presents a novel dataset for accurate 6-DoF motion tracklet estimation with the inertial sensors on commodity smartphones. The dataset consists of around 100 minutes of handheld motion with 3 predominant types of motion track-lets and accurate ground truth using the Vicon systems. With the presented dataset, we further benchmarked the trajectory estimation using a lightweight neural odometry model, showcasing how the dataset can be used while providing quantitative performance for downstream tasks. Our dataset, toolkit and source code available at https://github.com/MAPS-Lab/smartphone-tracking-dataset. Peize Li, Xiaoxuan Lu 0001 |
SenSys | 1 |
| 2019 | Production service system enabled by cloud-based smart resource hierarchy for a highly dynamic synchronized production process
Ting Qu 0002, Hongfei Jiang, Peize Li, Jinjie Xiang, George Q. Huang |
Adv. Eng. Informatics | 5 |