VLDB 2026 Research / reviewers in the wild / expert
Peiming Li
dblp:227/2099
· DBLP profile ↗
16ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MP1: MeanFlow Tames Policy Learning in 1-step for Robotic ManipulationabstractIn robot manipulation, robot learning has become a prevailing approach. However, generative models within this field face a fundamental trade-off between the slow, iterative sampling of diffusion models and the architectural constraints of faster Flow-based methods, which often rely on explicit consistency losses. To address these limitations, we introduce MP1, which pairs 3D point-cloud inputs with the MeanFlow paradigm to generate action trajectories in one network function evaluation (1-NFE). By directly learning the interval-averaged velocity via the "MeanFlow Identity", our policy avoids any additional consistency constraints. This formulation eliminates numerical ODE-solver errors during inference, yielding more precise trajectories. MP1 further incorporates CFG for improved trajectory controllability while retaining 1-NFE inference without reintroducing structural constraints. Because subtle scene-context variations are critical for robot learning, especially in few-shot learning, we introduce a lightweight Dispersive Loss that repels state embeddings during training, boosting generalization without slowing inference. We validate our method on the Adroit and Meta-World benchmarks, as well as in real-world scenarios. Experimental results show MP1 achieves superior average task success rates, outperforming DP3 by 10.2% and FlowPolicy by 7.3%. Its average inference time is only 6.8 ms—19 times faster than DP3 and nearly 2 times faster than FlowPolicy. Juyi Sheng, Peiming Li, Mengyuan Liu 0001 |
AAAI | 3 |
| 2026 | Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent ReasoningabstractChain-of-Thought (CoT) prompting has achieved remarkable success in unlocking the reasoning capabilities of Large Language Models (LLMs).Although CoT prompting enhances reasoning, its verbosity imposes substantial computational overhead.Recent works often focus exclusively on outcome alignment and lack supervision on the intermediate reasoning process.These deficiencies obscure the analyzability of the latent reasoning chain.To address these challenges, we introduce Render-of-Thought (RoT), the first framework to reify the reasoning chain by rendering textual steps into images, making the latent rationale explicit and traceable.Specifically, we leverage the vision encoders of existing Vision Language Models (VLMs) as semantic anchors to align the vision embeddings with the textual space.This design ensures plug-and-play implementation without incurring additional pre-training overhead.Extensive experiments on mathematical and logical reasoning benchmarks demonstrate that our method achieves 3-4× token compression and substantial inference acceleration compared to explicit CoT.Furthermore, it demonstrates a competitive efficiency-accuracy Pareto exploration compared to other methods, validating the feasibility of this paradigm.Our code is available at https://github.com/TencentBAC/RoT Peiming Li |
ACL (1) | 3 |
| 2026 | FiLM-DiffRec: Lightweight feature-wise modulation for enhanced timestep conditioning in diffusion recommender systems
Peiming Li, Shaohua Tuo, Zhixin Luo |
Expert Syst. Appl. | 2 |
| 2026 | CDG-Rec: Contrastive-Regularized diffusion with a Synergistic Dual-view Decoder for recommendation
Shaohua Tuo, Sibeier Chen, Peiming Li |
Expert Syst. Appl. | 4 |
| 2026 | PePNet: Pose-Enhanced Point Cloud Network for LiDAR-Based Human Action Recognition in Outdoor Long-Range ScenariosabstractWith potential applications in robotics and autonomous vehicles, LiDAR-based human action recognition (HAR) in outdoor long-range scenarios is challenging due to the degradation of point cloud density with distance and the simultaneous motion of humans and sensors. To address these issues, we propose the Pose-Enhanced Point Cloud Network (PePNet), a distance-aware framework for long-range HAR. As the core component, the Pose-Enhanced Point Cloud Block (PeP Block) integrates three modules: a Dynamic Enhancement Module that mitigates point cloud sparsity at long distances by generating supplementary points from motion cues, a Pose Prompter Module that introduces pose priors, and an Adaptive Point Selection Module that suppresses irrelevant body-part movements. We further design a Spatiotemporal Tube Embedding (ST-Tube), combined with the Mamba state space model, to capture long-range dependencies and complex motion dynamics. In addition, we construct Momo, a large-scale LiDAR-based HAR dataset that focuses on long-range (2-30 m) outdoor scenarios where sparse point clouds and simultaneous human-sensor motion pose prominent challenges, complementing existing benchmarks by providing a dedicated evaluation platform for long-range outdoor HAR. Experimental results show that PePNet achieves consistent performance gains over existing methods on Momo. Moreover, the proposed PeP Block can serve as a plug-and-play module to enhance other point cloud action recognition frameworks in long-range outdoor settings. The code is available at https://github.com/Shark0-0/PePNet. Mengyuan Liu 0001, Zhichao Deng, Peiming Li, Jun Liu 0036 |
IEEE Trans. Image Process. | 5 |
| 2026 | Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action RecognitionabstractRGB camera-based surveillance systems enable human action recognition for public safety and healthcare, yet raise serious privacy concerns. Existing methods rely on post-capture algorithms, which fail to protect privacy during data acquisition. We propose Lens Privacy Sealing (LPS), a simple hardware solution that physically obscures camera lenses with adjustable laminating film, providing pre-sensor privacy protection at minimal cost. Unlike software methods or expensive engineered optics, LPS achieves strong privacy through stochastic multi-layer scattering that is physically irreversible. We introduce the P3AR dataset for privacy-preserving action recognition, featuring both large-scale replay-captured (P3AR-NTU, 114K videos) and real-world collected (P3AR-PKU) subsets with privacy attribute annotations. To handle video degradation from LPS, we propose MSPNet, a single-stage framework incorporating Inter-Frame Noise Suppressor (IFNS) and Cross-Frame Semantic Aggregator (CFSA), enhanced by contrastive language-image pre-training for robust semantic extraction. Extensive experiments demonstrate that MSPNet with IFNS and CFSA nearly doubles action recognition accuracy compared to baseline methods while suppressing identity recognition to low levels. Comprehensive validation shows LPS achieves a superior privacy-utility trade-off compared to state-of-the-art hardware methods, resists reconstruction attacks including PSF inversion and data-driven recovery, and generalizes robustly across optical configurations and challenging environments. Code is available at https://github.com/wangzy01/MSPNet. Mengyuan Liu 0001, Peiming Li, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | EVENS: Equality versus Equity Notion Spectrum of LLMsabstractThe controversy surrounding COMPAS exposes a significant gap between computer science and social science in understanding bias, highlighting the need to align computational fairness metrics with humanistic interpretations. In response, we introduce the EVENS benchmark to align the Equality versus Equity Notion Spectrum in LLMs. Our contributions include constructing an equality–equity notion spectrum and generating a corresponding dataset of key fairness scenarios, evaluating models’ initial stances and test stance adjustments under external legal regulations and internal organizational regulations using Retrieval-Augmented Generation (RAG), introducing Chain-of-Thought (CoT) prompting to guide fairness reasoning, and adding an uncertain choice to assess its impact. Our findings indicate that LLMs initially favor equality over equity. Incorporating legal and organizational regulations of equity through RAG can reduce proportional equality in most models and enhance equity recognition in GPT4o significantly. CoT improves the equity reasoning of Chinese models but may also rationalize existing biases, and the uncertain option promotes more cautious responses. The links to the code and datasets: https://github.com/CrexCheng/EVEN. Qingjing Chen, Rongxin Cheng 0002, Ziheng Xie, Kangxin Zhao, Peiming Li, Antonino Rotolo, Yun Liu 0033, Weixing Shen |
ICAIL | 6 |
| 2025 | UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video ModelingabstractPoint cloud videos capture dynamic 3D motion while reducing the effects of lighting and viewpoint variations, making them highly effective for recognizing subtle and continuous human actions. Although Selective State Space Models (SSMs) have shown good performance in sequence modeling with linear complexity, the spatio-temporal disorder of point cloud videos hinders their unidirectional modeling when directly unfolding the point cloud video into a 1D sequence through temporally sequential scanning. To address this challenge, we propose the Unified Spatio-Temporal State Space Model (UST-SSM), which extends the latest advancements in SSMs to point cloud videos. Specifically, we introduce Spatial-Temporal Selection Scanning (STSS), which reorganizes unordered points into semantic-aware sequences through prompt-guided clustering, thereby enabling the effective utilization of points that are spatially and temporally distant yet similar within the sequence. For missing 4D geometric and motion details, Spatio-Temporal Structure Aggregation (STSA) aggregates spatio-temporal features and compensates. To improve temporal interaction within the sampled sequence, Temporal Interaction Sampling (TIS) enhances fine-grained temporal dependencies through non-anchor frame utilization and expanded receptive fields. Experimental results on the MSR-Action3D, NTU RGB+D, and Synthia 4D datasets validate the effectiveness of our method. Our code is available at https://github.com/wangzy01/UST-SSM. Peiming Li, Yulin Yuan, Hong Liu 0008, Xiangming Meng, Junsong Yuan 0001, Mengyuan Liu 0004 |
ICCV | 1 |
| 2025 | Recognizing Actions From Robotic View for Natural Human-Robot Interaction
Peiming Li, Hong Liu 0008, Zhichao Deng, Can Wang 0006, Jun Liu 0036, Junsong Yuan 0001, Mengyuan Liu 0001 |
ICCV | 2 |
| 2025 | MER-CLIP: AU-Guided Vision-Language Alignment for Micro-Expression RecognitionabstractAs a critical psychological stress response, micro-expressions (MEs) are fleeting and subtle facial movements revealing genuine emotions. Automatic ME recognition (MER) holds valuable applications in fields such as criminal investigation and psychological diagnosis. The Facial Action Coding System (FACS) encodes expressions by identifying activations of specific facial action units (AUs), serving as a key reference for ME analysis. However, current MER methods typically limit AU utilization to defining regions of interest (ROIs) or relying on specific prior knowledge, often resulting in limited performance and poor generalization. To address this, we integrate the CLIP model's powerful cross-modal semantic alignment capability into MER and propose a novel approach namely MER-CLIP. Specifically, we convert AU labels into detailed textual descriptions of facial muscle movements, guiding fine-grained spatiotemporal ME learning by aligning visual dynamics and textual AU-based representations. Additionally, we introduce an Emotion Inference Module to capture the nuanced relationships between ME patterns and emotions with higher-level semantic understanding. To mitigate overfitting caused by the scarcity of ME data, we put forward LocalStaticFaceMix, an effective data augmentation strategy blending facial images to enhance facial diversity while preserving critical ME features. Finally, comprehensive experiments on four benchmark ME datasets confirm the superiority of MER-CLIP. Notably, UF1 scores on CAS(ME)$^{3}$reach 0.7832, 0.6544, and 0.4997 for 3-, 4-, and 7-class classification tasks, significantly outperforming previous methods. Xinglong Mao, Sirui Zhao, Peiming Li, Tong Xu 0001, Enhong Chen |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | ClickDiff: Click to Induce Semantic Contact Map for Controllable Grasp Generation with Diffusion ModelsabstractGrasp generation aims to create complex hand-object interactions with a specified object. While traditional approaches for hand generation have primarily focused on visibility and diversity under scene constraints, they tend to overlook the fine-grained hand-object interactions such as contacts, resulting in inaccurate and undesired grasps. To address these challenges, we propose a controllable grasp generation task and introduce ClickDiff, a controllable conditional generation model that leverages a fine-grained Semantic Contact Map (SCM). Particularly when synthesizing interactive grasps, the method enables the precise control of grasp synthesis through either user-specified or algorithmically predicted Semantic Contact Map. Specifically, to optimally utilize contact supervision constraints and to accurately model the complex physical structure of hands, we propose a Dual Generation Framework. Within this framework, the Semantic Conditional Module generates reasonable contact maps based on fine-grained contact information, while the Contact Conditional Module utilizes contact maps alongside object point clouds to generate realistic grasps. We evaluate the evaluation criteria applicable to controllable grasp generation. Both unimanual and bimanual generation experiments on GRAB and ARCTIC datasets verify the validity of our proposed method, demonstrating the efficacy and robustness of ClickDiff, even with previously unseen objects. Our code is available at https://github.com/adventurer-w/ClickDiff. Peiming Li, Mengyuan Liu 0004, Hong Liu 0008, Chen Chen 0001 |
ACM Multimedia | 1 |
| 2024 | Breaking the water dilemma: Transmission-guided bilevel adaptive learning for underwater imagery
Sihan Xie, Peiming Li, Jiaxin Gao 0001, Ziyu Yue, Xin Fan 0001, Risheng Liu |
Neurocomputing | 2 |
| 2022 | Channel Knowledge Map for Environment-Aware Communications: EM Algorithm for Map ConstructionabstractChannel knowledge map (CKM) is an emerging technique to enable environment-aware wireless communications, in which databases with location-specific channel knowledge are used to facilitate or even obviate real-time channel state information acquisition. One fundamental problem for CKM-enabled communication is how to efficiently construct the CKM based on finite measurement data points at limited user locations. Towards this end, this paper proposes a novel map construction method based on the expectation maximization (EM) algorithm, by utilizing the available measurement data, jointly with the expert knowledge of well-established statistic channel models. The key idea is to partition the available data points into different groups, where each group shares the same modelling parameter values to be determined. We show that determining the modelling parameter values can be formulated as a maximum likelihood estimation problem with latent variables, which is then efficiently solved by the classic EM algorithm. Compared to the pure data-driven methods such as the nearest neighbor based interpolation, the proposed method is more efficient since only a small number of modelling parameters need to be determined and stored. Furthermore, the proposed method is extended for constructing a specific type of CKM, namely, the channel gain map (CGM), where closed-form expressions are derived for the E-step and M-step of the EM algorithm. Numerical results are provided to show the effectiveness of the proposed map construction method as compared to the benchmark curve fitting method with one single model. Peiming Li, Yong Zeng 0001, Jie Xu 0002 |
WCNC | 2 |
| 2022 | Distributed adaptive finite-time tracking for multi-agent systems and its application
Peiming Li, Xiangyong Chen, Jianlong Qiu |
Neurocomputing | 1 |
| 2021 | Asymmetric Interference Cancellation for 5G Non-Public Network with Uplink-Downlink Spectrum SharingabstractDifferent from public 4G/5G networks that are dominated by downlink (DL) traffic, emerging 5G non-public networks (NPNs) need to support significant uplink (UL) traffic to enable emerging applications such as industrial Internet of things (IIoT). The UL-DL spectrum sharing is becoming a viable solution to enhance the UL throughput of NPNs, which allows NPNs to perform the UL transmission over the time-frequency resources configured for DL transmission in coexisting public networks. To deal with the severe interference from the DL public base station (BS) transmitter to the coexisting UL non-public BS receiver, we propose an adaptive asymmetric successive interference cancellation (SIC) approach, in which the non-public BS is enabled to have the capability of decoding the DL signals transmitted from the public BS and cancelling them for interference mitigation. In particular, this paper studies a basic UL-DL spectrum sharing scenario when a UL non-public BS and a DL public BS coexist in the same area, each communicating with multiple users via orthogonal frequency-division multiple access (OFDMA). Under this setup, we aim to maximize the common UL throughput of all non-public users, under the condition that the DL throughput of each public user is above a certain threshold. The decision variables include the subcarrier allocation and user scheduling for both non-public and public BSs, the receiver mode of the non-public BS over subcarriers, as well as the rate and power control. Numerical results show that the proposed design significantly improves the common UL throughput as compared to benchmark schemes without such consideration. Peiming Li, Lifeng Xie, Jianping Yao, Jie Xu 0002, Shuguang Cui, Ping Zhang 0003 |
ICC | 1 |
| 2020 | Fundamental Rate Limits of UAV-Enabled Multiple Access Channel With Trajectory OptimizationabstractThis paper studies an unmanned aerial vehicle (UAV)-enabled multiple access channel (MAC), in which multiple ground users transmit individual messages to a mobile UAV in the sky. We consider a linear topology scenario, where these users locate in a straight line and the UAV flies at a fixed altitude above the line connecting them. Under this setup, we jointly optimize the one-dimensional (1D) UAV trajectory and wireless resource allocation to reveal the fundamental rate limits of the UAV-enabled MAC, under the users' individual maximum power constraints and the UAV's maximum flight speed constraints. First, we consider the capacity-achieving non-orthogonal multiple access (NOMA) transmission with successive interference cancellation (SIC) at the UAV receiver. In this case, we characterize the capacity region by maximizing the average sum-rate of all users subject to a set of rate profile constraints. To optimally solve this highly non-convex problem with infinitely many UAV location variables over time, we show that any speed-constrained UAV trajectory is equivalent to the combination of a maximum-speed flying trajectory and a speed-free trajectory, and accordingly transform the original speed-constrained trajectory optimization problem into a speed-free problem that is optimally solvable via the Lagrange dual decomposition. It is rigorously proved that the optimal 1D trajectory solution follows the successive hover-and-fly (SHF) structure, i.e., the UAV successively hovers above a number of optimized locations, and flies unidirectionally among them at the maximum speed. Next, we consider two orthogonal multiple access (OMA) transmission schemes, i.e., frequency-division multiple access (FDMA) and time-division multiple access (TDMA). We maximize the achievable rate regions in the two cases by jointly optimizing the 1D trajectory design and wireless resource (frequency/time) allocation. It is shown that the optimal trajectory solutions still follow the SHF structure but with different hovering locations for each scheme. Finally, numerical results show that the proposed optimal trajectory designs achieve considerable rate gains over other benchmark schemes, and the capacity region achieved by NOMA significantly outperforms the rate regions by FDMA and TDMA. Peiming Li, Jie Xu 0002 |
IEEE Trans. Wirel. Commun. | 1 |