Peiyi Wang

dblp:236/6569 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 7 first-author · 22 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ADHDPulse: A Smartphone-Based Approach to Monitor ADHD Symptom Improvement and ADHD Symptom Levels Prediction
Shweta Ware, Allison Baun, Caleb Kwakye, Ethan Swift, Sofia Dimotsi, Peiyi Wang, Nikoloz Gvelesiani, Laura E. Knouse
COMPSAC6
2026 Strain-Based Shape and 3-D Force Estimation for Rod-Driven Continuum Robots With Stretch Sensors
abstract
Soft robots' ability to safely navigate complex environments motivates the development of algorithms for accurate environmental interaction assessment, enabling greater autonomy. Specifically, strain-based shape and force estimation of continuum robots with embedded soft sensors poses an open challenge mainly owing to continuous softness, anisotropic deformation, and non-linear properties. Mathematical description of deformable soft bodies and accurate estimation of external forces are crucial for achieving controllable and intelligent behaviors of these robots. In this paper, a kinetostatic strain-based modeling for rod-driven soft robots (RDSR) with embedded stretch sensors is proposed, which incorporates local strains, actuation variables, and external interactions. The strain model enables full shape estimation of the robot and prediction of strain variations in soft bodies. Building on this, we develop a force estimator based on predicted and measured sensor and actuator lengths to evaluate 3D external forces, accounting for both orthogonal and tangential components relative to the backbone. Moreover, we introduce a methodology using a novel ellipsoid representation to handle tangential forces that may become insensitive in certain singular configurations. This estimator allows us to either disregard such forces when they do not influence deformation or estimate them when they become observable. Our simulations and experiments demonstrate how this approach can be used to analyze the robot's configuration and successfully estimate external forces. Finally, it is demonstrated that when the continuum arm follows trajectories with higher strain sensitivity, tangential force estimation is significantly improved.
Peiyi Wang, Daniel Feliú-Talegon, Zhexin Xie, Wenci Xin, Muhammad Sunny Nazeer, Cosimo Della Santina, Cecilia Laschi, Federico Renda
IEEE Trans. Robotics1
2025 Towards Harmonized Uncertainty Estimation for Large Language Models
abstract
To facilitate robust and trustworthy deployment of large language models (LLMs), it is essential to quantify the reliability of their generations through uncertainty estimation.While recent efforts have made significant advancements by leveraging the internal logic and linguistic features of LLMs to estimate uncertainty scores, our empirical analysis highlights the pitfalls of these methods to strike a harmonized estimation between indication, balance, and calibration, which hinders their broader capability for accurate uncertainty estimation.To address this challenge, we propose CUE (Corrector for Uncertainty Estimation): A straightforward yet effective method that employs a lightweight model trained on data aligned with the target LLM's performance to adjust uncertainty scores.Comprehensive experiments across diverse models and tasks demonstrate its effectiveness, which achieves consistent improvements of up to 60% over existing methods.Resources are available at https://github. com/O-L1RU1/Corrector4UE.
Rui Li 0094, Jing Long, Muge Qi, Heming Xia, Lei Sha, Peiyi Wang, Zhifang Sui
ACL (1)6
2025 CCAgent: Coordinating Collaborative Data Scaling for Operating System Agents via Web3
Liang Chen 0024, Haozhe Zhao, Yinzhen Huang, Tsekai Lin, Weichu Xie, Peiyi Wang, Runxin Xu, Ming Wu 0007, Baobao Chang
CIKM8
2025 VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
abstract
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench’s effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson’s r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io.
Lei Li 0039, Yuancheng Wei, Zhihui Xie 0002, Xuqing Yang, Yifan Song 0002, Peiyi Wang, Chenxin An, Tianyu Liu 0001, Sujian Li, Bill Y. Lin, Lingpeng Kong, Qi Liu 0049
CVPR6
2025 Origami-Inspired Soft Gripper with Tunable Constant Force Output
abstract
Soft robotic grippers gently and safely manipulate delicate objects due to their inherent adaptability and softness. Limited by insufficient stiffness and imprecise force control, conventional soft grippers are not suitable for applications that require stable grasping force. In this work, we propose a soft gripper that utilizes an origami-inspired structure to achieve tunable constant force output over a wide strain range. The geometry of each taper panel is established to provide necessary parameters such as protrusion distance, taper angle, and crease thickness required for 3D modeling and FEA analysis. Simulations and experiments show that by optimizing these parameters, our design can achieve a tunable constant force output. Moreover, the origami-inspired soft gripper dynamically adapts to different shapes while preventing excessive forces, with potential applications in logistics, manufacturing, and other industrial settings that require stable and adaptive operations.
Zhenwei Ni, Zhihang Qin, Ceng Zhang, Peiyi Wang, Cecilia Laschi
IROS6
2025 Instantly Learning Preference Alignment via In-context DPO
abstract
Feifan Song, Yuxuan Fan, Xin Zhang, Peiyi Wang, Houfeng Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Feifan Song 0001, Yuxuan Fan, Xin Zhang 0099, Peiyi Wang, Houfeng Wang
NAACL (Long Papers)4
2024 Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
abstract
Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes.However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains limited due to a scarcity of training datasets in scientific domains.To fill this gap, we introduce Multimodal ArXiv, consisting of ArXivCap and ArXivQA, for enhancing LVLMs scientific comprehension.ArXivCap is a figure-caption dataset comprising 6.4M images and 3.9M captions, sourced from 572K ArXiv papers spanning various scientific domains.Drawing from ArXivCap, we introduce ArXivQA, a questionanswering dataset generated by prompting GPT-4V based on scientific figures.ArXivQA greatly enhances open-sourced LVLMs' mathematical reasoning capabilities, achieving a 10.4% absolute accuracy gain on a multimodal mathematical reasoning benchmark.Furthermore, employing ArXivCap, we devise four vision-to-text tasks for benchmarking LVLMs.Evaluation results with state-of-the-art LVLMs underscore their struggle with the nuanced semantics of academic figures, while domainspecific training yields substantial performance gains.Our error analysis uncovers misinterpretations of visual context, recognition errors, and the production of overly simplified captions by current LVLMs, shedding light on future improvements.
Lei Li 0039, Yuqi Wang 0003, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, Qi Liu 0049
ACL (1)4
2024 Large Language Models are not Fair Evaluators
abstract
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peiyi Wang, Lei Li 0039, Liang Chen 0024, Zefan Cai, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu 0049, Tianyu Liu 0001, Zhifang Sui
ACL (1)1
2024 Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
abstract
Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peiyi Wang, Lei Li 0039, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li 0005, Deli Chen, Zhifang Sui
ACL (1)1
2024 Utilizing Local Hierarchy with Adversarial Training for Hierarchical Text Classification
abstract
Hierarchical text classification (HTC) is a challenging subtask of multi-label classification due to its complex taxonomic structure. Nearly all recent HTC works focus on how the labels are structured but ignore the sub-structure of ground-truth labels according to each input text which contains fruitful label co-occurrence information. In this work, we introduce this local hierarchy with an adversarial framework. We propose a HiAdv framework that can fit in nearly all HTC models and optimize them with the local hierarchy as auxiliary information. We test on two typical HTC models and find that HiAdv is effective in all scenarios and is adept at dealing with complex taxonomic hierarchies. Further experiments demonstrate that the promotion of our framework indeed comes from the local hierarchy and the local hierarchy is beneficial for rare classes which have insufficient training data.
Peiyi Wang, Houfeng Wang
LREC/COLING2
2024 AssistGUI: Task-Oriented PC Graphical User Interface Automation
abstract
Graphical User Interface (GUI) automation holds significant promise for assisting users with complex tasks, thereby boosting human productivity. Existing works leveraging Large Language Model (LLM) or LLM-based AI agents have shown capabilities in automating tasks on Android and Web platforms. However, these tasks are primarily aimed at simple device usage and entertainment operations. This paper presents a novel benchmark, Assistgui, to evaluate whether models are capable of manipulating the mouse and keyboard on the Windows platform in response to user-requested tasks. We carefully collected a set of 100 tasks from nine widely-used software applications, such as, After Effects and MS Word, each accompanied by the necessary project files for better evaluation. Moreover, we propose a multi-agent collaboration framework, which incorporates four agents to perform task decomposition, GUI parsing, action generation, and reflection. Our experimental results reveal that our multi-agent collaboration mechanism outshines existing methods in performance. Nevertheless, the potential remains substantial, with the best model attaining only a 46% success rate on our benchmark. We conclude with a thorough analysis of the current methods' limitations, setting the stage for future breakthroughs in this domain.
Difei Gao, Lei Ji 0001, Zechen Bai, Mingyu Ouyang, Dongxing Mao, Qinchen Wu, Peiyi Wang, Xiangwu Guo, Hengxu Wang, Luowei Zhou, Zheng Shou 0001
CVPR9
2024 VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
abstract
Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lei Li 0039, Zhihui Xie 0002, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen 0024, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu 0049
EMNLP5
2024 Strain-based Modeling of Rod-driven Soft Continuum Robots with Co-located Embedded Sensors
abstract
Rod-driven soft robots (RDSR) with a well-balanced performance in terms of perception, precision, and intelligence have a great potential for application. Mathematical description and predicted sensing of deformable soft bodies are crucial to achieve controllable and intelligent behaviors of these robots. In this work, we propose a kinetostatic model for RDSR embedded with co-located sensors based on the Geometric Variable Strain (GVS) approach where local deformations, actuation lengths and external interactions are included. This approach allows us to estimate the shape of RDSR and predict the strain variation of soft bodies under internal and external interactions. Simulations and experimental results show that tip position errors are not greater than 1.8% with respect to the whole body length under different loads (0, 100, 200, 300 gf). The maximum error of predicted sensor length change is up to 2 mm and its percentage relative to the actual length does not exceed 4%. The results demonstrate the accuracy and effectiveness of the proposed model.
Peiyi Wang, Daniel Feliú-Talegon, Sheng Guo 0001, Federico Renda, Cecilia Laschi
IROS1
2023 Rationale-Enhanced Language Models are Better Continual Relation Learners
abstract
Continual relation extraction (CRE) aims to solve the problem of catastrophic forgetting when learning a sequence of newly emerging relations.Recent CRE studies have found that catastrophic forgetting arises from the model's lack of robustness against future analogous relations.To address the issue, we introduce rationale, i.e., the explanations of relation classification results generated by large language models (LLM), into CRE task.Specifically, we design the multi-task rationale tuning strategy to help the model learn current relations robustly.We also conduct contrastive rationale replay to further distinguish analogous relations.Experimental results on two standard benchmarks demonstrate that our method outperforms the state-of-the-art CRE models.Our code is available at https://github.com/WeiminXiong/ RationaleCL
Weimin Xiong, Yifan Song 0002, Peiyi Wang, Sujian Li
EMNLP3
2023 Meta-Learning-Based Optimal Control for Soft Robotic Manipulators to Interact with Unknown Environments
abstract
Safe and efficient robot-environment interaction is a critical but challenging problem as robots are being increasingly employed to operate in unstructured and unpredictable environments. Soft robots are inherently compliant to safely interact with environments but their high nonlinearity exacerbates control difficulties. Meta-learning provides a powerful tool for fast online model adaptation because it can learn an efficient model from data across different environments. Thus, this work applies the idea of meta-learning for the control of soft robotics. In particular, a target-oriented proactive search strategy is firstly performed to collect environment-specific data efficiently when a new interaction environment occurs. Then meta-learning exploits past experience to train a data-driven probabilistic model prior, and the model prior is online updated to be fast adapted to the new environment. Lastly, a model-based optimal control policy is utilized to drive the robot to desired performance. Our approach controls a soft robotic manipulator to achieve the desired position and contact force simultaneously when interacting with unknown changing environments. Overall, this work provides a viable control approach for soft robots to interact with unknown environments.
Peiyi Wang, Wenci Xin, Zhexin Xie, Longxin Kan, Muralidharan Mohanakrishnan, Cecilia Laschi
ICRA2
2022 Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification
abstract
Hierarchical text classification is a challenging subtask of multi-label classification due to its complex label hierarchy.Existing methods encode text and label hierarchy separately and mix their representations for classification, where the hierarchy remains unchanged for all input text.Instead of modeling them separately, in this work, we propose Hierarchyguided Contrastive Learning (HGCLR) to directly embed the hierarchy into a text encoder.During training, HGCLR constructs positive samples for input text under the guidance of the label hierarchy.By pulling together the input text and its positive sample, the text encoder can learn to generate the hierarchy-aware text representation independently.Therefore, after training, the HGCLR enhanced text encoder can dispense with the redundant hierarchy.Extensive experiments on three benchmark datasets verify the effectiveness of HGCLR.
Peiyi Wang, Lianzhe Huang, Xin Sun 0013, Houfeng Wang
ACL (1)2
2022 Learning Robust Representations for Continual Relation Extraction via Adversarial Class Augmentation
abstract
Continual relation extraction (CRE) aims to continually learn new relations from a classincremental data stream.CRE model usually suffers from catastrophic forgetting problem, i.e., the performance of old relations seriously degrades when the model learns new relations.Most previous work attributes catastrophic forgetting to the corruption of the learned representations as new relations come, with an implicit assumption that the CRE models have adequately learned the old relations.In this paper, through empirical studies we argue that this assumption may not hold, and an important reason for catastrophic forgetting is that the learned representations do not have good robustness against the appearance of analogous relations in the subsequent learning process.To address this issue, we encourage the model to learn more precise and robust representations through a simple yet effective adversarial class augmentation mechanism (ACA), which is easy to implement and model-agnostic.Experimental results show that ACA can consistently improve the performance of state-of-theart CRE models on two popular benchmarks.
Peiyi Wang, Yifan Song 0002, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Sujian Li, Zhifang Sui
EMNLP1
2022 HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification
abstract
Hierarchical text classification (HTC) is a challenging subtask of multi-label classification due to its complex label hierarchy.Recently, the pretrained language models (PLM) have been widely adopted in HTC through a finetuning paradigm.However, in this paradigm, there exists a huge gap between the classification tasks with sophisticated label hierarchy and the masked language model (MLM) pretraining tasks of PLMs and thus the potential of PLMs cannot be fully tapped.To bridge the gap, in this paper, we propose HPT, a Hierarchy-aware Prompt Tuning method to handle HTC from a multi-label MLM perspective.Specifically, we construct a dynamic virtual template and label words that take the form of soft prompts to fuse the label hierarchy knowledge and introduce a zero-bounded multi-label cross-entropy loss to harmonize the objectives of HTC and MLM.Extensive experiments show HPT achieves state-of-the-art performances on 3 popular HTC datasets and is adept at handling the imbalance and low resource situations.
Peiyi Wang, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui, Houfeng Wang
EMNLP2
2022 Few-Shot Learning with Self-supervised Classifier for Complex Knowledge Base Question Answering
Peiyi Wang
KSEM (2)3
2022 An Enhanced Span-based Decomposition Method for Few-Shot Sequence Labeling
abstract
Peiyi Wang, Runxin Xu, Tianyu Liu, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui
NAACL-HLT1
2022 A Two-Stream AMR-enhanced Model for Document-level Event Argument Extraction
abstract
Runxin Xu, Peiyi Wang, Tianyu Liu, Shuang Zeng, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Runxin Xu, Peiyi Wang, Tianyu Liu 0001, Shuang Zeng, Baobao Chang, Zhifang Sui
NAACL-HLT2
2021 Behind the Scenes: An Exploration of Trigger Biases Problem in Few-Shot Event Classification
abstract
Few-Shot Event Classification (FSEC) aims at developing a model for event prediction, which can generalize to new event types with a limited number of annotated data. Existing FSEC studies have achieved high accuracy on different benchmarks. However, we find they suffer from trigger biases that signify the statistical homogeneity between some trigger words and target event types, which we summarize as trigger overlapping and trigger separability. The biases can result in context-bypassing problem, i.e., correct classifications can be gained by looking at only the trigger words while ignoring the entire context. Therefore, existing models can be weak in generalizing to unseen data in real scenarios. To further uncover the trigger biases and assess the generalization ability of the models, we propose two new sampling methods, Trigger-Uniform Sampling (TUS) and COnfusion Sampling (COS), for the meta tasks construction during evaluation. Besides, to cope with the context-bypassing problem in FSEC models, we introduce adversarial training and trigger reconstruction techniques. Experiments show these techniques help not only improve the performance, but also enhance the generalization ability of models.
Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Damai Dai, Baobao Chang, Zhifang Sui
CIKM1
2021 Biometric key generation based on generated intervals and two-layer error correcting technique
Peiyi Wang, Lin You, Gengran Hu, Liqin Hu, Zhihua Jian, Chaoping Xing
Pattern Recognit.1
2019 NRSA: Neural Recommendation with Summary-Aware Attention
Qiyao Peng 0001, Peiyi Wang, Wenjun Wang 0002, Hongtao Liu 0008, Yueheng Sun, Pengfei Jiao
KSEM (1)2
2019 REET: Joint Relation Extraction and Entity Typing via Multi-task Learning
Hongtao Liu 0008, Peiyi Wang, Fangzhao Wu, Pengfei Jiao, Wenjun Wang 0002, Xing Xie 0001, Yueheng Sun
NLPCC (1)2