Pengzhi Gao

dblp:157/1201 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0009-9392-6657ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mobile GUI Agents under Real-world Threats: Are We There Yet?
abstract
Recent years have witnessed a rapid development of mobile GUI agents powered by large language models (LLMs), which can autonomously execute diverse device-control tasks based on natural language instructions. The increasing accuracy of these agents on standard benchmarks has raised expectations for large-scale real-world deployment, and there are already several commercial agents released and used by early adopters. However, are we really ready for GUI agents integrated into our daily devices as system building blocks? We argue that an important pre-deployment validation is missing to examine whether the agents can maintain their performance under real-world threats. Specifically, unlike existing common benchmarks that are based on simple static app contents (they have to do so to ensure environment consistency between different tests), real-world apps are filled with contents from untrustworthy third parties, such as advertisement emails, user-generated posts and medias, etc. These contents may inevitably appear in the agents' observation space and influence the task execution process. Systematic investigation of this problem is challenging since the real-world app contents are significantly skewed—testing on normal real-world apps usually cannot uncover any potential risk since most app contents are benign. To this end, we introduce a scalable app content instrumentation framework to enable flexible and targeted content modifications within existing applications. Leveraging this framework, we create a test suite comprising both a dynamic task execution environment and a static dataset of challenging GUI states. The dynamic environment encompasses 122 reproducible tasks, and the static dataset consists of over 3,000 scenarios constructed from commercial apps. We perform experiments on both open-source and commercial GUI agents. Our findings reveal that all examined agents can be significantly degraded due to third-party contents, with an average misleading rate of 42.0% and 36.1% in dynamic and static environments respectively. The framework and benchmark has been released at https://agenthazard.github.io.
Guohong Liu 0002, Jialei Ye, Wei Liu 0302, Pengzhi Gao, Jian Luan 0001, Yuanchun Li 0003, Yunxin Liu 0001
MobiSys5
2025 Doubly Constrained Fair Clustering for General p-Norms
abstract
Fairness in clustering has received significant attention. Dickerson et al. in 2023 first proposed the doubly constrained fair clustering problem that aims two fairness constraints, namely, (1) the Group Fairness (GF), which requires that different groups within each cluster have a certain degree of representation, and (2) the Diversity in Center Selection fairness (DS), which requires that the selected centers represent a diverse range of different groups. However, their algorithm only focuses on the k-center objective. In this paper, we generalize the doubly constrained fair clustering to $$\ell _p$$ norm objectives with general p, thus including k-Center, k-Median, and k-Means as special cases. We propose the first approximation algorithm for the doubly constrained fair clustering problem with general p-norms. In polynomial time, our algorithm finds an $$O(\Delta ^{\frac{1}{p}})$$ -approximate clustering that violates the GF constraint by an additive factor of 5 and satisfies the DS constraint, where $$\Delta $$ is the largest size of clusters in the solution. Our main contribution is a novel method to select centers using the min cost network flow approach. Finally, we conduct experiments to validate our algorithm. The experimental results show that the clustering cost of our algorithm, while simultaneously considering both of the GF and DS constraints, is nearly identical to that of the clustering algorithm which only considers the GF constraint.
Lunhao Zhang, Pengzhi Gao, Peng Zhang 0008
COCOON (1)2
2025 BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
abstract
Graphical User Interface (GUI) agents have gained substantial attention due to their impressive capabilities to complete tasks through multiple interactions within GUI environments.However, existing agents primarily focus on enhancing the accuracy of individual actions and often lack effective mechanisms for detecting and recovering from errors.To address these shortcomings, we propose the BacktrackAgent, a robust framework that incorporates a backtracking mechanism to improve task completion efficiency.BacktrackAgent includes verifier, judger, and reflector components as modules for error detection and recovery, while also applying judgment rewards to further enhance the agent's performance.Additionally, we develop a training dataset specifically designed for the backtracking mechanism, which considers the outcome pages after action executions.Experimental results show that BacktrackAgent has achieved performance improvements in both task success rate and step accuracy on Mobile3M and Auto-UI benchmarks.Our data and code will be released upon acceptance.
Qinzhuo Wu, Pengzhi Gao, Wei Liu 0302, Jian Luan 0001
EMNLP2
2025 Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
abstract
Menglong Cui, Pengzhi Gao, Wei Liu, Jian Luan, Bin Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Menglong Cui, Pengzhi Gao, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004
NAACL (Long Papers)2
2025 A Simple Heuristic Finding Connectivity Bottleneck in Networks with Shared Risk Resource Groups
Yizhe Tong, Pengzhi Gao, Peng Zhang 0008
WASA (2)2
2024 LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
abstract
Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Sanjay Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip S. Yu, Wenpeng Yin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Jiangshu Du, Yibo Wang 0001, Wenting Zhao 0006, Zhongfen Deng, Shuaiqi Liu 0002, Renze Lou, Henry Peng Zou, Pranav Venkit, Mukund Srinath, Ranran Haoran Zhang, Tao Li 0039, Fei Wang 0060, Qin Liu 0010, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang 0003, Raj Sanjay Shah, Ruohao Guo, Haoran Li 0003, Kangda Wei, Zihao Wang 0001, Lu Cheng 0001, Surangika Ranathunga, Fei Liu 0004, Ruihong Huang, Eduardo Blanco 0002, Yixin Cao 0002, Rui Zhang 0037, Philip S. Yu, Wenpeng Yin 0001
EMNLP18
2024 An Empirical Study of Consistency Regularization for End-to-End Speech-to-Text Translation
abstract
Pengzhi Gao, Ruiqing Zhang, Zhongjun He, Hua Wu, Haifeng Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Pengzhi Gao, Ruiqing Zhang, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001
NAACL-HLT1
2022 Multimodal Emotion Recognition Using CNN-SVM with Data Augmentation
abstract
With the development of human-computer interaction and mobile sensors, emotion recognition based on physiological signals has aroused a lively discussion among scholars. The main difficulty faced is the small amount of data which leads to poor training results. In this paper, we proposed a multimodal emotion recognition using CNN-SVM and data augmentation (CSDAMER). Electrocardiography (ECG), galvanic skin response (GSR) and respiration (RSP) are utilized as input data, which are less requiring on the collection environment and can be collected by mobile sensors. To improve the training effect of model, data augmentation is performed by transformations, such as inversion, recombination and noise injection. Moreover, the convolutional layer of the convolutional neural network (CNN) is leveraged to extract the high-level features of the physiological signals, and then the features are input into the support vector machine (SVM) classifier to obtain the recognition results. The experimental results show that CSDAMER achieves 80.7% and 79.92% accuracy in arousal and valance, respectively. Compared with CNN alone, the accuracy of arousal and valance is increased by 12.87% and 9.95%. Meanwhile, the addition of the data augmentation improves the accuracy in arousal and valance by 21.94% and 25.73%.
Gengyuan Guo, Pengzhi Gao, Xiangwei Zheng 0001, Cun Ji
BIBM2
2022 Bi-SimCut: A Simple Strategy for Boosting Neural Machine Translation
abstract
We introduce Bi-SimCut: a simple but effective training strategy to boost neural machine translation (NMT) performance.It consists of two procedures: bidirectional pretraining and unidirectional finetuning.Both procedures utilize SimCut, a simple regularization method that forces the consistency between the output distributions of the original and the cutoff sentence pairs.Without leveraging extra dataset via back-translation or integrating large-scale pretrained model, Bi-SimCut achieves strong translation performance across five translation benchmarks (data sizes range from 160K to 20.2M): BLEU scores of 31.16 for en → de and 38.37 for de → en on the IWSLT14 dataset, 30.78 for en → de and 35.15 for de → en on the WMT14 dataset, and 27.17 for zh → en on the WMT17 dataset.Sim-Cut is not a new method, but a version of Cutoff (Shen et al., 2020) simplified and adapted for NMT, and it could be considered as a perturbation-based method.Given the universality and simplicity of SimCut and Bi-SimCut, we believe they can serve as strong baselines for future NMT research.
Pengzhi Gao, Zhongjun He, Hua Wu 0003, Haifeng Wang 0001
NAACL-HLT1
2018 Dynamic Matrix Recovery from Partially Observed and Erroneous Measurements
abstract
This paper studies the low-rank matrix recovery problem from partially lost and partially corrupted measurements. It shows both analytically and numerically that the recovery performance can be greatly enhanced if one further exploits the temporal correlations among a sequence of low-rank matrices. The matrix recovery problem is formulated as a non-convex optimization problem, and the recovery error is quantified analytically. A fast iterative algorithm is proposed to solve the non-convex problem, and every sequence generated by the algorithm converges to a critical point of the optimization problem. The method is numerically evaluated on the synthetic datasets.
Pengzhi Gao, Meng Wang 0003
ICASSP1