Fei Tang 0005

dblp:60/406-5 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0008-0019-1034ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 52% Vision and language · 28% Planning, search and constraint satisfaction · 13%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
GUI automation
1.012026
UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization · ACL (1) 2026
Computer vision › Vision and language › visual grounding › instruction grounding
GUI grounding
1.012026
Test-Time Reinforcement Learning for GUI Grounding via Region Consistency · AAAI 2026
Machine learning › Reinforcement learning
multi-turn reinforcement learning
1.012026
Experience-driven Multi-turn Reinforcement Learning for GUI Agents · ACL (1) 2026
Machine learning › Reinforcement learning
policy optimization
1.012026
UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization · ACL (1) 2026
Machine learning › Reinforcement learning › reward learning
reward modeling
1.012026
GUI-G²: Gaussian Reward Modeling for GUI Grounding · AAAI 2026
Machine learning › Reinforcement learning › online decision making › online reinforcement learning
test-time reinforcement learning
1.012026
Test-Time Reinforcement Learning for GUI Grounding via Region Consistency · AAAI 2026
Human-AI interaction › GUI agent
GUI grounding
1.012026
GUI-G²: Gaussian Reward Modeling for GUI Grounding · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation · ACM Multimedia 2025
Computing education
large language model evaluation
0.912025
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation · ACM Multimedia 2025
Visual content generation and editing
vector graphics generation
0.912025
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation · ACM Multimedia 2025
Machine learning › Efficient and distributed learning
data-efficient learning
0.312026
Test-Time Reinforcement Learning for GUI Grounding via Region Consistency · AAAI 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
GUI agent
0.312026
Experience-driven Multi-turn Reinforcement Learning for GUI Agents · ACL (1) 2026
Natural language and speech › Language models and text generation
LLM agents
0.312026
UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 3.0gaussian distribution modeling · 2.0tool-integrated policy optimization · 1.0test-time scaling · 1.0spatial voting · 1.0policy optimization · 1.0experience-driven reinforcement learning · 1.0
YearPublicationVenuePosition
2026 Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
abstract
Graphical User Interface (GUI) grounding, the task of mapping natural language instructions to precise screen coordinates, is fundamental to autonomous GUI agents. While existing methods achieve strong performance through extensive supervised training or reinforcement learning with labeled rewards, they remain constrained by the cost and availability of pixel-level annotations. We observe that when models generate multiple predictions for the same GUI element, the spatial overlap patterns reveal implicit confidence signals that can guide more accurate localization. Leveraging this insight, GUI-RC (Region Consistency), a test-time scaling method that constructs spatial voting grids from multiple sampled predictions to identify consensus regions where models show highest agreement. Without any training, GUI-RC improves accuracy by 2-3% across various architectures on ScreenSpot benchmarks. We further introduce GUI-RCPO (Region Consistency Policy Optimization), transforming these consistency patterns into rewards for test-time reinforcement learning. By computing how well each prediction aligns with the collective consensus, GUI-RCPO enables models to iteratively refine their outputs on unlabeled data during inference. Extensive experiments demonstrate the generality of our approach: using only 1,272 unlabeled data, GUI-RCPO achieves 3-6% accuracy improvements across various architectures on ScreenSpot benchmarks. Our approach reveals the untapped potential of test-time scaling and test-time reinforcement learning for GUI grounding, offering a promising path toward more data-efficient GUI agents.
Fei Tang 0005, Zhengxi Lu, Chang Zong, Weiming Lu 0001, Shengpei Jiang, Yongliang Shen 0001
AAAI3
2026 GUI-G²: Gaussian Reward Modeling for GUI Grounding
abstract
Graphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of spatial interactions. Motivated by human clicking behavior that naturally forms Gaussian distributions centered on target elements, we introduce GUI Gaussian Grounding Rewards (GUI-G2), a principled reward framework that models GUI elements as continuous Gaussian distributions across the interface plane. GUI-G2 incorporates two synergistic mechanisms: Gaussian point rewards model precise localization through exponentially decaying distributions centered on element centroids, while coverage rewards assess spatial alignment by measuring the overlap between predicted Gaussian distributions and target regions. To handle diverse element scales, we develop an adaptive variance mechanism that calibrates reward distributions based on element dimensions. This framework transforms GUI grounding from sparse binary classification to dense continuous optimization, where Gaussian distributions generate rich gradient signals that guide models toward optimal interaction positions. Extensive experiments across ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro benchmarks demonstrate that GUI-G2, substantially outperforms state-of-the-art method UI-TARS-72B, with the most significant improvement of 24.7% on ScreenSpot-Pro. Our analysis reveals that continuous modeling provides superior robustness to interface variations and enhanced generalization to unseen layouts, establishing a new paradigm for spatial reasoning in GUI interaction tasks.
Fei Tang 0005, Zhangxuan Gu, Zhengxi Lu, Shuheng Shen, Changhua Meng, Wen Wang 0009, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
AAAI1
2026 UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization
abstract
Zhengxi Lu, Fei Tang, Guangyi Liu, Jin Ma, Kaitao Song, Xu Tan, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengxi Lu, Fei Tang 0005, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang, Yongliang Shen 0001
ACL (1)2
2026 Experience-driven Multi-turn Reinforcement Learning for GUI Agents
abstract
Zhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen, Haiyang Xu, Ziwei Zheng, Weiming Lu, Ming Yan, Fei Huang, Jun Xiao, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengxi Lu, Jiabo Ye, Fei Tang 0005, Yongliang Shen 0001, Haiyang Xu 0001, Ziwei Zheng, Weiming Lu 0001, Ming Yan 0008, Fei Huang 0002, Jun Xiao 0001, Yueting Zhuang
ACL (1)3
2025 SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
abstract
Large Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of complexity stratification, and fragmented evaluation paradigms. We introduce SVGenius, a comprehensive benchmark comprising 2,377 queries across three progressive dimensions: understanding, editing, and generation. Built on real-world data from 24 application domains with systematic complexity stratification, SVGenius evaluates models through 8 task categories and 18 metrics. We assess 22 mainstream models spanning different scales, architectures, training paradigms, and accessibility levels. Our analysis reveals that while proprietary models significantly outperform open-source counterparts, all models exhibit systematic performance degradation with increasing complexity, indicating fundamental limitations in current approaches; however, reasoning-enhanced training proves more effective than pure scaling for overcoming these limitations, though style transfer remains the most challenging capability across all model types. SVGenius establishes the first systematic evaluation framework for SVG processing, providing crucial insights for developing more capable vector graphics models and advancing automated graphic design applications. Appendix and supplementary materials (including all data and code) are available at https://zju-real.github.io/SVGenius.
Haolei Xu, Fei Tang 0005, Linjuan Wu, Wenqi Zhang 0001, Guiyang Hou, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang
ACM Multimedia5