Yuzhe Guo

dblp:388/1679 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0000-7533-1478ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Image recognition and object detection · 70% Information extraction and text analysis · 30%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
0.912025
Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans? · CVPR 2025
Computer vision › Image recognition and object detection › object detection
prohibited item detection
0.912025
Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans? · CVPR 2025
Software testing
GUI testing
0.912025
Element-Aware Fine-Tuning of Vision-Language Models for Cost-Efficient GUI Testing in an Industrial Setting · ASE 2025
Natural language and speech › Information extraction and text analysis
emotion recognition
0.812024
Multi-level Disentangling Network for Cross-Subject Emotion Recognition Based on Multimodal Physiological Signals · IJCAI 2024
Software testing
mobile application testing
0.312025
Element-Aware Fine-Tuning of Vision-Language Models for Cost-Efficient GUI Testing in an Industrial Setting · ASE 2025
Medical and health informatics › biomedical signal processing
physiological signal analysis
0.212024
Multi-level Disentangling Network for Cross-Subject Emotion Recognition Based on Multimodal Physiological Signals · IJCAI 2024

Methods — techniques the papers use, named apart from their topics

multimodal physiological signals · 1.5disentangling network · 1.5vision-language model · 0.9omniparser · 0.9fine-tuning · 0.9auxiliary-view enhanced network · 0.9UI element detection · 0.9
YearPublicationVenuePosition
2026 From User Operations to Agentic Automation: Toward Intent-Oriented Software in the LLM Era
Tao Xie 0001, Dezhi Ran, Mengzhou Wu, Yuzhe Guo, Wei Yang 0013
J. Comput. Sci. Technol.5
2025 Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans?
abstract
To detect prohibited items in challenging categories, human inspectors typically rely on images from two distinct views (vertical and side). Can AI detect prohibited items from dual-view X-ray images in the same way humans do? Existing X-ray datasets often suffer from limitations, such as single-view imaging or insufficient sample diversity. To address these gaps, we introduce the Large-scale Dual-view X-ray (LDXray), which consists of 353,646 instances across 12 categories, providing a diverse and comprehensive resource for training and evaluating models. To emulate human intelligence in dual-view detection, we propose the Auxiliary-view Enhanced Network (AENet), a novel detection framework that leverages both the main and auxiliary views of the same object. The main-view pipeline focuses on detecting common categories, while the auxiliary-view pipeline handles more challenging categories using "expert models" learned from the main view. Extensive experiments on the LDXray dataset demonstrate that the dual-view mechanism significantly enhances detection performance, e.g., achieving improvements of up to +24.7% for the challenging category of umbrellas. Furthermore, our results show that AENet exhibits strong generalization across seven different detection models for X-ray Inspection1.
Renshuai Tao, Yuzhe Guo, Hairong Chen, Li Zhang 0023, Xianglong Liu 0001, Yunchao Wei, Yao Zhao 0001
CVPR3
2025 SECC-Stega: Generative Linguistic Steganographic Framework Based on Error Correcting Codes
abstract
With the rise and maturation of neural network technology, generative text steganography based on language models is gradually becoming the mainstream technique in text steganography. However, homomorphic extraction attacks and text modification attacks from third parties pose serious threats to the usability of generative text steganography. To address this issue, this paper proposes a generative text steganography algorithm framework based on error correction codes. This framework enhances the robustness and security of steganography by encoding the secret information. Experimental results verify that the proposed framework achieves the expected outcomes and exhibits a certain degree of generality.
Yuzhe Guo, Zhongliang Yang, Zhili Zhou 0001, Linna Zhou
ICASSP1
2025 Element-Aware Fine-Tuning of Vision-Language Models for Cost-Efficient GUI Testing in an Industrial Setting
abstract
User Interface (UI) testing is crucial for quality assurance of industrial mobile applications, and yet it remains labor-intensive and challenging to automate effectively. Recent advances in Vision-Language Models (VLMs) present a promising solution for automating GUI testing by mapping natural language instructions to pixel-level actions, significantly reducing the manual effort required for writing test scripts and even designing test cases. While numerous VLMs have been proposed and evaluated for GUI testing, they often fail to meet two critical industrial requirements: (1) effectiveness when handling complex, multi-step workflows in industrial applications, and (2) efficiency for large-scale, high-frequency testing environments typical in industrial settings. Toward addressing the preceding industrial requirements, in this paper, we report our experiences in developing and deploying RePeek, a novel approach employing a unified three-stage pipeline for both training and inference, enables a VLM to explicitly detect and reason over discrete GUI elements, thereby overcoming the limitations of pixel-based reasoning for both efficiency and effectiveness improvements. In the first stage, RePeek integrates a lightweight UI-element detector named OmniParser to decompose UI screenshots into a structured element list. In the second stage, RePeek adopts the vision encoder of the VLM to generate the embedding for each element. In the third stage, RePeek fuses these element embeddings with the textual instruction to reason and perform classification directly on the UI elements, empowering efficient small models to achieve superior performance against expensive large models. Comprehensive evaluations on public benchmarks and deployment at WeChat show that RePeek consistently achieves superior accuracy and efficiency compared to state-of-the-art VLMs. Specifically, RePeek enables a fine-tuned Qwen2.5-VL-3B model to outperform a 72B model with 75% less training data, validating the effectiveness of incorporating domain knowledge into VLM-based GUI testing. We conclude by summarizing three key lessons from developing and deploying RePeek, offering insights for both researchers and practitioners working on industrial-strength UI testing.
Mengzhou Wu, Yuzhe Guo, Haochuan Lu, Xia Zeng, Liangchao Yao, Yuetang Deng, Dezhi Ran, Wei Yang 0013, Tao Xie 0001
ASE2
2025 LS-PRISM: A layer-selective pruning method via low-rank approximation and sparsification for efficient large language model compression
Renshuai Tao, Hairong Chen, Yuzhe Guo, Jiakai Wang, Boying Wang, Yao Zhao 0001
Neural Networks3
2024 Multi-level Disentangling Network for Cross-Subject Emotion Recognition Based on Multimodal Physiological Signals
Ziyu Jia, Fengming Zhao, Yuzhe Guo, Hairong Chen, Tianzi Jiang
IJCAI3