Hailiang Jin

dblp:79/7839 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Software testing · 94% Program synthesis and code generation · 6%
Artificial intelligence
1 paper
Vision and language · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software testing
mobile application testing
1.722025
From Redundancy to Efficiency: Exploiting Shared UI Interactions towards Efficient LLM-Based Testing · ASE 2025
Minuku: Detecting Diverse Display Issues in Mobile Apps with Small-scale Dataset · ASE 2025
Software testing › UI testing
UI display issue detection
0.912025
Minuku: Detecting Diverse Display Issues in Mobile Apps with Small-scale Dataset · ASE 2025
Software testing
UI testing
0.912025
From Redundancy to Efficiency: Exploiting Shared UI Interactions towards Efficient LLM-Based Testing · ASE 2025
Computer vision › Vision and language › vision-language model › vision-language model adaptation
vision-language model fine-tuning
0.312025
Minuku: Detecting Diverse Display Issues in Mobile Apps with Small-scale Dataset · ASE 2025
Program synthesis and code generation
code generation with language models
0.312025
From Redundancy to Efficiency: Exploiting Shared UI Interactions towards Efficient LLM-Based Testing · ASE 2025

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.7fine-tuning · 1.7commonsense simulation · 1.7large language model · 0.9UI transition graph · 0.9
YearPublicationVenuePosition
2025 Minuku: Detecting Diverse Display Issues in Mobile Apps with Small-scale Dataset
abstract
User interface (UI) display issues, such as widgets occlusion, missing elements, and screen overflow, are emerging as a non-negligible source of user complaints in commercial mobile apps. However, existing automated testing tools typically rely on a vast amount of high-quality training data, making them cost-ineffective for industrial practice. Given that display issues are intuitively recognizable by humans, their diverse appearances can be abstracted by the violation of human commonsense expectations of UI appearance. Therefore, this paper proposes to reduce data requirements in display issue detection through commonsense simulation. Although leveraging large vision-language models (VLMs) to replicate human visual ability looks straightforward, off-the-shelf VLMs lack task-specific knowledge of UI designs and display correctness. To address this, we fine-tune a VLM to learn what constitutes an expected display and to reason potential display issues. This approach is termed as Minuku, an industrial data-efficient UI display issue detector. We evaluate the design effectiveness of Minuku via a set of ablation experiments. Moreover, real-world deployments in one of the largest E-commerce app providers further demonstrate that Minuku can effectively detect 40 previously unknown UI display issues and significantly reduce manual effort in industrial settings.
Yongxiang Hu 0003, Hailiang Jin, Juxing Yuan, Xin Wang 0002, Yangfan Zhou 0002
ASE3
2025 From Redundancy to Efficiency: Exploiting Shared UI Interactions towards Efficient LLM-Based Testing
abstract
Redundant test cases, although well-studied in software engineering, are previously underexplored in UI testing of mobile apps. Our study of real-world test suites shows that, equipped with large-scale testing suites, redundancy in UI testing often manifests as redundant UI interactions. Although negligible in traditional script-based workflows, such redundancy severely impacts the efficiency of emerging Large Language Model (LLM)-based UI agents, which incur substantial decision latency and token costs from repeated LLM queries for the same interactions. To this end, based on the idea of reusing LLMs’ former decisions, we present TestWeaver, a cost-effective LLM-based testing framework. Leveraging a semantic annotated UI Transition Graph (UTG), TestWeaver is capable of detecting shared interactions across test cases. It processes each interaction with a single LLM query, and reuses the result whenever the same interaction occurs. We evaluate TestWeaver on real-world test suites from Meituan. It achieves a 92% success rate with an average cost of $0.11 and 89.7 seconds per case, outperforming the state-of-the-art. We have also deployed TestWeaver in a real-world testing workflow at Meituan for over six months. TestWeaver has executed nearly 2,000 test cases and uncovered 10 previously undetected bugs, while reducing manual testing effort by 75%.
Yingchuan Wang, Yongxiang Hu 0003, Yu Zhang 0165, Hailiang Jin, Juxing Yuan, Yangfan Zhou 0002
ASE5