Biniam Fisseha Demissie

dblp:166/5793 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
2since 2021 · last 2026
0000-0002-5369-5235ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 VLM-Fuzz: Vision language model assisted recursive depth-first search exploration for effective GUI testing of android apps
abstract
Abstract Testing Android apps effectively requires a systematic exploration of the app’s possible states by simulating user interactions and system events. While existing approaches have proposed several fuzzing techniques to generate various text inputs and trigger user and system events for GUI state exploration, achieving high code coverage remains a significant challenge in Android app testing. The main challenges are (1) reasoning about the complex and dynamic layout of GUI screens; (2) generating required inputs/events to deal with certain widgets like pop-ups; and (3) coordination between current test inputs and previous inputs to avoid getting stuck in the same GUI screen without improving test coverage. To address these problems, we propose VLM-Fuzz , a novel automated approach for Android GUI testing. At its foundation, VLM-Fuzz utilizes a heuristic-based, recursive depth-first search (DFS) strategy that is intelligently guided by a Vision Language Model (VLM) to effectively explore the app’s complex GUI states. The core innovation of VLM-Fuzz is not simply the use of a VLM, but its strategic, on-demand integration within a hybrid exploration framework. Our approach combines a fast, heuristic-based DFS for standard GUI interactions with targeted, VLM-assisted analysis for visually complex screens. We use static analysis to analyze the Android Manifest file and the runtime GUI hierarchy XML to extract the list of components, intent-filters and interactive GUI widgets. VLM is used to reason about complex GUI layout and widgets on an on-demand basis. Based on the inputs from static analysis, VLM, and the current GUI state, we use some heuristics to deal with the above-mentioned challenges. We evaluated VLM-Fuzz based on a benchmark containing 59 apps obtained from a recent work and compared it against two state-of-the-art approaches: APE and DeepGUI . VLM-Fuzz outperforms the best baseline by 9.0% , 3.7% , and 2.1% in terms of class coverage, method coverage, and line coverage, respectively. We also ran VLM-Fuzz on 80 recent Google Play apps (i.e., updated in 2024). VLM-Fuzz detected 52 unique crashes in 12 apps, which have been reported to respective developers.
Biniam Fisseha Demissie, Yan Naing Tun, Lwin Khin Shar, Mariano Ceccato
Empir. Softw. Eng.1
2023 Experimental comparison of features, analyses, and classifiers for Android malware detection
Lwin Khin Shar, Biniam Fisseha Demissie, Mariano Ceccato, Yan Naing Tun, David Lo 0001, Lingxiao Jiang, Christoph Bienert
Empir. Softw. Eng.2
2020 Security analysis of permission re-delegation vulnerabilities in Android apps
abstract
Abstract The Android platform facilitates reuse of app functionalities by allowing an app to request an action from another app through inter-process communication mechanism. This feature is one of the reasons for the popularity of Android, but it also poses security risks to the end users because malicious, unprivileged apps could exploit this feature to make privileged apps perform privileged actions on behalf of them. In this paper, we investigate the hybrid use of program analysis, genetic algorithm based test generation, natural language processing, machine learning techniques for precise detection of permission re-delegation vulnerabilities in Android apps. Our approach first groups a large set of benign and non-vulnerable apps into different clusters, based on their similarities in terms of functional descriptions. It then generates permission re-delegation model for each cluster, which characterizes common permission re-delegation behaviors of the apps in the cluster. Given an app under test, our approach checks whether it has permission re-delegation behaviors that deviate from the model of the cluster it belongs to. If that is the case, it generates test cases to detect the vulnerabilities. We evaluated the vulnerability detection capability of our approach based on 1,258 official apps and 20 mutated apps. Our approach achieved 81.8% recall and 100% precision. We also compared our approach with two static analysis-based approaches — Covert and IccTA — based on 595 open source apps. Our approach detected 30 vulnerable apps whereas Covert detected one of them and IccTA did not detect any. Executable proof-of-concept attacks generated by our approach were reported to the corresponding app developers.
Biniam Fisseha Demissie, Mariano Ceccato, Lwin Khin Shar
Empir. Softw. Eng.1