Jinhan Dong

dblp:408/3481 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0002-8575-6378ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Vision and language · 39% Legged, aerial and field robots · 30% Robot navigation and mapping · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping › mobile robot navigation › 3d navigation
aerial robot navigation
0.912025
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025
Robotics › Legged, aerial and field robots › aerial robots
UAV navigation
0.912025
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025
Computer vision › Vision and language
vision-and-language navigation
0.912025
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025
Computer vision › Vision and language
vision-language model
0.312025
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

vision-language model · 0.9reinforcement learning · 0.9
YearPublicationVenuePosition
2025 FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
abstract
Hengxing Cai, Jinhan Dong, Jingjun Tan, Jingcheng Deng, Sihang Li, Zhifeng Gao, Haidong Wang, Zicheng Su, Agachai Sumalee, Renxin Zhong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Hengxing Cai, Jinhan Dong, Jingjun Tan, Jingcheng Deng, Sihang Li 0002, Zhifeng Gao, Zicheng Su, Agachai Sumalee, Renxin Zhong
EMNLP2
2025 MT-Agent: Constructing a GUI Agent via Modality Enhancement and Text-Guided Fusion
abstract
Graphical User Interfaces (GUIs) play a crucial role in facilitating user-computer interactions, making them an essential focus of research. However, current automated GUI agents face significant challenges in effectively associating task implementations with specific visual elements, and the resolution constraints of Vision-Language Models (VLMs) also limit the richness of visual information. To this end, we propose a novel multi-modal agent named MT-Agent, which enhances both textual and visual input modalities to enable the model to perceive visual elements in GUIs more effectively. Specifically, Textual Modality Enhancement improves the semantic richness of input text by capturing task-specific details via an external VLM, while Visual Modality Enhancement incorporates fine-grained visual details to better represent critical GUI elements. In addition, we introduce an innovative text-guided directional feature fusion mechanism, which leverages enriched text features to guide the integration with visual information. In experiments, MT-Agent demonstrated exceptional performance on AITZ dataset, achieving an action type prediction accuracy of 84.80% and a step prediction accuracy of 58.07%, surpassing previous state-of-the-art models. Furthermore, on the GUI Odyssey benchmark, MT-Agent achieves performance comparable to previous state-of-the-art models while using only about 1/20 of their trainable parameters. Our codes, demos, and relevant data will be released to facilitate further research and validation within the scientific community.
Jinhan Dong, Lei Jin 0003, Zhihong Zhang 0006, Runqing Zhang, Liqiang Xu, Junliang Xing
IEEE Internet Things J.1