Kunyu Yang

dblp:280/2487 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0007-6004-9147ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LGVLM-mIoT: A Lightweight Generative Visual-Language Model for Multilingual IoT Applications
abstract
The demand for edge device models equipped with multilingual visual capabilities is rapidly increasing in complex IoT application scenarios. While many studies have endowed models with strong visual sensory and language analysis capabilities, these models are often large and require substantial amounts of data. Moreover, multilingual parallel corpora are extremely scarce. Although large parameter sizes can enhance a model’s visual-language processing capabilities, the high training and inference costs make them unsuitable for edge devices and result in suboptimal performance in multilingual contexts. To address these challenges, this article proposes an generative visual language model, that is, cross-lingual, lightweight, data-efficient, easy to train, and easy to infer. We map both English and non-English features into the same space and align them with a visually distilled model while leveraging the inherent similarity information of languages to increase the supervision coverage of the dataset. Through extensive experiments, we demonstrate that our model achieves state-of-the-art performance across three downstream tasks: 1) image captioning; 2) machine translation; and 3) visual question answering, surpassing existing methods.
Kunyu Yang, Xiangyun Tang
IEEE Internet Things J.2
2025 SCAG: Semantic Co-occurring Attention Guided Alignment for Knowledge-based Visual Question Answering
abstract
In the realm of Knowledge-based Visual Question Answering (KB-VQA), the intricacy of the task lies in adeptly retrieving pertinent information from external sources and seamlessly aligning and amalgamating multimodal features. While numerous studies have effectively leveraged external knowledge to enrich factual connections among entities, there exists a tendency to overlook the significant reservoir of implicit information inherent in the visual-textual dimension. This oversight often results in suboptimal alignment and an undue reliance on the knowledge base. To address these challenges, this article introduces a novel strategy called SCAG. This approach aggregates the semantic co-occurring attention from diverse regions within images and various tokens within textual inputs using guidance weights to construct joint probabilistic representations grounded in the visual and textual dimensions, respectively. By employing this alignment strategy, the goal is to substantially mitigate information loss, reinforce inter-feature constraints within the model, reduce reliance on external knowledge sources, and enhance self-reasoning capabilities. The efficacy of our proposed model is comprehensively evaluated on the VQAv2 and OK-VQA datasets, with comparative analyses against multiple models conducted on the Ambiguous Knowledge (AK) dataset. Notably, our model exhibits a noteworthy 4.62% improvement over the state-of-the-art in addressing the knowledge dependency problem.
Kunyu Yang, Xuan Liu 0008, Honghao Gao
ACM Trans. Multim. Comput. Commun. Appl.2
2022 VAEPass: A lightweight passwords guessing model based on variational auto-encoder
Kunyu Yang, Xuexian Hu, Qihui Zhang, Jianghong Wei, Wenfen Liu
Comput. Secur.1
2021 Studies of Keyboard Patterns in Passwords: Recognition, Characteristics and Strength Evolution
Kunyu Yang, Xuexian Hu, Qihui Zhang, Jianghong Wei, Wenfen Liu
ICICS (1)1