VLDB 2026 Research / reviewers in the wild / expert
Kunyu Yang
dblp:280/2487
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0007-6004-9147ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LGVLM-mIoT: A Lightweight Generative Visual-Language Model for Multilingual IoT ApplicationsabstractThe demand for edge device models equipped with multilingual visual capabilities is rapidly increasing in complex IoT application scenarios. While many studies have endowed models with strong visual sensory and language analysis capabilities, these models are often large and require substantial amounts of data. Moreover, multilingual parallel corpora are extremely scarce. Although large parameter sizes can enhance a model’s visual-language processing capabilities, the high training and inference costs make them unsuitable for edge devices and result in suboptimal performance in multilingual contexts. To address these challenges, this article proposes an generative visual language model, that is, cross-lingual, lightweight, data-efficient, easy to train, and easy to infer. We map both English and non-English features into the same space and align them with a visually distilled model while leveraging the inherent similarity information of languages to increase the supervision coverage of the dataset. Through extensive experiments, we demonstrate that our model achieves state-of-the-art performance across three downstream tasks: 1) image captioning; 2) machine translation; and 3) visual question answering, surpassing existing methods. Kunyu Yang, Xiangyun Tang |
IEEE Internet Things J. | 2 |
| 2025 | SCAG: Semantic Co-occurring Attention Guided Alignment for Knowledge-based Visual Question AnsweringabstractIn the realm of Knowledge-based Visual Question Answering (KB-VQA), the intricacy of the task lies in adeptly retrieving pertinent information from external sources and seamlessly aligning and amalgamating multimodal features. While numerous studies have effectively leveraged external knowledge to enrich factual connections among entities, there exists a tendency to overlook the significant reservoir of implicit information inherent in the visual-textual dimension. This oversight often results in suboptimal alignment and an undue reliance on the knowledge base. To address these challenges, this article introduces a novel strategy called SCAG. This approach aggregates the semantic co-occurring attention from diverse regions within images and various tokens within textual inputs using guidance weights to construct joint probabilistic representations grounded in the visual and textual dimensions, respectively. By employing this alignment strategy, the goal is to substantially mitigate information loss, reinforce inter-feature constraints within the model, reduce reliance on external knowledge sources, and enhance self-reasoning capabilities. The efficacy of our proposed model is comprehensively evaluated on the VQAv2 and OK-VQA datasets, with comparative analyses against multiple models conducted on the Ambiguous Knowledge (AK) dataset. Notably, our model exhibits a noteworthy 4.62% improvement over the state-of-the-art in addressing the knowledge dependency problem. Kunyu Yang, Xuan Liu 0008, Honghao Gao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | VAEPass: A lightweight passwords guessing model based on variational auto-encoder
Kunyu Yang, Xuexian Hu, Qihui Zhang, Jianghong Wei, Wenfen Liu |
Comput. Secur. | 1 |
| 2021 | Studies of Keyboard Patterns in Passwords: Recognition, Characteristics and Strength Evolution
Kunyu Yang, Xuexian Hu, Qihui Zhang, Jianghong Wei, Wenfen Liu |
ICICS (1) | 1 |