Phu-Vinh Nguyen

dblp:384/9697 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Vision and language · 44% Image recognition and object detection · 44% Speech recognition and synthesis · 13%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object localization
0.912025
SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization · EMNLP 2025
Computer vision › Vision and language
visual question answering
0.912025
SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 0.9chain-of-thought reasoning · 0.9
YearPublicationVenuePosition
2025 SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
abstract
Visual Language Models have demonstrated remarkable capabilities across various tasks, including visual question answering and image captioning.However, most models rely on textbased instructions, limiting their effectiveness in natural human-machine interactions.Moreover, the quality of language models primarily depends on reasoning and prompting techniques, such as chain-of-thought, which remain underexplored when using speech instructions.To address these challenges, we propose SilVar, an end-to-end multimodal model that leverages speech instructions for reasoning-based visual question answering.Additionally, we investigate reasoning techniques at different levels, including conversational, simple, and complex speech instructions.SilVar is built upon CLIP, Whisper, and LLaMA 3.1-8B, enabling more intuitive interactions by allowing users to provide verbal or text-based instructions.To this end, we introduce a new dataset designed to challenge models with speech-based reasoning tasks for object localization.This dataset enhances the model's ability to process and explain visual scenes from spoken input, moving beyond simple object recognition to reasoningbased interactions.To our knowledge, SilVar is the first open-source, speech-driven VLM.We believe SilVar will inspire the next generation of multimodal reasoning models, advancing toward expert artificial general intelligence.Our code and dataset are publicly available here.
Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen, Chris Ngo, Truong Son Hy
EMNLP3
2024 Advancing Vietnamese Information Retrieval with Learning Objective and Benchmark
Phu-Vinh Nguyen, Minh-Nam Tran, Long H. B. Nguyen, Dinh Dien
PACLIC1