VLDB 2026 Research / reviewers in the wild / expert
Phu-Vinh Nguyen
dblp:384/9697
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Vision and language · 44% Image recognition and object detection · 44% Speech recognition and synthesis · 13% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
object localization |
0.9 | 1 | 2025 | SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization · EMNLP 2025 |
Computer vision › Vision and language
visual question answering |
0.9 | 1 | 2025 | SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
multimodal fusion · 0.9chain-of-thought reasoning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object LocalizationabstractVisual Language Models have demonstrated remarkable capabilities across various tasks, including visual question answering and image captioning.However, most models rely on textbased instructions, limiting their effectiveness in natural human-machine interactions.Moreover, the quality of language models primarily depends on reasoning and prompting techniques, such as chain-of-thought, which remain underexplored when using speech instructions.To address these challenges, we propose SilVar, an end-to-end multimodal model that leverages speech instructions for reasoning-based visual question answering.Additionally, we investigate reasoning techniques at different levels, including conversational, simple, and complex speech instructions.SilVar is built upon CLIP, Whisper, and LLaMA 3.1-8B, enabling more intuitive interactions by allowing users to provide verbal or text-based instructions.To this end, we introduce a new dataset designed to challenge models with speech-based reasoning tasks for object localization.This dataset enhances the model's ability to process and explain visual scenes from spoken input, moving beyond simple object recognition to reasoningbased interactions.To our knowledge, SilVar is the first open-source, speech-driven VLM.We believe SilVar will inspire the next generation of multimodal reasoning models, advancing toward expert artificial general intelligence.Our code and dataset are publicly available here. Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen, Chris Ngo, Truong Son Hy |
EMNLP | 3 |
| 2024 | Advancing Vietnamese Information Retrieval with Learning Objective and Benchmark
Phu-Vinh Nguyen, Minh-Nam Tran, Long H. B. Nguyen, Dinh Dien |
PACLIC | 1 |