Chris Ngo

dblp:392/3888 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Machine translation · 30% Vision and language · 30% Image recognition and object detection · 30%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation › speech translation
multilingual speech translation
0.912025
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation · EMNLP 2025
Computer vision › Image recognition and object detection
object localization
0.912025
SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization · EMNLP 2025
Computer vision › Vision and language
visual question answering
0.912025
SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization · EMNLP 2025
Medical and health informatics › clinical text processing
clinical dialogue
0.312025
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

speech translation · 1.7multimodal fusion · 0.9chain-of-thought reasoning · 0.9
YearPublicationVenuePosition
2025 MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
abstract
Khai Le-Duc, Tuyen Tran, Bach Phan Tat, Nguyen Kim Hai Bui, Quan Dang Anh, Hung-Phong Tran, Thanh Thuy Nguyen, Ly Nguyen, Tuan Minh Phan, Thi Thu Phuong Tran, Chris Ngo, Khanh Xuan Nguyen, Thanh Nguyen-Tang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Khai Le-Duc, Tuyen Tran, Bach Phan Tat, Nguyen Kim Hai Bui, Quan Dang Anh, Hung-Phong Tran, Thanh Thuy Nguyen, Ly Nguyen, Tuan-Minh Phan, Thi Thu Phuong Tran, Chris Ngo, Nguyen X. Khanh, Thanh Nguyen-Tang
EMNLP11
2025 SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
abstract
Visual Language Models have demonstrated remarkable capabilities across various tasks, including visual question answering and image captioning.However, most models rely on textbased instructions, limiting their effectiveness in natural human-machine interactions.Moreover, the quality of language models primarily depends on reasoning and prompting techniques, such as chain-of-thought, which remain underexplored when using speech instructions.To address these challenges, we propose SilVar, an end-to-end multimodal model that leverages speech instructions for reasoning-based visual question answering.Additionally, we investigate reasoning techniques at different levels, including conversational, simple, and complex speech instructions.SilVar is built upon CLIP, Whisper, and LLaMA 3.1-8B, enabling more intuitive interactions by allowing users to provide verbal or text-based instructions.To this end, we introduce a new dataset designed to challenge models with speech-based reasoning tasks for object localization.This dataset enhances the model's ability to process and explain visual scenes from spoken input, moving beyond simple object recognition to reasoningbased interactions.To our knowledge, SilVar is the first open-source, speech-driven VLM.We believe SilVar will inspire the next generation of multimodal reasoning models, advancing toward expert artificial general intelligence.Our code and dataset are publicly available here.
Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen, Chris Ngo, Truong Son Hy
EMNLP4