EDBT 2026 Demo / reviewers in the wild / expert
Steven Quoc Hung Truong
dblp:293/7483 · also Steven Q. H. Truong
· DBLP profile ↗
14ranked-venue papers
0as first author
14since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GIIM: Graph-based Learning of Inter- and Intra-view Dependencies for Multi-view Medical Image DiagnosisabstractComputer-aided diagnosis (CADx) has become vital in medical imaging, but automated systems often struggle to replicate the nuanced process of clinical interpretation. Expert diagnosis requires a comprehensive analysis of how abnormalities relate to each other across various views and time points, but current multi-view CADx methods frequently overlook these complex dependencies. Specifically, they fail to model the crucial relationships within a single view and the dynamic changes lesions exhibit across different views. This limitation, combined with the common challenge of incomplete data, greatly reduces their predictive reliability. To address these gaps, we reframe the diagnostic task as one of relationship modeling and propose GIIM, a novel graph-based approach. Our framework is uniquely designed to simultaneously capture both critical intra-view dependencies between abnormalities and inter-view dynamics. Furthermore, it ensures diagnostic robustness by incorporating specific techniques to effectively handle missing data, a common clinical issue. We demonstrate the generality of this approach through extensive evaluations on diverse imaging modalities, including CT, MRI, and mammography. The results confirm that our GIIM model significantly enhances diagnostic accuracy and robustness over existing methods, establishing a more effective framework for future CADx systems. Tran Bao Sam, Hung Vu, Trung Kien Dao, Tran Dat Dang, Van Ha Tang, Steven Quoc Hung Truong |
AAAI | 6 |
| 2025 | AdaCS: Adaptive Normalization for Enhanced Code-Switching ASRabstractIntra-sentential code-switching (CS) refers to the alternation between languages that happens within a single utterance and is a significant challenge for Automatic Speech Recognition (ASR) systems. For example, when a Vietnamese speaker uses foreign proper names or specialized terms within their speech. ASR systems often struggle to accurately transcribe intra-sentential CS due to their training on monolingual data and the unpredictable nature of CS. This issue is even more pronounced for low-resource languages, where limited data availability hinders the development of robust models. In this study, we propose AdaCS, a normalization model integrates an adaptive bias attention module (BAM) into encoder-decoder network. This novel approach provides a robust solution to CS ASR in unseen domains, thereby significantly enhancing our contribution to the field. By utilizing BAM to both identify and normalize CS phrases, AdaCS enhances its adaptive capabilities with a biased list of words provided during inference. Our method demonstrates impressive performance and the ability to handle unseen CS phrases across various domains. Experiments show that AdaCS outperforms previous state-of-the-art method on Vietnamese CS ASR normalization by considerable WER reduction of 56.2% and 36.8% on the two proposed test sets. The Chuong Chu, Pham Vu Tuan Dat, Trung Kien Dao, Ngoc Hoang Nguyen 0001, Steven Quoc Hung Truong |
ICASSP | 5 |
| 2024 | A Multi-phase Multi-graph Approach for Focal Liver Lesion Classification on CT Scans
Tran Bao Sam, Ta Duc Huy, Cong Tuyen Dao, Thanh Tin Lam, Van Ha Tang, Steven Quoc Hung Truong |
ACCV (10) | 6 |
| 2023 | Revisiting Reverse Distillation for Anomaly DetectionabstractAnomaly detection is an important application in large-scale industrial manufacturing. Recent methods for this task have demonstrated excellent accuracy but come with a latency trade-off. Memory based approaches with dominant performances like PatchCore or Coupled-hypersphere-based Feature Adaptation (CFA) require an external memory bank, which significantly lengthens the execution time. Another approach that employs Reversed Distillation (RD) can perform well while maintaining low latency. In this paper, we revisit this idea to improve its performance, establishing a new state-of-the-art benchmark on the challenging MVTec dataset for both anomaly detection and localization. The proposed method, called RD++, runs six times faster than PatchCore, and two times faster than CFA but introduces a negligible latency compared to RD. We also experiment on the BTAD and Retinal OCT datasets to demonstrate our method's generalizability and conduct important ablation experiments to provide insights into its configurations. Source code will be available at https://github.com/tientrandinh/Revisiting-Reverse-Distillation. Tran Dinh Tien, Nguyen Hoang Tran, Ta Duc Huy, Soan Thi Minh Duong, Chanh D. Tr. Nguyen, Steven Quoc Hung Truong |
CVPR | 7 |
| 2023 | An efficient approach for real-time abnormal human behavior recognition on surveillance camerasabstractIn recent years, abnormal human behavior recognition has become an attractive research topic of computer vision due to the rapid growth of demand to monitor human activities on closed-circuit television (CCTV) cameras. However, developing a deep learning-based model for abnormal/violent behavior recognition in surveillance systems is still quite challenging and costly due to inadequate data and model complexity. This paper presents an efficient approach to recognize violent behavior such as fighting, sexual harassment, and climbing fence in real-time on a multi-camera-one-edge-device system. Our approach develops a lightweight 3DCNN model trained by an effective optimization process to recognize human behavior from sequence frames of CCTV video signal input. In the optimization method, we utilize two advantages of deep learning techniques of knowledge distillation and contrastive learning to enhance the quality of the lightweight model on recognizing recorded human behaviors, which can help the student network learn distilled information from both the bigger model and contrastive object representations. We also establish a large CCTV human behavior video dataset containing 4,200 abnormal and 24,000 normal videos. The effectiveness of the proposed approach is shown by the high inference performance and impressive results evaluated on both public datasets the RWF-2000 dataset, the UCF101 dataset, and our collected datasets. Ngoc Hoang Nguyen 0001, Nhat Nguyen Xuan, Trung H. Bui, Dao Huu Hung, Steven Quoc Hung Truong, Vu Hoang |
FG | 5 |
| 2023 | Logovit: Local-Global Vision Transformer for Object Re-IdentificationabstractObject re-identification (ReID) is prone to errors under variations in scale, illumination, complex background, and object occlusion scenarios. To overcome these challenges, attention mechanisms are employed to focus on the object's characteristics, thereby extracting better discriminative features. This paper introduces a local-global vision transformer (LoGoViT) for object re-identification by learning a hierarchical-level representation from fine-grained (local) to general (global) context features. It comprises two components: (i) shift and shuffle operations to generate robust local features and (ii) local-global module to aggregate the multi-level hierarchy features of an object. Extensive experiments show that our method achieves state-of-the-art on the ReID benchmarks. We further investigate effective augmentation operations and discuss how the patch modifications improve the proposed model's generalization under occlusion scenarios. The source code is available at https://github.com/nguyenphan99/LoGoViT. Nguyen Phan, Ta Duc Huy, Soan Thi Minh Duong, Nguyen Hoang Tran, Sam Tran, Dao Huu Hung, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong |
ICASSP | 9 |
| 2022 | Dual consistency assisted multi-confident learning for the hepatic vessel segmentation using noisy labels
Nam Nguyen Phuong, Tuan Van Vo, Soan Thi Minh Duong, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong |
BMVC | 6 |
| 2022 | Improving Local Features with Relevant Spatial Information by Vision Transformer for Crowd Counting
Nguyen Hoang Tran, Ta Duc Huy, Soan Thi Minh Duong, Nguyen Phan, Dao Huu Hung, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong |
BMVC | 8 |
| 2022 | Adaptive Proxy Anchor Loss for Deep Metric LearningabstractDeep metric learning (or simply called metric learning) uses the deep neural network to learn the representation of images, leading to widely used in many applications, e.g. image retrieval and face recognition. In the metric learning approaches, proxy anchor takes advantage of proxy-based and pair-based approaches to enable fast convergence time and robustness to noisy labels. However, in training the proxy anchor, selecting the hyperparameter margin is important to achieve a good performance. This selection requires expertise and is time-consuming. This paper proposes a novel method to learn the margin while training the proxy anchor approach adaptively. The proposed adaptive proxy anchor simplifies the hyperparameter tuning process while advancing the proxy anchor. We achieve state of the art on three public datasets with a noticeably faster convergence time. Our code is available at https: //github.com/tks1998/Adaptive-Proxy-Anchor Nguyen Phan, Sen Tran, Ta Duc Huy, Soan Thi Minh Duong, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong |
ICIP | 7 |
| 2022 | An Efficient and High Fidelity Vietnamese Streaming End-to-End Speech Synthesis
Tho Nguyen Duc Tran, The Chuong Chu, Vu Hoang, Trung Huu Bui, Steven Quoc Hung Truong |
INTERSPEECH | 5 |
| 2022 | ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text MiningabstractPre-trained language models have become crucial to achieving competitive results across many Natural Language Processing (NLP) problems. For monolingual pre-trained models in low-resource languages, the quantity has been significantly increased. However, most of them relate to the general domain, and there are limited strong baseline language models for domain-specific. We introduce ViHealthBERT, the first domain-specific pre-trained language model for Vietnamese healthcare. The performance of our model shows strong results while outperforming the general domain language models in all health-related datasets. Moreover, we also present Vietnamese datasets for the healthcare domain for two tasks are Acronym Disambiguation (AD) and Frequently Asked Questions (FAQ) Summarization. We release our ViHealthBERT to facilitate future research and downstream application for Vietnamese NLP in domain-specific. Our dataset and code are available in https://github.com/demdecuong/vihealthbert. Nguyen Phuc Minh, Tran Hoang Vu, Vu Hoang, Ta Duc Huy, Trung Huu Bui, Steven Quoc Hung Truong |
LREC | 6 |
| 2021 | Diffeomorphism Matching for Fast Unsupervised Pretraining on Radiographs
Huynh Minh Thanh, Chanh D. Tr. Nguyen, Ta Duc Huy, Hoang Cao Huyen, Trung H. Bui, Steven Quoc Hung Truong |
BMVC | 6 |
| 2021 | ReSORT: an ID-recovery multi-face tracking method for surveillance camerasabstractAs an improvement over the standard simple online real-time tracking (SORT) method, DeepSORT introduces a cascade matching mechanism to track objects during a certain period of occlusion, effectively reducing the number of identity (ID) switches. However, DeepSORT lacks the capability of lost-identities recovery, which enables robustness and performance in face recognition systems. To address the issue, we propose a novel multi-face tracking method, named ReSORT, that can recover lost identities. Our method removes the cascade matching block in DeepSORT and extends a similarity matching (SM) block after the Kalman filter to assign uncertain tracks to their probable tracking IDs. Such arrangement significantly reduces the processing time while maintaining the longevity of tracking IDs. The SM block functions by storing existing facial features and comparing the similarity between the new and the existing facial features, enabling ReSORT to recover the lost IDs or IDs from other cameras. To benchmark the ID-recovery ability, we introduce three new metrics, calling IDnew, TIDRate, and TReRate. We also produce face tracking annotations for three public surveillance camera datasets, i.e., LAB, MSU-AVIS, and ChokePoint. Extensive experiments conducted on the three datasets with various resolutions and frame-rates settings demonstrate the superiority of ReSORT over DeepSORT, i.e. reducing the identity switches by average 36.38%, and the processing time by 5.19 times. Source code and annotations of all three datasets are available at https://github.com/tantm97/ReSORT. Tan M. Tran, Nguyen Hoang Tran, Soan Thi Minh Duong, Ta Duc Huy, Chanh D. Tr. Nguyen, Trung H. Bui, Steven Quoc Hung Truong |
FG | 7 |
| 2021 | ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development
Ta Duc Huy, Nguyen Anh Tu, Tran Hoang Vu, Nguyen Phuc Minh, Nguyen Phan, Trung H. Bui, Steven Quoc Hung Truong |
ICONIP (6) | 7 |