Minh-Hao Van

dblp:304/3234 · DBLP profile ↗
← Back
7ranked-venue papers in the field
3as first author
7since 2021 · last 2025
0000-0001-7342-6801ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 5 (1 first)Database Systems & Data Management · 1 (1 first)Data Mining & Knowledge Discovery · 1 (1 first)
YearPublicationVenuePosition
2025 Fair In-Context Learning via Latent Concept Variables
Karuna Bhaila, Minh-Hao Van, Kennedy Edemacu, Chen Zhao 0010, Feng Chen 0001, Xintao Wu
IEEE Big Data2
2025 A Machine Learning Framework for Automated Computational Ethology Using Markerless Pose Estimation
Minh-Hao Van, Christopher McEnaney, Ashtyn Le, Amy R. Poe, Xintao Wu
IEEE Big Data2
2025 Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
An Vuong, Minh-Hao Van, Chen Zhao 0010, Xintao Wu
IEEE Big Data2
2024 Beyond Human Vision: The Role of Large Vision Language Models in Microscope Image Analysis
abstract
Vision language models (VLMs) such as LLaVA, ChatGPT-4, and Gemini have recently emerged and gained the spotlight for their ability to comprehend the dual modality of image and textual data showing impressive performance on tasks such as natural image captioning, visual question answering, and spatial reasoning. Additionally, a universal segmentation model by Meta AI, Segment Anything Model (SAM) shows unprecedented performance at isolating objects from unforeseen images. Because medical experts, biologists, and materials scientists routinely examine microscopy or medical images in conjunction with textual information in the form of captions, literature, or reports, and draw conclusions of great importance and merit, it is essential to evaluate their performance on these images. In this study, we charge ChatGPT, LLaVA, Gemini, and SAM quantitatively with classification, segmentation and counting tasks. We observed that ChatGPT and Gemini were impressively able to comprehend the visual features in microscopy images, while SAM was quite capable at isolating artifacts in a general sense. However, the performance was not close to that of a domain expert – the models were readily encumbered by the introduction of impurities, defects, object overlaps and diversity present in the images.
Minh-Hao Van, Xintao Wu
IEEE Big Data2
2023 HINT: Healthy Influential-Noise based Training to Defend against Data Poisoning Attacks
abstract
While numerous defense methods have been proposed to prohibit potential poisoning attacks from untrusted data sources, most research works only defend against specific attacks, which leaves many avenues for an adversary to exploit. In this work, we propose an efficient and robust training approach to defend against data poisoning attacks based on influence functions, named Healthy Influential-Noise based Training. Using influence functions, we craft healthy noise that helps to harden the classification model against poisoning attacks without significantly affecting the generalization ability on test data. In addition, our method can perform effectively when only a subset of the training data is modified, instead of the current method of adding noise to all examples that has been used in several previous works. We conduct comprehensive evaluations over two image datasets with state-of-the-art poisoning attacks under different realistic attack scenarios. Our empirical results show that HINT can efficiently protect deep learning models against the effect of both untargeted and targeted poisoning attacks.
Minh-Hao Van, Alycia N. Carey, Xintao Wu
ICDM1
2022 Defending Evasion Attacks via Adversarially Adaptive Training
abstract
Adversarial machine learning has been extensively studied from perspectives of attack settings and defense strategies. However, existing adversarial training models fail to be adaptive and robust against new attacks during test time. In this paper, we propose a novel adversarially adaptive defense (AAD) framework based on adaptive training such that the trained prediction and detection models adapt at test time to new attacks. Our AAD structures the training data into groups and each group represents one attack scenario. Different from empirical risk minimization that trains a single robust model or learns an invariant feature space, our AAD learns a context vector from features of each batch during training and incorporates the learned context vector into both prediction and detection models. Thus, AAD can adapt at test time to new adversarial attacks. We formulate our problem by optimizing a joint loss from prediction, detection, and regularization via a multi-task learning framework. We conduct comprehensive empirical evaluations with popular adversarial attacks and defense strategies on two real-world datasets under different attack settings. Empirical results show that AAD achieves both high prediction and detection accuracy and significantly outperforms baselines.
Minh-Hao Van, Wei Du 0009, Xintao Wu, Feng Chen 0001, Aidong Lu
IEEE Big Data1
2022 Poisoning Attacks on Fair Machine Learning
Minh-Hao Van, Wei Du 0009, Xintao Wu, Aidong Lu
DASFAA (1)1