EDBT 2026 Demo / reviewers in the wild / expert
Tri Cao
dblp:217/3502
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-7865-8476ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 36% Language models and text generation · 19% Representation and self-supervised learning · 15% | |
| Databases, data mining, and information retrieval
3 papers |
Machine learning and data management · 41% Knowledge graphs · 31% Data mining · 27% | |
| Network and information security
2 papers |
Web and mobile security · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 100% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
1.7 | 2 | 2025 | Automating Steering for Safe Multimodal Large Language Models · EMNLP 2025 Words or Vision: Do Vision-Language Models Have Blind Faith in Text? · CVPR 2025 |
Web and mobile security
phishing detection |
1.6 | 2 | 2025 | PhishAgent: A Robust Multimodal Agent for Phishing Webpage Detection · AAAI 2025 KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection · USENIX Security Symposium 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
1.3 | 2 | 2023 | LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph Matching · NeurIPS 2023 Joint Self-Supervised Image-Volume Representation Learning with Intra-inter Contrastive Clustering · AAAI 2023 |
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering |
0.9 | 1 | 2025 | Automating Steering for Safe Multimodal Large Language Models · EMNLP 2025 |
Natural language and speech › Language models and text generation
inference-time intervention |
0.9 | 1 | 2025 | Automating Steering for Safe Multimodal Large Language Models · EMNLP 2025 |
Knowledge, reasoning and agents › Multi-agent systems
multimodal agent |
0.9 | 1 | 2025 | PhishAgent: A Robust Multimodal Agent for Phishing Webpage Detection · AAAI 2025 |
Machine learning › Trustworthy machine learning › generative model safety
multimodal large language model safety |
0.9 | 1 | 2025 | Automating Steering for Safe Multimodal Large Language Models · EMNLP 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Words or Vision: Do Vision-Language Models Have Blind Faith in Text? · CVPR 2025 |
Knowledge graphs
multimodal knowledge graph |
0.8 | 1 | 2024 | KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection · USENIX Security Symposium 2024 |
Web and mobile security › phishing detection
reference-based phishing detection |
0.8 | 1 | 2024 | KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection · USENIX Security Symposium 2024 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.7 | 1 | 2023 | Anomaly Detection under Distribution Shift · ICCV 2023 |
Medical and health informatics
medical imaging |
0.7 | 1 | 2023 | LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph Matching · NeurIPS 2023 |
Data mining
anomaly detection |
0.7 | 1 | 2023 | Anomaly Detection under Distribution Shift · ICCV 2023 |
Machine learning › Optimization for machine learning
optimal transport |
0.3 | 1 | 2026 | MELCOT: A Hybrid Learning Architecture with Marginal Preservation for Matrix-Valued Regression · WSDM 2026 |
Machine learning › Transfer learning and domain adaptation
domain shift |
0.2 | 1 | 2023 | LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph Matching · NeurIPS 2023 |
Computer vision › Image recognition and object detection › medical image analysis
medical image classification |
0.2 | 1 | 2023 | LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph Matching · NeurIPS 2023 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
0.2 | 1 | 2023 | LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph Matching · NeurIPS 2023 |
Medical and health informatics › medical imaging › medical image analysis › medical image segmentation
3d medical image segmentation |
0.2 | 1 | 2023 | Joint Self-Supervised Image-Volume Representation Learning with Intra-inter Contrastive Clustering · AAAI 2023 |
Medical and health informatics › medical imaging
medical image analysis |
0.2 | 1 | 2023 | Joint Self-Supervised Image-Volume Representation Learning with Intra-inter Contrastive Clustering · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
optimal transport · 2.0marginal estimation · 2.0deep learning · 2.0multimodal large language model · 1.7multimodal information retrieval · 1.7large language model · 1.5text augmentation · 0.9supervised fine-tuning · 0.9safety probing · 0.9representation engineering · 0.9refusal head · 0.9multimodal knowledge graphs · 0.8multimodal knowledge graph · 0.8unsupervised adaptation · 0.7second-order graph matching · 0.7masked embedding prediction · 0.7distribution alignment · 0.7deformable self-attention · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MELCOT: A Hybrid Learning Architecture with Marginal Preservation for Matrix-Valued RegressionabstractRegression is essential across many domains but remains challenging in high-dimensional settings, where existing methods often lose spatial structure or demand heavy storage. In this work, we address the problem of matrix-valued regression, where each sample is naturally represented as a matrix. We propose MELCOT, a hybrid model that integrates a classical machine–learning–based Marginal Estimation (ME) block with a deep-learning–based Learnable-Cost Optimal Transport (LCOT) block. The ME block estimates data marginals to preserve spatial information, while the LCOT block learns complex global features. This design enables MELCOT to inherit the strengths of both classical and deep learning methods. Extensive experiments across diverse datasets and domains demonstrate that MELCOT consistently outperforms all baselines while remaining highly efficient and provides potential applications in various domains, including high-content imaging (HCI). Khang Tran, Hieu Cao, Thinh Pham, Nghiem Diep, Tri Cao |
WSDM | 5 |
| 2025 | PhishAgent: A Robust Multimodal Agent for Phishing Webpage DetectionabstractPhishing attacks are a major threat to online security, exploiting user vulnerabilities to steal sensitive information. Various methods have been developed to counteract phishing, each with varying levels of accuracy, but they also face notable limitations. In this study, we introduce PhishAgent, a multimodal agent that combines a wide range of tools, integrating both online and offline knowledge bases with Multimodal Large Language Models (MLLMs). This combination leads to broader brand coverage, which enhances brand recognition and recall. Furthermore, we propose a multimodal information retrieval framework designed to extract the relevant top k items from offline knowledge bases, using available information from a webpage, including logos and HTML. Our empirical results, based on three real-world datasets, demonstrate that the proposed framework significantly enhances detection accuracy and reduces both false positives and false negatives, while maintaining model efficiency. Additionally, PhishAgent shows strong resilience against various types of adversarial attacks. Tri Cao, Chengyu Huang 0003, Yuexin Li, Huilin Wang, Amy He, Nay Oo, Bryan Hooi |
AAAI | 1 |
| 2025 | Words or Vision: Do Vision-Language Models Have Blind Faith in Text?abstractVision-Language Models (VLMs) excel in integrating visual and textual information for vision-centric tasks, but their handling of inconsistencies between modalities is underexplored. We investigate VLMs’ modality preferences when faced with visual data and varied textual inputs in vision-centered settings. By introducing textual variations to four vision-centric tasks and evaluating ten Vision-Language Models (VLMs), we discover a "blind faith in text" phenomenon: VLMs disproportionately trust textual data over visual data when inconsistencies arise, leading to significant performance drops under corrupted text and raising safety concerns. We analyze factors influencing this text bias, including instruction prompts, language model size, text relevance, token order, and the interplay between visual and textual certainty. While certain factors, such as scaling up the language model size, slightly mitigate text bias, others like token order can exacerbate it due to positional biases inherited from language models. To address this issue, we explore supervised fine-tuning with text augmentation and demonstrate its effectiveness in reducing text bias. Additionally, we provide a theoretical analysis suggesting that the blind faith in text phenomenon may stem from an imbalance of pure text and multi-modal data during training. Our findings highlight the need for balanced training and careful consideration of modality interactions in VLMs to enhance their robustness and reliability in handling multi-modal data inconsistencies. Ailin Deng, Tri Cao, Bryan Hooi |
CVPR | 2 |
| 2025 | Automating Steering for Safe Multimodal Large Language ModelsabstractRecent progress in Multimodal Large Language Models (MLLMs) has unlocked powerful cross-modal reasoning abilities, but also raised new safety concerns, particularly when faced with adversarial multimodal inputs. To improve the safety of MLLMs during inference, we introduce a modular and adaptive inference-time intervention technology, AutoSteer, without requiring any fine-tuning of the underlying model. AutoSteer incorporates three core components: (1) a novel Safety Awareness Score (SAS) that automatically identifies the most safety-relevant distinctions among the model’s internal layers; (2) an adaptive safety prober trained to estimate the likelihood of toxic outputs from intermediate representations; and (3) a lightweight Refusal Head that selectively intervenes to modulate generation when safety risks are detected. Experiments on LLaVA-OV and Chameleon across diverse safety-critical benchmarks demonstrate that AutoSteer significantly reduces the Attack Success Rate (ASR) for textual, visual, and cross-modal threats, while maintaining general abilities. These findings position AutoSteer as a practical, interpretable, and effective framework for safer deployment of multimodal AI systems. Lyucheng Wu, Ziwen Xu, Tri Cao, Nay Oo, Bryan Hooi, Shumin Deng |
EMNLP | 4 |
| 2025 | GuardReasoner-VL: Safeguarding VLMs via Reinforced ReasoningabstractTo enhance the safety of VLMs, this paper introduces a novel reasoning-based VLM guard model dubbed GuardReasoner-VL. The core idea is to incentivize the guard model to deliberatively reason before making moderation decisions via online RL.
First, we construct GuardReasoner-VLTrain, a reasoning corpus with 123K samples and 631K reasoning steps, spanning text, image, and text-image inputs.
Then, based on it, we cold-start our model's reasoning ability via SFT.
In addition, we further enhance reasoning regarding moderation through online RL.
Concretely, to enhance diversity and difficulty of samples, we conduct rejection sampling followed by data augmentation via the proposed safety-aware data concatenation.
Besides, we use a dynamic clipping parameter to encourage exploration in early stages and exploitation in later stages.
To balance performance and token efficiency, we design a length-aware safety reward that integrates accuracy, format, and token cost.
Extensive experiments demonstrate the superiority of our model.
Remarkably, it surpasses the runner-up by 19.27% F1 score on average, as shown in Figure 1.
We release data, code, and models (3B/7B) of GuardReasoner-VL: https://github.com/yueliu1999/GuardReasoner-VL. Yue Liu 0008, Shengfang Zhai, Mingzhe Du, Tri Cao, Hongcheng Gao, Xinfeng Li, Kun Wang 0056, Junfeng Fang, Jiaheng Zhang, Bryan Hooi |
NeurIPS | 5 |
| 2024 | KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection
Yuexin Li, Chengyu Huang 0003, Shumin Deng, Mei Lin Lock, Tri Cao, Nay Oo, Hoon Wei Lim, Bryan Hooi |
USENIX Security Symposium | 5 |
| 2023 | Joint Self-Supervised Image-Volume Representation Learning with Intra-inter Contrastive ClusteringabstractCollecting large-scale medical datasets with fully annotated samples for training of deep networks is prohibitively expensive, especially for 3D volume data. Recent breakthroughs in self-supervised learning (SSL) offer the ability to overcome the lack of labeled training samples by learning feature representations from unlabeled data. However, most current SSL techniques in the medical field have been designed for either 2D images or 3D volumes. In practice, this restricts the capability to fully leverage unlabeled data from numerous sources, which may include both 2D and 3D data. Additionally, the use of these pre-trained networks is constrained to downstream tasks with compatible data dimensions. In this paper, we propose a novel framework for unsupervised joint learning on 2D and 3D data modalities. Given a set of 2D images or 2D slices extracted from 3D volumes, we construct an SSL task based on a 2D contrastive clustering problem for distinct classes. The 3D volumes are exploited by computing vectored embedding at each slice and then assembling a holistic feature through deformable self-attention mechanisms in Transformer, allowing incorporating long-range dependencies between slices inside 3D volumes. These holistic features are further utilized to define a novel 3D clustering agreement-based SSL task and masking embedding prediction inspired by pre-trained language models. Experiments on downstream tasks, such as 3D brain segmentation, lung nodule detection, 3D heart structures segmentation, and abnormal chest X-ray detection, demonstrate the effectiveness of our joint 2D and 3D SSL approach. We improve plain 2D Deep-ClusterV2 and SwAV by a significant margin and also surpass various modern 2D and 3D SSL approaches. Duy M. H. Nguyen, Truong Thanh Nhat Mai, Tri Cao, Binh T. Nguyen 0001, Nhat Ho, Paul Swoboda, Shadi Albarqouni, Pengtao Xie, Daniel Sonntag |
AAAI | 4 |
| 2023 | Anomaly Detection under Distribution ShiftabstractAnomaly detection (AD) is a crucial machine learning task that aims to learn patterns from a set of normal training samples to identify abnormal samples in test data. Most existing AD studiesassume that the training and test data are drawn from the same data distribution, but the test data can have large distribution shifts arising in many real-world applications due to different natural variations such as new lighting conditions, object poses, or background appearances, rendering existing AD methods ineffective in such cases. In this paper, we consider the problem of anomaly detection under distribution shift and establish performance benchmarks on four widely-used AD and out-of-distribution (OOD) generalization datasets. We demonstrate that simple adaptation of state-of-the-art OOD generalization methods to AD settings fails to work effectively due to the lack of labeled anomaly data. We further introduce a novel robust AD approach to diverse distribution shifts by minimizing the distribution gap between in-distribution and OOD normal samples in both the training and inference stages in an unsupervised way. Our extensive empirical results on the four datasets show that our approach substantially outperforms state-of-the-art AD methods and OOD generalization methods on data with various distribution shifts, while maintaining the detection accuracy on in-distribution data. Code and data are available at https://github.com/mala-lab/ADShift. Tri Cao, Jiawen Zhu 0001, Guansong Pang |
ICCV | 1 |
| 2023 | LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph MatchingabstractObtaining large pre-trained models that can be fine-tuned to new tasks with limited annotated samples has remained an open challenge for medical imaging data. While pre-trained networks on ImageNet and vision-language foundation models trained on web-scale data are the prevailing approaches, their effectiveness on medical tasks is limited due to the significant domain shift between natural and medical images. To bridge this gap, we introduce LVM-Med, the first family of deep networks trained on large-scale medical datasets. We have collected approximately 1.3 million medical images from 55 publicly available datasets, covering a large number of organs and modalities such as CT, MRI, X-ray, and Ultrasound. We benchmark several state-of-the-art self-supervised algorithms on this dataset and propose a novel self-supervised contrastive learning algorithm using a graph-matching formulation. The proposed approach makes three contributions: (i) it integrates prior pair-wise image similarity metrics based on local and global information; (ii) it captures the structural constraints of feature embeddings through a loss function constructed through a combinatorial graph-matching objective, and (iii) it can be trained efficiently end-to-end using modern gradient-estimation techniques for black-box solvers. We thoroughly evaluate the proposed LVM-Med on 15 downstream medical tasks ranging from segmentation and classification to object detection, and both for the in and out-of-distribution settings. LVM-Med empirically outperforms a number of state-of-the-art supervised, self-supervised, and foundation models. For challenging tasks such as Brain Tumor Classification or Diabetic Retinopathy Grading, LVM-Med improves previous vision-language models trained on 1 billion masks by 6-7% while using only a ResNet-50. Duy M. H. Nguyen, Nghiem Tuong Diep, Tan Ngoc Pham, Tri Cao, Binh T. Nguyen 0001, Paul Swoboda, Nhat Ho, Shadi Albarqouni, Pengtao Xie, Daniel Sonntag, Mathias Niepert |
NeurIPS | 5 |
| 2022 | Ensemble approaches for Test Case Prioritization in UI testingabstractTest case prioritization, which focuses on ranking test cases, is an important activity in software engineering given a large number of test cases to be executed within a short period of time.Recent approaches use test execution history and test coverage as the key information for ranking prediction while reinforcement learning has the potential for improving the accuracy of prioritization.Still, each approach has its own advantages and limitations.This paper proposes a ensemble method to take advantages of several existing models by combining different them into a single one.We evaluate our ensemble models on the data sets, including sixteen projects.The results show that one of our proposed models outperforms all single models on 12 over 16 data sets. Tri Cao, Tuan Vu, Huyen Le, Vu Nguyen 0003 |
SEKE | 1 |
| 2017 | Front-end-of-line attacks in split manufacturingabstractBy splitting the manufacturing of integrated circuits into back-end-of-line (BEOL) and front-end-of-line (FEOL) at different foundries, the vulnerabilities to attacks by a untrusted foundry is considerably alleviated. Most previous works focus on the scenario of only BEOL attacks at untrusted FEOL foundries. In this work, we study a largely unexplored scenario, where FEOL attacks are launched by an untrusted BEOL foundry. A geometric pattern match attack and a machine learningbased attack technique are investigated. Defense techniques against the FEOL attacks are also discussed. The effectiveness of these techniques is demonstrated by experiments on benchmark circuits. Tri Cao, Jiang Hu 0001, Jeyavijayan Rajendran |
ICCAD | 2 |