EDBT 2026 Demo / reviewers in the wild / expert
Son Tran
dblp:19/2438
· DBLP profile ↗
19ranked-venue papers
0as first author
15since 2021 · last 2026
0009-0001-9206-2916ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SMPRO: Self-Supervised Visual Preference Alignment via Differentiable Multi-Preference Multi-Group RankingabstractDirect Preference Optimization (DPO) has emerged as a simple and effective approach for aligning models with human preferences. However, existing DPO-based methods suffer from 3 key drawbacks: they rely on only a single positive-negative preference pair per question, restricting the diversity and richness of feedback; they often emphasize minimizing negative preference scores while neglecting to strengthen the positive preferences; and they depend on either human-annotated preferences or expert model outputs - both expensive and difficult to scale. Moreover, the deterministic ranking assumptions of recent Group-based preference optimization methods break down in open-ended tasks such as Visual Question Answering (VQA), where multiple answers can be equally plausible but differ subtly in relevance or specificity. Given this subtle variance in preferences, we propose to perform ranking over groups of preferences rather than relying on fine-grained ranking of individual ones, which is often noisy and subjective. To address these challenges, we introduce Self-Supervised Visual Preference Alignment via Differentiable Multi-Preference Multi-Group Ranking (SMPRO), a novel framework that (1) self-generates rich, diverse preference groups while eliminating the need for external annotations, (2) employs a fully differentiable ranking objective based on sorting networks to capture nuanced preference gradients across arbitrary numbers of preferences both within and across these groups, and (3) incorporates multiple positive preferences to enrich the positive preference group, capturing subtle distinctions among high-quality preferences. Extensive experiments across diverse visual tasks show that our approach achieves state-of-the-art performance in self-supervised setting. Specifically, our model surpasses existing baselines, achieving notable gains such as 82.4% on MM-Bench, 63.2% on MMStar, 94.6% on LLaVA-W, and 81.9% on AI2D. These results underscore the effectiveness of our approach in capturing richer preference signals and demonstrate its scalability for open-ended, ambiguous VQA tasks. Sirnam Swetha, Shwetha Ram, Tal Neiman, Son Tran, Mubarak Shah |
AAAI | 5 |
| 2026 | ViLL-E: Video LLM Embeddings for RetrievalabstractRohit Gupta, Jayakrishnan Unnikrishnan, Fan Fei, Sheng Liu, Son Tran, Mubarak Shah. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rohit Gupta 0012, Jayakrishnan Unnikrishnan, Fei Fan 0002, Son Tran, Mubarak Shah |
ACL (1) | 5 |
| 2025 | M-LLM Based Video Frame Selection for Efficient Video UnderstandingabstractRecent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames that are fed into an M-LLM, particularly for long context videos. However, it could lose crucial context in certain periods of a video, so that the downstream M-LLM may not have sufficient visual information to answer a question. To attack this pain point, we propose a light-weight M-LLM-based frame selection method that adaptively select frames that are more relevant to users’ queries. In order to train the proposed frame selector, we introduce two supervision signals (i) Spatial signal, where single frame importance score by prompting an M-LLM; (ii) Temporal signal, in which multiple frames selection by prompting Large Language Model (LLM) using the captions of all frame candidates. The selected frames are then digested by a frozen downstream video M-LLM for visual reasoning and question answering. Empirical results show that the proposed M-LLM video frame selector improves the performances various downstream video Large Language Model (video-LLM) across medium (ActivityNet, NExT-QA) and long (EgoSchema, LongVideoBench) context video question answering benchmarks. Kai Hu 0010, Xiaohan Nie, Son Tran, Tal Neiman, Lingyun Wang 0005, Mubarak Shah, Raffay Hamid, Trishul Chilimbi |
CVPR | 5 |
| 2025 | CoLLM: A Large Language Model for Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a complex task that aims to retrieve images based on a multimodal query. Typical training data consists of triplets containing a reference image, a textual description of desired modifications, and the target image, which are expensive and time-consuming to acquire. The scarcity of CIR datasets has led to zero-shot approaches utilizing synthetic triplets or leveraging vision-language models (VLMs) with ubiquitous web-crawled image-caption pairs. However, these methods have significant limitations: synthetic triplets suffer from limited scale, lack of diversity, and unnatural modification text, while image-caption pairs hinder joint embedding learning of the multimodal query due to the absence of triplet data. Moreover, existing approaches struggle with complex and nuanced modification texts that demand sophisticated fusion and understanding of vision and language modalities. We present CoLLM, a one-stop framework that effectively addresses these limitations. Our approach generates triplets on-the-fly from image-caption pairs, enabling supervised training without manual annotation. We leverage Large Language Models (LLMs) to generate joint embeddings of reference images and modification texts, facilitating deeper multimodal fusion. Additionally, we introduce Multi-Text CIR (MTCIR), a large-scale dataset comprising 3.4M samples, and refine existing CIR benchmarks (CIRR and Fashion-IQ) to enhance evaluation reliability. Experimental results demonstrate that CoLLM achieves state-of-the-art performance across multiple CIR benchmarks and settings. MTCIR yields competitive results, with up to 15% performance improvement. Our refined benchmarks provide more reliable evaluation metrics for CIR models, contributing to the advancement of this important field. Project page is at collm-cvpr25.github.io. Chuong Huynh, Ashish Tawari, Mubarak Shah, Son Tran, Raffay Hamid, Trishul Chilimbi, Abhinav Shrivastava |
CVPR | 5 |
| 2025 | A Methodology for Incompleteness-Tolerant and Modular Gradual Semantics for Argumentative Statement GraphsabstractGradual semantics (GS) have demonstrated great potential in argumentation, in particular for deploying quantitative bipolar argumentation frameworks (QBAFs) in a number of real-world settings, from judgmental forecasting to explainable AI. In this paper, we provide a novel methodology for obtaining GS for statement graphs, a form of structured argumentation framework, where arguments and relations between them are built from logical statements. Our methodology differs from existing approaches in the literature in two main ways. First, it naturally accommodates incomplete information, so that arguments with partially specified premises can play a meaningful role in the evaluation. Second, it is modularly defined to leverage on any GS for QBAFs. We also define a set of novel properties for our GS and study their suitability alongside a set of existing properties (adapted to our setting) for two instantiations of our GS, demonstrating their advantages over existing approaches. Antonio Rago 0001, Stylianos Loukas Vasileiou, Son Tran, Francesca Toni, William Yeoh 0001 |
KR | 3 |
| 2025 | DreamBlend: Advancing Personalized Fine-Tuning of Text-to-Image Diffusion ModelsabstractGiven a small number of images of a subject, personalized image generation techniques can fine-tune large pretrained text-to-image diffusion models to generate images of the subject in novel contexts, conditioned on text prompts. In doing so, a tradeoff is made between prompt fidelity, subject fidelity and diversity. As the pretrained model is fine-tuned, earlier checkpoints synthesize images with low subject fidelity but high prompt fidelity and diversity. In contrast, later checkpoints generate images with low prompt fidelity and diversity but high subject fidelity. This inherent tradeoff limits the prompt fidelity, subject fidelity and diversity of generated images. In this work, we propose DreamBlend to combine the prompt fidelity from earlier checkpoints and the subject fidelity from later checkpoints during inference. We perform a cross attention guided image synthesis from a later checkpoint, guided by an image Shwetha Ram, Tal Neiman, Qianli Feng, Son Tran, Trishul Chilimbi |
WACV | 5 |
| 2024 | VidLA: Video-Language Alignment at ScaleabstractIn this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do not capture both short-range and long-range temporal dependencies and typically employ complex hierarchical deep network architectures that are hard to integrate with existing pretrained image-text foundation models. To effectively address this limitation, we instead keep the network architecture simple and use a set of data tokens that operate at different temporal resolutions in a hierarchical manner, accounting for the temporally hierarchical nature of videos. By employing a simple two-tower architecture, we are able to initialize our video-language model with pretrained image-text foundation models, thereby boosting the final performance. Second, existing video-language alignment works struggle due to the lack of semantically aligned large-scale training data. To overcome it, we leverage recent LLMs to curate the largest video-language dataset to date with better visual grounding. Furthermore, unlike existing video-text datasets which only contain short clips, our dataset is enriched with video clips of varying durations to aid our temporally hierarchical data to-kens in extracting better representations at varying temporal scales. Overall, empirical results show that our proposed approach surpasses state-of-the-art methods on Multiple retrieval benchmarks, especially on longer videos, and performs competitively on classification benchmarks. Mamshad Nayeem Rizve, Fei Fan 0002, Jayakrishnan Unnikrishnan, Son Tran, Benjamin Z. Yao, Belinda Zeng, Mubarak Shah, Trishul Chilimbi |
CVPR | 4 |
| 2024 | Open Vocabulary Multi-label Video Classification
Rohit Gupta 0012, Mamshad Nayeem Rizve, Jayakrishnan Unnikrishnan, Ashish Tawari, Son Tran, Mubarak Shah, Benjamin Z. Yao, Trishul Chilimbi |
ECCV (39) | 5 |
| 2024 | X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
Sirnam Swetha, Tal Neiman, Mamshad Nayeem Rizve, Son Tran, Benjamin Z. Yao, Trishul Chilimbi, Mubarak Shah |
ECCV (6) | 5 |
| 2024 | Bringing Multimodality to Amazon Visual Search SystemabstractImage to image matching has been well studied in the computer vision community. Previous studies mainly focus on training a deep metric learning model matching visual patterns between the query image and gallery images. In this study, we show that pure image-to- image matching suffers from false positives caused by matching to local visual patterns. To alleviate this issue, we propose to leverage recent advances in vision-language pretraining research. Specifically, we introduce additional image-text alignment losses into deep metric learning, which serve as constraints to the image-to-image matching loss. With additional alignments between the text (e.g., product title) and image pairs, the model can learn concepts from both modalities explicitly, which avoids matching low-level visual features. We progressively develop two variants, a 3-tower and a 4-tower model, where the latter takes one more short text query input. Through extensive experiments, we show that this change leads to a substantial improvement to the image to image matching problem. We further leveraged this model for multimodal search, which takes both image and reformulation text queries to improve search quality. Both offline and online experiments show strong improvements on the main metrics. Specifically, we see 4.95% relative improvement on image matching click through rate with the 3-tower model and 1.13% further improvement from the 4-tower model. Xinliang Zhu, Sheng-Wei Huang, Han Ding 0004, Kelvin Chen, Tal Neiman, Ouye Xie, Son Tran, Benjamin Z. Yao, Douglas Gray 0001, Anuj Bindal, Arnab Dhua |
KDD | 9 |
| 2023 | A Logic-based Explanation Generation Framework for Classical and Hybrid Planning Problems (Extended Abstract)abstractIn human-aware planning systems, a planning agent might need to explain its plan to a human user when that plan appears to be non-feasible or sub-optimal. A popular approach, called model reconciliation, has been proposed as a way to bring the model of the human user closer to the agent's model. In this paper, we approach the model reconciliation problem from a different perspective, that of knowledge representation and reasoning, and demonstrate that our approach can be applied not only to classical planning problems but also hybrid systems planning problems with durative actions and events/processes. Stylianos Loukas Vasileiou, William Yeoh 0001, Son Tran, Ashwin Kumar, Michael Cashmore, Daniele Magazzeni |
IJCAI | 3 |
| 2022 | Multi-modal Alignment using Representation CodebookabstractAligning signals from different modalities is an important step in vision-language representation learning as it affects the performance of later stages such as cross-modality fusion. Since image and text typically reside in different regions of the feature space, directly aligning them at instance level is challenging especially when features are still evolving during training. In this paper, we propose to align at a higher and more stable level using cluster representation. Specifically, we treat image and text as two “views” of the same entity, and encode them into a joint vision-language coding space spanned by a dictionary of cluster centers (codebook). We contrast positive and negative samples via their cluster assignments while simultaneously optimizing the cluster centers. To further smooth out the learning process, we adopt a teacher-student distillation paradigm, where the momentum teacher of one view guides the student learning of the other. We evaluated our approach on common vision language benchmarks and obtain new SoTA on zero-shot cross modality retrieval while being competitive on various other transfer tasks. Jiali Duan, Liqun Chen 0001, Son Tran, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi |
CVPR | 3 |
| 2022 | Vision-Language Pre-Training with Triple Contrastive LearningabstractVision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information (MI) between an image and its matched text. However, simply performing cross-modal alignment (CMA) ignores data potential within each modality, which may result in degraded representations. For instance, although CMA-based models are able to map image-text pairs close together in the embedding space, they fail to ensure that similar inputs from the same modality stay close by. This problem can get even worse when the pre-training data is noisy. In this paper, we propose triple contrastive learning (TCL) for vision-language pre-training by leveraging both cross-modal and intra-modal self-supervision. Besides CMA, TCL introduces an intra-modal contrastive objective to provide complementary benefits in representation learning. To take advantage of localized and structural information from image and text input, TCL further maximizes the average MI between local regions of image/text and their global summary. To the best of our knowledge, ours is the first work that takes into account local structure information for multi-modality representation learning. Experimental evaluations show that our approach is competitive and achieves the new state of the art on various common downstream vision-language tasks such as image-text retrieval and visual question answering. Jiali Duan, Son Tran, Yi Xu 0011, Sampath Chanda, Liqun Chen 0001, Belinda Zeng, Trishul Chilimbi, Junzhou Huang |
CVPR | 3 |
| 2022 | Amazon Shop the Look: A Visual Search System for Fashion and HomeabstractIn this paper, we introduce Shop the Look, a web-scale fashion and home product visual search system deployed at Amazon. Building such a system poses great challenges to both science and engineering practices. We leverage large-scale image data from the Amazon product catalog and adopt effective strategies to reduce the human effort required to annotate data. By employing state-of-the-art computer vision techniques, we train detection, recognition, and feature extraction models to bridge the domain gap between in-the-wild query images and product images which are taken under controlled settings. Our system is designed to achieve a balance between result accuracy and efficiency. The run-time service is optimized to provide retrieval results to users with low-latency. The scalable offline index-building pipeline adapts to the dynamic Amazon catalog that contains billions of products. We present both quantitative and qualitative evaluation results to demonstrate the performance of our system. We believe that the fast-growing Shop the Look service is shaping the way that customers shop on Amazon. Arnau Ramisa, Amit Kumar K. C, Sampath Chanda, Mengjiao Wang 0002, Neelakandan Rajesh, Shasha Li 0001, Yingchuan Hu, Nagashri Lakshminarayana, Son Tran, Douglas Gray 0001 |
KDD | 11 |
| 2022 | Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias PerspectiveabstractContrastive learning (CL) has been the de facto technique for self-supervised representation learning (SSL), with impressive empirical success such as multi-modal representation learning. However, traditional CL loss only considers negative samples from a minibatch, which could cause biased gradients due to the non-decomposibility of the loss. For the first time, we consider optimizing a more generalized contrastive loss, where each data sample is associated with an infinite number of negative samples. We show that directly using minibatch stochastic optimization could lead to gradient bias. To remedy this, we propose an efficient Bayesian data augmentation technique to augment the contrastive loss into a decomposable one, where standard stochastic optimization can be directly applied without gradient bias. Specifically, our augmented loss defines a joint distribution over the model parameters and the augmented parameters, which can be conveniently optimized by a proposed stochastic expectation-maximization algorithm. Our framework is more general and is related to several popular SSL algorithms. We verify our framework on both small scale models and several large foundation models, including SSL of ImageNet and SSL for vision-language representation learning. Experiment results indicate the existence of gradient bias in all cases, and demonstrate the effectiveness of the proposed method on improving previous state of the arts. Remarkably, our method can outperform the strong MoCo-v3 under the same hyper-parameter setting with only around half of the minibatch size; and also obtains strong results in the recent public benchmark ELEVATER for few-shot image classification. Changyou Chen, Yi Xu 0011, Liqun Chen 0001, Jiali Duan, Yiran Chen 0001, Son Tran, Belinda Zeng, Trishul Chilimbi |
NeurIPS | 7 |
| 2019 | Reinforcement Learning Framework to Identify Cause of Diseases - Predicting Asthma Attack CaseabstractAsthma attack prediction is a highly challenging problem because of the dynamic and multi-factor nature of its etiology. Disease severity level, physiological measurements, patient behaviors and characteristics, environmental triggers, and Personal Risk Scores (RS) are among the strong predictors of an asthma attack. In this paper, we propose a Deep Reinforcement Leaning framework to predict asthma attacks using historical data linking the severity level of the disease and the personalized risk scores of triggers. Deep Q-learning based prediction framework can model future reward explicitly. Besides, the risk scores of triggers which are calculated using Additive Interaction Analysis of Exposures technique helps increase the prediction performance. The main purpose of this study is to investigate the ability of using Q-learning method to create a prediction model that would help asthmatic individuals to take evasive action when the probability of an attack was at their personal threshold levels. Quan T. Do, Son Tran, Alexa K. Doig |
IEEE BigData | 2 |
| 2019 | A Taxonomy for Selecting Wearable Input Devices for Mixed RealityabstractComposite wearable computers consist of multiple wearable devices connected together and working as a cohesive whole. These composite wearable computers are promising for augmenting our interaction with the physical, virtual, and mixed play spaces (e.g., mixed reality games). Yet little research has directly addressed how mixed reality system designers can select wearable input devices and how these devices can be assembled together to form a cohesive wearable computer. We present an initial taxonomy of wearable input devices to aid designers in deciding which devices to select and assemble together to support different mixed reality systems. We undertook a grounded theory analysis of 84 different wearable input devices resulting in a design taxonomy for composite wearable computers. The taxonomy consists of two axes: TYPE OF INTERACTIVITY and BODY LOCATION. These axes enable designers to identify which devices fill particular needs in the system development process and how these devices can be assembled together to form a cohesive wearable computer. Ahmed S. Khalaf, Sultan A. Alharthi, Son Tran, Igor Dolgov, Phoebe O. Toups Dugas |
ISS | 3 |
| 2018 | Bidding in Periodic Double Auctions Using Heuristics and Dynamic Monte Carlo Tree SearchabstractIn a Periodic Double Auction (PDA), there are multiple discrete trading periods for a single type of good. PDAs are commonly used in real-world energy markets to trade energy in specific time slots to balance demand on the power grid. Strategically, bidding in a PDA is complicated because the bidder must predict and plan for future auctions that may influence the bidding strategy for the current auction. We present a general bidding strategy for PDAs based on forecasting clearing prices and using Monte Carlo Tree Search (MCTS) to plan a bidding strategy across multiple time periods. In addition, we present a fast heuristic strategy that can be used either as a standalone method or as an initial set of bids to seed the MCTS policy. We evaluate our bidding strategies using a PDA simulator based on the wholesale market implemented in the Power Trading Agent Competition (PowerTAC) competition. We demonstrate that our strategies outperform state-of-the-art bidding strategies designed for that competition. Moinul Morshed Porag Chowdhury, Christopher Kiekintveld, Son Tran, William Yeoh 0001 |
IJCAI | 3 |
| 2013 | Object-Oriented Knowledge Bases in Logic Programming
Vinay K. Chaudhri, Stijn Heymans, Son Tran, Michael A. Wessel |
Theory Pract. Log. Program. | 3 |