Quan Tran

dblp:28/8930 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
6since 2021 · last 2026
0009-0005-7360-9639ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 TriaGS: Differentiable Triangulation-Guided Geometric Consistency for 3D Gaussian Splatting
abstract
3D Gaussian Splatting is crucial for real-time novel view synthesis due to its efficiency and ability to render photorealistic images. However, building a 3D Gaussian is guided solely by photometric loss, which can result in inconsistencies in reconstruction. This under-constrained process often results in "floater" artifacts and unstructured geometry, preventing the extraction of high-fidelity surfaces. To address this issue, our paper introduces a novel method that improves reconstruction by enforcing global geometry consistency through constrained multi-view triangulation. Our approach aims to achieve a consensus on 3D representation in the physical world by utilizing various estimated views. We optimize this process by penalizing the deviation of a rendered 3D point from a robust consensus point, which is re-triangulated from a bundle of neighboring views in a self-supervised fashion. We demonstrate the effectiveness of our method across multiple datasets, achieving state-of-the-art results. On the DTU dataset, our method attains a mean Chamfer Distance of 0.50 mm, outperforming comparable explicit methods. We will make our code open-source to facilitate community validation and ensure reproducibility.
Quan Tran, Tuan Dang
WACV1
2024 Large Language Model Prompting with Episodic Memory
abstract
Prompt optimization is essential for enhancing the performance of Large Language Models (LLMs) in a range of Natural Language Processing (NLP) tasks, particularly in scenarios of few-shot learning where training examples are incorporated directly into the prompt. Despite the growing interest in optimizing prompts with few-shot examples, existing methods for prompt optimization are often resource-intensive or perform inadequately. In this work, we propose PrOmpting with Episodic Memory (POEM), a novel prompt optimization technique that is simple, efficient, and demonstrates strong generalization capabilities. We approach prompt optimization as a Reinforcement Learning (RL) challenge, using episodic memory to archive combinations of input data, permutations of few-shot examples, and the rewards observed during training. In the testing phase, we optimize the sequence of examples for each test query by selecting the sequence that yields the highest total rewards from the top-k most similar training examples in the episodic memory. Our results show that POEM outperforms recent techniques like TEMPERA and RLPrompt by over 5.3% in various text classification tasks. Furthermore, our approach adapts well to broader language understanding tasks, consistently outperforming conventional heuristic methods for ordering examples.
Van Dai Do, Quan Tran, Svetha Venkatesh, Hung Le 0002
ECAI2
2023 Boosting Punctuation Restoration with Data Generation and Reinforcement Learning
Viet Dac Lai, Abel Salinas, Hao Tan 0002, Trung Bui, Quan Tran, Seunghyun Yoon 0002, Hanieh Deilamsalehy, Franck Dernoncourt, Thien Huu Nguyen
INTERSPEECH5
2022 Improving Closed and Open-Vocabulary Attribute Prediction Using Transformers
Khoi Pham, Kushal Kafle, Zhe Lin 0001, Zhihong Ding, Scott Cohen, Quan Tran, Abhinav Shrivastava
ECCV (25)6
2021 Learning To Predict Visual Attributes in the Wild
abstract
Visual attributes constitute a large portion of information contained in a scene. Objects can be described using a wide variety of attributes which portray their visual appearance (color, texture), geometry (shape, size, posture), and other intrinsic properties (state, action). Existing work is mostly limited to study of attribute prediction in specific domains. In this paper, we introduce a large-scale in-the-wild visual attribute prediction dataset consisting of over 927K attribute annotations for over 260K object instances. Formally, object attribute prediction is a multi-label classification problem where all attributes that apply to an object must be predicted. Our dataset poses significant challenges to existing methods due to large number of attributes, label sparsity, data imbalance, and object occlusion. To this end, we propose several techniques that systematically tackle these challenges, including a base model that utilizes both low- and high-level CNN features with multi-hop attention, reweighting and resampling techniques, a novel negative label expansion scheme, and a novel supervised attribute-aware contrastive learning algorithm. Using these techniques, we achieve near 3.7 mAP and 5.7 overall F1 points improvement over the current state of the art. Further details about the VAW dataset can be found at https://vawdataset.com/
Khoi Pham, Kushal Kafle, Zhe Lin 0001, Zhihong Ding, Scott Cohen, Quan Tran, Abhinav Shrivastava
CVPR6
2021 Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real Images
abstract
While neural symbolic methods demonstrate impressive performance in visual question answering on synthetic images, their performance suffers on real images. We identify that the long-tail distribution of visual concepts and unequal importance of reasoning steps in real data are the two key obstacles that limit the models’ real-world potentials. To address these challenges, we propose a new paradigm, Calibrating Concepts and Operations (CCO), which enables neural symbolic models to capture underlying data characteristics and to reason with hierarchical importance. Specifically, we introduce an executor with learnable concept embedding magnitudes for handling distribution imbalance, and an operation calibrator for highlighting important operations and suppressing redundant ones.Our experiments show CCO substantially boosts the performance of neural symbolic methods on real images. By evaluating models on the real world dataset GQA, CCO helps the neural symbolic method NSCL outperforms its vanilla counterpart by 9.1% (from 47.0% to 56.1%); this result also largely reduces the performance gap between symbolic and non-symbolic methods. Additionally, we create a perturbed test set for better understanding and analyzing model performance on real images. Code is available at https://lizw14.github.io/project/ccosr.
Zhuowan Li, Elias Stengel-Eskin, Yixiao Zhang 0001, Cihang Xie, Quan Tran, Benjamin Van Durme, Alan L. Yuille
ICCV5
2020 Context-Aware Group Captioning via Self-Attention and Contrastive Features
abstract
While image captioning has progressed rapidly, existing works focus mainly on describing single images. In this paper, we introduce a new task, context-aware group captioning, which aims to describe a group of target images in the context of another group of related reference images. Context-aware group captioning requires not only summarizing information from both the target and reference image group but also contrasting between them. To solve this problem, we propose a framework combining self-attention mechanism with contrastive feature construction to effectively summarize common information from each image group while capturing discriminative information between them. To build the dataset for this task, we propose to group the images and generate the group captions based on single image captions using scene graphs matching. Our datasets are constructed on top of the public Conceptual Captions dataset and our new Stock Captions dataset. Experiments on the two datasets show the effectiveness of our method on this new task.
Zhuowan Li, Quan Tran, Long Mai, Zhe Lin 0001, Alan L. Yuille
CVPR2
2020 Open-Edit: Open-Domain Image Manipulation with Open-Vocabulary Instructions
Xihui Liu, Zhe Lin 0001, Jianming Zhang 0001, Handong Zhao, Quan Tran, Xiaogang Wang 0001, Hongsheng Li 0001
ECCV (11)5
2017 Named Entity Recognition with Stack Residual LSTM and Trainable Bias Decoding
abstract
Recurrent Neural Network models are the state-of-the-art for Named Entity Recognition (NER). We present two innovations to improve the performance of these models. The first innovation is the introduction of residual connections between the Stacked Recurrent Neural Network model to address the degradation problem of deep neural networks. The second innovation is a bias decoding mechanism that allows the trained system to adapt to non-differentiable and externally computed objectives, such as the entity-based F-measure. Our work improves the state-of-the-art results for both Spanish and English languages on the standard train/development/test split of the CoNLL 2003 Shared Task NER dataset.
Quan Tran, Andrew MacKinlay, Antonio Jimeno-Yepes
IJCNLP(1)1
2014 Online maneuver recognition and multimodal trajectory prediction for intersection assistance using non-parametric regression
abstract
Maneuver recognition and trajectory prediction of moving vehicles are two important and challenging tasks of advanced driver assistance systems (ADAS) at urban intersections. This paper presents a continuing work to handle these two problems in a consistent framework using non-parametric regression models. We provide a feature normalization scheme and present a strategy for constructing three-dimensional Gaussian process regression models from two-dimensional trajectory patterns These models can capture spatio-temporal characteristics of traffic situations. Given a new, partially observed and unlabeled trajectory, the maneuver can be recognized online by comparing the likelihoods of the observation data for each individual regression model. Furthermore, we take advantage of our representation for trajectory prediction. Because predicting possible trajectories at urban intersection involves obvious multimodalities and non-linearities, we employ the Monte Carlo method to handle these difficulties. This approach allows the incremental prediction of possible trajectories in situations where unimodal estimators such as Kalman Filters would not work well. The proposed framework is evaluated experimentally in urban intersection scenarios using real-world data.
Quan Tran, Jonas Firl
Intelligent Vehicles Symposium1
2013 Modelling of traffic situations at urban intersections with probabilistic non-parametric regression
abstract
Driving intention recognition and trajectory prediction of moving vehicles are two important requirements of future advanced driver assistance systems (ADAS) for urban intersections. In this paper, we present a consistent framework for solving these two problems. The key idea is to model the spatio-temporal dependencies of traffic situations with a two-dimensional Gaussian process regression. With this representation the driving intention can be recognized by evaluating the data likelihood for each individual regression model. For the trajectory prediction purpose, we transform these regression models into the corresponding dynamical models and combine them with Unscented Kalman Filters (UKF) to overcome the non-linear issue. We evaluate our framework with data collected from real traffic scenarios and show that our approach can be used for recognition of different driving intentions and for long-term trajectory prediction of traffic situations occurring at urban intersections.
Quan Tran, Jonas Firl
Intelligent Vehicles Symposium1
2012 A probabilistic discriminative approach for situation recognition in traffic scenarios
abstract
Understanding of traffic situations is an essential part of future advanced driver assistance systems (ADAS). This has to handle spatio-temporal dependencies of multiple traffic participants and uncertainties from different sources. Most existing approaches use probabilistic generative joint structures like Hidden Markov Models (HMM), which have long been used for dealing with activity recognition problems. Two significant limitations of these models are the assumption of conditional independence of observations and the availability of prior information. In this study, we present a probabilistic discriminative approach based on undirected probabilistic graphical models (Markov Networks). We combine two well-studied models: the log-linear model and the Conditional Random Field (CRF), which use dynamic programming for efficient, exact inference and their parameters can be learned via convex optimization. Since CRF conditions on entire observation sequences, we can avoid the requirement of independence between observations. Additionally, with discriminative models prior information of each activity is not necessary when performing a classification step. These two advantages of the discriminative models are very useful for our focusing problem of traffic scene understanding. We evaluate our approach with real data and show that it is able to recognize different driving maneuvers occurring at an urban intersection.
Quan Tran, Jonas Firl
Intelligent Vehicles Symposium1