Hieu Le 0001

dblp:123/2117-1 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0001-7855-2778ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
abstract
Reasoning segmentation seeks pixel-accurate masks for targets referenced by complex, often implicit instructions, requiring context-dependent reasoning over the scene. Recent multimodal language models have advanced instruction following segmentation, yet generalization remains limited. The key bottleneck is the high cost of curating diverse, high-quality pixel annotations paired with rich linguistic supervision leading to brittle performance under distribution shift. Therefore, we present CORA, a semi-supervised reasoning segmentation framework that jointly learns from limited labeled data and a large corpus of unlabeled images. CORA introduces three main components: 1) conditional visual instructions that encode spatial and contextual relationships between objects; 2) a noisy pseudo-label filter based on the consistency of Multimodal LLM’s outputs across semantically equivalent queries; and 3) a token-level contrastive alignment between labeled and pseudo-labeled samples to enhance feature consistency. These components enable CORA to perform robust reasoning segmentation with minimal supervision, outperforming existing baselines under constrained annotation settings. CORA achieves state-of-the-art results, requiring as few as 100 labeled images on Cityscapes, a benchmark dataset for urban scene understanding, surpassing the baseline by +2.3%. Similarly, CORA improves performance by +2.4% with only 180 labeled images on PanNuke, a histopathology dataset.
Prantik Howlader, Hoang Nguyen-Canh, Srijan Das, Hieu Le 0001, Dimitris Samaras
WACV5
2025 Few-shot Personalized Scanpath Prediction
abstract
A personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training scanpath prediction models are data-intensive and cannot be effectively personalized to new individuals with only a few available examples. In this paper, we propose few-shot personalized scanpath prediction task (FS-PSP) and a novel method to address it, which aims to predict scan-paths for an unseen subject using minimal support data of that subject’s scanpath behavior. The key to our method’s adaptability is the Subject-Embedding Network (SE-Net), specifically designed to capture unique, individualized representations for each subject’s scanpaths. SE-Net generates subject embeddings that effectively distinguish between subjects while minimizing variability among scanpaths from the same individual. The personalized scanpath prediction model is then conditioned on these subject embeddings to produce accurate, personalized results. Experiments on multiple eye-tracking datasets demonstrate that our method excels in FS-PSP settings and does not require any fine-tuning steps at test time. Code is available at: https://github.com/cvlab-stonybrook/few-shot-scanpath
Ruoyu Xue, Sounak Mondal, Hieu Le 0001, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras
CVPR4
2025 Counting Stacked Objects
Corentin Dumery, Noa Etté, Aoxiang Fan, Hieu Le 0001, Pascal Fua
ICCV6
2025 Importance-Based Token Merging for Efficient Image and Video Generation
Hieu Le 0001, Dimitris Samaras
ICCV3
2025 Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds
abstract
Text-to-image diffusion models have demonstrated remarkable capability in generating realistic images from arbitrary text prompts. However, they often produce inconsistent results for compositional prompts such as "two dogs" or "a penguin on the right of a bowl". Understanding these inconsistencies is crucial for reliable image generation. In this paper, we highlight the significant role of initial noise in these inconsistencies, where certain noise patterns are more reliable for compositional prompts than others. Our analyses reveal that different initial random seeds tend to guide the model to place objects in distinct image areas, potentially adhering to specific patterns of camera angles and image composition associated with the seed. To improve the model's compositional ability, we propose a method for mining these reliable cases, resulting in a curated training set of generated images without requiring any manual annotation. By fine-tuning text-to-image models on these generated images, we significantly enhance their compositional capabilities. For numerical composition, we observe relative increases of 29.3\% and 19.5\% for Stable Diffusion and PixArt-$\alpha$, respectively. Spatial composition sees even larger gains, with 60.7\% for Stable Diffusion and 21.1\% for PixArt-$\alpha$.
Shuangqi Li, Hieu Le 0001, Mathieu Salzmann
ICLR2
2025 QT-DoG: Quantization-Aware Training for Domain Generalization
abstract
A key challenge in Domain Generalization (DG) is preventing overfitting to source domains, which can be mitigated by finding flatter minima in the loss landscape. In this work, we propose Quantization-aware Training for Domain Generalization (QT-DoG) and demonstrate that weight quantization effectively leads to flatter minima in the loss landscape, thereby enhancing domain generalization. Unlike traditional quantization methods focused on model compression, QT-DoG exploits quantization as an implicit regularizer by inducing noise in model weights, guiding the optimization process toward flatter minima that are less sensitive to perturbations and overfitting. We provide both an analytical perspective and empirical evidence demonstrating that quantization inherently encourages flatter minima, leading to better generalization across domains. Moreover, with the benefit of reducing the model size through quantization, we demonstrate that an ensemble of multiple quantized models further yields superior accuracy than the state-of-the-art DG approaches with no computational or memory overheads. Code is released at: https://saqibjaved1.github.io/QT_DoG/.
Saqib Javed, Hieu Le 0001, Mathieu Salzmann
ICML2
2025 Pairwise-Constrained Implicit Functions for 3D Human Heart Modeling
Hieu Le 0001, Nicolas Talabot, Jiancheng Yang, Pascal Fua
MICCAI (16)1
2025 High Resolution UDF Meshing via Iterative Networks
abstract
Unsigned Distance Fields (UDFs) are a natural implicit representation for open surfaces but, unlike Signed Distance Fields (SDFs), are challenging to triangulate into explicit meshes. This is especially true at high resolutions where neural UDFs exhibit higher noise levels, which makes it hard to capture fine details. Most current techniques perform within single voxels without reference to their neighborhood, resulting in missing surface and holes where the UDF is ambiguous or noisy. We show that this can be remedied by performing several passes and by reasoning on previously extracted surface elements to incorporate neighborhood information. Our key contribution is an iterative neural network that does this and progressively improves surface recovery within each voxel by spatially propagating information from increasingly distant neighbors. Unlike single-pass methods, our approach integrates newly detected surfaces, distance values, and gradients across multiple iterations, effectively correcting errors and stabilizing extraction in challenging regions. Experiments on diverse 3D models demonstrate that our method produces significantly more accurate and complete meshes than existing approaches, particularly for complex geometries, enabling UDF surface extraction at higher resolutions where traditional methods fail.
Federico Stella, Nicolas Talabot, Hieu Le 0001, Pascal Fua
NeurIPS3
2025 Shadow Removal Refinement via Material-Consistent Shadow Edges
abstract
Shadow boundaries can be confused with material boundaries as both exhibit sharp changes in luminance or contrast within a scene. However, shadows do not modify the intrinsic color or texture of surfaces. Therefore, on both sides of shadow edges traversing regions with the same material, the original color and texture should be the same if the shadow is removed properly. These shadow/shadow-free pairs are very useful but difficult-to-collect supervision signals. The crucial contribution of this paper is to learn how to identify those shadow edges that traverse material-consistent regions and how to use them as self-supervision for shadow removal refinement during test time. To achieve this, we fine-tune SAM, an image segmentation foundation model, to produce a shadow-invariant segmentation and then extract material-consistent shadow edges by comparing the SAM segmentation with the shadow mask. Utilizing these shadow edges, we in-troduce color- and texture-consistency losses to enhance the shadow removal process. We demonstrate the effectiveness of our method in improving shadow removal results on more challenging, in-the-wild images, outperforming the state-of-the-art shadow removal methods. Addition-ally, we propose a new metric and an annotated dataset for evaluating the performance of shadow removal methods without the need for paired shadow/shadow-free data. Our code and dataset are available at: https://github.com/cvlab-stonybrook/ShadowRemovalRefine
Shilin Hu, Hieu Le 0001, Shahrukh Athar, Sagnik Das, Dimitris Samaras
WACV2
2025 Learning to Count from Pseudo-Labeled Segmentation
abstract
Class-agnostic counting (CAC) has numerous potential applications across various domains. The goal is to count objects of an arbitrary category during testing, based on only a few annotated exemplars. However, existing methods often count all objects in the image, including those from different categories than the exemplars. To address this issue, we propose localizing the area containing the objects of interest via an exemplar-based segmentation model before counting them. To train this model, we propose a novel method to obtain pseudo-labeled segmentation masks. Specifically, we use an unsupervised image clustering method to generate a set of candidate pseudo object masks, from which we select the optimal one using a pretrained CAC model. We show that the trained segmentation model can effectively localize objects of interest based on the exemplars and prevent the model from counting everything. To properly evaluate the performance of CAC methods in real-world scenarios, we introduce two new benchmarks: a synthetic test set and a new test set of real images containing countable objects from multiple classes. Our proposed method shows a significant advantage over previous CAC methods on these two benchmarks.
Hieu Le 0001, Dimitris Samaras
WACV2
2024 Beyond Pixels: Semi-supervised Semantic Segmentation with a Multi-scale Patch-Based Multi-label Classifier
Prantik Howlader, Srijan Das, Hieu Le 0001, Dimitris Samaras
ECCV (75)3
2024 Weighting Pseudo-labels via High-Activation Feature Index Similarity and Object Detection for Semi-supervised Segmentation
Prantik Howlader, Hieu Le 0001, Dimitris Samaras
ECCV (75)2
2024 Neural Surface Detection for Unsigned Distance Fields
Federico Stella, Nicolas Talabot, Hieu Le 0001, Pascal Fua
ECCV (58)3
2024 Assessing Sample Quality via the Latent Space of Generative Models
Hieu Le 0001, Dimitris Samaras
ECCV (59)2
2024 Enabling Uncertainty Estimation in Iterative Neural Networks
abstract
Turning pass-through network architectures into iterative ones, which use their own output as input, is a well-known approach for boosting performance. In this paper, we argue that such architectures offer an additional benefit: The convergence rate of their successive outputs is highly correlated with the accuracy of the value to which they converge. Thus, we can use the convergence rate as a useful proxy for uncertainty. This results in an approach to uncertainty estimation that provides state-of-the-art estimates at a much lower computational cost than techniques like Ensembles, and without requiring any modifications to the original iterative model. We demonstrate its practical value by embedding it in two application domains: road detection in aerial images and the estimation of aerodynamic properties of 2D and 3D shapes.
Nikita Durasov, Doruk Öner, Jonathan Donier, Hieu Le 0001, Pascal Fua
ICML4
2024 Generating Anatomically Accurate Heart Structures via Neural Implicit Fields
Jiancheng Yang, Ekaterina Sedykh, Jason Ken Adhinarta, Hieu Le 0001, Pascal Fua
MICCAI (1)4
2023 Zero-Shot Object Counting
abstract
Class-agnostic object counting aims to count object instances of an arbitrary class at test time. Current methods for this challenging problem require human-annotated exemplars as inputs, which are often unavailable for novel categories, especially for autonomous systems. Thus, we propose zero-shot object counting (ZSC), a new setting where only the class name is available during test time. Such a counting system does not require human annotators in the loop and can operate automatically. Starting from a class name, we propose a method that can accurately identify the optimal patches which can then be used as counting exemplars. Specifically, we first construct a class prototype to select the patches that are likely to contain the objects of interest, namely class-relevant patches. Furthermore, we introduce a model that can quantitatively measure how suitable an arbitrary patch is as a counting exemplar. By applying this model to all the candidate patches, we can select the most suitable patches as exemplars for counting. Experimental results on a recent class-agnostic counting dataset, FSC-147, validate the effectiveness of our method. Code is available at https://github.com/cvlabstonybrook/zero-shot-counting.
Hieu Le 0001, Vu Nguyen 0004, Viresh Ranjan, Dimitris Samaras
CVPR2
2023 Generating Features with Increased Crop-Related Diversity for Few-Shot Object Detection
abstract
Two-stage object detectors generate object proposals and classify them to detect objects in images. These proposals often do not contain the objects perfectly but overlap with them in many possible ways, exhibiting great variability in the difficulty levels of the proposals. Training a robust classifier against this crop-related variability requires abundant training data, which is not available in few-shot settings. To mitigate this issue, we propose a novel variational autoencoder (VAE) based data generation model, which is capable of generating data with increased croprelated diversity. The main idea is to transform the latent space such latent codes with different norms represent different crop-related variations. This allows us to generate features with increased crop-related diversity in difficulty levels by simply varying the latent norm. In particular, each latent code is rescaled such that its norm linearly correlates with the IoU score of the input crop w.r.t. the ground-truth box. Here the IoU score is a proxy that represents the difficulty level of the crop. We train this VAE model on base classes conditioned on the semantic code of each class and then use the trained model to generate features for novel classes. In our experiments our generated features consistently improve state-of-the-art few-shot object detection methods on the PASCAL VOC and MS COCO datasets.
Hieu Le 0001, Dimitris Samaras
CVPR2
2022 Generating Representative Samples for Few-Shot Classification
abstract
Few-shot learning (FSL) aims to learn new categories with a few visual samples per class. Few-shot class representations are often biased due to data scarcity. To mitigate this issue, we propose to generate visual samples based on semantic embeddings using a conditional variational autoencoder (CVAE) model. We train this CVAE model on base classes and use it to generate features for novel classes. More importantly, we guide this VAE to strictly generate representative samples by removing non-representative samples from the base training set when training the CVAE model. We show that this training scheme enhances the representativeness of the generated samples and therefore, improves the few-shot classification results. Experimental results show that our method improves three FSL baseline methods by substantial margins, achieving state-of-the-art few-shot classification performance on miniImageNet and tieredImageNet datasets for both 1-shot and 5-shot settings. Code is available at: https://github.com/cvlab-stonybrook/fsl-rsvae.
Hieu Le 0001
CVPR2
2022 Physics-Based Shadow Image Decomposition for Shadow Removal
abstract
We propose a novel deep learning method for shadow removal. Inspired by physical models of shadow formation, we use a linear illumination transformation to model the shadow effects in the image that allows the shadow image to be expressed as a combination of the shadow-free image, the shadow parameters, and a matte layer. We use two deep networks, namely SP-Net and M-Net, to predict the shadow parameters and the shadow matte respectively. This system allows us to remove the shadow effects from images. We then employ an inpainting network, I-Net, to further refine the results. We train and test our framework on the most challenging shadow removal dataset (ISTD). Our method improves the state-of-the-art in terms of mean absolute error (MAE) for the shadow area by 20%. Furthermore, this decomposition allows us to formulate a patch-based weakly-supervised shadow removal method. This model can be trained without any shadow- free images (that are cumbersome to acquire) and achieves competitive shadow removal results compared to state-of-the-art methods that are trained with fully paired shadow and shadow-free images. Last, we introduce SBU-Timelapse, a video shadow removal dataset for evaluating shadow removal methods.
Hieu Le 0001, Dimitris Samaras
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Variational Feature Disentangling for Fine-Grained Few-Shot Classification
abstract
Data augmentation is an intuitive step towards solving the problem of few-shot classification. However, ensuring both discriminability and diversity in the augmented samples is challenging. To address this, we propose a feature disentanglement framework that allows us to augment features with randomly sampled intra-class variations while preserving their class-discriminative features. Specifically, we disentangle a feature representation into two components: one represents the intra-class variance and the other encodes the class-discriminative information. We assume that the intra-class variance induced by variations in poses, backgrounds, or illumination conditions is shared across all classes and can be modelled via a common distribution. Then we sample features repeatedly from the learned intra-class variability distribution and add them to the class-discriminative features to get the augmented features. Such a data augmentation scheme ensures that the augmented features inherit crucial class-discriminative features while exhibiting large intra-class variance. Our method significantly outperforms the state-of-the-art methods on multiple challenging fine-grained few-shot image classification benchmarks. Code is available at: https://github.com/cvlab-stonybrook/vfd-iccv21
Hieu Le 0001, Mingzhen Huang, Shahrukh Athar, Dimitris Samaras
ICCV2
2020 From Shadow Segmentation to Shadow Removal
Hieu Le 0001, Dimitris Samaras
ECCV (11)1
2019 Shadow Removal via Shadow Image Decomposition
abstract
We propose a novel deep learning method for shadow removal. Inspired by physical models of shadow formation, we use a linear illumination transformation to model the shadow effects in the image that allows the shadow image to be expressed as a combination of the shadow-free image, the shadow parameters, and a matte layer. We use two deep networks, namely SP-Net and M-Net, to predict the shadow parameters and the shadow matte respectively. This system allows us to remove the shadow effects on the images. We train and test our framework on the most challenging shadow removal dataset (ISTD). Compared to the state-of-the-art method, our model achieves a 40% error reduction in terms of root mean square error (RMSE) for the shadow area, reducing RMSE from 13.3 to 7.9. Moreover, we create an augmented ISTD dataset based on an image decomposition system by modifying the shadow parameters to generate new synthetic shadow images. Training our model on this new augmented ISTD dataset further lowers the RMSE on the shadow area to 7.4.
Hieu Le 0001, Dimitris Samaras
ICCV1
2019 Efficient and Extensible Policy Mining for Relationship-Based Access Control
abstract
Relationship-based access control (ReBAC) is a flexible and expressive framework that allows policies to be expressed in terms of chains of relationship between entities as well as attributes of entities. ReBAC policy mining algorithms have a potential to significantly reduce the cost of migration from legacy access control systems to ReBAC, by partially automating the development of a ReBAC policy. Existing ReBAC policy mining algorithms support a policy language with a limited set of operators; this limits their applicability.
Thang Bui, Scott D. Stoller, Hieu Le 0001
SACMAT3
2018 A+D Net: Training a Shadow Detector with Adversarial Shadow Attenuation
Hieu Le 0001, Tomás F. Yago Vicente, Vu Nguyen 0004, Minh Hoai, Dimitris Samaras
ECCV (2)1
2018 Iterative Crowd Counting
Viresh Ranjan, Hieu Le 0001, Minh Hoai
ECCV (7)2
2016 Geodesic Distance Histogram Feature for Video Segmentation
Hieu Le 0001, Vu Nguyen 0004, Chen-Ping Yu, Dimitris Samaras
ACCV (1)1
2015 Efficient Video Segmentation Using Parametric Graph Partitioning
abstract
Video segmentation is the task of grouping similar pixels in the spatio-temporal domain, and has become an important preprocessing step for subsequent video analysis. Most video segmentation and supervoxel methods output a hierarchy of segmentations, but while this provides useful multiscale information, it also adds difficulty in selecting the appropriate level for a task. In this work, we propose an efficient and robust video segmentation framework based on parametric graph partitioning (PGP), a fast, almost parameter free graph partitioning method that identifies and removes between-cluster edges to form node clusters. Apart from its computational efficiency, PGP performs clustering of the spatio-temporal volume without requiring a pre-specified cluster number or bandwidth parameters, thus making video segmentation more practical to use in applications. The PGP framework also allows processing sub-volumes, which further improves performance, contrary to other streaming video segmentation methods where sub-volume processing reduces performance. We evaluate the PGP method using the SegTrack v2 and Chen Xiph.org datasets, and show that it outperforms related state-of-the-art algorithms in 3D segmentation metrics and running time.
Chen-Ping Yu, Hieu Le 0001, Gregory J. Zelinsky, Dimitris Samaras
ICCV2