VLDB 2026 Research / reviewers in the wild / expert
Minh Hoai
dblp:135/4935 · also Minh Hoai Nguyen
· DBLP profile ↗
92ranked-venue papers
18as first author
42since 2021 · last 2025
0000-0002-2415-6048ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 80 · 17 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 77 · 13 first-author · 39 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Few-shot Personalized Scanpath PredictionabstractA personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training scanpath prediction models are data-intensive and cannot be effectively personalized to new individuals with only a few available examples. In this paper, we propose few-shot personalized scanpath prediction task (FS-PSP) and a novel method to address it, which aims to predict scan-paths for an unseen subject using minimal support data of that subject’s scanpath behavior. The key to our method’s adaptability is the Subject-Embedding Network (SE-Net), specifically designed to capture unique, individualized representations for each subject’s scanpaths. SE-Net generates subject embeddings that effectively distinguish between subjects while minimizing variability among scanpaths from the same individual. The personalized scanpath prediction model is then conditioned on these subject embeddings to produce accurate, personalized results. Experiments on multiple eye-tracking datasets demonstrate that our method excels in FS-PSP settings and does not require any fine-tuning steps at test time. Code is available at: https://github.com/cvlab-stonybrook/few-shot-scanpath Ruoyu Xue, Sounak Mondal, Hieu Le 0001, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras |
CVPR | 6 |
| 2025 | Multi-View Gaze Target EstimationabstractThis paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to improve accuracy and expand applicability, addressing limitations in existing single-view methods that face challenges such as face occlusion, target ambiguity, and out-of-view targets. Our method processes a pair of camera views as input, incorporating a Head Information Aggregation (HIA) module for leveraging head information from both views for more accurate gaze estimation, an Uncertainty-based Gaze Selection (UGS) for identifying the most reliable gaze output, and an Epipolar-based Scene Attention (ESA) module for cross-view background information sharing. This approach significantly outperforms single-view baselines, especially when the second camera provides a clear view of the person's face. Additionally, our method can estimate the gaze target in the first view using the image of the person in the second view only, a capability not possessed by single-view GTE methods. Furthermore, the paper introduces a multi-view dataset for developing and evaluating multi-view GTE methods. Data and code are available at https://www3.cs.stonybrook.edu/~cvl/multiview_gte.html Qiaomu Miao, Vivek Raju Golani, Progga Paromita Dutta, Minh Hoai, Dimitris Samaras |
ICCV | 5 |
| 2025 | Region-Level Data Attribution for Text-To-Image Generative Models
Trong Bang Nguyen, Phi-Le Nguyen, Simon Lucey, Minh Hoai |
ICCV | 4 |
| 2025 | DualMat: PBR Material Estimation via Coherent Dual-Path DiffusionabstractWe present DualMat, a novel dual-path diffusion framework for estimating Physically Based Rendering (PBR) materials from single images under complex lighting conditions. Our approach operates in two distinct latent spaces: an albedo-optimized path leveraging pretrained visual knowledge through RGB latent space, and a material-specialized path operating in a compact latent space designed for precise metallic and roughness estimation. To ensure coherent predictions between the albedo-optimized and material-specialized paths, we introduce feature distillation during training. We employ rectified flow to enhance efficiency by reducing inference steps while maintaining quality. Our framework extends to high-resolution and multi-view inputs through patch-based estimation and cross-view attention, enabling seamless integration into image-to-3D pipelines. DualMat achieves state-of-the-art performance on both Objaverse and real-world data, significantly outperforming existing methods with up to 28% improvement in albedo estimation and 39% reduction in metallic-roughness prediction errors. Our project can be found at yifehuang97.github.io/DualMatProjPage/. Yi Xu 0002, Minh Hoai, Zhong Li 0007 |
ACM Multimedia | 4 |
| 2025 | Class-Agnostic Repetitive Action Counting Using Wearable DevicesabstractWe present Class-agnostic Repetitive action Counting (CaRaCount), a novel approach to count repetitive human actions in the wild using wearable devices time series data. CaRaCount is the first few-shot class-agnostic method, being able to count repetitions of any action class with only a short exemplar data sequence containing a few examples from the action class of interest. To develop and evaluate this method, we collect a large-scale time series dataset of repetitive human actions in various context, containing smartwatch data from 10 subjects performing 50 different activities. Experiments on this dataset and three other activity counting datasets namely Crossfit, Recofit, and MM-Fit show that CaRaCount can count repetitive actions with low error, and it outperforms other baselines and state-of-the-art action counting methods. Finally, with a user experience study, we evaluate the usability of our real-time implementation. Our results highlight the efficiency and effectiveness of our approach when deployed outside the laboratory environments. Duc Duy Nguyen, Lam Thanh Nguyen, Cuong Pham 0001, Minh Hoai |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Count What You Want: Exemplar Identification and Few-Shot Counting of Human Actions in the WildabstractThis paper addresses the task of counting human actions of interest using sensor data from wearable devices. We propose a novel exemplar-based framework, allowing users to provide exemplars of the actions they want to count by vocalizing predefined sounds ``one'', ``two'', and ``three''. Our method first localizes temporal positions of these utterances from the audio sequence. These positions serve as the basis for identifying exemplars representing the action class of interest. A similarity map is then computed between the exemplars and the entire sensor data sequence, which is further fed into a density estimation module to generate a sequence of estimated density values. Summing these density values provides the final count. To develop and evaluate our approach, we introduce a diverse and realistic dataset consisting of real-world data from 37 subjects and 50 action categories, encompassing both sensor and audio data. The experiments on this dataset demonstrate the viability of the proposed method in counting instances of actions from new classes and subjects that were not part of the training data. On average, the discrepancy between the predicted count and the ground truth value is 7.47, significantly lower than the errors of the frequency-based and transformer-based methods. Our project, code and dataset can be found at https://github.com/cvlab-stonybrook/ExRAC. Duc Duy Nguyen, Cuong Pham 0001, Minh Hoai |
AAAI | 5 |
| 2024 | Efficiency-preserving Scene-adaptive Object Detection
Vu Quang Truong, Minh Hoai |
BMVC | 3 |
| 2024 | Unifying Top-Down and Bottom-Up Scanpath Prediction Using TransformersabstractMost models of visual attention aim at predicting either top-down or bottom-up control, as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT), a single model that predicts both forms of attention control. HAT uses a novel transformer-based architecture and a simplified foveated retina that collectively create a spatio-temporal awareness akin to the dynamic visual working memory of humans. HAT not only establishes a new state-of-the-art in predicting the scanpath of fixations made during target-present and target-absent visual search and "taskless" free viewing, but also makes human gaze behavior interpretable. Unlike previous methods that rely on a coarse grid of fixation cells and experience information loss due to fixation discretization, HAT features a sequential dense prediction architecture and outputs a dense heatmap for each fixation, thus avoiding discretizing fixations. HAT sets a new standard in computational attention, which emphasizes effectiveness, generality, and interpretability. HAT's demonstrated scope and applicability will likely inspire the development of new attention models that can better predict human behavior in various attention-demanding scenarios. Code is available at https://github.com/cvlab-stonybrook/HAT. Zhibo Yang 0002, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras |
CVPR | 6 |
| 2024 | Error Detection in Egocentric Procedural Task VideosabstractWe present a new egocentric procedural error dataset containing videos with various types of errors as well as normal videos and propose a new framework for procedural error detection using error-free training videos only. Our framework consists of an action segmentation model and a contrastive step prototype learning module to segment actions and learn useful features for error detection. Based on the observation that interactions between hands and objects often inform action and error understanding, we propose to combine holistic frame features with relations features, which we learn by building a graph using active object detection followed by a Graph Convolutional Network. To handle errors, unseen during training, we use our contrastive step prototype learning to learn multiple prototypes for each step, capturing variations of error-free step executions. At inference time, we use feature-prototype similarities for error detection. By experiments on three datasets, we show that our proposed framework outperforms state-of-the-art video anomaly detection methods for error detection and provides smooth action and error predictions.11Code and data is available at https://github.com/robert80203/EgoPER_official Shih-Po Lee, Zijia Lu, Minh Hoai, Ehsan Elhamifar |
CVPR | 4 |
| 2024 | HanDiffuser: Text-to-Image Generation with Realistic Hand AppearancesabstractText-to-image generative models can generate high-quality humans, but realism is lost when generating hands. Common artifacts include irregular hand poses, shapes, incorrect numbers of fingers, and physically implausible finger orientations. To generate images with realistic hands, we propose a novel diffusion-based architecture called HanDiffuser that achieves realism by injecting hand embeddings in the generative process. HanDiffuser consists of two components: a Text-to-Hand-Params diffusion model to generate SMPL-Body and MANO-Hand parameters from input text prompts, and a Text-Guided Hand-Params-to-Image diffusion model to synthesize images by conditioning on the prompts and hand parameters generated by the previous component. We incorporate multiple aspects of hand representation, including 3D shapes and joint-level finger positions, orientations and articulations, for robust learning and reliable performance during inference. We conduct extensive quantitative and qualitative experiments and perform user studies to demonstrate the efficacy of our method in generating images with high-quality hands. Project page: https://supreethn.github.io/research/handiffuser/index.html Supreeth Narasimhaswamy, Uttaran Bhattacharya, Xiang Chen 0010, Ishita Dasgupta 0002, Saayan Mitra, Minh Hoai |
CVPR | 6 |
| 2024 | HOIST-Former: Hand-Held Objects Identification, Segmentation, and Tracking in the WildabstractWe address the challenging task of identifying, segmenting, and tracking hand-held objects, which is crucial for applications such as human action segmentation and per-formance evaluation. This task is particularly challenging due to heavy occlusion, rapid motion, and the transitory nature of objects being hand-held, where an object may be held, released, and subsequently picked up again. To tackle these challenges, we have developed a novel transformer-based architecture called HOIST-Former. HOIST-Former is adept at spatially and temporally segmenting hands and objects by iteratively pooling features from each other, ensuring that the processes of identification, segmentation, and tracking of hand-held objects depend on the hands' positions and their contextual appearance. We further refine HOIST-Former with a contact loss that focuses on areas where hands are in contact with objects. Moreover, we also contribute an in-the-wild video dataset called HOIST, which comprises 4,125 videos complete with bounding boxes, segmentation masks, and tracking IDs for hand-held objects. Through experiments on the HOIST dataset and two ad-ditional public datasets, we demonstrate the efficacy of HOIST-Former in segmenting and tracking hand-held objects. Project page: https://supreethn.github.io/research/hoistformer/index.html Supreeth Narasimhaswamy, Huy Anh Nguyen, Lihan Huang, Minh Hoai |
CVPR | 4 |
| 2024 | Blur2Blur: Blur Conversion for Unsupervised Image Deblurring on Unknown DomainsabstractThis paper presents an innovative framework designed to train an image deblurring algorithm tailored to a specific camera device. This algorithm works by transforming a blurry input image, which is challenging to deblur, into another blurry image that is more amenable to deblurring. The transformation process, from one blurry state to another, leverages unpaired data consisting of sharp and blurry images captured by the target camera device. Learning this blur-to-blur transformation is inherently simpler than direct blur-to-sharp conversion, as it primarily involves modifying blur patterns rather than the intricate task of reconstructing fine image details. The efficacy of the proposed approach has been demonstrated through comprehensive experiments on various benchmarks, where it significantly outperforms state-of-the-art methods both quantitatively and qualitatively. Our code and data are available at https://github.com/VinAIResearch/Blur2Blur Bang-Dang Pham, Anh Tuan Tran 0001, Cuong Pham 0001, Rang Nguyen, Minh Hoai |
CVPR | 6 |
| 2024 | Diffusion-Refined VQA Annotations for Semi-supervised Gaze Following
Qiaomu Miao, Alexandros Graikos, Sounak Mondal, Minh Hoai, Dimitris Samaras |
ECCV (39) | 5 |
| 2024 | Look Hear: Gaze Prediction for Speech-Directed Human Attention
Sounak Mondal, Seoyoung Ahn, Zhibo Yang 0002, Niranjan Balasubramanian, Dimitris Samaras, Gregory J. Zelinsky, Minh Hoai |
ECCV (42) | 7 |
| 2024 | Characterizing Learners' Complex Attentional States During Online Multimedia Learning Using Eye-tracking, Egocentric Camera, Webcam, and Retrospective recallsabstractAs online learning becomes increasingly ubiquitous, a key challenge is maintaining learners' sustained attention. Using eye-tracking, together with observing and interviewing learners, we can characterize both 1) whether they are looking at their learning materials, and 2) whether they are thinking about them. Critically, eye-tracking only speaks to the first distinction, not the second. To overcome this limitation, we supplemented eye-tracking with an egocentric camera, a webcam, a retrospective recall, and mind-wandering probes to capture a 2×2 matrix of attentional/cognitive states. We then categorized N=101 learners' attentional/cognitive states while they completed a multimedia physics module. This meets two goals: 1) allowing basic research to understand the relationship between attentional/cognitive states and behavioral outcomes; and 2) facilitating applied research by generating rich ground truth for future use in training machine learning to categorize this 2×2 set of attentional states, for which eye-tracking is necessary, but not sufficient. Prasanth Chandran, Jeremy Munsell, Brian Howatt, Brayden Wallace, Lindsey Wilson, Sidney K. D'Mello, Minh Hoai, N. Sanjay Rebello, Lester C. Loschky |
ETRA | 8 |
| 2023 | Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human AttentionabstractPredicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent scanpath prediction models focus on goaldirected attention (search). Such models are limited in their application due to a common approach relying on trained target detectors for all possible objects, and the availability of human gaze data for their training (both not scalable). In response, we pose a new task called ZeroGaze, a new variant of zero-shot learning where gaze is predicted for never-before-searched objects, and we develop a novel model, Gazeformer; to solve the ZeroGaze problem. In contrast to existing methods using object detector modules, Gazeformer encodes the target using a natural language model, thus leveraging semantic similarities in scanpath prediction. We use a transformer-based encoder-decoder architecture because transformers are particularly useful for generating contextual representations. Gazeformer surpasses other models by a large margin (19%-70%) on the ZeroGaze setting. It also outperforms existing target-detection models on standard gaze prediction for both target-present and target-absent search tasks. In addition to its improved performance, Gazeformer is more than five times faster than the state-of-the-art target-present visual search model. Code can be found at https://github.com/cvlab-stonybrook/Gazeformer/ Sounak Mondal, Zhibo Yang 0002, Seoyoung Ahn, Dimitris Samaras, Gregory J. Zelinsky, Minh Hoai |
CVPR | 6 |
| 2023 | HyperCUT: Video Sequence from a Single Blurry Image using Unsupervised OrderingabstractWe consider the challenging task of training models for image-to-video deblurring, which aims to recover a sequence of sharp images corresponding to a given blurry image input. A critical issue disturbing the training of an image-to-video model is the ambiguity of the frame ordering since both the forward and backward sequences are plausible solutions. This paper proposes an effective self-supervised ordering scheme that allows training highquality image-to-video deblurring models. Unlike previous methods that rely on order-invariant losses, we assign an explicit order for each video sequence, thus avoiding the order-ambiguity issue. Specifically, we map each video sequence to a vector in a latent high-dimensional space so that there exists a hyperplane such that for every video sequence, the vectors extracted from it and its reversed sequence are on different sides of the hyperplane. The side of the vectors will be used to define the order of the corresponding sequence. Last but not least, we propose a real-image dataset for the image-to-video deblurring problem that covers a variety of popular domains, including face, hand, and street. Extensive experimental results confirm the effectiveness of our method. Code and data are available at https://github.com/VinAIResearch/HyperCUT.git Bang-Dang Pham, Anh Tuan Tran 0001, Cuong Pham 0001, Rang Nguyen, Minh Hoai |
CVPR | 6 |
| 2023 | Object Detection with Self-Supervised Scene AdaptationabstractThis paper proposes a novel method to improve the performance of a trained object detector on scenes with fixed camera perspectives based on self-supervised adaptation. Given a specific scene, the trained detector is adapted using pseudo-ground truth labels generated by the detector itself and an object tracker in a cross-teaching manner. When the camera perspective is fixed, our method can utilize the background equivariance by proposing artifact-free object mixup as a means of data augmentation, and utilize accurate background extraction as an additional input modality. We also introduce a large-scale and diverse dataset for the development and evaluation of scene-adaptive object detection. Experiments on this dataset show that our method can improve the average precision of the original detector, outperforming the previous state-of-the-art selfsupervised domain adaptive object detection methods by a large margin. Our dataset and code are published at https://github.com/cvlab-stonybrook/scenes100. Minh Hoai |
CVPR | 2 |
| 2023 | Interactive Class-Agnostic Object CountingabstractWe propose a novel framework for interactive class-agnostic object counting, where a human user can interactively provide feedback to improve the accuracy of a counter. Our framework consists of two main components: a user-friendly visualizer to gather feedback and an efficient mechanism to incorporate it. In each iteration, we produce a density map to show the current prediction result, and we segment it into non-overlapping regions with an easily verifiable number of objects. The user can provide feedback by selecting a region with obvious counting errors and specifying the range for the estimated number of objects within it. To improve the counting result, we develop a novel adaptation loss to force the visual counter to output the predicted count within the user-specified range. For effective and efficient adaptation, we propose a refinement module that can be used with any density-based visual counter, and only the parameters in the refinement module will be updated during adaptation. Our experiments on two challenging class-agnostic object counting benchmarks, FSCD-LVIS and FSC-147, show that our method can reduce the mean absolute error of multiple state-of-the-art visual counters by roughly 30% to 40% with minimal user input. Our project can be found at https://yifehuang97.github.io/ICACountProjectPage/. Viresh Ranjan, Minh Hoai |
ICCV | 3 |
| 2023 | Patch-level Gaze Distribution Prediction for Gaze FollowingabstractGaze following aims to predict where a person is looking in a scene, by predicting the target location, or indicating that the target is located outside the image. Recent works detect the gaze target by training a heatmap regression task with a pixel-wise mean-square error (MSE) loss, while formulating the in/out prediction task as a binary classification task. This training formulation puts a strict, pixel-level constraint in higher resolution on the single annotation available in training, and does not consider annotation variance and the correlation between the two subtasks. To address these issues, we introduce the patch distribution prediction (PDP) method. We replace the in/out prediction branch in previous models with the PDP branch, by predicting a patch-level gaze distribution that also considers the outside cases. Experiments show that our model regularizes the MSE loss by predicting better heatmap distributions on images with larger annotation variances, meanwhile bridging the gap between the target prediction and in/out prediction subtasks, showing a significant improvement in performance on both subtasks on public gaze following datasets. Qiaomu Miao, Minh Hoai, Dimitris Samaras |
WACV | 2 |
| 2022 | Exemplar Free Class Agnostic Counting
Viresh Ranjan, Minh Hoai |
ACCV (4) | 2 |
| 2022 | From Within to Between: Knowledge Distillation for Cross Modality Retrieval
Vinh Tran 0005, Niranjan Balasubramanian, Minh Hoai |
ACCV (4) | 3 |
| 2022 | Self-supervised Learning with Multi-view Rendering for 3D Point Cloud Analysis
Bach Tran, Binh-Son Hua, Anh Tuan Tran 0001, Minh Hoai |
ACCV (1) | 4 |
| 2022 | Forward Propagation, Backward Regression, and Pose Association for Hand Tracking in the WildabstractWe propose HandLer, a novel convolutional architecture that can jointly detect and track hands online in unconstrained videos. HandLer is based on Cascade-RCNN with additional three novel stages. The first stage is Forward Propagation, where the features from frame t −1 are propagated to frame t based on previously detected hands and their estimated motion. The second stage is the Detection and Backward Regression, which uses outputs from the forward propagation to detect hands for frame t and their relative offset in frame t −1. The third stage uses an off-the-shelf human pose method to link any fragmented hand tracklets. We train the forward propagation and backward regression and detection stages end-to-end together with the other Cascade-RCNN components. To train and evaluate HandLer, we also contribute YouTube-Hand, the first challenging large-scale dataset of unconstrained videos annotated with hand locations and their trajectories. Experiments on this dataset and other benchmarks show that HandLer outperforms the existing state-of-the-art tracking algorithms by a large margin. Code and data are available at https://vision.cs.stonybrook.edu/~mingzhen/handler/. Mingzhen Huang, Supreeth Narasimhaswamy, Saif Vazir, Haibin Ling, Minh Hoai |
CVPR | 5 |
| 2022 | Whose Hands are These? Hand Detection and Hand-Body Association in the WildabstractWe study a new problem of detecting hands and finding the location of the corresponding person for each detected hand. This task is helpful for many downstream tasks such as hand tracking and hand contact estimation. Associating hands with people is challenging in unconstrained conditions since multiple people can be present in the scene with varying overlaps and occlusions. We propose a novel end-to-end trainable convolutional network that can Jointly detect hands and the body location for the corresponding person. Our method first detects a set of hands and bodies and uses a novel Hand-Body Association Network to predict association scores between them. We use these association scores to find the body location for each detected hand. We also introduce a new challenging dataset called BodyHands containing uncon-strained images with hand and their corresponding body locations annotations. We conduct extensive experiments on BodyHands and another public dataset to show the effectiveness of our method. Finally, we demonstrate the benefits of hand-body association in two critical applications: hand tracking and hand contact estimation. Our experiments show that hand tracking and hand contact estimation methods can be improved significantly by reasoning about the hand-body association. Code and data can be found at http://vision.cs.stonybrook.edu/~supreeth/BodyHands/. Supreeth Narasimhaswamy, Mingzhen Huang, Minh Hoai |
CVPR | 4 |
| 2022 | Few-Shot Object Counting and Detection
Chau Pham 0002, Khoi Nguyen 0001, Minh Hoai |
ECCV (20) | 4 |
| 2022 | Target-Absent Human Attention
Zhibo Yang 0002, Sounak Mondal, Seoyoung Ahn, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras |
ECCV (4) | 5 |
| 2021 | Localization in the Crowd with Topological ConstraintsabstractWe address the problem of crowd localization, i.e., the prediction of dots corresponding to people in a crowded scene. Due to various challenges, a localization method is prone to spatial semantic errors, i.e., predicting multiple dots within a same person or collapsing multiple dots in a cluttered region. We propose a topological approach targeting these semantic errors. We introduce a topological constraint that teaches the model to reason about the spatial arrangement of dots. To enforce this constraint, we define a persistence loss based on the theory of persistent homology. The loss compares the topographic landscape of the likelihood map and the topology of the ground truth. Topological reasoning improves the quality of the localization algorithm especially near cluttered regions. On multiple public benchmarks, our method outperforms previous localization methods. Additionally, we demonstrate the potential of our method in improving the performance in the crowd counting task. Shahira Abousamra, Minh Hoai, Dimitris Samaras, Chao Chen 0012 |
AAAI | 2 |
| 2021 | Exemplar-Based Early Event Prediction in Video
Farrukh M. Koraishy, Minh Hoai |
BMVC | 3 |
| 2021 | Progressive Semantic SegmentationabstractThe objective of this work is to segment high-resolution images without overloading GPU memory usage or losing the fine details in the output segmentation map. The memory constraint means that we must either downsample the big image or divide the image into local patches for separate processing. However, the former approach would lose the fine details, while the latter can be ambiguous due to the lack of a global picture. In this work, we present MagNet, a multi-scale framework that resolves local ambiguity by looking at the image at multiple magnification levels. MagNet has multiple processing stages, where each stage corresponds to a magnification level, and the output of one stage is fed into the next stage for coarse-to-fine information propagation. Each stage analyzes the image at a higher resolution than the previous stage, recovering the previously lost details due to the lossy downsampling step, and the segmentation output is progressively refined through the processing stages. Experiments on three high-resolution datasets of urban views, aerial scenes, and medical images show that MagNet consistently outperforms the state-of-the-art methods by a significant margin. Code is available at https://github.com/VinAIResearch/MagNet. Chuong Huynh, Anh Tuan Tran 0001, Khoa Luu, Minh Hoai |
CVPR | 4 |
| 2021 | Dictionary-Guided Scene Text RecognitionabstractLanguage prior plays an important role in the way humans detect and recognize text in the wild. Current scene text recognition methods do use lexicons to improve recognition performance, but their naive approach of casting the output into a dictionary word based purely on the edit distance has many limitations. In this paper, we present a novel approach to incorporate a dictionary in both the training and inference stage of a scene text recognition system. We use the dictionary to generate a list of possible outcomes and find the one that is most compatible with the visual appearance of the text. The proposed method leads to a robust scene text recognition model, which is better at handling ambiguous cases encountered in the wild, and improves the overall performance of state-of-the-art scene text spotting frameworks. Our work suggests that incorporating language prior is a potential approach to advance scene text detection and recognition methods. Besides, we contribute VinText, a challenging scene text dataset for Vietnamese, where some characters are equivocal in the visual form due to accent symbols. This dataset will serve as a challenging benchmark for measuring the applicability and robustness of scene text detection and recognition algorithms. Code and dataset are available at https://github.com/VinAIResearch/dict-guided. Thu Nguyen 0003, Vinh Tran 0005, Minh-Triet Tran, Thanh Duc Ngo, Thien Huu Nguyen, Minh Hoai |
CVPR | 7 |
| 2021 | Lipstick Ain't Enough: Beyond Color Matching for In-the-Wild Makeup TransferabstractMakeup transfer is the task of applying on a source face the makeup style from a reference image. Real-life makeups are diverse and wild, which cover not only color-changing but also patterns, such as stickers, blushes, and jewelries. However, existing works overlooked the latter components and confined makeup transfer to color manipulation, focusing only on light makeup styles. In this work, we propose a holistic makeup transfer framework that can handle all the mentioned makeup components. It consists of an improved color transfer branch and a novel pattern transfer branch to learn all makeup properties, including color, shape, texture, and location. To train and evaluate such a system, we also introduce new makeup datasets for real and synthetic extreme makeup. Experimental results show that our framework achieves the state of the art performance on both light and extreme makeup styles. Code is available at https://github.com/VinAIResearch/CPM. Anh Tuan Tran 0001, Minh Hoai |
CVPR | 3 |
| 2021 | Learning To Count EverythingabstractExisting works on visual counting primarily focus on one specific category at a time, such as people, animals, and cells. In this paper, we are interested in counting everything, that is to count objects from any category given only a few annotated instances from that category. To this end, we pose counting as a few-shot regression task. To tackle this task, we present a novel method that takes a query image together with a few exemplar objects from the query image and predicts a density map for the presence of all objects of interest in the query image. We also present a novel adaptation strategy to adapt our network to any novel visual category at test time, using only a few exemplar objects from the novel category. We also introduce a dataset of 147 object categories containing over 6000 images that are suitable for the few-shot counting task. The images are annotated with two types of annotation, dots and bounding boxes, and they can be used for developing few-shot counting models. Experiments on this dataset shows that our method outperforms several state-of-the-art object detectors and few-shot counting approaches. Our code and dataset can be found at https://github.com/cvlab-stonybrook/LearningToCountEverything. Viresh Ranjan, Udbhav Sharma, Thu Nguyen 0003, Minh Hoai |
CVPR | 4 |
| 2021 | Explore Image Deblurring via Encoded Blur Kernel SpaceabstractThis paper introduces a method to encode the blur operators of an arbitrary dataset of sharp-blur image pairs into a blur kernel space. Assuming the encoded kernel space is close enough to in-the-wild blur operators, we propose an alternating optimization algorithm for blind image deblurring. It approximates an unseen blur operator by a kernel in the encoded space and searches for the corresponding sharp image. Unlike recent deep-learning-based methods, our system can handle unseen blur kernel, while avoiding using complicated handcrafted priors on the blur operator often found in classical methods. Due to the method’s design, the encoded kernel space is fully differentiable, thus can be easily adopted in deep neural network models. Moreover, our method can be used for blur synthesis by transferring existing blur operators from a given dataset into a new domain. Finally, we provide experimental results to confirm the effectiveness of the proposed method. The code is available at https://github.com/VinAIResearch/blur-kernelspace-exploring. Anh Tuan Tran 0001, Quynh Phung, Minh Hoai |
CVPR | 4 |
| 2021 | Toward Realistic Single-View 3D Object Reconstruction with Unsupervised Learning from Multiple ImagesabstractRecovering the 3D structure of an object from a single image is a challenging task due to its ill-posed nature. One approach is to utilize the plentiful photos of the same object category to learn a strong 3D shape prior for the object. This approach has successfully been demonstrated by a recent work of Wu et al. (2020), which obtained impressive 3D reconstruction networks with unsupervised learning. However, their algorithm is only applicable to symmetric objects. In this paper, we eliminate the symmetry requirement with a novel unsupervised algorithm that can learn a 3D reconstruction network from a multi-image dataset. Our algorithm is more general and covers the symmetry-required scenario as a special case. Besides, we employ a novel albedo loss that improves the reconstructed details and realisticity. Our method surpasses the previous work in both quality and robustness, as shown in experiments on datasets of various structures, including single-view, multiview, image-collection, and video sets. Code is available at: https://github.com/VinAIResearch/LeMul. Long-Nhat Ho, Anh Tuan Tran 0001, Quynh Phung, Minh Hoai |
ICCV | 4 |
| 2021 | Knowledge Distillation for Human Action AnticipationabstractWe consider the task of training a neural network to anticipate human actions in video. This task is challenging given the complexity of video data, the stochastic nature of the future, and the limited amount of annotated training data. In this paper, we propose a novel knowledge distillation framework that uses an action recognition network to supervise the training of an action anticipation network, guiding the latter to attend to the relevant information needed for correctly anticipating the future actions. This framework is possible thanks to a novel loss function to account for positional shifts of semantic concepts in a dynamic video. The knowledge distillation framework is a form of self-supervised learning, and it takes advantage of unlabeled data. Experimental results on JHMDB and EPIC-KITCHENS dataset show the effectiveness of our approach. Vinh Tran 0005, Yang Wang 0097, Minh Hoai |
ICIP | 4 |
| 2021 | Progressive Knowledge Distillation For Early Action RecognitionabstractWe present a novel framework to train a recurrent neural network for early recognition of human actions, which is an important but challenging task given the need to recognize an on-going action based on partial observation. Our framework is based on knowledge distillation, where the network for early recognition is viewed as a student model. The student is trained using knowledge distilled from a more knowledgeable teacher model that can peek into the future and incorporate extra observations about the action in consideration. This framework can be used in both supervised and semi-supervised learning settings, being able to utilize both the labeled and unlabeled training data. Experiments on the UCF101, SYSU 3DHOI, and NTU RGB-D datasets show the effectiveness of knowledge distillation for early recognition, including when we only have a small amount of annotated training data. Vinh Tran 0005, Niranjan Balasubramanian, Minh Hoai |
ICIP | 3 |
| 2021 | Adaptive Streaming of 360-Degree Videos with Reinforcement LearningabstractFor bandwidth-efficient streaming of 360-degree videos, the streaming technique must adapt both to the changing viewport of the user and variations of the available network bandwidth. The state-of-the-art streaming techniques for this problem attempt to solve an optimization using simplified rules that do not adapt very well to the uncertainties related to the viewport or network. We adopt a 3D-Convolutional Neural Networks (3DCNN) model to extract spatio-temporal features of videos and predict the viewport. Given the sequential decision-making nature of such streaming technique, we then apply a Reinforcement Learning (RL) based adaptive streaming approach. We address the challenges of using RL in this scenario, such as large action space and delayed reward evaluation. Comprehensive evaluations with real network traces show that the proposed method outperforms three tile-based streaming techniques for 360-degree videos. Compared to the tile-based streaming techniques, the average user-perceived bitrate of the proposed method is 1.3-1.7 times higher and the average quality of experience of the proposed method is also 1.6-3.4 times higher. Subjective user studies further confirm the superiority of the proposed approach. Sohee Kim Park, Minh Hoai, Arani Bhattacharya, Samir Ranjan Das |
WACV | 2 |
| 2021 | Supervoxel Attention Graphs for Long-Range Video ModelingabstractA significant challenge in video understanding is posed by the high dimensionality of the input, which induces large computational cost and high memory footprints. Deep convolutional models operating on video apply pooling and striding to reduce feature dimensionality and to increase the receptive field. However, despite these strategies, modern approaches cannot effectively leverage spatiotemporal structure over long temporal extents. In this paper we introduce an approach that reduces a video of 10 seconds to a sparse graph of only 160 feature nodes such that efficient inference in this graph produces state-of-the-art accuracy on challenging action recognition datasets. The nodes of our graph are semantic supervoxels that capture the spatiotemporal structure of objects and motion cues in the video, while edges between nodes encode spatiotemporal relations and feature similarity. We demonstrate that a shallow network that interleaves graph convolution and graph pooling on this compact representation implements an effective mechanism of relational reasoning yielding strong recognition results on both Charades and Something-Something. Yang Wang 0097, Gedas Bertasius, Tae-Hyun Oh, Minh Hoai, Lorenzo Torresani |
WACV | 5 |
| 2021 | Large Scale Shadow Annotation and Detection Using Lazy Annotation and Stacked CNNsabstractRecent shadow detection algorithms have shown initial success on small datasets of images from specific domains. However, shadow detection on broader image domains is still challenging due to the lack of annotated training data, caused by the intense manual labor required for annotating shadow data. In this paper we propose "lazy annotation", an efficient annotation method where an annotator only needs to mark the important shadow areas and some non-shadow areas. This yields data with noisy labels that are not yet useful for training a shadow detector. We address the problem of label noise by jointly learning a shadow region classifier and recovering the labels in the training set. We consider the training labels as unknowns and formulate label recovery as the minimization of the sum of squared leave-one-out errors of a Least Squares SVM, which can be efficiently optimized. Experimental results show that a classifier trained with recovered labels achieves comparable performance to a classifier trained on the properly annotated data. These results motivated us to collect a new dataset that is 20 times larger than existing datasets and contains a large variety of scenes and image types. Naturally, such a large dataset is appropriate for training deep learning methods. Thus, we propose a stacked Convolutional Neural Network architecture that efficiently trains on patch level shadow examples while incorporating image level semantic information. This means that the detected shadow patches are refined based on image semantics. Our proposed pipeline, trained on recovered labels, performs at state-of-the art level. Furthermore, the proposed model performs exceptionally well on a cross dataset task, proving the generalization power of the proposed architecture and dataset. Le Hou, Tomás F. Yago Vicente, Minh Hoai, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Sequence-to-Segments Networks for Detecting Segments in VideosabstractDetecting segments of interest from videos is a common problem for many applications. And yet it is a challenging problem as it often requires not only knowledge of individual target segments, but also contextual understanding of the entire video and the relationships between the target segments. To address this problem, we propose the Sequence-to-Segments Network (S2N), a novel and general end-to-end sequential encoder-decoder architecture. S2N first encodes the input video into a sequence of hidden states that capture information progressively, as it appears in the video. It then employs the Segment Detection Unit (SDU), a novel decoding architecture, that sequentially detects segments. At each decoding step, the SDU integrates the decoder state and encoder hidden states to detect a target segment. During training, we address the problem of finding the best assignment of predicted segments to ground truth using the Hungarian Matching Algorithm with Lexicographic Cost. Additionally we propose to use the squared Earth Mover's Distance to optimize the localization errors of the segments. We show the state-of-the-art performance of S2N across numerous tasks, including video highlighting, video summarization, and human action proposal generation. Zijun Wei, Boyu Wang 0001, Minh Hoai, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Radomír Mech, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Interactive Visual Study of Multiple Attributes Learning Model of X-Ray Scattering ImagesabstractExisting interactive visualization tools for deep learning are mostly applied to the training, debugging, and refinement of neural network models working on natural images. However, visual analytics tools are lacking for the specific application of x-ray image classification with multiple structural attributes. In this paper, we present an interactive system for domain scientists to visually study the multiple attributes learning models applied to x-ray scattering images. It allows domain scientists to interactively explore this important type of scientific images in embedded spaces that are defined on the model prediction output, the actual labels, and the discovered feature space of neural networks. Users are allowed to flexibly select instance images, their clusters, and compare them regarding the specified visual representation of attributes. The exploration is guided by the manifestation of model performance related to mutual relationships among attributes, which often affect the learning accuracy and effectiveness. The system thus supports domain scientists to improve the training dataset and model, find questionable attributes labels, and identify outlier images or spurious data clusters. Case studies and scientists feedback demonstrate its functionalities and usefulness. Xinyi Huang 0003, Suphanut Jamonnak, Ye Zhao 0003, Boyu Wang 0001, Minh Hoai, Kevin G. Yager, Wei Xu 0020 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Uncertainty Estimation and Sample Selection for Crowd Counting
Viresh Ranjan, Boyu Wang 0001, Mubarak Shah, Minh Hoai |
ACCV (5) | 4 |
| 2020 | Attentive Action and Context Factorization
Yang Wang 0097, Vinh Tran 0005, Gedas Bertasius, Lorenzo Torresani, Minh Hoai |
BMVC | 5 |
| 2020 | Active Vision for Early Recognition of Human ActionsabstractWe propose a method for early recognition of human actions, one that can take advantages of multiple cameras while satisfying the constraints due to limited communication bandwidth and processing power. Our method considers multiple cameras, and at each time step, it will decide the best camera to use so that a confident recognition decision can be reached as soon as possible. We formulate the camera selection problem as a sequential decision process, and learn a view selection policy based on reinforcement learning. We also develop a novel recurrent neural network architecture to account for the unobserved video frames and the irregular intervals between the observed frames. Experiments on three datasets demonstrate the effectiveness of our approach for early recognition of human actions. Boyu Wang 0001, Lihan Huang, Minh Hoai |
CVPR | 3 |
| 2020 | Learning Visual Emotion Representations From Web DataabstractWe present a scalable approach for learning powerful visual features for emotion recognition. A critical bottleneck in emotion recognition is the lack of large scale datasets that can be used for learning visual emotion features. To this end, we curate a webly derived large scale dataset, StockEmotion, which has more than a million images. StockEmotion uses 690 emotion related tags as labels giving us a fine-grained and diverse set of emotion labels, circumventing the difficulty in manually obtaining emotion annotations. We use this dataset to train a feature extraction network, EmotionNet, which we further regularize using joint text and visual embedding and text distillation. Our experimental results establish that EmotionNet trained on the StockEmotion dataset outperforms SOTA models on four different visual emotion tasks. An aded benefit of our joint embedding training approach is that EmotionNet achieves competitive zero-shot recognition performance against fully supervised baselines on a challenging visual emotion dataset, EMOTIC, which further highlights the generalizability of the learned emotion features. Zijun Wei, Jianming Zhang 0001, Zhe Lin 0001, Joon-Young Lee, Niranjan Balasubramanian, Minh Hoai, Dimitris Samaras |
CVPR | 6 |
| 2020 | Predicting Goal-Directed Human Attention Using Inverse Reinforcement LearningabstractHuman gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We propose the first inverse reinforcement learning (IRL) model to learn the internal reward function and policy used by humans during visual search. We modeled the viewer's internal belief states as dynamic contextual belief maps of object locations. These maps were learned and then used to predict behavioral scanpaths for multiple target categories. To train and evaluate our IRL model we created COCO-Search18, which is now the largest dataset of high-quality search fixations in existence. COCO-Search18 has 10 participants searching for each of 18 target-object categories in 6202 images, making about 300,000 goal-directed fixations. When trained and evaluated on COCO-Search18, the IRL model outperformed baseline models in predicting search fixation scanpaths, both in terms of similarity to human search behavior and search efficiency. Finally, reward maps recovered by the IRL model reveal distinctive target-dependent patterns of object prioritization, which we interpret as a learned object context. Zhibo Yang 0002, Lihan Huang, Yupei Chen, Zijun Wei, Seoyoung Ahn, Gregory J. Zelinsky, Dimitris Samaras, Minh Hoai |
CVPR | 8 |
| 2020 | Distribution Matching for Crowd CountingabstractIn crowd counting, each training image contains multiple people, where each person is annotated by a dot. Existing crowd counting methods need to use a Gaussian to smooth each annotated dot or to estimate the likelihood of every pixel given the annotated point. In this paper, we show that imposing Gaussians to annotations hurts generalization performance. Instead, we propose to use Distribution Matching for crowd COUNTing (DM-Count). In DM-Count, we use Optimal Transport (OT) to measure the similarity between the normalized predicted density map and the normalized ground truth density map. To stabilize OT computation, we include a Total Variation loss in our model. We show that the generalization error bound of DM-Count is tighter than that of the Gaussian smoothed methods. In terms of Mean Absolute Error, DM-Count outperforms the previous state-of-the-art methods by a large margin on two large-scale counting datasets, UCF-QNRF and NWPU, and achieves the state-of-the-art results on the ShanghaiTech and UCF-CC50 datasets. DM-Count reduced the error of the state-of-the-art published result by approximately 16%. Code is available at https://github.com/cvlab-stonybrook/DM-Count. Boyu Wang 0001, Huidong Liu, Dimitris Samaras, Minh Hoai |
NeurIPS | 4 |
| 2020 | Detecting Hands and Recognizing Physical Contact in the WildabstractWe investigate a new problem of detecting hands and recognizing their physical contact state in unconstrained conditions. This is a challenging inference task given the need to reason beyond the local appearance of hands. The lack of training annotations indicating which object or parts of an object the hand is in contact with further complicates the task. We propose a novel convolutional network based on Mask-RCNN that can jointly learn to localize hands and predict their physical contact to address this problem. The network uses outputs from another object detector to obtain locations of objects present in the scene. It uses these outputs and hand locations to recognize the hand's contact state using two attention mechanisms. The first attention mechanism is based on the hand and a region's affinity, enclosing the hand and the object, and densely pools features from this region to the hand region. The second attention module adaptively selects salient features from this plausible region of contact. To develop and evaluate our method's performance, we introduce a large-scale dataset called ContactHands, containing unconstrained images annotated with hand locations and contact states. The proposed network, including the parameters of attention modules, is end-to-end trainable. This network achieves approximately 7% relative improvement over a baseline network that was built on the vanilla Mask-RCNN architecture and trained for recognizing hand contact states. Supreeth Narasimhaswamy, Minh Hoai |
NeurIPS | 3 |
| 2019 | WorkingHands: A Hand-Tool Assembly Dataset for Image Segmentation and Activity Mining
Roy Shilkrot, Supreeth Narasimhaswamy, Saif Vazir, Minh Hoai |
BMVC | 4 |
| 2019 | GIF2Video: Color Dequantization and Temporal Interpolation of GIF ImagesabstractGraphics Interchange Format (GIF) is a highly portable graphics format that is ubiquitous on the Internet. Despite their small sizes, GIF images often contain undesirable visual artifacts such as flat color regions, false contours, color shift, and dotted patterns. In this paper, we propose GIF2Video, the first learning-based method for enhancing the visual quality of GIFs in the wild. We focus on the challenging task of GIF restoration by recovering information lost in the three steps of GIF creation: frame sampling, color quantization, and color dithering. We first propose a novel CNN architecture for color dequantization. It is built upon a compositional architecture for multi-step color correction, with a comprehensive loss function designed to handle large quantization errors. We then adapt the SuperSlomo network for temporal interpolation of GIF frames. We introduce two large datasets, namely GIF-Faces and GIF-Moments, for both training and evaluation. Experimental results show that our method can significantly improve the visual quality of GIFs, and outperforms direct baseline and state-of-the-art approaches. Yang Wang 0097, Chuan Wang 0001, Tong He 0002, Jue Wang 0001, Minh Hoai |
CVPR | 6 |
| 2019 | Contextual Attention for Hand Detection in the WildabstractWe present Hand-CNN, a novel convolutional network architecture for detecting hand masks and predicting hand orientations in unconstrained images. Hand-CNN extends MaskRCNN with a novel attention mechanism to incorporate contextual cues in the detection process. This attention mechanism can be implemented as an efficient network module that captures non-local dependencies between features. This network module can be inserted at different stages of an object detection network, and the entire detector can be trained end-to-end. We also introduce large-scale annotated hand datasets containing hands in unconstrained images for training and evaluation. We show that Hand-CNN outperforms existing methods on the newly collected datasets and the publicly available PASCAL VOC human layout dataset. Data and code: https://www3.cs.stonybrook.edu/~cvl/projects/hand_det_attention/. Supreeth Narasimhaswamy, Zhengwei Wei, Yang Wang 0097, Justin Zhang 0003, Minh Hoai |
ICCV | 5 |
| 2018 | Pulling Actions out of Context: Explicit Separation for Effective CombinationabstractThe ability to recognize human actions in video has many potential applications. Human action recognition, however, is tremendously challenging for computers due to the complexity of video data and the subtlety of human actions. Most current recognition systems flounder on the inability to separate human actions from co-occurring factors that usually dominate subtle human actions. In this paper, we propose a novel approach for training a human action recognizer, one that can: (1) explicitly factorize human actions from the co-occurring factors; (2) deliberately build a model for human actions and a separate model for all correlated contextual elements; and (3) effectively combine the models for human action recognition. Our approach exploits the benefits of conjugate samples of human actions, which are video clips that are contextually similar to human action samples, but do not contain the action. Experiments on ActionThread, PASCAL VOC, UCF101, and Hollywood2 datasets demonstrate the ability to separate action from context of the proposed approach. Yang Wang 0097, Minh Hoai |
CVPR | 2 |
| 2018 | Good View Hunting: Learning Photo Composition From Dense View PairsabstractFinding views with good photo composition is a challenging task for machine learning methods. A key difficulty is the lack of well annotated large scale datasets. Most existing datasets only provide a limited number of annotations for good views, while ignoring the comparative nature of view selection. In this work, we present the first large scale Comparative Photo Composition dataset, which contains over one million comparative view pairs annotated using a cost-effective crowdsourcing workflow. We show that these comparative view annotations are essential for training a robust neural network model for composition. In addition, we propose a novel knowledge transfer framework to train a fast view proposal network, which runs at 75+ FPS and achieves state-of-the-art performance in image cropping and thumbnail generation tasks on three benchmark datasets. The superiority of our method is also demonstrated in a user study on a challenging experiment, where our method significantly outperforms the baseline methods in producing diversified well-composed views. Zijun Wei, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Radomír Mech, Minh Hoai, Dimitris Samaras |
CVPR | 6 |
| 2018 | A+D Net: Training a Shadow Detector with Adversarial Shadow Attenuation
Hieu Le 0001, Tomás F. Yago Vicente, Vu Nguyen 0004, Minh Hoai, Dimitris Samaras |
ECCV (2) | 4 |
| 2018 | Iterative Crowd Counting
Viresh Ranjan, Hieu Le 0001, Minh Hoai |
ECCV (7) | 3 |
| 2018 | Predicting Body Movement and Recognizing Actions: An Integrated Framework for Mutual BenefitsabstractHuman action recognition and body movement prediction are important tasks. They are different and have traditionally been addressed separately. These tasks, however, provide mutual benefits to each other, and existing methods fail to capture these benefits. In this paper, we propose a method for jointly recognizing the action and predicting the movement of a person. Our method is based on two Long-Short Term Memory (LSTM) recurrent neural networks, but extend them to provide and receive benefits of each other. In particular, we design two LSTM architectures. One LSTM can generate a sequence of body movement conditioned on the past movement and the predicted class of the action, and the other LSTM can recognize the human action based on the predicted sequence of body movement. Experiments on Montalbano and MSR Action 3D datasets show that movement prediction provides benefits to early recognition of human action, which in turn improves the quality of the predicted movement. Boyu Wang 0001, Minh Hoai |
FG | 2 |
| 2018 | Eigen-Evolution Dense Trajectory DescriptorsabstractTrajectory-pooled Deep-learning Descriptors have been the state-of-the-art feature descriptors for human action recognition in video on many datasets. This paper improves their performance by applying the proposed eigen-evolution pooling to each trajectory, encoding the temporal evolution of deep learning features computed along the trajectory. This leads to Eigen-Evolution Trajectory (EET) descriptors, a novel type of video descriptor that significantly outperforms Trajectory-pooled Deep-learning Descriptors. EET descriptors are defined based on dense trajectories, and they provide complimentary benefits to video descriptors that are not based on trajectories. Empirically, we observe that the combination of EET descriptors and VideoDarwin outperforms the state-of-the-art methods on the Hollywood2 dataset, and its performance on the UCF101 dataset is close to the state-of-the-art. Yang Wang 0097, Vinh Quang Tran, Minh Hoai |
FG | 3 |
| 2018 | Sequence-to-Segment Networks for Segment DetectionabstractDetecting segments of interest from an input sequence is a challenging problem which often requires not only good knowledge of individual target segments, but also contextual understanding of the entire input sequence and the relationships between the target segments. To address this problem, we propose the Sequence-to-Segment Network (S$^2$N), a novel end-to-end sequential encoder-decoder architecture. S$^2$N first encodes the input into a sequence of hidden states that progressively capture both local and holistic information. It then employs a novel decoding architecture, called Segment Detection Unit (SDU), that integrates the decoder state and encoder hidden states to detect segments sequentially. During training, we formulate the assignment of predicted segments to ground truth as bipartite matching and use the Earth Mover's Distance to calculate the localization errors. We experiment with S$^2$N on temporal action proposal generation and video summarization and show that S$^2$N achieves state-of-the-art performance on both tasks. Zijun Wei, Boyu Wang 0001, Minh Hoai, Jianming Zhang 0001, Zhe Lin 0001, Xiaohui Shen, Radomír Mech, Dimitris Samaras |
NeurIPS | 3 |
| 2018 | Back to the beginning: Starting point detection for early recognition of ongoing human actions
Boyu Wang 0001, Minh Hoai |
Comput. Vis. Image Underst. | 2 |
| 2018 | Leave-One-Out Kernel Optimization for Shadow Detection and RemovalabstractThe objective of this work is to detect shadows in images. We pose this as the problem of labeling image regions, where each region corresponds to a group of superpixels. To predict the label of each region, we train a kernel Least-Squares Support Vector Machine (LSSVM) for separating shadow and non-shadow regions. The parameters of the kernel and the classifier are jointly learned to minimize the leave-one-out cross validation error. Optimizing the leave-one-out cross validation error is typically difficult, but it can be done efficiently in our framework. Experiments on two challenging shadow datasets, UCF and UIUC, show that our region classifier outperforms more complex methods. We further enhance the performance of the region classifier by embedding it in a Markov Random Field (MRF) framework and adding pairwise contextual cues. This leads to a method that outperforms the state-of-the-art for shadow detection. In addition we propose a new method for shadow removal based on region relighting. For each shadow region we use a trained classifier to identify a neighboring lit region of the same material. Given a pair of lit-shadow regions we perform a region relighting transformation based on histogram matching of luminance values between the shadow region and the lit region. Once a shadow is detected, we demonstrate that our shadow removal approach produces results that outperform the state of the art by evaluating our method using a publicly available benchmark dataset. Tomás F. Yago Vicente, Minh Hoai, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Latent Bi-Constraint SVM for Video-Based Object RecognitionabstractWe address the task of recognizing objects from video input. This important problem is relatively unexplored, compared with image-based object recognition. To this end, we make the following contributions. First, we introduce two comprehensive data sets for video-based object recognition. Second, we propose latent bi-constraint SVM (LBSVM), a maximum-margin framework for video-based object recognition. LBSVM is based on structured-output SVM, but extends it to handle noisy video data and ensure consistency of the output decision throughout time. We apply LBSVM to recognize office objects and museum sculptures, and we demonstrate its benefits over image-based, set-based, and other video-based object recognition. Yang Liu 0020, Minh Hoai, Mang Shao, Tae-Kyun Kim 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Large-scale Continual Road Inspection: Visual Infrastructure Assessment in the Wild
Ke Ma 0005, Minh Hoai, Dimitris Samaras |
BMVC | 2 |
| 2017 | Shadow Detection with Conditional Generative Adversarial NetworksabstractWe introduce scGAN, a novel extension of conditional Generative Adversarial Networks (GAN) tailored for the challenging problem of shadow detection in images. Previous methods for shadow detection focus on learning the local appearance of shadow regions, while using limited local context reasoning in the form of pairwise potentials in a Conditional Random Field. In contrast, the proposed adversarial approach is able to model higher level relationships and global scene characteristics. We train a shadow detector that corresponds to the generator of a conditional GAN, and augment its shadow accuracy by combining the typical GAN loss with a data loss term. Due to the unbalanced distribution of the shadow labels, we use weighted cross entropy. With the standard GAN architecture, properly setting the weight for the cross entropy would require training multiple GANs, a computationally expensive grid procedure. In scGAN, we introduce an additional sensitivity parameter w to the generator. The proposed approach effectively parameterizes the loss of the trained detector. The resulting shadow detector is a single network that can generate shadow maps corresponding to different sensitivity levels, obviating the need for multiple models and a costly training procedure. We evaluate our method on the large-scale SBU and UCF shadow datasets, and observe up to 17% error reduction with respect to the previous state-of-the-art method. Vu Nguyen 0004, Tomás F. Yago Vicente, Maozheng Zhao, Minh Hoai, Dimitris Samaras |
ICCV | 4 |
| 2017 | X-Ray Scattering Image Classification Using Deep LearningabstractVisual inspection of x-ray scattering images is a powerful technique for probing the physical structure of materials at the molecular scale. In this paper, we explore the use of deep learning to develop methods for automatically analyzing x-ray scattering images. In particular, we apply Convolutional Neural Networks and Convolutional Autoencoders for x-ray scattering image classification. To acquire enough training data for deep learning, we use simulation software to generate synthetic x-ray scattering images. Experiments show that deep learning methods outperform previously published methods by 10% on synthetic and real datasets. Boyu Wang 0001, Kevin G. Yager, Dantong Yu, Minh Hoai |
WACV | 4 |
| 2016 | Noisy Label Recovery for Shadow Detection in Unfamiliar DomainsabstractRecent shadow detection algorithms have shown initial success on small datasets of images from specific domains. However, shadow detection on broader image domains is still challenging due to the lack of annotated training data. This is due to the intense manual labor in annotating shadow data. In this paper we propose "lazy annotation", an efficient annotation method where an annotator only needs to mark the important shadow areas and some non-shadow areas. This yields data with noisy labels that are not yet useful for training a shadow detector. We address the problem of label noise by jointly learning a shadow region classifier and recovering the labels in the training set. We consider the training labels as unknowns and formulate the label recovery problem as the minimization of the sum of squared leave-one-out errors of a Least Squares SVM, which can be efficiently optimized. Experimental results show that a classifier trained with recovered labels achieves comparable performance to a classifier trained on the properly annotated data. These results suggest a feasible approach to address the task of detecting shadows in an unfamiliar domain: collecting and lazily annotating some images from the new domain for training. As will be demonstrated, this approach outperforms methods that rely on precisely annotated but less relevant datasets. Initial results suggest more general applicability. Tomás F. Yago Vicente, Minh Hoai, Dimitris Samaras |
CVPR | 2 |
| 2016 | Improving Human Action Recognition by Non-action ClassificationabstractIn this paper we consider the task of recognizing human actions in realistic video where human actions are dominated by irrelevant factors. We first study the benefits of removing non-action video segments, which are the ones that do not portray any human action. We then learn a nonaction classifier and use it to down-weight irrelevant video segments. The non-action classifier is trained using Action-Thread, a dataset with shot-level annotation for the occurrence or absence of a human action. The non-action classifier can be used to identify non-action shots with high precision and subsequently used to improve the performance of action recognition systems. Yang Wang 0097, Minh Hoai |
CVPR | 2 |
| 2016 | Region Ranking SVM for Image ClassificationabstractThe success of an image classification algorithm largely depends on how it incorporates local information in the global decision. Popular approaches such as average-pooling and max-pooling are suboptimal in many situations. In this paper we propose Region Ranking SVM (RRSVM), a novel method for pooling local information from multiple regions. RRSVM exploits the correlation of local regions in an image, and it jointly learns a region evaluation function and a scheme for integrating multiple regions. Experiments on PASCAL VOC 2007, VOC 2012, and ILSVRC2014 datasets show that RRSVM outperforms the methods that use the same feature type and extract features from the same set of local regions. RRSVM achieves similar to or better than the state-of-the-art performance on all datasets. Zijun Wei, Minh Hoai |
CVPR | 2 |
| 2016 | Large-Scale Training of Shadow Detectors with Noisily-Annotated Shadow Examples
Tomás F. Yago Vicente, Le Hou, Chen-Ping Yu, Minh Hoai, Dimitris Samaras |
ECCV (6) | 4 |
| 2016 | Learned Region Sparsity and Diversity Also Predicts Visual AttentionabstractLearned region sparsity has achieved state-of-the-art performance in classification tasks by exploiting and integrating a sparse set of local information into global decisions. The underlying mechanism resembles how people sample information from an image with their eye movements when making similar decisions. In this paper we incorporate the biologically plausible mechanism of Inhibition of Return into the learned region sparsity model, thereby imposing diversity on the selected regions. We investigate how these mechanisms of sparsity and diversity relate to visual attention by testing our model on three different types of visual search tasks. We report state-of-the-art results in predicting the locations of human gaze fixations, even though our model is trained only on image-level labels without object location annotations. Notably, the classification performance of the extended model remains the same as the original. This work suggests a new computational perspective on visual attention mechanisms and shows how the inclusion of attention-based mechanisms can improve computer vision techniques. Zijun Wei, Hossein Adeli, Minh Hoai, Gregory J. Zelinsky, Dimitris Samaras |
NIPS | 3 |
| 2015 | Leave-One-Out Kernel Optimization for Shadow DetectionabstractThe objective of this work is to detect shadows in images. We pose this as the problem of labeling image regions, where each region corresponds to a group of superpixels. To predict the label of each region, we train a kernel Least-Squares SVM for separating shadow and non-shadow regions. The parameters of the kernel and the classifier are jointly learned to minimize the leave-one-out cross validation error. Optimizing the leave-one-out cross validation error is typically difficult, but it can be done efficiently in our framework. Experiments on two challenging shadow datasets, UCF and UIUC, show that our region classifier outperforms more complex methods. We further enhance the performance of the region classifier by embedding it in an MRF framework and adding pairwise contextual cues. This leads to a method that significantly outperforms the state-of-the-art. Tomás F. Yago Vicente, Minh Hoai, Dimitris Samaras |
ICCV | 2 |
| 2014 | Thread-Safe: Towards Recognizing Human Actions Across Shot Boundaries
Minh Hoai, Andrew Zisserman |
ACCV (4) | 1 |
| 2014 | Improving Human Action Recognition Using Score Distribution and Ranking
Minh Hoai, Andrew Zisserman |
ACCV (5) | 1 |
| 2014 | Regularized Max Pooling for Image Categorization
Minh Hoai |
BMVC | 1 |
| 2014 | Action Recognition From Weak Alignment of Body Parts
Minh Hoai, Lubor Ladicky, Andrew Zisserman |
BMVC | 1 |
| 2014 | Talking Heads: Detecting Humans and Recognizing Their InteractionsabstractThe objective of this work is to accurately and efficiently detect configurations of one or more people in edited TV material. Such configurations often appear in standard arrangements due to cinematic style, and we take advantage of this to provide scene context. We make the following contributions: first, we introduce a new learnable context aware configuration model for detecting sets of people in TV material that predicts the scale and location of each upper body in the configuration, second, we show that inference of the model can be solved globally and efficiently using dynamic programming, and implement a maximum margin learning framework, and third, we show that the configuration model substantially outperforms a Deformable Part Model (DPM) for predicting upper body locations in video frames, even when the DPM is equipped with the context of other upper bodies. Experiments are performed over two datasets: the TV Human Interaction dataset, and 150 episodes from four different TV shows. We also demonstrate the benefits of the model in recognizing interactions in TV shows. Minh Hoai, Andrew Zisserman |
CVPR | 1 |
| 2014 | Max-Margin Early Event Detectors
Minh Hoai, Fernando De la Torre |
Int. J. Comput. Vis. | 1 |
| 2014 | Learning discriminative localization from weakly labeled data
Minh Hoai, Lorenzo Torresani, Fernando De la Torre, Carsten Rother |
Pattern Recognit. | 1 |
| 2013 | Discriminative Sub-categorizationabstractThe objective of this work is to learn sub-categories. Rather than casting this as a problem of unsupervised clustering, we investigate a weakly supervised approach using both positive and negative samples of the category. We make the following contributions: (i) we introduce a new model for discriminative sub-categorization which determines cluster membership for positive samples whilst simultaneously learning a max-margin classifier to separate each cluster from the negative samples, (ii) we show that this model does not suffer from the degenerate cluster problem that afflicts several competing methods (e.g., Latent SVM and Max-Margin Clustering), (iii) we show that the method is able to discover interpretable sub-categories in various datasets. The model is evaluated experimentally over various datasets, and its performance advantages over k-means and Latent SVM are demonstrated. We also stress test the model and show its resilience in discovering sub-categories as the parameters are varied. Minh Hoai, Andrew Zisserman |
CVPR | 1 |
| 2012 | Max-margin early event detectorsabstractThe need for early detection of temporal events from sequential data arises in a wide spectrum of applications ranging from human-robot interaction to video security. While temporal event detection has been extensively studied, early detection is a relatively unexplored problem. This paper proposes a maximum-margin framework for training temporal event detectors to recognize partial events, enabling early detection. Our method is based on Structured Output SVM, but extends it to accommodate sequential data. Experiments on datasets of varying complexity, for detecting facial expressions, hand gestures, and human activities, demonstrate the benefits of our approach. To the best of our knowledge, this is the first paper in the literature of computer vision that proposes a learning formulation for early event detection. Minh Hoai, Fernando De la Torre |
CVPR | 1 |
| 2011 | Joint segmentation and classification of human actions in videoabstractAutomatic video segmentation and action recognition has been a long-standing problem in computer vision. Much work in the literature treats video segmentation and action recognition as two independent problems; while segmentation is often done without a temporal model of the activity, action recognition is usually performed on pre-segmented clips. In this paper we propose a novel method that avoids the limitations of the above approaches by jointly performing video segmentation and action recognition. Unlike standard approaches based on extensions of dynamic Bayesian networks, our method is based on a discriminative temporal extension of the spatial bag-of-words model that has been very popular in object recognition. The classification is performed robustly within a multi-class SVM framework whereas the inference over the segments is done efficiently with dynamic programming. Experimental results on honeybee, Weizmann, and Hollywood datasets illustrate the benefits of our approach compared to state-of-the-art methods. Minh Hoai, Zhen-Zhong Lan, Fernando De la Torre |
CVPR | 1 |
| 2010 | Action unit detection with segment-based SVMsabstractAutomatic facial action unit (AU) detection from video is a long-standing problem in computer vision. Two main approaches have been pursued: (1) static modeling - typically posed as a discriminative classification problem in which each video frame is evaluated independently; (2) temporal modeling - frames are segmented into sequences and typically modeled with a variant of dynamic Bayesian networks. We propose a segment-based approach, kSeg-SVM, that incorporates benefits of both approaches and avoids their limitations. kSeg-SVM is a temporal extension of the spatial bag-of-words. kSeg-SVM is trained within a structured output SVM framework that formulates AU detection as a problem of detecting temporal events in a time series of visual features. Each segment is modeled by a variant of the BoW representation with soft assignment of the words based on similarity. Our framework has several benefits for AU detection: (1) both dependencies between features and the length of action units are modeled; (2) all possible segments of the video may be used for training; and (3) no assumptions are required about the underlying structure of the action unit events (e.g., i.i.d.). Our algorithm finds the best k-or-fewer segments that maximize the SVM score. Experimental results suggest that the proposed method outperforms state-of-the-art static methods for AU detection. Tomas Simon, Minh Hoai, Fernando De la Torre, Jeffrey F. Cohn |
CVPR | 2 |
| 2010 | Metric Learning for Image Alignment
Minh Hoai, Fernando De la Torre |
Int. J. Comput. Vis. | 1 |
| 2010 | Optimal feature selection for support vector machines
Minh Hoai, Fernando De la Torre |
Pattern Recognit. | 1 |
| 2009 | Weakly supervised discriminative localization and classification: a joint learning processabstractVisual categorization problems, such as object classification or action recognition, are increasingly often approached using a detection strategy: a classifier function is first applied to candidate subwindows of the image or the video, and then the maximum classifier score is used for class decision. Traditionally, the subwindow classifiers are trained on a large collection of examples manually annotated with masks or bounding boxes. The reliance on time-consuming human labeling effectively limits the application of these methods to problems involving very few categories. Furthermore, the human selection of the masks introduces arbitrary biases (e.g. in terms of window size and location) which may be suboptimal for classification. In this paper we propose a novel method for learning a discriminative subwindow classifier from examples annotated with binary labels indicating the presence of an object or action of interest, but not its location. During training, our approach simultaneously localizes the instances of the positive class and learns a subwindow SVM to recognize them. We extend our method to classification of time series by presenting an algorithm that localizes the most discriminative set of temporal segments in the signal. We evaluate our approach on several datasets for object and action recognition and show that it achieves results similar and in many cases superior to those obtained with full supervision. Minh Hoai, Lorenzo Torresani, Lorenzo de la Torre, Carsten Rother |
ICCV | 1 |
| 2008 | Local minima free Parameterized Appearance ModelsabstractParameterized Appearance Models (PAMs) (e.g. Eigentracking, Active Appearance Models, Morphable Models) are commonly used to model the appearance and shape variation of objects in images. While PAMs have numerous advantages relative to alternate approaches, they have at least two drawbacks. First, they are especially prone to local minima in the fitting process. Second, often few if any of the local minima of the cost function correspond to acceptable solutions. To solve these problems, this paper proposes a method to learn a cost function by explicitly optimizing that the local minima occur at and only at the places corresponding to the correct fitting parameters. To the best of our knowledge, this is the first paper to address the problem of learning a cost function to explicitly model local properties of the error surface to fit PAMs. Synthetic and real examples show improvement in alignment performance in comparison with traditional approaches. Minh Hoai, Fernando De la Torre |
CVPR | 1 |
| 2008 | Parameterized Kernel Principal Component Analysis: Theory and applications to supervised and unsupervised image alignmentabstractParameterized Appearance Models (PAMs) (e.g. eigen-tracking, active appearance models, morphable models) use Principal Component Analysis (PCA) to model the shape and appearance of objects in images. Given a new image with an unknown appearance/shape configuration, PAMs can detect and track the object by optimizing the model’s parameters that best match the image. While PAMs have numerous advantages for image registration relative to alternative approaches, they suffer from two major limitations: First, PCA cannot model non-linear structure in the data. Second, learning PAMs requires precise manually labeled training data. This paper proposes Parameterized Kernel Principal Component Analysis (PKPCA), an extension of PAMs that uses Kernel PCA (KPCA) for learning a non-linear appearance model invariant to rigid and/or non-rigid deformations. We demonstrate improved performance in supervised and unsupervised image registration, and present a novel application to improve the quality of manual landmarks in faces. In addition, we suggest a clean and effective matrix formulation for PKPCA. Fernando De la Torre, Minh Hoai |
CVPR | 2 |
| 2008 | Facial feature detection with optimal pixel reduction SVMabstractAutomatic facial feature localization has been a long-standing challenge in the field of computer vision for several decades. This can be explained by the large variation a face in an image can have due to factors such as position, facial expression, pose, illumination, and background clutter. Support Vector Machines (SVMs) have been a popular statistical tool for facial feature detection. Traditional SVM approaches to facial feature detection typically extract features from images (e.g. multiband filter, SIFT features) and learn the SVM parameters. Independently learning features and SVM parameters might result in a loss of information related to the classification process. This paper proposes an energy-based framework to jointly perform relevant feature weighting and SVM parameter learning. Preliminary experiments on standard face databases have shown significant improvement in speed with our approach. Minh Hoai, Joan Perez, Fernando De la Torre |
FG | 1 |
| 2008 | Learning image alignment without local minima for face detection and trackingabstractActive appearance models (AAMs) have been extensively used for face alignment during the last 20 years. While AAMs have numerous advantages relative to alternate approaches, they suffer from two major drawbacks: (i) AAMs are especially prone to local minima in the fitting process; (ii) few if any of the local minima of the cost function correspond to acceptable solutions. To minimize these problems, this paper proposes a method to learn the fitting cost function that explicitly optimizes that the local minima occur at and only at the places corresponding to the correct fitting parameters. The paper explores two methods to parameterize the cost function: pixel weighting and subspace learning. Experiments on synthetic and real data show the effectiveness of our approach for face alignment. Minh Hoai, Fernando De la Torre |
FG | 1 |
| 2008 | Robust Kernel Principal Component AnalysisabstractKernel Principal Component Analysis (KPCA) is a popular generalization of linear PCA that allows non-linear feature extraction. In KPCA, data in the input space is mapped to higher (usually) dimensional feature space where the data can be linearly modeled. The feature space is typically induced implicitly by a kernel function, and linear PCA in the feature space is performed via the kernel trick. However, due to the implicitness of the feature space, some extensions of PCA such as robust PCA cannot be directly generalized to KPCA. This paper presents a technique to overcome this problem, and extends it to a unified framework for treating noise, missing data, and outliers in KPCA. Our method is based on a novel cost function to perform inference in KPCA. Extensive experiments, in both synthetic and real data, show that our algorithm outperforms existing methods. Minh Hoai, Fernando De la Torre |
NIPS | 1 |
| 2008 | Image-based ShavingabstractAbstract Many categories of objects, such as human faces, can be naturally viewed as a composition of several different layers. For example, a bearded face with glasses can be decomposed into three layers: a layer for glasses, a layer for the beard and a layer for other permanent facial features. While modeling such a face with a linear subspace model could be very difficult, layer separation allows for easy modeling and modification of some certain structures while leaving others unchanged. In this paper, we present a method for automatic layer extraction and its applications to face synthesis and editing. Layers are automatically extracted by utilizing the differences between subspaces and modeled separately. We show that our method can be used for tasks such beard removal (virtual shaving), beard synthesis, and beard transfer, among others. Minh Hoai, Jean-François Lalonde, Alexei A. Efros, Fernando De la Torre |
Comput. Graph. Forum | 1 |
| 2003 | DRT: A Tool for Design Recovery of Interactive Graphical ApplicationsabstractNowadays, the majority of productivity applications are interactive and graphical in nature. In this demonstration, we explore the possibility of taking advantage of these two characteristics in a design recovery tool. Specifically, the fact that an application is interactive means that we can identify distinct execution bursts corresponding closely to "actions" performed by the user. The fact that the application is graphical means that we can describe those actions visually from a fragment of the application display itself. Combining these two ideas, we obtain an explicit mapping from high-level actions performed by a user (similar to use case scenarios/specification fragments) to their low-level implementation. This mapping can be used for design recovery of interactive graphical applications. We demonstrate our approach using L/sub Y/X, a scientific word processor. Keith Chan, Annie Chen, Zhi Cong Leo Liang, Amir Michail, Minh Hoai, Nicholas Seow |
ICSE | 5 |