Tim K. Marks

dblp:62/4853 · DBLP profile ↗
← Back
41ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-3517-4468ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 10 since 2021Systems, architecture and hardware · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Recovering Pulse Waves From Video Using Deep Unrolling and Deep Equilibrium Models
abstract
Camera-based contactless monitoring of vital signs, also known as imaging photoplethysmography (iPPG), has seen applications in driver-monitoring, perfusion assessment, affective computing, and more. iPPG involves sensing the underlying cardiac pulse from video of the skin and estimating vital signs such as the pulse rate or a full pulse waveform. Some previous iPPG methods impose model-based sparse priors on the pulse signals and use iterative optimization for pulse wave recovery, while others use end-to-end black-box deep learning methods. In contrast, we introduce methods that combine signal processing and deep learning methods in an inverse problem framework. Our methods estimate the underlying pulse signal, pulse rate, and pulse rate variability from facial video by learning deep-network-based denoising operators that leverage deep algorithm unfolding and deep equilibrium models. Experiments show that our methods can denoise an acquired signal from the face and infer the correct underlying pulse rate and pulse rate variability, achieving pulse rate estimation performance consistent with the state-of-the-art on well-known benchmarks, all with less than one-fifth the number of learnable parameters as the closest competing method.
Vineet R. Shenoy, Suhas Lohit, Hassan Mansour, Rama Chellappa, Tim K. Marks
IEEE Trans. Image Process.5
2024 TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models
abstract
Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., “a woman is drinking water.”). Existing TI2V frameworks often require costly training on video-text datasets and spe-cific model designs for text and image conditioning. In this paper, we propose TI2V-Zero, a zero-shot, tuning-free method that empowers a pretrained text-to-video (T2V) diffusion model to be conditioned on a provided image, enabling TI2V generation without any optimization, fine-tuning, or introducing external modules. Our approach leverages a pretrained T2V diffusion foundation model as the generative prior. To guide video generation with the additional image input, we propose a “repeat-and-slide” strategy that modulates the reverse denoising process, al-lowing the frozen diffusion model to synthesize a video frame-by-frame starting from the provided image. To ensure temporal continuity, we employ a DDPM inversion strategy to initialize Gaussian noise for each newly synthesized frame and a resampling technique to help preserve visual details. We conduct comprehensive experiments on both domain-specific and open-domain datasets, where TI2V-Zero consistently outperforms a recent open-domain TI2V model. Furthermore, we show that TI2V-Zero can seam-lessly extend to other tasks such as video infilling and pre-diction when provided with more images. Its autoregressive design also supports long video generation.
Haomiao Ni, Bernhard Egger 0001, Suhas Lohit, Anoop Cherian, Ye Wang 0001, Toshiaki Koike-Akino, Sharon X. Huang, Tim K. Marks
CVPR8
2024 RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual Navigation
abstract
We leverage Large Language Models (LLM) for zero-shot Semantic Audio Visual Navigation (SAVN). Existing methods utilize extensive training demonstrations for rein-forcement learning, yet achieve relatively low success rates and lack generalizability. The intermittent nature of au-ditory signals further poses additional obstacles to infer-ring the goal information. To address this challenge, we present the Reflective and Imaginative Language Agent (RILA). By employing multi-modal models to process sen-sory data, we instruct an LLM-based planner to actively ex-plore the environment. During the exploration, our agent adaptively evaluates and dismisses inaccurate perceptual descriptions. Additionally, we introduce an auxiliary LLM-based assistant to enhance global environmental compre-hension by mapping room layouts and providing strategic insights. Through comprehensive experiments and analy-sis, we show that our method outperforms relevant base-lines without training demonstrations from the environment and complementary semantic information.
Zeyuan Yang 0002, Jiageng Lin, Peihao Chen, Anoop Cherian, Tim K. Marks, Jonathan Le Roux, Chuang Gan 0001
CVPR5
2024 Few-shot Transparent Instance Segmentation for Bin Picking
abstract
In this paper, we consider the problem of segmenting multiple instances of a transparent object from RGB or gray scale camera images in a robotic bin picking setting. Prior methods for solving this task are usually built on the Mask-RCNN framework, but they require large annotated datasets for fine-tuning. Instead, we consider the task in a few-shot setting and present TrInSeg, a data-efficient and robust instance segmentation method for transparent objects based on Mask-RCNN. Our key innovations in TrInSeg are twofold: i) a novel method, dubbed TransMixup, for producing new training images using synthetic transparent object instances created by spatially transforming annotated examples; and ii) a method for scoring the consistency between the predicted segments and rotations of an ideal object template. In our new scoring method, the spatial transformations are produced by an auxiliary neural network, and the scores are then used to filter inconsistent instance predictions. To demonstrate the effectiveness of our method, we present experiments on a new few-shot dataset consisting of seven categories of non-opaque (transparent and translucent) objects, each category varying in the size, shape, and degree of transparency of the objects. Our results show that TrInSeg achieves state-of-the-art performance, improving fine-tuned Mask-RCNN by more than 14% in mIoU, while requiring very few annotated training samples.
Anoop Cherian, Siddarth Jain, Tim K. Marks
IROS3
2024 Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
abstract
We study the problem of multimodal physical scene understanding, where an embodied agent needs to find fallen objects by inferring object properties, direction, and distance of an impact sound source. Previous works adopt feed-forward neural networks to directly regress the variables from sound, leading to poor generalization and domain adaptation issues. In this paper, we illustrate that learning a disentangled model of acoustic formation, referred to as disentangled acoustic field (DAF), to capture the sound generation and propagation process, enables the embodied agent to construct a spatial uncertainty map over where the objects may have fallen. We demonstrate that our analysis-by-synthesis framework can jointly infer sound properties by explicitly decomposing and factorizing the latent space of the disentangled model. We further show that the spatial uncertainty map can significantly improve the success rate for the localization of fallen objects by proposing multiple plausible exploration locations.
Andrew Luo 0001, Yilun Du, Anoop Cherian, Tim K. Marks, Jonathan Le Roux, Chuang Gan 0001
IROS5
2023 Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image Synthesis
abstract
Conditional generative models typically demand large annotated training sets to achieve high-quality synthesis. As a result, there has been significant interest in designing models that perform plug-and-play generation, i.e., to use a predefined or pretrained model, which is not explicitly trained on the generative task, to guide the generative process (e.g., using language). However, such guidance is typically useful only towards synthesizing high-level semantics rather than editing fine-grained details as in image-to-image translation tasks. To this end, and capitalizing on the powerful fine-grained generative control offered by the recent diffusion-based generative models, we introduce Steered Diffusion, a generalized framework for photorealistic zero-shot conditional image generation using a diffusion model trained for unconditional generation. The key idea is to steer the image generation of the diffusion model at inference time via designing a loss using a pre-trained inverse model that characterizes the conditional task. This loss modulates the sampling trajectory of the diffusion process. Our framework allows for easy incorporation of multiple conditions during inference. We present experiments using steered diffusion on several tasks including inpainting, colorization, text-guided semantic editing, and image super-resolution. Our results demonstrate clear qualitative and quantitative improvements over state-of-the-art diffusion-based plug-and-play models while adding negligible additional computational cost.
Nithin Gopalakrishnan Nair, Anoop Cherian, Suhas Lohit, Ye Wang 0001, Toshiaki Koike-Akino, Vishal M. Patel, Tim K. Marks
ICCV7
2023 Unrolled iPPG: Video Heart Rate Estimation via Unrolling Proximal Gradient Descent
abstract
Imaging photoplethysmography (iPPG) is the process of estimating a person’s heart rate from video. In this work, we propose Unrolled iPPG, in which we integrate iterative optimization updates with deep learning-based signal priors to estimate the pulse waveform and heart rate from facial videos. We model the signal extracted from video as the sum of an underlying pulse signal and noise, but instead of explicitly imposing a handcrafted prior (e.g., sparsity in the frequency domain) on the signal, we learn priors on the signal and noise using neural networks. We solve for the underlying pulse signal by unrolling proximal gradient descent; the algorithm alternates between gradient descent steps and application of learned denoisers, which replace handcrafted priors and their proximal operators. Using this method, we achieve state-of-the-art heart rate estimation on the challenging MMSE-HR dataset.
Vineet R. Shenoy, Tim K. Marks, Hassan Mansour, Suhas Lohit
ICIP2
2023 Discriminative 3D Shape Modeling for Few-Shot Instance Segmentation
abstract
In this paper, we present a simple and efficient scheme for segmenting approximately convex 3D object instances in depth images in a few-shot setting via discriminatively modeling the 3D shape of the object using a neural network. Our key idea is to select pairs of 3D points on the depth image between which we compute surface geodesics. As the number of such geodesics is quadratic in the number of image pixels, we can create a large training set of geodesics using only very limited ground truth instance annotations. These annotations are used to create a binary label for each geodesic, which indicates whether or not that geodesic belongs entirely to one instance segment. A neural network is then trained to classify the geodesics using these labels. During inference, we create geodesics from selected seed points in the test depth image, then produce a convex hull of the points that are classified by the neural network as belonging to the same instance, thereby achieving instance segmentation. We present experiments applying our method to segmenting instances of food items in real-world depth images. Our results demonstrate promising performances compared to prior methods in accuracy and computational efficiency.
Anoop Cherian, Siddarth Jain, Tim K. Marks, Alan Sullivan
ICRA3
2023 H-SAUR: Hypothesize, Simulate, Act, Update, and Repeat for Understanding Object Articulations from Interactions
abstract
The world is filled with articulated objects that are difficult to determine how to use from vision alone, e.g., a door might open inwards or outwards. Humans handle these objects with strategic trial-and-error: first pushing a door then pulling if that doesn't work. We enable these capabilities in autonomous agents by proposing “Hypothesize, Simulate, Act, Update, and Repeat” (H-SAUR), a probabilistic generative framework that simultaneously generates a distribution of hypotheses about how objects articulate given input observations, captures certainty over hypotheses over time, and infer plausible actions for exploration and goal-conditioned manipulation. We compare our model with existing work in manipulating objects after a handful of exploration actions, on the PartNet-Mobility dataset. We further propose a novel PuzzleBoxes benchmark that contains locked boxes that require multiple steps to solve. We show that the proposed model significantly outperforms the current state-of-the-art articulated object manipulation framework, despite using zero training data. We further improve the test-time efficiency of H-SAUR by integrating a learned prior from learning-based vision models.
Kei Ota, Hsiao-Yu Fish Tung, Kevin A. Smith 0001, Anoop Cherian, Tim K. Marks, Alan Sullivan, Asako Kanezaki, Josh Tenenbaum
ICRA5
2022 (2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering
abstract
Spatio-temporal scene-graph approaches to video-based reasoning tasks, such as video question-answering (QA), typically construct such graphs for every video frame. These approaches often ignore the fact that videos are essentially sequences of 2D ``views'' of events happening in a 3D space, and that the semantics of the 3D scene can thus be carried over from frame to frame. Leveraging this insight, we propose a (2.5+1)D scene graph representation to better capture the spatio-temporal information flows inside the videos. Specifically, we first create a 2.5D (pseudo-3D) scene graph by transforming every 2D frame to have an inferred 3D structure using an off-the-shelf 2D-to-3D transformation module, following which we register the video frames into a shared (2.5+1)D spatio-temporal space and ground each 2D scene graph within it. Such a (2.5+1)D graph is then segregated into a static sub-graph and a dynamic sub-graph, corresponding to whether the objects within them usually move in the world. The nodes in the dynamic graph are enriched with motion features capturing their interactions with other graph nodes. Next, for the video QA task, we present a novel transformer-based reasoning pipeline that embeds the (2.5+1)D graph into a spatio-temporal hierarchical latent space, where the sub-graphs and their interactions are captured at varied granularity. To demonstrate the effectiveness of our approach, we present experiments on the NExT-QA and AVSD-QA datasets. Our results show that our proposed (2.5+1)D representation leads to faster training and inference, while our hierarchical model showcases superior performance on the video QA task versus the state of the art.
Anoop Cherian, Chiori Hori, Tim K. Marks, Jonathan Le Roux
AAAI3
2022 MOST-GAN: 3D Morphable StyleGAN for Disentangled Face Image Manipulation
abstract
Recent advances in generative adversarial networks (GANs) have led to remarkable achievements in face image synthesis. While methods that use style-based GANs can generate strikingly photorealistic face images, it is often difficult to control the characteristics of the generated faces in a meaningful and disentangled way. Prior approaches aim to achieve such semantic control and disentanglement within the latent space of a previously trained GAN. In contrast, we propose a framework that a priori models physical attributes of the face such as 3D shape, albedo, pose, and lighting explicitly, thus providing disentanglement by design. Our method, MOST-GAN, integrates the expressive power and photorealism of style-based GANs with the physical disentanglement and flexibility of nonlinear 3D morphable models, which we couple with a state-of-the-art 2D hair manipulation network. MOST-GAN achieves photorealistic manipulation of portrait images with fully disentangled 3D control over their physical attributes, enabling extreme manipulation of lighting, facial expression, and pose variations up to full profile view.
Safa C. Medin, Bernhard Egger 0001, Anoop Cherian, Ye Wang 0001, Josh Tenenbaum, Xiaoming Liu 0002, Tim K. Marks
AAAI7
2022 Audio-Visual Scene-Aware Dialog and Reasoning Using Audio-Visual Transformers with Joint Student-Teacher Learning
abstract
In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges (DSTC7, DSTC8). In these challenges, the best-performing systems relied heavily on human-generated descriptions of the video content, which were available in the datasets but would be unavailable in real-world applications. To promote further advancements for real-world applications, we proposed a third AVSD challenge, at DSTC10, with two modifications: 1) the human-created description is unavailable at inference time, and 2) systems must demonstrate temporal reasoning by finding evidence from the video to support each answer. This paper introduces the new task that includes temporal reasoning and our new extension of the AVSD dataset for DSTC10, for which we collected human-generated temporal reasoning data. We also introduce a baseline system built using an AV-transformer, which we released along with the new dataset. Finally, this paper introduces a new system that extends our baseline system with attentional multimodal fusion, joint student-teacher learning (JSTL), and model combination techniques, achieving state-of-the-art performances on the AVSD datasets for DSTC7, DSTC8, and DSTC10. We also propose two temporal reasoning methods for AVSD: one attention-based, and one based on a time-domain region proposal network.
Ankit P. Shah, Shijie Geng, Peng Gao 0007, Anoop Cherian, Takaaki Hori, Tim K. Marks, Jonathan Le Roux, Chiori Hori
ICASSP6
2022 Near-Infrared Imaging Photoplethysmography During Driving
abstract
Imaging photoplethysmography (iPPG) could greatly improve driver safety systems by enabling capabilities ranging from identifying driver fatigue to unobtrusive early heart failure detection. Unfortunately, the driving context poses unique challenges to iPPG, including illumination and motion. First, drastic illumination variations present during driving can overwhelm the small intensity-based iPPG signals. Second, significant driver head motion during driving, as well as camera motion (e.g., vibration) make it challenging to recover iPPG signals. To address these two challenges, we present two innovations. First, we demonstrate that we can reduce most outside light variations using narrow-band near-infrared (NIR) video recordings and obtain reliable heart rate estimates. Second, we present a novel optimization algorithm, which we call AutoSparsePPG, that leverages the quasi-periodicity of iPPG signals and achieves better performance than the state-of-the-art methods. In addition, we release the first publicly available driving dataset that contains both NIR and RGB video recordings of a passenger’s face with simultaneous ground truth pulse oximeter recordings.
Ewa Magdalena Nowara, Tim K. Marks, Hassan Mansour, Ashok Veeraraghavan
IEEE Trans. Intell. Transp. Syst.2
2021 InSeGAN: A Generative Approach to Segmenting Identical Instances in Depth Images
abstract
In this paper, we present InSeGAN, an unsupervised 3D generative adversarial network (GAN) for segmenting (nearly) identical instances of rigid objects in depth images. Using an analysis-by-synthesis approach, we design a novel GAN architecture to synthesize a multiple-instance depth image with independent control over each instance. InSeGAN takes in a set of code vectors (e.g., random noise vectors), each encoding the 3D pose of an object that is represented by a learned implicit object template. The generator has two distinct modules. The first module, the instance feature generator, uses each encoded pose to transform the implicit template into a feature map representation of each object instance. The second module, the depth image renderer, aggregates all of the single-instance feature maps output by the first module and generates a multiple-instance depth image. A discriminator distinguishes the generated multiple-instance depth images from the distribution of true depth images. To use our model for instance segmentation, we propose an instance pose encoder that learns to take in a generated depth image and reproduce the pose code vectors for all of the object instances. To evaluate our approach, we introduce a new synthetic dataset, "Insta-10," consisting of 100,000 depth images, each with 5 instances of an object from one of 10 classes. Our experiments on Insta-10, as well as on real-world noisy depth images, show that InSeGAN achieves state-of-the-art performance, often outperforming prior methods by large margins.
Anoop Cherian, Gonçalo Dias Pais, Siddarth Jain, Tim K. Marks, Alan Sullivan
ICCV4
2021 Turnip: Time-Series U-Net With Recurrence For NIR Imaging PPG
abstract
Imaging photoplethysmography (iPPG) is the process of estimating the waveform of a person’s pulse by processing a video of their face to detect minute color or intensity changes in the skin. Typically, iPPG methods use three-channel RGB video to address challenges due to motion. In situations such as driving, however, illumination in the visible spectrum is often quickly varying (e.g., daytime driving through shadows of trees and buildings) or insufficient (e.g., night driving). In such cases, a practical alternative is to use active illumination and bandpass-filtering from a monochromatic near-infrared (NIR) light source and camera. Contrary to learning-based iPPG solutions designed for multi-channel RGB, previous work in single-channel NIR iPPG has been based on hand-crafted models (with only a few manually tuned parameters), exploiting the sparsity of the PPG signal in the frequency domain. In contrast, we propose a modular framework for iPPG estimation of the heartbeat signal, in which the first module extracts a time-series signal from monochromatic NIR face video. The second module consists of a novel time-series U-net architecture in which a GRU (gated recurrent unit) network has been added to the passthrough layers. We test our approach on the challenging MR-NIRP Car Dataset, which consists of monochromatic NIR videos taken in both stationary and driving conditions. Our model’s iPPG estimation performance on NIR video outperforms both the state-of-the-art model-based method and a recent end-to-end deep learning method that we adapted to monochromatic video.
Armand Comas, Tim K. Marks, Hassan Mansour, Suhas Lohit, Yechi Ma, Xiaoming Liu 0002
ICIP2
2021 Overview of the Eighth Dialog System Technology Challenge: DSTC8
abstract
This paper introduces the Eighth Dialog System Technology Challenge. In line with recent challenges, the eighth edition focuses on applying end-to-end dialog technologies in a pragmatic way for multi-domain task-completion, noetic response selection, audio visual scene-aware dialog, and schema-guided dialog state tracking tasks. This paper describes the task definition, provided datasets, baselines and evaluation set-up for each track. We also summarize the results of the submitted systems to highlight the overall trends of the state-of-the-art technologies for the tasks.
Seokhwan Kim, Michel Galley, R. Chulaka Gunasekara, Adam Atkinson, Baolin Peng, Hannes Schulz, Jianfeng Gao 0001, Jinchao Li, Mahmoud Adada, Minlie Huang, Luis A. Lastras, Jonathan K. Kummerfeld, Walter S. Lasecki, Chiori Hori, Anoop Cherian, Tim K. Marks, Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara
IEEE ACM Trans. Audio Speech Lang. Process.17
2020 LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood
abstract
Modern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted locations nor predict whether landmarks are visible. In this paper, we present a novel framework for jointly predicting landmark locations, associated uncertainties of these predicted locations, and landmark visibilities. We model these as mixed random variables and estimate them using a deep network trained using our proposed Location, Uncertainty, and Visibility Likelihood (LUVLi) loss. In addition, we release an entirely new labeling of a large face alignment dataset with over 19,000 face images in a full range of head poses. Each face is manually labeled with the ground-truth locations of 68 landmarks, with the additional information of whether each landmarks is visible, self-occluded (due to extreme head poses), or externally occluded. Not only does our joint estimation yield accurate estimates of the uncertainty of predicted landmark locations, but it also yields state-of-the-art estimates for the landmark locations themselves on mulitple standard face alignment datasets. Our method's estimates of the uncertainty of predicted landmark locations could be used to automatically identify input images on which face alignment fails, which can be critical for downstream tasks.
Abhinav Kumar 0004, Tim K. Marks, Wenxuan Mou, Ye Wang 0001, Michael J. Jones 0001, Anoop Cherian, Toshiaki Koike-Akino, Xiaoming Liu 0002, Chen Feng 0002
CVPR2
2020 Spatio-Temporal Ranked-Attention Networks for Video Captioning
abstract
Generating video descriptions automatically is a challenging task that involves a complex interplay between spatio-temporal visual features and language models. Given that videos consist of spatial (frame-level) features and their temporal evolutions, an effective captioning model should be able to attend to these different cues selectively. To this end, we propose a Spatio-Temporal and Temporo-Spatial (STaTS) attention model which, conditioned on the language state, hierarchically combines spatial and temporal attention to videos in two different orders: (i) a spatiotemporal (ST) sub-model, which first attends to regions that have temporal evolution, then temporally pools the features from these regions; and (ii) a temporo-spatial (TS) sub-model, which first decides a single frame to attend to, then applies spatial attention within that frame. We propose a novel LSTM-based temporal ranking function, which we call ranked attention, for the ST model to capture action dynamics. Our entire framework is trained end-to-end. We provide experiments on two benchmark datasets: MSVD and MSR-VTT. Our results demonstrate the synergy between the ST and TS modules, outperforming recent state-of-the-art methods.
Anoop Cherian, Jue Wang 0010, Chiori Hori, Tim K. Marks
WACV4
2020 FX-GAN: Self-Supervised GAN Learning via Feature Exchange
abstract
We propose a self-supervised approach to improve the training of Generative Adversarial Networks (GANs) via inducing the discriminator to examine the structural consistency of images. Although natural image samples provide ideal examples of both valid structure and valid texture, learning to reproduce both together remains an open challenge. In our approach, we augment the training set of natural images with modified examples that have degraded structural consistency. These degraded examples are automatically created by randomly exchanging pairs of patches in an image’s convolutional feature map. We call this approach feature exchange. With this setup, we propose a novel GAN formulation, termed Feature eXchange GAN (FX-GAN), in which the discriminator is trained not only to distinguish real versus generated images, but also to perform the auxiliary task of distinguishing between real images and structurally corrupted (feature-exchanged) real images. This auxiliary task causes the discriminator to learn the proper feature structure of natural images, which in turn guides the generator to produce images with more realistic structure. Compared with strong GAN baselines, our proposed self-supervision approach improves generated image quality, diversity, and training stability for both the unconditional and class-conditional settings.
Wenju Xu, Teng-Yok Lee, Anoop Cherian, Ye Wang 0001, Tim K. Marks
WACV6
2020 Overview of the seventh Dialog System Technology Challenge: DSTC7
Luis Fernando D'Haro, Koichiro Yoshino, Chiori Hori, Tim K. Marks, Lazaros Polymenakos, Jonathan K. Kummerfeld, Michel Galley, Xiang Gao 0011
Comput. Speech Lang.4
2019 Audio Visual Scene-Aware Dialog
abstract
We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully, agents must ground concepts from the question in the video while leveraging contextual cues from the dialog history. To benchmark this task, we introduce the Audio Visual Scene-Aware Dialog (AVSD) Dataset. For each of more than 11,000 videos of human actions from the Charades dataset, our dataset contains a dialog about the video, plus a final summary of the video by one of the dialog participants. We train several baseline systems for this task and evaluate the performance of the trained models using both qualitative and quantitative metrics. Our results indicate that models must utilize all the available inputs (video, audio, question, and dialog history) to perform best on this dataset.
Huda AlAmri, Vincent Cartillier, Abhishek Das 0002, Jue Wang 0010, Anoop Cherian, Irfan A. Essa, Dhruv Batra, Tim K. Marks, Chiori Hori, Stefan Lee, Devi Parikh
CVPR8
2019 End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video Features
abstract
In order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differentiable network. As a result, Audio Visual Scene-Aware Dialog (AVSD) systems for real-world applications can be developed by integrating state-of-the-art technologies from multiple research areas, including end-to-end dialog technologies, visual question answering (VQA) technologies, and video description technologies. In this paper, we introduce a new data set of dialogs about videos of human behaviors, as well as an end-to-end Audio Visual Scene-Aware Dialog (AVSD) model, trained using this new data set, that generates responses in a dialog about a video. By using features that were developed for multimodal attention-based video description, our system improves the quality of generated dialog about dynamic video scenes.
Chiori Hori, Huda AlAmri, Jue Wang 0010, Gordon Wichern, Takaaki Hori, Anoop Cherian, Tim K. Marks, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das 0002, Irfan A. Essa, Dhruv Batra, Devi Parikh
ICASSP7
2019 Joint Student-Teacher Learning for Audio-Visual Scene-Aware Dialog
Chiori Hori, Anoop Cherian, Tim K. Marks, Takaaki Hori
INTERSPEECH3
2017 Early and late integration of audio features for automatic video description
abstract
This paper presents our approach to improve video captioning by integrating audio and video features. Video captioning is the task of generating a textual description to describe the content of a video. State-of-the-art approaches to video captioning are based on sequence-to-sequence models, in which a single neural network accepts sequential images and audio data, and outputs a sequence of words that best describe the input data in natural language. The network thus learns to encode the video input into an intermediate semantic representation, which can be useful in applications such as multimedia indexing, automatic narration, and audio-visual question answering. In our prior work, we proposed an attention-based multi-modal fusion mechanism to integrate image, motion, and audio features, where the multiple features are integrated in the network. Here, we apply hypothesis-level integration based on minimum Bayes-risk (MBR) decoding to further improve the caption quality, focusing on well-known evaluation metrics (BLEU and METEOR scores). Experiments with the YouTube2Text and MSR-VTT datasets demonstrate that combinations of early and late integration of multimodal features significantly improve the audio-visual semantic representation, as measured by the resulting caption quality. In addition, we compared the performance of our method using two different types of audio features: MFCC features, and the audio features extracted using SoundNet, which was trained to recognize objects and scenes from videos using only the audio signals.
Chiori Hori, Takaaki Hori, Tim K. Marks, John R. Hershey
ASRU3
2017 Attention-Based Multimodal Fusion for Video Description
Chiori Hori, Takaaki Hori, Teng-Yok Lee, Bret Harsham, John R. Hershey, Tim K. Marks, Kazuhiro Sumi
ICCV7
2016 A Multi-stream Bi-directional Recurrent Neural Network for Fine-Grained Action Detection
abstract
We present a multi-stream bi-directional recurrent neural network for fine-grained action detection. Recently, twostream convolutional neural networks (CNNs) trained on stacked optical flow and image frames have been successful for action recognition in videos. Our system uses a tracking algorithm to locate a bounding box around the person, which provides a frame of reference for appearance and motion and also suppresses background noise that is not within the bounding box. We train two additional streams on motion and appearance cropped to the tracked bounding box, along with full-frame streams. Our motion streams use pixel trajectories of a frame as raw features, in which the displacement values corresponding to a moving scene point are at the same spatial position across several frames. To model long-term temporal dynamics within and between actions, the multi-stream CNN is followed by a bi-directional Long Short-Term Memory (LSTM) layer. We show that our bi-directional LSTM network utilizes about 8 seconds of the video sequence to predict an action label. We test on two action detection datasets: the MPII Cooking 2 Dataset, and a new MERL Shopping Dataset that we introduce and make available to the community with this paper. The results demonstrate that our method significantly outperforms state-of-the-art action detection methods on both datasets.
Tim K. Marks, Michael J. Jones 0001, Oncel Tuzel, Ming Shao
CVPR2
2016 Robust Face Alignment Using a Mixture of Invariant Experts
Oncel Tuzel, Tim K. Marks, Salil Tambe
ECCV (5)2
2015 An improved deep learning architecture for person re-identification
abstract
In this work, we propose a method for simultaneously learning features and a corresponding similarity metric for person re-identification. We present a deep convolutional architecture with layers specially designed to address the problem of re-identification. Given a pair of images as input, our network outputs a similarity value indicating whether the two input images depict the same person. Novel elements of our architecture include a layer that computes cross-input neighborhood differences, which capture local relationships between the two input images based on mid-level features from each input image. A high-level summary of the outputs of this layer is computed by a layer of patch summary features, which are then spatially integrated in subsequent layers. Our method significantly outperforms the state of the art on both a large data set (CUHK03) and a medium-sized data set (CUHK01), and is resistant to over-fitting. We also demonstrate that by initially training on an unrelated large data set before fine-tuning on a small target data set, our network can achieve results comparable to the state of the art even on a small data set (VIPeR).
Ejaz Ahmed 0002, Michael J. Jones 0001, Tim K. Marks
CVPR3
2015 Real-time 3D head pose and facial landmark estimation from depth images using triangular surface patch features
abstract
We present a real-time system for 3D head pose estimation and facial landmark localization using a commodity depth sensor. We introduce a novel triangular surface patch (TSP) descriptor, which encodes the shape of the 3D surface of the face within a triangular area. The proposed descriptor is viewpoint invariant, and it is robust to noise and to variations in the data resolution. Using a fast nearest neighbor lookup, TSP descriptors from an input depth map are matched to the most similar ones that were computed from synthetic head models in a training phase. The matched triangular surface patches in the training set are used to compute estimates of the 3D head pose and facial landmark positions in the input depth map. By sampling many TSP descriptors, many votes for pose and landmark positions are generated which together yield robust final estimates. We evaluate our approach on the publicly available Biwi Kinect Head Pose Database to compare it against state-of-the-art methods. Our results show a significant improvement in the accuracy of both pose and landmark location estimates while maintaining real-time speed.
Chavdar Papazov, Tim K. Marks, Michael J. Jones 0001
CVPR2
2014 Log-linear dialog manager
abstract
We design a log-linear probabilistic model for solving the dialog management task. In both planning and learning we optimize the same objective function: the expected reward. Rather than performing full policy optimization, we perform on-line estimation of the optimal action as a belief-propagation inference step. We employ context-free grammars to describe our variable spaces, which enables us to define rich features. To scale our approach to large variable spaces, we use particle belief propagation. Experiments show that the model is able to choose system actions that yield a high expected reward, outperforming its POMDP-like log-linear counterpart and a hand-crafted rule-based system.
Hao Tang 0002, Shinji Watanabe 0001, Tim K. Marks, John R. Hershey
ICASSP3
2014 Cost-level integration of statistical and rule-based dialog managers
Shinji Watanabe 0001, John R. Hershey, Tim K. Marks, Youichi Fujii, Yusuke Koji
INTERSPEECH3
2014 Detecting 3D geometric boundaries of indoor scenes under varying lighting
abstract
The goal of this research is to identify 3D geometric boundaries in a set of 2D photographs of a static indoor scene under unknown, changing lighting conditions. A 3D geometric boundary is a contour located at a 3D depth discontinuity or a discontinuity in the surface normal. These boundaries can be used effectively for reasoning about the 3D layout of a scene. To distinguish 3D geometric boundaries from 2D texture edges, we analyze the illumination subspace of local appearance at each image location. In indoor time-lapse photography and surveillance video, we frequently see images that are lit by unknown combinations of uncalibrated light sources. We introduce an algorithm for semi-binary nonnegative matrix factorization (SBNMF) to decompose such images into a set of lighting basis images, each of which shows the scene lit by a single light source. These basis images provide a natural, succinct representation of the scene, enabling tasks such as scene editing (e.g., relighting) and shadow edge identification.
Jie Ni, Tim K. Marks, Oncel Tuzel, Fatih Porikli
WACV2
2011 Pose Normalization via Learned 2D Warping for Fully Automatic Face Recognition
abstract
We present a novel approach to pose-invariant face recognition that handles continuous pose variations, is not database-specific, and achieves high accuracy without any manual intervention. Our method uses multidimensional Gaussian process regression to learn a nonlinear mapping function from the 2D shapes of faces at any non-frontal pose to the corresponding 2D frontal face shapes. We use this mapping to take an input image of a new face at an arbitrary pose and pose-normalize it, generating a synthetic frontal image of the face that is then used for recognition. Our fully automatic system for face recognition includes automatic methods for extracting 2D facial feature points and accurately estimating 3D head pose, and this information is used as input to the 2D pose-normalization algorithm. The current system can handle pose variation up to 45 degrees to the left or right (yaw angle) and up to 30 degrees up or down (pitch angle). The system demonstrates high accuracy in recognition experiments on the CMU-PIE, USF 3D, and Multi-PIE databases, showing excellent generalization across databases and convincingly outperforming other automatic methods.
Akshay Asthana, Michael J. Jones 0001, Tim K. Marks, Kinh H. Tieu, Roland Göcke
BMVC3
2011 Fully automatic pose-invariant face recognition via 3D pose normalization
abstract
An ideal approach to the problem of pose-invariant face recognition would handle continuous pose variations, would not be database specific, and would achieve high accuracy without any manual intervention. Most of the existing approaches fail to match one or more of these goals. In this paper, we present a fully automatic system for pose-invariant face recognition that not only meets these requirements but also outperforms other comparable methods. We propose a 3D pose normalization method that is completely automatic and leverages the accurate 2D facial feature points found by the system. The current system can handle 3D pose variation up to ±45° in yaw and ±30° in pitch angles. Recognition experiments were conducted on the USF 3D, Multi-PIE, CMU-PIE, FERET, and FacePix databases. Our system not only shows excellent generalization by achieving high accuracy on all 5 databases but also outperforms other methods convincingly.
Akshay Asthana, Tim K. Marks, Michael J. Jones 0001, Kinh H. Tieu, M. V. Rohith
ICCV2
2011 Entropy-based motion selection for touch-based registration using Rao-Blackwellized particle filtering
abstract
To achieve versatile locomotion in complex amphibious environments, a robot should be capable of performing different gaits. In this paper we present such a versatile amphibious robot based on a novel eccentric paddle mechanism (ePaddle). We first illustrate the concept of the ePaddle with five major possible gaits and conceptual gait sequences. We then summarize five types of configurations from these gaits. Based on these configurations, two motion behaviors are found and modeled by using kinematic equations for the future gait planning tasks. To verify the proposed ideas, we develop an ePaddle prototype module. Several simulations on these gaits are performed to verify the conceptual locomotion gait and the developed kinematic models. Experiments on five possible configurations demonstrate the valid of the ePaddle concept and the prototype design.
Yuichi Taguchi, Tim K. Marks, John R. Hershey
IROS2
2010 Morphable Reflectance Fields for enhancing face recognition
abstract
In this paper, we present a novel framework to address the confounding effects of illumination variation in face recognition. By augmenting the gallery set with realistically relit images, we enhance recognition performance in a classifier-independent way. We describe a novel method for single-image relighting, Morphable Reflectance Fields (MoRF), which does not require manual intervention and provides relighting superior to that of existing automatic methods. We test our framework through face recognition experiments using various state-of-the-art classifiers and popular benchmark datasets: CMU PIE, Multi-PIE, and MERL Dome. We demonstrate that our MoRF relighting and gallery augmentation framework achieves improvements in terms of both rank-1 recognition rates and ROC curves. We also compare our model with other automatic relighting methods to confirm its advantage. Finally, we show that the recognition rates achieved using our framework exceed those of state-of-the-art recognizers on the aforementioned databases.
Ritwik Kumar, Michael J. Jones 0001, Tim K. Marks
CVPR3
2010 P2Pi: A Minimal Solution for Registration of 3D Points to 3D Planes
Srikumar Ramalingam, Yuichi Taguchi, Tim K. Marks, Oncel Tuzel
ECCV (5)3
2010 Rao-Blackwellized particle filtering for probing-based 6-DOF localization in robotic assembly
abstract
This paper presents a probing-based method for probabilistic localization in automated robotic assembly. We consider peg-in-hole problems in which a needle-like peg has a single point of contact with the object that contains the hole, and in which the initial uncertainty in the relative pose (3D position and 3D angle) between the peg and the object is much greater than the required accuracy (assembly clearance). We solve this 6 degree-of-freedom (6-DOF) localization problem using a Rao-Blackwellized particle filter, in which the probability distribution over the peg's pose is factorized into two components: The distribution over position (3-DOF) is represented by particles, while the distribution over angle (3-DOF) is approximated as a Gaussian distribution for each particle, updated using an extended Kalman filter. This factorization reduces the number of particles required for localization by orders of magnitude, enabling real-time online 6-DOF pose estimation. Each measurement is simply the contact position obtained by randomly repositioning the peg and moving towards the object until there is contact. To compute the likelihood of each measurement, we use as a map a mesh model of the object that is based on the CAD model but also explicitly models the uncertainty in the map. The mesh uncertainty model makes our system robust to cases in which the actual measurement is different from the expected one. We demonstrate the advantages of our approach over previous methods using simulations as well as physical experiments with a robotic arm and a metal peg and object.
Yuichi Taguchi, Tim K. Marks, Haruhisa Okuda
ICRA2
2010 Tracking Motion, Deformation, and Texture Using Conditionally Gaussian Processes
abstract
We present a generative model and inference algorithm for 3D nonrigid object tracking. The model, which we call G-flow, enables the joint inference of 3D position, orientation, and nonrigid deformations, as well as object texture and background texture. Optimal inference under G-flow reduces to a conditionally Gaussian stochastic filtering problem. The optimal solution to this problem reveals a new space of computer vision algorithms, of which classic approaches such as optic flow and template matching are special cases that are optimal only under special circumstances. We evaluate G-flow on the problem of tracking facial expressions and head motion in 3D from single-camera video. Previously, the lack of realistic video data with ground truth nonrigid position information has hampered the rigorous evaluation of nonrigid tracking. We introduce a practical method of obtaining such ground truth data and present a new face video data set that was created using this technique. Results on this data set show that G-flow is much more robust and accurate than current deterministic optic-flow-based approaches.
Tim K. Marks, John R. Hershey, Javier R. Movellan
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Gamma-SLAM: Using stereo vision and variance grid maps for SLAM in unstructured environments
abstract
We introduce a new method for stereo visual SLAM (simultaneous localization and mapping) that works in unstructured, outdoor environments. Unlike other grid-based SLAM algorithms, which use occupancy grid maps, our algorithm uses a new mapping technique that maintains a posterior distribution over the height variance in each cell. This idea was motivated by our experience with outdoor navigation tasks, which has shown height variance to be a useful measure of traversability. To obtain a joint posterior over poses and maps, we use a Rao-Blackwellized particle filter: the pose distribution is estimated using a particle filter, and each particle has its own map that is obtained through exact filtering conditioned on the particle's pose. Visual odometry provides good proposal distributions for the particle pose. In the analytical (exact) filter for the map, we update the sufficient statistics of a gamma distribution over the precision (inverse variance) of heights in each grid cell. We verify the algorithm's accuracy on two outdoor courses by comparing with ground truth data obtained using electronic surveying equipment. In addition, we solve for the optimal transformation from the SLAM map to georeferenced coordinates, based on a noisy GPS signal. We derive an online version of this alignment process, which can be used to maintain a running estimate of the robot's global position that is much more accurate than the GPS readings.
Tim K. Marks, Max Bajracharya, Garrison W. Cottrell, Larry H. Matthies
ICRA1
2004 Joint Tracking of Pose, Expression, and Texture using Conditionally Gaussian Filters
abstract
We present a generative model and stochastic filtering algorithm for si- multaneous tracking of 3D position and orientation, non-rigid motion, object texture, and background texture using a single camera. We show that the solution to this problem is formally equivalent to stochastic fil- tering of conditionally Gaussian processes, a problem for which well known approaches exist [3, 8]. We propose an approach based on Monte Carlo sampling of the nonlinear component of the process (object mo- tion) and exact filtering of the object and background textures given the sampled motion. The smoothness of image sequences in time and space is exploited by using Laplace's method to generate proposal distributions for importance sampling [7]. The resulting inference algorithm encom- passes both optic flow and template-based tracking as special cases, and elucidates the conditions under which these methods are optimal. We demonstrate an application of the system to 3D non-rigid face tracking.
Tim K. Marks, John R. Hershey, J. Cooper Roddey, Javier R. Movellan
NIPS1