VLDB 2026 Research / reviewers in the wild / expert
Richard J. Radke
dblp:21/356
· DBLP profile ↗
53ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0001-5064-7775ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 23 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Utterance: Understanding Group Problem Solving through Discussion Sequences
Zhuoxu Duan, Zhengye Yang, Brooke Foucault Welles, Richard J. Radke |
ICMI | 4 |
| 2024 | Few-Shot 3D Volumetric Segmentation with Multi-surrogate Fusion
Meng Zheng 0002, Benjamin Planche, Zhongpai Gao, Terrence Chen, Richard J. Radke, Ziyan Wu 0001 |
MICCAI (9) | 5 |
| 2023 | Self-supervised Learning with Local Contrastive Loss for Detection and Semantic SegmentationabstractWe present a self-supervised learning (SSL) method suitable for semi-global tasks such as object detection and semantic segmentation. We enforce local consistency between self-learned features that represent corresponding image locations of transformed versions of the same image, by minimizing a pixel-level local contrastive (LC) loss during training. LC-loss can be added to existing self-supervised learning methods with minimal overhead. We evaluate our SSL approach on two downstream tasks – object detection and semantic segmentation, using COCO, PASCAL VOC, and CityScapes datasets. Our method outperforms the existing state-of-the-art SSL approaches by 1.9% on COCO object detection, 1.4% on PASCAL VOC detection, and 0.6% on CityScapes segmentation. Ashraful Islam, Ben Lundell, Harpreet Sawhney, Sudipta Sinha, Peter Morales, Richard J. Radke |
WACV | 6 |
| 2022 | Visual Similarity AttentionabstractWhile there has been substantial progress in learning suitable distance metrics, these techniques in general lack transparency and decision reasoning, i.e., explaining why the input set of images is similar or dissimilar. In this work, we solve this key problem by proposing the first method to generate generic visual similarity explanations with gradient-based attention. We demonstrate that our technique is agnostic to the specific similarity model type, e.g., we show applicability to Siamese, triplet, and quadruplet models. Furthermore, we make our proposed similarity attention a principled part of the learning process, resulting in a new paradigm for learning similarity functions. We demonstrate that our learning mechanism results in more generalizable, as well as explainable, similarity models. Finally, we demonstrate the generality of our framework by means of experiments on a variety of tasks, including image retrieval, person re-identification, and low-shot semantic segmentation. Meng Zheng 0002, Srikrishna Karanam, Terrence Chen, Richard J. Radke, Ziyan Wu 0001 |
IJCAI | 4 |
| 2022 | Natural Language Video Moment Localization Through Query-Controlled Temporal ConvolutionabstractThe goal of natural language video moment localization is to locate a short segment of a long, untrimmed video that corresponds to a description presented as natural text. The description may contain several pieces of key information, including subjects/objects, sequential actions, and locations. Here, we propose a novel video moment localization framework based on the convolutional response between multimodal signals, i.e., the video sequence, the text query, and subtitles for the video if they are available. We emphasize the effect of the language sequence as a query about the video content, by converting the query sentence into a boundary detector with a filter kernel size and stride. We convolve the video sequence with the query detector to locate the start and end boundaries of the target video segment. When subtitles are available, we blend the boundary heatmaps from the visual and subtitle branches together using an LSTM to capture asynchronous dependencies across two modalities in the video. We perform extensive experiments on the TVR, Charades-STA, and TACoS benchmark datasets, demonstrating that our model achieves state-of-the-art results on all three. Lingyu Zhang 0002, Richard J. Radke |
WACV | 2 |
| 2021 | A Hybrid Attention Mechanism for Weakly-Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization is a challenging vision task due to the absence of ground-truth temporal locations of actions in the training videos. With only video-level supervision during training, most existing methods rely on a Multiple Instance Learning (MIL) framework to predict the start and end frame of each action category in a video. However, the existing MIL-based approach has a major limitation of only capturing the most discriminative frames of an action, ignoring the full extent of an activity. Moreover, these methods cannot model background activity effectively, which plays an important role in localizing foreground activities. In this paper, we present a novel framework named HAM-Net with a hybrid attention mechanism which includes temporal soft, semi-soft and hard attentions to address these issues. Our temporal soft attention module, guided by an auxiliary background class in the classification module, models the background activity by introducing an ``action-ness'' score for each video snippet. Moreover, our temporal semi-soft and hard attention modules, calculating two attention scores for each video snippet, help to focus on the less discriminative frames of an action to capture the full action boundary. Our proposed approach outperforms recent state-of-the-art methods by at least 2.2% mAP at IoU threshold 0.5 on the THUMOS14 dataset, and by at least 1.3% mAP at IoU threshold 0.75 on the ActivityNet1.2 dataset. Ashraful Islam, Chengjiang Long, Richard J. Radke |
AAAI | 3 |
| 2021 | A Broad Study on the Transferability of Visual Representations with Contrastive LearningabstractTremendous progress has been made in visual representation learning, notably with the recent success of self-supervised contrastive learning methods. Supervised contrastive learning has also been shown to outperform its cross-entropy counterparts by leveraging labels for choosing where to contrast. However, there has been little work to explore the transfer capability of contrastive learning to a different domain. In this paper, we conduct a comprehensive study on the transferability of learned representations of different contrastive approaches for linear evaluation, full-network transfer, and few-shot recognition on 12 downstream datasets from different domains, and object detection tasks on MSCOCO and VOC0712. The results show that the contrastive approaches learn representations that are easily transferable to a different downstream task. We further observe that the joint objective of self-supervised contrastive loss with cross-entropy/supervised-contrastive loss leads to better transferability of these models over their supervised counterparts. Our analysis reveals that the representations learned from the contrastive approaches contain more low/mid-level semantics than cross-entropy models, which enables them to quickly adapt to a new task. Our codes and models will be publicly available to facilitate future research on transferability of visual representations.1 Ashraful Islam, Chun-Fu Chen 0001, Rameswar Panda, Leonid Karlinsky, Richard J. Radke, Rogério Feris |
ICCV | 5 |
| 2021 | Dynamic Distillation Network for Cross-Domain Few-Shot Recognition with Unlabeled DataabstractMost existing works in few-shot learning rely on meta-learning the network on a large base dataset which is typically from the same domain as the target dataset. We tackle the problem of cross-domain few-shot learning where there is a large shift between the base and target domain. The problem of cross-domain few-shot recognition with unlabeled target data is largely unaddressed in the literature. STARTUP was the first method that tackles this problem using self-training. However, it uses a fixed teacher pretrained on a labeled base dataset to create soft labels for the unlabeled target samples. As the base dataset and unlabeled dataset are from different domains, projecting the target images in the class-domain of the base dataset with a fixed pretrained model might be sub-optimal. We propose a simple dynamic distillation-based approach to facilitate unlabeled images from the novel/base dataset. We impose consistency regularization by calculating predictions from the weakly-augmented versions of the unlabeled images from a teacher network and matching it with the strongly augmented versions of the same images from a student network. The parameters of the teacher network are updated as exponential moving average of the parameters of the student network. We show that the proposed network learns representation that can be easily adapted to the target domain even though it has not been trained with target-specific classes during the pretraining phase. Our model outperforms the current state-of-the art method by 4.4% for 1-shot and 3.6% for 5-shot classification in the BSCD-FSL benchmark, and also shows competitive performance on traditional in-domain few-shot learning task. Ashraful Islam, Chun-Fu Chen 0001, Rameswar Panda, Leonid Karlinsky, Rogério Feris, Richard J. Radke |
NeurIPS | 6 |
| 2020 | Towards Visually Explaining Variational AutoencodersabstractRecent advances in Convolutional Neural Network (CNN) model interpretability have led to impressive progress in visualizing and understanding model predictions. In particular, gradient-based visual attention methods have driven much recent effort in using visual attention maps as a means for visual explanations. A key problem, however, is these methods are designed for classification and categorization tasks, and their extension to explaining generative models, e.g., variational autoencoders (VAE) is not trivial. In this work, we take a step towards bridging this crucial gap, proposing the first technique to visually explain VAEs by means of gradient-based attention. We present methods to generate visual attention from the learned latent space, and also demonstrate such attention explanations serve more than just explaining VAE predictions. We show how these attention maps can be used to localize anomalies in images, demonstrating state-of-the-art performance on the MVTec-AD dataset. We also show how they can be infused into model training, helping bootstrap the VAE into learning improved latent space disentanglement, demonstrated on the Dsprites dataset. WenQian Liu, Runze Li 0003, Meng Zheng 0002, Srikrishna Karanam, Ziyan Wu 0001, Bir Bhanu, Richard J. Radke, Octavia I. Camps |
CVPR | 7 |
| 2020 | Temporal Attention and Consistency Measuring for Video Question AnsweringabstractSocial signal processing algorithms have become increasingly better at solving well-defined prediction and estimation problems in audiovisual recordings of group discussion. However, much human behavior and communication is less structured and more subtle. In this paper, we address the problem of generic question answering from diverse audiovisual recordings of human interaction. The goal is to select the correct free-text answer to a free-text question about human interaction in a video. We propose an RNN-based model with two novel ideas: a temporal attention module that highlights key words and phrases in the question and candidate answers, and a consistency measurement module that scores the similarity between the multimodal data, the question, and the candidate answers. This small set of consistency scores forms the input to the final question-answering stage, resulting in a lightweight model. We demonstrate that our model achieves state of the art accuracy on the Social-IQ dataset containing hundreds of videos and question/answer pairs. Lingyu Zhang 0002, Richard J. Radke |
ICMI | 2 |
| 2020 | Weakly Supervised Temporal Action Localization Using Deep Metric LearningabstractTemporal action localization is an important step towards video understanding. Most current action localization methods depend on untrimmed videos with full temporal annotations of action instances. However, it is expensive and time-consuming to annotate both action labels and temporal boundaries of videos. To this end, we propose a weakly supervised temporal action localization method that only requires video-level action instances as supervision during training. We propose a classification module to generate action labels for each segment in the video, and a deep metric learning module to learn the similarity between different action instances. We jointly optimize a balanced binary cross-entropy loss and a metric loss using a standard backpropagation algorithm. Extensive experiments demonstrate the effectiveness of both of these components in temporal localization. We evaluate our algorithm on two challenging untrimmed video datasets: THUMOS14 and ActivityNet1.2. Our approach improves the current state-of-the-art result for THUMOS14 by 6.5% mAP at IoU threshold 0.5, and achieves competitive performance for ActivityNet1.2. Ashraful Islam, Richard J. Radke |
WACV | 2 |
| 2020 | Multiparty Visual Co-Occurrences for Estimating Personality Traits in Group MeetingsabstractParticipants’ body language during interactions with others in a group meeting can reveal important information about their individual personalities, as well as their contribution to a team. Here, we focus on the automatic extraction of visual features from each person, including her/his facial activity, body movement, and hand position, and how these features co-occur among team members (e.g., howfre- quently a person moves her/his arms or makes eye contact when she/he is the focus of attention of the group). We correlate these features with user questionnaires to reveal relationships with the "Big Five" personality traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroti- cism), as well as with team judgements about the leader and dominant contributor in a conversation. We demonstrate that our algorithms achieve state-of-the-art accuracy with an average of 80% for Big-Five personality trait prediction, potentially enabling integration into automatic group meeting understanding systems. Lingyu Zhang 0002, Indrani Bhattacharya, Mallory Morgan, Michael Foley, Christoph Riedl, Brooke Foucault Welles, Richard J. Radke |
WACV | 7 |
| 2019 | Keep Meeting Summaries on Topic: Abstractive Multi-Modal Meeting SummarizationabstractTranscripts of natural, multi-person meetings differ significantly from documents like news articles, which can make Natural Language Generation models generate unfocused summaries.We develop an abstractive meeting summarizer from both videos and audios of meeting recordings.Specifically, we propose a multi-modal hierarchical attention mechanism across three levels: topic segment, utterance and word.To narrow down the focus into topically-relevant segments, we jointly model topic segmentation and summarization.In addition to traditional textual features, we introduce new multi-modal features derived from visual focus of attention, based on the assumption that an utterance is more important if its speaker receives more attention.Experiments show that our model significantly outperforms the state-of-the-art with both BLEU and ROUGE measures. Manling Li, Lingyu Zhang 0002, Heng Ji 0001, Richard J. Radke |
ACL (1) | 4 |
| 2019 | Re-Identification With Consistent Attentive Siamese NetworksabstractWe propose a new deep architecture for person re-identification (re-id). While re-id has seen much recent progress, spatial localization and view-invariant representation learning for robust cross-view matching remain key, unsolved problems. We address these questions by means of a new attention-driven Siamese learning architecture, called the Consistent Attentive Siamese Network. Our key innovations compared to existing, competing methods include (a) a flexible framework design that produces attention with only identity labels as supervision, (b) explicit mechanisms to enforce attention consistency among images of the same person, and (c) a new Siamese framework that integrates attention and attention consistency, producing principled supervisory signals as well as the first mechanism that can explain the reasoning behind the Siamese framework's predictions. We conduct extensive evaluations on the CUHK03-NP, DukeMTMC-ReID, and Market-1501 datasets and report competitive performance. Meng Zheng 0002, Srikrishna Karanam, Ziyan Wu 0001, Richard J. Radke |
CVPR | 4 |
| 2019 | Improved Visual Focus of Attention Estimation and Prosodic Features for Analyzing Group InteractionsabstractCollaborative group tasks require efficient and productive verbal and non-verbal interactions among the participants. Studying such interaction patterns could help groups perform more efficiently, but the detection and measurement of human behavior is challenging since it is inherently multimodal and changes on a millisecond time frame. In this paper, we present a method to study groups performing a collaborative decision-making task using non-verbal behavioral cues. First, we present a novel algorithm to estimate the visual focus of attention (VFOA) of participants using frontal cameras. The algorithm can be used in various group settings, and performs with a state-of-the-art accuracy of 90%. Secondly, we present prosodic features for non-verbal speech analysis. These features are commonly used in speech/music classification tasks, but are rarely used in human group interaction analysis. We validate our algorithms on a multimodal dataset of 14 group meetings with 45 participants, and show that a combination of VFOA-based visual metrics and prosodic-feature-based metrics can predict emergent group leaders with 64% accuracy and dominant contributors with 86% accuracy. We also report our findings on the correlations between the non-verbal behavioral metrics with gender, emotional intelligence, and the Big 5 personality traits. Lingyu Zhang 0002, Mallory Morgan, Indrani Bhattacharya, Michael Foley, Jonas Braasch, Christoph Riedl, Brooke Foucault Welles, Richard J. Radke |
ICMI | 8 |
| 2019 | The unobtrusive group interaction (UGI) corpusabstractStudying group dynamics requires fine-grained spatial and temporal understanding of human behavior. Social psychologists studying human interaction patterns in face-to-face group meetings often find themselves struggling with huge volumes of data that require many hours of tedious manual coding. There are only a few publicly available multi-modal datasets of face-to-face group meetings that enable the development of automated methods to study verbal and non-verbal human behavior. In this paper, we present a new, publicly available multi-modal dataset for group dynamics study that differs from previous datasets in its use of ceiling-mounted, unobtrusive depth sensors. These can be used for fine-grained analysis of head and body pose and gestures, without any concerns about participants' privacy or inhibited behavior. The dataset is complemented by synchronized and time-stamped meeting transcripts that allow analysis of spoken content. The dataset comprises 22 group meetings in which participants perform a standard collaborative group task designed to measure leadership and productivity. Participants' post-task questionnaires, including demographic information, are also provided as part of the dataset. We show the utility of the dataset in analyzing perceived leadership, contribution, and performance, by presenting results of multi-modal analysis using our sensor-fusion algorithms designed to automatically understand audio-visual interactions. Indrani Bhattacharya, Michael Foley, Christine Ku, Tongtao Zhang, Cameron Mine, Manling Li, Heng Ji 0001, Christoph Riedl, Brooke Foucault Welles, Richard J. Radke |
MMSys | 11 |
| 2019 | A Systematic Evaluation and Benchmark for Person Re-Identification: Features, Metrics, and DatasetsabstractPerson re-identification (re-id) is a critical problem in video analytics applications such as security and surveillance. The public release of several datasets and code for vision algorithms has facilitated rapid progress in this area over the last few years. However, directly comparing re-id algorithms reported in the literature has become difficult since a wide variety of features, experimental protocols, and evaluation metrics are employed. In order to address this need, we present an extensive review and performance evaluation of single- and multi-shot re-id algorithms. The experimental protocol incorporates the most recent advances in both feature extraction and metric learning. To ensure a fair comparison, all of the approaches were implemented using a unified code library that includes 11 feature extraction algorithms and 22 metric learning and ranking techniques. All approaches were evaluated using a new large-scale dataset that closely mimics a real-world problem setting, in addition to 16 other publicly available datasets: VIPeR, GRID, CAVIAR, DukeMTMC4ReID, 3DPeS, PRID, V47, WARD, SAIVT-SoftBio, CUHK01, CHUK02, CUHK03, RAiD, iLIDSVID, HDA+, and Market1501. The evaluation codebase and results will be made publicly available for community use. Srikrishna Karanam, Mengran Gou, Ziyan Wu 0001, Angels Rates-Borras, Octavia I. Camps, Richard J. Radke |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2018 | A Multimodal-Sensor-Enabled Room for Unobtrusive Group Meeting AnalysisabstractGroup meetings can suffer from serious problems that undermine performance, including bias, "groupthink", fear of speaking, and unfocused discussion. To better understand these issues, propose interventions, and thus improve team performance, we need to study human dynamics in group meetings. However, this process currently heavily depends on manual coding and video cameras. Manual coding is tedious, inaccurate, and subjective, while active video cameras can affect the natural behavior of meeting participants. Here, we present a smart meeting room that combines microphones and unobtrusive ceiling-mounted Time-of-Flight (ToF) sensors to understand group dynamics in team meetings. We automatically process the multimodal sensor outputs with signal, image, and natural language processing algorithms to estimate participant head pose, visual focus of attention (VFOA), non-verbal speech patterns, and discussion content. We derive metrics from these automatic estimates and correlate them with user-reported rankings of emergent group leaders and major contributors to produce accurate predictors. We validate our algorithms and report results on a new dataset of lunar survival tasks of 36 individuals across 10 groups collected in the multimodal-sensor-enabled smart room. Indrani Bhattacharya, Michael Foley, Tongtao Zhang, Christine Ku, Cameron Mine, Heng Ji 0001, Christoph Riedl, Brooke Foucault Welles, Richard J. Radke |
ICMI | 10 |
| 2018 | Learning Affine Hull Representations for Multi-Shot Person Re-IdentificationabstractWe consider the person re-identification problem, assuming the availability of a sequence of images for each person, commonly referred to as video-based or multi-shot re-identification. We approach this problem from the perspective of learning discriminative distance metric functions. While existing distance metric learning methods typically employ the average feature vector as the data exemplar, this discards the inherent structure of the data. To overcome this issue, we describe the image sequence data using affine hulls. We show that directly computing the distance between the closest points on these affine hulls as in existing recognition algorithms is not sufficiently discriminative in the context of person re-identification. To this end, we incorporate affine hull data modeling into the traditional distance metric learning framework, learning discriminative feature representations directly using affine hulls. We perform extensive experiments on several publicly available data sets to show that the proposed approach improves the performance of existing metric learning algorithms irrespective of the feature space employed to perform metric learning. Furthermore, we advance the state of the art on iLIDS-VID, PRID, and SAIVT, with absolute rank-1 performance improvements of 6.0%, 11.4%, and 6.0% respectively. Srikrishna Karanam, Ziyan Wu 0001, Richard J. Radke |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Person re-identification with block sparse recovery
Srikrishna Karanam, Yang Li 0056, Richard J. Radke |
Image Vis. Comput. | 3 |
| 2017 | From the Lab to the Real World: Re-identification in an Airport Camera NetworkabstractOver the past ten years, human re-identification has received increased attention from the computer vision research community. However, for the most part, these research papers are divorced from the context of how such algorithms would be used in a real-world system. This paper describes the unique opportunity our group of academic researchers had to design and deploy a human re-identification system in a demanding real-world environment: a busy airport. The system had to be designed from the ground up, including robust modules for real-time human detection and tracking, a distributed, low-latency software architecture, and a front-end user interface designed for a specific scenario. None of these issues are typically addressed in re-identification research papers, but all are critical to an effective system that end users would actually be willing to adopt. We detail the challenges of the real-world airport environment, the computer vision algorithms underlying our human detection and re-identification algorithms, our robust software architecture, and the ground-truthing system required to provide the training and validation data for the algorithms. Our initial results show that despite the challenges and constraints of the airport environment, the proposed system achieves very good performance while operating in real time. Octavia I. Camps, Mengran Gou, Tom Hebble, Srikrishna Karanam, Oliver Lehmann, Yang Li 0056, Richard J. Radke, Ziyan Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2016 | Sensor fusion for occupancy detection and activity recognition using time-of-flight sensors
Tianna-Kaye Woodstock, Richard J. Radke, Arthur C. Sanderson |
FUSION | 2 |
| 2016 | Arrays of single pixel time-of-flight sensors for privacy preserving tracking and coarse pose estimationabstractWe present a method for real-time person tracking and coarse pose estimation in a smart room using a sparse array of single pixel time-of flight (ToF) sensors mounted in the ceiling of the room. The single pixel sensors are relatively inexpensive compared to commercial ToF cameras and are privacy preserving in that they only return the range to a small set of hit points. The tracking algorithm includes higher level logic about how people move and interact in a room and makes estimates about the locations of people even in the absence of direct measurements. A maximum likelihood classifier based on features extracted from the time series of ToF measurements is used for robust pose classification into sitting, standing and walking states. We use both computer simulation and real-world experiments to show that the algorithms are capable of robust person tracking and pose estimation even with a sensor spacing of 60 cm (i.e., 1 sensor per ceiling tile). Indrani Bhattacharya, Richard J. Radke |
WACV | 2 |
| 2015 | Particle dynamics and multi-channel feature dictionaries for robust visual trackingabstractWe present a novel approach to solve the visual tracking problem in a particle filter framework based on sparse visual representations. Current state-of-the-art trackers use low-resolution image intensity features in target appearance modeling. Such features of-ten fail to capture sufficient visual information about the target. Here, we demonstrate the efficacy of visually richer representation schemes by employing multi-channel fea-ture dictionaries as part of the appearance model. To further mitigate the tracking drift problem, we propose a novel dynamic adaptive state transition model, taking into account the dynamics of the past states. Finally, we demonstrate the computational tractability of using richer appearance modeling schemes by adaptively pruning candidate particles during each sampling step, and using a fast augmented Lagrangian technique to solve the associated optimization problem. Extensive quantitative evaluations and robustness tests on several challenging video sequences demonstrate that our approach substantially outperforms the state of the art, and achieves stable results. 1 Srikrishna Karanam, Yang Li 0056, Richard J. Radke |
BMVC | 3 |
| 2015 | Multi-Shot Human Re-Identification Using Adaptive Fisher Discriminant Analysis
Yang Li 0056, Ziyan Wu 0001, Srikrishna Karanam, Richard J. Radke |
BMVC | 4 |
| 2015 | Person Re-Identification with Discriminatively Trained Viewpoint Invariant DictionariesabstractThis paper introduces a new approach to address the person re-identification problem in cameras with non-overlapping fields of view. Unlike previous approaches that learn Mahalanobis-like distance metrics in some transformed feature space, we propose to learn a dictionary that is capable of discriminatively and sparsely encoding features representing different people. Our approach directly addresses two key challenges in person re-identification: viewpoint variations and discriminability. First, to tackle viewpoint and associated appearance changes, we learn a single dictionary to represent both gallery and probe images in the training phase. We then discriminatively train the dictionary by enforcing explicit constraints on the associated sparse representations of the feature vectors. In the testing phase, we re-identify a probe image by simply determining the gallery image that has the closest sparse representation to that of the probe image in the Euclidean sense. Extensive performance evaluations on three publicly available multi-shot re-identification datasets demonstrate the advantages of our algorithm over several state-of-the-art dictionary learning, temporal sequence matching, and spatial appearance and metric learning based techniques. Srikrishna Karanam, Yang Li 0056, Richard J. Radke |
ICCV | 3 |
| 2015 | Collaborative human-robot manipulation of highly deformable materialsabstractRobotic manipulation of highly deformable materials is inherently challenging due to the need to maintain tension and the high dimensionality of the state of the material. Past work in this area mostly focuses on generating a detailed model for the material and its interaction with the robot, then using the model to construct a motion plan. In this paper, we take a different approach by using only sensor feedback to dictate the robot motion. We consider the collaborative manipulation of a deformable sheet between a person and a dual-armed robot (Baxter by Rethink Robotics). The robot is capable of contact sensing via joint torque sensors and is equipped with a head-mounted RGBd sensor. The robot senses contact force to maintain tension of the sheet, and in turn comply to the human motion. This is akin to handling a tablecloth with a partner but with one's eyes closed. To improve the response, we use the RGBd sensor to detect folds, and command the robot to move in an orthogonal direction to smooth them out. This is like handling cloth by looking at the cloth itself. Both controllers are able to follow human motion without excessive crimps in the sheet, but as expected, the hybrid controller combining force and vision outperforms the force controller alone in terms of tension force transient. The ability to quickly detect the state of the deformable material also enables more complex manipulation strategies in the future. Daniel Kruse, Richard J. Radke, John T. Wen |
ICRA | 2 |
| 2015 | Multi-shot Re-identification with Random-Projection-Based Random ForestsabstractHuman re-identification remains one of the fundamental, difficult problems in video surveillance and analysis. Current metric learning algorithms mainly focus on finding an optimized vector space such that observations of the same person in this space have a smaller distance than observations of two different people. In this paper, we propose a novel metric learning approach to the human reidentification problem, with an emphasis on the multi-shot scenario. First, we perform dimensionality reduction on image feature vectors through random projection. Next, a random forest is trained based on pair wise constraints in the projected subspace. This procedure repeats with a number of random projection bases, so that a series of random forests are trained in various feature subspaces. Finally, we select personalized random forests for each subject using their multi-shot appearances. We evaluate the performance of our algorithm on three benchmark datasets. Yang Li 0056, Ziyan Wu 0001, Richard J. Radke |
WACV | 3 |
| 2015 | Viewpoint Invariant Human Re-Identification in Camera Networks Using Pose Priors and Subject-Discriminative FeaturesabstractHuman re-identification across cameras with non-overlapping fields of view is one of the most important and difficult problems in video surveillance and analysis. However, current algorithms are likely to fail in real-world scenarios for several reasons. For example, surveillance cameras are typically mounted high above the ground plane, causing serious perspective changes. Also, most algorithms approach matching across images using the same descriptors, regardless of camera viewpoint or human pose. Here, we introduce a re-identification algorithm that addresses both problems. We build a model for human appearance as a function of pose, using training data gathered from a calibrated camera. We then apply this "pose prior" in online re-identification to make matching and identification more robust to viewpoint. We further integrate person-specific features learned over the course of tracking to improve the algorithm's performance. We evaluate the performance of the proposed algorithm and compare it to several state-of-the-art algorithms, demonstrating superior performance on standard benchmarking datasets as well as a challenging new airport surveillance scenario. Ziyan Wu 0001, Yang Li 0056, Richard J. Radke |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | A Sensor-Based Dual-Arm Tele-Robotic SystemabstractWe present a novel system to achieve coordinated task-based control on a dual-arm industrial robot for the general tasks of visual servoing and bimanual hybrid motion/force control. The industrial robot, consisting of a rotating torso and two seven degree-of-freedom arms, performs autonomous vision-based target alignment of both arms with the aid of fiducial markers, two-handed grasping and force control, and robust object manipulation in a tele-robotic framework. The operator uses hand motions to command the desired position for the object via Microsoft Kinect while the autonomous force controller maintains a stable grasp. Gestures detected by the Kinect are also used to dictate different operation modes. We demonstrate the effectiveness of our approach using a variety of common objects with different sizes, shapes, weights, and surface compliances. Daniel Kruse, John T. Wen, Richard J. Radke |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2014 | Improving counterflow detection in dense crowds with scene features
Ziyan Wu 0001, Richard J. Radke |
Pattern Recognit. Lett. | 2 |
| 2014 | Using Time-of-Flight Measurements for Privacy-Preserving Tracking in a Smart RoomabstractWe present a method for real-time person tracking and coarse pose recognition in a smart room using time-of-flight measurements. The time-of-flight images are severely downsampled to preserve the privacy of the occupants and simulate future applications that use single-pixel sensors in “smart” ceiling panels. The tracking algorithms use grayscale morphological image reconstruction to avoid false detections and are designed to not mistakenly detect pieces of furniture as people. A maximum-likelihood estimation method using a simple Markov model was implemented for robust pose classification. We show that the algorithms work effectively even when the sensors are spaced apart by 25 cm, using both real-world experiments and environmental simulation. Richard J. Radke |
IEEE Trans. Ind. Informatics | 2 |
| 2013 | A Reduced Order Memetic Algorithm for Constraint Optimization in Radiation Therapy Treatment PlanningabstractIn this paper, a novel hybrid genetic algorithm is presented for optimization in radiation therapy treatment planning. The proposed Reduced Order Memetic Algorithm (ROMA) is a combination of an evolutionary multi-objective optimization algorithm and gradient-based local search in a reduced order space. The gradient-based optimizer is used for a fast local search and is a variant of the sequential quadratic programming method. The execution time of the local search is improved by applying dynamically a principal component analysis to the solutions generated by the genetic optimizer and reducing the high-dimensionality search-space. In particular, for intensity modulated radiation therapy (IMRT) we observed reduction of the search-space dimensionality from several hundreds to less than twenty. Latin hypercube sampling was used to define the weights of the scalarization scheme for the local search fitness function for each individual solution. The proposed hybrid algorithm obtains efficiently a set of diverse non-dominated solutions for a large scale multi-objective problem such as in radiation treatment planning optimization. The applicability of the proposed algorithm is demonstrated for IMRT optimization for a case of prostate cancer. Georgios Kalantzis 0001, Aditya P. Apte, Richard J. Radke |
SNPD | 3 |
| 2013 | Keeping a Pan-Tilt-Zoom Camera CalibratedabstractPan-tilt-zoom (PTZ) cameras are pervasive in modern surveillance systems. However, we demonstrate that the (pan, tilt) coordinates reported by PTZ cameras become inaccurate after many hours of operation, endangering tracking and 3D localization algorithms that rely on the accuracy of such values. To solve this problem, we propose a complete model for a PTZ camera that explicitly reflects how focal length and lens distortion vary as a function of zoom scale. We show how the parameters of this model can be quickly and accurately estimated using a series of simple initialization steps followed by a nonlinear optimization. Our method requires only 10 images to achieve accurate calibration results. Next, we show how the calibration parameters can be maintained using a one-shot dynamic correction process; this ensures that the camera returns the same field of view every time the user requests a given (pan, tilt, zoom), even after hundreds of hours of operation. The dynamic calibration algorithm is based on matching the current image against a stored feature library created at the time the PTZ camera is mounted. We evaluate the calibration and dynamic correction algorithms on both experimental and real-world datasets, demonstrating the effectiveness of the techniques. Ziyan Wu 0001, Richard J. Radke |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Physical Scale Keypoints: Matching and Registration for Combined Intensity/Range Images
Eric R. Smith, Richard J. Radke, Charles V. Stewart |
Int. J. Comput. Vis. | 2 |
| 2012 | Image segmentation with one shape prior - A template-based formulation
Daniel Cremers, Richard J. Radke |
Image Vis. Comput. | 3 |
| 2011 | Segmenting the prostate and rectum in CT imagery using anatomical constraints
D. Michael Lovelock, Richard J. Radke |
Medical Image Anal. | 3 |
| 2011 | Automated kymograph analysis for profiling axonal transport of secretory granules
Amit Mukherjee, Brian Jenkins, Richard J. Radke, Gary Banker, Badrinath Roysam |
Medical Image Anal. | 4 |
| 2010 | Converting Level Set Gradients to Shape Gradients
Guillaume Charpiat, Richard J. Radke |
ECCV (5) | 3 |
| 2009 | Level set segmentation with both shape and intensity priorsabstractWe present a new variational level-set-based segmentation formulation that uses both shape and intensity prior information learned from a training set. By applying Bayes' rule to the segmentation problem, the cost function decomposes into shape and image energy parts. The shape energy is based on recently proposed nonparametric shape distributions, and we propose a new image energy model that incorporates learned intensity information from both foreground and background objects. The proposed variational level set segmentation framework has two main advantages. First, by characterizing image information with regional intensity distributions, there is no need to balance image energy and shape energy using a heuristic weighting factor. Second, by incorporating learned intensity information into the image model using a nonparametric density estimation method and an appropriate distance measure, our segmentation framework can handle problems where the interior/exterior of the shape has a highly inhomogeneous intensity distribution. We demonstrate our segmentation algorithm using challenging pelvis CT scans. Richard J. Radke |
ICCV | 2 |
| 2009 | Non-negative matrix factorization of partial track data for motion segmentationabstractThis paper addresses the problem of segmenting low-level partial feature point tracks belonging to multiple motions. We show that the local velocity vectors at each instant of the trajectory are an effective basis for motion segmentation. We decompose the velocity profiles of point tracks into different motion components and corresponding non-negative weights using non-negative matrix factorization (NNMF). We then segment the different motions using spectral clustering on the derived weights. We test our algorithm on the Hopkins 155 benchmarking database and several new sequences, demonstrating that the proposed algorithm can accurately segment multiple motions at a speed of a few seconds per frame. We show that our algorithm is particularly successful on low-level tracks from real-world video that are fragmented, noisy and inaccurate. Anil M. Cheriyadat, Richard J. Radke |
ICCV | 2 |
| 2008 | Registration of combined range-intensity scans: Initialization through verification
Eric R. Smith, Bradford J. King, Charles V. Stewart, Richard J. Radke |
Comput. Vis. Image Underst. | 4 |
| 2008 | Calibrating Distributed Camera NetworksabstractRecent developments in wireless sensor networks have made feasible distributed camera networks, in which cameras and processing nodes may be spread over a wide geographical area, with no centralized processor and limited ability to communicate a large amount of information over long distances. This paper overviews distributed algorithms for the calibration of such camera networks- that is, the automatic estimation of each camera's position, orientation, and focal length. In particular, we discuss a decentralized method for obtaining the vision graph for a distributed camera network, in which each edge of the graph represents two cameras that image a sufficiently large part of the same environment. We next describe a distributed algorithm in which each camera performs a local, robust nonlinear optimization over the camera parameters and scene points of its vision graph neighbors in order to obtain an initial calibration estimate. We then show how a distributed inference algorithm based on belief propagation can refine the initial estimate to be both accurate and globally consistent. Dhanya Devarajan, Zhaolin Cheng, Richard J. Radke |
Proc. IEEE | 3 |
| 2007 | Reslicing axially sampled 3D shapes using elliptic Fourier descriptors
Yongwon Jeong, Richard J. Radke |
Medical Image Anal. | 2 |
| 2006 | Distributed metric calibration of ad hoc camera networksabstractWe discuss how to automatically obtain the metric calibration of an ad hoc network of cameras with no centralized processor. We model the set of uncalibrated cameras as nodes in a communication network, and propose a distributed algorithm in which each camera performs a local, robust bundle adjustment over the camera parameters and scene points of its neighbors in an overlay “vision graph.” We analyze the performance of the algorithm on both simulated and real data, and show that the distributed algorithm results in a fairer allocation of messages per node while achieving comparable calibration accuracy to centralized bundle adjustment. Dhanya Devarajan, Richard J. Radke, Haeyong Chung |
ACM Trans. Sens. Networks | 2 |
| 2005 | Image change detection algorithms: a systematic surveyabstractDetecting regions of change in multiple images of the same scene taken at different times is of widespread interest due to a large number of applications in diverse disciplines, including remote sensing, surveillance, medical diagnosis and treatment, civil infrastructure, and underwater sensing. This paper presents a systematic survey of the common processing steps and core decision rules in modern change detection algorithms, including significance and hypothesis testing, predictive models, the shading model, and background modeling. We also discuss important preprocessing methods, approaches to enforcing the consistency of the change mask, and principles for evaluating and comparing the performance of change detection algorithms. It is hoped that our classification of algorithms into a relatively small number of categories will provide useful guidance to the algorithm designer. Richard J. Radke, Srinivas Andra, Omar Al-Kofahi, Badrinath Roysam |
IEEE Trans. Image Process. | 1 |
| 2005 | Model-based segmentation of medical imagery by matching distributionsabstractThe segmentation of deformable objects from three-dimensional (3-D) images is an important and challenging problem, especially in the context of medical imagery. We present a new segmentation algorithm based on matching probability distributions of photometric variables that incorporates learned shape and appearance models for the objects of interest. The main innovation over similar approaches is that there is no need to compute a pixelwise correspondence between the model and the image. This allows for a fast, principled algorithm. We present promising results on difficult imagery for 3-D computed tomography images of the male pelvis for the purpose of image-guided radiotherapy of the prostate. Daniel Freedman, Richard J. Radke, Yongwon Jeong, D. Michael Lovelock, George T. Y. Chen |
IEEE Trans. Medical Imaging | 2 |
| 2004 | Geometric-model-based segmentation of the prostate and surrounding structures for image-guided radiotherapyabstractWe present a computer vision tool to improve the clinical outcome of patients undergoing radiation therapy for prostate cancer by improving irradiation technique. While intensity modulated radiotherapy (IMRT) allows one to irradiate a specific region in the body with high accuracy, it is still difficult to know exactly where to aim the radiation beam on every day of the 30~40 treatments that are necessary. This paper presents a geometric model-based technique to accurately segment the prostate and other surrounding structures in a daily serial CT image, compensating for daily motion and shape variation. We first acquire a collection of serial CT scans of patients undergoing external beam radiotherapy, and manual segmentation of the prostate and other nearby structures by radiation oncologists. Then we train shape and local appearance models for the structures of interest. When new images are available, an iterative algorithm is applied to locate the prostate and surrounding structures automatically. Our experimental results show that excellent matches can be given to the prostate and surrounding structure. Convergence is declared after 10 iterations. For 256 x 256 images, the mean distance between the hand-segmented contour and the automatically estimated contour is about 1.5 pixels (2.44 mm), with variance about 0.6 pixel (1.24 mm). Yongwon Jeong, Richard J. Radke, George T. Y. Chen |
VCIP | 3 |
| 2003 | Efficiently synthesizing virtual videoabstractGiven a set of synchronized video sequences of a dynamic scene taken by different cameras, we address the problem of creating a virtual video of the scene from a novel viewpoint. A key aspect of our algorithm is a method for recursively propagating dense and physically accurate correspondences between the two video sources. By exploiting temporal continuity and suitably constraining the correspondences, we provide an efficient framework for synthesizing realistic virtual video. The stability of the propagation algorithm is analyzed, and experimental results are presented. Richard J. Radke, Peter J. Ramadge, Sanjeev R. Kulkarni, Tomio Echigo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | Using view interpolation for low bit rate videoabstractWe demonstrate that in some situations, perceptual quality can be maintained using an approach based on synthesizing "virtual" images of a scene that match frames from a source video clip. We use this algorithm for interpolation of video frames in the time domain, using a small amount of information to construct an approximation of the original video. Our algorithm is well-suited for the limitations in bandwidth and complexity characteristic of wireless multimedia channels. Since the approach is based on estimating functions of the underlying camera motion parameters, it can capture relationships between image correspondences that extend across many (perhaps hundreds) of video frames. Each interpolated image can be rendered using only a few tens of bytes of side information, and the rendering process itself has low computational requirements. We present experimental results to demonstrate that for certain types of video, our algorithm can give significant perceptual improvement over MPEG-4 coded video at the same low bit rate. Richard J. Radke, Peter J. Ramadge, Sanjeev R. Kulkarni, Tomio Echigo |
ICIP (1) | 1 |
| 2000 | Efficiently Estimating Projective TransformationsabstractThe estimation of the parameters of a projective transformation that relates the coordinates of two image planes is a standard problem that arises in image and video mosaicking, virtual video, and problems in computer vision. This problem is often posed as a least squares minimization problem based on a finite set of noisy point samples of the underlying transformation. While in some special cases this problem can be solved using a linear approximation, in general, it results in an 8-dimensional nonquadratic minimization problem that is solved numerically using an 'off-the-shelf' procedure such as the Levenberg-Marquardt algorithm. We show that the general least squares problem for estimating a projective transformation can be analytically reduced to a 2-dimensional nonquadratic minimization problem. Moreover, we provide both analytical and experimental evidence that the minimization of this function is computationally attractive. We propose a particular algorithm that is a combination of a projection and an approximate Gauss-Newton scheme, and experimentally verify that this algorithm efficiently solves the least squares problem. Richard J. Radke, Peter J. Ramadge, Tomio Echigo, Shun-ichi Iisaku |
ICIP | 1 |
| 2000 | Recursive Propagation of Correspondences with Applications to the Creation of Virtual VideoabstractThis paper is concerned with the efficient temporal propagation of correspondences between frames of two video sequences, an integral component of many video processing tasks. The main contribution is a framework for the recursive propagation of these correspondences. The propagation consists of a time update step and a measurement update step. The time update depends only on the dynamics of the rotating source cameras, while the measurement update can be tailored to any member of a general class of image correspondence algorithms. Using these results, the correspondence between points of each frame pair can be propagated and updated in a fraction of the time required to estimate correspondences anew at every frame. We discuss an application of the recursive correspondence propagation framework to the creation of virtual video. Previous virtual view algorithms have been used to generate synthetic video of a static scene, in which objects seem frozen in time. In contrast, the algorithms described here allow the creation of "true" virtual video, in the sense that the synthetic video evolves dynamically along with the scene. While virtual video is our motivating application, the recursive correspondence propagation framework applies to any two-camera video application in which correspondence is difficult and prohibitively time-consuming to estimate by processing frame pairs independently. Richard J. Radke, Peter J. Ramadge, Sanjeev R. Kulkarni, Tomio Echigo, Shun-ichi Iisaku |
ICIP | 1 |
| 1999 | Ghost Error Elimination and Superimposition of Moving Objects in Video MosaicingabstractThe paper presents an approach for region based video mosaicing, treating moving objects separately from the background, and with improved ghost-like noise elimination. The mosaic images show the moving objects superimposed over a stationary background. Conventional technologies can reduce the ghost-like noise that occurs from moving objects by using temporal median filtering, but its efficiency depends on the ratio between the speeds of the camera and the moving object. Our technology eliminates these noises more efficiently by using segmented images of a spatio-temporal video sequence. Segmentation is performed using a novel technique that uses different configurations of quad-trees for the initial separation in the split-and-merge process. The segmented images are also used to display tracked moving objects on the panoramic image. Tomio Echigo, Richard J. Radke, Peter J. Ramadge, Hisashi Miyamori, Shun-ichi Iisaku |
ICIP (4) | 2 |