VLDB 2026 Research / reviewers in the wild / expert
Quanfu Fan
dblp:66/3950
· DBLP profile ↗
49ranked-venue papers
15as first author
11since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 12 first-author · 6 since 2021Artificial intelligence and machine learning · 27 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Grafting Vision TransformersabstractVision Transformers (ViTs) have recently become the state-of-the-art across many computer vision tasks. In contrast to convolutional networks (CNNs), ViTs enable global information sharing even within shallow layers of a network, i.e., among high-resolution features. However, this perk was later overlooked with the success of pyramid architectures such as Swin Transformer, which show better performance-complexity trade-offs. In this paper, we present a simple and efficient add-on component (termed GrafT) that considers global dependencies and multi-scale information throughout the network, in both high- and low-resolution features alike. It has the flexibility of branching out at arbitrary depths and shares most of the parameters and computations of the backbone. GrafT shows consistent gains over various well-known models which includes both hybrid and pure Transformer types, both homogeneous and pyramid structures, and various self-attention methods. In particular, it largely benefits mobile-size models by providing high-level semantics. On the ImageNet-1k dataset, GrafT delivers +3.9%, +1.4%, and +1.9% top-1 accuracy improvement to DeiT-T, Swin-T, and MobViTXXS, respectively. Our code and models are at https://github.com/jongwoopark7978/Grafting-Vision-Transformer. Jongwoo Park 0003, Kumara Kahatapitiya, Donghyun Kim 0006, Shivchander Sudalairaj, Quanfu Fan, Michael S. Ryoo |
WACV | 5 |
| 2024 | CryoRL: Reinforcement Learning Enables Efficient Cryo-EM Data CollectionabstractSingle-particle cryo-electron microscopy (cryo-EM) has become one of the mainstream structural biology techniques because of its ability to determine high-resolution structures of dynamic bio-molecules. However, cryo-EM data acquisition remains expensive and labor-intensive, requiring substantial expertise. Structural biologists need a more efficient and objective method to collect the best data in a limited time frame. We formulate the cryo-EM data collection task as an optimization problem in this work. The goal is to maximize the total number of good images taken within a specified period. We show that reinforcement learning offers an effective way to plan cryo-EM data collection, successfully navigating heterogenous cryo-EM grids. The approach we developed, cryoRL, demonstrates better performance than average users for data collection under similar settings. Quanfu Fan, Yilai Li, Yuguang Yao, John Cohn, Sijia Liu 0001, Ziping Xu, Seychelle M. Vos, Michael A. Cianfrocco |
WACV | 1 |
| 2023 | Improve Video Representation with Temporal Adversarial AugmentationabstractRecent works reveal that adversarial augmentation benefits the generalization of neural networks (NNs) if used in an appropriate manner. In this paper, we introduce Temporal Adversarial Augmentation (TA), a novel video augmentation technique that utilizes temporal attention. Unlike conventional adversarial augmentation, TA is specifically designed to shift the attention distributions of neural networks with respect to video clips by maximizing a temporal-related loss function. We demonstrate that TA will obtain diverse temporal views, which significantly affect the focus of neural networks. Training with these examples remedies the flaw of unbalanced temporal information perception and enhances the ability to defend against temporal shifts, ultimately leading to better generalization. To leverage TA, we propose Temporal Video Adversarial Fine-tuning (TAF) framework for improving video representations. TAF is a model-agnostic, generic, and interpretability-friendly training strategy. We evaluate TAF with four powerful models (TSM, GST, TAM, and TPN) over three challenging temporal-related benchmarks (Something-something V1&V2 and diving48). Experimental results demonstrate that TAF effectively improves the test accuracy of these models with notable margins without introducing additional parameters or computational costs. As a byproduct, TAF also improves the robustness under out-of-distribution (OOD) settings. Code is available at https://github.com/jinhaoduan/TAF. Jinhao Duan, Quanfu Fan, Hao Cheng 0015, Xiaoshuang Shi, Kaidi Xu |
IJCAI | 2 |
| 2022 | RegionViT: Regional-to-Local Attention for Vision Transformers
Chun-Fu Chen 0001, Rameswar Panda, Quanfu Fan |
ICLR | 3 |
| 2022 | Can an Image Classifier Suffice For Action Recognition?
Quanfu Fan, Chun-Fu Chen 0001, Rameswar Panda |
ICLR | 1 |
| 2022 | Distributed adversarial training to robustify deep neural networks at scaleabstractCurrent deep neural networks (DNNs) are vulnerable to adversarial attacks, where adversarial perturbations to the inputs can change or manipulate classification. To defend against such attacks, an effective and popular approach, known as adversarial training (AT), has been shown to mitigate the negative impact of adversarial attacks by virtue of a min-max robust training method. While effective, it remains unclear whether it can successfully be adapted to the distributed learning context. The power of distributed optimization over multiple machines enables us to scale up robust training over large models and datasets. Spurred by that, we propose distributed adversarial training (DAT), a large-batch adversarial training framework implemented over multiple machines. We show that DAT is general, which supports training over labeled and unlabeled data, multiple types of attack generation methods, and gradient compression operations favored for distributed optimization. Theoretically, we provide, under standard conditions in the optimization theory, the convergence rate of DAT to the first-order stationary points in general non-convex settings. Empirically, we demonstrate that DAT either matches or outperforms state-of-the-art robust accuracies and achieves a graceful training speedup (e.g., on ResNet-50 under ImageNet). Codes are available at https://github.com/dat-2022/dat. Gaoyuan Zhang, Songtao Lu, Xiangyi Chen, Quanfu Fan, Lee Martie, Lior Horesh, Mingyi Hong 0001, Sijia Liu 0001 |
UAI | 6 |
| 2022 | Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video UnderstandingabstractVideos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a single label per video. Consequently, models can be incorrectly penalized for classifying actions that exist in the videos but are not explicitly labeled and do not learn the full spectrum of information present in each video in training. Towards this goal, we present the Multi-Moments in Time dataset (M-MiT) which includes over two million action labels for over one million three second videos. This multi-label dataset introduces novel challenges on how to train and analyze models for multi-action detection. Here, we present baseline results for multi-action recognition using loss functions adapted for long tail multi-label learning, provide improved methods for visualizing and interpreting models trained for multi-label action detection and show the strength of transferring models trained on M-MiT to smaller datasets. Mathew Monfort, Bowen Pan, Kandan Ramakrishnan, Alex Andonian, Barry A. McNamara, Alex Lascelles, Quanfu Fan, Dan Gutfreund, Rogério Feris, Aude Oliva |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | Deep Analysis of CNN-Based Spatio-Temporal Representations for Action RecognitionabstractIn recent years, a number of approaches based on 2D or 3D convolutional neural networks (CNN) have emerged for video action recognition, achieving state-of-the-art results on several large-scale benchmark datasets. In this paper, we carry out in-depth comparative analysis to better understand the differences between these approaches and the progress made by them. To this end, we develop an unified framework for both 2D-CNN and 3D-CNN action models, which enables us to remove bells and whistles and provides a common ground for fair comparison. We then conduct an effort towards a large-scale analysis involving over 300 action recognition models. Our comprehensive analysis reveals that a) a significant leap is made in efficiency for action recognition, but not in accuracy; b) 2D-CNN and 3D-CNN models behave similarly in terms of spatio-temporal representation abilities and transferability. Our codes are available at https://github.com/IBM/action-recognition-pytorch. Chun-Fu Chen 0001, Rameswar Panda, Kandan Ramakrishnan, Rogério Feris, John Cohn, Aude Oliva, Quanfu Fan |
CVPR | 7 |
| 2021 | CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationabstractThe recently developed vision transformer (ViT) has achieved promising results on image classification compared to convolutional neural networks. Inspired by this, in this paper, we study how to learn multi-scale feature representations in transformer models for image classification. To this end, we propose a dual-branch transformer to com-bine image patches (i.e., tokens in a transformer) of different sizes to produce stronger image features. Our approach processes small-patch and large-patch tokens with two separate branches of different computational complexity and these tokens are then fused purely by attention multiple times to complement each other. Furthermore, to reduce computation, we develop a simple yet effective token fusion module based on cross attention, which uses a single token for each branch as a query to exchange information with other branches. Our proposed cross-attention only requires linear time for both computational and memory complexity instead of quadratic time otherwise. Extensive experiments demonstrate that our approach performs better than or on par with several concurrent works on vision transformer, in addition to efficient CNN models. For example, on the ImageNet1K dataset, with some architectural changes, our approach outperforms the recent DeiT by a large margin of 2% with a small to moderate increase in FLOPs and model parameters. Our source codes and models are available at https://github.com/IBM/CrossViT. Chun-Fu Chen 0001, Quanfu Fan, Rameswar Panda |
ICCV | 2 |
| 2021 | AdaMML: Adaptive Multi-Modal Learning for Efficient Video RecognitionabstractMulti-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition. While traditional multi-modal learning offers excellent recognition results, its computational expense limits its impact for many real-world applications. In this paper, we propose an adaptive multi-modal learning framework, called AdaMML, that selects on-the-fly the optimal modalities for each segment conditioned on the input for efficient video recognition. Specifically, given a video segment, a multi-modal policy net-work is used to decide what modalities should be used for processing by the recognition model, with the goal of improving both accuracy and efficiency. We efficiently train the policy network jointly with the recognition model using standard back-propagation. Extensive experiments on four challenging diverse datasets demonstrate that our proposed adaptive approach yields 35% − 55% reduction in computation when compared to the traditional baseline that simply uses all the modalities irrespective of the in-put, while also achieving consistent improvements in accuracy over the state-of-the-art methods. Project page: https://rpand002.github.io/adamml.html. Rameswar Panda, Chun-Fu Chen 0001, Quanfu Fan, Ximeng Sun, Kate Saenko, Aude Oliva, Rogério Feris |
ICCV | 3 |
| 2021 | Generating Adversarial Computer Programs using Optimized Obfuscations
Shashank Srikant, Sijia Liu 0001, Tamara Mitrovska, Shiyu Chang, Quanfu Fan, Gaoyuan Zhang, Una-May O'Reilly |
ICLR | 5 |
| 2020 | Adversarial T-Shirt! Evading Person Detectors in a Physical World
Kaidi Xu, Gaoyuan Zhang, Sijia Liu 0001, Quanfu Fan, Mengshu Sun, Hongge Chen, Yanzhi Wang 0001, Xue Lin 0001 |
ECCV (5) | 4 |
| 2020 | Moments in Time Dataset: One Million Videos for Event UnderstandingabstractWe present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions occurring in 3 second videos poses many challenges: meaningful events do not include only people, but also objects, animals, and natural phenomena; visual and auditory events can be symmetrical in time ("opening" is "closing" in reverse), and either transient or sustained. We describe the annotation process of our dataset (each video is tagged with one action or activity label among 339 different classes), analyze its scale and diversity in comparison to other large-scale video datasets for action recognition, and report results of several baseline models addressing separately, and jointly, three modalities: spatial, temporal and auditory. The Moments in Time dataset, designed to have a large coverage and diversity of events in both visual and auditory modalities, can serve as a new challenge to develop models that scale to the level of complexity and abstract reasoning that a human processes on a daily basis. Mathew Monfort, Carl Vondrick, Aude Oliva, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa M. Brown, Quanfu Fan, Dan Gutfreund |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2019 | Reasoning About Human-Object Interactions Through Dual Attention NetworksabstractObjects are entities we act upon, where the functionality of an object is determined by how we interact with it. In this work we propose a Dual Attention Network model which reasons about human-object interactions. The dual-attentional framework weights the important features for objects and actions respectively. As a result, the recognition of objects and actions mutually benefit each other. The proposed model shows competitive classification performance on the human-object interaction dataset Something-Something. Besides, it can perform weak spatiotemporal localization and affordance segmentation, despite being trained only with video-level labels. The model not only finds when an action is happening and which object is being manipulated, but also identifies which part of the object is being interacted with. Tete Xiao, Quanfu Fan, Dan Gutfreund, Mathew Monfort, Aude Oliva, Bolei Zhou |
ICCV | 2 |
| 2019 | Big-Little Net: An Efficient Multi-Scale Feature Representation for Visual and Speech Recognition
Chun-Fu Chen 0001, Quanfu Fan, Neil Mallinar, Tom Sercu, Rogério Feris |
ICLR (Poster) | 2 |
| 2019 | Structured Adversarial Attack: Towards General Implementation and Better Interpretability
Kaidi Xu, Sijia Liu 0001, Pu Zhao 0001, Huan Zhang 0001, Quanfu Fan, Deniz Erdogmus, Yanzhi Wang 0001, Xue Lin 0001 |
ICLR (Poster) | 6 |
| 2019 | More Is Less: Learning Efficient Video Representations by Big-Little Network and Depthwise Temporal AggregationabstractCurrent state-of-the-art models for video action recognition are mostly based on expensive 3D ConvNets. This results in a need for large GPU clusters to train and evaluate such architectures. To address this problem, we present an lightweight and memory-friendly architecture for action recognition that performs on par with or better than current architectures by using only a fraction of resources. The proposed architecture is based on a combination of a deep subnet operating on low-resolution frames with a compact subnet operating on high-resolution frames, allowing for high efficiency and accuracy at the same time. We demonstrate that our approach achieves a reduction by 3~4 times in FLOPs and ~2 times in memory usage compared to the baseline. This enables training deeper models with more input frames under the same computational budget. To further obviate the need for large-scale 3D convolutions, a temporal aggregation module is proposed to model temporal dependencies in a video at very small additional computational costs. Our models achieve strong performance on several action recognition benchmarks including Kinetics, Something-Something and Moments-in-time. The code and models are available at \url{https://github.com/IBM/bLVNet-TAM}. Quanfu Fan, Chun-Fu Chen 0001, Hilde Kuehne, Marco Pistoia, David D. Cox |
NeurIPS | 1 |
| 2018 | SC-Conv: Sparse-Complementary Convolution for Efficient Model Utilization on CNNsabstractWe propose sparse-complementary convolution (SC-Conv) to improve model utilization of convolution neural networks (CNNs). The networks with SC-Conv achieve better accuracy than the regular convolution under similar computations and parameters. The proposed SC-Conv is paired with two deterministic sparse kernels, and one of kernels is complementary to the other one at in either spatial or channel domain or both; the deterministic sparsity increases the computational speed theoretically and practically; furthermore, by having the complementary characteristic, SC-Conv retains the same receptive field to the conventional convolution. This insightful but straightforward SC-Conv reuses of modern network architectures (ResNet and DenseNet), and at the same FLOPs and parameters, SC-Conv improves top-1 classification accuracy on ImageNet by 0.6 points for ResNet-101 and keep the same model complexity. Furthermore, SC-Conv also outperforms recent sparse networks by 1.3 points at top-1 accuracy for ImageNet, and after integrating SC-Conv with the sparse network, we further improve another 1.8 points accuracy at similar FLOPs and parameters. Chun-Fu Chen 0001, Jinwook Oh, Quanfu Fan, Marco Pistoia |
ISM | 3 |
| 2018 | Semantically Guided Visual Question AnsweringabstractWe present a novel approach to enhance the challenging task of Visual Question Answering (VQA) by incorporating and enriching semantic knowledge in a VQA model. We first apply Multiple Instance Learning (MIL) to extract a richer visual representation addressing concepts beyond objects such as actions and colors. Motivated by the observation that semantically related answers often appear together in prediction, we further develop a new semantically-guided loss function for model learning which has the potential to drive weakly-scored but correct answers to the top while suppressing wrong answers. We show that these two ideas contribute to performance improvement in a complementary way. We demonstrate competitive results comparable to the state of the art on two VQA benchmark datasets. Handong Zhao, Quanfu Fan, Dan Gutfreund, Yun Fu 0001 |
WACV | 2 |
| 2017 | Sparse Deep Feature Representation for Object Detection from Wearable Cameras
Quanfu Fan |
BMVC | 1 |
| 2016 | A Unified Multi-scale Deep Convolutional Neural Network for Fast Object Detection
Zhaowei Cai, Quanfu Fan, Rogério Feris, Nuno Vasconcelos |
ECCV (4) | 2 |
| 2016 | Enhanced face detection using body part detections for wearable camerasabstractWith the recent broad acceptance of body worn cameras (BWC) for police departments, there is an increased need to perform video analytics for this domain. However, body worn cameras pose several challenges including severe motion blur, barrel camera distortion from wide angle lenses, close proximity and odd viewing angles, and poor lighting conditions. In this paper, we evaluate the performance of several state of the art face detection approaches including Aggregate Channel Features [1] and Faster R-CNN [2] and show their limitations in this domain. We then describe how face detection can be improved for BWC by corroborating information from body parts detection. We design a system using 0–1 linear integer programming to optimize the matching of body parts detections for each person in the scene and to maximize the hit rate of faces with supporting evidence. By leveraging information from body parts detection, we are able to improve the average precision by nearly two percent. Lisa M. Brown, Quanfu Fan |
ICPR | 2 |
| 2016 | A closer look at Faster R-CNN for vehicle detectionabstractFaster R-CNN achieves state-of-the-art performance on generic object detection. However, a simple application of this method to a large vehicle dataset performs unimpressively. In this paper, we take a closer look at this approach as it applies to vehicle detection. We conduct a wide range of experiments and provide a comprehensive analysis of the underlying structure of this model. We show that through suitable parameter tuning and algorithmic modification, we can significantly improve the performance of Faster R-CNN on vehicle detection and achieve competitive results on the KITTI vehicle dataset. We believe our studies are instructive for other researchers investigating the application of Faster R-CNN to their problems and datasets. Quanfu Fan, Lisa M. Brown, John R. Smith |
Intelligent Vehicles Symposium | 1 |
| 2016 | People detection in crowded scenes by context-driven label propagationabstractExploiting contextual cues has been a key idea to improve people detection in crowded scenes. Along this line we present a novel context-driven approach to detect people in crowded scenes. Based on a context graph that incorporates both geometric and social contextual patterns in crowds, we apply label propagation to discover weak detections contextually compatible with true detections while suppressing irrelevant false alarms. Compared to previous approaches for context modeling limited to only pairwise spatial interactions between local object neighbors, our approach provides a more effective way to model people interactions in a global context. Our approach achieves performance comparable to state of the art on two challenging datasets for people and pedestrian detection. Jingjing Liu 0001, Quanfu Fan, Sharath Pankanti, Dimitris N. Metaxas |
WACV | 2 |
| 2015 | Self-calibration from vehicle informationabstractIn this paper, we describe a method to automatically calibrate a static camera using only information regarding the vehicles in the scene. We use standard Adaboost vehicle detection combined with feature-based tracking to determine the size and direction of travel of vehicles at various locations in the scene. We then infer the scene geometry including the height of the camera, the camera focal length, and the angle of the camera with respect to the ground. Given the calibration, we can estimate the size and speed of objects and the distances travelled in the scene. We evaluate our system on 10 cameras with manually annotated calibration based on a 3D cube with user specified locations in 2D. We show our system can be used to accurately estimate vehicle sizes to within 2 feet for a range of camera configurations. Lisa M. Brown, Quanfu Fan, Yun Zhai |
AVSS | 2 |
| 2014 | Long-term object tracking for parked vehicle detectionabstractWe develop a robust approach to detect parked vehicles in real time. Our approach particularly focuses on tracking vehicles in long term under challenging conditions such as lighting changes and occlusions. Vehicle tracking is performed by template matching based on fast-computed corner points. The template model is made self-adaptive over time to accommodate lighting changes. We also present an effective way to manage and track multiple vehicles when they are parked close together and occlude one another. We demonstrate the effectiveness of our approach on the challenging i-LIDs data set and another large one collected from real-world scenarios. Quanfu Fan, Sharath Pankanti, Lisa M. Brown |
AVSS | 1 |
| 2014 | Temporal Sequence Modeling for Video Event DetectionabstractWe present a novel approach for event detection in video by temporal sequence modeling. Exploiting temporal information has lain at the core of many approaches for video analysis (i.e., action, activity and event recognition). Unlike previous works doing temporal modeling at semantic event level, we propose to model temporal dependencies in the data at sub-event level without using event annotations. This frees our model from ground truth and addresses several limitations in previous work on temporal modeling. Based on this idea, we represent a video by a sequence of visual words learnt from the video, and apply the Sequence Memoizer [21] to capture long-range dependencies in a temporal context in the visual sequence. This data-driven temporal model is further integrated with event classification for jointly performing segmentation and classification of events in a video. We demonstrate the efficacy of our approach on two challenging datasets for visual recognition. Yu Cheng 0001, Quanfu Fan, Sharath Pankanti, Alok N. Choudhary |
CVPR | 2 |
| 2014 | Random Laplace Feature Maps for Semigroup Kernels on HistogramsabstractWith the goal of accelerating the training and testing complexity of nonlinear kernel methods, several recent papers have proposed explicit embeddings of the input data into low-dimensional feature spaces, where fast linear methods can instead be used to generate approximate solutions. Analogous to random Fourier feature maps to approximate shift-invariant kernels, such as the Gaussian kernel, on Rd, we develop a new randomized technique called random Laplace features, to approximate a family of kernel functions adapted to the semigroup structure of R+d. This is the natural algebraic structure on the set of histograms and other non-negative data representations. We provide theoretical results on the uniform convergence of random Laplace features. Empirical analyses on image classification and surveillance event detection tasks demonstrate the attractiveness of using random Laplace features relative to several other feature maps proposed in the literature. Jiyan Yang, Vikas Sindhwani, Quanfu Fan, Haim Avron, Michael W. Mahoney |
CVPR | 3 |
| 2013 | Relative Attributes for Large-Scale Abandoned Object DetectionabstractEffective reduction of false alarms in large-scale video surveillance is rather challenging, especially for applications where abnormal events of interest rarely occur, such as abandoned object detection. We develop an approach to prioritize alerts by ranking them, and demonstrate its great effectiveness in reducing false positives while keeping good detection accuracy. Our approach benefits from a novel representation of abandoned object alerts by relative attributes, namely static ness, foreground ness and abandonment. The relative strengths of these attributes are quantified using a ranking function[19] learnt on suitably designed low-level spatial and temporal features. These attributes of varying strengths are not only powerful in distinguishing abandoned objects from false alarms such as people and light artifacts, but also computationally efficient for large-scale deployment. With these features, we apply a linear ranking algorithm to sort alerts according to their relevance to the end-user. We test the effectiveness of our approach on both public data sets and large ones collected from the real world. Quanfu Fan, Prasad Gabbur, Sharath Pankanti |
ICCV | 1 |
| 2013 | Spatio-temporal fisher vector coding for surveillance event detectionabstractWe present a generic event detection system evaluated in the Surveillance Event Detection (SED) task of TRECVID 2012. We investigate a statistical approach with spatio-temporal features applied to seven event classes, which were defined by the SED task. This approach is based on local spatio-temporal descriptors, called MoSIFT and generated by pair-wise video frames. A Gaussian Mixture Model(GMM) is learned to model the distribution of the low level features. Then for each sliding window, the Fisher vector encoding [improvedFV] is used to generate the sample representation. The model is learnt using a Linear SVM for each event. The main novelty of our system is the introduction of Fisher vector encoding into video event detection. Fisher vector encoding has demonstrated great success in image classification. The key idea is to model the low level visual features as a Gaussian Mixture Model and to generate an intermediate vector representation for bag of features. FV encoding uses higher order statistics in place of histograms in the standard BoW. FV has several good properties: (a) it can naturally separate the video specific information from the noisy local features and (b) we can use a linear model for this representation. We build an efficient implementation for FV encoding which can attain a 10 times speed-up over real-time. We also take advantage of non-trivial object localization techniques to feed into the video event detection, e.g. multi-scale detection and non-maximum suppression. This approach outperformed the results of all other teams submissions in TRECVID SED 2012 on four of the seven event types. Qiang Chen 0007, Yang Cai 0002, Lisa M. Brown, Ankur Datta, Quanfu Fan, Rogério Feris, Shuicheng Yan, Alex Hauptmann 0001, Sharath Pankanti |
ACM Multimedia | 5 |
| 2012 | Robust Foreground and Abandonment Analysis for Large-Scale Abandoned Object Detection in Complex Surveillance VideosabstractWe present a robust system for large-scale abandoned object detection (AOD) with low false positive rates and good detection accuracy under complex realistic scenarios. The robustness of our system is largely attributed to an approach we develop for foreground analysis, which can effectively differentiate foreground objects from background under challenging conditions such as lighting changes, low textureness and low contrast as well as cluttered background. This significantly eliminates false positives caused by lighting changes while retaining true drops better. We further perform abandonment analysis to reduce more false positives including those related to people, at a small cost of accuracy (≤ 2%). We demonstrate the effectiveness of our approach on two large data sets collected in various challenging scenes, providing detailed analysis of experiments. Quanfu Fan, Sharath Pankanti |
AVSS | 1 |
| 2012 | Hand tracking by binary quadratic programming and its application to retail activity recognitionabstractSubstantial ambiguities arise in hand tracking due to issues such as small hand size, deformable hand shapes and similar hand appearances. These issues have greatly limited the capability of current multi-target tracking techniques in hand tracking. As an example, state-of-the-art approaches for people tracking handle indentity switching by exploiting the appearance cues using advanced object detectors. For hand tracking, such approaches will fail due to similar, or even identical hand appearances. The main contribution of our work is a global optimization framework based on binary quadratic programming (BQP) that seamlessly integrates appearance, motion and complex interactions between hands. Our approach effectively handles key challenges such as occlusion, detection failure, identity switching, and robustly tracks both hands in two challenging real-life scenarios: retail surveillance and sign languages. In addition, we demonstrate that an automatic method based on hand trajectory analysis outperforms state-of-the-art on checkout-related activity recognition in grocery stores. Hoang Trinh, Quanfu Fan, Prasad Gabbur, Sharath Pankanti |
CVPR | 2 |
| 2012 | Multimodal ranking for non-compliance detection in retail surveillanceabstractIn retail stores, cashier non-compliance activities at the Point of Sale (POS) are one of the prevalent sources of retail loss. In this paper, we propose a novel approach to reliably rank the list of detected non-compliance activities of a given retail surveillance system, thereby provide a means of significantly reducing the false alarms and improving the precision in non-compliance detection. Our approach represents each detected non-compliance activity using multi-modal features coming from video data, transaction logs (TLog) data and intermediate results of the video analytics. We then learn a binary classifier that successfully separate true positives and false positives in a labeled training set. A confidence score for each detection can then be computed using the decision value of the trained classifier, and a ranked list of detections can be formed based on this score. The benefit from having this ranked list is two-fold. First, a large number of false alarms can be avoided by simply keeping the top part of the list and discarding the rest. Second, a trade off between precision and recall can easily be performed by sliding the discarding threshold along this ranked list. Experimental results on a large scale dataset captured from real stores demonstrate that our approach achieves better precision than a state-of-the-art system at the same recall. Our approach can also reach an operating point that exceeds the retailers' expectation in terms of precision, while retaining an acceptable recall of more than 60%. Hoang Trinh, Sharath Pankanti, Quanfu Fan |
WACV | 3 |
| 2011 | Modeling of temporarily static objects for robust abandoned object detection in urban surveillanceabstractWe propose a robust approach for abandoned object detection in urban surveillance with over thousands of cameras. For such a large-scale monitoring based on intelligent video analysis, it is critical that a system be designed with careful control of false alarms. Our approach is based on proactive modeling of temporally static objects (TSO) such as cars stopping at red light and still pedestrians in the street. We develop a finite state machine to track the entire life cycles of TSOs from creation to termination. The semantically meaningful object information provided by the state machine in turn allows adaptive region-level updating of the background model without using any sophisticated object classification techniques. We demonstrate that our approach significantly mitigates the problematic issue of false alarm related to people in city surveillance, using both a small publicly available data set and a large one collected from various realistic urban scenarios. Quanfu Fan, Sharath Pankanti |
AVSS | 1 |
| 2011 | Detecting human activities in retail surveillance using hierarchical finite state machineabstractCashiers in retail stores usually exhibit certain repetitive and periodic activities when processing items. Detecting such activities plays a key role in most retail fraud detection systems. In this paper, we propose a highly efficient, effective and robust vision technique to detect checkout-related primitive activities, based on a hierarchical finite state machine (FSM). Our deterministic approach uses visual features and prior spatial constraints on the hand motion to capture particular motion patterns performed in primitive activities. We also apply our approach to the problem of retail fraud detection. Experimental results on a large set of video data captured from retail stores show that our approach, while much simpler and faster, achieves significantly better results than state-of-the-art machine learning-based techniques both in detecting checkout-related activities and in detecting checkout-related fraudulent incidents. Hoang Trinh, Quanfu Fan, Jiyan Pan, Prasad Gabbur, Sachiko Miyazawa, Sharath Pankanti |
ICASSP | 2 |
| 2011 | Robust abandoned object detection using region-level analysisabstractWe propose a robust abandoned object detection algorithm for real-time video surveillance. Different from conventional approaches that mostly rely on pixel-level processing, we perform region-level analysis in both background maintenance and static foreground object detection. In background maintenance, region-level information is fed back to adaptively control the learning rate. In static foreground object detection, region-level analysis double-checks the validity of candidate abandoned blobs. Attributed to such analysis, our algorithm is robust against illumination change, "ghosts" left by removed objects, distractions from partially static objects, and occlusions. Experiments on nearly 130,000 frames of i-LIDS dataset show the superior performance of our approach. Jiyan Pan, Quanfu Fan, Sharath Pankanti |
ICIP | 2 |
| 2011 | A pattern discovery approach to retail fraud detectionabstractA major source of revenue shrink in retail stores is the intentional or unintentional failure of proper checking out of items by the cashier. More recently, a few automated surveillance systems have been developed to monitor cashier lanes and detect non-compliant activities such as fake item checkouts or scans done with the intention of deriving monetary benefit. These systems use data from surveillance video cameras and transaction logs (TLog) recorded at the Point-of-Sale (POS). In this paper, we present a pattern discovery based approach to detect fraudulent events at the POS. Our approach is based on mining time-ordered text streams, representing retail transactions, formed from a combination of visually detected checkout related activities called primitives and barcodes from TLog data. Patterns representing single item checkouts, i.e. anchored around a single barcode, are discovered from these text streams using an efficient pattern discovery technique called Teiresias. The discovered patterns are used to build models for true and fake item scans by retaining or discarding the anchoring barcodes in those patterns respectively. A pattern matching and classification scheme is designed to robustly detect non-compliant cashier activities in the presence of noise in either the TLog or the video data. Different weighting schemes for quantifying the relative importance of the discovered patterns are explored: Frequency, Support Vector Machine (SVM) and Frequency+SVM. Using a large scale dataset recorded from retail stores, our approach discovers semantically meaningful cashier scan patterns. Our experiments also suggest that different weighting schemes result in varied false and true positive performances on the task of fake scan detection. Prasad Gabbur, Sharath Pankanti, Quanfu Fan, Hoang Trinh |
KDD | 3 |
| 2011 | Soft margin keyframe comparison: Enhancing precision of fraud detection in retail surveillanceabstractWe propose a novel approach for enhancing precision in a leading video analytics system that detects cashier fraud in grocery stores for loss prevention. While intelligent video analytics has recently become a promising means of loss prevention for retailers, most of the real-world systems suffer from a large number of false alarms, resulting in a significant waste of human labor during manual verification. Our proposed approach starts with the candidate fraudulent events detected by a state-of-the-art system. Such fraudulent events are a set of visually recognized checkout-related activities of the cashier without barcode associations. Instead of conducting costly video analysis, we extract a few keyframes to represent the essence of each candidate fraudulent event, and compare those keyframes to identify whether or not the event is a valid check-out process that involves consistent appearance changes on the lead-in belt, the scan area and the take-away belt. Our approach also performs a margin-based soft classification so that the user could trade off between saving human labor and preserving high recall. Experiments on days of surveillance videos collected from real grocery stores show that our algorithm can save about 50% of human labor while preserving over 90% of true alarms with small computational overhead. Jiyan Pan, Quanfu Fan, Sharath Pankanti, Hoang Trinh, Prasad Gabbur, Sachiko Miyazawa |
WACV | 2 |
| 2011 | Robust Spatiotemporal Matching of Electronic Slides to Presentation VideosabstractWe describe a robust and efficient method for automatically matching and time-aligning electronic slides to videos of corresponding presentations. Matching electronic slides to videos provides new methods for indexing, searching, and browsing videos in distance-learning applications. However, robust automatic matching is challenging due to varied frame composition, slide distortion, camera movement, low-quality video capture, and arbitrary slides sequence. Our fully automatic approach combines image-based matching of slide to video frames with a temporal model for slide changes and camera events. To address these challenges, we begin by extracting scale-invariant feature-transformation (SIFT) keypoints from both slides and video frames, and matching them subject to a consistent projective transformation (homography) by using random sample consensus (RANSAC). We use the initial set of matches to construct a background model and a binary classifier for separating video frames showing slides from those without. We then introduce a new matching scheme for exploiting less distinctive SIFT keypoints that enables us to tackle more difficult images. Finally, we improve upon the matching based on visual information by using estimated matching probabilities as part of a hidden Markov model (HMM) that integrates temporal information and detected camera operations. Detailed quantitative experiments characterize each part of our approach and demonstrate an average accuracy of over 95% in 13 presentation videos. Quanfu Fan, Kobus Barnard, Arnon Amir, Alon Efrat |
IEEE Trans. Image Process. | 1 |
| 2010 | Graph based event detection from realistic videos using weak feature correspondenceabstractWe study the problem of event detection from realistic videos with repetitive sequential human activities. Despite the large body of work on event detection and recognition, very few have addressed low-quality videos captured from realistic environments. Our framework is based on solving the shortest path on a temporal-event graph constructed from the video content. Graph vertices correspond to detected event primitives, and edge weights are set according to generic knowledge of the event patterns and the discrepancy between event primitives based on a greedy matching of their visual features. Experimental results on videos collected from a retail environment validate the usefulness of the proposed approach. Lei Ding 0002, Quanfu Fan, Jen-Hao Hsiao, Sharath Pankanti |
ICASSP | 2 |
| 2010 | An integer programming approach to visual complianceabstractVisual compliance has emerged as a new paradigm to ensure that employees comply with processes and policies in a business context. In this paper, we focus on videos from retail stores and formulate a binary integer program for detecting checkout events, which enables enforcing visual compliance in such an environment. The proposed integer program maximizes the essential quantities that characterize true events of interest, subject to an array of constraints. In particular, the binary decision variables correspond to the presence of a set of hypothesized visual events. In the objective function, the binary variables are weighted by quality measures derived from infinite Gaussian mixture modeling of the video content, such that maximizing the overall quality measure is expected to uncover the meaningful visual events. Our framework is tested and validated on videos recorded at checkout lanes, and leads to better performance than previous methods. Lei Ding 0002, Quanfu Fan, Sharath Pankanti |
ICIP | 2 |
| 2009 | Recognition of repetitive sequential human activityabstractWe present a novel framework for recognizing repetitive sequential events performed by human actors with strong temporal dependencies and potential parallel overlap. Our solution incorporates sub-event (or primitive) detectors and a spatiotemporal model for sequential event changes. We develop an effective and efficient method to integrate primitives into a set of sequential events where strong temporal constraints are imposed on the ordering of the primitives. In particular, the combination process is approached as an optimization problem. A specialized Viterbi algorithm is designed to learn and infer the target sequential events and handle the event overlap simultaneously. To demonstrate the effectiveness of the proposed framework, we report detailed quantitative analysis on a large set of cashier checkout activities in a retail store. Quanfu Fan, Russell Bobbitt, Yun Zhai, Akira Yanagawa, Sharath Pankanti, Arun Hampapur |
CVPR | 1 |
| 2009 | Detecting sweethearting in retail surveillance videosabstractA significant portion of retail shrink is attributed to employees and occurs around the point of sale (POS). In this paper, we target a major type of retail fraud in surveillance videos, known as sweethearting (or fake scan), where a cashier intentionally fails to enter one or more items into the transaction in an attempt to get free merchandise for the customer. We first develop a motion-based algorithm to identify video segments as candidates for primitive events at the POS. We then apply spatio-temporal features to recognize true primitive events from the candidates and prune those falsely alarmed. In particular, we learn location-aware event models by Multiple-Instance Learning to address the location-sensitive issues that appear in our problem. Finally, we validate the entire transaction by combining primitive events according to temporal ordering constraints. We demonstrate the effectiveness of our approach on data captured from a real grocery store. Quanfu Fan, Akira Yanagawa, Russell Bobbitt, Yun Zhai, Rick Kjeldsen, Sharath Pankanti, Arun Hampapur |
ICASSP | 1 |
| 2009 | Accurate alignment of presentation slides with educational videoabstractSpatio-temporal alignment of electronic slides with corresponding presentation video opens up a number of possibilities for making the instructional content more accessible and understandable, such as video quality improvement, better content analysis and novel compression approaches for low bandwidth access. However, these applications need finding accurate transformations between slides and video frames, which is quite challenging in capture settings using pan-tilt-zoom (PTZ) cameras. In this paper we present a nonlinear optimization approach for accurate registration of slide images to video frames. Instead of estimating the projective transformation (i.e., homography) between a single pair of slide and frame images, we solve a set of homographies jointly in a frame sequence that is associated with a given slide. Quantitative evaluation confirms that this substantively improves alignment accuracy. Quanfu Fan, Kobus Barnard, Arnon Amir, Alon Efrat |
ICME | 1 |
| 2009 | Fast detection of retail fraud using polar touch buttonsabstractVideo analytics have recently emerged as a promising technique of retail fraud detection for loss prevention. Efficient video analytic algorithms are highly desired for a practical fraud detection system. In this paper, we present a real-time algorithm for recognizing a cashier's actions at the point of sale (POS), which can be further used to analyze cashier behaviors for identifying fraudulent incidents. The algorithm uses a set of simple but effective features derived from a global representation of motion energy called polar motion map (PMM). These features capture the motion patterns exhibited in a cashier's actions as a focused beam of motion energy, characterizing the actions as the extension and retraction movement of the cashier's arm with respect to a prespecified region. Our algorithm demonstrates comparable accuracy against one of the state-of-the-art event recognition techniques while running significantly faster. Quanfu Fan, Akira Yanagawa, Russell Bobbitt, Yun Zhai, Rick Kjeldsen, Sharath Pankanti, Arun Hampapur |
ICME | 1 |
| 2009 | Multi-media compliance: A practical paradigm for managing business integrityabstractIn virtually every business context there is a need to establish some form of monitoring system to ensure that employees comply with business processes and policies. Compliance failures range from organized theft to gaps in procedure that can be easily remedied through retraining. It is clearly important for businesses to capture and record these deviations to minimize loss prevention and maximize workplace safety and efficiency. In this workshop, we discuss the growing problem of compliance failure and how our system addresses this problem in a retail context to detect checkout-related fraud through the integration of visual and non-visual data. Sharath Pankanti, Quanfu Fan, Yun Zhai, Russell Bobbitt, Akira Yanagawa, Sachiko Miyazawa, Frederik Kjeldsen, Arun Hampapur |
ICME | 2 |
| 2008 | Evaluation of Localized Semantics: Data, Methodology, and Experiments
Kobus Barnard, Quanfu Fan, Ranjini Swaminathan, Anthony Hoogs, Roderic Collins, Pascale Rondot, John P. Kaufhold |
Int. J. Comput. Vis. | 2 |
| 2007 | Reducing correspondence ambiguity in loosely labeled training dataabstractWe develop an approach to reduce correspondence ambiguity in training data where data items are associated with sets of plausible labels. Our domain is images annotated with keywords where it is not known which part of the image a keyword refers to. In contrast to earlier approaches that build predictive models or classifiers despite the ambiguity, we argue that that it is better to first address the correspondence ambiguity, and then build more complex models from the improved training data. This addresses difficulties of fitting complex models in the face of ambiguity while exploiting all the constraints available from the training data. We contribute a simple and flexible formulation of the problem, and show results validated by a recently developed comprehensive evaluation data set and corresponding evaluation methodology. Kobus Barnard, Quanfu Fan |
CVPR | 2 |
| 2007 | Temporal Modeling of Slide Change in Presentation VideosabstractWe develop a general framework to automatically match electronic slides to the videos of the corresponding presentations. The synchronized slides support indexing and browsing of educational and corporate digital video libraries. Our approach extends previous work that matches slides based on visual features alone, and integrates multiple cues to further improve performance in more difficult cases. We model slide change in a presentation with a dynamic hidden Markov model (HMM) that captures the temporal notion of slide change and whose transition probabilities are adapted locally by using the camera events in the inference process. Our results show that combining multiple cues in a state model can greatly improve the performance in ambiguous cases. Quanfu Fan, Arnon Amir, Kobus Barnard, Ranjini Swaminathan, Alon Efrat |
ICASSP (1) | 1 |