Hassan Foroosh

dblp:55/1119 · also Hassan Shekarforoush · DBLP profile ↗
← Back
133ranked-venue papers
14as first author
17since 2021 · last 2024
0000-0002-7601-8165ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 103 · 13 first-author · 7 since 2021Artificial intelligence and machine learning · 66 · 3 first-author · 13 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs
abstract
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Hassan Foroosh, Dong Yu 0001, Fei Liu 0004
ACL (1)5
2024 When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives
abstract
Reasoning is most powerful when an LLM accurately aggregates relevant information.We examine the critical role of information aggregation in reasoning by requiring the LLM to analyze sports narratives.To succeed at this task, an LLM must infer points from actions, identify related entities, attribute points accurately to players and teams, and compile key statistics to draw conclusions.We conduct comprehensive experiments with real NBA basketball data and present SPORTSGEN, a new method to synthesize game narratives.By synthesizing data, we can rigorously evaluate LLMs' reasoning capabilities under complex scenarios with varying narrative lengths and density of information.Our findings show that most models, including GPT-4o, often fail to accurately aggregate basketball scores due to frequent scoring patterns.Open-source models like Llama-3 further suffer from significant score hallucinations.Finally, the effectiveness of reasoning is influenced by narrative complexity, information density, and domain-specific terms, highlighting the challenges in analytical reasoning tasks. 1 * Work done during Yebowen Hu's internship; Kaiqiang Song and Sangwoo Cho were full-time researchers at Tencent AI Lab, Seattle, USA at the time of this work. https://github.com/YebowenHu/SportsGenAnalyze the team-player affiliations and play-by-play descriptions below to determine the total points scored by each team (player).Please explain your reasoning step by step and provide the final results in JSON format.Start with: {New York Knicks: 0, Denver Nuggets: 0} ({Andrea Bargnani: 0, Timofey Mozgov: 0, ...
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Wenlin Yao, Hassan Foroosh, Dong Yu 0001, Fei Liu 0004
EMNLP6
2024 LPFormer: LiDAR Pose Estimation Transformer with Multi-Task Network
abstract
Due to the difficulty of acquiring large-scale 3D human keypoint annotation, previous methods for 3D human pose estimation (HPE) have often relied on 2D image features and sequential 2D annotations. Furthermore, the training of these networks typically assumes the prediction of a human bounding box and the accurate alignment of 3D point clouds with 2D images, making direct application in real-world scenarios challenging. In this paper, we present the 1stframework for end-to-end 3D human pose estimation, named LPFormer, which uses only LiDAR as its input along with its corresponding 3D annotations. LPFormer consists of two stages: firstly, it identifies the human bounding box and extracts multi-level feature representations, and secondly, it utilizes a transformer-based network to predict human keypoints based on these features. Our method demonstrates that 3D HPE can be seamlessly integrated into a strong LiDAR perception network and benefit from the features extracted by the network. Experimental results on the Waymo Open Dataset demonstrate the state-of-the-art performance, and improvements even compared to previous multi-modal solutions.
Dongqiangzi Ye, Yufei Xie, Weijia Chen, Lingting Ge, Hassan Foroosh
ICRA6
2024 LiDARFormer: A Unified Transformer-based Multi-task Network for LiDAR Perception
abstract
There is a recent need in the LiDAR perception field for unifying multiple tasks in a single strong network with improved performance, as opposed to using separate networks for each task. In this paper, we introduce a new LiDAR multi-task learning paradigm based on the transformer. The proposed LiDARFormer utilizes cross-space global contextual feature information and exploits cross-task synergy to boost the performance of LiDAR perception tasks across multiple large-scale datasets and benchmarks. Our novel transformer-based framework includes a cross-space transformer module that learns attentive features between the 2D dense Bird’s Eye View (BEV) and 3D sparse voxel feature maps. Additionally, we propose a transformer decoder for the segmentation task to dynamically adjust the learned features by leveraging the categorical feature representations. Furthermore, we combine the segmentation and detection features in a shared transformer decoder with cross-task attention layers to enhance and integrate the object-level and class-level features. LiDARFormer is evaluated on the large-scale nuScenes and the Waymo Open datasets for both 3D detection and semantic segmentation tasks, and it achieves state-of-the-art performance on both tasks.
Dongqiangzi Ye, Weijia Chen, Yufei Xie, Panqu Wang, Hassan Foroosh
ICRA7
2024 Multiple Adverse Weather Conditions Adaptation for Object Detection via Causal Intervention
abstract
Most state-of-the-art object detection methods have achieved impressive perfomrace on several public benchmarks, which are trained with high definition images. However, existing detectors are often sensitive to the visual variations and out-of-distribution data due to the domain gap caused by various confounders, e.g. the adverse weathre conditions. To bridge the gap, previous methods have been mainly exploring domain alignment, which requires to collect an amount of domain-specific training samples. In this paper, we introduce a novel domain adaptation model to discover a weather condition invariant feature representation. Specifically, we first employ a memory network to develop a confounder dictionary, which stores prototypes of object features under various scenarios. To guarantee the representativeness of each prototype in the dictionary, a dynamic item extraction strategy is used to update the memory dictionary. After that, we introduce a causal intervention reasoning module to explore the invariant representation of a specific object under different weather conditions. Finally, a categorical consistency regularization is used to constrain the similarities between categories in order to automatically search for the aligned instances among distinct domains. Experiments are conducted on several public benchmarks (RTTS, Foggy-Cityscapes, RID, and BDD 100K) with state-of-the-art performance achieved under multiple weather conditions.
Hua Zhang 0008, Xiaohong Li 0001, Xiaochun Cao, Hassan Foroosh
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Robust registration and learning using multi-radii spherical polar Fourier transform
Alam Abbas Syed, Hassan Foroosh
Signal Process.2
2023 Channel Regeneration: Improving Channel Utilization for Compact DNNs
Ankit Kumar Sharma, Hassan Foroosh
AAAI2
2023 LidarMultiNet: Towards a Unified Multi-Task Network for LiDAR Perception
abstract
LiDAR-based 3D object detection, semantic segmentation, and panoptic segmentation are usually implemented in specialized networks with distinctive architectures that are difficult to adapt to each other. This paper presents LidarMultiNet, a LiDAR-based multi-task network that unifies these three major LiDAR perception tasks. Among its many benefits, a multi-task network can reduce the overall cost by sharing weights and computation among multiple tasks. However, it typically underperforms compared to independently combined single-task models. The proposed LidarMultiNet aims to bridge the performance gap between the multi-task network and multiple single-task networks. At the core of LidarMultiNet is a strong 3D voxel-based encoder-decoder architecture with a Global Context Pooling (GCP) module extracting global contextual features from a LiDAR frame. Task-specific heads are added on top of the network to perform the three LiDAR perception tasks. More tasks can be implemented simply by adding new task-specific heads while introducing little additional cost. A second stage is also proposed to refine the first-stage segmentation and generate accurate panoptic segmentation results. LidarMultiNet is extensively tested on both Waymo Open Dataset and nuScenes dataset, demonstrating for the first time that major LiDAR perception tasks can be unified in a single strong network that is trained end-to-end and achieves state-of-the-art performance. Notably, LidarMultiNet reaches the official 1 place in the Waymo Open Dataset 3D semantic segmentation challenge 2022 with the highest mIoU and the best accuracy for most of the 22 classes on the test set, using only LiDAR points as input. It also sets the new state-of-the-art for a single model on the Waymo 3D object detection benchmark and three nuScenes benchmarks.
Dongqiangzi Ye, Weijia Chen, Yufei Xie, Panqu Wang, Hassan Foroosh
AAAI7
2023 MeetingBank: A Benchmark Dataset for Meeting Summarization
abstract
Yebowen Hu, Timothy Ganter, Hanieh Deilamsalehy, Franck Dernoncourt, Hassan Foroosh, Fei Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yebowen Hu, Tim Ganter, Hanieh Deilamsalehy, Franck Dernoncourt, Hassan Foroosh, Fei Liu 0004
ACL (1)5
2023 DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4
abstract
Human preference judgments are pivotal in guiding large language models (LLMs) to produce outputs that align with human values.Human evaluations are also used in summarization tasks to compare outputs from various systems, complementing existing automatic metrics.Despite their significance, however, there has been limited research probing these pairwise or kwise comparisons.The collective impact and relative importance of factors such as output length, informativeness, fluency, and factual consistency are still not well understood.It is also unclear if there are other hidden factors influencing human judgments.In this paper, we conduct an in-depth examination of a collection of pairwise human judgments released by Ope-nAI.Utilizing the Bradley-Terry-Luce (BTL) model, we reveal the inherent preferences embedded in these human judgments.We find that the most favored factors vary across tasks and genres, whereas the least favored factors tend to be consistent, e.g., outputs are too brief, contain excessive off-focus content or hallucinated facts.Our findings have implications on the construction of balanced datasets in human preference evaluations, which is a crucial step in shaping the behaviors of future LLMs.
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Hassan Foroosh, Fei Liu 0004
EMNLP5
2023 MobileRec: A Large Scale Dataset for Mobile Apps Recommendation
abstract
Recommender systems have become ubiquitous in our digital lives, from recommending products on e-commerce websites to suggesting movies and music on streaming platforms. Existing recommendation datasets, such as Amazon Product Reviews and MovieLens, greatly facilitated the research and development of recommender systems in their respective domains. While the number of mobile users and applications (aka apps) has increased exponentially over the past decade, research in mobile app recommender systems has been significantly constrained, primarily due to the lack of high-quality benchmark datasets, as opposed to recommendations for products, movies, and news. To facilitate research for app recommendation systems, we introduce a large-scale dataset, called MobileRec. We constructed MobileRec from users' activity on the Google play store. MobileRec contains 19.3 million user interactions (i.e., user reviews on apps) with over 10K unique apps across 48 categories. MobileRec records the sequential activity of a total of 0.7 million distinct users. Each of these users has interacted with no fewer than five distinct apps, which stands in contrast to previous datasets on mobile apps that recorded only a single interaction per user. Furthermore, MobileRec presents users' ratings as well as sentiments on installed apps, and each app contains rich metadata such as app name, category, description, and overall rating, among others. We demonstrate that MobileRec can serve as an excellent testbed for app recommendation through a comparative study of several state-of-the-art recommendation approaches. The MobileRec dataset is available at https://huggingface.co/datasets/recmeapp/mobilerec.
Muhammad Hasan Maqbool, Umar Farooq 0002, Adib Mosharrof, A. B. Siddique 0001, Hassan Foroosh
SIGIR5
2022 Personalizing Task-oriented Dialog Systems via Zero-shot Generalizable Reward Function
abstract
Task-oriented dialog systems enable users to accomplish tasks using natural language. State-of-the-art systems respond to users in the same way regardless of their personalities, although personalizing dialogues can lead to higher levels of adoption and better user experiences. Building personalized dialog systems is an important, yet challenging endeavor, and only a handful of works took on the challenge. Most existing works rely on supervised learning approaches and require laborious and expensive labeled training data for each user profile. Additionally, collecting and labeling data for each user profile is virtually impossible. In this work, we propose a novel framework, P-ToD, to personalize task-oriented dialog systems capable of adapting to a wide range of user profiles in an unsupervised fashion using a zero-shot generalizable reward function. P-ToD uses a pre-trained GPT-2 as a backbone model and works in three phases. Phase one performs task-specific training. Phase two kicks off unsupervised personalization by leveraging the proximal policy optimization algorithm that performs policy gradients guided by the zero-shot generalizable reward function. Our novel reward function can quantify the quality of the generated responses even for unseen profiles. The optional final phase fine-tunes the personalized model using a few labeled training examples. We conduct extensive experimental analysis using the personalized bAbI dialogue benchmark for five tasks and up to 180 diverse user profiles. The experimental results demonstrate that P-ToD, even when it had access to zero labeled examples, outperforms state-of-the-art supervised personalization models and achieves competitive performance on BLEU and ROUGE metrics when compared to a strong fully-supervised GPT-2 baseline.
A. B. Siddique 0001, Muhammad Hasan Maqbool, Kshitija Taywade, Hassan Foroosh
CIKM4
2022 CenterFormer: Center-Based Transformer for 3D Object Detection
Xiangchen Zhao, Panqu Wang, Hassan Foroosh
ECCV (38)5
2022 Haar Wavelet-Based Attention Network for Image Dehazing
abstract
Weather conditions such as haze, mist, and fog can degrade the image clarity in outdoor scenes. Although various image dehazing techniques exist in the literature, it is very challenging to remove non-homogeneous and/or densely concentrated homogeneous fog. We propose an end-to-end dehazing network1that utilizes attention-based learning from Haar wavelet coefficients. The model uses multi-scale wavelet transformation to break down feature maps into low- and high-frequency coefficients. The channel attention layer learns haze features while the spatial attention layer focuses on the feature location from these coefficients to further refine the output. Quantitative and visual analyses demonstrate that the proposed framework outperforms recent state-of-the-art methods in removing haze from dense and non-homogeneous hazy images while maintaining color accuracy.
Sumit Laha, Hassan Foroosh
ICIP2
2022 RAPID: A Single Stage Pruning Framework
abstract
Conventional network pruning methods require multiple stages to identify and train a single compact pruned model. This approach has a high computational overhead and requires multiple iterations to train multiple pruned models increasing the total computational cost. In this work, we present a single-stage Random Pruning Online Distillation (RAPID) framework to identify and train multiple pruned models at once. We randomly prune the original network at scratch and then leverage distillation at the later training epochs to transfer information from the original network(teacher) to multiple pruned models(students) online. Extensive experiments demonstrate the effectiveness of the proposed RAPID framework on several datasets and architectures when com-pared with other pruning methods.
Ankit Kumar Sharma, Hassan Foroosh
ICIP2
2021 Panoptic-PolarNet: Proposal-Free LiDAR Point Cloud Panoptic Segmentation
abstract
Panoptic segmentation presents a new challenge in exploiting the merits of both detection and segmentation, with the aim of unifying instance segmentation and semantic segmentation in a single framework. However, an efficient solution for panoptic segmentation in the emerging domain of LiDAR point cloud is still an open research problem and is very much under-explored. In this paper, we present a fast and robust LiDAR point cloud panoptic segmentation framework, referred to as Panoptic-PolarNet. We learn both semantic segmentation and class-agnostic instance clustering in a single inference network using a polar Bird's Eye View (BEV) representation, enabling us to circumvent the issue of occlusion among instances in urban street scenes. To improve our network's learnability, we also propose an adapted instance augmentation technique and a novel adversarial point cloud pruning method. Our experiments show that Panoptic-PolarNet outperforms the baseline methods on SemanticKITTI and nuScenes datasets with an almost real-time inference speed. Panoptic-PolarNet achieved 54.1% PQ in the public SemanticKITTI panoptic segmentation leaderboard and leading performance for the validation set of nuScenes.
Yang Zhang 0035, Hassan Foroosh
CVPR3
2021 StreamHover: Livestream Transcript Summarization and Annotation
abstract
Sangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui, Nedim Lipka, Walter Chang, Hailin Jin, Jonathan Brandt, Hassan Foroosh, Fei Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Sangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui, Nedim Lipka, Walter Chang, Hailin Jin, Jonathan Brandt, Hassan Foroosh, Fei Liu 0004
EMNLP (1)9
2020 UCF-STAR: A Large Scale Still Image Dataset for Understanding Human Actions
abstract
Action recognition in still images poses a great challenge due to (i) fewer available training data, (ii) absence of temporal information. To address the first challenge, we introduce a dataset for STill image Action Recognition (STAR), containing over $1M$ images across 50 different human body-motion action categories. UCF-STAR is the largest dataset in the literature for action recognition in still images. The key characteristics of UCF-STAR include (1) focusing on human body-motion rather than relatively static human-object interaction categories, (2) collecting images from the wild to benefit from a varied set of action representations, (3) appending multiple human-annotated labels per image rather than just the action label, and (4) inclusion of rich, structured and multi-modal set of metadata for each image. This departs from existing datasets, which typically provide single annotation in a smaller number of images and categories, with no metadata. UCF-STAR exposes the intrinsic difficulty of action recognition through its realistic scene and action complexity. To benchmark and demonstrate the benefits of UCF-STAR as a large-scale dataset, and to show the role of “latent” motion information in recognizing human actions in still images, we present a novel approach relying on predicting temporal information, yielding higher accuracy on 5 widely-used datasets.
Marjaneh Safaei, Pooyan Balouchian, Hassan Foroosh
AAAI3
2020 PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation
abstract
The requirement of fine-grained perception by autonomous driving systems has resulted in recently increased research in the online semantic segmentation of single-scan LiDAR. Emerging datasets and technological advancements have enabled researchers to benchmark this problem and improve the applicable semantic segmentation algorithms. Still, online semantic segmentation of LiDAR scans in autonomous driving applications remains challenging due to three reasons: (1) the need for near-real-time latency with limited hardware, (2) points are distributed unevenly across space, and (3) an increasing number of more fine-grained semantic classes. The combination of the aforementioned challenges motivates us to propose a new LiDAR-specific, KNN-free segmentation algorithm - PolarNet. Instead of using common spherical or bird's-eye-view projection, our polar bird's-eye-view representation balances the points per grid and thus indirectly redistributes the network's attention over the long-tailed points distribution over the radial axis in polar coordination. We find that our encoding scheme greatly increases the mIoU in three drastically different real urban LiDAR single-scan segmentation datasets while retaining ultra low latency and near real-time throughput.
Yang Zhang 0035, Philip David, Xiangyu Yue 0001, Zerong Xi, Boqing Gong, Hassan Foroosh
CVPR7
2020 Better Highlighting: Creating Sub-Sentence Summary Highlights
abstract
Amongst the best means to summarize is highlighting. In this paper, we aim to generate summary highlights to be overlaid on the original documents to make it easier for readers to sift through a large amount of text. The method allows summaries to be understood in context to prevent a summarizer from distorting the original meaning, of which abstractive summarizers usually fall short. In particular, we present a new method to produce self-contained highlights that are understandable on their own to avoid confusion. Our method combines determinantal point processes and deep contextualized representations to identify an optimal set of sub-sentence segments that are both important and non-redundant to form summary highlights. To demonstrate the flexibility and modeling power of our method, we conduct extensive experiments on summarization datasets. Our analysis provides evidence that highlighting is a promising avenue of research towards future summarization.
Sangwoo Cho, Kaiqiang Song, Chen Li 0003, Dong Yu 0001, Hassan Foroosh, Fei Liu 0004
EMNLP (1)5
2020 Slim-CNN: A Light-Weight CNN for Face Attribute Prediction
abstract
We introduce a computationally-efficient CNN micro-architecture Slim Module to design a lightweight deep neural network, Slim-CNN, for face attribute prediction. Slim Modules are constructed by assembling depthwise separable convolutions with pointwise convolution to produce a computationally efficient module. The problem of facial attribute prediction is challenging because of the large variations in pose, background, illumination, and dataset imbalance. We stack multiple Slim Modules to devise a compact CNN, which still maintains very high accuracy. Additionally, Slim-CNN has a very low memory footprint, which makes it suitable for mobile and embedded applications. Experiments on the CelebA dataset show that Slim-CNN achieve an accuracy of 91.24% with 25x fewer parameters compared to MCNN-AUX and 100x fewer parameters when compared to DTML. This reduces the memory storage requirement of Slim-CNN by at least 87%. Furthermore, we compare Slim Modules with other well-known micro-architectures, such as Inception modules, residual blocks, Shuffle-Unit, and Inverted Residual units, and show it outperforms them in performance and in memory size, making it suitable for face-related tasks in embedded applications.
Ankit Kumar Sharma, Hassan Foroosh
FG2
2020 Improving Image Matching with Varied Illumination
abstract
We present a method to maximize feature matching performance across stereo image pairs by varying illumination. We perform matching between views per lighting condition, finding unique SIFT correspondences for each condition. These feature matches are then collected together into a single set, selecting those features which present the highest quality match. Instead of capturing each view under each illumination, we approximate lighting changes with a pretrained relighting convolutional neural network which only requires each view captured under a single specified lighting condition. We then collect the best of these feature matches over all lighting conditions offered by the relighting network. We further present an optimization to limit the number of lighting conditions evaluated to gain a specified number of matches. Our method is evaluated on a set of indoor scenes excluded from training the network with comparison to features extracted from pretrained VGG16. Our method offers an average 5.5× improvement in number of correct matches while retaining similar precision than by the original lit image pair per scene alone.
Sarah Braeger, Hassan Foroosh
ICPR2
2020 CCA: Exploring the Possibility of Contextual Camouflage Attack on Object Detection
abstract
Deep neural network based object detection has become the cornerstone of many real-world applications. Along with this success comes concerns about its vulnerability to malicious attacks. To gain more insight into this issue, we propose a contextual camouflage attack (CCA for short) algorithm to influence the performance of object detectors. In this paper, we use an evolutionary search strategy and adversarial machine learning in interactions with a photo-realistic simulated environment to find camouflage patterns that are effective over a huge variety of object locations, camera poses, and lighting conditions. The proposed camouflages are validated effective to most of the state-of-the-art object detectors.
Shengnan Hu, Yang Zhang 0035, Sumit Laha, Ankit Kumar Sharma, Hassan Foroosh
ICPR5
2020 Near-Infrared Depth-Independent Image Dehazing using Haar Wavelets
abstract
We propose a fusion algorithm for haze removal that combines color information from an RGB image and edge information extracted from its corresponding NIR image using Haar wavelets. The proposed algorithm is based on the key observation that NIR edge features are more prominent in the hazy regions of the image than the RGB edge features in those same regions. To combine the color and edge information, we introduce a haze-weight map which proportionately distributes the color and edge information during the fusion process. Because NIR images are, intrinsically, nearly haze-free, our work makes no assumptions like existing works that rely on a scattering model and essentially designing a depth-independent method. This helps in minimizing artifacts and gives a more realistic sense to the restored haze-free image. Extensive experiments show that the proposed algorithm is both qualitatively and quantitatively better on several key metrics when compared to existing state-of-the-art methods.
Sumit Laha, Ankit Kumar Sharma, Shengnan Hu, Hassan Foroosh
ICPR4
2020 Multi-Scale Keypoint Matching
abstract
We propose a new hierarchical method to match keypoints by exploiting information across multiple scales. Traditionally, for each keypoint a single scale is detected and the matching process is done in the specific scale. We replace this approach with matching across scale-space. The holistic information from higher scales are used for early rejection of candidates that are far away in the feature space. The more localized and finer details of lower scale are then used to decide between remaining possible points. The proposed multi-scale solution is more consistent with the multi-scale processing that is present in the human visual system and is therefore biologically plausible. We evaluate our method on several datasets and achieve state of the art accuracy, while significantly outperforming others in extraction time.
Sina Lotfian, Hassan Foroosh
ICPR2
2020 Self-Attention Network for Skeleton-based Human Action Recognition
abstract
Skeleton-based action recognition has recently attracted a lot of attention. Researchers are coming up with new approaches for extracting spatio-temporal relations and making considerable progress on large-scale skeleton-based datasets. Most of the architectures being proposed are based upon recurrent neural networks (RNNs), convolutional neural networks (CNNs) and graph-based CNNs. When it comes to skeleton-based action recognition, the importance of long term contextual information is central which is not captured by the current architectures. In order to come up with a better representation and capturing of long term spatio-temporal relationships, we propose three variants of Self-Attention Network (SAN), namely, SAN-V1, SAN-V2 and SAN-V3. Our SAN variants has the impressive capability of extracting high-level semantics by capturing long-range correlations. We have also integrated the Temporal Segment Network (TSN) with our SAN variants which resulted in improved overall performance. Different configurations of Self-Attention Network (SAN) variants and Temporal Segment Network (TSN) are explored with extensive experiments. Our chosen configuration outperforms state-of-the-art Top-1 and Top-5 by 4.4% and 7.9% respectively on Kinetics and shows consistently better performance than state-of-the-art methods on NTU RGB+D.
Sangwoo Cho, Muhammad Hasan Maqbool, Fei Liu 0004, Hassan Foroosh
WACV4
2020 Wide Hidden Expansion Layer for Deep Convolutional Neural Networks
abstract
Non-linearity is an essential factor contributing to the success of deep convolutional neural networks. Increasing the non-linearity in the network will enhance the network's learning capability, attributing to better performance. We present a novel Wide Hidden Expansion (WHE) layer that can significantly increase (by an order of magnitude) the number of activation functions in the network, with very little increase of computational complexity and memory consumption. It can be flexibly embedded with different network architectures to boost the performance of the original networks. The WHE layer is composed of a wide hidden layer, in which each channel only connects with two input channels and one output channel. Before connecting to the output channel, each intermediate channel in the WHE layer is followed by one activation function. In this manner, the number of activation functions can grow along with the number of channels in the hidden layer. We apply the WHE layer to ResNet, WideResNet, SENet, and MobileNet architectures and evaluate on ImageNet, CIFAR-100, and Tiny ImageNet dataset. On the ImageNet dataset, models with the WHE layer can achieve up to 2.01% higher Top-1 accuracy than baseline models, with less than 4% computation increase and less than 2% more parameters. On CIFAR-100 and Tiny ImageNet, when applying the WHE layer to ResNet models, it demonstrates consistent improvement in the accuracy of the networks. Applying the WHE layer to ResNet backbone of the CenterNet object detection model can also boost its performance on COCO and Pascal VOC datasets.
Baoyuan Liu, Hassan Foroosh
WACV3
2020 A Curriculum Domain Adaptation Approach to the Semantic Segmentation of Urban Scenes
abstract
During the last half decade, convolutional neural networks (CNNs) have triumphed over semantic segmentation, which is one of the core tasks in many applications such as autonomous driving and augmented reality. However, to train CNNs requires a considerable amount of data, which is difficult to collect and laborious to annotate. Recent advances in computer graphics make it possible to train CNNs on photo-realistic synthetic imagery with computer-generated annotations. Despite this, the domain mismatch between real images and the synthetic data hinders the models' performance. Hence, we propose a curriculum-style learning approach to minimizing the domain gap in urban scene semantic segmentation. The curriculum domain adaptation solves easy tasks first to infer necessary properties about the target domain; in particular, the first task is to learn global label distributions over images and local distributions over landmark superpixels. These are easy to estimate because images of urban scenes have strong idiosyncrasies (e.g., the size and spatial relations of buildings, streets, cars, etc.). We then train a segmentation network, while regularizing its predictions in the target domain to follow those inferred properties. In experiments, our method outperforms the baselines on two datasets and three backbone networks. We also report extensive ablation studies about our approach.
Yang Zhang 0035, Philip David, Hassan Foroosh, Boqing Gong
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Lattice-Constrained Stratified Sampling for Point Cloud Levels of Detail
abstract
In the last several years, airborne topographic light detection and ranging (LiDAR) has emerged as an important remote-sensing technology supporting a wide variety of applications. In that time, researchers have conducted studies to determine the optimal sampling densities required to support their respective applications. This natural progression of experimentation and analysis has resulted in several recommended sampling density stratifications for LiDAR point cloud products. Recognizing the need for point clouds of varying sample densities provides at least two motivations for creating level of details (LOD) for high-resolution point clouds. First, from the perspective of LiDAR data consumers, there is a desire to use the coarsest sampling that supports the application to reduce procurement costs, storage constraints, and processing times. Second, from the perspective of LiDAR data providers, there is a desire to collect data once at the highest supported fidelity to minimize recollection costs and redundancy in data holdings. In this article, we present an approach for generating point cloud LODs by constraining samples to regular discrete lattices to optimize the coverage of the volumes represented by each sample. We compare our approach to the two most common point cloud sampling methods: random sampling and rectangular lattice sampling. We discuss two approaches for representing the point cloud LODs. Finally, we propose an extension of our sampling approach for processing single-photon and Geiger-mode avalanche photodiode (GmAPD) LiDAR.
Kristian L. Damkjer, Hassan Foroosh
IEEE Trans. Geosci. Remote. Sens.2
2019 Improving the Similarity Measure of Determinantal Point Processes for Extractive Multi-Document Summarization
abstract
The most important obstacles facing multidocument summarization include excessive redundancy in source descriptions and the looming shortage of training data.These obstacles prevent encoder-decoder models from being used directly, but optimization-based methods such as determinantal point processes (DPPs) are known to handle them well.In this paper we seek to strengthen a DPP-based method for extractive multi-document summarization by presenting a novel similarity measure inspired by capsule networks.The approach measures redundancy between a pair of sentences based on surface form and semantic information.We show that our DPP system with improved similarity measure performs competitively, outperforming strong summarization baselines on benchmark datasets.Our findings are particularly meaningful for summarizing documents created by multiple authors containing redundant yet lexically diverse expressions. 1
Sangwoo Cho, Logan Lebanoff, Hassan Foroosh, Fei Liu 0004
ACL (1)3
2019 An Unsupervised Subspace Ranking Method for Continuous Emotions in Face Images
Pooyan Balouchian, Marjaneh Safaei, Xiaochun Cao, Hassan Foroosh
BMVC4
2019 ComDefend: An Efficient Image Compression Model to Defend Adversarial Examples
abstract
Deep neural networks (DNNs) have been demonstrated to be vulnerable to adversarial examples. Specifically, adding imperceptible perturbations to clean images can fool the well trained deep neural networks. In this paper, we propose an end-to-end image compression model to defend adversarial examples: ComDefend. The proposed model consists of a compression convolutional neural network (ComCNN) and a reconstruction convolutional neural network (ResCNN). The ComCNN is used to maintain the structure information of the original image and purify adversarial perturbations. And the ResCNN is used to reconstruct the original image with high quality. In other words, ComDefend can transform the adversarial image to its clean version, which is then fed to the trained classifier. Our method is a pre-processing module, and does not modify the classifier's structure during the whole process. Therefore it can be combined with other model-specific defense models to jointly improve the classifier's robustness. A series of experiments conducted on MNIST, CIFAR10 and ImageNet show that the proposed method outperforms the state-of-the-art defense methods, and is consistently effective to protect classifiers against adversarial attacks.
Xiaojun Jia, Xingxing Wei 0001, Xiaochun Cao, Hassan Foroosh
CVPR4
2019 CAMOU: Learning Physical Vehicle Camouflages to Adversarially Attack Detectors in the Wild
Yang Zhang 0035, Hassan Foroosh, Philip David, Boqing Gong
ICLR (Poster)2
2019 LUCFER: A Large-Scale Context-Sensitive Image Dataset for Deep Learning of Visual Emotions
abstract
Still image emotion recognition has been receiving increasing attention in recent years due to the tremendous amount of social media content available on the Web. Opinion mining, visual emotion analysis, search and retrieval are among the application areas, to name a few. While there exist works on the subject, offering methods to detect image sentiment; i.e. recognizing the polarity of the image, less efforts focus on emotion analysis; i.e. dealing with recognizing the exact emotion aroused when exposed to certain visual stimuli. Main gaps tackled in this work include (1) lack of large-scale image datasets for deep learning of visual emotions and (2) lack of context-sensitive single-modality approaches in emotion analysis in the still image domain. In this paper, we introduce LUCFER (Pronounced LU-CI-FER), a dataset containing over 3.6M images, with 3-dimensional labels; i.e. emotion, context and valence. LUCFER, the largest dataset of the kind currently available, is collected using a novel data collection pipeline, proposed and implemented in this work. Moreover, we train a context-sensitive deep classifier using a novel multinomial classification technique proposed here via adding a dimensionality reduction layer to the CNN. Relying on our categorical approach to emotion recognition, we claim and show empirically that injecting context to our unified training process helps (1) achieve a more balanced precision and recall, and (2) boost performance, yielding an overall classification accuracy of 73.12% compared to 58.3% achieved in the closest work in the literature.
Pooyan Balouchian, Marjaneh Safaei, Hassan Foroosh
WACV3
2019 Still Image Action Recognition by Predicting Spatial-Temporal Pixel Evolution
abstract
Both spatial and temporal patterns provide crucial information for recognizing human actions. However, lack of temporal information in still images is a major obstacle in single-image action recognition. In this paper, (i) We introduce a novel image representation domain, Ranked Saliency Map and Predicted Optical Flow or Rank_SM-POF for short. This domain captures both actor appearance and the future movement patterns of the actor. This is accomplished through capturing the temporal ordering of each pixel by training a linear ranking machine on the predicted tensor of spatial-temporal representation of images. (ii) We employ a transfer learning approach to propose a new spatial-temporal Convolutional Neural Network, named STCNN for the task of single image action classification, by fine-tuning a CNN which is pre-trained specifically for appearance based classification. (iii) Finally, extensive experiments on five benchmarks clearly demonstrate that appearance and motion are complementary sources of information and using both leads to significant performance improvement in single image action recognition, hence outperforming state-of-the-art methods.
Marjaneh Safaei, Hassan Foroosh
WACV2
2019 Sparse One-Grab Sampling with Probabilistic Guarantees
abstract
Sampling is an important and effective strategy in analyzing "big data," whereby a smaller subset of a dataset is used to estimate the characteristics of its entire population. The main goal in sampling is often to achieve a significant gain in the computational time. However, a major obstacle towards this goal is the assessment of the smallest sample size needed to ensure, with a high probability, a faithful representation of the entire dataset, especially when the data set is compiled of a large number of diverse structures (e.g., clusters). To address this problem, we propose a method referred to as the Sparse Withdrawal of Inliers in a First Trial (SWIFT) that determines the smallest sample size of a subset of a dataset sampled in one grab, with the guarantee that the subset provides a sufficient number of samples from each of the underlying structures necessary for the discovery and inference. The latter is established with high probability, and the lower bound of the smallest sample size depends on probabilistic guarantees. In addition, we derive an upper bound on the smallest sample size that allows for detection of the structures and show that the two bounds are very close to each other in a variety of scenarios. We show that the problem can be modeled using either a hypergeometric or a multinomial probability mass function (pmf), and derive accurate mathematical bounds to determine a tight approximation to the sample size, leading thus to a sparse sampling strategy. The key features of the proposed method are: (i) sparseness of the sampled subset for analyzing data, where the level of sparseness is independent of the population size; (ii) no prior knowledge of the distribution of data, or the number of underlying structures in the data; and (iii) robustness in the presence of overwhelming number of outliers. We evaluate the method thoroughly in terms of accuracy, its behavior against different parameters, and its effectiveness in reducing the computational cost in various applications of computer vision, such as subspace clustering and structure from motion.
Maryam Jaberi, Marianna Pensky, Hassan Foroosh
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Learning Structural Representations via Dynamic Object Landmarks Discovery for Sketch Recognition and Retrieval
abstract
State-of-the-art methods on sketch classification and retrieval are based on deep convolutional neural network to learn representations. Although deep neural networks have the ability to model images with hierarchical representations by convolution kernels, they can not automatically extract the structural representations of object categories in a human-perceptible way. Furthermore, sketch images usually have large scale visual variations caused by the styles of drawing or viewpoints, which make it difficult to develop generalized representations using the fixed computational mode of convolutional kernel. In this paper, our aim is to address the problem of fixed computational mode in feature extraction process without extra supervision. We propose a novel architecture to dynamically discover the object landmarks and learn the discriminative structural representations. Our model is composed of two components: a representative landmark discovering module that localizes the key points on the object, and a category-aware representation learning module that develops the category-specific features. Specifically, we develop a structure-aware offset layer to dynamically localize the representative landmarks, which is optimized based on the category labels without extra supervision. After that, a diversity branch is introduced to extract the global discriminative features for each category. Finally, we employ a multi-task loss function to develop an end-to-end trainable architecture. At testing time, we fuse all the predictions with different number of landmarks to achieve the final results. Through extensive experiments, we compare our model with several state-of-the-art methods on two challenging datasets TU-Berlin and Sketchy for sketch classification and retrieval, and the experimental results demonstrate the effectiveness of our proposed model.
Hua Zhang 0008, Peng She, Yong Liu 0018, Jianhou Gan, Xiaochun Cao, Hassan Foroosh
IEEE Trans. Image Process.6
2018 Spatio-Temporal Fusion Networks for Action Recognition
Sangwoo Cho, Hassan Foroosh
ACCV (1)2
2018 A Novel Multi-purpose Deep Architecture for Facial Attribute and Emotion Understanding
Ankit Kumar Sharma, Pooyan Balouchian, Hassan Foroosh
CIARP3
2018 Context-Sensitive Single-Modality Image Emotion Analysis: A Unified Architecture from Dataset Construction to CNN Classification
abstract
Still image emotion recognition is receiving increasing attention in recent years due to the tremendous amount of social media content on the Web. Opinion mining, visual emotion analysis, search and retrieval are among the application areas to name a few. Works are published on the subject, offering methods to detect image sentiments, while others focus on extracting the true social signals, such as happiness and anger, among others. However “context-sensitive” emotion recognition has been by and large discarded in the literature so far. Moreover, the problem in the single-modal domain; i.e. using only still images, remains less attended. In this work, we introduce the largest dataset of images collected from the wild, UCF ER, labeled with emotion and context. We train a context-sensitive classifier to classify images based on both emotion and context, hence introducing the first single-modal context-sensitive emotion recognition CNN model trained on our newly constructed dataset. Relying on our categorical approach to emotion recognition, we claim and show that including context as part of a unified training process helps boost performance, while reducing dependency on cross-modality approaches. Experimental results demonstrate considerable boost in performance compared to state-of-the-art.
Pooyan Balouchian, Hassan Foroosh
ICIP2
2018 Curvature Augmented Deep Learning for 3D Object Recognition
abstract
This paper presents a new method to incorporate shape information into convolutional neural network (CNN)s for 3D object recognition. Voxel CNNs have been very successful with the task of 3D object recognition. However, continuous shape information that is useful for recognition is often lost in their conversion to a voxel representation. We propose a single dimensional feature that can be applied to voxel CNNs. This paper presents a novel rotation-invariant feature based on mean curvature that improves shape recognition for voxel CNNs. We augment the recent voxel CNN Octnet architecture with our feature and demonstrate a 1 % overall accuracy increase on the ModelNet 10 dataset.
Sarah Braeger, Hassan Foroosh
ICIP2
2018 TICNN: A Hierarchical Deep Learning Framework for Still Image Action Recognition Using Temporal Image Prediction
abstract
Lack of temporal information in still images is a major obstacle in still image action recognition. Here, we propose TICNN, a novel deep learning framework, addressing this challenge. We first introduce the concept of a Temporal Image, a compact representation of hypothetical sequence of images. We then take advantage of this concept to architect TICNN, composed of two convolutional neural networks. The first CNN learns a novel model to predict temporal images using transfer learning. Our second CNN extracts temporal image features to classify human actions in still images. Unlike previous efforts, TICNN learns temporal information, the missing cue, in a still image rather than merely focusing on spatial features. To the best of our knowledge, this is the first attempt to predict temporal images, dynamic patterns of a still image, to alleviate the lack of motion information in still images. Extensive experiments on four benchmarks demonstrate the positive effect of temporal image prediction on the accuracy of still image action recognition, outperforming the state-of-the-art methods in the literature.
Marjaneh Safaei, Pooyan Balouchian, Hassan Foroosh
ICIP3
2018 A Zero-Shot Architecture for Action Recognition in Still Images
abstract
Motion is a missing information in an image, however, it is a valuable cue for action recognition. Therefore, not only actions depend on the spatial-salient pixels, but also the temporal patterns of those pixels are evidently crucial. In this paper, we propose a novel unsupervised zero-shot approach, employing both spatial and temporal patterns, to perform action recognition in still images through Tensor Decomposition. In the proposed model, (1) we devise a novel strategy to form tensors from individual images in a way that each tensor encodes useful spatial-temporal information regarding the action being performed in images. Tensor decomposition is then used to estimate the overall signature of the action, while action is encoded in the spatial-temporal descriptions of images. (2) We show that appearance and motion are complementary sources of information. Comprehensive experiments on four benchmarks: Stanford-40, Willow, WIDER and the newly introduced UCFSI -101 still images dataset clearly demonstrate that our method outperforms state-of-the-art approaches.
Marjaneh Safaei, Hassan Foroosh
ICIP2
2018 Probabilistic Sparse Subspace Clustering Using Delayed Association
abstract
Discovering and clustering subspaces in high-dimensional data is a fundamental problem of machine learning with a wide range of applications in data mining, computer vision, and pattern recognition. Earlier methods divided the problem into two separate stages of finding the similarity matrix and finding clusters. Similar to some recent works, we integrate these two steps using a joint optimization approach. We make the following contributions: (i) we estimate the reliability of the cluster assignment for each point before assigning a point to a subspace. We group the data points into two groups of “certain” and “uncertain”, with the assignment of latter group delayed until their subspace association certainty improves. (ii) We demonstrate that delayed association is better suited for clustering subspaces that have ambiguities, i.e. when subspaces intersect or data are contaminated with outliers/noise. (iii) We demonstrate experimentally that such delayed probabilistic association leads to a more accurate self-representation and final clusters. The proposed method has higher accuracy both for points that exclusively lie in one subspace, and those that are on the intersection of subspaces. (iv) We show that delayed association leads to huge reduction of computational cost, since it allows for incremental spectral clustering.
Maryam Jaberi, Marianna Pensky, Hassan Foroosh
ICPR3
2018 A Temporal Sequence Learning for Action Recognition and Prediction
abstract
In this work1, we present a method to represent a video with a sequence of words, and learn the temporal sequencing of such words as the key information for predicting and recognizing human actions. We leverage core concepts from the Natural Language Processing (NLP) literature used in sentence classification to solve the problems of action prediction and action recognition. Each frame is converted into a word that is represented as a vector using the Bag of Visual Words (BoW) encoding method. The words are then combined into a sentence to represent the video, as a sentence. The sequence of words in different actions are learned with a simple but effective Temporal Convolutional Neural Network (T-CNN) that captures the temporal sequencing of information in a video sentence. We demonstrate that a key characteristic of the proposed method is its low-latency, i.e. its ability to predict an action accurately with a partial sequence (sentence). Experiments on two datasets, UCF101 and HMDB51 show that the method on average reaches 95% of its accuracy within half the video frames. Results, also demonstrate that our method achieves compatible state-of-the-art performance in action recognition (i.e. at the completion of the sentence) in addition to action prediction.
Sangwoo Cho, Hassan Foroosh
WACV2
2018 Look-Up Table Unit Activation Function for Deep Convolutional Neural Networks
abstract
Activation functions provide deep neural networks the non-linearity that is necessary to learn complex distributions. It is still inconclusive what is the optimal shape for the activation function. In this work, we introduce a novel type of activation function of which the shape is learned with network training. The proposed Look-up Table Unit (LuTU) stores a set of anchor points in a look-up table like structure, and the activation function is generated from the anchor points by either linear interpolation or smoothing with a single period cosine mask function. LuTU is in theory able to approximate any univariate function. By observing the learned shapes of LuTU, we further propose a Mixture of Gaussian Unit (MoGU) that can learn similar non-linear shapes with much fewer parameters. Finally, we use a multiple activation function fusion framework that combines multiple types of functions to achieve better performance. The inference complexity of multiple activation function fusion is constant with linear interpolation approximation. Our experiments on a synthetic dataset, ImageNet, and CIFAR-10 demonstrate that the proposed method outperforms traditional ReLU family activation functions. On the ImageNet dataset, our method achieves 1.47% and 1.0% higher accuracy on ResNet-18 and ResNet-34 models, respectively. With the proposed activation function, we can design a network that has the same performance as ResNet-34 but 8 fewer convolutional layers.
Baoyuan Liu, Hassan Foroosh
WACV3
2018 Designing a symmetric classifier for image annotation using multi-layer sparse coding
Amara Tariq, Hassan Foroosh
Image Vis. Comput.2
2017 Improving RANSAC-Based Segmentation through CNN Encapsulation
abstract
In this work, we present a method for improving a random sample consensus (RANSAC) based image segmentation algorithm by encapsulating it within a convolutional neural network (CNN). The improvements are gained by gradient descent training on the set of pre-RANSAC filtering and thresholding operations using a novel RANSAC-based loss function, which is geared toward optimizing the strength of the correct model relative to the most convincing false model. Thus, it can be said that our loss function trains the network on metrics that directly dictate the success or failure of the final segmentation rather than metrics that are merely correlated to the success or failure. We demonstrate successful application of this method to a RANSAC method for identifying the pupil boundary in images from the CASIA-IrisV3 iris recognition data set, and we expect that this method could be successfully applied to any RANSAC-based segmentation algorithm.
Dustin Morley, Hassan Foroosh
CVPR2
2017 Motion compensation using critically sampled DWT subbands for low-bitrate video coding
abstract
In this paper, we propose a novel motion estimation/motion compensation (ME/MC) method for wavelet-based (i.e. in-band) motion compensated temporal filtering (MCTF), with application to low-bitrate video coding. Unlike the conventional in-band MCTF algorithms, which require redundancy to overcome the shift-variance problem of critically sampled (i.e. complete) discrete wavelet transforms (DWT), we perform ME/MC steps directly on DWT coefficients by avoiding the need of shift-invariance. We omit upsampling, inverse-DWT (IDWT), and calculation of redundant DWT coefficients, while achieving arbitrary subpixel accuracy without interpolation, and high video quality even at very low-bitrates, by deriving the exact relationships between DWT subbands of input image sequences. Experimental results demonstrate the accuracy of the proposed method, confirming that our model for ME/MC effectively improves video coding quality.
Vildan Atalay Aydin, Hassan Foroosh
ICIP2
2017 View-invariant object recognition using homography constraints
abstract
Change in viewpoint is one of the major factors for variation in object appearance across different images. Thus, view-invariant object recognition is a challenging and important image understanding task. In this paper, we propose a method that can match objects in images taken under different viewpoints. Unlike most methods in the literature, no restriction on camera orientations or internal camera parameters are imposed and no prior knowledge of 3D structure of the object is required. We prove that when two cameras take pictures of the same object from two different viewing angels, the relationship between every quadruple of points reduces to the special case of homography with two equal eigenvalues. Based on this property, we formulate the problem as an error function that indicates how likely two sets of 2D points are projections of the same set of 3D points under two different cameras. Comprehensive set of experiments were conducted to prove the robustness of the method to noise, and evaluate its performance on real-world applications, such as face and object recognition.
Sina Lotfian, Hassan Foroosh
ICIP2
2017 NELasso: Group-Sparse Modeling for Characterizing Relations Among Named Entities in News Articles
abstract
Named entities such as people, locations, and organizations play a vital role in characterizing online content. They often reflect information of interest and are frequently used in search queries. Although named entities can be detected reliably from textual content, extracting relations among them is more challenging, yet useful in various applications (e.g., news recommending systems). In this paper, we present a novel model and system for learning semantic relations among named entities from collections of news articles. We model each named entity occurrence with sparse structured logistic regression, and consider the words (predictors) to be grouped based on background semantics. This sparse group LASSO approach forces the weights of word groups that do not influence the prediction towards zero. The resulting sparse structure is utilized for defining the type and strength of relations. Our unsupervised system yields a named entities' network where each relation is typed, quantified, and characterized in context. These relations are the key to understanding news material over time and customizing newsfeeds for readers. Extensive evaluation of our system on articles from TIME magazine and BBC News shows that the learned relations correlate with static semantic relatedness measures like WLM, and capture the evolving relationships among named entities over time.
Amara Tariq, Asim Karim, Hassan Foroosh
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 A Context-Driven Extractive Framework for Generating Realistic Image Descriptions
abstract
Automatic image annotation methods are extremely beneficial for image search, retrieval, and organization systems. The lack of strict correlation between semantic concepts and visual features, referred to as the semantic gap, is a huge challenge for annotation systems. In this paper, we propose an image annotation model that incorporates contextual cues collected from sources both intrinsic and extrinsic to images, to bridge the semantic gap. The main focus of this paper is a large real-world data set of news images that we collected. Unlike standard image annotation benchmark data sets, our data set does not require human annotators to generate artificial ground truth descriptions after data collection, since our images already include contextually meaningful and real-world captions written by journalists. We thoroughly study the nature of image descriptions in this real-world data set. News image captions describe both visual contents and the contexts of images. Auxiliary information sources are also available with such images in the form of news article and metadata (e.g., keywords and categories). The proposed framework extracts contextual-cues from available sources of different data modalities and transforms them into a common representation space, i.e., the probability space. Predicted annotations are later transformed into sentence-like captions through an extractive framework applied over news articles. Our context-driven framework outperforms the state of the art on the collected data set of approximately 20 000 items, as well as on a previously available smaller news images data set.
Amara Tariq, Hassan Foroosh
IEEE Trans. Image Process.2
2016 Character recognition in natural scene images using rank-1 tensor decomposition
abstract
Many approaches to solve the problem of scene character recognition utilize local features such as histograms of oriented gradients (HoG), SIFT, Shape Contexts (SC), Geometric Blur (GB), etc. An issue associated with these methods is the ad hoc rasterization of the local features into a single vector which perturbs the global spatial correlations that carry crucial information for recognition. To address this issue, we propose a novel holistic solution by incorporating tensor decomposition to get image features and utilizing image-to-class distance metric learning (I2CDML) for classification. For each training image, we first form a 3-mode tensor by rotating it through a sequence of angles. Then we perform rank-1 decomposition on the tensor to get the descriptor for each image. Utilizing the I2CDML framework, we then learn metrics for each class that are finally used to classify test images. We report results on popular natural scene character datasets, namely Chars74K-Font, Chars74K-Image, and ICDAR2003. We achieve results better than several baseline methods based on local features (e.g. HoG) and show that leave-random-one-out-cross validation yield even better recognition performance.
Hassan Foroosh
ICIP2
2016 Frequency estimation of sinusoids from nonuniform samples
Alam Abbas Syed, Qiyu Sun, Hassan Foroosh
Signal Process.3
2016 3D Pose Tracking With Multitemplate Warping and SIFT Correspondences
abstract
Template warping is a popular technique in vision-based 3D motion tracking and 3D pose estimation due to its flexibility of being applicable to monocular video sequences. However, the method suffers from two major limitations that hamper its successful use in practice. First, it requires the camera to be calibrated prior to applying the method. Second, it may fail to provide good results if the inter-frame displacements are too large. To overcome the first problem, we propose to estimate the unknown focal length of the camera from several initial frames by an iterative optimization process. To alleviate the second problem, we propose a tracking method based on combining complementary information provided by dense optical flow and tracked scale-invariant feature transform (SIFT) features. While optical flow is good for small displacements and provides accurate local information, tracked SIFT features are better at handling larger displacements or global transformations. To combine these two pieces of complementary information, we introduce a forgetting factor to bootstrap the 3D pose estimates provided by SIFT features, and refine the final results using optical flow. Experiments are performed on three public databases, i.e., the Biwi Head Pose dataset, the BU dataset, and the McGill Faces datasets. The results illustrate that the proposed solution provides more accurate results than baseline methods that rely solely on either template warping or SIFT features. In addition, the approach can be applied in a larger variety of scenarios, due to circumventing the need for camera calibration, thus providing a more flexible solution to the problem than existing methods.
Luming Liang, Wenzhang Liang, Hassan Foroosh
IEEE Trans. Circuits Syst. Video Technol.4
2015 SWIFT: Sparse Withdrawal of Inliers in a First Trial
abstract
We study the simultaneous detection of multiple structures in the presence of overwhelming number of outliers in a large population of points. Our approach reduces the problem to sampling an extremely sparse subset of the original population of data in one grab, followed by an unsupervised clustering of the population based on a set of instantiated models from this sparse subset. We show that the problem can be modeled using a multivariate hypergeometric distribution, and derive accurate mathematical bounds to determine a tight approximation to the sample size, leading thus to a sparse sampling strategy. We evaluate the method thoroughly in terms of accuracy, its behavior against varying input parameters, and comparison against existing methods, including the state of the art. The key features of the proposed approach are: (i) sparseness of the sampled set, where the level of sparseness is independent of the population size and the distribution of data, (ii) robustness in the presence of overwhelming number of outliers, and (iii) unsupervised detection of all model instances, i.e. without requiring any prior knowledge of the number of embedded structures. To demonstrate the generic nature of the proposed method, we show experimental results on different computer vision problems, such as detection of physical structures e.g. lines, planes, etc., as well as more abstract structures such as fundamental matrices, and homographies in multi-body structure from motion.
Maryam Jaberi, Marianna Pensky, Hassan Foroosh
CVPR3
2015 Sparse Convolutional Neural Networks
abstract
Deep neural networks have achieved remarkable performance in both image classification and object detection problems, at the cost of a large number of parameters and computational complexity. In this work, we show how to reduce the redundancy in these parameters using a sparse decomposition. Maximum sparsity is obtained by exploiting both inter-channel and intra-channel redundancy, with a fine-tuning step that minimize the recognition loss caused by maximizing sparsity. This procedure zeros out more than 90% of parameters, with a drop of accuracy that is less than 1% on the ILSVRC2012 dataset. We also propose an efficient sparse matrix multiplication algorithm on CPU for Sparse Convolutional Neural Networks (SCNN) models. Our CPU implementation demonstrates much higher efficiency than the off-the-shelf sparse matrix libraries, with a significant speedup realized over the original dense network. In addition, we apply the SCNN model to the object detection problem, in conjunction with a cascade model and sparse fully connected layers, to achieve significant speedups.
Baoyuan Liu, Hassan Foroosh, Marshall F. Tappen, Marianna Pensky
CVPR3
2015 Feature-independent context estimation for automatic image annotation
abstract
Automatic image annotation is a highly valuable tool for image search, retrieval and archival systems. In the absence of an annotation tool, such systems have to rely on either users' input or large amount of text on the webpage of the image, to acquire its textual description. Users may provide insufficient/noisy tags and all the text on the webpage may not be a description or an explanation of the accompanying image. Therefore, it is of extreme importance to develop efficient tools for automatic annotation of images with correct and sufficient tags. The context of the image plays a significant role in this process, along with the content of the image. A suitable quantification of the context of the image may reduce the semantic gap between visual features and appropriate textual description of the image. In this paper, we present an unsupervised feature-independent quantification of the context of the image through tensor decomposition. We incorporate the estimated context as prior knowledge in the process of automatic image annotation. Evaluation of the predicted annotations provides evidence of the effectiveness of our feature-independent context estimation method.
Amara Tariq, Hassan Foroosh
CVPR2
2015 Motion retrieval using consistency of epipolar geometry
abstract
In this paper, we present an efficient method for motion retrieval method based on the consistency of the homographies with the epipolar geometry. We treat the body pose as body point triplets and use the fact that each homography obtained from corresponding body point triplets should be consistent with epipolar geometry to estimate the similarity of two poses. We show that our method is invariant to camera internal parameters and viewpoint. Experiments are performed on the CMU MoCap dataset, and IXMAS dataset testing testing view-invariance, and action recognition. The results demonstrate that our method can accurately identify human action from video sequences when they are observed from totally different viewpoints with different camera parameters.
Nazim Ashraf, Hassan Foroosh
ICIP2
2015 T-clustering: Image clustering by tensor decomposition
abstract
Image clustering is an important tool for organizing evergrowing image repositories for efficient search and retrieval. A variety of clustering algorithms have been employed to cluster images. In this paper, we present a clustering algorithm, named T-Clustering, especially tailored to suit image collections. T-Clustering is based on tensor decomposition and takes into account the spatial configuration of images. This algorithm is non-parametric and works very well with raw images, thus alleviating the need for transformation of images in any feature domain. Our experiments prove that this algorithm outperforms well-known non-parametric clustering algorithms for a variety of image collections.
Amara Tariq, Hassan Foroosh
ICIP2
2015 Scene Text Deblurring Using Text-Specific Multiscale Dictionaries
abstract
Texts in natural scenes carry critical semantic clues for understanding images. When capturing natural scene images, especially by handheld cameras, a common artifact, i.e., blur, frequently happens. To improve the visual quality of such images, deblurring techniques are desired, which also play an important role in character recognition and image understanding. In this paper, we study the problem of recovering the clear scene text by exploiting the text field characteristics. A series of text-specific multiscale dictionaries (TMD) and a natural scene dictionary is learned for separately modeling the priors on the text and nontext fields. The TMD-based text field reconstruction helps to deal with the different scales of strings in a blurry image effectively. Furthermore, an adaptive version of nonuniform deblurring method is proposed to efficiently solve the real-world spatially varying problem. Dictionary learning allows more flexible modeling with respect to the text field property, and the combination with the nonuniform method is more appropriate in real situations where blur kernel sizes are depth dependent. Experimental results show that the proposed method achieves the deblurring results with better visual quality than the state-of-the-art methods.
Xiaochun Cao, Wenqi Ren, Wangmeng Zuo, Xiaojie Guo 0001, Hassan Foroosh
IEEE Trans. Image Process.5
2015 Constrained Multi-View Video Face Clustering
abstract
In this paper, we focus on face clustering in videos. To promote the performance of video clustering by multiple intrinsic cues, i.e., pairwise constraints and multiple views, we propose a constrained multi-view video face clustering method under a unified graph-based model. First, unlike most existing video face clustering methods which only employ these constraints in the clustering step, we strengthen the pairwise constraints through the whole video face clustering framework, both in sparse subspace representation and spectral clustering. In the constrained sparse subspace representation, the sparse representation is forced to explore unknown relationships. In the constrained spectral clustering, the constraints are used to guide for learning more reasonable new representations. Second, our method considers both the video face pairwise constraints as well as the multi-view consistence simultaneously. In particular, the graph regularization enforces the pairwise constraints to be respected and the co-regularization penalizes the disagreement among different graphs of multiple views. Experiments on three real-world video benchmark data sets demonstrate the significant improvements of our method over the state-of-the-art methods.
Xiaochun Cao, Changqing Zhang 0002, Chengju Zhou, Huazhu Fu, Hassan Foroosh
IEEE Trans. Image Process.5
2015 Exploring Sparseness and Self-Similarity for Action Recognition
abstract
We propose that the dynamics of an action in video data forms a sparse self-similar manifold in the space-time volume, which can be fully characterized by a linear rank decomposition. Inspired by the recurrence plot theory, we introduce the concept of Joint Self-Similarity Volume (Joint-SSV) to model this sparse action manifold, and hence propose a new optimized rank-1 tensor approximation of the Joint-SSV to obtain compact low-dimensional descriptors that very accurately characterize an action in a video sequence. We show that these descriptor vectors make it possible to recognize actions without explicitly aligning the videos in time in order to compensate for speed of execution or differences in video frame rates. Moreover, we show that the proposed method is generic, in the sense that it can be applied using different low-level features, such as silhouettes, tracked points, histogram of oriented gradients, and so forth. Therefore, our method does not necessarily require explicit tracking of features in the space-time volume. Our experimental results on five public data sets demonstrate that our method produces promising results and outperforms many baseline methods.
Imran N. Junejo, Marshall F. Tappen, Hassan Foroosh
IEEE Trans. Image Process.4
2014 Action Recognition by Weakly-Supervised Discriminative Region Localization
Hakan Boyraz, Syed Zain Masood, Baoyuan Liu, Marshall F. Tappen, Hassan Foroosh
BMVC5
2014 Feature-Independent Action Spotting without Human Localization, Segmentation, or Frame-wise Tracking
abstract
In this paper, we propose an unsupervised framework for action spotting in videos that does not depend on any specific feature (e.g. HOG/HOF, STIP, silhouette, bag-of-words, etc.). Furthermore, our solution requires no human localization, segmentation, or framewise tracking. This is achieved by treating the problem holistically as that of extracting the internal dynamics of video cuboids by modeling them in their natural form as multilinear tensors. To extract their internal dynamics, we devised a novel Two-Phase Decomposition (TP-Decomp) of a tensor that generates very compact and discriminative representations that are robust to even heavily perturbed data. Technically, a Rank-based Tensor Core Pyramid (Rank-TCP) descriptor is generated by combining multiple tensor cores under multiple ranks, allowing to represent video cuboids in a hierarchical tensor pyramid. The problem then reduces to a template matching problem, which is solved efficiently by using two boosting strategies: (1) to reduce search space, we filter the dense trajectory cloud extracted from the target video, (2) to boost the matching speed, we perform matching in an iterative coarse-to-fine manner. Experiments on 5 benchmarks show that our method outperforms current state-of-the-art under various challenging conditions. We also created a challenging dataset called Heavily Perturbed Video Array (HPVA) to validate the robustness of our framework under heavily perturbed situations.
Marshall F. Tappen, Hassan Foroosh
CVPR3
2014 Mesh-free sparse representation of multidimensional LiDAR data
abstract
Modern LiDAR collection systems generate very large data sets approaching several million to billions of point samples per product. Compression techniques have been developed to help manage the large data sets. However, sparsifying LiDAR survey data by means other than random decimation remains largely unexplored. In contrast, surface model simplification algorithms are well-established, especially with respect to the complementary problem of surface reconstruction. Unfortunately, surface model simplification algorithms are often not directly applicable to LiDAR survey data due to the true 3D nature of the data sets. Further, LiDAR data is often attributed with additional user data that should be considered as potentially salient information. This paper makes the following main contributions in this area: (i) We generalize some features defined on spatial coordinates to arbitrary dimensions and extend these features to provide local multidimensional statistics. (ii) We propose an approach for sparsifying point clouds similar to mesh-free surface simplification that preserves saliency with respect to the multidimensional information content. (iii) We show direct application to LiDAR data and evaluate the benefits in terms of level of sparsity versus entropy.
Kristian L. Damkjer, Hassan Foroosh
ICIP2
2014 Should we discard sparse or incomplete videos?
abstract
In this paper, we determine whether incomplete videos that are often discarded carry useful information for action recognition, and if so, how one can represent such mixed collection of video data (complete versus incomplete, and labeled versus unlabeled) in a unified manner. We propose a novel framework to handle incomplete videos in action classification, and make three main contributions: (1) We cast the action classification problem for a mixture of complete and incomplete data as a semi-supervised learning problem of labeled and unlabeled data. (2) We introduce a two-step approach to convert the input mixed data into a uniform compact representation. (3) Exhaustively scrutinizing 280 configurations, we experimentally show on our two created benchmarks that, even the videos are extremely sparse and incomplete, it is still possible to recover useful information from them, and classify unknown actions by a graph based semi-supervised learning framework.
Hassan Foroosh
ICIP2
2014 Scene-based automatic image annotation
abstract
Image search and retrieval systems depend heavily on availability of descriptive textual annotations with images, to match them with textual queries of users. In most cases, such systems have to rely on users to provide tags or keywords with images. Users may add insufficient or noisy tags. A system to automatically generate descriptive tags for images can be extremely helpful for search and retrieval systems. Automatic image annotation has been explored widely in both image and text processing research communities. In this paper, we present a novel approach to tackle this problem by incorporating contextual information provided by scene analysis of image. Image can be represented by features which indicate type of scene shown in the image, instead of representing individual objects or local characteristics of that image. We have used such features to provide context in the process of predicting tags for images.
Amara Tariq, Hassan Foroosh
ICIP2
2014 View invariant action recognition using projective depth
Nazim Ashraf, Hassan Foroosh
Comput. Vis. Image Underst.3
2013 View invariant action recognition using weighted fundamental ratios
Nazim Ashraf, Yuping Shen, Xiaochun Cao, Hassan Foroosh
Comput. Vis. Image Underst.4
2012 Human action recognition in video data using invariant characteristic vectors
abstract
We introduce the concept of the “characteristic vector” as an invariant vector associated with a set of freely moving points relative to a plane. We show that if the motion of two sets of points differ only up to a similarity transformation, then the elements of the characteristic vector differ up to scale regardless of viewing directions and cameras. Furthermore, this invariant vector is given by any arbitrary homography that is consistent with epipolar geometry. The characteristic vector of moving points can thus be used to recognize the transitions of a set of points in an articulated body during the course of an action regardless of the camera orientation and parameters. Our extensive experimental results on both motion capture data and real data indicates very good performance.
Nazim Ashraf, Hassan Foroosh
ICIP2
2012 Motion retrival using low-rank decomposition of Fundamental Ratios
abstract
This paper proposes a novel framework for efficient retrieval of motion capture data. The method uses Fundamental Ratios to convert action sequences into compact representations of the action, greatly reducing the spatiotemporal dimensionality of the sequences. We propose a low-rank decomposition scheme that allows for converting the motion sequence volumes into compact lower dimensional representations, without losing the nonlinear dynamics of the motion manifold, and the proposed method performs well even when interclass differences are small or intra-class differences are large. We evaluate the performance of our retrieval framework on the CMU mocap dataset and Microsoft Kinect dataset, which demonstrate satisfying retrieval rates.
Nazim Ashraf, Hassan Foroosh
ICIP3
2012 Optimizing PTZ camera calibration from two images
Imran N. Junejo, Hassan Foroosh
Mach. Vis. Appl.2
2011 Action recognition using rank-1 approximation of Joint Self-Similarity Volume
abstract
In this paper, we make three main contributions in the area of action recognition: (i) We introduce the concept of Joint Self-Similarity Volume (Joint SSV) for modeling dynamical systems, and show that by using a new optimized rank-1 tensor approximation of Joint SSV one can obtain compact low-dimensional descriptors that very accurately preserve the dynamics of the original system, e.g. an action video sequence; (ii) The descriptor vectors derived from the optimized rank-1 approximation make it possible to recognize actions without explicitly aligning the action sequences of varying speed of execution or different frame rates; (iii) The method is generic and can be applied using different low-level features such as silhouettes, histogram of oriented gradients, etc. Hence, it does not necessarily require explicit tracking of features in the space-time volume. Our experimental results on three public datasets demonstrate that our method produces remarkably good results and outperforms all baseline methods.
Imran N. Junejo, Hassan Foroosh
ICCV3
2011 Selective subtraction when the scene cannot be learned
abstract
Background subtraction techniques model the background of the scene using the stationarity property and classify the scene into two classes of foreground and background. In doing so, most moving objects become foreground indiscriminately, except for perhaps some waving tree leaves, water ripples, or a water fountain, which are typically “learned” as part of the background using a large training set of video data. We introduce a novel concept of background as the objects other than the foreground, which may include moving objects in the scene that cannot be learned from a training set because they occur only irregularly and sporadically, e.g. a walking person. We propose a “selective subtraction” method as an alternative to standard background subtraction, and show that a reference plane in a scene viewed by two cameras can be used as the decision boundary between foreground and background. In our definition, the foreground may actually occur behind a moving object. Furthermore, the reference plane can be selected in a very flexible manner, using for example the actual moving objects in the scene, if needed. We present diverse set of examples to show that: (i) the technique performs better than standard background subtraction techniques without the need for training, camera calibration, disparity map estimation, or special camera configurations; (ii) it is potentially more powerful than standard methods because of its flexibility of making it possible to select in real-time what to filter out as background, regardless of whether the object is moving or not, or whether it is a rare event or a frequent one.
Adeel A. Bhutta, Imran N. Junejo, Hassan Foroosh
ICIP3
2011 Motion Retrieval Using Low-Rank Subspace Decomposition of Motion Volume
abstract
Abstract This paper proposes a novel framework that allows for a flexible and an efficient retrieval of motion capture data in huge databases. The method first converts an action sequence into a novel representation, i.e. the Self‐Similarity Matrix (SSM), which is based on the notion of self‐similarity. This conversion of the motion sequences into compact and low‐rank subspace representations greatly reduces the spatiotemporal dimensionality of the sequences. The SSMs are then used to construct order‐3 tensors, and we propose a low‐rank decomposition scheme that allows for converting the motion sequence volumes into compact lower dimensional representations, without losing the nonlinear dynamics of the motion manifold. Thus, unlike existing linear dimensionality reduction methods that distort the motion manifold and lose very critical and discriminative components, the proposed method performs well even when inter‐class differences are small or intra‐class differences are large. In addition, the method allows for an efficient retrieval and does not require the time‐alignment of the motion sequences. We evaluate the performance of our retrieval framework on the CMU mocap dataset under two experimental settings, both demonstrating promising retrieval rates.
Imran N. Junejo, Hassan Foroosh
Comput. Graph. Forum3
2010 View-Invariant Action Recognition Using Rank Constraint
abstract
We propose a new method for view-invariant action recognition based on the rank constraint on the family of planar homographies associated with triplets of body points. We represent action as a sequence of poses and we use the fact that the family of homographies associated with two identical poses would have rank 4 to gauge similarity of the pose between two subjects, observed by different perspective cameras and from different viewpoints. Extensive experimental results show that our method can accurately identify action from video sequences when they are observed from totally different viewpoints with different camera parameters.
Nazim Ashraf, Yuping Shen, Hassan Foroosh
ICPR3
2010 GPS coordinates estimation and camera calibration from solar shadows
Imran N. Junejo, Hassan Foroosh
Comput. Vis. Image Underst.2
2010 Camera calibration and geo-location estimation from two shadow trajectories
Lin Wu 0001, Xiaochun Cao, Hassan Foroosh
Comput. Vis. Image Underst.3
2010 Video synchronization and its application to object transfer
Xiaochun Cao, Lin Wu 0001, Jiangjian Xiao, Hassan Foroosh, Jigui Zhu, Xiaohong Li 0001
Image Vis. Comput.4
2009 View-Invariant Action Recognition from Point Triplets
abstract
We propose a new view-invariant measure for action recognition. For this purpose, we introduce the idea that the motion of an articulated body can be decomposed into rigid motions of planes defined by triplets of body points. Using the fact that the homography induced by the motion of a triplet of body points in two identical pose transitions reduces to the special case of a homology, we use the equality of two of its eigenvalues as a measure of the similarity of the pose transitions between two subjects, observed by different perspective cameras and from different viewpoints. Experimental results show that our method can accurately identify human pose transitions and actions even when they include dynamic timeline maps, and are obtained from totally different viewpoints with different unknown camera parameters.
Yuping Shen, Hassan Foroosh
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 View-invariant action recognition using fundamental ratios
abstract
A moving plane observed by a fixed camera induces a fundamental matrix F across multiple frames, where the ratios among the elements in the upper left 2×2 submatrix are herein referred to as the Fundamental Ratios. We show that fundamental ratios are invariant to camera parameters, and hence can be used to identify similar plane motions from varying viewpoints. For action recognition, we decompose a body posture into a set of point triplets (planes). The similarity between two actions is then determined by the motion of point triplets and hence by their associated fundamental ratios, providing thus view-invariant recognition of actions. Results evaluated over 255 semi-synthetic video data with 100 independent trials at a wide range of noise levels, and also on 56 real videos of 8 different classes of actions, confirm that our method can recognize actions under substantial amount of noise, even when they have dynamic timeline maps, and the viewpoints and camera parameters are unknown and totally different.
Yuping Shen, Hassan Foroosh
CVPR2
2008 View-invariant recognition of body pose from space-time templates
abstract
We propose a new template-based approach for view invariant recognition of body poses, based on geometric constraints derived from the motion of body point triplets. In addition to spatial information our templates encode temporal information of body pose transitions. Unlike existing methods that study a body pose as a whole, we decompose it into a number of body point triplets, and compare their motions to our templates. Using the fact that the homography induced by the motion of a triplet of body points in two identical body pose transitions reduces to the special case of a homology, we exploit the equality of two of its eigenvalues to impose constraints on the similarity of the pose transitions between two subjects, observed by different perspective cameras and from different viewpoints. Extensive experimental results show that our method can accurately identify human poses from video sequences when they are observed from totally different viewpoints with different camera parameters.
Yuping Shen, Hassan Foroosh
CVPR2
2008 Estimating Geo-temporal Location of Stationary Cameras Using Shadow Trajectories
Imran N. Junejo, Hassan Foroosh
ECCV (1)2
2008 Using solar shadow trajectories for camera calibration
abstract
In this paper we show that solar shadow trajectories can be used for a robust camera calibration. Expanding on the previous work on calibration from solar shadows, we relax the condition that the shadow casting object be visible in the image. This enables us to work with shadows of non-vertical objects as well. The important observation that we make in this work is that these shadow trajectories form an interesting geometry on the ground plane. Using properties of these cast shadows, the horizon line (or the line at infinity) of the ground plane is robustly estimated. This leads to pole-polar constraints on the image of the absolute conic (IAC), which we decompose for estimating the camera parameters. We show that our method performs well in presence of large noise. We perform experiments with synthetic data and real data captured from live webcams, demonstrating encouraging results.
Imran N. Junejo, Hassan Foroosh
ICIP2
2008 Practical PTZ camera calibration using Givens rotations
abstract
Due to their ease of use and availability, pan-tilt-zoom (PTZ) cameras are everywhere. However, in order to utilize these cameras for a meaningful application, it is necessary that these cameras be calibrated. In this paper, we address the problem of calibrating such a rotating and zooming camera and present a simple yet novel solution. We show that in the case of a general rotation the problem can be solved in closed- form by applying a series of Givens rotations, providing five intrinsic parameters from only a minimum set of two images. Whereas other methods use orthogonality of the rotation matrix as a constraint, our approach applies direct decomposition of the infinite homography As a result, we are able to obtain four constraints for estimating camera parameters. We rigourously test our method on synthetic data by introducing very large noise. We also compare our method with the state- of-the-art PTZ camera calibration method. Our solutions and analysis are thoroughly validated on both synthetic and real data.
Imran N. Junejo, Hassan Foroosh
ICIP2
2008 Robust auto-calibration of a PTZ camera with non-overlapping FOV
abstract
We consider the problem of auto-calibration of cameras, which are fixed in location but are free to rotate while changing their internal parameters by zooming. Our method is based on line correspondences between two views, which may have non-overlapping field of view. Camera calibration from images having non-overlapping field of view is the basic motivation behind this research. The key observation is that the planes formed by the optic center and the line correspondences are really the same plane. We use this fact together with the orthonormality constraint of the rotation matrix to estimate the unknown camera parameters. We show experimental results on synthetic and real data, and analyze the accuracy and stability of our method.
Nazim Ashraf, Hassan Foroosh
ICPR2
2008 Practical pure pan and pure tilt camera calibration
abstract
Often the deployed pan-tilt-zoom (PTZ) cameras undergo a pure pan or pure tilt rotation. This is a degenerate case for most of the PTZ camera calibration methods. That is, under this motion, the estimated camera parameters are not unique. In this regard, we present a novel camera calibration method to estimate five camera parameters from these pure pan/tilt cameras by using only two images. Our solution is based on using the infinite homography and performing its eigendecomposition. Our solutions and analyses are thoroughly validated and tested on both synthetic and real data, whereby the proposed method is shown to be accurate and noise resilient.
Imran N. Junejo, Hassan Foroosh
ICPR2
2008 GPS coordinate estimation from calibrated cameras
abstract
We present a novel application where, using only three points from the shadow trajectory of an object, one can accurately determine the geo-location of the camera, up to a longitude ambiguity, and also the date of image acquisition without using any GPS or other special instruments. We refer to this as ldquogeotemporal localizationrdquo. We consider possible cases where ambiguities can be removed if additional information is available. Our method does not require any knowledge of the date or the time when the pictures are taken, and geo-temporal information is recovered directly from the images. We demonstrate the accuracy of our geo-temporal localization method using synthetic and real data.
Imran N. Junejo, Hassan Foroosh
ICPR2
2008 Refining PTZ camera calibration
abstract
Due to the increased need for security and surveillance, PTZ cameras are now being widely used in many domains. Therefore, it is very important for the applications like video mosaic generation or automatic surveillance that these camera be accurately calibrated. In this paper, we address the problem of parameter refinement for such pan-tilt-zoom (PTZ) cameras. Use of bundle-adjustment for parameter refinement has widely been adopted in the computer vision field. However, as has been shown by researchers, in presence of noise, this Maximum Likelihood estimate looses its optimality. We propose a novel statistically optimal error function that is shown to experimentally outperform this ML estimate in presence of significant noise. We perform tests on synthetic as well as on real data to verify our method.
Imran N. Junejo, Hassan Foroosh
ICPR2
2008 Action recognition based on homography constraints
abstract
In this paper, we present a new approach for view-invariant action recognition using constraints derived from the eigenvalues of planar homographies associated with triplets of body points. Unlike existing methods that study an action as a whole, or break it down into individual poses, we represent an action as a sequence of pose transitions. Using the fact that the homography induced by the motion of a triplet of body points in two identical pose transitions reduces to the special case of a homology, we exploit the equality of two of its eigenvalues to impose constraints on the similarity of the pose transitions between two subjects, observed by different perspective cameras and from different viewpoints. Experimental results show that our method can accurately identify human pose transitions and actions even when they include dynamic timeline maps, and are obtained from totally different viewpoints with different camera parameters.
Yuping Shen, Nazim Ashraf, Hassan Foroosh
ICPR3
2008 Euclidean path modeling for video surveillance
Imran N. Junejo, Hassan Foroosh
Image Vis. Comput.2
2008 Phase-Shifting for Nonseparable 2-D Haar Wavelets
abstract
In this paper, we present a novel and efficient solution to phase-shifting 2-D nonseparable Haar wavelet coefficients. While other methods either modify existing wavelets or introduce new ones to handle the lack of shift-invariance, we derive the explicit relationships between the coefficients of the shifted signal and those of the unshifted one. We then establish their computational complexity, and compare and demonstrate the superior performance of the proposed approach against classical interpolation tools in terms of accumulation of errors under successive shifting.
Mais Alnasser, Hassan Foroosh
IEEE Trans. Image Process.2
2007 Near-Optimal Mosaic Selection for Rotating and Zooming Video Cameras
Nazim Ashraf, Imran N. Junejo, Hassan Foroosh
ACCV (2)3
2007 Euclidean Path Modeling from Ground and Aerial Views
abstract
We address the issue of Euclidean path modeling in a single camera for activity monitoring in a multi-camera video surveillance system. The paper proposes a novel linear solution to auto-calibrate any camera observing pedestrians and uses these calibrated cameras to detect unusual object behavior. The input trajectories are metric rectified and the input sequences are registered to the satellite imagery and prototype path models are constructed. During the testing phase, using our simple yet efficient similarity measures, we seek a relation between the input trajectories derived from a sequence and the prototype path models. Real-world pedestrian sequences are used to demonstrate the practicality of the proposed method.
Imran N. Junejo, Hassan Foroosh
CVPR2
2007 Trajectory Rectification and Path Modeling for Video Surveillance
abstract
Path modeling for video surveillance is an active area of research. We address the issue of Euclidean path modeling in a single camera for activity monitoring in a multi- camera video surveillance system. The paper proposes (i) a novel linear solution to auto-calibrate any camera observing pedestrians and (ii) to use these calibrated cameras to detect unusual object behavior. During the unsupervised training phase, after auto-calibrating a camera and metric rectifying the input trajectories, the input sequences are registered to the satellite imagery and prototype path models are constructed. This allows us to estimate metric information directly from the video sequences. During the testing phase, using our simple yet efficient similarity measures, we seek a relation between the input trajectories derived from a sequence and the prototype path models. We test the proposed method on synthetic as well as on real-world pedestrian sequences.
Imran N. Junejo, Hassan Foroosh
ICCV2
2007 Robust Auto-Calibration using Fundamental Matrices Induced by Pedestrians
abstract
The knowledge of camera intrinsic and extrinsic parameters is useful, as it allows us to make world measurements. Unfortunately, calibration information is rarely available in video surveillance systems and is difficult to obtain once the system is installed. Auto-calibrating cameras using moving objects (humans) has recently attracted a lot of interest. Two methods were proposed by Lv-Nevatia (2002) and Krahnstoever-Mendonca (2005). The inherent difficulty of the problem lies in the noise that is generally present in the data. We propose a robust and a general linear solution to the problem by adopting a formulation different from the existing methods. The uniqueness of our formulation lies in recognizing two fundamental matrices present in the geometry obtained by observing pedestrians, and then using their properties to impose linear constraints on the unknown camera parameters. Experiments with synthetic as well as real data are presented -indicating the practicality of the proposed system.
Imran N. Junejo, Nazim Ashraf, Yuping Shen, Hassan Foroosh
ICIP (3)4
2007 Using Calibrated Camera for Euclidean Path Modeling
abstract
In this paper, we address the issue of Euclidean path modeling in a single camera for activity monitoring in a multi-camera video surveillance system. The paper proposes to use calibrated cameras to detect unusual object behavior. During the unsupervised training phase, after metric rectifying the input trajectories, the input sequences are registered to the satellite imagery and prototype path models are constructed. During the testing phase, using our simple yet efficient similarity measures, we seek a relation between the input trajectories derived from a sequence and the prototype path models. Real-world pedestrian sequences are used to demonstrate the practicality of the proposed method.
Imran N. Junejo, Hassan Foroosh
ICIP (3)2
2007 Camera calibration and light source orientation from solar shadows
Xiaochun Cao, Hassan Foroosh
Comput. Vis. Image Underst.2
2007 Autoconfiguration of a Dynamic Nonoverlapping Camera Network
abstract
In order to monitor sufficiently large areas of interest for surveillance or any event detection, we need to look beyond stationary cameras and employ an automatically configurable network of nonoverlapping cameras. These cameras need not have an overlapping field of view and should be allowed to move freely in space. Moreover, features like zooming in/out, readily available in security cameras these days, should be exploited in order to focus on any particular area of interest if needed. In this paper, a practical framework is proposed to self-calibrate dynamically moving and zooming cameras and determine their absolute and relative orientations, assuming that their relative position is known. A global linear solution is presented for self-calibrating each zooming/focusing camera in the network. After self-calibration, it is shown that only one automatically computed vanishing point and a line lying on any plane orthogonal to the vertical direction is sufficient to infer the dynamic network configuration. Our method generalizes previous work which considers restricted camera motions. Using minimal assumptions, we are able to successfully demonstrate promising results on synthetic, as well as on real data.
Imran N. Junejo, Xiaochun Cao, Hassan Foroosh
IEEE Trans. Syst. Man Cybern. Part B3
2006 Geometry of a Non-Overlapping Multi-Camera Network
abstract
Moving beyond single or stationary camera surveillance systems, we employ an automatically configurable network of non-overlapping cameras. These cameras need not have an overlapping Field of View (FoV) and should be able to move freely in space. In this paper, a practical framework is proposed that determines the geometry of such a dynamic camera network. Assuming each camera in the network is calibrated, it is shown that only one automatically computed vanishing point and a line lying on any plane orthogonal to the vertical direction is sufficient to infer the dynamic network configuration. Our method generalizes previous work which considers restricted camera motions. Using minimal assumptions, we are able to successfully demonstrate promising results on synthetic as well as on real data.
Imran N. Junejo, Xiaochun Cao, Hassan Foroosh
AVSS3
2006 Dissecting the Image of the Absolute Conic
abstract
In this paper, we revisit the role of the image of the absolute conic (IAC) in recovering the camera geometry. We derive new constraints on IAC that advance our understanding of its underlying building blocks. The new constraints are shown to be intrinsic to IAC, rather than exploiting the scene geometry or the prior knowledge on the camera. We provide geometric interpretations for these new intrinsic constraints, and show their relations to the invariant properties of the IAC. This in turn provides a better insight into the role that IAC plays in determining the camera internal geometry. Since the new constraints are invariant properties of the IAC, they can be used to reparameterize its elements. We show that such reparameterization would allow to recover a more general camera geometry from a single view, compared to existing methods. We apply the new constraints to single view calibration using vanishing points, investigate the error resilience, and compare our results to Liebowitz-Zisserman (1999).
Imran N. Junejo, Hassan Foroosh
AVSS2
2006 Robust Auto-Calibration from Pedestrians
abstract
The knowledge of camera intrinsic and extrinsic parameters is useful, as it allows us to make world measurements. Unfortunately, calibration information is rarely available in video surveillance systems and it is difficult to obtain once the system is installed. Auto-calibrating cameras using moving objects (humans) has recently attracted a lot of interest. Two methods are proposed by Lv-Nevatia(2002) and Krahnstoever-Mendonca(2005). The inherent difficulty of the problem lies in the noise that is generally present in the data. We propose a robust and a general linear solution to the problem by adopting a formulation different from the existing methods. The uniqueness of formulation lies in recognizing two harmonic homologies present in the geometry obtained by observing pedestrians, and then using properties of these homologies to obtain linear constraints on the unknown camera parameters. Experiments with synthetic as well as on real data are presented - indicating the practicality of the proposed system.
Imran N. Junejo, Hassan Foroosh
AVSS2
2006 Rendering Synthetic Objects in Natural Scenes
abstract
We present a method for solving the light integral problem for synthetic diffuse objects rendered within a natural scene. The approach generates realistic shading using only a few images of the surroundings of the object to represent the ambient light. We model the global illumination integral using Chebyshev polynomials. We show that due to the orthogonality of 2D Chebyshev moments, the global illumination integral can reduce to the inner product of two vectors, representing the irradiance and the bidirectional reflectance distribution function (BRDF). The Chebyshev moments of these two functions are computed off-line and stored in the memory. The rendering of the object in the scene then becomes a simple problem of computing the inner product of the two vectors for each point.
Mais Alnasser, Hassan Foroosh
ICIP2
2006 Image-Based Simulation of Gaseous Material
abstract
We present a method for real-time image based simulation of gaseous material such as fire and smoke. We model the kinematics of these phenomena by a mass-spring system, and their turbulent visual dynamics by texture sequencing with variable speeds and transparencies that depend on the speed of the vaporized fuel. The approach allows for incorporating external forces such as the gravity, and the wind force for added realism. Also, a specific characteristic of our method is that any object inserted in a flame or inside smoke can be modeled simply as an external force in the mass-spring system, making real-time interactions with these phenomena easy to integrate. While the approach maintains an extremely low computational cost, these flexibilities increase the realism of our simulation by allowing for real-time control and the interactivity, which is a highly desirable requirement in applications such as augmented reality.
Murat Balci, Mais Alnasser, Hassan Foroosh
ICIP3
2006 Real-time 3D fire simulation using a spring-mass model
abstract
We present a method for real-time simulation of 3D fire inspired by an old mechanical trick known as the "silk torch". Motivated by the proven illusive effect of silk torch, we model the kinematics of flames by a mass-spring system and its turbulent visual dynamics by texture sequencing with variable speeds and transparencies that depend on the speed of the vaporized fuel. The approach allows for incorporating external forces such as the gravity, and the wind force for added realism. Also, a specific characteristic of our method is that any object inserted in aflame can be modeled simply as an external force in the mass-spring system, making real-time interactions with fire a simple addition to the overall system. While the approach maintains an extremely low computational cost, these flexibilities increase the realism of our 3D fire by allowing for real-time control and the interactivity, which is a highly desirable requirement in applications such as augmented reality
Murat Balci, Hassan Foroosh
MMM2
2006 A new framework for video cut and paste
abstract
In this paper, we describe a novel framework to cut and paste objects among different video shots. First, based on geometric analysis of the camera motion of a shot, we classify the shot into either simple camera motion shot, e.g. panning, tracking and zooming, or complex camera motion shot, e.g. hand-held shots. Next, for the simple camera motion shots, we temporally align them captured by cameras undergoing the same motions based on the quantization in the geometric analysis. In the case of complex camera motion, we recover the camera poses for each frame of both the source and target shots and, thus, for each frame in the target shot, we find the corresponding source frame with the closest viewing direction. Then, the foreground objects in the source shots are automatically cut by combining the merits of motion layer segmentation and alpha matting techniques. Finally, the extracted foreground mattes can be directly blended into the corresponding target frames for simple motion shots. For complex shots, combining the estimated rough depths of the foreground objects, foreground layers are rendered and blended into the target frames. Results on various content types, e.g. home videos and feature films, and different camera motions are reported for validating the proposed framework.
Jiangjian Xiao, Xiaochun Cao, Hassan Foroosh
MMM3
2006 3D Object Transfer Between Non-Overlapping Videos
abstract
Given two video sequences of different scenes acquired with moving cameras, it is interesting to seamlessly transfer a 3D object from one sequence to the other. In this paper, we present a video-based approach to extract the alpha mattes of rigid or approximately rigid 3D objects from one or more source videos, and then geometrycorrectly transfer them into another target video of a different scene. Our framework builds upon techniques in camera pose estimation, 3D spatiotemporal video alignment, depth recovery, key-frame editing, natural video matting, and image-based rendering. Based on the explicit camera pose estimation, the camera trajectories of the source and target videos are aligned in 3D space. Combinied with the estimated dense depth information, this allows us to significantly relieve the burdens of key-frame editing and efficiently improve the quality of video matting. During the transfer, our approach not only correctly restores the geometric deformation of the 3D object due to the different camera trajectories, but also effectively retains the soft shadow and environmental lighting properties of the object to ensure that the augmenting object is in harmony with the target scene.
Jiangjian Xiao, Xiaochun Cao, Hassan Foroosh
VR3
2006 Self-calibration from turn-table sequences in presence of zoom and focus
Xiaochun Cao, Jiangjian Xiao, Hassan Foroosh, Mubarak Shah
Comput. Vis. Image Underst.3
2006 Subpixel Estimation of Shifts Directly in the Fourier Domain
abstract
In this paper, we establish the exact relationship between the continuous and the discrete phase difference of two shifted images, and show that their discrete phase difference is a two-dimensional sawtooth signal. Subpixel registration can, thus, be performed directly in the Fourier domain by counting number of cycles of the phase difference matrix along each frequency axis. The subpixel portion is given by the noninteger fraction of the last cycle along each axis. The problem is formulated as an overdetermined homogeneous quadratic cost function under rank constraint for the phase difference, and the shape constraint for the filter that computes the group delay. The optimal tradeoff for imposing the constraints is determined using the method of generalized cross validation. Also, in order to robustify the solution, we assume a mixture model of inlying and outlying estimated shifts and truncate our quadratic cost function using expectation maximization.
Murat Balci, Hassan Foroosh
IEEE Trans. Image Process.2
2006 Camera Calibration Using Symmetric Objects
abstract
This paper proposes a novel method for camera calibration using images of a mirror symmetric object. Assuming unit aspect ratio and zero skew, we show that interimage homographies can be expressed as a function of only the principal point. By minimizing symmetric transfer errors, we thus obtain an accurate solution for the camera parameters. We also extend our approach to a calibration technique using images of a 1-D object with a fixed pivoting point. Unlike existing methods that rely on orthogonality or pole-polar relationship, our approach utilizes new inter-image constraints and does not require knowledge of the 3-D coordinates of feature points. To demonstrate the effectiveness of the approach, we present results for both synthetic and real images.
Xiaochun Cao, Hassan Foroosh
IEEE Trans. Image Process.2
2005 Inferring motion from the rank constraint of the phase matrix
abstract
We investigate the rank constraint of the discrete phase difference, and derive its exact parametric model. We show that the discrete phase difference of two shifted images, or their subregions, is a 2-dimensional sawtooth signal. This allows us to determine the motion parameters to subpixel accuracy by simply counting the number of cycles of the phase difference along each frequency axis. The subpixel portion is given by the non-integer fraction of the last cycle along each axis. The problem is formulated as a homogeneous cost function under rank constraint for the phase matrix, and the shape constraint for the filter that computes the group delay, and is solved using a robust technique.
Murat Balci, Hassan Foroosh
ICASSP (2)2
2005 Self-calibrated reconstruction of partially viewed symmetric objects
abstract
Traditional stereo reconstruction techniques, based on point correspondences and estimation of the cameras from the fundamental matrix, introduce a four-fold ambiguity. Moreover, there is a projective ambiguity inherent in the fundamental matrix. We show that a symmetric object can be modeled, even under partial occlusion, by a pair of uncalibrated stereo images. This implies that, unlike traditional stereo algorithms, we can extract 3D information from two arbitrary viewpoints, even when there is no left-to-right point correspondences. To demonstrate the effectiveness of the method, we present experimental results on both synthetic and real images.
Hassan Foroosh, Murat Balci, Xiaochun Cao
ICASSP (2)1
2005 Estimating sub-pixel shifts directly from the phase difference
abstract
In this paper, we establish the exact relationship between the continuous and the discrete phase-difference of two shifted images, and show that their discrete phase difference is a 2-dimensional sawtooth signal. As a result, the exact shifts between two images can be determined to sub-pixel accuracy by counting the number of cycles of the phase difference matrix along each frequency axis. The sub-pixel portion is represented by a fraction of a cycle corresponding to the non-integer part of the shift. The problem is formulated as an over-determined system of equations, and is solved by imposing a regularity constraint, using the method of generalized cross validation (GCV).
Murat Balci, Hassan Foroosh
ICIP (1)2
2005 Metrology in uncalibrated images given one vanishing point
abstract
In this paper, we describe how 3D Euclidean measurements can be made in a pair of uncalibrated images, when only minimal geometric information are available in the image planes. This minimal information consists of a line in a reference plane, and the vanishing point orthogonal to it. Given such limited information, we show that the length ratio of two objects perpendicular to the reference plane can be expressed as a function of the camera intrinsic parameters. Assuming that the camera intrinsic parameters remain invariant between two views, we perform Euclidean metric measurements directly in the perspective images.
Hassan Foroosh, Xiaochun Cao, Murat Balci
ICIP (3)1
2005 Pixelwise-adaptive blind optical flow assuming nonstationary statistics
abstract
In this paper, we address some of the major issues in optical flow within a new framework assuming nonstationary statistics for the motion field and for the errors. Problems addressed include the preservation of discontinuities, model/data errors, outliers, confidence measures, and performance evaluation. In solving these problems, we assume that the statistics of the motion field and the errors are not only spatially varying, but also unknown. We, thus, derive a blind adaptive technique based on generalized cross validation for estimating an independent regularization parameter for each pixel. Our formulation is pixelwise and combines existing first- and second-order constraints with a new second-order temporal constraint. We derive a new confidence measure for an adaptive rejection of erroneous and outlying motion vectors, and compare our results to other techniques in the literature. A new performance measure is also derived for estimating the signal-to-noise ratio for real sequences when the ground truth is unknown.
Hassan Foroosh
IEEE Trans. Image Process.1
2005 Single view compositing with shadows
Xiaochun Cao, Yuping Shen, Mubarak Shah, Hassan Foroosh
Vis. Comput.4
2004 Metrology from Vertical Objects
abstract
In this paper, we describe how 3D Euclidean measurements can be made in a pair of perspective images, when only minimal geometric information are available in the image planes. This minimal information consists of one line on a reference plane and one vanishing point for a direction perpendicular to the plane. Given these information, we show that the length ratio of two objects perpendicular to the reference plane can be expressed as a function of the camera principal point. Assuming that the camera intrinsic parameters remain invariant between the two views, we recover the principal point and the camera focal length by minimizing the symmetric transfer error of geometric distances. Euclidean metric measurements can then be made directly from the images. To demonstrate the effectiveness of the approach, we present the processing results for synthetic and natural images, including measurements along both parallel and non-parallel lines. 1
Xiaochun Cao, Hassan Foroosh
BMVC2
2004 Camera calibration without metric information using 1D objects
abstract
This paper addresses the problem of calibrating a pin-hole camera from images of 1D objects. Assuming a unit aspect ratio and zero skew, we introduce a novel and simple approach that uses four observations of a 1D object and requires no information about the distances between the points on the object. This is in contrast to existing methods that use two images, but impose more restrictive configurations that require measured distances on the calibrating object. The key features of the proposed technique are its simplicity and ease of use due to the lack of need for any metric information. To demonstrate the effectiveness of the algorithm, we present the processing results on synthetic and real images.
Xiaochun Cao, Hassan Foroosh
ICIP2
2004 An adaptive scheme for estimating motion
Hassan Foroosh
ICIP1
2004 Sub-pixel registration and estimation of local shifts directly in the fourier domain
Hassan Foroosh, Murat Balci
ICIP1
2004 Expression morphing from distant viewpoints
abstract
In this paper, we propose an image-based approach to photo-realistic view synthesis by integrating field morphing and view morphing in a single framework. We thus provide a unified technique for synthesizing new images that include both viewpoint changes and object deformations. For view morphing, we relax the requirement of monotonicity along epipolar lines to piecewise monotonicity, by incorporating a segmentation stage prior to interpolation. This allows for dealing with occlusions and visibility issues, and hence alleviates the "ghosting effects" that typically occur when morphing is performed between distant viewpoints. We have particularly applied our approach to the synthesis of human facial expressions, while allowing for wide change of viewing positions and directions.
Hassan Foroosh
ICIP2
2002 Extension of phase correlation to subpixel registration
abstract
In this paper, we have derived analytic expressions for the phase correlation of downsampled images. We have shown that for downsampled images the signal power in the phase correlation is not concentrated in a single peak, but rather in several coherent peaks mostly adjacent to each other. These coherent peaks correspond to the polyphase transform of a filtered unit impulse centered at the point of registration. The analytic results provide a closed-form solution to subpixel translation estimation, and are used for detailed error analysis. Excellent results have been obtained for subpixel translation estimation of images of different nature and across different spectral bands.
Hassan Foroosh, Josiane Zerubia, Marc Berthod
IEEE Trans. Image Process.1
2001 A closed-form solution for optical flow by imposing temporal constraints
abstract
We present a well-constrained formulation for estimation of optical flow in a video sequence, which yields a closed-form solution. Unlike other existing closed-form approaches, our solution does not require second order spatial derivatives. It relies only on first order and cross-space-time derivatives which can be computed more reliably. Instead of the commonly used spatial smoothness constraint, we propose a temporal smoothness constraint which has a clear physical interpretation. Our formulation also allows for accurate analysis of sources of error and ill-conditioning, leading to a more tractable implementation and mathematically justifiable thresholding values.
Hassan Foroosh
ICIP (3)1
2000 A Multi-Fractal Formalism for Stabilization, Object Detection and Tracking in Flir Sequences
abstract
In this paper, we investigate the problem of stabilization, and detection and tracking of moving or stationary objects in a forward-looking infrared (FLIR) sequence. A multifractal formalism is proposed for stabilization and activity detection and the method has been compared to three other classical techniques in image processing. 1.
Hassan Foroosh, Rama Chellappa
ICIP1
1998 Denoising by extracting fractional order singularities
abstract
In this paper we introduce a method of isolating and extracting a certain class of local singular behaviours of signals/images which in turn leads to a method of pointwise noise estimation and suppression. The underlying motivation is to decompose functions directly in terms of components which would naturally represent different orders of regular or singular behaviours defined by the local Holder exponents. We have shown that such a decomposition can lead to a factorization of the spectrum of the singular portion of the signal in terms of the spectrum of the original signal and that of a denoising filter.
Hassan Foroosh, Josiane Zerubia, Marc Berthod
ICASSP1
1998 Blind Estimation of PSF for Out of Focus Video Data
Hassan Foroosh, Rama Chellappa
ICIP (3)1
1997 Super-Resolution with Adaptive Regularization
abstract
Multi-channel super-resolution is a means of recovering high frequency information by trading off the temporal bandwidth. Almost all the methods proposed in the literature are based on optimizing a cost function. But since the problem is usually ill-posed, one needs to impose some regularity constraints. However, regularity constraints tend to attenuate the high frequency contents of the data (usually present in the form of discontinuities). This inherent contradiction between regularization and super-resolution has not been addressed in the literature, despite the availability of off the shelf tools. W have investigated this issue in the context of adaptive regularization, using /spl phi/-functions (convex, non-convex, bounded, unbounded).
Anne Lorette, Hassan Foroosh, Josiane Zerubia
ICIP (1)2
1996 Subpixel Image Registration by Estimating the Polyphase Decomposition of Cross Power Spectrum
abstract
A method of registering images at subpixel accuracy has been proposed, which does not resort to interpolation. The method is based on the phase correlation method and is remarkably robust to correlated noise and uniform variations of luminance. We have shown that the cross power spectrum of two images, containing subpixel shifts, is a polyphase decomposition of a Dirac delta function. By estimating the sum of polyphase components one can then determine sub-pixel shifts along each axis.
Hassan Foroosh, Marc Berthod, Josiane Zerubia
CVPR1
1996 Sub-pixel Bayesian estimation of albedo and height
Hassan Foroosh, Marc Berthod, Josiane Zerubia, Michael Werman
Int. J. Comput. Vis.1
1995 Sub-pixel Reconstruction of a Variable Albedo Lambertian Surface
abstract
Using a probabilistic interpretation of an n dimensional extension of Papoulis's Generalized Sampling Theorem, an iterative algorithm has been devised for 3D reconstruction of a Lambertian surface at subpixel accuracy. The problem has been formulated as an optimization one in a Bayesian framework. The latter allows for introducing a priori information on the solution, using Markov Random Fields (MRF). The estimated 3D features of the surface are the albedo and the height which are obtained simultaneously using a set of low resolution images. keywords: 3D Super resolution, Generalized Sampling Expansion, Low level image processing, Markov Random Fields (MRF).
Hassan Foroosh, Marc Berthod, Josiane Zerubia
BMVC1
1995 3D super-resolution using generalized sampling expansion
abstract
Using a set of low resolution images it is possible to reconstruct high resolution information by merging low resolution data on a finer grid. A 3D super-resolution algorithm is proposed, based on a probabilistic interpretation of the n-dimensional version of Papoulis' (1977) generalized sampling theorem. The algorithm is devised for recovering the albedo and the height map of a Lambertian surface in a Bayesian framework, using Markov random fields for modeling the a priori knowledge.
Hassan Foroosh, Marc Berthod, Josiane Zerubia
ICIP1
1994 Reconstruction of high resolution 3D visual information
abstract
Given a set of low resolution camera images, it is possible to reconstruct high resolution luminance and depth information, specially if the relative displacements of the image frames are known. We propose iterative algorithms for recovering hash resolution albedo and depth maps that require no a priori knowledge of the scene, and therefore do not depend on other methods, as regards boundary and initial conditions. The problem of surface reconstruction has been formulated as one of expectation maximization (EM) and has been tackled in a probabilistic framework using Markov random fields (MRF). As for the depth map, our method directly recovers surface heights without refering to surface orientations, while increasing the resolution by camera jittering. Conventional statistical models have been coupled with geometrical techniques to construct a general model of the world and the imaging process.>
Marc Berthod, Hassan Foroosh, Michael Werman, Josiane Zerubia
CVPR2