Hamid K. Aghajan

dblp:23/1305 · also Hamid Aghajan 0001 · DBLP profile ↗
← Back
48ranked-venue papers
10as first author
1since 2021 · last 2021
0000-0001-9054-5642ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 11 · 2 first-authorComputer networks · 6Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 38% Deep learning architectures and training · 31% Video understanding and tracking · 16%
Computer graphics and multimedia
7 papers
Multimedia analysis and retrieval · 63% Image and video processing · 28% Multimedia systems and quality of experience · 9%
Human-computer interaction and pervasive computing
3 papers
Ubiquitous computing and smart environments · 51% Haptics and multimodal interaction · 49%
Computer networks
3 papers
Internet of things and sensor networks · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Embedded and real-time systems · 74% Distributed systems · 26%

Topics — the 21 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
convolutional neural network
0.412019
Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks · NeurIPS 2019
Computer vision › 3D vision › biological vision modeling
visual cortex modeling
0.412019
Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks · NeurIPS 2019
Multimedia analysis and retrieval
video analysis
0.122010
Pervasive video analysis: workshop overview · ACM Multimedia 2010
Estimation of multiple 2-D uniform motions by SLIDE: subspace-based line detection · IEEE Trans. Image Process. 1999
Machine learning › Representation and self-supervised learning › computational neuroscience
neural coding
0.112019
Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks · NeurIPS 2019
Computer vision › Video understanding and tracking
multi-camera tracking
0.112010
Pervasive video analysis: workshop overview · ACM Multimedia 2010
Haptics and multimodal interaction
multimodal interaction
0.112010
Human-centered multimedia systems: tutorial overview · ACM Multimedia 2010
Computer vision › Video understanding and tracking › video analytics
behavior analysis
0.112008
ACM multimedia 2008: 1st workshop on vision networks for behavior analysis (VNBA 2008) · ACM Multimedia 2008
Computer vision › Face, body and person analysis
human pose estimation
0.112008
Real-Time Human Pose Estimation: A Case Study in Algorithm Design for Smart Camera Networks · Proc. IEEE 2008
Computer vision › 3D vision › pose estimation
multi-view pose estimation
0.112008
Real-Time Human Pose Estimation: A Case Study in Algorithm Design for Smart Camera Networks · Proc. IEEE 2008
Internet of things and sensor networks › camera sensor networks › camera networks
smart camera networks
0.112008
Real-Time Human Posture Reconstruction in Wireless Smart Camera Networks · IPSN 2008
Internet of things and sensor networks
wireless sensor network
0.112008
Real-Time Human Posture Reconstruction in Wireless Smart Camera Networks · IPSN 2008
Multimedia analysis and retrieval
image indexing and retrieval
0.012010
Human-centered multimedia systems: tutorial overview · ACM Multimedia 2010
Ubiquitous computing and smart environments › pervasive computing infrastructure
distributed camera networks
0.012010
Pervasive video analysis: workshop overview · ACM Multimedia 2010
Multimedia analysis and retrieval
video surveillance
0.012008
ACM multimedia 2008: 1st workshop on vision networks for behavior analysis (VNBA 2008) · ACM Multimedia 2008
Internet of things and sensor networks
camera sensor networks
0.012008
ACM multimedia 2008: 1st workshop on vision networks for behavior analysis (VNBA 2008) · ACM Multimedia 2008
Distributed systems › distributed data processing
distributed video processing
0.012008
Real-Time Human Pose Estimation: A Case Study in Algorithm Design for Smart Camera Networks · Proc. IEEE 2008
Image and video processing
motion estimation
0.011999
Estimation of multiple 2-D uniform motions by SLIDE: subspace-based line detection · IEEE Trans. Image Process. 1999
Image and video processing › motion estimation
multiple motion estimation
0.011999
Estimation of multiple 2-D uniform motions by SLIDE: subspace-based line detection · IEEE Trans. Image Process. 1999
Image and video processing › edge detection
line detection
0.021994
SLIDE: Subspace-Based Line Detection · IEEE Trans. Pattern Anal. Mach. Intell. 1994
Sensor array processing techniques for super resolution multi-line-fitting and straight edge detection · IEEE Trans. Image Process. 1993
Image and video processing
edge detection
0.011993
Sensor array processing techniques for super resolution multi-line-fitting and straight edge detection · IEEE Trans. Image Process. 1993
Image and video processing › edge detection
straight edge detection
0.011993
Sensor array processing techniques for super resolution multi-line-fitting and straight edge detection · IEEE Trans. Image Process. 1993

Methods — techniques the papers use, named apart from their topics

surround modulation · 0.4excitatory-inhibitory connections · 0.4distributed observation fusion · 0.3multi-view fusion · 0.2data fusion · 0.2machine learning · 0.2context modeling · 0.2SIMD processing · 0.23d pose reconstruction · 0.2tracking · 0.1stereo vision · 0.1object detection · 0.1subspace-based line detection · 0.0multiline fitting · 0.0subspace estimation · 0.0hough transform · 0.0direction-of-arrival estimation · 0.0
YearPublicationVenuePosition
2021 Guest Editorial Introduction to the Special Issue on Large-Scale Visual Sensor Networks: Architectures and Applications
abstract
Large–scale visual sensor networks have become progressively an essential part of our daily lives underpinning many technological, financial, and social advancements today, with applications in smart cities, traffic monitoring, environmental pollution control, public safety, and crime prevention.
Paolo Spagnolo, Hamid K. Aghajan, George Bebis, Shaogang Gong, Amy Loutfi, Leonid Sigal, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2019 Surround Modulation: A Bio-inspired Connectivity Structure for Convolutional Neural Networks
abstract
Numerous neurophysiological studies have revealed that a large number of the primary visual cortex neurons operate in a regime called surround modulation. Surround modulation has a substantial effect on various perceptual tasks, and it also plays a crucial role in the efficient neural coding of the visual cortex. Inspired by the notion of surround modulation, we designed new excitatory-inhibitory connections between a unit and its surrounding units in the convolutional neural network (CNN) to achieve a more biologically plausible network. Our experiments show that this simple mechanism can considerably improve both the performance and training speed of traditional CNNs in visual tasks. We further explore additional outcomes of the proposed structure. We first evaluate the model under several visual challenges, such as the presence of clutter or change in lighting conditions and show its superior generalization capability in handling these challenging situations. We then study possible changes in the statistics of neural activities such as sparsity and decorrelation and provide further insight into the underlying efficiencies of surround modulation. Experimental results show that importing surround modulation into the convolutional layers ensues various effects analogous to those derived by surround modulation in the visual cortex.
Hosein Hasani, Mahdieh Soleymani Baghshah, Hamid K. Aghajan
NeurIPS3
2017 Election Vote Share Prediction using a Sentiment-based Fusion of Twitter Data with Google Trends and Online Polls
Parnian Kassraie, Alireza Modirshanechi, Hamid K. Aghajan
DATA3
2017 Structured prediction with short/long-range dependencies for human activity recognition from depth skeleton data
abstract
One of the main abilities that the robots need to maintain is to efficiently communicate with people in a humanly manner. Thus, human activity recognition (HAR) would be an integral part of such a human-robot interaction system. One of the major challenges in HAR is that the individuals perform their activities in different manners. Furthermore, there is a very wide range of different types of activities that the robots would require to understand. Some activities are simple, quick and short (e.g., sit down), while many others are complex, have many details and span through a long range of time (e.g., wearing contact lens). In this paper, we model the recognition of activities into a sequence-labeling problem and propose a new probabilistic graphical model (PGM) that can recognize both short/long-range activities, by introducing a hierarchical classification model and including extra links and loopy conditions in our PGM. To optimize the PGM and obtain its parameters during training, we use a structured prediction technique, a general framework that involves latent structured support vector machines (LSSVM) and hidden-state conditional random fields (HCRF). We evaluate our method on two widely used datasets (CAD-60 & UT-Kinect) that contain both activity types. Our obtained results are promising and show that our method can recognize both types of activities effectively, while most of the previous works only focused on one of these two major types. We further explore distributed processing techniques, since our method can easily be distributed over processing nodes. We also propose an efficient divide-and-merge technique to further speedup the training step.
Mohammad M. Arzani, Mahmood Fathy, Hamid K. Aghajan, A. Akbariazirani, Kaamran Raahemifar, Ehsan Adeli-Mosabbeb
IROS3
2016 A novel approach for detecting intersections from GPS traces
abstract
Intersection detection is a critical aspect for both route planning and path optimization. In literature, intersections are detected indirectly using the road users' turning behaviors. This paper proposes a novel approach to detect intersections directly using their definition of connecting road segments. We first detect the Longest Common Sub-Sequences (LCSS) between each pair of GPS traces using dynamic programming approach. Second, we partition the longest nonconsecutive subsequences into consecutive substrings. The starting and ending points of each common substring are connecting points where two GPS traces split to different directions after they share a series of common locations. At last, we estimate Kernel Density (KD) of the connecting points and find the local maximas on the density map as intersections. Experimental results show our proposed method outperforms the state-of-the-art work with a high accuracy for intersection detection.
Xingzhe Xie, Wenzi Liao, Hamid K. Aghajan, Peter Veelaert, Wilfried Philips
IGARSS3
2014 A low resolution multi-camera system for person tracking
abstract
The current multi-camera systems have not studied the problem of person tracking under low resolution constraints. In this paper, we propose a low resolution sensor network for person tracking. The network is composed of cameras with a resolution of 30×30 pixels. The multi-camera system is used to evaluate probability occupancy mapping and maximum likelihood trackers against ground truth collected by ultra-wideband (UWB) testbed. Performance evaluation is performed on two video sequences of 30 minutes. The experimental results show that maximum likelihood estimation based tracker outperforms the state-of-the-art on low resolution cameras.
Mohamed Y. Eldib, Nyan Bo Bo, Francis Deboeverie, Jorge Oswaldo Niño Castañeda, Junzhi Guan, Samuel Van de Velde, Heidi Steendam, Hamid K. Aghajan, Wilfried Philips
ICIP8
2014 Learning routines over long-term sensor data using topic models
abstract
Abstract Recent advances on sensor network technology provide the infrastructure to create intelligent environments on physical places. One of the main issues of sensor networks is the large amount of data they generate. Therefore, it is necessary to have good data analysis techniques with the aim of learning and discovering what is happening on the monitored environment. The problem becomes even more challenging if this process is performed following an unsupervised way (without having any a priori information) and applied over a long‐term timeline with many sensors. In this work, topic models are employed to learn the latent structure and dynamics of sensor network data. Experimental results using two realistic datasets, having over 50 weeks of data, have shown the ability to find routines of activity over sensor network data in office environments.
Federico Castanedo, Diego López-de-Ipiña, Hamid K. Aghajan, Richard P. Kleihorst
Expert Syst. J. Knowl. Eng.3
2014 Camera selection for tracking in distributed smart camera networks
abstract
Tracking persons with multiple cameras with overlapping fields of view instead of with one camera leads to more robust decisions. However, operating multiple cameras instead of one requires more processing power and communication bandwidth, which are limited resources in practical networks. When the fields of view of different cameras overlap, not all cameras are equally needed for localizing a tracking target. When only a selected set of cameras do processing and transmit data to track the target, a substantial saving of resources is achieved. The recent introduction of smart cameras with on-board image processing and communication hardware makes such a distributed implementation of tracking feasible. We present a novel framework for selecting cameras to track people in a distributed smart camera network that is based on generalized information-theory. By quantifying the contribution of one or more cameras to the tracking task, the limited network resources can be allocated appropriately, such that the best possible tracking performance is achieved. With the proposed method, we dynamically assign a subset of all available cameras to each target and track it in difficult circumstances of occlusions and limited fields of view with the same accuracy as when using all cameras.
Linda Tessens, Marleen Morbée, Hamid K. Aghajan, Wilfried Philips
ACM Trans. Sens. Networks3
2013 Riemannian manifold-based support vector machine for human activity classification in images
abstract
This paper addresses the issue of classification of human activities in still images. We propose a novel method where part-based features focusing on human and object interaction are utilized for activity representation, and classification is designed on manifolds by exploiting underlying Riemannian geometry. The main contributions of the paper include: (a) represent human activity by appearance features from image patches containing hands, and by structural features formed from the distances between the torso and patch centers; (b) formulate SVM kernel function based on the geodesics on Riemannian manifolds under the log-Euclidean metric; (c) apply multi-class SVM classifier on the manifold under the one-against-all strategy. Experiments were conducted on a dataset containing 2750 images in 7 classes of activities from 10 subjects. Results have shown good performance (average classification rate of 95.83%, false positive rate of 0.71%). Comparisons with three other related classifiers provide further support to the proposed method.
Yixiao Yun, Irene Y. H. Gu, Hamid K. Aghajan
ICIP3
2013 Guest Editorial: Human-Computer Interaction: Real-Time Vision Aspects of Natural User Interfaces
Zoran Zivkovic, Nicu Sebe, Hamid K. Aghajan, Branislav Kisacanin
Int. J. Comput. Vis.3
2012 Integrating video and accelerometer signals for nocturnal epileptic seizure detection
abstract
Epileptic seizure detection is traditionally done using video/electroencephalogram (EEG) monitoring, which is not applicable in a home situation. In recent years, attempts have been made to detect the seizures using other modalities. In this paper we investigate if a combined usage of accelerometers attached to the limbs and video data would increase the performance compared to a single modality approach. Therefore, we used two existing approaches for seizure detection in accelerometers and video and combined them using a linear discriminant analysis (LDA) classifier. The results for a combined detection have a better positive predictive value (PPV) of 95.00% compared to the single modality detection and reached a sensitivity of 83.33%.
Kris Cuppens, Chih-Wei Chen, Kevin Bing-Yung Wong, Anouk Van de Vel, Lieven Lagae, Berten Ceulemans, Tinne Tuytelaars, Sabine Van Huffel, Bart Vanrumste, Hamid K. Aghajan
ICMI10
2011 Discovering social interactions in real work environments
abstract
The goal of this work is to detect pairwise primitive interactions in groups for social interaction analysis in real environments. We propose a system that extracts locations and head poses of people from videos captured in an unconstrained environment, namely a research lab. Our system is designed to work with realistic data capturing natural human interactions. An efficient tracking method based on Chamfer matching finds the head and shoulder silhouettes of people in real-time, and a head orientation classifier estimates their head poses. The location, relative distance and head orientation of people capture the use of space by individuals and their interactive behavioral patterns which are inferred with a probabilistic model. We present quantitative evaluation and experimental results of our system, demonstrating the effectiveness of our proposed approach on challenging real-world data.
Chih-Wei Chen, Rodrigo Cilla, Chen Wu 0002, Hamid K. Aghajan
FG4
2011 Real-time social interaction analysis
abstract
This demonstration presents a social interaction analysis system designed to operate in real-time and under real environment conditions. A webcam is used to capture videos of a group of people interacting in an unconstrained environment. Locations of multiple people and their head poses are extracted from the videos. Direct pairwise interactions are detected based on the relative distance and head orientation using a probabilistic model. Communicative behavior analysis is carried out by grouping people based on their interaction patterns, i.e. frequency and duration, among the social group.
Chih-Wei Chen, Chen Wu 0002, Hamid K. Aghajan
FG3
2011 Vision-based user-centric light control for smart environments
Huang Lee, Chen Wu 0002, Hamid K. Aghajan
Pervasive Mob. Comput.3
2011 User-Centric Environment Discovery With Camera Networks in Smart Homes
abstract
We propose a data integration and reasoning technique for automatic environment discovery in smart homes based on observations of user interactions with objects. This approach is complementary to traditional appearance-based object recognition, which often demands large training sets. In our approach, object recognition is achieved in a semantic way by linking object behaviors to the pose and activity of the person using them. The complex relations between objects and user activities are modeled with a Markov logic network. The embodiments of the proposed approach in two multicamera smart environments are described, and the experimental results are presented.
Chen Wu 0002, Hamid K. Aghajan
IEEE Trans. Syst. Man Cybern. Part A2
2010 Recognizing Objects in Smart Homes Based on Human Interaction
Chen Wu 0002, Hamid K. Aghajan
ACIVS (2)2
2010 Pervasive video analysis: workshop overview
abstract
This workshop aims at tackling the novel challenging scenarios in pervasive video analysis which require not only to address specific problems (e.g., tracking, recognition) on a single view, but to deal with a set of distributed observations, eventually integrated with subjective mobile video streams. Accepted papers cover a wide range of subjects going from the joint analysis of video sequences, taken from fixed location and mobile cameras, to situation awareness and understanding.
Hamid K. Aghajan, Marco Cristani, Vittorio Murino, Nicu Sebe
ACM Multimedia1
2010 Human-centered multimedia systems: tutorial overview
abstract
This tutorial will focus on technical analysis and interaction techniques formulated from the perspective of key human factors in a user-centered approach to developing multimedia systems. The tutorial will take a holistic view on the research issues and applications of Human-Centered Systems, focusing on four main areas: (1) multimodal interaction: visual (body, gaze, gesture); (2) image indexing and retrieval: user behavior, context modeling, cultural issues, and machine learning for user-centric approaches; (3) multimedia data: conceptual analysis at different levels (feature, cognitive, and affective); and (4) sources of contextual information and case studies in multi-camera networks.
Nicu Sebe, Alejandro Jaimes, Hamid K. Aghajan
ACM Multimedia3
2010 Special issue on multi-camera and multi-modal sensor fusion
Andrea Cavallaro, Hamid K. Aghajan
Comput. Vis. Image Underst.2
2010 Special Issue on Multimodal Affective Interaction
abstract
The 11 papers in this special issue can be categorized into five groups: Emotional speech synthesis and recognition; affective video content analysis; facial expressions and head movements; affect analysis in small groups; and audio-visual affective corpus.
Nicu Sebe, Hamid K. Aghajan, Thomas S. Huang, Nadia Magnenat-Thalmann, Caifeng Shan
IEEE Trans. Multim.2
2010 Near-lifetime-optimal data collection in wireless sensor networks via spatio-temporal load balancing
abstract
In wireless sensor networks, periodic data collection appears in many applications. During data collection, messages from sensor nodes are periodically collected and sent back to a set of base stations for processing. In this article, we present and analyze a near-lifetime-optimal and scalable solution for data collection in stationary wireless sensor networks and an energy-efficient packet exchange mechanism. In our solution, instead of using a fixed network topology, we construct a set of communication topologies and apply each topology to different data collection cycles. We not only use the flexibility in distributing the traffic load across different routes in the network (spatial load balancing), but also balance the energy consumption in the time domain (temporal load balancing). We show that this method achieves an average energy consumption rate very close to the optimal value found by network flow optimization techniques. To increase the scalability, we further extend our solution such that it can be applied to networks with multiple base stations where each base station only stores part of the network configuration, cooperating with each other to find a global solution in a distributed manner. The proposed methods are analyzed and evaluated by simulations.
Huang Lee, Abtin Keshavarzian, Hamid K. Aghajan
ACM Trans. Sens. Networks3
2008 Sub-optimal Camera Selection in Practical Vision Networks through Shape Approximation
Huang Lee, Linda Tessens, Marleen Morbée, Hamid K. Aghajan, Wilfried Philips
ACIVS4
2008 Human Pose Estimation in Vision Networks Via Distributed Local Processing and Nonparametric Belief Propagation
Chen Wu 0002, Hamid K. Aghajan
ACIVS2
2008 Multi-Cluster Multi-Parent Wake-Up Scheduling in Delay-Sensitive Wireless Sensor Networks
abstract
Immediate notification of urgent but rare events and delivery of time sensitive actuation commands appear in many practical wireless sensor and actuator network applications. Multi-parent wake-up scheduling was presented as a technique which can provide bi-directional end-to-end latency guarantees while optimizing the node battery lifetime. This method takes a cross-layer approach where multiple routes for transfer of messages and wake-up schedules for nodes are crafted in synergy to reduce overall message latencies. In this paper, we generalize the multi-parent method to support a multi-cluster model for the network where we assume that the network has multiple central points called cluster-head (CH) that are in charge of scheduling the nodes in the network. A key step in multi-parent method is to divide the nodes in network into disjoint groups such that each node has at least one link to a node in each group. We formulate this step as a graph coloring problem which is shown to be NP-complete. We propose an algorithm where all the cluster-heads cooperate to find a heuristic solution for the graph coloring optimization problem in a distributed manner. We show that each cluster-head requires less memory and computational power compared to the case where one cluster-head finds the global solution, therefore the solution is very scalable.
Huang Lee, Abtin Keshavarzian, Hamid K. Aghajan
GLOBECOM3
2008 Real-Time Human Posture Reconstruction in Wireless Smart Camera Networks
abstract
While providing a variety of intriguing application opportunities, a vision sensor network poses three key challenges. High computation capacity is required for early vision functions to enable real-time performance. Wireless links limit image transmission in the network due to both bandwidth and energy concerns. Last but not least, there is a lack of established vision-based fusion mechanisms when a network of cameras is available. In this paper a distributed vision processing implementation of human pose interpretation on a wireless smart camera network is presented. The motivation for employing distributed processing is to both achieve real-time vision and provide scalability for developing more complex vision algorithms. The distributed processing operation includes two levels. One is that each smart camera processes its local vision data, achieving spatial parallelism. The other is that different functionalities of the whole line of vision processing are assigned to early vision and object-level processors, achieving functional parallelism based on the processor capabilities. Aiming for low power consumption and high image processing performance, the wireless smart camera is based on an SIMD (single-instruction multiple-data) video analysis processor, an 8051 micro-controller as the local host, and wireless communication through the IEEE 802.15.4 standard. The vision algorithm implements 3D human pose reconstruction. From the live image data from the sensor the smart camera acquires critical joints of the subject in the scene through local processing. The results obtained by multiple smart cameras are then transmitted through the wireless channel to a central PC where the 3D pose is recovered and demonstrated in a virtual reality gaming application. The system operates in real time with a 30 frames/sec rate.
Chen Wu 0002, Hamid K. Aghajan, Richard P. Kleihorst
IPSN2
2008 ACM multimedia 2008: 1st workshop on vision networks for behavior analysis (VNBA 2008)
abstract
The VNBA workshop marks a new era of the successful series of the Video Surveillance and Sensor Networks (VSSN) workshops, held until 2006 in conjunction with the ACM Multimedia conference. This new version of the workshop inherits from VSSN the experienced Technical Program Committee as well as the interests of its community, but shifts the focus to cover higher level topics and applications under the common framework of "behaviour analysis", hence aiming to adapt to the evolved directions of interest in the field, and reaching out to other research communities with overlapping interests.
Hamid K. Aghajan, Andrea Prati 0001
ACM Multimedia1
2008 Optimal camera selection in vision networks for shape approximation
abstract
Within a camera network, the contribution of a camera to the observation of a scene depends on its viewpoint and on the scene configuration. This is a dynamic property, as the scene content is subject to change over time. An automatic selection of a subset of cameras that significantly contributes to the desired observation of a scene can be of great value for the reduction of the amount of transmitted or stored image data. In this work, we propose low data rate schemes to select from a vision network a subset of cameras that provides a good frontal observation of the persons in the scene and allows for the best approximation of their 3D shape. We also investigate to what degree low data rates trade off quality of reconstructed 3D shapes.
Marleen Morbée, Linda Tessens, Huang Lee, Wilfried Philips, Hamid K. Aghajan
MMSP5
2008 Real-Time Human Pose Estimation: A Case Study in Algorithm Design for Smart Camera Networks
abstract
Monitoring human activities finds novel applications in smart environment settings. Examples include immersive multimedia and virtual reality, smart buildings and occupancy-based services, assisted living and patient monitoring, and interactive classrooms and teleconferencing. A network of cameras can enable detection and interpretation of human events by utilizing multiple views and collaborative processing. Distributed processing of acquired videos at the source camera facilitates operation of scalable vision networks by avoiding transfer of raw images. This allows for efficient collaboration between the cameras under the communication and latency constraints, as well as being motivated by aiming to preserve the privacy of the network users (no image transfer out of the camera) while offering services in applications such as assisted living or virtual placement. In this paper, collaborative processing and data fusion techniques in a multicamera setting are examined in the context of human pose estimation. Multiple mechanisms for information fusion across the space (multiple views), time, and different feature levels are introduced to meet system constraints and are described through examples.
Chen Wu 0002, Hamid K. Aghajan
Proc. IEEE2
2007 Spatiotemporal Fusion Framework for Multi-camera Face Orientation Analysis
Chung-Ching Chang, Hamid K. Aghajan
ACIVS2
2007 A Multi-touch Surface Using Multiple Cameras
Itai Katz, Kevin Gabayan, Hamid K. Aghajan
ACIVS3
2007 Model-Based Image Segmentation for Multi-view Human Gesture Analysis
Chen Wu 0002, Hamid K. Aghajan
ACIVS2
2007 A LQR spatiotemporal fusion technique for face profile collection in smart camera surveillance
abstract
In this paper, we propose a joint face orientation estimation technique for face profile collection in smart camera networks. The system is composed of in-node coarse estimation and joint refined estimation between cameras. Innode signal processing algorithms are designed to be lightweight to reduce computation load, yielding coarse estimates which may be erroneous. The proposed model-based technique determines the orientation and the angular motion of the face using two features, namely the hair-face ratio and the head optical flow. These features yield an estimate of the face orientation and the angular velocity through Least Squares (LS) analysis. In the joint refined estimation step, a discrete-time linear dynamical model is defined. Spatiotemporal consistency between cameras is measured by a cost function, which is minimized through Linear Quadratic Regulation (LQR) to yield a robust closed-loop feedback system that estimates the face orientation, angular motion, and relative angular difference to the face between cameras. Based on the face orientation estimates, a collection of face profile are accumulated over time as the human subject moves around. The proposed technique does not require camera locations to be known in prior, and hence is applicable to vision networks deployed casually without localization.
Chung-Ching Chang, Hamid K. Aghajan
AVSS2
2007 Model-based human posture estimation for gesture analysis in an opportunistic fusion smart camera network
abstract
In multi-camera networks rich visual data is provided both spatially and temporally. In this paper a method of human posture estimation is described incorporating the concept of an opportunistic fusion framework aiming to employ manifold sources of visual information across space, time, and feature levels. One motivation for the proposed method is to reduce raw visual data in a single camera to elliptical parameterized segments for efficient communication between cameras. A 3D human body model is employed as the convergence point of spatiotemporal and feature fusion. It maintains both geometric parameters of the human posture and the adaptively learned appearance attributes, all of which are updated from the three dimensions of space, time and features of the opportunistic fusion. In sufficient confidence levels parameters of the 3D human body model are again used as feedback to aid subsequent in-node vision analysis. Color distribution registered in the model is used to initialize segmentation. Perceptually Organized Expectation Maximization (POEM) is then applied to refine color segments with observations from a single camera. Geometric configuration of the 3D skeleton is estimated by Particle Swarm Optimization (PSO).
Chen Wu 0002, Hamid K. Aghajan
AVSS2
2007 Layered and Collaborative Gesture Analysis in Multi-Camera Networks
abstract
A layered and collaborative architecture for gesture recognition in a multi-camera network is presented in this paper. The proposed approach is motivated by the diversity of gestures expressed in passive monitoring applications. It is based on the concept of opportunistic fusion of simple features within a single camera and active collaboration between multiple cameras in the decision making process. The decision process is pursued through mutual, assisted, and self correspondences, using features available at different cameras. The dynamics employed by the opportunistic fusion of different features within a single camera as well as those from multiple cameras offer the potential to address gesture recognition problems more efficiently and accurately across a variety of different applications.
Hamid K. Aghajan, Chen Wu 0002
ICASSP (4)1
2007 Distributed Vision-Based Accident Management for Assisted Living
Hamid K. Aghajan, Juan Carlos Augusto, Chen Wu 0002, Paul J. McCullagh, Julie-Ann Augusto-Walkden
ICOST1
2007 MeshEye: a hybrid-resolution smart camera mote for applications in distributed intelligent surveillance
abstract
Surveillance is one of the promising applications to which smart camera motes forming a vision-enabled network can add increasing levels of intelligence. We see a high degree of in-node processing in combination with distributed reasoning algorithms as the key enablers for such intelligent surveillance systems. To put these systems into practice still requires a considerable amount of research ranging from mote architectures, pixel-processing algorithms, up to distributed reasoning engines. This paper introduces MeshEye, an energy-efficient smart camera mote architecture that has been designed with intelligent surveillance as the target application in mind. Special attention is given to MeshEye's unique vision system: a low-resolution stereo vision system continuously determines position, range, and size of moving objects entering its field of view. This information triggers a color camera module to acquire a high-resolution image sub-array containing the object, which can be efficiently processed in subsequent stages. It offers reduced complexity, response time, and power consumption over conventional solutions. Basic vision algorithms for object detection, acquisition, and tracking are described and illustrated on real-world data. The paper also presents a basic power model that estimates lifetime of our smart camera mote in battery-powered operation for intelligent surveillance event processing.
Stephan Hengstler, Daniel Prashanth, Sufen Fong, Hamid K. Aghajan
IPSN4
2006 Color-Based Multiple Agent Tracking for Wireless Image Sensor Networks
Emre Oto, Frances Lau, Hamid K. Aghajan
ACIVS3
2006 Subspace Techniques for Vision-Based Node Localization in Wireless Sensor Networks
abstract
We present novel techniques for localization of nodes in a wireless image sensor network. Based on visual observations of a moving object by the network nodes, the proposed techniques employ simple image processing functions to produce equations that contain the node positions and orientation angles as the unknown parameters. Observations made at the nodes relate the position of the observed object to the physical coordinates of the node via the mapped position of the object in the node's image plane. In one formulation of the problem, multiple observations by a network node from a moving beacon with known coordinates result in a system of equations with a rank-deficient matrix. Hence, the solution for the desired node coordinates lies in the null space of the data matrix. In a second formulation, a different configuration of image sensor deployment with more degrees of freedom results in a least-squares solution for the unknown parameters. In a third formulation, multiple observations are made at each node from a target which moves at a fixed velocity vector. The solution to this problem formulation is also shown to correspond to the null space of the data matrix. The proposed algorithms are based on in-node processing and hence are scalable to large networks. Simulation and experimental results are provided in the paper.
Huang Lee, Laura Savidge, Hamid K. Aghajan
ICASSP (4)3
2006 Wireless symbolic positioning using support vector machines
abstract
This paper introduces a novel symbolic positioning system based on wireless access points and Support Vector Machines. The system works both indoors and outdoors and is cost-effective since it can even work with widely deployed 802.11 access points as infrastructure. The system requires minimal setup time, which makes it readily available for real-world applications.
C. Philipp Schloter, Hamid K. Aghajan
IWCMC2
2006 Robot-Assisted Localization Techniques for Wireless Image Sensor Networks
abstract
We present a vision-based solution to the problem of topology discovery and localization of wireless sensor networks. In the proposed model, a robot controlled by the network is introduced to assist with localization of a network of image sensors, which are assumed to have image planes parallel to the agent's motion plane. The localization algorithm for the scenario where the moving agent has knowledge of its global coordinates is first studied. This baseline scenario is then used to build more complex localization algorithms in which the robot has no knowledge of its global positions. Two cases where the sensors have overlapping and non-overlapping fields of view (FOVs) are investigated. In order to implement the discovery algorithms for these two different cases, a forest structure is introduced to represent the topology of the network. We consider the collection of sensors with overlapping FOVs as a tree in the forest. The robot searches for nodes in each tree through boundary patrolling, while it searches for other trees by a radial pattern motion. Numerical analyses are provided to verify the proposed algorithms. Finally, experiment results show that the sensor coordinates estimated by the proposed algorithms accurately reflect the results found by manual methods
Huang Lee, Hattie Dong, Hamid K. Aghajan
SECON3
1999 Estimation of multiple 2-D uniform motions by SLIDE: subspace-based line detection
abstract
A technique is proposed for estimating the parameters of two-dimensional (2-D) uniform motion of multiple moving objects in a scene, based on long-sequence image processing and the application of a multiline fitting algorithm. Plots of the vertical and horizontal projections versus frame number give new images in which uniformly moving objects are represented by skewed band regions, with the angles of the skew from the vertical being a measure of the velocities of the moving objects. For example, vertical bands will correspond to objects with zero velocity. An algorithm called subspace-based line detection (SLIDE) can be used to efficiently determine the skew angles. SLIDE exploits the temporal coherence between the contributions of each of the moving patterns in the frame projections to enhance and distinguish a signal subspace that is defined by the desired motion parameters. A similar procedure can be used to determine the vertical velocities. Some further steps must then be taken to properly associate the horizontal and vertical velocities.
Hamid K. Aghajan, Babak Hossein Khalaj, Thomas Kailath
IEEE Trans. Image Process.1
1994 Estimation of skew angle in text-image analysis bySLIDE: Subspace-based line detection
Hamid K. Aghajan, Babak Hossein Khalaj, Thomas Kailath
Mach. Vis. Appl.1
1994 Patterned wafer inspection by high resolution spectral estimation techniques
Babak Hossein Khalaj, Hamid K. Aghajan, Thomas Kailath
Mach. Vis. Appl.2
1994 SLIDE: Subspace-Based Line Detection
abstract
An analogy is made between each straight line in an image and a planar propagating wavefront impinging on an array of sensors so as to obtain a mathematical model exploited in recent high resolution methods for direction-of-arrival estimation in sensor array processing. The new so-called SLIDE (subspace-based line detection) algorithm then exploits the spatial coherence between the contributions of each line in different rows of the image to enhance and distinguish a signal subspace that is defined by the desired line parameters. SLIDE yields closed-form and high resolution estimates for line parameters, and its computational complexity and storage requirements are far less than those of the standard method of the Hough transform. If unknown a priori, the number of lines is also estimated in the proposed technique. The signal representation employed in this formulation is also generalized to handle grey-scale images as well. The technique has also been generalized to fitting planes in 3-D images. Some practical issues of the proposed technique are given.>
Hamid K. Aghajan, Thomas Kailath
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 SLIDE: subspace-based line detection
Hamid K. Aghajan, Thomas Kailath
ICASSP (5)1
1993 Sensor array processing techniques for super resolution multi-line-fitting and straight edge detection
abstract
A signal processing method is developed for solving the problem of fitting multiple lines in a two-dimensional image. It formulates the multi-line-fitting problem in a special parameter estimation framework such that a signal structure similar to the sensor array processing signal representation is obtained. Then the recently developed algorithms in that formalism can be exploited to produce super-resolution estimates for line parameters. The number of lines may also be estimated in this framework. The signal representation used can be generalized to handle problems of line fitting and of straight edge detection. Details of the proposed algorithm and several experimental results are presented. The method exhibits considerable computational speed superiority over existing single- and multiple-line-fitting algorithms such as the Hough transform method. Potential applications include road tracking in robotic vision, mask wafer alignment in semiconductor manufacturing, aerial image analysis, text alignment in document analysis, and particle tracking in bubble chambers.
Hamid K. Aghajan, Thomas Kailath
IEEE Trans. Image Process.1
1992 A subspace fitting approach to super resolution multi-line fitting and straight edge detection
abstract
A new fundamental signal processing method is developed for solving the problem of fitting multiple lines in a two-dimensional image. The proposed technique formulates the multiline fitting problem in a special parameter estimation framework such that a signal structure similar to the sensor array processing signal representation is obtained. Then, recently developed algorithms in that formalism (e.g., the ESPRIT technique) are exploited to produce superresolution estimates in this framework. The signal representation used in this formulation can be generalized in a fashion to handle both problems of line fitting (in which a set of binary-valued discrete pixels is given) and of straight edge detection (in which one starts with a gray-scale image). The proposed method possesses extensive computational speed superiority over previous single- and multiple-line fitting algorithms such as the Hough transform method. Details of the new formulation are explained, and several experimental results are presented.>
Hamid K. Aghajan, Thomas Kailath
ICASSP1
1992 Automated direct patterned wafer inspection
abstract
A self-reference technique is developed for detecting the location of defects in repeated pattern wafers and masks. The application area of the proposed method includes inspection of memory chips, shift registers, switch capacitors, and CCD arrays. Using high resolution spectral estimation algorithms, the proposed technique first extracts the period and structure of repeated patterns from the image to sub-pixel resolution, and then produces a defect-free reference image for making comparison with the actual image. Since the technique acquires all its needed information from a single image, there is no need for a database image, a scaling procedure, or any a-priori knowledge about the repetition period of the patterns.>
Babak Hossein Khalaj, Hamid K. Aghajan, Thomas Kailath
WACV2