EDBT 2026 Demo / reviewers in the wild / expert
Mukesh Saini
dblp:83/3852 · also Mukesh Kumar Saini
· DBLP profile ↗
41ranked-venue papers
14as first author
14since 2021 · last 2026
0000-0003-2215-9365ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 14 first-author · 10 since 2021Computer networks · 4 · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpinVision: An end-to-end volleyball spin estimation with Siamese-based deep classification
Shreya Bansal, Anterpreet Kaur Bedi, Pratibha Kumari 0001, Rishi Kumar Soni, Narayanan Chatapuram Krishnan, Mukesh Saini |
Comput. Vis. Image Underst. | 6 |
| 2026 | Graph-Based Event and Sub-Event Grouping in User-Generated Videos with Distortion-Aware Keyframe ClusteringabstractUser-generated content (UGC) videos recorded in uncontrolled environments often exhibit blur, camera shake, lighting fluctuations, and large viewpoint differences, making it difficult to organize multiple recordings of the same event. This work proposes an integrated pipeline that groups UGC videos into coherent events and sub-events by combining distortion-aware keyframe selection, adaptive audio–visual fusion, and confidence-weighted graph construction. The method first filters and clusters segment-level representations to obtain reliable keyframes, then fuses audio and visual cues through a lightweight gating module to produce robust multimodal descriptors. These descriptors populate a similarity graph whose strong and weak edges reveal sub-event and event structure without requiring shot boundaries or manual segmentation. Although distortion modeling, keyframe extraction, and multimodal similarity have been studied separately, existing approaches do not integrate them for hierarchical UGC video grouping. Experiments on the JIKU dataset and a curated YouTube dataset show consistent improvements in fidelity, diversity, and clustering metrics, demonstrating the applicability of the approach to video summarization, multi-view organization, and other UGC analysis tasks. Malya Singh, Wei Tsang Ooi, Abdulmotaleb El Saddik, Mukesh Saini |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | A Step Closer Towards the Digital Twin of the PlantabstractDigital twins can provide vital insights into agricultural products and processes. There have been a lot of documented attempts at digital twins in agriculture. However, majority of these attempts build synthetic models and ignore the temporal dimension of the plant growth. Therefore, the existing models fail to depict actual plant details and growth. Our work replicates the actual growth of a real plant in the digital world by acquiring 3D meshes of the plant at various instants. It focuses on the transition between those acquired meshes by approximating all the consecutive pairs into approximate mesh pairs that have a common topology. The quality of these common approximate mesh pairs is quantitatively measured by an Energy term, which is minimized during the optimization process. Later, the meshes with the common topology are interpolated (morphing) to build the final digital twin of the plant. Experimental results show that the proposed methodology to attain the final morph has the potential to be a vital module, which could be responsible for the visual updates in the digital replica of the digital twin of the plant. Karanvir Singh, Abdulmotaleb El Saddik, Mukesh Saini |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | ASTAnet: Transformer-based Siamese Network for Robust Audio-to-Audio Alignment in Amateur User Generated Audio ClipsabstractAudio alignment involves synchronizing two or more audio recordings. Existing methods depend on handcrafted features and struggle with precision in lengthy or noisy recordings. Deep learning techniques have proven effective across various domains; however, their application in audio-to-audio alignment is still in its infancy. We propose ASTAnet, a framework that integrates the Vision Transformer for feature extraction with the Siamese network for similarity estimation. With timestamp positional encoding, ASTAnet improves temporal precision and reduces alignment errors using a contrastive learning objective based on Euclidean distance. Our experiments achieved an overall mean absolute error value of 0.005, a 1.8X improvement compared to the previous works. Extensive evaluations demonstrate its effectiveness, particularly for varying and longer audio recordings. Malya Singh, Priyankar Choudhary, Abdulmotaleb El Saddik, Mukesh Saini |
ICME | 4 |
| 2025 | GroMo25: ACM Multimedia 2025 Grand Challenge for Plant Growth Modeling with Multiview Images
Shreya Bansal, Ruchi Bhatt, Amanpreet Chander, Malya Singh, Mohan Kankanhalli, Abdulmotaleb El Saddik, Mukesh Saini |
ACM Multimedia | 8 |
| 2024 | Multimedia datasets for anomaly detection: a review
Pratibha Kumari 0001, Anterpreet Kaur Bedi, Mukesh Saini |
Multim. Tools Appl. | 3 |
| 2024 | A highly robust deep learning technique for overlap detection using audio fingerprinting
Akash Uikey, Anterpreet Kaur Bedi, Priyankar Choudhary, Wei Tsang Ooi, Mukesh Saini |
Multim. Tools Appl. | 5 |
| 2024 | A deep learning approach for early detection of drought stress in maize using proximal scale digital images
Pooja Goyal, Rakesh Sharda, Mukesh Saini, Mukesh Siag |
Neural Comput. Appl. | 3 |
| 2024 | Concept drift challenge in multimedia anomaly detection: A case study with facial datasets
Pratibha Kumari 0001, Priyankar Choudhary, Vinit Kujur, Pradeep K. Atrey, Mukesh Saini |
Signal Process. Image Commun. | 5 |
| 2023 | Towards Digital Twin of Crops for Growth Modelling using Virtual RealityabstractA major problem that a farmer faces, while adopting a new crop variety in the farm; is the uncertainty associated with its growth. Farmers working on real farms are not aware of the growth models, even for the existing crops. Hence, there is a need for more accessible and intuitive models. This work is a step towards the realization of another promising model, which is the digital twin of a crop. A primary requirement of the digital twin is the digital representation of the crop itself. Extending that notion, the work discusses the development of 3D assets of crops and their temporal alignment. It also describes the methodology involved in the development of a VR framework, which stores the ideal growth of a crop. This framework could be useful to farmers who want to confirm the growth of their crops. Furthermore, it also proposes a quantitative metric to evaluate the VR framework. The consistency of this proposed metric is further backed by a user study which is based on a qualitative method. Karanvir Singh, Mukesh Saini |
MMAsia | 2 |
| 2023 | Mapi-Pro: An Energy Efficient Memory Mapping Technique for Intermittent ComputingabstractBattery-less technology evolved to replace battery usage in space, deep mines, and other environments to reduce cost and pollution. Non-volatile memory (NVM) based processors were explored for saving the system state during a power failure. Such devices have a small SRAM and large non-volatile memory. To make the system energy efficient, we need to use SRAM efficiently. So we must select some portions of the application and map them to either SRAM or FRAM. This paper proposes an ILP-based memory mapping technique for intermittently powered IoT devices. Our proposed technique gives an optimal mapping choice that reduces the system’s Energy-Delay Product (EDP). We validated our system using TI-based MSP430FR6989 and MSP430F5529 development boards. Our proposed memory configuration consumes 38.10% less EDP than the baseline configuration and 9.30% less EDP than the existing work under stable power. Our proposed configuration achieves 20.15% less EDP than the baseline configuration and 26.87% less EDP than the existing work under unstable power. This work supports intermittent computing and works efficiently during frequent power failures. Satya Jaswanth Badri, Mukesh Saini, Neeraj Goel |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | An Efficient NVM-Based Architecture for Intermittent Computing Under Energy ConstraintsabstractBatteryless technology evolved to replace battery technology. Nonvolatile memory (NVM)-based processors were explored to store the program state during a power failure. The energy stored in a capacitor is used for a backup during a power failure. Since the size of a capacitor is fixed and limited, the available energy in a capacitor is also limited and fixed. Thus, the capacitor energy is insufficient to store the entire program state during frequent power failures. This article proposes an architecture that assures safe backup of volatile contents during a power failure under energy constraints. Using a proposed dirty block table (DBT) and writeback queue (WBQ), this work limits the number of dirty blocks in the L1 cache at any given time. We further conducted a set of experiments by varying the parameter sizes to help the user make appropriate design decisions concerning their energy requirements. The proposed architecture decreases energy consumption by 17.56%, the number of writes to NVM by 18.97% at last level cache (LLC), and 10.66% at a main-memory level compared to baseline architecture. Satya Jaswanth Badri, Mukesh Saini, Neeraj Goel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | An Experimental Study of the Concept Drift Challenge in Farm Intrusion Detection using AudioabstractIntrusion detection in farm settings is a challenging task. The data distribution suffers significant drift due to variations in environmental sounds. In this paper, we study the effect of this drift on state-of-the-art deep models. We experimentally found that the traditional models fail to deal with such variations and exhibit performance degradation. VGG16 turns out to be the best deep model, which shows an improvement of 2.76% over the best-performing state-of-the-art deep model on parameter F1-score. Consequently, we make an initial attempt to overcome this drift by integrating an unsupervised background noise component with standard models. On denoised signals, we obtained an average improvement of 50% on parameter accuracy for VGG16. The experimental analysis demonstrates the need for adaptive learning to handle the Spatio-temporal drifts in outdoor farm settings. Ruchi Bhatt, Simrandeep Singh, Priyankar Choudhary, Mukesh Saini |
AVSS | 4 |
| 2021 | A Seismic Sensor based Human Activity Recognition Framework using Deep LearningabstractActivity recognition has gained attention due to the rapid development of microelectromechanical sensors. Numerous human-centric applications in healthcare, security, and smart environments can benefit from an efficient human activity recognition system. In this paper, we demonstrate the use of a seismic sensor for human activity recognition. Traditionally, researchers have relied on handcrafted features to identify the target activity, but these features may be inefficient in complex and noisy environments. The proposed framework employs an autoencoder to map the activity into a compact representative descriptor. Further, an Artificial Neural Network (ANN) classifier is trained on the extracted descriptors. We compare the proposed framework with multiple machine learning classifiers and a state-of-the-art framework on different evaluation metrics. On 5-fold cross-validation, the proposed approach outperforms the state-of-the-art in terms of precision and recall by an average of 10.68 and 23.36%, respectively. We also collected a dataset to assess the efficacy of the proposed seismic sensor-based activity recognition. The dataset is collected in a variety of challenging environments, such as variable grass length, soil moisture content, and the passing of unwanted vehicles nearby. Priyankar Choudhary, Neeraj Goel, Mukesh Saini |
AVSS | 3 |
| 2020 | Dynamic Scheduling of an Autonomous PTZ Camera for Effective SurveillanceabstractPTZ cameras can be an effective replacement for multiple camera networks with their pan-tilt-zoom capability. However, the state of the art scheduling method for the PTZ cameras focuses mainly on tracking, not on coverage. In this paper, we aim to maximize coverage as well as information gain, thus, leading to effective surveillance. Towards this goal, we define an information map that represents the sensitivity of a region. We propose a scheduling algorithm in which the camera visits those states more often that are likely to be more important than others, thus, maximizing information gain. A probabilistic framework is used to maximize information gain and coverage simultaneously. Currently, there are no existing datasets and methods to evaluate PTZ camera scheduling methods. We build a real multi-camera dataset and develop a performance measure for this purpose. Experimental results show that the proposed stochastic scheduling algorithm based on adaptive information gain probability is better than traditional as well as other variants proposed in the paper in terms of information gain as well as coverage. Pratibha Kumari 0001, Nikhil Nandyala, Allu Krishna Sai Teja, Neeraj Goel, Mukesh Saini |
MASS | 5 |
| 2019 | Watch Me from Distance (WMD): A Privacy-Preserving Long-Distance Video Surveillance SystemabstractPreserving the privacy of people in video surveillance systems is quite challenging, and a significant amount of research has been done to solve this problem in recent times. Majority of existing techniques are based on detecting bodily cues such as face and/or silhouette and obscuring them so that people in the videos cannot be identified. We observe that merely hiding bodily cues is not enough for protecting identities of the individuals in the videos. An adversary, who has prior contextual knowledge about the surveilled area, can identify people in the video by exploiting the implicit inference channels such as behavior, place, and time. This article presents an anonymous surveillance system, called Watch Me from Distance (WMD), which advocates for outsourcing of surveillance video monitoring (similar to call centers) to the long-distance sites where professional security operators watch the video and alert the local site when any suspicious or abnormal event takes place. We find that long-distance monitoring helps in decoupling the contextual knowledge of security operators. Since security operators at the remote site could turn into adversaries, a trust computation model to determine the credibility of the operators is presented as an integral part of the proposed system. The feasibility study and experiments suggest that the proposed system provides more robust measures of privacy yet maintains surveillance effectiveness. Pradeep K. Atrey, Bakul Trehan, Mukesh Saini |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | Multimodal Drunk Density Estimation for Safety AssessmentabstractDrinking alcohol in excess leads to lower self-consciousness, damaging a persons judgment and thus enhances risk of aggressive behavior. It leads to various problems like social abuse, violence, crime, and road accidents. Hence, density of drunk people in a given area is one of the indicators of safety risk. In this work we propose a novel framework to determine density of drunk people in a smart city scenario. Smart cities provide multiple sources of information such as audio, video, and text (online social networks). We detect presence of drunk persons along with time and location by analyzing these information sources individually and then fuse this information to obtain a single drunk index for a given location. We put special focus text analysis and propose a more accurate method to detect drunk event (person) with an accuracy of 84.2%. Experimental results demonstrate the functionality and efficacy of the proposed framework. Pratibha Kumari 0001, Mandhatya Singh, Mukesh Saini |
AVSS | 3 |
| 2017 | A Crowd-Sourced Adaptive Safe Navigation for Smart CitiesabstractThere has been a lot of work on finding safest path from source to destination. However, all these works mainly rely on the crime data available through government sources. The works are missing many important factor that affect safety of a route. In this work, we first propose safety measurement score and then show how it can be used for safe navigation. We proposed a comprehensive safety measure that adapts with time according to the user feedback and online crowd-sourced information. We have developed a complete end-to-end system to demonstrate efficacy of the proposed framework. Neeraj Goel, Rajat Sharma, N. Nikhil, S. D. Mahanoor, Mukesh Saini |
ISM | 5 |
| 2017 | The One Man ShowabstractTraditional video mashup and summarization methods assume that all video clips have common audio, but with varying quality. Hence, selecting the best quality audio is sufficient. In this work we explore a new scenario in which a single person plays each instrument one by one, leading to multi-view video clips, but each video clip having only partial audio, e.g. a single instrument or vocal. To get the complete audio, we need to merge all partial audios. In this way, although the videos are recorded at different times, they correspond to a common timeline in the final mashup. The proposed framework automatically recognizes the type of instrument for a given audio clip and employs instrument specific enhancements before merging. To select video segments for the final mashup, the framework automatically recognizes the dominant instrument from given audio clips and selects the video capturing that instrument. The proposed framework enables a single artist to play instruments, add vocals, and act in the mashup video. The complete framework has been implemented in the form of an android application. Sai Samarth R. Phaye, Love Mehta, Mukesh Saini |
ISM | 3 |
| 2017 | Shall IoT User Interfaces Start Recommending Multimedia Devices as Well?abstractRecommendation of media objects, such as audio and video clips, has been there for a while. Most user interfaces include a list of favourites or most popular media objects. On the other hand, recommendation of multimedia devices is limited to shopping websites. With the evolution of IoT, however, users these days are surrounded by many interconnected devices that can be used to accomplish the same task at any given time. For example, while at home, a user can choose to play a media le on a smartphone, tablet, laptop, TV, or home theatre. In this article we investigate the question of whether or not users are ready to accept automatic recommendation of physical things with a case study of media playback devices. We further investigate various factors that a ect user's choice of media playback device with a user study. The analysis shows that users like device recommendation in general. In addition, many users even prefer the playback to be automatically transferred directly to the most appropriate device, while other users just want a noti cation. We also found that user's gender, profession, age, and the duration of media le a ect the choice of playback device. Mukesh Saini, Ali Danesh, Abdulmotaleb El Saddik |
ISM | 1 |
| 2017 | InCloud: a cloud-based middleware for vehicular infotainment systems
Mukesh Saini, Kazi Masudul Alam, Haolin Guo, Abdulhameed Alelaiwi, Abdulmotaleb El Saddik |
Multim. Tools Appl. | 1 |
| 2017 | DST: days spent together using soft sensory information on OSNs - a case study on Facebook
Fatimah Al-Zamzami, Mukesh Saini, Abdulmotaleb El Saddik |
Soft Comput. | 2 |
| 2016 | MUDVA: A multi-sensory dataset for the vehicular CPS applicationsabstractVehicular Cyber-Physical System (VCPS) is a new trend in the research of the intelligent transport systems (ITS). In VCPS, vehicles work as a hub of sensors to collect interior and exterior information about the vehicle. Vehicles can use ad-hoc networking or 3G/LTE communication technology to share useful information with their neighboring vehicles or with the infrastructures to accomplish user safety, comfort, and entertainment tasks. In order to facilitate efficient sensor-services fusion in the VCPS applications, we need real life vehicular sensory datasets. While there has been many datasets containing vehicle mobility traces, there is hardly any that contains sensory information to be shared on the network. In this paper, we present a scenario specific modular dataset architecture along with some multi-sensory dataset modules. One of the dataset modules provides time synchronized multi-vehicle data including multi-view video, multi-directional sound, GPS, accelerometer, gyroscope, and magnetic field sensors. Each of the three vehicles recorded front, back, left, and right videos while moving closely in the suburban areas to let explore vehicular cooperative applications. Another module presents necessary tools and datasets to identify vehicular events such as acceleration, deceleration, turn, and no-turn events. We also present development details of a safety application using the presented datasets along with a list of other possible applications. Kazi Masudul Alam, Mohammed Bin Hariz, Seyed Vahid Hosseinioun, Mukesh Saini, Abdulmotaleb El Saddik |
MMSP | 4 |
| 2016 | Personality assessment using multiple online social networks
Shally Bhardwaj, Pradeep K. Atrey, Mukesh Saini, Abdulmotaleb El Saddik |
Multim. Tools Appl. | 3 |
| 2015 | A Proxemic Multimedia Interaction over the Internet of Things
Ali Danesh, Mukesh Saini, Abdulmotaleb El Saddik |
MMM (2) | 2 |
| 2014 | A Real-Time Smart Assistant for Video Surveillance Through Handheld DevicesabstractIn a remote surveillance system, a high resolution surveillance camera streams its video to a user's handheld device. Such devices are unable to make use of the high resolution video due to their limited display size and bandwidth. In this paper, we propose a method to assist the mobile operator of the surveillance camera in focusing on sensitive regions of the video. Our system automatically identifies relevant regions. We introduce a pan and zoom strategy to ensure that the operator is able to see fine details in these areas while maintaining contextual knowledge. Regions of interest are identified using foreground detection as well as face and body detection. The efficacy of the proposed method is demonstrated through a user study. Our proposed method was reported to be more useful than two comparable approaches for getting an understanding of the activities in a surveillance scene while maintaining context. Hao Kuang, Benjamin Guthier, Mukesh Saini, Dwarikanath Mahapatra, Abdulmotaleb El Saddik |
ACM Multimedia | 3 |
| 2014 | W3-privacy: understanding what, when, and where inference channels in multi-camera surveillance video
Mukesh Saini, Pradeep K. Atrey, Sharad Mehrotra, Mohan Kankanhalli |
Multim. Tools Appl. | 1 |
| 2013 | Jiku director: a mobile video mashup systemabstractIn this technical demonstration, we demonstrate a Web-based application called Jiku Director that automatically creates a mashup video from event videos uploaded by users. The system runs an algorithm that considers view quality (shakiness, tilt, occlusion), video quality (blockiness, contrast, sharpness, illumination, burned pixels), and spatial-temporal diversity (shot angles, shot lengths) to create a mashup video with smooth shot transitions while covering the event from different perspectives. Duong-Trung-Dung Nguyen, Mukesh Saini, Vu-Thanh Nguyen, Wei Tsang Ooi |
ACM Multimedia | 2 |
| 2013 | The jiku mobile video datasetabstractProliferation of mobile devices with video recording capability has lead to a tremendous growth in the amount of user-generated mobile videos. Researchers have embarked on developing new interesting applications and enhancement algorithms for mobile video. There is, however, no standard dataset with videos that could represent characteristics of mobile videos captured in realistic scenarios. In this paper, we present our effort to create one such dataset, consisting of videos simultaneously recorded using mobile devices in an unconstrained manner by multiple users attending performance events. Each video is accompanied by concurrent readings from accelerometer and compass sensors. At the time of writing, the dataset contains 473 video clips, with a total length of 30 hours 41 minutes and total size of 122.8 GB. We believe this dataset is useful as a common benchmark dataset for a variety of different research topics on mobile videos, including video analytics, video quality enhancement, and automatic video mashups. Mukesh Saini, Padmanabha Venkatagiri Seshadri, Wei Tsang Ooi, Mun Choon Chan |
MMSys | 1 |
| 2012 | MoViMash: online mobile video mashupabstractWith the proliferation of mobile video cameras, it is becoming easier for users to capture videos of live performances and socially share them with friends and public. As an attendee of such live performances typically has limited mobility, each video camera is able to capture only from a range of restricted viewing angles and distance, producing a rather monotonous video clip. At such performances, however, multiple video clips can be captured by different users, likely from different angles and distances. These videos can be combined to produce a more interesting and representative mashup of the live performances for broadcasting and sharing. The earlier works select video shots merely based on the quality of currently available videos. In real video editing process, however, recent selection history plays an important role in choosing future shots. In this work, we present MoViMash, a framework for automatic online video mashup that makes smooth shot transitions to cover the performance from diverse perspectives. Shot transition and shot length distributions are learned from professionally edited videos. Further, we introduce view quality assessment in the framework to filter out shaky, occluded, and tilted videos. To the best of our knowledge, this is the first attempt to incorporate history-based diversity measurement, state-based video editing rules, and view quality in automated video mashup generations. Experimental results have been provided to demonstrate the effectiveness of MoViMash framework. Mukesh Saini, Raghudeep Gadde, Shuicheng Yan, Wei Tsang Ooi |
ACM Multimedia | 1 |
| 2012 | Jiku live: a live zoomable video streaming systemabstractWe present Jiku Live, a client-server system that supports zoom and pan operations in live video streaming from network cameras. The client is an Android mobile application that plays back live video from a selected camera and supports multi-touch zoom and pan interaction. The server acquires video streams from network cameras and transcodes the video feeds into one-second video segments at multiple resolutions. The transcoded video supports random access into any region-of-interest (RoI) within the video. Upon receiving zoom or pan requests, the server transmits the RoIs from the corresponding video segments to the client. Arash Shafiei 0001, Ngo Quang Minh Khiem, Guntur Ravindra, Mukesh Saini, Cong Pang, Wei Tsang Ooi |
ACM Multimedia | 4 |
| 2012 | Adaptive Workload Equalization in Multi-Camera Surveillance SystemsabstractSurveillance and monitoring systems generally employ a large number of cameras to capture people's activities in the environment. These activities are analyzed by hosts (human operators and/or computers) for threat detection. Threat detection is a target centric task in which the behavior of each target is analyzed separately, which requires a significant amount of human attention and is a computationally intensive task for automatic analysis. In order to meet the real-time requirements of surveillance, it is necessary to distribute the video processing load over multiple hosts. In general, cameras are statically assigned to the hosts; we show that this is not a desirable solution as the workload for a particular camera may vary over time depending on the number of targets in its view. In the future, this uneven distribution of workload will become more critical as the sensing infrastructures are being deployed on the cloud. In this paper, we model the camera workload as a function of the number of targets, and use that to dynamically assign video feeds to the hosts. Experimental results show that the proposed model successfully captures the variability of the workload, and that the dynamic workload assignment provides better results than a static assignment. Mukesh Saini, Xiangyu Wang 0002, Pradeep K. Atrey, Mohan Kankanhalli |
IEEE Trans. Multim. | 1 |
| 2011 | Anonymous surveillanceabstractVideo surveillance is a very effective tool of surveillance that enables a single security agent to monitor wide areas. However, it compromises the privacy of the individuals. There have been attempts to obfuscate face and silhouette regions of the images to hide the identity of individuals. We recognize that in traditional surveillance systems, the viewer generally has sufficient contextual knowledge about location of the camera, time, and activity patterns; which can lead to identity leakage even when the visual cues (face and appearance) are not present. In this way, the viewer can relate the identity of individuals to the sensitive information in the video causing privacy loss. In order to provide robust privacy preservation, the context knowledge needs to be decoupled from the video; however, human monitoring of the videos is also necessary for the assessment of the situation. In this paper we propose anonymous surveillance framework that decouples the contextual knowledge and video to the minimal extent required for situation assessment. The experimental results confirm that the proposed framework is very effective in protecting the privacy, yet does not affect much of the surveillance utility of the data. Mukesh Saini, Pradeep K. Atrey, Sharad Mehrotra, Mohan Kankanhalli |
ICME | 1 |
| 2011 | Dynamic workload assignment in video surveillance systemsabstractCurrent surveillance systems consist of large numbers of cameras. The video feeds from cameras are automatically processed for threat detection, which is a computationally intensive task. In order to meet the real-time requirements of surveillance, we need to distribute the video processing over multiple computers. Generally the cameras are statically assigned to the processors; we show that this is not a desirable solution as the workload for a particular camera may vary over time depending on the number of the targets in its view. In future, this uneven distribution of workload will become more critical as the sensing infrastructures are being deployed on the cloud. In this work, we model the camera workload as a function of the number of targets, and use that to dynamically assign video feeds to the processors. Experimental results show that the proposed model successfully captures the variability of the workload, and that dynamic workload assignment provides better results than a static assignment. Mukesh Saini, Xiangyu Wang 0002, Pradeep K. Atrey, Mohan Kankanhalli |
ICME | 1 |
| 2010 | Functionality Delegation in Distributed Surveillance SystemsabstractThe utilization of multimedia devices is growing rapidly in surveillance and monitoring applications. These multimedia surveillance systems need to process large amounts of multimodal sensor data in order to detect events and objects. While processing this large amount of data, the system faces many processing and network bottlenecks. The design of efficient multimedia surveillance system requires intelligent architectural decisions and performance evaluation to cope with these resource demands. One critical issue among all these architectures is task assignment among processing units. To study the effect of this task assignment on system performance with quantifiable performance measures is very useful and challenging. We define a Functionality Delegation Coefficient which abstracts the delegation of functionality among processing units of a distributed surveillance system and show its effect on event blocking probability and response time. Simulation and real implementation results are provided to validate the model. Mukesh Saini, Pradeep K. Atrey, Sabu Emmanuel, Mohan Kankanhalli |
AVSS | 1 |
| 2010 | Privacy modeling for video data publicationabstractVideo cameras are being extensively used in many applications. Huge amounts of video are being recorded and stored everyday by surveillance systems. Any proposed application of this data raises severe privacy concerns. An assessment of privacy loss is necessary before any potential application of the data. In traditional methods of privacy modeling, researchers have focused on explicit means of identity leakage like facial information, etc. However, other implicit inference channels through which individual's an identity can be learned have not been considered. For example, an adversary can observe the behavior, look at the places visited and combine that with the temporal information to infer the identity of the person in the video. In this work, we thoroughly investigate privacy issues involved with the video data considering both implicit and explicit channels. We first establish an analogy with the statistical databases and then propose a model to calculate the privacy loss that might occur due to publication of the video data. The experimental results demonstrate the utility of the proposed model. Mukesh Saini, Pradeep K. Atrey, Sharad Mehrotra, Sabu Emmanuel, Mohan Kankanhalli |
ICME | 1 |
| 2009 | Context-Based Multimedia Sensor Selection MethodabstractModern multimedia systems have large number of sensors spread across a wide area. In a time-shared multimedia system, many people will be making queries to the system simultaneously which requires sharing of computing resources. In such scenarios, processing information from all the sensors for each query will make the system inefficient. Considering the fact that only few sensors provide information relevant to the query, we can reduce the cost incurred in query evaluation by efficiently selecting a subset of sensors to be processed without compromising the system performance. This paper demonstrates a two-stage sensor selection method which uses contextual information and confidence in individual sensors to select sensors which provide more reliable answers to the queries. Mukesh Saini, Mohan Kankanhalli |
AVSS | 1 |
| 2009 | A Flexible Surveillance System ArchitectureabstractTraditional multimedia surveillance systems are task specific and tightly coupled to the environment. Moreover, system designs generally start with the assumption that the environment, context, and sensors always remain static. With such a tight coupling, it becomes very difficult to port the system to new environments. Furthermore, for most of the systems, there is no straightforward way to upgrade the existing system to incorporate technological advancements such as new sensors or novel feature extraction techniques. We propose a flexible surveillance system architecture which can be easily ported in different environments, is dynamic without any significant compromise in system performance, and can be extended to integrate newer technological developments. We also introduce the notion of environment model (EM), which completely defines the coupling between system and the physical environment. The isolation of environment specific variables in EM makes the system easily portable in different environments. We present results of a prototype implementation of the system that highlights our design goals. Mukesh Saini, Mohan Kankanhalli, Ramesh Jain 0001 |
AVSS | 1 |
| 2009 | Performance Modeling of Multimedia Surveillance SystemsabstractAutomated surveillance is critically important in the current scenario of heightened security concerns. Therefore, there has been a surge in the development of surveillance systems. Surveillance systems employ sensors to capture various environmental aspects in order to reason about the dynamically changing situation. Due to their cheap availability, the number of sensors used in modern systems is quite large. While processing the large amount of sensor data, the system faces many bottlenecks in terms of processor and memory requirements. Despite many efforts to propose efficient system architectures, they all ignore the study of the dynamic behavior of the system and the impact of various factors on system performance. In this work, we develop an analytical model to evaluate the performance of surveillance systems. Using the proposed model, we obtain closed form equations for event miss probability and response time as a function of system parameters. The results obtained from the model are validated with those obtained from the simulator. Finally we explore the different trade-offs among the system parameters and performance metrics. Mukesh Saini, Yashas Natraj, Mohan Kankanhalli |
ISM | 1 |
| 2008 | Illumination invariant tracking in office environments using neurobiology-saliency based particle filterabstractBackground subtraction is a commonly employed approach for tracking in scenarios where the ambience is more or less constant in terms of illumination and number of objects. However in office environments, where the illumination can very easily change by switching off or on lights, the background subtraction method can lead to erroneous tracking. In this paper we propose a neurobiology-saliency based particle filter approach that uses low-level features like color, luminance and edge information along with motion cues to track a single person. We have tested our method on clips showing a single person carrying out a range of activities expected in an office environment. Our method performs better than a background subtraction method using a Kalman filter, in terms of the number of frames showing correct tracking and change detection for automatic initialization of tracks. Dwarikanath Mahapatra, Mukesh Saini, Ying Sun 0001 |
ICME | 2 |
| 2008 | Multimodal observation systemsabstractIn recent years, we have seen a significant research interest in a number of multimodal sensing applications like surveillance, video ethnography, tele-presence, assisted living, life blogging etc. However, these applications are currently evolving as separate silos with no interconnection. Further, the individual application-centric architectures typically tend to focus on specific sensors, specific (hardwired) queries and deal with specific environments. We present a generic sensing architecture 'Observation System', which allows multiple users to undertake different applications through abstracted interaction with a common set of sensors. The observation system observes behavior of various objects in an environment and keeps a record of important events and activities in an eventbase. In this system, multifarious data collected from disparate sensors and other sources are correlated to understand and gain insights in the environment. The observation system has applications in many areas including but not limited to surveillance, traffic monitoring, ethnography, marketing, and healthcare. In this paper, we present the architecture and functionality of such a system and present details of activity detection using multiple sensor streams in a distributed sensing environment. We also present results of such an approach and potential extensions to the analysis of more complex activities and events. Mukesh Saini, Vivek K. Singh 0001, Ramesh Jain 0001, Mohan Kankanhalli |
ACM Multimedia | 1 |