Zeeshan Rasheed 0002

dblp:50/6581-2 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-2369-9753ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-authorDatabases, data management, data science and information retrieval · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2025 Uncertainty-aware Spatio-Temporal Human Mobility Modeling and Anomaly Detection
abstract
Given the temporal GPS coordinates from a large set of human agents, how can we model their mobility behavior toward effective anomaly (e.g., bad-actor or malicious behavior) detection without any labeled data? Human mobility and trajectory modeling have been extensively studied, showcasing varying abilities to manage complex inputs and balance performance-efficiency trade-offs. In this work, we formulate anomaly detection in complex human behavior by modeling raw GPS data as a sequence of stay-point events, each characterized by spatio-temporal features, along with trips (i.e., commutes) between the stay-points. Our problem formulation allows us to leverage modern sequence models for unsupervised training and anomaly detection. Notably, we equip our proposed model USTAD (for Uncertainty-aware Spatio-Temporal Anomaly Detection) with aleatoric (i.e., data) uncertainty estimation to account for inherent stochasticity in certain individuals' behavior, as well as epistemic (i.e., model) uncertainty to handle data sparsity under a large variety of human behaviors. Together, aleatoric and epistemic uncertainties unlock a robust loss function as well as uncertainty-aware decision-making in anomaly scoring. Extensive experiments show that USTAD significantly outperforms baselines in industry-scale data. We open-source all code at https://github.com/wenhaomin/USTAD.
Haomin Wen, Shurui Cao, Zeeshan Rasheed 0002, Khurram Shafique, Leman Akoglu
SIGSPATIAL/GIS3
2024 Generating Trajectories from Implicit Neural Models
abstract
Modeling human mobility under uncertain conditions and individual preferences remains a difficult and unsolved problem. Data-driven deep learning approaches require extensive trajectory data for training, while more traditional methods often assume deterministic conditions or simple minimum-cost paths. We propose an implicit neural representation (INR) to learn continuous, latent fields of stochastic traffic properties over space and time. We successfully impute speeds on a road network with hundreds of thousands of edges from only a few hundred vehicles, then illustrate the quality of these representations on a trajectory generation task. A near-shortest-path algorithm weighted by the INR’s predictions produces plausible real-world routing choices, showing potential for applications in route planning and anomaly detection.
Mark Tenzer, Emmanuel Tung, Zeeshan Rasheed 0002, Khurram Shafique
MDM3
2023 The Geospatial Generalization Problem: When Mobility Isn't Mobile
abstract
Human mobility research has significantly benefited from recent advances in machine learning, as have numerous other industries. Aided by the ever-increasing availability of geospatial and mobility data, machine learning models have enabled large-scale systems for simulating city-wide macro and micro mobility behaviors, urban planning, transportation management, and disaster relief optimization. However, while many fields have invested significant effort in solving the model transferability and generalization problem, the inability of machine learning-based human mobility models to generalize to new locations has come to be implicitly accepted in most geospatial research. In this vision paper, we focus on this geospatial generalization problem, its root causes, and how it is restricting the applications of otherwise-promising research. Most importantly, we argue for several data- and modeling-driven innovations which could help remedy this problem, spanning mega-scale simulations, large foundation models, and multi-task, transfer, and meta-learning. We also spotlight a handful of promising ideas which have recently emerged from the community. We hope that these proposals take root and help develop more capable, flexible, and generalizable models in research and industry.
Mark Tenzer, Zeeshan Rasheed 0002, Khurram Shafique
SIGSPATIAL/GIS2
2022 Learning citywide patterns of life from trajectory monitoring
abstract
The recent proliferation of real-world human mobility datasets has catalyzed geospatial and transportation research in trajectory prediction, demand forecasting, travel time estimation, and anomaly detection. However, these datasets also enable, more broadly, a descriptive analysis of intricate systems of human mobility. We formally define patterns of life analysis as a natural, explainable extension of online unsupervised anomaly detection, where we not only monitor a data stream for anomalies but also explicitly extract normal patterns over time. To learn patterns of life, we adapt Grow When Required (GWR) episodic memory from research in computational biology and neurorobotics to a new domain of geospatial analysis. This biologically-inspired neural network, related to self-organizing maps (SOM), constructs a set of "memories" or prototype traffic patterns incrementally as it iterates over the GPS stream. It then compares each new observation to its prior experiences, inducing an online, unsupervised clustering and anomaly detection on the data. We mine patterns-of-interest from the Porto taxi dataset, including both major public holidays and newly-discovered transportation anomalies, such as festivals and concerts which, to our knowledge, have not been previously acknowledged or reported in prior work. We anticipate that the capability to incrementally learn normal and abnormal road transportation behavior will be useful in many domains, including smart cities, autonomous vehicles, and urban planning and management.
Mark Tenzer, Zeeshan Rasheed 0002, Khurram Shafique
SIGSPATIAL/GIS2
2022 Meta-learning over time for destination prediction tasks
abstract
A need to understand and predict vehicles' behavior underlies both public and private goals in the transportation domain, including urban planning and management, ride-sharing services, and intelligent transportation systems. Individuals' preferences and intended destinations vary throughout the day, week, and year: for example, bars are most popular in the evenings, and beaches are most popular in the summer. Despite this principle, we note that recent studies on a popular benchmark dataset from Porto, Portugal have found, at best, only marginal improvements in predictive performance from incorporating temporal information. We propose an approach based on hypernetworks, a variant of meta-learning ("learning to learn") in which a neural network learns to change its own weights in response to an input. In our case, the weights responsible for destination prediction vary with the metadata, in particular the time, of the input trajectory. The time-conditioned weights notably improve the model's error relative to ablation studies and comparable prior work, and we confirm our hypothesis that knowledge of time should improve prediction of a vehicle's intended destination.
Mark Tenzer, Zeeshan Rasheed 0002, Khurram Shafique, Nuno Vasconcelos
SIGSPATIAL/GIS2
2021 DIVINIA: Rare Object Localization and Search in Overhead Imagery
abstract
This work introduces DIVINIA, a feature extractor and novel training objective for content-based image retrieval. DIVINIA combines a semantic matching objective with a ranking objective to produce a feature extractor that is able to retrieve semantically relevant regions from a large search corpus. It further ranks them appropriately according to visual similarity. Furthermore, DIVINIA provides a mechanism for performing one-shot and even zero-shot object localization without the need to fine-tune the feature extraction model or re-index the corpus of search features. We demonstrate the capabilities of the DIVINIA system in the context of object localization in satellite imagery. We present quantitative and qualitative results that show robust domain transfer between satellite image optics and sensor modalities. We show good precision and search relevance ordering when returning areas of interest to specific object classes.
Jonathan Amazon, Khurram Shafique, Zeeshan Rasheed 0002, Aaron Reite
ICDM3
2010 Automatic Geo-Registration for Port Surveillance
abstract
This paper proposes a new solution to geo-register the nearly feature-less maritime video feeds. We detect the horizon using sizable or uniformly moving vessels, and estimate the vertical apex using water reflections of the street lamps. The computed horizon and apex provide a metric rectification that removes the affine distortions and reduces the searching space for geo-registration. Geo-registration is obtained by searching the best orientation where the estimated water masks on satellite images and camera views are matched. The proposed solution has the following contributions: first, water and coastlines are used as features for registration between horizontally looking maritime views and satellite images. Second, water reflections are proposed to estimate the vertical vanishing point. Third, we give algorithms for the detection of water areas in both satellite images and camera views. Experimental results and applications on cross camera tracking are demonstrated. We also discuss several observations, as well as limitations of the proposed approach.
Xiaochun Cao, Lin Wu 0001, Zeeshan Rasheed 0002, Tae Eun Choe, Feng Guo 0006, Niels Haering
Int. J. Pattern Recognit. Artif. Intell.3
2008 Automatic geo-registration of maritime video feeds
abstract
We propose an automatic method to geo-register maritime video feeds to satellite images. The method first detects horizon during the day time and apex during the night time for similarity rectification. It then finds water pixels on both satellite images and camera views. Finally, geo-registration is obtained by searching the best orientation where the water masks are matched. The current state-art-of-state video surveillance system benefits with this method by providing situational awareness on maps with physical data including object size, location, and velocity in the world coordinates. Experimental results and applications such as cross camera tracking are demonstrated.
Xiaochun Cao, Zeeshan Rasheed 0002, Niels Haering
ICPR2
2008 Modeling inter-camera space-time and appearance relationships for tracking across non-overlapping views
Omar Javed, Khurram Shafique, Zeeshan Rasheed 0002, Mubarak Shah
Comput. Vis. Image Underst.3
2005 On the use of computable features for film classification
Zeeshan Rasheed 0002, Yaser Sheikh, Mubarak Shah
IEEE Trans. Circuits Syst. Video Technol.1
2005 Detection and representation of scenes in videos
abstract
This paper presents a method to perform a high-level segmentation of videos into scenes. A scene can be defined as a subdivision of a play in which either the setting is fixed, or when it presents continuous action in one place. We exploit this fact and propose a novel approach for clustering shots into scenes by transforming this task into a graph partitioning problem. This is achieved by constructing a weighted undirected graph called a shot similarity graph (SSG), where each node represents a shot and the edges between the shots are weighted by their similarity based on color and motion information. The SSG is then split into subgraphs by applying the normalized cuts for graph partitioning. The partitions so obtained represent individual scenes in the video. When clustering the shots, we consider the global similarities of shots rather than the individual shot pairs. We also propose a method to describe the content of each scene by selecting one representative image from the video as a scene key-frame. Recently, DVDs have become available with a chapter selection option where each chapter is represented by one image. Our algorithm automates this objective which is useful for applications such as video-on-demand, digital libraries, and the Internet. Experiments are presented with promising results on several Hollywood movies and one sitcom.
Zeeshan Rasheed 0002, Mubarak Shah
IEEE Trans. Multim.1
2003 Scene Detection In Hollywood Movies and TV Shows
abstract
A scene can be defined as one of the subdivisions of a play in which the setting is fixed, or when it presents continuous action in one place. We propose a novel two-pass algorithm for scene boundary detection, which utilizes the motion content, shot length and color properties of shots as the features. In our approach, shots are first clustered by computing Backward Shot Coherence (BSC) - a shot color similarity measure that detects Potential Scene Boundaries (PSBs) in the videos. In the second pass we compute Scene Dynamics (SD), a function of shot length and the motion content in the potential scenes. In this pass, a scene merging criteria has been developed to remove weak PSBs in order to reduce over segmentation. We also propose a method to describe the content of each scene by selecting one representative image. The segmentation of video data into number of scenes facilitates an improved browsing of videos in electronic form, such as video on demand, digital libraries, Internet. The proposed algorithm has been tested on a variety of videos that include five Hollywood movies, one sitcom, and one interview program and promising results have been obtained.
Zeeshan Rasheed 0002, Mubarak Shah
CVPR (2)1
2003 Tracking Across Multiple Cameras With Disjoint Views
abstract
Conventional tracking approaches assume proximity in space, time and appearance of objects in successive observations. However, observations of objects are often widely separated in time and space when viewed from multiple non-overlapping cameras. To address this problem, we present a novel approach for establishing object correspondence across non-overlapping cameras. Our multicamera tracking algorithm exploits the redundance in paths that people and cars tend to follow, e.g. roads, walk-ways or corridors, by using motion trends and appearance of objects, to establish correspondence. Our system does not require any inter-camera calibration, instead the system learns the camera topology and path probabilities of objects using Parzen windows, during a training phase. Once the training is complete, correspondences are assigned using the maximum a posteriori (MAP) estimation framework. The learned parameters are updated with changing trajectory patterns. Experiments with real world videos are reported, which validate the proposed approach.
Omar Javed, Zeeshan Rasheed 0002, Khurram Shafique, Mubarak Shah
ICCV2
2003 KNIGHT™: a real time surveillance system for multiple and non-overlapping cameras
abstract
In this paper, we present a wide area surveillance system that detects, tracks and classifies moving objects across multiple cameras. At the single camera level, tracking is performed using a voting based approach that utilizes color and shape cues to establish correspondence. The system uses the single camera tracking results along with the relationship between camera field of view (FOV) boundaries to establish correspondence between views of the same object in multiple cameras. To this end, a novel approach is described to find the relationships between the FOV lines of cameras. The proposed approach combines tracking in cameras with overlapping and/or non-overlapping FOVs in a unified framework, without requiring explicit calibration. The proposed algorithm has been implemented in a real time system. The system uses a client-server architecture and runs at 10 Hz with three cameras.
Omar Javed, Zeeshan Rasheed 0002, Orkun Alatas, Mubarak Shah
ICME2
2001 A Framework for Segmentation of Talk and Game Shows
abstract
In this paper, we present a method to remove commercials from talk and game show videos and to segment these videos into host and guest shots. In our approach, we mainly rely on information contained in shot transitions, rather than analyzing the scene content of individual frames. We utilize the inherent differences in scene structure of commercials and talk shows to differentiate between them. Similarly, we make use of the well-defined structure of talk shows, which can be exploited to classify shots as host or guest shots. The entire show is first segmented into camera shots based on color histogram. Then, we construct a data-structure (shot connectivity graph) which links similar shots over time. Analysis of the shot connectivity graph helps us to automatically separate commercials from program segments. This is done by first detecting stories, and then assigning a weight to each story based on its likelihood of being a commercial. Further analysis on stories is done to distinguish shots of the hosts from shots of the guests. We have tested our approach on several full-length shows (including commercials) and have achieved video segmentation with high accuracy. The whole scheme is fast and works even on low quality video (160/spl times/120 pixel images at 5 Hz).
Omar Javed, Zeeshan Rasheed 0002, Mubarak Shah
ICCV2
2001 Human Tracking in Multiple Cameras
Sohaib Khan, Omar Javed, Zeeshan Rasheed 0002, Mubarak Shah
ICCV3