Lance M. Kaplan

dblp:47/4107 · DBLP profile ↗
← Back
145ranked-venue papers
25as first author
30since 2021 · last 2026
0000-0002-3627-4471ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 62 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 26 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 11 first-author · 4 since 2021Computer networks · 22 · 1 first-author · 1 since 2021Systems, architecture and hardware · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021Human-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 X-MAP: eXplainable Misclassification Analysis and Profiling for Spam and Phishing Detection
Qi Zhang 0104, Dian Chen 0007, Lance M. Kaplan, Audun Jøsang, Dong Hyun Jeong, Feng Chen 0001, Jin-Hee Cho
PAKDD (3)3
2025 Preliminary Insights Into Resource-Constrained Neuro-Symbolic Causal Complex Event Processing
abstract
We propose a neuro-symbolic approach for learning causal complex event models from multi-source data, integrating causal discovery and temporal logic. Given resource constraints, we employ signal-level fusion by averaging the data from different antennas of the same WiFi receiver, followed by downsampling to reduce computational overhead. We consider a dataset of WiFi Channel State Information capturing human activities alongside video data from which we extract atomic symbolic activities such as “moving the upper arm.” The extracted symbolic information is processed through LPCMCI (Latent PCMCI). This causal discovery method extends PCMCI (Peter and Clark Momentary Conditional Independence) to handle latent dependencies across multiple time steps while mitigating false discoveries due to auto-correlations. The resulting causal structure is then translated into a temporal logic formula, which serves as a symbolic constraint in a neuro-symbolic learning pipeline. To efficiently process and learn from these structured constraints under resource limitations, we leverage Spiking Neural Networks, which offer energy-efficient computation while preserving temporal dynamics.
Christian Bresciani, Luca Lavazza, Marco Cominelli, Liying Han, Gaofeng Dong, Francesco Gringoli, Lance M. Kaplan, Mani Srivastava 0001, Trevor J. Bihl, Erik Blasch, Felix J. Knutson, Federico Cerutti 0001
FUSION7
2025 The Effect of the Prior on Asymptotic Performance of Uncertain Naïve Bayesian Networks
abstract
In this work, we analyze limited knowledge about likelihoods in traditional fusion as an uncertain Naïve Bayesian network whose conditional probabilities are known within posterior Dirichlet distributions reflected of limited training data. In prior work, we showed that inferences from these uncertain networks are confidently precise despite finite training data as the number of features goes to infinity for a uniform prior. The inference is usually correct except for pathological cases that diminish as the training data size goes to infinity. This work extends the analysis by studying how various priors affect asymptotic inference when the priors do or do not match the generative process for the conditional probabilities. Furthermore, the work analyzes how the prior affects the convergence rate of the inference.
Lance M. Kaplan, James Zachary Hare, Parth Paritosh
FUSION1
2025 Evaluation of LLM Reasoning Under Uncertainty: An Atomic Comparison to Normative Approaches
abstract
Evaluating uncertainty in large language model (LLM) reasoning is challenging due to their vast parameter space, abstract knowledge representation, and limited transparency regarding training data. While normative formalisms, such as deductive logic, clearly define sound reasoning in the absence of uncertainty, reasoning under uncertainty admits multiple approaches, including probabilistic reasoning (e.g. Bayesian), belief function reasoning (e.g. Dempster-Shafer), or fuzzy logic, to name a few. This paper examines how LLMs handle uncertainty by analyzing outcomes based on an atomic fusion and reasoning problem. We establish a point of reference using the simplest of fusion topologies to facilitate transparency and understanding of how LLMs align with established theories. The reasoning approaches of different LLMs with varying complexities are compared to established normative frameworks, providing insights into which formalism best aligns with LLM reasoning and assessing its soundness and consistency. A deviation function for assessment is developed, and the results indicate that the tested LLMs' reasoning under uncertainty does not consistently align with established theories, even for the simplest information fusion topologies. These preliminary results form the basis for further investigations and LLM refinements.
J. P. de Villiers, Allan De Freitas, Anne-Laure Jousselme, Lance M. Kaplan, Erik Blasch, Claire Laudy, P. C. Costa
FUSION4
2025 Keeping the Best: The K-Best rule for Efficient Quickest Change Detection with Unknown Post-Change Distribution
abstract
We study the problem of quickest change detection (QCD) when the post-change distribution has parametric uncertainty. The generalized likelihood ratio (GLR) cumulative sum (CuSum) procedure is known to be asymptotically optimum in this setting. However, this rule requires significant memory and computational resources, making it difficult to implement in practice. To overcome this limitation, sliding window approaches, such as the window-limited GLR CuSum and window-limited adaptive CuSum tests, have been employed, where the test statistic is computed over a fixed window of the latest observations. We propose the K-Best rule which instead keeps track of K hypothesized change points that have the largest test statistic. This allows the hypothesized change points to reduce epistemic uncertainty over time, while restricting the number of hypothesized change points considered. We characterize the growth rate of the K-Best window necessary to achieve the detection performance of the GLR-CuSum rule and quantify the computational benefits over the existing windowing approaches.
James Zachary Hare, Lance M. Kaplan, Venugopal V. Veeravalli, Don Towsley
ICASSP2
2025 fair-LDP: Uncertainty-Guided Fairness and Privacy for Federated Healthcare Learning
abstract
Federated Learning (FL) offers a promising approach for collaborative model training in healthcare while preserving data privacy. However, existing FL methods often fall short in addressing two critical challenges: client-level fairness and compounded uncertainty from data heterogeneity and privacy-preserving mechanisms. We propose fair-LDP, a fairness-aware Local Differential Privacy framework that promotes fairness and privacy via uncertainty-guided aggregation in federated healthcare AI. fair-LDP leverages evidential neural networks (ENNs) to quantify predictive uncertainty and introduces a novel strategy that uses uncertainty-driven local differential privacy to guide fairness-aware updates while preserving data privacy. This ensures equitable performance across clients with varying data quality while mitigating the influence of unreliable or outlier updates. fair-LDP incorporates an adaptive mechanism that adjusts each client's privacy budget based on model performance, balancing fairness, privacy, and accuracy. We evaluate fair-LDP on real-world healthcare datasets under both IID and non-IID settings. Our experimental results show that it consistently outperforms state-of-the-art fairness-aware and privacy-preserving FL baselines, with no added computational overhead, while maintaining privacy guarantees comparable to homomorphic encryption and secure multiparty computation. By integrating uncertainty modeling, fairness-aware aggregation, and adaptive local differential privacy, fair-LDP provides a practical and principled solution for responsible, equitable, and privacy-preserving federated learning in healthcare.
Dian Chen 0007, Qi Zhang 0104, Lance M. Kaplan, Audun Jøsang, Dong Hyun Jeong, Feng Chen 0001, Jin-Hee Cho
ICDM3
2025 ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources
abstract
Multimodal deep learning systems are deployed in dynamic scenarios due to the robustness afforded by multiple sensing modalities. Nevertheless, they struggle with varying compute resource availability (due to multi-tenancy, device heterogeneity, etc.) and fluctuating quality of inputs (from sensor feed corruption, environmental noise, etc.). Statically provisioned multimodal systems cannot adapt when compute resources change over time, while existing dynamic networks struggle with strict compute budgets. Additionally, both systems often neglect the impact of variations in modality quality. Consequently, modalities suffering substantial corruption may needlessly consume resources better allocated towards other modalities. We propose ADMN, a layer-wise Adaptive Depth Multimodal Network capable of tackling both challenges - it adjusts the total number of active layers across all modalities to meet compute resource constraints, and continually reallocates layers across input modalities according to their modality quality. Our evaluations showcase ADMN can match the accuracy of state-of-the-art networks while reducing up to 75% of their floating-point operations.
Yuyang Yuan, Kang Yang 0005, Lance M. Kaplan, Mani Srivastava 0001
NeurIPS4
2025 Risk-aware classification via uncertainty quantification
abstract
Autonomous and semi-autonomous systems are using deep learning models to improve decision-making. However, deep classifiers can be overly confident in their incorrect predictions, a major issue especially in safety-critical domains. The present study introduces three foundational desiderata for developing real-world risk-aware classification systems. Expanding upon the previously proposed Evidential Deep Learning ( EDL ), we demonstrate the unity between these principles and EDL ’s operational attributes. We then augment EDL empowering autonomous agents to exercise discretion during structured decision-making when uncertainty and risks are inherent. We rigorously examine empirical scenarios to substantiate these theoretical innovations. In contrast to existing risk-aware classifiers, our proposed methodologies consistently exhibit superior performance, underscoring their transformative potential in risk-conscious classification strategies. • Evidential deep learning uses Dirichlet distributions to represent the predictive uncertainty of neural classifiers. • Pignistic probabilities can be used to model rational decision-making under uncertainty. • Risk awareness can be integrated into evidential classifiers using pignistic Dirichlet priors.
Murat Sensoy, Lance M. Kaplan, Simon J. Julier, Maryam Saleki, Federico Cerutti 0001
Expert Syst. Appl.2
2024 Uncertainty-Aware Influence Maximization: Enhancing Propagation in Competitive Social Networks with Subjective Logic
abstract
The Competitive Influence Maximization (CIM) problem involves entities competing to maximize influence in online social networks (OSNs). While Deep Reinforcement Learning (DRL) methods have shown promise, most assume binary user opinions and overlook behavioral factors. We introduce DRIM, a novel DRL-based CIM framework using Subjective Logic (SL) to incorporate user preferences and uncertainty, optimizing seed selection to spread true information while countering false information. DRIM’s Uncertainty-based Opinion Model (UOM) provides a realistic representation of user opinions. Results demonstrate that UOM maintains over 80% true influence against advanced misinformation, and DRIM outperforms state-of-the-art methods by up to 45% in influence and 77% in speed. DRIM also excels in limited-resource scenarios, networks with 10% invisibility, and when users are inclined to doubt true information.
Qi Zhang 0104, Lance M. Kaplan, Audun Jøsang, Dong Hyun Jeong, Feng Chen 0001, Jin-Hee Cho
IEEE Big Data2
2024 TeamCollab: A Framework for Collaborative Perception-Cognition-Communication-Action
abstract
Teams of embodied AI-enabled agents are critical for applications in extreme and highly dynamic environments. Developing robust controllers for such agents requires a deep understanding of the challenges encountered when attempting to coordinate and synchronize their individual perception-cognition-communication-action (PCCA) loops for team-wide mission objectives. We introduce a framework to explore the coordination of the PCCA loops across multiple agents in a new simulated physical environment designed to explore collaboration in each PCCA stage. This environment tasks teams of agents with the correct disposal of dangerous objects in an area and forces careful coordination of sensing, communication, movement, and manipulation actions by providing spatially-bounded communication, incorporating situations that require concerted effort by groups of agents, and introducing uncertainty into agents’ sensing capabilities. We provide a set of heuristic controllers, an offline oracle model, and an initial exploration of a Reward Machine-based controller that learns its policies from training. Together these approaches serve to provide insights into the complexity of the multi-agent PCCA loop coordination problem. The multiagent PCCA simulation environment, which supports AI and human-controlled agents, and the code for various agent controllers are available at https://github.com/nesl/AI-Collab.
Julian de Gortari Briseno, Roko Parac, Leo Ardon, Marc Roig Vilamala, Daniel Furelos-Blanco, Lance M. Kaplan, Vinod K. Mishra, Federico Cerutti 0001, Alun D. Preece, Alessandra Russo, Mani Srivastava 0001
FUSION6
2024 Neuro-Symbolic Fusion of Wi-Fi Sensing Data for Passive Radar with Inter-Modal Knowledge Transfer
abstract
Wi-Fi devices, akin to passive radars, can discern human activities within indoor settings due to the human body’s interaction with electromagnetic signals. Current Wi-Fi sensing applications predominantly employ data-driven learning techniques to associate the fluctuations in the physical properties of the communication channel with the human activity causing them. However, these techniques often lack the desired flexibility and transparency. This paper introduces DeepProbHAR, a neuro-symbolic architecture for Wi-Fi sensing, providing initial evidence that Wi-Fi signals can differentiate between simple movements, such as leg or arm movements, which are integral to human activities like running or walking. The neuro-symbolic approach affords gathering such evidence without needing additional specialised data collection or labelling. The training of DeepProbHAR is facilitated by declarative domain knowledge obtained from a camera feed and by fusing signals from various antennas of the Wi-Fi receivers. DeepProbHAR achieves results comparable to the state-of-the-art in human activity recognition. Moreover, as a by-product of the learning process, DeepProbHAR generates specialised classifiers for simple movements that match the accuracy of models trained on finely labelled datasets, which would be particularly costly.
Marco Cominelli, Francesco Gringoli, Lance M. Kaplan, Mani Srivastava 0001, Trevor J. Bihl, Erik Blasch, Nandini Iyer, Federico Cerutti 0001
FUSION3
2024 On Network Quickest Change Detection with Uncertain Models: An Experimental Study
abstract
We study the problem of Quickest Change Detection (QCD) in a complex networked system consisting of a set of heterogeneous agents that sequentially feed information to a central fusion center. At any unknown deterministic time, a persistent anomaly occurs, causing the distribution of observations from an unknown distinguishable subset of agents to simultaneously change from a nominal (pre-change) distribution to an anomalous (post-change) distribution, and the goal of the fusion center is to detect the change as quickly as possible subject to a false alarm constraint. Traditionally, various fusion rules have been proposed that assume that the distributions at each agent are either completely known or unknown and are locally solved using the Cumulative Sum (CuSum) and Generalized Likelihood Ratio (GLR) statistics, respectively. When an agent has access to training data, the Uncertain Likelihood Ratio (ULR) test generalizes distributional assumptions using uncertain distributions. However, the ULR has not been implemented for network change detection. This paper empirically studies incorporating the ULR statistics into the existing fusion rules for QCD and compares the average detection delay. Our results show that the ULR test can improve the average detection delay over the GLR tests using certain fusion techniques, while approaching the detection delay of the CuSum tests as the training data increases. Our results provide insights into future theoretical analysis to improve network QCD with imprecise knowledge of the distributions.
James Zachary Hare, Lance M. Kaplan, Venugopal V. Veeravalli
FUSION3
2024 Asymptotic Analysis of Uncertain Naïve Bayes via Second-Order Probabilities
abstract
Likelihood fusion is a special case of Bayesian networks known as naïve Bayes. It is well known that as the number of observations goes to infinity with known likelihoods, the aleatoric uncertainty of the queried (or parent) variable goes to zero, and furthermore, the declared values are guaranteed to match the ground truth. This work considers the case that the conditional probabilities are learned with limited training data leading to uncertain likelihoods, and second-order probabilistic reasoning is incorporated to characterize the aleatoric and epistemic uncertainty. Remarkably, it is shown that both the aleatoric and epistemic uncertainty goes to zero despite limited knowledge of the likelihoods. The rate of convergence is dictated by a quasi-divergence value that is related to the Kullback-Liebler (KL) divergence. However, the quasi-divergence can be negative leading to false declarations. This paper investigates when false declarations can emerge and shows how such cases diminish as the amount of training data for the likelihoods increases.
Lance M. Kaplan, James Zachary Hare
FUSION1
2024 Hyper Evidential Deep Learning to Quantify Composite Classification Uncertainty
abstract
Deep neural networks (DNNs) have been shown to perform well on exclusive, multi-class classification tasks. However, when different classes have similar visual features, it becomes challenging for human annotators to differentiate them. When an image is ambiguous, such as a blurry one where an annotator can't distinguish between a husky and a wolf, it may be labeled with both classes: {husky, wolf}. This scenario necessitates the use of composite set labels. In this paper, we propose a novel framework called Hyper-Evidential Neural Network (HENN) that explicitly models predictive uncertainty caused by composite set labels in training data in the context of the belief theory called Subjective Logic (SL). By placing a Grouped Dirichlet distribution on the class probabilities, we treat predictions of a neural network as parameters of hyper-subjective opinions and learn the network that collects both single and composite evidence leading to these hyper-opinions by a deterministic DNN from data. We introduce a new uncertainty type called vagueness originally designed for hyper-opinions in SL to quantify composite classification uncertainty for DNNs. Our experiments prove that HENN outperforms its state-of-the-art counterparts based on four image datasets. The code and datasets are available at: https://shorturl.at/dhoqx.
Changbin Li, Kangshuo Li, Yuzhe Ou, Lance M. Kaplan, Audun Jøsang, Jin-Hee Cho, Dong Hyun Jeong, Feng Chen 0001
ICLR4
2024 FlexLoc: Conditional Neural Networks for Zero-Shot Sensor Perspective Invariance in Object Localization with Distributed Multimodal Sensors
abstract
Localization is a critical technology for various applications ranging from navigation and surveillance to assisted living. Localization systems typically fuse information from sensors viewing the scene from different perspectives to estimate the target location while also employing multiple modalities for enhanced robustness and accuracy. Recently, such systems have employed end-to-end deep neural models trained on large datasets due to their superior performance and ability to handle data from diverse sensor modalities. However, such neural models are often trained on data collected from a particular set of sensor poses (i.e., locations and orientations). During real-world deployments, slight deviations from these sensor poses can result in extreme inaccuracies. To address this challenge, we introduce FlexLoc, which employs conditional neural networks to inject node perspective information to adapt the localization pipeline. Specifically, a small subset of model weights are derived from node poses at run time, enabling accurate generalization to unseen perspectives with minimal additional overhead. Our evaluations on a multimodal, multi-view indoor tracking dataset showcase that FlexLoc improves the localization accuracy by almost 50% in the zero-shot case (no calibration data available) compared to the baselines. The source code of FlexLoc is available in https://github.com/nesl/FlexLoc.
Ziqi Wang 0001, Xiaomin Ouyang, Ho Lyun Jeong, Colin Samplawski, Lance M. Kaplan, Benjamin M. Marlin, Mani Srivastava 0001
IROS6
2023 Accurate Passive Radar via an Uncertainty-Aware Fusion of Wi-Fi Sensing Data
abstract
Wi-Fi devices can effectively be used as passive radar systems that sense what happens in the surroundings and can even discern human activity. We propose, for the first time, a principled architecture which employs Variational Auto-Encoders for estimating a latent distribution responsible for generating the data, and Evidential Deep Learning for its ability to sense out-of-distribution activities. We verify that the fused data processed by different antennas of the same Wi-Fi receiver results in increased accuracy of human activity recognition compared with the most recent benchmarks, while still being informative when facing out-of-distribution samples and enabling semantic interpretation of latent variables in terms of physical phenomena. The results of this paper are a first contribution toward the ultimate goal of providing a flexible, semantic characterisation of black-swan events, i.e., events for which we have limited to no training data.
Marco Cominelli, Francesco Gringoli, Lance M. Kaplan, Mani Srivastava 0001, Federico Cerutti 0001
FUSION3
2023 Bounded Subjective Opinions
abstract
The shifted Dirichlet distribution is utilized to extend the definition of subjective opinions to include upper and lower bounds on belief and base rates. The characteristics of the bounded subjective opinions are examined and contrasted with unbounded subjective opinions.
Michael McDonald 0001, Lance M. Kaplan, Van Nguyen 0002
FUSION2
2023 Improved Small Sample Hypothesis Testing Using the Uncertain Likelihood Ratio
abstract
This paper studies the problem of event detection in highly dynamic and uncertain environments where training data is limited and the number of observations is small. The objective is to determine if the training data and observations are drawn from the same or different distributions. The traditional approach is to use the Generalized Likelihood Ratio (GLR) test statistic. However, with limited training data and observations, the GLR is suboptimal. To overcome this, we propose to use the Uncertain Likelihood Ratio (ULR) test statistic to accurately account for the uncertainty. We show that the ULR is the mean most powerful test and that it is asymptotically equivalent to the GLR under a Jeffreys Prior. Furthermore, a simulation study validates the theory that the performance of the ULR is superior to the GLR when testing in the small sample regime.
James Zachary Hare, Lance M. Kaplan
ICASSP2
2023 Multi-Label Temporal Evidential Neural Networks for Early Event Detection
abstract
Early event detection aims to detect events even before the event is complete. However, most of the existing methods focus on an event with a single label but fail to be applied to cases with multiple labels. Another non-negligible issue for early event detection is a prediction with overconfidence due to the high vacuity uncertainty that exists in the early time series. It results in an over-confidence estimation and hence unreliable predictions. To this end, technically, we propose a novel framework, Multi-Label Temporal Evidential Neural Network (MTENN), for multi-label uncertainty estimation in temporal data. MTENN is able to quality predictive uncertainty due to the lack of evidence for multi-label classifications at each time stamp based on belief/evidence theory. In addition, we introduce a novel uncertainty estimation head (weighted binomial comultiplication (WBC)) to quantify the fused uncertainty of a sub-sequence for early event detection. We validate the performance of our approach with state-of-the-art techniques on real-world audio datasets.
Xujiang Zhao, Xuchao Zhang, Chen Zhao 0010, Jin-Hee Cho, Lance M. Kaplan, Dong Hyun Jeong, Audun Jøsang, Feng Chen 0001
ICASSP5
2023 Detecting Intents of Fake News Using Uncertainty-Aware Deep Reinforcement Learning
abstract
Intent mining is critical for controlling the spread of false information across online social networks (OSNs). To this end, we develop deep reinforcement learning (DRL) agents guided by a delayed reward based on intent prediction using a classifier of long short-term memory (LSTM). Additionally, we incorporate an uncertainty-aware function that leverages subjective opinions derived from Subjective Logic (SL). Through evaluation using an annotated fake news tweet dataset, our results demonstrate that our intent classification framework surpasses competing methods in terms of intent accuracy. Our intent mining solutions using DRL algorithms can support effective and efficient intervention strategies for fake news spreading on OSNs.
Zhen Guo 0002, Qi Zhang 0104, Qisheng Zhang, Lance M. Kaplan, Audun Jøsang, Feng Chen 0001, Dong Hyun Jeong, Jin-Hee Cho
ICWS4
2023 DeepProbCEP: A neuro-symbolic approach for complex event processing in adversarial settings
abstract
Detecting complex events from subsymbolic data streams (such as images, audio recordings or videos) is a challenging problem, as traditional symbolic approaches cannot be used to process subsymbolic data, and neural-only approaches usually require larger amounts of training data than available. In this paper, we present DeepProbCEP, a Complex Event Processing (CEP) approach designed with four objectives: (i) allowing the use of subsymbolic data as an input, (ii) retaining flexibility and modularity in the definition of complex event rules, (iii) limiting the cost of obtaining training data and (iv) being robust against adversarial conditions. DeepProbCEP archives this by using a neuro-symbolic approach, which combines the neural and symbolic approaches to allow training with sparse data. This is made possible through the injection of human knowledge. In this paper, we demonstrate that DeepProbCEP outperforms other state-of-the-art approaches when training using sparse data. We also show that DeepProbCEP is robust in different adversarial settings. Finally, DeepProbCEP’s flexibility is demonstrated by showing it can be used to process both images and audio as input.
Marc Roig Vilamala, Tianwei Xing, Harrison Taylor, Luis Garcia 0001, Mani Srivastava 0001, Lance M. Kaplan, Alun D. Preece, Angelika Kimmig, Federico Cerutti 0001
Expert Syst. Appl.6
2023 AcTrak: Controlling a Steerable Surveillance Camera using Reinforcement Learning
abstract
Steerable cameras that can be controlled via a network, to retrieve telemetries of interest have become popular. In this paper, we develop a framework called AcTrak , to automate a camera’s motion to appropriately switch between (a) zoom ins on existing targets in a scene to track their activities, and (b) zoom out to search for new targets arriving to the area of interest. Specifically, we seek to achieve a good trade-off between the two tasks, i.e., we want to ensure that new targets are observed by the camera before they leave the scene, while also zooming in on existing targets frequently enough to monitor their activities. There exist prior control algorithms for steering cameras to optimize certain objectives; however, to the best of our knowledge, none have considered this problem, and do not perform well when target activity tracking is required. AcTrak automatically controls the camera’s PTZ configurations using reinforcement learning (RL ), to select the best camera position given the current state. Via simulations using real datasets, we show that AcTrak detects newly arriving targets 30% faster than a non-adaptive baseline and rarely misses targets, unlike the baseline which can miss up to 5% of the targets. We also implement AcTrak to control a real camera and demonstrate that in comparison with the baseline, it acquires about 2× more high resolution images of targets.
Abdulrahman Fahim, Evangelos E. Papalexakis, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury, Lance M. Kaplan, Tarek F. Abdelzaher
ACM Trans. Cyber Phys. Syst.5
2022 Uncertainty-Aware Quickest Change Detection: An Experimental Study
James Zachary Hare, Lance M. Kaplan
FUSION2
2022 SOLBP: Second-Order Loopy Belief Propagation for Inference in Uncertain Bayesian Networks
Conrad D. Hougen, Lance M. Kaplan, Magdalena Ivanovska, Federico Cerutti 0001, Kumar Vijay Mishra, Alfred O. Hero III
FUSION2
2022 Evidential Reasoning and Learning: a Survey
abstract
When collaborating with an artificial intelligence (AI) system, we need to assess when to trust its recommendations. Suppose we mistakenly trust it in regions where it is likely to err. In that case, catastrophic failures may occur, hence the need for Bayesian approaches for reasoning and learning to determine the confidence (or epistemic uncertainty) in the probabilities of the queried outcome. Pure Bayesian methods, however, suffer from high computational costs. To overcome them, we revert to efficient and effective approximations. In this paper, we focus on techniques that take the name of evidential reasoning and learning from the process of Bayesian update of given hypotheses based on additional evidence. This paper provides the reader with a gentle introduction to the area of investigation, the up-to-date research outcomes, and the open questions still left unanswered.
Federico Cerutti 0001, Lance M. Kaplan, Murat Sensoy
IJCAI2
2022 BigEye: Detection and Summarization of Key Global Events From Distributed Crowdsensed Data
abstract
Social media postings using smartphones (referred to as crowd-sensed data) can often facilitate real-time detection of key physical events in applications like disaster recovery or in smart cities. These postings also often contain visual content (e.g., images) that can be used to obtain zoomed-in views of such events. These crowd-sensed data are likely to be of large volume and distributed across a plurality of producers (e.g., cloudlets). Blindly transferring this large volume of raw data from the producers to a consumer will induce information overload and consume very high bandwidth. The problem is exacerbated in scenarios with limited bandwidth (e.g., after a disaster). In this article, we designBigEye, a novel framework that only transfers very limited data from the distributed producers to a central summarizer, and yet supports: 1) highly accurate detection and 2) concise visual summarization of key events of global interest. In realizingBigEye, we address several challenges including: 1) identifying events that have the highest global interest via the transfer of appropriate limited metadata from the producers to the summarizer; 2) reconciling metadata that could be inconsistent across the producers; and 3) the timely retrieval of visual summaries of the key events given bandwidth constraints. We show thatBigEyeachieves the same accuracy in detecting key events, as a system, where all data are available centrally while transferring only 1% of the raw data volume. Compared to the baseline approaches,BigEye’s parallelized transfer of visual content reduces the average delay by 67%.
Abdulrahman Fahim, Ajaya Neupane, Evangelos E. Papalexakis, Lance M. Kaplan, Srikanth V. Krishnamurthy, Tarek F. Abdelzaher
IEEE Internet Things J.4
2022 Handling epistemic and aleatory uncertainties in probabilistic circuits
Federico Cerutti 0001, Lance M. Kaplan, Angelika Kimmig, Murat Sensoy
Mach. Learn.2
2021 Towards Neural-Symbolic Learning to support Human-Agent Operations
Daniel Cunnington, Mark Law, Alessandra Russo, Jorge Lobo 0001, Lance M. Kaplan
FUSION5
2021 Toward Uncertainty Aware Quickest Change Detection
James Zachary Hare, Lance M. Kaplan, Venugopal V. Veeravalli
FUSION2
2021 Truth Discovery With Multi-Modal Data in Social Sensing
abstract
This article proposes unsupervised truth-finding algorithms that combine consideration of multi-modal content features with analysis of propagation patterns to evaluate the veracity of observations in social sensing applications. A key social sensing challenge is to develop effective algorithms for estimating both the reliability of sources and the veracity of their observations without prior knowledge. In contrast to prior solutions that use labeled examples to learn content features that are correlated with veracity, our approach is entirely unsupervised. Hence, given no prior training data, we jointly learn the importance of different content features together with the veracity of observations using propagation patterns as an indicator of perceived content reliability. A novel penalized expectation maximization (PEM) algorithm is proposed to improve the quality of estimation results for observations bolstered by multiple features. In addition, we develop a constrained expectation maximum likelihood with multiple features (CEM-MultiF) that introduces a novel constraint to boost the probability of correctness of some claims. Finally, we evaluate the performance of the proposed algorithms, called EM-Multi, CEM-Multi and PEM-MultiF, respectively, on real-world data sets collected from Twitter. The evaluation results demonstrate that the proposed algorithms outperform the existing fact-finding approaches, and offer tunable knobs for controlling robustness/performance trade-offs in the presence of malicious sources.
Huajie Shao, Dachun Sun, Shuochao Yao, Lu Su 0001, Zhibo Wang 0001, Dongxin Liu, Shengzhong Liu, Lance M. Kaplan, Tarek F. Abdelzaher
IEEE Trans. Computers8
2020 Uncertainty-Aware Deep Classifiers Using Generative Models
abstract
Deep neural networks are often ignorant about what they do not know and overconfident when they make uninformed predictions. Some recent approaches quantify classification uncertainty directly by training the model to output high uncertainty for the data samples close to class boundaries or from the outside of the training distribution. These approaches use an auxiliary data set during training to represent out-of-distribution samples. However, selection or creation of such an auxiliary data set is non-trivial, especially for high dimensional data such as images. In this work we develop a novel neural network model that is able to express both aleatoric and epistemic uncertainty to distinguish decision boundary and out-of-distribution regions of the feature space. To this end, variational autoencoders and generative adversarial networks are incorporated to automatically generate out-of-distribution exemplars for training. Through extensive analysis, we demonstrate that the proposed approach provides better estimates of uncertainty for in- and out-of-distribution samples, and adversarial examples on well-known data sets against state-of-the-art approaches including recent Bayesian approaches for neural networks and anomaly detection methods.
Murat Sensoy, Lance M. Kaplan, Federico Cerutti 0001, Maryam Saleki
AAAI2
2020 Second-Order Learning and Inference using Incomplete Data for Uncertain Bayesian Networks: A Two Node Example
abstract
Efficient second-order probabilistic inference in uncertain Bayesian networks was recently introduced. However, such second -order inference methods presume training over complete training data. While the expectation-maximization framework is well-established for learning Bayesian network parameters for incomplete training data, the framework does not determine the covariance of the parameters. This paper introduces two methods to compute the covariances for the parameters of Bayesian networks or Markov random fields due to incomplete data for two-node networks. The first method computes the covariances directly from the posterior distribution of parameters, and the second method more efficiently estimates the covariances from the Fisher information matrix. Finally, the implications and effectiveness of these covariances is theoretically and empirically evaluated.
Lance M. Kaplan, Federico Cerutti 0001, Murat Sensoy, Kumar Vijay Mishra
FUSION1
2020 Communication Constrained Learning with Uncertain Models
abstract
We consider the problem of distributed inference of a group of agents in a social network, where the agents construct, share, and update beliefs in a non-Bayesian framework to identify the underlying true state of the world. We build upon the concept of uncertain models that accurately represents each agents knowledge of the distribution of each hypothesis based on the amount of training data collected. Then, we propose an event-triggered communication protocol that only transmits a belief for a hypothesis if new information has been incorporated since the previous communication time. We show that the proposed solution allows the agents to achieve beliefs within the neighborhood of a full communication network, while significantly reducing the amount of transmissions.
James Zachary Hare, César A. Uribe, Lance M. Kaplan, Ali Jadbabaie
ICASSP3
2020 Neuroplex: learning to detect complex events in sensor networks through knowledge injection
abstract
Despite the remarkable success in a broad set of sensing applications, state-of-the-art deep learning techniques struggle with complex reasoning tasks across a distributed set of sensors. Unlike recognizing transient complex activities (e.g., human activities such as walking or running) from a single sensor, detecting more complex events with larger spatial and temporal dependencies across multiple sensors is extremely difficult, e.g., utilizing a hospital's sensor network to detect whether a nurse is following a sanitary protocol as they traverse from patient to patient. Training a more complicated model requires a larger amount of data-which is unrealistic considering complex events rarely happen in nature. Moreover, neural networks struggle with reasoning about serial, aperiodic events separated by large quantities in the spatial-temporal dimensions.
Tianwei Xing, Luis Garcia 0001, Marc Roig Vilamala, Federico Cerutti 0001, Lance M. Kaplan, Alun D. Preece, Mani Srivastava 0001
SenSys5
2019 Probabilistic Logic Programming with Beta-Distributed Random Variables
abstract
We enable aProbLog—a probabilistic logical programming approach—to reason in presence of uncertain probabilities represented as Beta-distributed random variables. We achieve the same performance of state-of-the-art algorithms for highly specified and engineered domains, while simultaneously we maintain the flexibility offered by aProbLog in handling complex relational domains. Our motivation is that faithfully capturing the distribution of probabilities is necessary to compute an expected utility for effective decision making under uncertainty: unfortunately, these probability distributions can be highly uncertain due to sparse data. To understand and accurately manipulate such probability distributions we need a well-defined theoretical framework that is provided by the Beta distribution, which specifies a distribution of probabilities representing all the possible values of a probability when the exact value is unknown.
Federico Cerutti 0001, Lance M. Kaplan, Angelika Kimmig, Murat Sensoy
AAAI2
2019 On Malicious Agents in Non-Bayesian Social Learning with Uncertain Models
James Zachary Hare, César A. Uribe, Lance M. Kaplan, Ali Jadbabaie
FUSION3
2019 Unsupervised Fact-finding with Multi-modal Data in Social Sensing
Huajie Shao, Shuochao Yao, Yiran Zhao 0001, Lu Su 0001, Zhibo Wang 0001, Dongxin Liu, Shengzhong Liu, Lance M. Kaplan, Tarek F. Abdelzaher
FUSION8
2019 Edge-Assisted Detection and Summarization of Key Global Events from Distributed Crowd-Sensed Data
abstract
This paper introduces a novel service for distributed detection and summarization of crowd-sensed events. The work is motivated by the proliferation of microblogging media, such as Twitter, that can be used to detect and describe events in the physical world, such as protests, disasters, or civil unrest. Since crowd-sensed data is likely to be distributed, we consider an architecture, where the data first accumulates across a plurality of edge servers (e.g. cloudlets or repositories) and is then summarized, rather than being shipped directly to its ultimate destination (e.g., in a remote cloud). The architecture allows graceful handling of overload and bandwidth limitations (e.g., in scenarios where capacity is impaired, as the case might be after a disaster). When bandwidth is scarce, our service, BigEye, only transfers very limited metadata from the distributed edge repositories to the central summarizer and yet supports highly accurate detection and concise summarization of key events of global interest. These summaries can then be sent to consumers (e.g., rescue personnel). Our emulations show that BigEye achieves the same precision and recall values in detecting key events as a system where all data is available centrally, while consuming only 1% of the bandwidth needed to transmit all raw data.
Abdelrahman Fahim, Ajaya Neupane, Evangelos E. Papalexakis, Lance M. Kaplan, Srikanth V. Krishnamurthy, Tarek F. Abdelzaher
IC2E4
2019 Spherical Text Embedding
abstract
Unsupervised text embedding has shown great power in a wide range of NLP tasks. While text embeddings are typically learned in the Euclidean space, directional similarity is often more effective in tasks such as word similarity and document clustering, which creates a gap between the training stage and usage stage of text embedding. To close this gap, we propose a spherical generative model based on which unsupervised word and paragraph embeddings are jointly learned. To learn text embeddings in the spherical space, we develop an efficient optimization algorithm with convergence guarantee based on Riemannian optimization. Our model enjoys high efficiency and achieves state-of-the-art performances on various text embedding tasks including word similarity and document clustering.
Yu Meng 0001, Jiaxin Huang 0001, Guangyuan Wang, Chao Zhang 0014, Honglei Zhuang, Lance M. Kaplan, Jiawei Han 0001
NeurIPS6
2019 DeepCEP: Deep Complex Event Processing Using Distributed Multimodal Information
abstract
Deep learning models typically make inferences over transient features of the latent space, i.e., they learn data representations to make decisions based on the current state of the inputs over short periods of time. Such models would struggle with state-based events, or complex events, that are composed of simple events with complex spatial and temporal dependencies. In this paper, we propose DeepCEP, a framework that integrates the concepts of deep learning models with complex event processing engines to make inferences across distributed, multimodal information streams with complex spatial and temporal dependencies. DeepCEP utilizes deep learning to detect primitive events. A user can define a complex event to be detected as a particular sequence or pattern of primitive events as well as any other logical predicates that constrain the definition of such an event. The integration of human logic not only increases robustness and interpretability, but also greatly reduces the amount of training data required. Further, we demonstrate how the uncertainty of a model can be propagated throughout the complex event detection pipeline. Finally, we enumerate the future directions of research enabled by DeepCEP. In particular, we detail how an end-to-end training model for complex event processing with deep learning may be realized.
Tianwei Xing, Marc Roig Vilamala, Luis Garcia 0001, Federico Cerutti 0001, Lance M. Kaplan, Alun D. Preece, Mani Srivastava 0001
SMARTCOMP5
2019 A Semi-Supervised Active-learning Truth Estimator for Social Networks
abstract
This paper introduces an active-learning-based truth estimator for social networks, such as Twitter, that enhances estimation accuracy significantly by requesting a well-selected (small) fraction of data to be labeled. Data assessment and truth discovery from arbitrary open online sources are a hard problem due to uncertainty regarding source reliability. Multiple truth finding systems were developed to solve this problem. Their accuracy is limited by the noisy nature of the data, where distortions, fabrications, omissions, and duplication are introduced. This paper presents a semi-supervised truth estimator for social networks, in which a portion of inputs are carefully selected to be reliably verified. The challenge is to find the subset of observations to verify that would maximally enhance the overall fact-finding accuracy. This work extends previous passive approaches to recursive truth estimation, as well as semi-supervised approaches where the estimator has no control over the choice of data to be labeled. Results show that by optimally selecting claims to be verified, we improve estimated accuracy by 12% over unsupervised baseline, and by 5% over previous semi-supervised approaches.
Hang Cui 0001, Tarek F. Abdelzaher, Lance M. Kaplan
WWW3
2018 Recursive Truth Estimation of Time-Varying Sensing Data from Online Open Sources
abstract
This paper is motivated by prospective Internet of Things (IoT) applications that exploit inputs from online open sources whose reliability may be uncertain. Unlike physical signal fusion (that can leverage solid analytic foundations derived from physical properties of fused signals), data reliability assessment from arbitrary online open sources is a harder problem. At least two difficulties arise. First, source reliability is harder to estimate from first principles due to lack of visibility into the sensing and subsequent processing stages for published data. Second, by virtue of being open, some sources can be copied by others, leading to correlated errors at a large scale. This paper presents a recursive truth estimator for online public data streams that addresses the above two problems. We focus on categorical data. Many truth-finding systems were developed to cope with unreliable categorical data. Most of them are designed for batch analysis of bulk datasets. This work extends previous efforts by developing an online recursive estimator. Unlike previous recursive fact-finders, ours is the first that can jointly handle (i) changes in the population of sources over time, (ii) changes in the ground-truth state of the physical phenomenon being observed (that result in the appearance of conflicting claims), and (iii) correlated errors due to potential copying among sources. Results show that our algorithm not only outperforms other recursive fact-finders in the case of changing ground-truth state, but also improves estimation accuracy of static state.
Hang Cui 0001, Tarek F. Abdelzaher, Lance M. Kaplan
DCOSS3
2018 Learning and Reasoning in Complex Coalition Information Environments: A Critical Analysis
abstract
In this paper we provide a critical analysis with metrics that will inform guidelines for designing distributed systems for Collective Situational Understanding (CSU). CSU requires both collective insight-i.e., accurate and deep understanding of a situation derived from uncertain and often sparse data and collective foresight-i.e., the ability to predict what will happen in the future. When it comes to complex scenarios, the need for a distributed CSU naturally emerges, as a single monolithic approach not only is unfeasible: it is also undesirable. We therefore propose a principled, critical analysis of AI techniques that can support specific tasks for CSU to derive guidelines for designing distributed systems for CSU.
Federico Cerutti 0001, Moustafa Farid Alzantot, Tianwei Xing, Dan Harborne, Jonathan Z. Bakdash, Dave Braines, Supriyo Chakraborty, Lance M. Kaplan, Angelika Kimmig, Alun D. Preece, Ramya Raghavendra, Murat Sensoy, Mani Srivastava 0001
FUSION8
2018 Trust Estimation of Sources Over Correlated Propositions
abstract
This work analyzes the impact of correlated propositions when estimating the reporting behavior of information sources. These behavior estimates are critical for fusion, and traditional methods assume the propositions are statistically independent. A new source behavior estimation methods is presented that accounts for statistical dependencies between the training propositions. Simulations seem to indicate that the potential performance gains for accounting for the correlations is small relative to the increased computational complexity. One may conclude that the traditional independence assumption in source behavior estimation methods is reasonable even in cases where it is actually violated.
Lance M. Kaplan, Murat Sensoy
FUSION1
2018 Doc2Cube: Allocating Documents to Text Cube Without Labeled Data
abstract
Data cube is a cornerstone architecture in multidimensional analysis of structured datasets. It is highly desirable to conduct multidimensional analysis on text corpora with cube structures for various text-intensive applications in healthcare, business intelligence, and social media analysis. However, one bottleneck to constructing text cube is to automatically put millions of documents into the right cube cells so that quality multidimensional analysis can be conducted afterwards-it is too expensive to allocate documents manually or rely on massively labeled data. We propose Doc2Cube, a method that constructs a text cube from a given text corpus in an unsupervised way. Initially, only the label names (e.g., USA, China) of each dimension (e.g., location) are provided instead of any labeled data. Doc2Cube leverages label names as weak supervision signals and iteratively performs joint embedding of labels, terms, and documents to uncover their semantic similarities. To generate joint embeddings that are discriminative for cube construction, Doc2Cube learns dimension-tailored document representations by selectively focusing on terms that are highly label-indicative in each dimension. Furthermore, Doc2Cube alleviates label sparsity by propagating the information from label names to other terms and enriching the labeled term set. Our experiments on real data demonstrate the superiority of Doc2Cube over existing methods.
Fangbo Tao, Chao Zhang 0014, Xiusi Chen, Meng Jiang 0001, Tim Hanratty, Lance M. Kaplan, Jiawei Han 0001
ICDM6
2018 A Constrained Maximum Likelihood Estimator for Unguided Social Sensing
abstract
This paper develops a constrained expectation maximization algorithm (CEM) that improves the accuracy of truth estimation in unguided social sensing applications. Unguided social sensing refers to the act of leveraging naturally occurring observations on social media as “sensor measurements”, when the sources post at will and not in response to specific sensing campaigns or surveys. A key challenge in social sensing, in general, lies in estimating the veracity of reported observations, when the sources reporting these observations are of unknown reliability and their observations themselves cannot be readily verified. This problem is known as fact-finding. Unsupervised solutions have been proposed to the fact-finding problem that explore notions of internal data consistency in order to estimate observation veracity. This paper observes that unguided social sensing gives rise to a new (and very simple) constraint that dramatically reduces the space of feasible fact-finding solutions, hence significantly improving the quality of fact-finding results. The constraint relies on a simple approximate test of source independence, applicable to unguided sensing, and incorporates information about the number of independent sources of an observation to constrain the posterior estimate of its probability of correctness. Two different approaches are developed to test the independence of sources for purposes of applying this constraint, leading to two flavors of the CEM algorithm, we call CEM and CEM-Jaccard. We show using both simulation and real data sets collected from Twitter that by forcing the algorithm to converge to a solution in which the constraint is satisfied, the quality of solutions is significantly improved.
Huajie Shao, Shuochao Yao, Yiran Zhao 0001, Chao Zhang 0014, Jinda Han, Lance M. Kaplan, Lu Su 0001, Tarek F. Abdelzaher
INFOCOM6
2018 Evidential Deep Learning to Quantify Classification Uncertainty
abstract
Deterministic neural nets have been shown to learn effective predictors on a wide range of machine learning problems. However, as the standard approach is to train the network to minimize a prediction loss, the resultant model remains ignorant to its prediction confidence. Orthogonally to Bayesian neural nets that indirectly infer prediction uncertainty through weight uncertainties, we propose explicit modeling of the same using the theory of subjective logic. By placing a Dirichlet distribution on the class probabilities, we treat predictions of a neural net as subjective opinions and learn the function that collects the evidence leading to these opinions by a deterministic neural net from data. The resultant predictor for a multi-class classification problem is another Dirichlet distribution whose parameters are set by the continuous output of a neural net. We provide a preliminary analysis on how the peculiarities of our new loss function drive improved uncertainty estimation. We observe that our method achieves unprecedented success on detection of out-of-distribution queries and endurance against adversarial perturbations.
Murat Sensoy, Lance M. Kaplan, Melih Kandemir
NeurIPS2
2018 AspEm: Embedding Learning by Aspects in Heterogeneous Information Networks
abstract
Heterogeneous information networks (HINs) are ubiquitous in real-world applications. Due to the heterogeneity in HINs, the typed edges may not fully align with each other. In order to capture the semantic subtlety, we propose the concept of aspects with each aspect being a unit representing one underlying semantic facet. Meanwhile, network embedding has emerged as a powerful method for learning network representation, where the learned embedding can be used as features in various downstream applications. Therefore, we are motivated to propose a novel embedding learning framework-ASPEM-to preserve the semantic information in HINs based on multiple aspects. Instead of preserving information of the network in one semantic space, ASPEM encapsulates information regarding each aspect individually. In order to select aspects for embedding purpose, we further devise a solution for ASPEM based on dataset-wide statistics. To corroborate the efficacy of ASPEM, we conducted experiments on two real-words datasets with two types of applications-classification and link prediction. Experiment results demonstrate that ASPEM can outperform baseline network embedding learning methods by considering multiple aspects, where the aspects can be selected from the given HIN in an unsupervised manner.
Yu Shi 0002, Huan Gui, Qi Zhu 0008, Lance M. Kaplan, Jiawei Han 0001
SDM4
2018 Efficient belief propagation in second-order Bayesian networks for singly-connected graphs
Lance M. Kaplan, Magdalena Ivanovska
Int. J. Approx. Reason.1
2018 GeoBurst+: Effective and Real-Time Local Event Detection in Geo-Tagged Tweet Streams
abstract
The real-time discovery of local events (e.g., protests, disasters) has been widely recognized as a fundamental socioeconomic task. Recent studies have demonstrated that the geo-tagged tweet stream serves as an unprecedentedly valuable source for local event detection. Nevertheless, how to effectively extract local events from massive geo-tagged tweet streams in real time remains challenging. To bridge the gap, we propose a method for effective and real-time local event detection from geo-tagged tweet streams. Our method, named G eo B urst+ , first leverages a novel cross-modal authority measure to identify several pivots in the query window. Such pivots reveal different geo-topical activities and naturally attract similar tweets to form candidate events. G eo B urst+ further summarizes the continuous stream and compares the candidates against the historical summaries to pinpoint truly interesting local events. Better still, as the query window shifts, G eo B urst+ is capable of updating the event list with little time cost, thus achieving continuous monitoring of the stream. We used crowdsourcing to evaluate G eo B urst+ on two million-scale datasets and found it significantly more effective than existing methods while being orders of magnitude faster.
Chao Zhang 0014, Dongming Lei, Quan Yuan 0001, Honglei Zhuang, Lance M. Kaplan, Shaowen Wang 0001, Jiawei Han 0001
ACM Trans. Intell. Syst. Technol.5
2018 A Fundamental Limitation on Maximum Parameter Dimension for Accurate Estimation With Quantized Data
abstract
It is revealed that there is a link between the quantization approach employed and the dimension of the vector parameter which can be accurately estimated by a quantized estimation system. A critical quantity called inestimable dimension for quantized data (IDQD) is introduced, which does not depend on the quantization regions and the statistical models of the observations but instead depends only on the number of sensors and on the precision of the vector quantizers employed by the system. It is shown that the IDQD describes a quantization-induced fundamental limitation on the estimation capabilities of the system. To be specific, if the dimension of the desired vector parameter is larger than the IDQD of the quantized estimation system, then the Fisher information matrix for estimating the desired vector parameter is singular, and, moreover, there exist infinitely many nonidentifiable vector parameter points in the vector parameter space. Furthermore, it is shown that under some common assumptions on the statistical models of the observations and the quantization system, a smaller IDQD can be obtained, which can specify an even more limiting quantization induced fundamental limitation on the estimation capabilities of the system.
Jiangfan Zhang, Rick S. Blum, Lance M. Kaplan, Xuanxuan Lu
IEEE Trans. Inf. Theory3
2018 Towards Quality Aware Information Integration in Distributed Sensing Systems
abstract
In this paper, we present GDA, a generalized decision aggregation framework that integrates information from distributed sensor nodes for decision making in a resource efficient manner. Different from traditional approaches, our proposed GDA framework is able to not only estimate the reliability of each sensor, but also take advantage of its confidence information, and thus achieves higher decision accuracy. Targeting generalized problem domains, our framework can naturally handle the scenarios where different sensor nodes observe different sets of events whose numbers of possible classes may also be different. GDA also makes no assumption about the availability level of ground truth label information, while being able to take advantage of any if present. For these reasons, our approach can be applied to a much broader spectrum of sensing scenarios. In this paper, we also propose two extensions of the GDA framework, i.e., incremental GDA (I-GDA) and parallel GDA (P-GDA) to deal with streaming and large-scale data. The advantages of our proposed methods are demonstrated through both theoretic analysis and extensive experiments.
Chenglin Miao, Lu Su 0001, Qi Li 0012, Shaohan Hu, Shiguang Wang, Jing Gao 0004, Hengchang Liu, Tarek F. Abdelzaher, Jiawei Han 0001, Xue (Steve) Liu, Yan Gao 0010, Lance M. Kaplan
IEEE Trans. Parallel Distributed Syst.13
2017 Graphical approach for influence maximization in social networks under generic threshold-based non-submodular model
abstract
As a widely observable social effect, influence diffusion refers to a process where innovations, trends, awareness, etc. spread across the network via the social impact among individuals. Motivated by such social effect, the concept of influence maximization is coined, where the goal is to select a bounded number of the most influential nodes (seed nodes) from a social network so that they can jointly trigger the maximal influence diffusion. A rich body of research in this area is performed under statistical diffusion models with provable submodularity, which essentially simplifies the problem as the optimal result can be approximated by the simple greedy search. When the diffusion models are non-submodular, however, the research community mostly focuses on how to bound/approximate them by tractable submodular functions. In other words, there is still a lack of efficient methods that can directly resolve non-submodular influence maximization problems. In this regard, we fill the gap by proposing seed selection strategies using network graphical properties in a generalized non-submodular threshold-based model, called influence barricade model. Under this model, we first establish theories to reveal graphical conditions that ensure the network generated by node removals has the same optimal seed set as that in the original network. We then exploit these theoretical conditions to develop efficient algorithms by strategically removing less-important nodes and selecting seeds only in the remaining network. To the best of our knowledge, this is the first graph-based approach that directly tackles non-submodular influence maximization. Evaluations on both synthetic and real-world Facebook/Twitter datasets confirm the superior efficiency of the proposed algorithms, which are orders of magnitude faster than benchmarks for large networks.
Guohong Cao, Lance M. Kaplan
IEEE BigData3
2017 Identifying Semantically Deviating Outlier Documents
abstract
A document outlier is a document that substantially deviates in semantics from the majority ones in a corpus.Automatic identification of document outliers can be valuable in many applications, such as screening health records for medical mistakes.In this paper, we study the problem of mining semantically deviating document outliers in a given corpus.We develop a generative model to identify frequent and characteristic semantic regions in the word embedding space to represent the given corpus, and a robust outlierness measure which is resistant to noisy content in documents.Experiments conducted on two real-world textual data sets show that our method can achieve an up to 135% improvement over baselines in terms of recall at top-1% of the outlier ranking.
Honglei Zhuang, Chi Wang 0001, Fangbo Tao, Lance M. Kaplan, Jiawei Han 0001
EMNLP4
2017 Joint tracking and power level estimation of multiple targets using a proximity sensor network
abstract
This work documents our investigation of multiple target tracking filters in proximity sensor networks when the target power levels are not known. The challenge is that when the targets are close, it is hard to determine if the sensor reports are the results of a loud target or multiple quiet targets. Given the binary measurements:1 for detection of targets and 0 for nondetection of targets, the works studies the feasibility of using the joint multitarget particle filter to estimate the number of targets, and at the same time estimate the 5D target states including the positions, velocities and power levels.
Qiang Le, Lance M. Kaplan
FUSION2
2017 Cyber attacks on estimation sensor networks and iots: Impact, mitigation and implications to unattacked systems
abstract
Estimation of an unknown deterministic vector from quantized sensor data is considered in the presence of spoofing and man-in-the-middle attacks. First, asymptotically optimum processing, which identifies and categorizes the attacked sensors into different groups according to distinct types of attacks, is outlined in the face of man-in-the-middle attacks. Necessary and sufficient conditions are provided under which utilizing the attacked sensor data will lead to better estimation performance when compared to approaches where the attacked sensors are ignored. Next, necessary and sufficient conditions are provided under which spoofing attacks provide a guaranteed attack performance in terms of the Cramer-Rao Bound regardless of the processing the estimation system employs. It is shown that it is always possible to construct such a highly desirable attack by properly employing an attack vector parameter having a sufficiently large dimension relative to the number of quantization levels employed, which was not observed previously. For unattacked quantized estimation systems, a general limitation on the dimension of a vector parameter which can be accurately estimated is uncovered.
Jiangfan Zhang, Rick S. Blum, Lance M. Kaplan
ICASSP3
2017 Energy Efficient Object Detection in Camera Sensor Networks
abstract
A wireless camera network can provide situation awareness information (e.g., humans in distress) in scenarios such as disaster recovery. If such camera sensors are battery operated, sending raw video feeds back to a central controller can be expensive in terms of energy consumption. Further, if all cameras were to use the optimal processing algorithm for object decision, they may also expend unnecessary energy. Stated otherwise, cameras that capture the same objects may not all have to use the optimal algorithm to achieve a desired accuracy, and this can save processing energy costs. In this paper, our objective is to design and implement a framework that can support coordination among cameras to deliver highly accurate detection of objects in an energy efficient way. The framework, which we call EECS (for energy efficient camera sensors), estimates the detection accuracy and energy costs incurred (both the processing and communication costs are taken into account) with each detection algorithm for each camera, and comes up with a choice of cameras for sending information pertaining to the object of interest. This set of cameras and the video processing algorithms that they must use, are chosen so as to minimize the energy expenditures, given a desired detection accuracy. We implement EECS on a camera network built with smartphones, and demonstrate that it reduces the energy consumption by up to 40% while ensuring a object detection accuracy of over 86%.
Tuan Dao, Karim Khalil, Amit K. Roy-Chowdhury, Srikanth V. Krishnamurthy, Lance M. Kaplan
ICDCS5
2017 Optimizing Source Selection in Social Sensing in the Presence of Influence Graphs
abstract
This paper addresses the problem of choosing the right sources to solicit data from in sensing applications involving broadcast channels, such as those crowdsensing applications where sources share their observations on social media. The goal is to select sources such that expected fusion error is minimized. We assume that soliciting data from a source incurs a cost and that the cost budget is limited. Contrary to other formulations of this problem, we focus on the case where some sources influence others. Hence, asking a source to make a claim affects the behavior of other sources as well, according to an influence model. The paper makes two contributions. First, we develop an analytic model for estimating expected fusion error, given a particular influence graph and solution to the source selection problem. Second, we use that model to search for a solution that minimizes expected fusion error, formulating it as a zero-one integer non-linear programming (INLP) problem. To scale the approach, the paper further proposes a novel reliability-based pruning heuristic (RPH) and a similarity-based lossy estimation (SLE) algorithm that significantly reduce the complexity of the INLP algorithm at the cost of a modest approximation. The analytically computed expected fusion error is validated using both simulations and real-world data from Twitter, demonstrating a good match between analytic predictions and empirical measurements. It is also shown that our method outperforms baselines in terms of resulting fusion error.
Huajie Shao, Shiguang Wang, Shen Li 0002, Shuochao Yao, Yiran Zhao 0001, Md. Tanvir Al Amin, Tarek F. Abdelzaher, Lance M. Kaplan
ICDCS8
2017 Unveiling polarization in social networks: A matrix factorization approach
abstract
This paper presents unsupervised algorithms to uncover polarization in social networks (namely, Twitter) and identify polarized groups. The approach is language-agnostic and thus broadly applicable to global and multilingual media. In cases of conflict, dispute, or situations involving multiple parties with contrasting interests, opinions get divided into different camps. Previous manual inspection of tweets has shown that such situations produce distinguishable signatures on Twitter, as people take sides leading to clusters that preferentially propagate information confirming their individual cluster-specific bias. We propose a model for polarized social networks, and show that approaches based on factorizing the matrix of sources and their claims can automate the discovery of polarized clusters with no need for prior training or natural language processing. In turn, identifying such clusters offers insights into prevalent social conflicts and helps automate the generation of less biased descriptions of ongoing events. We evaluate our factorization algorithms and their results on multiple Twitter datasets involving polarization of opinions, demonstrating the efficacy of our approach. Experiments show that our method is almost always correct in identifying the polarized information from real-world twitter traces, and outperforms the baseline mechanisms by a large margin.
Md. Tanvir Al Amin, Charu C. Aggarwal, Shuochao Yao, Tarek F. Abdelzaher, Lance M. Kaplan
INFOCOM5
2017 On localizing urban events with Instagram
abstract
This paper develops an algorithm that exploits picture-oriented social networks to localize urban events. We choose picture-oriented networks because taking a picture requires physical proximity, thereby revealing the location of the photographed event. Furthermore, most modern cell phones are equipped with GPS, making picture location, and time metadata commonly available. We consider Instagram as the social network of choice and limit ourselves to urban events (noting that the majority of the world population lives in cities). The paper introduces a new adaptive localization algorithm that does not require the user to specify manually tunable parameters. We evaluate the performance of our algorithm for various real-world datasets, comparing it against a few baseline methods. The results show that our method achieves the best recall, the fewest false positives, and the lowest average error in localizing urban events.
Prasanna Giridhar, Shiguang Wang, Tarek F. Abdelzaher, Raghu K. Ganti, Lance M. Kaplan, Jemin George
INFOCOM5
2017 MetaPAD: Meta Pattern Discovery from Massive Text Corpora
abstract
Mining textual patterns in news, tweets, papers, and many other kinds of text corpora has been an active theme in text mining and NLP research. Previous studies adopt a dependency parsing-based pattern discovery approach. However, the parsing results lose rich context around entities in the patterns, and the process is costly for a corpus of large scale. In this study, we propose a novel typed textual pattern structure, called meta pattern, which is extended to a frequent, informative, and precise subsequence pattern in certain context. We propose an efficient framework, called MetaPAD, which discovers meta patterns from massive corpora with three techniques: (1) it develops a context-aware segmentation method to carefully determine the boundaries of patterns with a learnt pattern quality assessment function, which avoids costly dependency parsing and generates high-quality patterns; (2) it identifies and groups synonymous meta patterns from multiple facets---their types, contexts, and extractions; and (3) it examines type distributions of entities in the instances extracted by each group of patterns, and looks for appropriate type levels to make discovered patterns precise. Experiments demonstrate that our proposed framework discovers high-quality typed textual patterns efficiently from different genres of massive corpora and facilitates information extraction.
Meng Jiang 0001, Jingbo Shang, Taylor Cassidy, Xiang Ren 0001, Lance M. Kaplan, Tim Hanratty, Jiawei Han 0001
KDD5
2017 Accurate and Timely Situation Awareness Retrieval from a Bandwidth Constrained Camera Network
abstract
Wireless cameras can be used to gather situation awareness information (e.g., humans in distress) in disaster recovery scenarios. However, blindly sending raw video streams from such cameras, to an operations center or controller can be prohibitive in terms of bandwidth. Further, these raw streams could contain either redundant or irrelevant information. Thus, we ask "how do we extract accurate situation awareness information from such camera nodes and send it in a timely manner, back to the operations center?" Towards this, we design ACTION, a framework that (a) detects objects of interest (e.g., humans) from the video streams, (b) combines these streams intelligently to eliminate redundancies and (c) transmits only parts of the feeds that are sufficient in achieving a desired detection accuracy to the controller. ACTION uses small amounts of metadata to determine if the objects from different camera feeds are the same. A resource-aware greedy algorithm is used to select a subset of video feeds that are associated with the same object, so as to provide a desired accuracy, for being sent to the operations center. Our evaluations show that ACTION helps reduce the network usage up to threefold, and yet achieves a high detection accuracy of ≈ 90%.
Tuan Dao, Amit K. Roy-Chowdhury, Nasser M. Nasrabadi, Srikanth V. Krishnamurthy, Prasant Mohapatra, Lance M. Kaplan
MASS6
2017 ClariSense+: An enhanced traffic anomaly explanation service using social network feeds
Prasanna Giridhar, Md. Tanvir Al Amin, Tarek F. Abdelzaher, Dong Wang 0002, Lance M. Kaplan, Jemin George, Raghu K. Ganti
Pervasive Mob. Comput.5
2017 Embedding Learning with Events in Heterogeneous Information Networks
abstract
In real-world applications, objects of multiple types are interconnected, formingHeterogeneous Information Networks. In such heterogeneous information networks, we make the key observation that many interactions happen due to someeventand the objects in each event form a complete semantic unit. By taking advantage of such a property, we propose a generic framework calledHyperEdge-BasedEmbedding(Hebe) to learn object embeddings with events in heterogeneous information networks, where ahyperedgeencompasses the objects participating in one event. TheHebeframework models the proximity among objects in each event with two methods: (1) predicting a target object given other participating objects in the event, and (2) predicting if the event can be observed given all the participating objects. Since each hyperedge encapsulates more information of a given event,Hebeis robust to data sparseness and noise. In addition,Hebeis scalable when the data size spirals. Extensive experiments on large-scale real-world datasets show the efficacy and robustness of the proposed framework.
Huan Gui, Fangbo Tao, Meng Jiang 0001, Brandon Norick, Lance M. Kaplan, Jiawei Han 0001
IEEE Trans. Knowl. Data Eng.6
2016 On predicting social unrest using social media
abstract
We study the possibility of predicting a social protest (planned, or unplanned) based on social media messaging. We consider the process called mobilization, described in the literature as the precursor of participation. Mobilization includes four stages: being sympathetic to the cause, being aware of the movement, motivation to take part and ability to participate. We suggest that expressions of mobilization in communications of individuals may be used to predict the approaching protest. We have utilized several Natural Language Processing techniques to create a methodology to identify mobilization in social media communication. Results of experimentation with Twitter data collected before and during the 2015 Baltimore events and the information on actual protests taken from news media show a correlation over time between volume of Twitter communications related to mobilization and occurrences of protest at certain geographical locations. We conclude with discussion of possible theoretical explanations and practical applications of these results.
Rostyslav Korolov, Di Lu 0003, Claire Bonial, Clare R. Voss, Lance M. Kaplan, William A. Wallace, Jiawei Han 0001, Heng Ji 0001
ASONAM7
2016 Principles of subjective networks
Audun Jøsang, Lance M. Kaplan
FUSION2
2016 Efficient Subjective Bayesian network belief propagation for trees
Lance M. Kaplan, Magdalena Ivanovska
FUSION1
2016 Source behavior discovery for fusion of subjective opinions
Murat Sensoy, Lance M. Kaplan, Geeth de Mel, Taha D. Gunes
FUSION2
2016 On Source Dependency Models for Reliable Social Sensing: Algorithms and Fundamental Error Bounds
abstract
This paper develops a simplified dependency model for sources on social networks that is shown to improve the quality of fact-finding -- assessing veracity of observations shared on social media. Recent literature developed a mathematical approach for exploiting social networks, such as Twitter, as noisy sensor networks that report observations on the state of the physical world. It was shown that the quality of state estimation from such noisy data, known as fact-finding, was a function of assumptions made regarding the independence of sources or lack thereof. When sources propagate information they hear from others (without verification), correlated errors may arise that degrade fact-finding performance. This work advances the state of the art by developing a simplified model of dependencies between sources and designing an improved dependency-aware estimator to assess veracity of observations, taking into account the observed dependency structure. A fundamental error bound is derived for this estimator to understand the gap in its performance from optimal. It is shown that the new estimator outperforms state of the art fact-finders and, in some cases, yields an accuracy close to the fundamental error bound.
Shuochao Yao, Shaohan Hu, Shen Li 0002, Yiran Zhao 0001, Lu Su 0001, Lance M. Kaplan, Aylin Yener, Tarek F. Abdelzaher
ICDCS6
2016 Recursive Ground Truth Estimator for Social Data Streams
abstract
The paper develops a recursive state estimator for social network data streams that allows exploitation of social networks, such as Twitter, as sensor networks to reliably observe physical events. Recent literature suggested using social networks as sensor networks leveraging the fact that much of the information upload on the former constitutes acts of sensing. A significant challenge identified in that context was that source reliability is often unknown, leading to uncertainty regarding the veracity of reported observations. Multiple truth finding systems were developed to solve this problem, generally geared towards batch analysis of offline datasets. This work complements the present batch approaches by developing an online recursive state estimator that recovers ground truth from streaming data. In this paper, we model physical world state by a set of binary signals (propositions, called assertions, about world state) and the social network as a noisy medium, where distortion, fabrication, omissions, and duplication are introduced. Our recursive state estimator is designed to recover the original binary signal (the true propositions) from the received noisy signal, essentially decoding the unreliable social network output to obtain the best estimate of ground truth in the physical world. Results show that the estimator is both effective and efficient at recovering the original signal with a high degree of accuracy. The estimator gives rise to a novel situation awareness tool that can be used for reliably following unfolding events in real time, using dynamically arriving social network data.
Shuochao Yao, Md. Tanvir Al Amin, Lu Su 0001, Shaohan Hu, Shen Li 0002, Shiguang Wang, Yiran Zhao 0001, Tarek F. Abdelzaher, Lance M. Kaplan, Charu C. Aggarwal, Aylin Yener
IPSN9
2016 From Truth Discovery to Trustworthy Opinion Discovery: An Uncertainty-Aware Quantitative Modeling Approach
abstract
In this era of information explosion, conflicts are often encountered when information is provided by multiple sources. Traditional truth discovery task aims to identify the truth the most trustworthy information, from conflicting sources in different scenarios. In this kind of tasks, truth is regarded as a fixed value or a set of fixed values. However, in a number of real-world cases, objective truth existence cannot be ensured and we can only identify single or multiple reliable facts from opinions. Different from traditional truth discovery task, we address this uncertainty and introduce the concept of trustworthy opinion of an entity, treat it as a random variable, and use its distribution to describe consistency or controversy, which is particularly difficult for data which can be numerically measured, i.e. quantitative information. In this study, we focus on the quantitative opinion, propose an uncertainty-aware approach called Kernel Density Estimation from Multiple Sources (KDEm) to estimate its probability distribution, and summarize trustworthy information based on this distribution. Experiments indicate that KDEm not only has outstanding performance on the classical numeric truth discovery task, but also shows good performance on multi-modality detection and anomaly detection in the uncertain-opinion setting.
Mengting Wan, Lance M. Kaplan, Jiawei Han 0001, Jing Gao 0004, Bo Zhao 0001
KDD3
2016 Semantic Reasoning with Uncertain Information from Unreliable Sources
Murat Sensoy, Lance M. Kaplan, Geeth de Mel
PRIMA2
2016 GeoBurst: Real-Time Local Event Detection in Geo-Tagged Tweet Streams
abstract
The real-time discovery of local events (e.g., protests, crimes, disasters) is of great importance to various applications, such as crime monitoring, disaster alarming, and activity recommendation. While this task was nearly impossible years ago due to the lack of timely and reliable data sources, the recent explosive growth in geo-tagged tweet data brings new opportunities to it. That said, how to extract quality local events from geo-tagged tweet streams in real time remains largely unsolved so far.
Chao Zhang 0014, Quan Yuan 0001, Honglei Zhuang, Yu Zheng 0004, Lance M. Kaplan, Shaowen Wang 0001, Jiawei Han 0001
SIGIR6
2016 ClariSense+: An enhanced traffic anomaly explanation service using social network feeds
Prasanna Giridhar, Md. Tanvir Al Amin, Tarek F. Abdelzaher, Dong Wang 0002, Lance M. Kaplan, Jemin George, Raghu K. Ganti
Pervasive Mob. Comput.5
2016 Aggregating Crowdsourced Quantitative Claims: Additive and Multiplicative Models
abstract
Truth discovery is an important technique for enabling reliable crowdsourcing applications. It aims to automatically discover the truths from possibly conflicting crowdsourced claims. Most existing truth discovery approaches focus oncategoricalapplications, such as image classification. They use the accuracy, i.e., rate of exactly correct claims, to capture the reliability of participants. As a consequence, they are not effective for truth discovery inquantitativeapplications, such as percentage annotation and object counting, where similarity rather than exact matching between crowdsourced claims and latent truths should be considered. In this paper, we propose two unsupervised Quantitative Truth Finders (QTFs) for truth discovery in quantitative crowdsourcing applications. One QTF explores an additive model and the other explores a multiplicative model to capture different relationships between crowdsourced claims and latent truths in different classes of quantitative tasks. These QTFs naturally incorporate the similarity between variables. Moreover, they use the bias and the confidence instead of the accuracy to capture participants’ abilities in quantity estimation. These QTFs are thus capable of accurately discovering quantitative truths in particular domains. Through extensive experiments, we demonstrate that these QTFs outperform other state-of-the-art approaches for truth discovery in quantitative crowdsourcing applications and they are also quite efficient.
Wentao Robin Ouyang, Lance M. Kaplan, Alice Toniolo, Mani Srivastava 0001, Timothy J. Norman
IEEE Trans. Knowl. Data Eng.2
2016 Parallel and Streaming Truth Discovery in Large-Scale Quantitative Crowdsourcing
abstract
To enable reliable crowdsourcing applications, it is of great importance to develop algorithms that can automatically discover the truths from possibly noisy and conflicting claims provided by various information sources. In order to handle crowdsourcing applications involving big or streaming data, a desirable truth discovery algorithm should not only beeffective, but also bescalable. However, with respect to quantitative crowdsourcing applications such as object counting and percentage annotation, existing truth discovery algorithms are not simultaneously effective and scalable. They either address truth discovery in categorical crowdsourcing or perform batch processing that does not scale. In this paper, we propose new parallel and streaming truth discovery algorithms for quantitative crowdsourcing applications. Through extensive experiments on real-world and synthetic datasets, we demonstrate that 1) both of them are quite effective, 2) the parallel algorithm can efficiently perform truth discovery on large datasets, and 3) the streaming algorithm processes data incrementally, and it can efficiently perform truth discovery both on large datasets and in data streams.
Wentao Robin Ouyang, Lance M. Kaplan, Alice Toniolo, Mani Srivastava 0001, Timothy J. Norman
IEEE Trans. Parallel Distributed Syst.2
2015 Joint Localization of Events and Sources in Social Networks
abstract
Recent sensor network literature investigated the use of social networks as sensor networks, and formulated a physical event localization problem from social network data. This paper improves on the above results by formulating a joint localization problem of events and sources, leveraging the fact that sources on social networks often have a location affinity: They tend to comment more on events in their locations of interest. While social networks, such as Twitter, do not offer source location information for the majority of sources, we show that our algorithms for jointly inferring source and event location significantly improve localization quality by mutually enhancing location estimation of both events and sources. We evaluate the performance of our algorithm both in simulation and using Twitter data about current events. The results show that joint inference of source and event location allows us to localize many more of the events identified in real-world datasets.
Prasanna Giridhar, Shiguang Wang, Tarek F. Abdelzaher, Jemin George, Lance M. Kaplan, Raghu K. Ganti
DCOSS5
2015 Towards subjective networks: Extending conditional reasoning in subjective logic
Lance M. Kaplan, Magdalena Ivanovska, Audun Jøsang, Francesco Sambo
FUSION1
2015 FUSE-BEE: Fusion of subjective opinions through behavior estimation
Murat Sensoy, Lance M. Kaplan, Gonul Ayci, Geeth de Mel
FUSION2
2015 Debiasing crowdsourced quantitative characteristics in local businesses and services
abstract
Information about quantitative characteristics in local businesses and services, such as the number of people waiting in line in a cafe and the number of available fitness machines in a gym, is important for informed decision, crowd management and event detection. In this paper, we investigate the potential of leveraging crowds as sensors to report such quantitative characteristics and investigate how to recover the true quantity values from noisy crowdsourced information. Through experiments, we find that crowd sensors have both bias and variance in quantity sensing, and task difficulties impact the sensing accuracy. Based on these findings, we propose an unsupervised probabilistic model to jointly assess task difficulties, ability of crowd sensors and true quantity values. Our model differs from existing categorical truth finding models as ours is specifically designed to tackle quantitative truth. In addition to devising an efficient model inference algorithm in a batch mode, we also design an even faster online version for handling streaming data. Experimental results in various scenarios demonstrate the effectiveness of our model.
Wentao Robin Ouyang, Lance M. Kaplan, Paul Martin 0008, Alice Toniolo, Mani Srivastava 0001, Timothy J. Norman
IPSN2
2015 Scalable social sensing of interdependent phenomena
abstract
The proliferation of mobile sensing and communication devices in the possession of the average individual generated much recent interest in social sensing applications. Significant advances were made on the problem of uncovering ground truth from observations made by participants of unknown reliability. The problem, also called fact-finding commonly arises in applications where unvetted individuals may opt in to report phenomena of interest. For example, reliability of individuals might be unknown when they can join a participatory sensing campaign simply by downloading a smartphone app. This paper extends past social sensing literature by offering a scalable approach for exploiting dependencies between observed variables to increase fact-finding accuracy. Prior work assumed that reported facts are independent, or incurred exponential complexity when dependencies were present. In contrast, this paper presents the first scalable approach for accommodating dependency graphs between observed states. The approach is tested using real-life data collected in the aftermath of hurricane Sandy on availability of gas, food, and medical supplies, as well as extensive simulations. Evaluation shows that combining expected correlation graphs (of outages) with reported observations of unknown reliability, results in a much more reliable reconstruction of ground truth from the noisy social sensing data. We also show that correlation graphs can help test hypotheses regarding underlying causes, when different hypotheses are associated with different correlation patterns. For example, an observed outage profile can be attributed to a supplier outage or to excessive local demand. The two differ in expected correlations in observed outages, enabling joint identification of both the actual outages and their underlying causes.
Shiguang Wang, Lu Su 0001, Shen Li 0002, Shaohan Hu, Md. Tanvir Al Amin, Shuochao Yao, Lance M. Kaplan, Tarek F. Abdelzaher
IPSN8
2015 GIN: A Clustering Model for Capturing Dual Heterogeneity in Networked Data
abstract
Networked data often consists of interconnected multi-typed nodes and links. A common assumption behind such heterogeneity is the shared clustering structure. However, existing network clustering approaches over-simplify the heterogeneity by either treating nodes or links in a homogeneous fashion, resulting in massive loss of information. In addition, these studies are more or less restricted to specific network schemas or applications, losing generality. In this paper, we introduce a flexible model to explain the process of forming heterogeneous links based on shared clustering information of heterogeneous nodes. Specifically, we categorize the link generation process into binary and weighted cases and model them respectively. We show these two cases can be seamlessly integrated into a unified model. We propose to maximize a joint log-likelihood function to infer the model efficiently with Expectation Maximization (EM) algorithms. Experiments on real-world networked data sets demonstrate the effectiveness and flexibility of the proposed method in fully capturing the dual heterogeneity of both nodes and links.
Chi Wang 0001, Jing Gao 0004, Quanquan Gu, Charu C. Aggarwal, Lance M. Kaplan, Jiawei Han 0001
SDM6
2015 Graph Regularized Meta-path Based Transductive Regression in Heterogeneous Information Network
abstract
A number of real-world networks are heterogeneous information networks, which are composed of different types of nodes and links. Numerical prediction in heterogeneous information networks is a challenging but significant area because network based information for unlabeled objects is usually limited to make precise estimations. In this paper, we consider a graph regularized meta-path based transductive regression model (Grempt), which combines the principal philosophies of typical graph-based transductive classification methods and transductive regression models designed for homogeneous networks. The computation of our method is time and space efficient and the precision of our model can be verified by numerical experiments.
Mengting Wan, Yunbo Ouyang, Lance M. Kaplan, Jiawei Han 0001
SDM3
2015 Reliable social sensing with physical constraints: analytic bounds and performance evaluation
Dong Wang 0002, Tarek F. Abdelzaher, Lance M. Kaplan, Raghu K. Ganti, Shaohan Hu, Hengchang Liu
Real Time Syst.3
2015 Fusion of Quantized and Unquantized Sensor Data for Estimation
abstract
This letter investigates the usefulness of quantized data for estimation problems in which unquantized data is already available. A worst case scenario is considered in which a fusion center has access to continuous and binary-valued measurements of the same uniformly distributed parameter observed in Gaussian noise. The difference in mean squared error between a minimum mean squared error estimate using unquantized data and a minimum mean squared error estimate using both quantized and unquantized data is used to quantify the value of fusing the two kinds of data. Discussion of the Cramér-Rao Bound predicts how noise in the quantized data affects the reduction in estimate mean squared error from fusing the data types. It is then determined that the maximum reduction in estimate mean squared error from fusion can be approximated as a rational function of the ratio of the standard deviations of the measurement noise in the two data types. Finally, similarities between the approximation to the reduction in estimate mean squared error for the most favorable uniform prior width and a closed form expression based on the Cramér-Rao Bound are discussed.
David Saska, Rick S. Blum, Lance M. Kaplan
IEEE Signal Process. Lett.3
2015 You can't get there from here: sensor scheduling with refocusing delays
Yosef Alayev, Amotz Bar-Noy, Matthew P. Johnson 0001, Lance M. Kaplan, Thomas La Porta
Wirel. Networks4
2014 Multi-shooter localization using finite point process
Jemin George, Lance M. Kaplan, Socrates Deligeorges, George Cakiades
FUSION2
2014 Trust estimation and fusion of uncertain information by exploiting consistency
Lance M. Kaplan, Murat Sensoy, Geeth de Mel
FUSION1
2014 Using humans as sensors: an estimation-theoretic perspective
Dong Wang 0002, Md. Tanvir Al Amin, Shen Li 0002, Tarek F. Abdelzaher, Lance M. Kaplan, Siyu Gu, Chenji Pan, Hengchang Liu, Charu C. Aggarwal, Raghu K. Ganti, Xinlei (Oscar) Wang, Prasant Mohapatra, Boleslaw K. Szymanski, Hieu Khac Le
IPSN5
2014 Analyzing expert behaviors in collaborative networks
abstract
Collaborative networks are composed of experts who cooperate with each other to complete specific tasks, such as resolving problems reported by customers. A task is posted and subsequently routed in the network from an expert to another until being resolved. When an expert cannot solve a task, his routing decision (i.e., where to transfer a task) is critical since it can significantly affect the completion time of a task. In this work, we attempt to deduce the cognitive process of task routing, and model the decision making of experts as a generative process where a routing decision is made based on mixed routing patterns.
Huan Sun 0001, Mudhakar Srivatsa, Shulong Tan, Yang Li 0150, Lance M. Kaplan, Shu Tao, Xifeng Yan
KDD5
2014 Generalized Decision Aggregation in Distributed Sensing Systems
abstract
In this paper, we present GDA, a generalized decision aggregation framework that integrates information from distributed sensor nodes for decision making in a resource efficient manner. Traditional approaches that target similar problems only take as input the discrete label information from individual sensors that observe the same events. Different from them, our proposed GDA framework is able to take advantage of the confidence information of each sensor about its decision, and thus achieves higher decision accuracy. Targeting generalized problem domains, our framework can naturally handle the scenarios where different sensor nodes observe different sets of events whose numbers of possible classes may also be different. GDA also makes no assumption about the availability level of ground truth label information, while being able to take advantage of any if present. For these reasons, our approach can be applied to a much broader spectrum of sensing scenarios. The advantages of our proposed framework are demonstrated through both theoretic analysis and extensive experiments.
Lu Su 0001, Qi Li 0012, Shaohan Hu, Shiguang Wang, Jing Gao 0004, Hengchang Liu, Tarek F. Abdelzaher, Jiawei Han 0001, Xue (Steve) Liu, Yan Gao 0010, Lance M. Kaplan
RTSS11
2014 Towards Cyber-Physical Systems in Social Spaces: The Data Reliability Challenge
abstract
Today's cyber-physical systems (CPS) increasingly operate in social spaces. Examples include transportation systems, disaster response systems, and the smart grid, where humans are the drivers, survivors, or users. Much information about the evolving system can be collected from humans in the loop, a practice that is often called crowd-sensing. Crowd-sensing has not traditionally been considered a CPS topic, largely due to the difficulty in rigorously assessing its reliability. This paper aims to change that status quo by developing a mathematical approach for quantitatively assessing the probability of correctness of collected observations (about an evolving physical system), when the observations are reported by sources whose reliability is unknown. The paper extends prior literature on state estimation from noisy inputs, that often assumed unreliable sources that fall into one or a small number of categories, each with the same (possibly unknown) background noise distribution. In contrast, in the case of crowd-sensing, not only do we assume that the error distribution is unknown but also that each (human) sensor has its own possibly different error distribution. Given the above assumptions, we rigorously estimate data reliability in crowd-sensing systems, hence enabling their exploitation as state estimators in CPS feedback loops. We first consider applications where state is described by a number of binary variables, then extend the approach trivially to multivalued variables. The approach also extends prior work that addressed the problem in the special case of systems whose state does not change over time. Evaluation results, using both simulation and a real-life case-study, demonstrate the accuracy of the approach.
Shiguang Wang, Dong Wang 0002, Lu Su 0001, Lance M. Kaplan, Tarek F. Abdelzaher
RTSS4
2014 Efficient Solutions Framework for Optimal Multitask Resource Assignments for Data Fusion in Wireless Sensor Networks
abstract
Motivated by the need to judiciously allocate scarce sensing resources to attain the highest benefit for the applications that sensor networks serve, in this article we develop a flexible solutions methodology for maximizing the overall reward attained, subject to constraints on the resource demands under fairly general reward or demand functions. We map a broad class of related problems for data fusion in wireless sensor networks into an integer programming problem and provide an iterative Lagrangian relaxation technique to solve it. Each iteration step involves solving for a maximum-weight independent set of an appropriately constructed graph, which, in many cases, can be obtained in polynomial time. We apply our methodology to the problem of tracking targets moving over a period of time through a nonhomogeneous, energy-constrained sensor field. With rewards represented by the quality of information attained in tracking, we study its trade-offs and relationship with energy consumption and periodic measurement taking. We finally illustrate other applications of our framework in sensor networks.
Srikanth Hariharan, Chatschik Bisdikian, Lance M. Kaplan, Tien Pham
ACM Trans. Sens. Networks3
2014 Maximum likelihood analysis of conflicting observations in social sensing
abstract
This article addresses the challenge of truth discovery from noisy social sensing data. The work is motivated by the emergence of social sensing as a data collection paradigm of growing interest, where humans perform sensory data collection tasks. Unlike the case with well-calibrated and well-tested infrastructure sensors, humans are less reliable, and the likelihood that participants' measurements are correct is often unknown a priori. Given a set of human participants of unknown trustworthiness together with their sensory measurements, we pose the question of whether one can use this information alone to determine, in an analytically founded manner, the probability that a given measurement is true. In our previous conference paper, we offered the first maximum likelihood solution to the aforesaid truth discovery problem for corroborating observations only. In contrast, this article extends the conference paper and provides the first maximum likelihood solution to handle the cases where measurements from different participants may be conflicting . The article focuses on binary measurements. The approach is shown to outperform our previous work used for corroborating observations, the state-of-the-art fact-finding baselines, as well as simple heuristics such as majority voting.
Dong Wang 0002, Lance M. Kaplan, Tarek F. Abdelzaher
ACM Trans. Sens. Networks2
2013 A fusion architecture for tracking a group of people using a distributed sensor network
Thyagaraju R. Damarla, Lance M. Kaplan
FUSION2
2013 Reasoning under uncertainty: Variations of subjective logic deduction
Lance M. Kaplan, Murat Sensoy, Supriyo Chakraborty, Chatschik Bisdikian, Geeth de Mel
FUSION1
2013 TRIBE: Trust revision for information based on evidence
Murat Sensoy, Geeth de Mel, Lance M. Kaplan, Tien Pham, Timothy J. Norman
FUSION3
2013 Recursive Fact-Finding: A Streaming Approach to Truth Estimation in Crowdsourcing Applications
abstract
This paper presents a streaming approach to solve the truth estimation problem in crowdsourcing applications. We consider a category of crowdsourcing applications where a group of individuals volunteer (or are recruited to) share certain observations or measurements about the physical world. Examples include reporting locations of gas stations that remain operational after a natural disaster or reporting locations of potholes on city streets. We call such applications social sensing. Ascertaining the correctness of reported observations is a key challenge in such applications, referred to as the truth estimation problem. This problem is made difficult by the fact that the reliability of individual sources is usually unknown a priori, since any concerned citizen may, in principle, participate. Moreover, the timescales of crowdsourcing campaigns of interest can be as small as a few hours or days, which does not offer enough history for a reputation system to converge. Instead, recent prior work, including our own, developed fact-finding algorithms to solve this problem by iteratively assessing the credibility of sources and their claims in the absence of reputation scores. Such algorithms, however, operate on the entire dataset of reported observations in a batch fashion, which makes them less suited to applications where new observations arrive continually. In this paper, we describe a streaming fact-finder that recursively updates previous estimates based on new data. The recursive algorithm solves an expectation maximization (EM) problem to determine the odds of correctness of different observations. We compare the performance of our recursive EM algorithm to a batch EM algorithm, as well as to several state-of-art fact-finders through extensive simulations. We also demonstrate convergence of the recursive algorithm to the results of the batch version through a real social sensing experiment. Our evaluation shows that the proposed approach can process data streams much more efficiently while keeping the truth estimation accuracy close to that of the (much slower) batch algorithm. Ours is therefore the first fact-finder developed with explicit consideration to the continuous update needs of crowd-sourcing applications.
Dong Wang 0002, Tarek F. Abdelzaher, Lance M. Kaplan, Charu C. Aggarwal
ICDCS3
2013 Exploitation of Physical Constraints for Reliable Social Sensing
abstract
This paper develops and evaluates algorithms for exploiting physical constraints to improve the reliability of social sensing. Social sensing refers to applications where a group of sources (e.g., individuals and their mobile devices) volunteer to collect observations about the physical world. A key challenge in social sensing is that the reliability of sources and their devices is generally unknown, which makes it non-trivial to assess the correctness of collected observations. To solve this problem, the paper adopts a cyber-physical approach, where assessment of correctness of individual observations is aided by knowledge of physical constraints on both sources and observed variables to compensate for the lack of information on source reliability. We cast the problem as one of maximum likelihood estimation. The goal is to jointly estimate both (i) the latent physical state of the observed environment, and (ii) the inferred reliability of individual sources such that they are maximally consistent with both provenance information (who claimed what) and physical constraints. We evaluate the new framework through a real-world social sensing application. The results demonstrate significant performance gains in estimation accuracy of both source reliability and observation correctness.
Dong Wang 0002, Tarek F. Abdelzaher, Lance M. Kaplan, Raghu K. Ganti, Shaohan Hu, Hengchang Liu
RTSS3
2013 On Credibility Estimation Tradeoffs in Assured Social Sensing
abstract
Two goals of network science are to (i) uncover fundamental properties of phenomena modeled as networks, and to (ii) explore novel use of networks as models for a diverse range of systems and phenomena in order to improve our understanding of such systems and phenomena. This paper advances the latter direction by casting credibility estimation in social sensing applications as a network science problem, and by presenting a network model that helps understand the fundamental accuracy trade-offs of a credibility estimator. Social sensing refers to data collection scenarios, where observations are collected from (possibly unvetted) human sources. We call such observations claims to emphasize that we do not know whether or not they are factually correct. Predictable, scalable and robust estimation of both source reliability and claim correctness, given neither in advance, becomes a key challenge given the unvetted nature of sources and lack of means to verify their claims. In a previous conference publication, we proposed a maximum likelihood approach to jointly estimate both source reliability and claim correctness. We also derived confidence bounds to quantify the accuracy of such estimation. In this paper, we cast credibility estimation as a network science problem and offer systematic sensitivity analysis of the optimal estimator to understand its fundamental accuracy trade-offs as a function of an underlying network topology that describes key problem space parameters. It enables assured social sensing, where not only source reliability and claim correctness are estimated, but also the accuracy of such estimates is correctly predicted for the problem at hand.
Dong Wang 0002, Lance M. Kaplan, Tarek F. Abdelzaher, Charu C. Aggarwal
IEEE J. Sel. Areas Commun.2
2013 Impact of In-Network Aggregation on Target Tracking Quality Under Network Delays
abstract
In this paper, we investigate how in-network aggregation approach impacts the target tracking quality in multi-hop wireless sensor networks under network delays. Specifically, we use the mean squared error (MSE) of the target location estimate to quantify the target tracking quality, and investigate how in-network aggregation affects the MSE. To obtain insights without being obscured by onerous mathematical details, we assume a Brownian motion mobility model for the target, Gaussian measurement noise for the sensors, and independent per-hop delays. Under the above assumptions, we first propose an aggregation scheme that preserves a sufficient statistic for optimal tracking under data aggregation at the intermediate nodes and arbitrary network delays. We then analytically study the impact of aggregation in three increasingly more complicated scenarios: single task tracking with only transmission delay, single task tracking with both transmission delay and queueing delay at intermediate nodes, and multi-task tracking. Our results demonstrate that in-network aggregation improves tracking quality in all three scenarios. Furthermore, our analysis provides guidelines on how to choose aggregation parameters in practice.
Wei Wei 0001, Ting He 0001, Chatschik Bisdikian, Dennis Goeckel, Bo Jiang 0003, Lance M. Kaplan, Don Towsley
IEEE J. Sel. Areas Commun.6
2013 On the quality and value of information in sensor networks
abstract
The increasing use of sensor-derived information from planned, ad-hoc, and/or opportunistically deployed sensor networks provides enhanced visibility to everyday activities and processes, enabling fast-paced data-to-decision in personal, social, civilian, military, and business contexts. The value that information brings to this visibility and ensuing decisions depends on the quality characteristics of the information gathered. In this article, we highlight, refine, and extend upon our past work in the areas of quality and value of information (QoI and VoI) for sensor networks. Specifically, we present and elaborate on our two-layer QoI/VoI definition, where the former relates to context-independent aspects and the latter to context-dependent aspects of an information product. Then, we refine our taxonomy of pertinent QoI and VoI attributes anchored around a simple ontological relationship between the two. Finally, we introduce a framework for scoring and ranking information products based on their VoI attributes using the analytic hierarchy multicriteria decision process, illustrated via a simple example.
Chatschik Bisdikian, Lance M. Kaplan, Mani Srivastava 0001
ACM Trans. Sens. Networks2
2012 Balancing value and risk in information sharing through obfuscation
Supriyo Chakraborty, Kasturi Rangan Raghavan, Mani Srivastava 0001, Chatschik Bisdikian, Lance M. Kaplan
FUSION5
2012 Subjective logic with uncertain partial observations
Lance M. Kaplan, Supriyo Chakraborty, Chatschik Bisdikian
FUSION1
2012 On truth discovery in social sensing: a maximum likelihood estimation approach
abstract
This paper addresses the challenge of truth discovery from noisy social sensing data. The work is motivated by the emergence of social sensing as a data collection paradigm of growing interest, where humans perform sensory data collection tasks. A challenge in social sensing applications lies in the noisy nature of data. Unlike the case with well-calibrated and well-tested infrastructure sensors, humans are less reliable, and the likelihood that participants' measurements are correct is often unknown a priori. Given a set of human participants of unknown reliability together with their sensory measurements, this paper poses the question of whether one can use this information alone to determine, in an analytically founded manner, the probability that a given measurement is true. The paper focuses on binary measurements. While some previous work approached the answer in a heuristic manner, we offer the first optimal solution to the above truth discovery problem. Optimality, in the sense of maximum likelihood estimation, is attained by solving an expectation maximization problem that returns the best guess regarding the correctness of each measurement. The approach is shown to outperform the state of the art fact-finding heuristics, as well as simple baselines such as majority voting.
Dong Wang 0002, Lance M. Kaplan, Hieu Khac Le, Tarek F. Abdelzaher
IPSN2
2012 IntruMine: Mining Intruders in Untrustworthy Data of Cyber-physical Systems
abstract
A Cyber-Physical System (CPS) integrates physical (i.e., sensor) devices with cyber (i.e., informational) components to form a situation-aware system that responds intelligently to dynamic changes in real-world. It has wide application to scenarios of traffic control, environment monitoring and battlefield surveillance. This study investigates the specific problem of intruder mining in CPS: With a large number of sensors deployed in a designated area, the task is real time detection of intruders who enter the area, based on untrustworthy data. We propose a method called IntruMine to detect and verify the intruders. IntruMine constructs monitoring graphs to model the relationships between sensors and possible intruders, and computes the position and energy of each intruder with the link information from these monitoring graphs. Finally, a confidence rating is calculated for each potential detection, reducing false positives in the results. IntruMine is a generalized approach. Two classical methods of intruder detection can be seen as special cases of IntruMine under certain conditions. We conduct extensive experiments to evaluate the performance of IntruMine on both synthetic and real datasets and the experimental results show that IntruMine has better effectiveness and efficiency than existing methods.
Lu-An Tang, Quanquan Gu, Xiao Yu 0007, Jiawei Han 0001, Thomas La Porta, Alice Leung, Tarek F. Abdelzaher, Lance M. Kaplan
SDM8
2012 On scalability and robustness limitations of real and asymptotic confidence bounds in social sensing
abstract
This paper estimates new confidence bounds on source reliability in social sensing applications. Scalable and robust estimation of source reliability is a key challenge in social sensing where humans or human-operated sensors act as data sources. In order to assess correctness of data, the reliability of sources must first be assessed, yet this is complicated when sources are not a priori known and vetted, but rather can opt in at will, for example, by downloading a sensing application on their mobile device. In our previous work, we developed a maximum likelihood source reliability estimator and approximately quantified confidence in its estimation based on an asymptotic Cramer-Rao lower bound (CRLB). In this paper we show that the asymptotic bound fails to track estimation performance when the number of sources is small. We derive the real CRLB to accurately characterize estimation performance for scenarios where the asymptotic bound fails. We study the limitations of the real and asymptotic CRLBs and show the trade-offs they offer between computational complexity and estimation scalability. We also evaluate the robustness of these bounds to changes in the number of sources. The results offer an understanding of attainable estimation accuracy of source reliability in social sensing applications that rely on un-vetted sources whose reliability is not known in advance.
Dong Wang 0002, Lance M. Kaplan, Tarek F. Abdelzaher, Charu C. Aggarwal
SECON2
2012 Measuring Two-Event Structural Correlations on Graphs
abstract
Real-life graphs usually have various kinds of events happening on them, e.g., product purchases in online social networks and intrusion alerts in computer networks. The occurrences of events on the same graph could be correlated, exhibiting either attraction or repulsion. Such structural correlations can reveal important relationships between different events. Unfortunately, correlation relationships on graph structures are not well studied and cannot be captured by traditional measures. In this work, we design a novel measure for assessing two-event structural correlations on graphs. Given the occurrences of two events, we choose uniformly a sample of "reference nodes" from the vicinity of all event nodes and employ the Kendall's τ rank correlation measure to compute the average concordance of event density changes. Significance can be efficiently assessed by τ's nice property of being asymptotically normal under the null hypothesis. In order to compute the measure in large scale networks, we develop a scalable framework using different sampling strategies. The complexity of these strategies is analyzed. Experiments on real graph datasets with both synthetic and real events demonstrate that the proposed framework is not only efficacious, but also efficient and scalable.
Ziyu Guan, Xifeng Yan, Lance M. Kaplan
Proc. VLDB Endow.3
2011 Multi-sensor data fusion: An unscented least squares approach
Jemin George, Lance M. Kaplan
FUSION2
2011 Shooter localization using soldier-worn gunfire detection systems
Jemin George, Lance M. Kaplan
FUSION2
2011 QoI-based resource allocation for multi-target tracking in energy constrained sensor networks
Srikanth Hariharan, Chatschik Bisdikian, Lance M. Kaplan, Tien Pham
FUSION3
2011 Target positioning and tracking in degenerate geometry
Lance M. Kaplan, Erik Blasch, Michael Bakich
FUSION2
2011 Optimal placement of heterogeneous sensors in target tracking
Lance M. Kaplan, Erik Blasch, Michael Bakich
FUSION2
2011 Diffuse Prior Monotonic Likelihood Ratio Test for Evaluation of Fused Image Quality Measures
abstract
This paper introduces a novel method to score how well proposed fused image quality measures (FIQMs) indicate the effectiveness of humans to detect targets in fused imagery. The human detection performance is measured via human perception experiments. A good FIQM should relate to perception results in a monotonic fashion. The method computes a new diffuse prior monotonic likelihood ratio (DPMLR) to facilitate the comparison of the H(1) hypothesis that the intrinsic human detection performance is related to the FIQM via a monotonic function against the null hypothesis that the detection and image quality relationship is random. The paper discusses many interesting properties of the DPMLR and demonstrates the effectiveness of the DPMLR test via Monte Carlo simulations. Finally, the DPMLR is used to score FIQMs with test cases considering over 35 scenes and various image fusion algorithms.
Chuanming Wei, Lance M. Kaplan, Stephen D. Burks, Rick S. Blum
IEEE Trans. Image Process.2
2010 Design of operation parameters to resolve two targets using proximity sensors
Qiang Le, Lance M. Kaplan
FUSION2
2010 You can't get there from here: Sensor scheduling with refocusing delays
abstract
We study a problem in which a single sensor is scheduled to observe sites periodically, motivated by applications in which the goal is to maintain up-to-date readings for all the observed sites. In the existing literature, it is typically assumed that the time for a sensor switching from one site to another is negligible. This may not be the case in applications such as camera surveillance of a border, however, in which the camera takes time to pan and tilt to refocus itself to a new geographical location. We formulate a problem with refocusing delay constraints. We prove the problem to be NP-hard and then study a special case in which refocusing is proportional to some Euclidian metric. We give a lower bound on the optimal cost for the scheduling problem. Finally, we provide and experimentally evaluate several heuristic algorithms, some of them based on this computed lower bound.
Yosef Alayev, Amotz Bar-Noy, Matthew P. Johnson 0001, Lance M. Kaplan, Thomas La Porta
MASS4
2010 An Ontology-Based Framework for Designing a Sensor Quality of Information Software Library
abstract
Assessing the quality of sensor-originated information is key for the effective and predictable operation sensor-enabled computerized applications. However, with increasing uncertainties due to alternative deployment scenarios, operational realities, and sensing resource use, it becomes very challenging in designing replicable software solutions that can be easily reused in a number of occasions with minimal customization. Leveraging semantic sensor web technologies, this paper presents an ontology-based design framework for organizing a library of quality of information (QoI) analysis algorithms specific to a data source, and interfacing to a library containing computational algorithms assessing quality of information. The ontology-based framework is broad enough to allow easy accommodation of new computational algorithms that domain experts may provide to the library as needed to reflect specific deployment, operational, sensing realities.
Jerrid Matthews, Chatschik Bisdikian, Lance M. Kaplan, Tien Pham
Mobile Data Management3
2010 Quality Tradeoffs in Object Tracking with Duty-Cycled Sensor Networks
abstract
Extending the lifetime of wireless sensor networks requires energy-conserving operations such as duty-cycling. However, such operations may impact the effectiveness of high fidelity real-time sensing tasks, such as object tracking, which require high accuracy and short response times. In this paper, we quantify the influence of different duty-cycle schemes on the efficiency of bearings-only object tracking. Specifically, we use the Maximum Likelihood localization technique to analyze the accuracy limits of object location estimates under different response latencies considering variable network density and duty-cycle parameters. Moreover, we study the tradeoffs between accuracy and response latency under various scenarios and motion patterns of the object. We have also investigated the effects of different duty-cycled schedules on the tracking accuracy using acoustic sensor data collected at Aberdeen Proving Ground, Maryland, by the U.S. Army Research Laboratory (ARL).
Sadaf Zahedi, Mani Srivastava 0001, Chatschik Bisdikian, Lance M. Kaplan
RTSS4
2009 Detection and Localization Sensor Assignment with Exact and Fuzzy Locations
Hosam Rowaihy, Matthew P. Johnson 0001, Diego Pizzocaro, Amotz Bar-Noy, Lance M. Kaplan, Thomas La Porta, Alun D. Preece
DCOSS5
2009 Building principles for a quality of information specification for sensor information
Chatschik Bisdikian, Lance M. Kaplan, Mani Srivastava 0001, David J. Thornley, Dinesh C. Verma, Robert I. Young
FUSION2
2009 Diffuse prior monotonic likelihood ratio test for evaluation of fused image quality metrics
Chuanming Wei, Lance M. Kaplan, Stephen D. Burks, Rick S. Blum
FUSION2
2009 Block Wiener-based image registration for moving target indication
Lance M. Kaplan, Nasser M. Nasrabadi
Image Vis. Comput.1
2009 Acoustic sensor network design for position estimation
abstract
In this article, we develop tractable mathematical models and approximate solution algorithms for a class of integer optimization problems with probabilistic and deterministic constraints, with applications to the design of distributed sensor networks that have limited connectivity. For a given deployment region size, we calculate the Pareto frontier of the sensor network utility at the desired probabilities for d -connectivity and k -coverage. As a result of our analysis, we determine (1) the number of sensors of different types to deploy from a sensor pool, which offers a cost vs. performance trade-off for each type of sensor, (2) the minimum required radio transmission ranges of the sensors to ensure connectivity, and (3) the lifetime of the sensor network. For generality, we consider randomly deployed sensor networks and formulate constrained optimization technique to obtain the localization performance. The approach is guided and validated using an unattended acoustic sensor network design. Finally, approximations of the complete statistical characterization of the acoustic sensor networks are given, which enable average network performance predictions of any combination of acoustic sensors.
Volkan Cevher, Lance M. Kaplan
ACM Trans. Sens. Networks2
2008 Pareto Frontiers of Sensor Networks for Localization
abstract
We develop a theory to predict the localization performance of randomly distributed sensor networks consisting of various sensor modalities when only a constant active subset of sensors that minimize localization error is used for estimation. The characteristics of the modalities include measurement type (bearing or range) and error, sensor reliability, FOV, sensing range, and mobility. We show that the localization performance of a sensor network is a function of a weighted sum of the total number of each sensor modality. We also show that optimization of this weighted sum is independent of how the sensor management strategy chooses the active sensors. We combine the utility objective with other objectives, such as lifetime, coverage and reliability to determine the best mix of sensors for an optimal sensor network design. The Pareto efficient frontier of the multi objectives are obtained with a dynamic program, which also accommodates additional convex constraints.
Volkan Cevher, Lance M. Kaplan
IPSN2
2008 QoI for passive acoustic gunfire localization
abstract
A network of acoustic arrays can localize a shooter. Each array can determine the direction of arrival (DOA) of the muzzle blast. Some arrays can also see the shockwave generated by the bullet as it travels at supersonic speeds. These arrays can also estimate the range to the shooter as a nonlinear function of the DOA and time difference of arrival (TDOA) between the shockwave and the muzzle blast, and these arrays can provide a rough estimate of the shooter location. The measurements from all the arrays that detect the gunfire can be fused to obtain a more accurate shooter location. This paper presents the accuracy component of the QoI for the entire sensor network by deriving the Cramer-Rao bound for the shooter location. The correspondence between QoI and actual root mean squared localization error is verified via Monte Carlo simulations.
Lance M. Kaplan, Thyagaraju R. Damarla, Tien Pham
MASS1
2007 Human infrastructure & human activity detection
abstract
Prior to committing personnel to investigate a building or suspicious site such as a cave, it is imperative to determine the importance and current danger of the site. To this end, sensors on a robotic platform can interrogate the site prior to sending in personnel. This paper investigates methods to exploit multiple sensor modalities in order to automatically 1) detect human presence, and 2) detect human infrastructure and recent human activity. The paper describes 10 experimental scenarios to support these two tasks, demonstrates what type of inference each modality can make, and shows how to fuse the information from all sensors. Experimental results are also provided for the detection of the presence of humans.
Thyagaraju R. Damarla, Lance M. Kaplan, LipChen Alex Chan
FUSION2
2006 Connectivity for the Frisbee Architecture
abstract
In this paper we investigate the k-connectivity threshold of distributed dense ad hoc heterogeneous wireless sensor network architecture. We consider the situation when sensors are deployed in the surveillance area according to a uniform distribution perturbed by a Gaussian noise. We derive analytically the minimum detection range which guarantees an emerging structure in the network, namely the connectivity, which becomes larger and larger as the number of sensors in the network increase. This allows the target track to be propagated almost surely throughout the network using the minimum possible amount of prime energy. We report the results of some simulation experiments which further support the theoretical results.
Agostino Capponi, Concetta Pilotto, Alfonso Farina, Giovanni Golino, Lance M. Kaplan
FUSION5
2006 Algorithms for the selection of the active sensors in distributed tracking: comparison between Frisbee and GNS methods
abstract
This paper compares two different approaches for sensor selection for distributed tracking: 1) The Frisbee method, and 2) Global Node Selection (GNS). The Frisbee method is based on the proximity of the nodes to the predicted location of the target; GNS is based on minimizing the unbiased Cramer Rao lower bound (CRLB). Both theoretical and experimental results indicate that the Frisbee method is as effective as GNS. Furthermore, the Frisbee method is attractive due to its very light computational load.
Agostino Capponi, Concetta Pilotto, Giovanni Golino, Alfonso Farina, Lance M. Kaplan
FUSION5
2004 Direction-of-arrival estimation of wideband sources using arbitrary shaped multidimensional arrays
abstract
We generalize a new incoherent wideband direction-of-arrival (DOA) estimation algorithm that provides higher resolution than traditional incoherent techniques. The new method is able to better adjust its beam response when multiple sources are present than incoherent MUSIC. The original algorithm was designed to work strictly for an uniform linear array. In this paper, we demonstrate how to generalize the algorithm to work over arbitrary 1D or 2D arrays. We demonstrate the higher resolution of the new algorithm against incoherent MUSIC for 1D and 2D arrays using simulations of two source signals.
Yeo-Sun Yoon, Lance M. Kaplan, James H. McClellan
ICASSP (2)2
2004 A comparison of the octave-band directional filter bank and gabor filters for texture classification
Paul S. Hong, Lance M. Kaplan, Mark J. T. Smith
ICIP2
2003 New signal subspace direction-of-arrival estimator for wideband sources
abstract
Signal subspace methods in DOA estimation for wideband sources require preprocessing to find initial values which are close enough to the true values or to convert sensor outputs into desired forms. The preprocessing procedure should be carefully done lest it introduce some distortion. Failure to find proper initial values may prevent convergence of the estimator or cause biases in the estimator. The proposed method detects uncorrelated wideband sources using the signal subspace and the noise subspace of decomposed wideband signals. It does not require any initial values. The only preprocessing is narrowband decomposition of the sensor output which is very common in other wideband methods. Computer simulation showed that the proposed method has less bias than and comparable variance to CSSM with small focusing errors.
Yeo-Sun Yoon, Lance M. Kaplan, James H. McClellan
ICASSP (5)2
2003 Pose estimation of SAR imagery using the two dimensional continuous wavelet transform
Lance M. Kaplan, Romain Murenzi
Pattern Recognit. Lett.1
2001 Maximum likelihood methods for bearings-only target localization
abstract
We develop four maximum likelihood (ML) methods to localize a moving target using a network of acoustical sensor arrays. Each array transmits a direction-of-arrival (DOA) estimate to a central processor, which employs one of the localization techniques. The four ML approaches use different target signal models where the time retardation factor for the target position and the degradation of the target signal through the air may or may not be included in the model. We compare these methods along with a linear least squares approach through a number of simulations at various signal to noise levels.
Lance M. Kaplan, Qiang Le, Péter Molnár
ICASSP1
2000 Error Analysis for Quadtree Image Formation
abstract
The quadtree image formation technique is a computationally efficient approximation to standard backprojection. Where the computational load of backprojection is O(N/sup 3/) for N sensors forming an N/spl times/N image, the quadtree method uses a divide-and-conquer strategy similar to the fast Fourier transform (FFT) to reduce the computational load down to O(N/sup 2/ log(N)). However, the quadtree introduces errors in the relative time shifts used to focus pulses. These errors reduce the signal gain in the mainlobe response for isotropic point-like targets. In addition, the oscillations of the sidelobes increase from stage to stage. This paper develops performance bounds for the mainlobe losses under far field conditions and relates these bounds to the slow-time Nyquist rate.
Lance M. Kaplan, Seung-Mok Oh, Matthew Cobb, James H. McClellan
ICIP1
2000 Image Metrics for Clutter Characterization
abstract
A large number of parameters affect the fundamental performance of automatic target recognition (ATR) systems including sensor, geometry, and scene parameters. Clutter complexity refers to the scene parameter that measures the extent that the objects in the background of the scene are target-like. This paper describes our initial work to understand how to measure clutter complexity and bound ATR performance as a function of this complexity, given that the other parameters are held constant. For this study, we use an unrealistically omnipotent ATR to approximate performance bounds and characterize the clutter complexity for each scene in the COMANCHE FLIR database. Using this characterization, we generate a clutter complexity metric from a number of image processing features and compare the new metric against the performance of an actual ATR.
Kamesh Namuduri, Karim Bouyoucef, Lance M. Kaplan
ICIP3
1999 Extended fractal analysis for texture classification and segmentation
abstract
The Hurst parameter for two-dimensional (2-D) fractional Brownian motion (fBm) provides a single number that completely characterizes isotropic textured surfaces whose roughness is scale-invariant. Extended self-similar (ESS) processes were previously introduced in order to provide a generalization of fBm. These new processes are described by a number of multiscale Hurst parameters. In contrast to the single Hurst parameter, the extended parameters are able to characterize a greater variety of natural textures where the roughness of these textures is not necessarily scale-invariant. In this work, we evaluate the effectiveness of multiscale Hurst parameters as features for texture classification and segmentation. For texture classification, the performance of the generalized Hurst features is compared to traditional Hurst and Gabor features. Our experiments show that classification accuracy for the generalized Hurst and Gabor features are comparable even though the generalized Hurst features lower the dimensionality by a factor of five. Next, the segmentation accuracy using generalized and standard Hurst features is evaluated on images of texture mosaics. For these experiments, the performance is evaluated with and without supplemental contrast and average grayscale features. Finally, we investigate the effectiveness of the Hurst features to segment real synthetic aperture radar (SAR) imagery.
Lance M. Kaplan
IEEE Trans. Image Process.1
1998 Evaluation of CFAR and texture based target detection statistics on SAR imagery
abstract
In this work, we evaluated the effectiveness of synthetic aperture radar (SAR) target detection algorithms that consist of any number of combinations of three statistics which include two-parameter CFAR, variance, and extended fractal features. The performance of these algorithms were tested at various threshold settings over the public domain MSTAR database. This database contains one foot resolution X-band SAR imagery. Receiver-operating-characteristic (ROC) curves were generated for the seven resulting algorithms. The results indicate that the CFAR statistic is the least effective detection statistic.
Lance M. Kaplan, Romain Murenzi
ICASSP1
1997 Texture Segmentation using Multiscale Hurst Features
abstract
We evaluate the effectiveness of multiscale Hurst parameters as features for texture segmentation. These extended Hurst features quantize texture roughness at various scales. The performance of these new features are compared against standard Hurst features using images of texture mosaics. For the experiments, the performance was evaluated with and without supplemental contrast and average grayscale features.
Lance M. Kaplan, Romain Murenzi
ICIP (3)1
1996 An improved method for 2-D self-similar image synthesis
abstract
We propose a new method called incremental Fourier synthesis to generate 2-D self-similar images based on a 2D fractional Brownian motion (fBm) model. With this method, the stationary increments of fBm are created by a Fourier synthesis method and the increments are added up to generate the nonstationary 2D fBm process. Since the new method takes advantage of the FFT, its computational complexity is only O(N(2)log(2)(N)), and its memory requirement is only O(N(2)) for a self-similar image of size NxN.
Lance M. Kaplan, C.-C. Jay Kuo
IEEE Trans. Image Process.1
1995 Fast image synthesis using an extended self-similar (ESS) model
abstract
Proposes a method called incremental Fourier synthesis to generate images based upon the 2-D extended self-similar (ESS) model. This algorithm creates the stationary increments of ESS processes by Fourier synthesis. Then, the increments are added up to generate the nonstationary 2-D ESS process. Because the new method can take advantage of the FFT, its computational complexity is only O(N/sup 2/ log/sub 2/(N)), and its memory requirement is O(N/sup 2/) for an image of size N/spl times/N.
Lance M. Kaplan, C.-C. Jay Kuo
ICASSP1
1995 Texture Segmentation via Haar Fractal Feature Estimation
Lance M. Kaplan, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.1
1995 Texture Roughness Analysis and Synthesis via Extended Self-Similar (ESS) Model
abstract
The 2D fractional Brownian motion (FBM) model provides a useful tool to model textured surfaces whose roughness is scale-invariant. To represent textures whose roughness is scale-dependent, we generalize the FBM model to the extended self-similar (ESS) model in this research. We present an estimation algorithm to extract the model parameters from real texture data. Furthermore, a new incremental Fourier synthesis algorithm is proposed to generate the 2D realizations of the ESS model. Finally, the estimation and rendering methods are combined to synthesize real textured surfaces.>
Lance M. Kaplan, C.-C. Jay Kuo
IEEE Trans. Pattern Anal. Mach. Intell.1
1994 Signal modeling using increments of extended self-similar processes
abstract
The article expands the fractional Brownian motion (fBm) model by investigating the idea of generalizing self-similarity to create extended self-similar (ESS) processes for which fBm processes are a subset. Properties of ESS processes are discussed, and an ESS increment model parameterized by variables controlling short and long term correlation effects is examined. We propose a fast parameter estimation algorithm for the new model which is based on the decorrelation effect of the Haar transform on the ESS increments, and we demonstrate the performance of this parameter estimation algorithm with numerical simulations.>
Lance M. Kaplan, C.-C. Jay Kuo
ICASSP (4)1
1994 Beyond Self-Similarity for Landscale Modeling
abstract
We present a model that generalizes fractional Brownian motion for the purpose or rendering realistic appearing landscapes. To describe the new model, we extend the concept of stochastic self-similarity. The advantage of the new model is that an artist can choose to alter the roughness of the landscape at different scales. We demonstrate the features of the new model through numerous examples.>
Lance M. Kaplan, C.-C. Jay Kuo
ICIP (1)1
1993 An improved algorithm for fractal estimation from noisy measurements
Lance M. Kaplan, C.-C. Jay Kuo
ICASSP (4)1