Matthai Philipose

dblp:p/MPhilipose · DBLP profile ↗
← Back
50ranked-venue papers
2as first author
2since 2021 · last 2026
0009-0005-8249-9435ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-authorSoftware engineering, systems software and programming languages · 10 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7Computer networks · 6Databases, data management, data science and information retrieval · 6Systems, architecture and hardware · 2Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Hardware accelerators and domain-specific architectures · 49% GPUs and heterogeneous computing · 18% Energy-efficient computing · 16%
Artificial intelligence
19 papers
Efficient and distributed learning · 36% Deep learning architectures and training · 22% Knowledge representation and reasoning · 10%
Human-computer interaction and pervasive computing
12 papers
Ubiquitous computing and smart environments · 60% Interaction techniques and input · 24% Wearable and physiological sensing · 8%
Databases, data mining, and information retrieval
5 papers
Query processing and optimization · 40% Data stream processing · 25% Recommender systems · 24%
Computer graphics and multimedia
5 papers
Multimedia analysis and retrieval · 62% Multimedia systems and quality of experience · 29% Computational photography and imaging · 10%
Computer networks
7 papers
Edge and fog computing · 75% Internet of things and sensor networks · 12% Content delivery and video streaming · 11%

Topics — the 30 heaviest of 79, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.832021
Boosting the Throughput and Accelerator Utilization of Specialized CNN Inference Beyond Increasing Batch Size · ICML 2021
MCDNN: An Approximation-Based Execution Framework for Deep Stream Processing Under Resource Constraints · MobiSys 2016
Heterogeneous Bitwidth Binarization in Convolutional Neural Networks · NeurIPS 2018
Machine learning › Efficient and distributed learning
model compression
0.622018
Heterogeneous Bitwidth Binarization in Convolutional Neural Networks · NeurIPS 2018
MCDNN: An Approximation-Based Execution Framework for Deep Stream Processing Under Resource Constraints · MobiSys 2016
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN inference accelerator
0.512021
Boosting the Throughput and Accelerator Utilization of Specialized CNN Inference Beyond Increasing Batch Size · ICML 2021
Ubiquitous computing and smart environments › context recognition
activity recognition
0.462009
Recognizing daily activities with RFID-based sensors · UbiComp 2009
Common Sense Based Joint Training of Human Activity Recognizers · IJCAI 2007
A Scalable Approach to Activity Recognition based on Object Use · ICCV 2007
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.412019
Nexus: a GPU cluster engine for accelerating DNN-based video analysis · SOSP 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
DNN inference scheduling
0.412019
Nexus: a GPU cluster engine for accelerating DNN-based video analysis · SOSP 2019
GPUs and heterogeneous computing › GPU resource management
GPU cluster resource management
0.412019
Nexus: a GPU cluster engine for accelerating DNN-based video analysis · SOSP 2019
GPUs and heterogeneous computing
GPU scheduling
0.412019
Nexus: a GPU cluster engine for accelerating DNN-based video analysis · SOSP 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.352008
Structure Learning on Large Scale Common Sense Statistical Models of Human State · AAAI 2008
Common Sense Based Joint Training of Human Activity Recognizers · IJCAI 2007
Learning Large Scale Common Sense Models of Everyday Life · AAAI 2007
Machine learning › Efficient and distributed learning › model compression
quantization
0.312018
Heterogeneous Bitwidth Binarization in Convolutional Neural Networks · NeurIPS 2018
Query processing and optimization › complex data query processing
video query processing
0.312018
Focus: Querying Large Video Datasets with Low Latency and Low Cost · OSDI 2018
Multimedia analysis and retrieval
video retrieval
0.312018
Focus: Querying Large Video Datasets with Low Latency and Low Cost · OSDI 2018
Energy-efficient computing › power management › device power management
image sensor power management
0.322013
Energy proportional image sensors for continuous mobile vision · MobiSys 2013
Energy characterization and optimization of image sensing toward continuous mobile vision · MobiSys 2013
Embedded and real-time systems › real-time embedded systems › multimedia embedded systems › embedded vision system
mobile vision systems
0.322013
Energy proportional image sensors for continuous mobile vision · MobiSys 2013
Energy characterization and optimization of image sensing toward continuous mobile vision · MobiSys 2013
Energy-efficient computing
power management
0.322013
Energy proportional image sensors for continuous mobile vision · MobiSys 2013
Energy characterization and optimization of image sensing toward continuous mobile vision · MobiSys 2013
Machine learning › Deep learning architectures and training
convolution
0.312017
Do Deep Convolutional Nets Really Need to be Deep and Convolutional? · ICLR (Poster) 2017
Machine learning › Deep learning architectures and training
convolutional neural network
0.312017
Do Deep Convolutional Nets Really Need to be Deep and Convolutional? · ICLR (Poster) 2017
Computer vision › Video understanding and tracking
video classification
0.312017
Fast Video Classification via Adaptive Cascading of Deep Models · CVPR 2017
Recommender systems
sequential decision making
0.312017
Fast Video Classification via Adaptive Cascading of Deep Models · CVPR 2017
Multimedia analysis and retrieval › event detection
video event detection
0.312017
Demo: Live Video Stream Triggers · MobiSys 2017
Hardware accelerators and domain-specific architectures
vision accelerator
0.312017
Glimpse: A Programmable Early-Discard Camera Architecture for Continuous Mobile Vision · MobiSys 2017
Machine learning › Graph learning
network analysis
0.212016
Analysis of Deep Neural Networks with Extended Data Jacobian Matrix · ICML 2016
Machine learning › Deep learning architectures and training
regularization
0.212016
Analysis of Deep Neural Networks with Extended Data Jacobian Matrix · ICML 2016
Edge and fog computing › mobile edge computing › computation offloading › inference offloading
DNN inference offloading
0.212016
MCDNN: An Approximation-Based Execution Framework for Deep Stream Processing Under Resource Constraints · MobiSys 2016
Edge and fog computing › mobile edge computing › computation offloading › mobile computation offloading
mobile cloud offloading
0.212016
MCDNN: An Approximation-Based Execution Framework for Deep Stream Processing Under Resource Constraints · MobiSys 2016
Computer vision › Image recognition and object detection
object detection
0.212014
Efficient Object Detection via Adaptive Online Selection of Sensor-Array Elements · AAAI 2014
Interaction techniques and input
voice interaction
0.212013
NLify: lightweight spoken natural language interfaces via exhaustive paraphrasing · UbiComp 2013
Query processing and optimization
approximate query processing
0.122010
Approximation trade-offs in Markovian stream processing: An empirical study · ICDE 2010
Lahar Demonstration: Warehousing Markovian Streams · Proc. VLDB Endow. 2009
Computational photography and imaging › image acquisition
image sensing
0.122013
Energy proportional image sensors for continuous mobile vision · MobiSys 2013
Energy characterization and optimization of image sensing toward continuous mobile vision · MobiSys 2013
Indexing and storage engines
access methods
0.112009
Access Methods for Markovian Streams · ICDE 2009

Methods — techniques the papers use, named apart from their topics

programmable camera pipeline · 0.9low-power vision modalities · 0.9co-scheduling · 0.8DNN fragment execution · 0.8heterogeneous binarization · 0.7bitwidth selection · 0.7standing query matching · 0.6convolutional neural network · 0.6bandit algorithms · 0.6approximation · 0.6optimizing compiler · 0.5folded CNN design · 0.5batch-size scaling · 0.5DNN approximation · 0.5shallow network training · 0.3knowledge distillation · 0.3runtime scheduler · 0.2hidden markov model · 0.2
YearPublicationVenuePosition
2026 Proofs Promptly: Proof-Oriented Programming with AI Agents (Experience Report)
abstract
The widespread adoption of AI-assisted coding is directly proportional to an increase in software bugs; can AI-assisted formal verification help reduce bugs at a comparable scale? In this experience report we give an anecdotal account of AI agents, equipped with a CLI and a proof assistant, producing thousands of lines of machine-checked code. We detail our experience across different proof-engineering tasks: implementing verified data structures for a standard library, translating unverified code into a formal language while inferring its specification, and porting and refactoring existing proofs to new frameworks. We present the techniques that made agentic proof-oriented programming (PoP) effective---or ineffective---and characterize the role of the human expert, whose contribution reduces to providing natural-language problem descriptions, reviewing auto-generated specifications, and occasionally supplying a key invariant. Our findings suggest that this division of labor provides substantial leverage to the human expert in the loop: three experts, over the course of two weeks, completed case studies whose manual proof-engineering cost we estimate at roughly half a year.
Eleftherios Ioannidis, Nikhil Swamy, Gabriel Ebner, Matthai Philipose, Tahina Ramananandro
Proc. ACM Program. Lang.4
2021 Boosting the Throughput and Accelerator Utilization of Specialized CNN Inference Beyond Increasing Batch Size
abstract
Datacenter vision systems widely use small, specialized convolutional neural networks (CNNs) trained on specific tasks for high-throughput inference. These settings employ accelerators with massive computational capacity, but which specialized CNNs underutilize due to having low arithmetic intensity. This results in suboptimal application-level throughput and poor returns on accelerator investment. Increasing batch size is the only known way to increase both application-level throughput and accelerator utilization for inference, but yields diminishing returns; specialized CNNs poorly utilize accelerators even with large batch size. We propose FoldedCNNs, a new approach to CNN design that increases inference throughput and utilization beyond large batch size. FoldedCNNs rethink the structure of inputs and layers of specialized CNNs to boost arithmetic intensity: in FoldedCNNs, f images with C channels each are concatenated into a single input with fC channels and jointly classified by a wider CNN. Increased arithmetic intensity in FoldedCNNs increases the throughput and GPU utilization of specialized CNN inference by up to 2.5x and 2.8x, with accuracy close to the original CNN in most cases.
Jack Kosaian, Amar Phanishayee, Matthai Philipose, Debadeepta Dey, K. V. Rashmi
ICML3
2019 Nexus: a GPU cluster engine for accelerating DNN-based video analysis
abstract
We address the problem of serving Deep Neural Networks (DNNs) efficiently from a cluster of GPUs. In order to realize the promise of very low-cost processing made by accelerators such as GPUs, it is essential to run them at sustained high utilization. Doing so requires cluster-scale resource management that performs detailed scheduling of GPUs, reasoning about groups of DNN invocations that need to be co-scheduled, and moving from the conventional whole-DNN execution model to executing fragments of DNNs. Nexus is a fully implemented system that includes these innovations. In large-scale case studies on 16 GPUs, when required to stay within latency constraints at least 99% of the time, Nexus can process requests at rates 1.8-12.7X higher than state of the art systems can. A long-running multi-application deployment stays within 84% of optimal utilization and, on a 100-GPU cluster, violates latency SLOs on 0.27% of requests.
Haichen Shen, Lequn Chen 0001, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, Ravi Sundaram
SOSP6
2018 Limiting Numerical Precision of Neural Networks to Achieve Real-Time Voice Activity Detection
abstract
Fast and robust voice-activity detection is critical to efficiently process speech. While deep-learning based methods to detect voice have shown competitive accuracies, the best models in the literature incur over a 100 ms latency on commodity processors. Such delays are unacceptable for real-time speech processing. In this paper, we study the impact of lowering the representation precision of the neural-network weights and neurons on both the accuracy and delay of voice-activity detection. Based on a design-space exploration, we not only determine the optimal scaling strategy but also adjust the network structure to accommodate the new quantization levels. Through experiments conducted with real user data, we demonstrate that optimized deep neural networks with lower bit precisions outperform the state-of-the-art WebRTC voice-activity detector with 87x lower delay and 6.8% lower error rate.
Jong Hwan Ko, Josh Fromm, Matthai Philipose, Ivan Tashev, Shuayb Zarar
ICASSP3
2018 Heterogeneous Bitwidth Binarization in Convolutional Neural Networks
abstract
Recent work has shown that fast, compact low-bitwidth neural networks can be surprisingly accurate. These networks use homogeneous binarization: all parameters in each layer or (more commonly) the whole model have the same low bitwidth (e.g., 2 bits). However, modern hardware allows efficient designs where each arithmetic instruction can have a custom bitwidth, motivating heterogeneous binarization, where every parameter in the network may have a different bitwidth. In this paper, we show that it is feasible and useful to select bitwidths at the parameter granularity during training. For instance a heterogeneously quantized version of modern networks such as AlexNet and MobileNet, with the right mix of 1-, 2- and 3-bit parameters that average to just 1.4 bits can equal the accuracy of homogeneous 2-bit versions of these networks. Further, we provide analyses to show that the heterogeneously binarized systems yield FPGA- and ASIC-based implementations that are correspondingly more efficient in both circuit area and energy efficiency than their homogeneous counterparts.
Josh Fromm, Shwetak N. Patel, Matthai Philipose
NeurIPS3
2018 Focus: Querying Large Video Datasets with Low Latency and Low Cost
Kevin Hsieh, Ganesh Ananthanarayanan, Peter Bodík, Shivaram Venkataraman, Paramvir Bahl, Matthai Philipose, Phillip B. Gibbons, Onur Mutlu
OSDI6
2017 Fast Video Classification via Adaptive Cascading of Deep Models
abstract
Recent advances have enabled oracle classifiers that can classify across many classes and input distributions with high accuracy without retraining. However, these classifiers are relatively heavyweight, so that applying them to classify video is costly. We show that day-to-day video exhibits highly skewed class distributions over the short term, and that these distributions can be classified by much simpler models. We formulate the problem of detecting the short-term skews online and exploiting models based on it as a new sequential decision making problem dubbed the Online Bandit Problem, and present a new algorithm to solve it. When applied to recognizing faces in TV shows and movies, we realize end-to-end classification speedups of 2.4-7.8x/2.6-11.2x (on GPU/CPU) relative to a state-of-the-art convolutional neural network, at competitive accuracy.
Haichen Shen, Seungyeop Han, Matthai Philipose, Arvind Krishnamurthy
CVPR3
2017 Do Deep Convolutional Nets Really Need to be Deep and Convolutional?
Gregor Urban, Krzysztof J. Geras, Samira Ebrahimi Kahou, Özlem Aslan, Shengjie Wang 0001, Abdel-rahman Mohamed, Matthai Philipose, Matthew Richardson, Rich Caruana
ICLR (Poster)7
2017 Glimpse: A Programmable Early-Discard Camera Architecture for Continuous Mobile Vision
abstract
We consider the problem of continuous computer-vision based analysis of video streams from mobile cameras over extended periods. Given high computational demands, general visual processing must currently be offloaded to the cloud. To reduce mobile battery and bandwidth consumption, recent proposals offload only "interesting" video frames, discarding the rest. However, determining what to discard is itself typically a power-hungry computer vision calculation, very often well beyond what most mobile devices can afford on a continuous basis. We present the Glimpse system, a re-design of the conventional mobile video processing pipeline to support such "early discard" flexibly, efficiently and accurately. Glimpse is a novel architecture that gates wearable vision using low-power vision modalities. Our proposed architecture adds novel sensing, processing, algorithmic and programming-system components to the camera pipeline to this end. We present a complete implementation and evaluation of our design. In common settings, Glimpse reduces mobile power and data usage by more than one order of magnitude relative to earlier designs, and moves continuous vision on lightweight wearables to the realm of the practical.
Saman Naderiparizi, Matthai Philipose, Bodhi Priyantha, Jie Liu 0001, Deepak Ganesan
MobiSys3
2017 Demo: Live Video Stream Triggers
abstract
Live streaming is an increasingly popular way to broadcast videos ranging from formal news channels to kitten cams to home security camera feeds. Live streaming marries the rich detail of video with the timeliness of live transmission and the ease of use of consumer cameras, thus promising to vastly increase the amount of detailed, up-to-the minute information available about the real world. The volume of potentially interesting footage brings up the question of how end-users can avoid being glued to one (or worse, many) streams of videos waiting for events of interest. In this demo, we present Lookout, a system that allows users to register standing queries, called triggers over live video streams. Lookout then notifies the user when events of interest to them occur in their streams of interest. For example, a user could point to a cat cam and write a trigger that sends a notification when the cat wakes up and starts moving. Users can also write triggers to look for certain news being covered in a live new channel, a gamer moving to a certain level in a Twitch stream, a stranger showing up in a outdoor surveillance camera, etc.
Lenin Ravindranath, Matthai Philipose, Peter Bodík, Paramvir Bahl
MobiSys2
2017 Live Video Analytics at Scale with Approximation and Delay-Tolerance
Ganesh Ananthanarayanan, Peter Bodík, Matthai Philipose, Paramvir Bahl, Michael J. Freedman
NSDI4
2016 Analysis of Deep Neural Networks with Extended Data Jacobian Matrix
abstract
Deep neural networks have achieved great successes on various machine learning tasks, however, there are many open fundamental questions to be answered. In this paper, we tackle the problem of quantifying the quality of learned wights of different networks with possibly different architectures, going beyond considering the final classification error as the only metric. We introduce \emphExtended Data Jacobian Matrix to help analyze properties of networks of various structures, finding that, the spectrum of the extended data jacobian matrix is a strong discriminating factor for networks of different structures and performance. Based on such observation, we propose a novel regularization method, which manages to improve the network performance comparably to dropout, which in turn verifies the observation.
Shengjie Wang 0001, Abdel-rahman Mohamed, Rich Caruana, Jeff A. Bilmes, Matthai Philipose, Matthew Richardson, Krzysztof J. Geras, Gregor Urban, Özlem Aslan
ICML5
2016 MCDNN: An Approximation-Based Execution Framework for Deep Stream Processing Under Resource Constraints
abstract
We consider applying computer vision to video on cloud-backed mobile devices using Deep Neural Networks (DNNs). The computational demands of DNNs are high enough that, without careful resource management, such applications strain device battery, wireless data, and cloud cost budgets. We pose the corresponding resource management problem, which we call Approximate Model Scheduling, as one of serving a stream of heterogeneous (i.e., solving multiple classification problems) requests under resource constraints. We present the design and implementation of an optimizing compiler and runtime scheduler to address this problem. Going beyond traditional resource allocators, we allow each request to be served approximately, by systematically trading off DNN classification accuracy for resource use, and remotely, by reasoning about on-device/cloud execution trade-offs. To inform the resource allocator, we characterize how several common DNNs, when subjected to state-of-the art optimizations, trade off accuracy for resource use such as memory, computation, and energy. The heterogeneous streaming setting is a novel one for DNN execution, and we introduce two new and powerful DNN optimizations that exploit it. Using the challenging continuous mobile vision domain as a case study, we show that our techniques yield significant reductions in resource usage and perform effectively over a broad range of operating conditions.
Seungyeop Han, Haichen Shen, Matthai Philipose, Sharad Agarwal, Alec Wolman, Arvind Krishnamurthy
MobiSys3
2015 SAPPHIRE: an always-on context-aware computer vision system for portable devices
Swagath Venkataramani, Paramvir Bahl, Xian-Sheng Hua 0001, Jie Liu 0001, Jin Li 0001, Matthai Philipose, Bodhi Priyantha, Mohammed Shoaib
DATE6
2014 Efficient Object Detection via Adaptive Online Selection of Sensor-Array Elements
abstract
We examine how to use emerging far-infrared imager ensembles to detect certain objects of interest (e.g., faces, hands, people and animals) in synchronized RGB video streams at very low power. We formulate the problem as one of selecting subsets of sensing elements (among many thousand possibilities) from the ensembles for tests. The subset selection problem is naturally adaptive and online: testing certain elements early can obviate the need for testing many others later, and selection policies must be updated at inference time. We pose the ensemble sensor selection problem as a structured extension of test-cost-sensitive classification, propose a principled suite of techniques to exploit ensemble structure to speed up processing and show how to re-estimate policies fast. We estimate reductions in power consumption of roughly 50x relative to even highly optimized implementations of face detection, a canonical object-detection problem. We also illustrate the benefits of adaptivity and online estimation.
Matthai Philipose
AAAI1
2014 Approximation trade-offs in a Markovian stream warehouse: An empirical study
Julie Letchner, Magdalena Balazinska, Christopher Ré, Matthai Philipose
Inf. Syst.4
2014 Recognizing object manipulation activities using depth and visual cues
Matthai Philipose, Martin Pettersson, Ming-Ting Sun
J. Vis. Commun. Image Represent.2
2014 Automatic objects segmentation with RGB-D cameras
Matthai Philipose, Ming-Ting Sun
J. Vis. Commun. Image Represent.2
2013 The Case for Onloading Continuous High-Datarate Perception to the Phone
Seungyeop Han, Matthai Philipose
HotOS2
2013 NLify: lightweight spoken natural language interfaces via exhaustive paraphrasing
abstract
This paper presents the design and implementation of a programming system that enables third-party developers to add spoken natural language (SNL) interfaces to standalone mobile applications. The central challenge is to create statistical recognition models that are accurate and resource-efficient in the face of the variety of natural language, while requiring little specialized knowledge from developers. We show that given a few examples from the developer, it is possible to elicit comprehensive sets of paraphrases of the examples using internet crowds. The exhaustive nature of these paraphrases allows us to use relatively simple, automatically derived statistical models for speech and language understanding that perform well without per-application tuning. We have realized our design fully as an extension to the Visual Studio IDE. Based on a new benchmark dataset with 3500 spoken instances of 27 commands from 20 subjects and a small developer study, we establish the promise of our approach and the impact of various design choices.
Seungyeop Han, Matthai Philipose, Yun-Cheng Ju
UbiComp2
2013 Energy characterization and optimization of image sensing toward continuous mobile vision
abstract
A major hurdle to frequently performing mobile computer vision tasks is the high power consumption of image sensing. In this work, we report the first publicly known experimental and analytical characterization of CMOS image sensors. We find that modern image sensors are not energy-proportional: energy per pixel is in fact inversely proportional to frame rate and resolution of image capture, and thus image sensor systems fail to provide an important principle of energy-aware system design: trading quality for energy efficiency. We reveal two energy-proportional mechanisms, supported by current image sensors but unused by mobile systems: (i) using an optimal clock frequency reduces the power up to 50% or 30% for low-quality single frame (photo) and sequential frame (video) capturing, respectively; (ii) by entering low-power standby mode between frames, an image sensor achieves almost constant energy per pixel for video capture at low frame rates, resulting in an additional 40% power reduction. We also propose architectural modifications to the image sensor that would further improve operational efficiency. Finally, we use computer vision benchmarks to show the performance and efficiency tradeoffs that can be achieved with existing image sensors. For image registration, a key primitive for image mosaicking and depth estimation, we can achieve a 96% success rate at 3 FPS and 0.1 MP resolution. At these quality metrics, an optimal clock frequency reduces image sensor power consumption by 36% and aggressive standby mode reduces power consumption by 95%.
Robert LiKamWa, Bodhi Priyantha, Matthai Philipose, Lin Zhong 0001, Paramvir Bahl
MobiSys3
2013 Energy proportional image sensors for continuous mobile vision
abstract
A hurdle to frequently performing mobile computer vision tasks is the high energy cost of image sensing. In particular, modern image sensors are not energy proportional; for low resolution and low frame rate capture, the image sensor consumes almost the same amount of energy as it does at high resolutions and high frame rates. We reveal two system-level energy proportional mechanisms: (i) using an optimal pixel clock frequency; (ii) entering low power standby mode between frames. These techniques can be implemented by the image sensor driver with minimal hardware adjustment. Further improvements can be made by designing sensors with heterogeneous hardware architectures. With energy proportionality, computer vision frameworks can be optimized for power consumption, continuously requesting low resolution frames with low energy while only occasionally using high energy to request high resolution frames. This will in turn enable low power continuous mobile vision applications.
Robert LiKamWa, Bodhi Priyantha, Matthai Philipose, Lin Zhong 0001, Paramvir Bahl
MobiSys3
2012 Automatic object segmentation with 3-D cameras
abstract
Recently, active 3-D cameras, which provide streams of depth and color images, have become widespread and popular. The depth data provides useful information for identifying object boundaries, making automatic object segmentation possible. However, the depth images are extremely noisy, and due to different response time of the color and depth sensors, the depth and color information often lose synchronization when the object is moving fast. In this work, we show how to combine depth and color information to clean up the depth maps and produce an accurate segmentation of the object. On a large dataset, we show that our proposed techniques are effective.
Matthai Philipose, Ming-Ting Sun
ICIP2
2011 Towards a Physical and Personal Math Coin Tutoring System
Georgios Theocharous, Nicholas J. Butko, Matthai Philipose
AIED3
2011 Automated facial affect analysis for one-on-one tutoring applications
abstract
In this paper, we explore the use of computer vision techniques to analyze students' moods during one-on-one teaching interactions. The eventual goal is to create automated tutoring systems that are sensitive to the student's mood and affective state. We find that the problem of accurately determining a child's mood from a single video frame is surprisingly difficult, even for humans. However when the system is allowed to make decisions based on information from 10 to 30 seconds of video, excellent performance may be obtained.
Nicholas J. Butko, Georgios Theocharous, Matthai Philipose, Javier R. Movellan
FG3
2011 Introduction to the special issue on intelligent systems for activity recognition
abstract
International audience
Daqing Zhang 0001, Matthai Philipose, Qiang Yang 0001
ACM Trans. Intell. Syst. Technol.2
2010 Approximation trade-offs in Markovian stream processing: An empirical study
abstract
A large amount of the world's data is both sequential and imprecise. Such data is commonly modeled as Markovian streams; examples include words/sentences inferred from raw audio signals, or discrete location sequences inferred from RFID or GPS data. The rich semantics and large volumes of these streams make them difficult to query efficiently. In this paper, we study the effects-on both efficiency and accuracy-of two common stream approximations. Through experiments on a realworld RFID data set, we identify conditions under which these approximations can improve performance by several orders of magnitude, with only minimal effects on query results. We also identify cases when the full rich semantics are necessary.
Julie Letchner, Christopher Ré, Magdalena Balazinska, Matthai Philipose
ICDE4
2009 Recognizing daily activities with RFID-based sensors
abstract
We explore a dense sensing approach that uses RFID sensor network technology to recognize human activities. In our setting, everyday objects are instrumented with UHF RFID tags called WISPs that are equipped with accelerometers. RFID readers detect when the objects are used by examining this sensor data, and daily activities are then inferred from the traces of object use via a Hidden Markov Model. In a study of 10 participants performing 14 activities in a model apartment, our approach yielded recognition rates with precision and recall both in the 90% range. This compares well to recognition with a more intrusive short-range RFID bracelet that detects objects in the proximity of the user; this approach saw roughly 95% precision and 60% recall in the same study. We conclude that RFID sensor networks are a promising approach for indoor activity monitoring.
Michael Buettner, Richa Prasad, Matthai Philipose, David Wetherall
UbiComp3
2009 Access Methods for Markovian Streams
abstract
Model-based views have recently been proposed as an effective method for querying noisy sensor data. Commonly used models from the AI literature (e.g., the hidden Markov model) expose to applications a stream of probabilistic and correlated state estimates computed from the sensor data. Many applications want to detect sophisticated patterns of states from these Markovian streams. Such queries are called event queries. In this paper, we present a new Markovian stream storage manager, Caldera. We develop and evaluate Caldera as a component of Lahar, a Markovian stream event query processing system developed in previous work. At the heart of Caldera is a set of access methods for Markovian streams that can improve event query performance by orders of magnitude compared to existing techniques, which must scan the entire stream. Our access methods use new adaptations of traditional B+ tree indexes, and a new index, called the Markov-chain index. They efficiently extract only the relevant timesteps from a stream, while retaining the stream's Markovian properties. We have implemented our prototype system on BDB and demonstrate its effectiveness on both synthetic data and real data from a building-wide RFID deployment.
Julie Letchner, Christopher Ré, Magdalena Balazinska, Matthai Philipose
ICDE4
2009 Bonfire: a nomadic system for hybrid laptop-tabletop interaction
abstract
We present Bonfire, a self-contained mobile computing system that uses two laptop-mounted laser micro-projectors to project an interactive display space to either side of a laptop keyboard. Coupled with each micro-projector is a camera to enable hand gesture tracking, object recognition, and information transfer within the projected space. Thus, Bonfire is neither a pure laptop system nor a pure tabletop system, but an integration of the two into one new nomadic computing platform. This integration (1) enables observing the periphery and responding appropriately, e.g., to the casual placement of objects within its field of view, (2) enables integration between physical and digital objects via computer vision, (3) provides a horizontal surface in tandem with the usual vertical laptop display, allowing direct pointing and gestures, and (4) enlarges the input/output space to enrich existing applications. We describe Bonfire's architecture, and offer scenarios that highlight Bonfire's advantages. We also include lessons learned and insights for further development and use.
Shaun K. Kane, Daniel Avrahami, Jacob O. Wobbrock, Beverly L. Harrison, Adam D. Rea, Matthai Philipose, Anthony LaMarca
UIST6
2009 Lahar Demonstration: Warehousing Markovian Streams
abstract
Lahar is a warehousing system for Markovian streams ---a common class of uncertain data streams produced via inference on probabilistic models. Example Markovian streams include text inferred from speech, location streams inferred from GPS or RFID readings, and human activity streams inferred from sensor data. Lahar supports OLAP-style queries on Markovian stream archives by leveraging novel approximation and indexing techniques that efficiently manipulate stream probabilities. This demonstration allows users to interactively query a warehouse of imprecise text streams inferred automatically from audio podcasts. Through this interaction, the demo introduces users to the challenges of Markovian stream processing as well as technical contributions developed to address these challenges.
Julie Letchner, Christopher Ré, Magdalena Balazinska, Matthai Philipose
Proc. VLDB Endow.4
2008 Structure Learning on Large Scale Common Sense Statistical Models of Human State
William Pentney, Matthai Philipose, Jeff A. Bilmes
AAAI2
2008 Relational Transformation-based Tagging for Activity Recognition
Niels Landwehr, Bernd Gutmann, Ingo Thon, Luc De Raedt, Matthai Philipose
Fundam. Informaticae5
2007 Learning Large Scale Common Sense Models of Everyday Life
William Pentney, Matthai Philipose, Jeff A. Bilmes, Henry A. Kautz
AAAI2
2007 A Long-Term Evaluation of Sensing Modalities for Activity Recognition
Beth Logan, Jennifer A. Healey, Matthai Philipose, Emmanuel Munguia Tapia, Stephen S. Intille
UbiComp3
2007 A Statistical Reasoning System for Medication Prompting
Sengul Vurgun, Matthai Philipose, Misha Pavel
UbiComp2
2007 A Scalable Approach to Activity Recognition based on Object Use
abstract
We propose an approach to activity recognition based on detecting and analyzing the sequence of objects that are being manipulated by the user. In domains such as cooking, where many activities involve similar actions, object-use information can be a valuable cue. In order for this approach to scale to many activities and objects, however, it is necessary to minimize the amount of human-labeled data that is required for modeling. We describe a method for automatically acquiring object models from video without any explicit human supervision. Our approach leverages sparse and noisy readings from RFID tagged objects, along with common-sense knowledge about which objects are likely to be used during a given activity, to bootstrap the learning process. We present a dynamic Bayesian network model which combines RFID and video data to jointly infer the most likely activity and object labels. We demonstrate that our approach can achieve activity recognition rates of more than 80% on a real-world dataset consisting of 16 household activities involving 33 objects with significant background clutter. We show that the combination of visual object recognition with RFID data is significantly more effective than the RFID sensor alone. Our work demonstrates that it is possible to automatically learn object models from video of household activities and employ these models for activity recognition, without requiring any explicit human labeling.
Jianxin Wu 0001, Adebola Osuntogun, Tanzeem Choudhury, Matthai Philipose, James M. Rehg
ICCV4
2007 Common Sense Based Joint Training of Human Activity Recognizers
Kai Wang 0059, William Pentney, Ana-Maria Popescu, Tanzeem Choudhury, Matthai Philipose
IJCAI5
2006 Sensor-Based Understanding of Daily Life via Large-Scale Use of Common Sense
William Pentney, Ana-Maria Popescu, Kai Wang 0059, Henry A. Kautz, Matthai Philipose
AAAI5
2005 Unsupervised Activity Recognition Using Automatically Mined Common Sense
Danny Wyatt, Matthai Philipose, Tanzeem Choudhury
AAAI2
2005 Maximum A Posteriori Path Estimation with Input Trace Perturbation: Algorithms and Application to Credible Rating of Human Routines
Daniel H. Wilson, Matthai Philipose
IJCAI2
2004 I Sense a Disturbance in the Force: Unobtrusive Detection of Interactions with RFID-tagged Objects
Kenneth P. Fishkin, Bing Jiang, Matthai Philipose, Sumit Roy 0001
UbiComp3
2004 Mapping and Localization with RFID Technology
abstract
We analyze whether radio frequency identification (RFID) technology can be used to improve the localization of mobile robots and persons in their environment. In particular we study the problem of localizing RFID tags with a mobile platform that is equipped with a pair of RFID antennas. We present a probabilistic measurement model for RFID readers that allow us to accurately localize RFID tags in the environment. We also demonstrate how such maps can be used to localize a robot and persons in their environment. Finally, we present experiments illustrating that the computational requirements for global robot localization can be reduced strongly by fusing RFID information with laser data.
Dirk Hähnel, Wolfram Burgard, Dieter Fox, Kenneth P. Fishkin, Matthai Philipose
ICRA5
2004 Mining models of human activities from the web
abstract
The ability to determine what day-to-day activity (such as cooking pasta, taking a pill, or watching a video) a person is performing is of interest in many application domains. A system that can do this requires models of the activities of interest, but model construction does not scale well: humans must specify low-level details, such as segmentation and feature selection of sensor data, and high-level structure, such as spatio-temporal relations between states of the model, for each and every activity. As a result, previous practical activity recognition systems have been content to model a tiny fraction of the thousands of human activities that are potentially useful to detect. In this paper, we present an approach to sensing and modeling activities that scales to a much larger class of activities than before. We show how a new class of sensors, based on Radio Frequency Identification (RFID) tags, can directly yield semantic terms that describe the state of the physical world. These sensors allow us to formulate activity models by translating labeled activities, such as 'cooking pasta', into probabilistic collections of object terms, such as 'pot'. Given this view of activity models as text translations, we show how to mine definitions of activities in an unsupervised manner from the web. We have used our technique to mine definitions for over 20,000 activities. We experimentally validate our approach using data gathered from actual human activity as well as simulated data.
Mike Perkowitz, Matthai Philipose, Kenneth P. Fishkin, Donald J. Patterson
WWW2
2002 Towards automatic construction of staged compilers
abstract
Some compilation systems, such as offline partial evaluators and selective dynamic compilation systems, support staged optimizations. A staged optimization is one where a logically single optimization is broken up into stages, with the early stage(s) performing preplanning set-up work, given any available partial knowledge about the program to be compiled, and the final stage completing the optimization. The final stage can be much faster than the original optimization by having much of its work performed by the early stages. A key limitation of current staged optimizers is that they are written by hand, sometimes in an ad hoc manner. We have developed a framework called the Staged Compilation Framework (SCF) for systematically and automatically converting single-stage optimizations into staged versions. The framework is based on a combination of aggressive partial evaluation and dead-assignment elimination. We have implemented SCF in Standard ML. A preliminary evaluation shows that SCF can speed up classical optimization of some commonly used C functions by up to 12× (and typically between 4.5× and 5.5×).
Matthai Philipose, Craig Chambers, Susan J. Eggers
POPL1
2000 DyC: an expressive annotation-directed dynamic compiler for C
Brian Grant, Markus Mock, Matthai Philipose, Craig Chambers, Susan J. Eggers
Theor. Comput. Sci.3
2000 The benefits and costs of DyC's run-time optimizations
abstract
DyC selectively dynamically compiles programs during their execution, utilizing the run-time-computed values of variables and data structures to apply optimizations that are based on partial evaluation. The dynamic optimizations are preplanned at static compile time in order to reduce their run-time cost; we call this staging . DyC's staged optimizations include (1) an advanced binding-time analysis that supports polyvariant specialization (enabling both single-way and multiway complete loop unrolling), polyvariant division, static loads, and static calls, (2) low-cost, dynamic versions of traditional global optimizations, such as zero and copy propagation and dead-assignment elimination, and (3) dynamic peephole optimizations, such as strength reduction. Because of this large suite of optimizations and its low dynamic compilation overhead, DyC achieves good performance improvements on programs that are larger and more complex than the kernels previously targeted by other dynamic compilation systems. This paper evaluates the benefits and costs of applying DyC's optimizations. We assess their impact on the performance of a variety of small to medium-sized programs, both for the regions of code that are actually transformed and for the entire application as a whole. Our study includes an analysis of the contribution to performance of individual optimizations, the performance effect of changing the applications' inputs, and a detailed accounting of dynamic compilation costs.
Brian Grant, Markus Mock, Matthai Philipose, Craig Chambers, Susan J. Eggers
ACM Trans. Program. Lang. Syst.3
1999 An Evaluation of Staged Run-Time Optimizations in DyC
abstract
Previous selective dynamic compilation systems have demonstrated that dynamic compilation can achieve performance improvements at low cost on small kernels, but they have had difficulty scaling to larger programs. To overcome this limitation, we developed DyC, a selective dynamic compilation system that includes more sophisticated and flexible analyses and transformations. DyC is able to achieve good performance improvements on programs that are much larger and more complex than the kernels. We analyze the individual optimizations of DyC and assess their impact on performance collectively and individually.
Brian Grant, Matthai Philipose, Markus Mock, Craig Chambers, Susan J. Eggers
PLDI2
1997 Annotation-Directed Run-Time Specialization in C
abstract
We present the design of a dynamic compilation system for C. Directed by a few declarative user annotations specifying where and on what dynamic compilation is to take place, a binding time analysis computes the set of run-time constants at each program point in each annotated procedure's control flow graph; the analysis supports program-point-specific polyvariant division and specialization. The analysis results guide the construction of a specialized run-time specializer for each dynamically compiled region; the specializer supports various caching strategies for managing dynamically generated code and supports mixes of speculative and demand-driven specialization of dynamic branch successors. Most of the key cost/benefit trade-offs in the binding time analysis and the run-time specialize are open to user control through declarative policy annotations. Our design is being implemented in the context of art existing optimizing compiler.
Brian Grant, Markus Mock, Matthai Philipose, Craig Chambers, Susan J. Eggers
PEPM3
1996 Fast, Effective Dynamic Compilation
abstract
Dynamic compilation enables optimization based on the values of invariant data computed at run-time. Using the values of these run-time constants, a dynamic compiler can eliminate their memory loads, perform constant propagation and folding, remove branches they determine, and fully unroll loops they bound. However, the performance benefits of the more efficient, dynamically-compiled code are offset by the run-time cost of the dynamic compile. Our approach to dynamic compilation strives for both fast dynamic compilation and high-quality dynamically-compiled code: the programmer annotates regions of the programs that should be compiled dynamically; a static, optimizing compiler automatically produces pre-optimized machine-code templates, using a pair of dataflow analyses that identify which variables will be constant at run-time; and a simple, dynamic compiler copies the templates, patching in the computed values of the run-time constants, to produce optimized, executable code. Our work targets general- purpose, imperative programming languages, initially C. Initial experiments applying dynamic compilation to C programs have produced speedups ranging from 1.2 to 1.8.
Joel Auslander, Matthai Philipose, Craig Chambers, Susan J. Eggers, Brian N. Bershad
PLDI2