EDBT 2026 Demo / reviewers in the wild / expert
Philip David
dblp:31/6188
· DBLP profile ↗
20ranked-venue papers
5as first author
5since 2021 · last 2023
0000-0003-4702-482XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-authorSystems, architecture and hardware · 9 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Generalized self-cueing real-time attention scheduling with intermittent inspection and image resizing
Shengzhong Liu, Xinzhe Fu, Yigong Hu, Maggie B. Wigness, Philip David, Shuochao Yao, Lui Sha, Tarek F. Abdelzaher |
Real Time Syst. | 5 |
| 2022 | Multi-View Scheduling of Onboard Live Video Analytics to Minimize Frame Processing LatencyabstractThis paper presents a real-time multi-view scheduling framework for DNN-based live video analytics at the edge to minimize frame processing latency. The work is motivated by applications where a higher frame rate is important, not to miss actions of interest. Examples include defense, border security, and intruder detection applications where sensors (in this paper, cameras) are deployed to monitor key roads, chokepoints, or passageways to identify events of interest (and intervene in real-time). Supporting a higher frame rate entails lowering frame processing latency. We assume that multiple cameras are deployed with partially overlapping views. Each camera has access to limited onboard computing capacity. Many targets cross the field of view of these cameras (but the great majority do not require action). We take advantage of the spatial-temporal correlations among multi-camera video streams to perform target-to-camera assignment such that the maximum frame processing time across cameras is minimized. Specifically, we use a data-driven approach to identify objects seen by multiple cameras, and propose a batch-aware latency-balanced (BALB) scheduling algorithm to drive the object-to-camera assignment. We empirically evaluate the proposed system with a real-world surveillance dataset on a testbed consisting of multiple NVIDIA Jetson boards. The results show that our system substantially improves the video processing speed, attaining multiplicative speedups of 2.45× to 6.85×, and consistently outperforms the competitive static region partitioning strategy. Shengzhong Liu, Tianshi Wang 0002, Hongpeng Guo, Xinzhe Fu, Philip David, Maggie B. Wigness, Archan Misra, Tarek F. Abdelzaher |
ICDCS | 5 |
| 2022 | Self-Cueing Real-Time Attention Scheduling in Criticality-Aware Visual Machine PerceptionabstractThis paper presents a self-cueing real-time frame-work for attention prioritization in AI-enabled visual perception systems that minimizes a notion of state uncertainty. By attention prioritization we refer to inspecting some parts of the scene before others in a criticality-aware fashion. By self-cueing, we refer to not needing external cueing sensors for prioritizing attention, thereby simplifying design. We show that attention prioritization saves resources, thus enabling more efficient and responsive real-time object tracking on resource-limited embedded platforms. The system consists of two components: First, an optical flow-based module decides on the regions to be viewed on a subframe level, as well as their criticality. Second, a novel batched proportional balancing (BPB) scheduling policy decides how to schedule these regions for inspection by a deep neural network (DNN), and how to parallelize execution on the GPU. We implement the system on an NVIDIA Jetson Xavier platform, and empirically demonstrate the superiority of the proposed architecture through an extensive evaluation using a real-word driving dataset. Shengzhong Liu, Xinzhe Fu, Maggie B. Wigness, Philip David, Shuochao Yao, Lui Sha, Tarek F. Abdelzaher |
RTAS | 4 |
| 2022 | Real-time task scheduling with image resizing for criticality-based machine perception
Yigong Hu, Shengzhong Liu, Tarek F. Abdelzaher, Maggie B. Wigness, Philip David |
Real Time Syst. | 5 |
| 2021 | On Exploring Image Resizing for Optimizing Criticality-based Machine PerceptionabstractOn-board computing capacity remains a key bottleneck in modern machine inference pipelines that run on embedded hardware, such as aboard autonomous drones or cars. To mitigate this bottleneck, recent work proposed an architecture for segmenting input frames of complex modalities, such as video, and prioritizing downstream machine perception tasks based on criticality of the respective segments of the perceived scene. Criticality-based prioritization allows limited machine resources (of lower-end embedded GPUs) to be spent more judiciously on tracking more important objects first. This paper explores a novel dimension in criticality-based prioritization of machine perception; namely, the role of criticality-dependent image resizing as a way to improve the trade-off between perception quality and timeliness. Given an assessment of criticality (e.g., an object’s distance from the autonomous car), the scheduler is allowed to choose from several image resizing options (and related inference models) before passing the resized images to the perception module. Experiments on an AI-powered embedded platform with a real-world driving dataset demonstrate significant improvements in the trade-off between perception accuracy and response time when the proposed resizing algorithm is used. The improvement is attributed to two advantages of the proposed scheme: (i) improved preferential treatment of more critical objects by reducing time spent on less critical ones, and (ii) improved image batching within the GPU, thanks to re-sizing, leading to better resource utilization. Yigong Hu, Shengzhong Liu, Tarek F. Abdelzaher, Maggie B. Wigness, Philip David |
RTCSA | 5 |
| 2020 | PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic SegmentationabstractThe requirement of fine-grained perception by autonomous driving systems has resulted in recently increased research in the online semantic segmentation of single-scan LiDAR. Emerging datasets and technological advancements have enabled researchers to benchmark this problem and improve the applicable semantic segmentation algorithms. Still, online semantic segmentation of LiDAR scans in autonomous driving applications remains challenging due to three reasons: (1) the need for near-real-time latency with limited hardware, (2) points are distributed unevenly across space, and (3) an increasing number of more fine-grained semantic classes. The combination of the aforementioned challenges motivates us to propose a new LiDAR-specific, KNN-free segmentation algorithm - PolarNet. Instead of using common spherical or bird's-eye-view projection, our polar bird's-eye-view representation balances the points per grid and thus indirectly redistributes the network's attention over the long-tailed points distribution over the radial axis in polar coordination. We find that our encoding scheme greatly increases the mIoU in three drastically different real urban LiDAR single-scan segmentation datasets while retaining ultra low latency and near real-time throughput. Yang Zhang 0035, Philip David, Xiangyu Yue 0001, Zerong Xi, Boqing Gong, Hassan Foroosh |
CVPR | 3 |
| 2020 | A Curriculum Domain Adaptation Approach to the Semantic Segmentation of Urban ScenesabstractDuring the last half decade, convolutional neural networks (CNNs) have triumphed over semantic segmentation, which is one of the core tasks in many applications such as autonomous driving and augmented reality. However, to train CNNs requires a considerable amount of data, which is difficult to collect and laborious to annotate. Recent advances in computer graphics make it possible to train CNNs on photo-realistic synthetic imagery with computer-generated annotations. Despite this, the domain mismatch between real images and the synthetic data hinders the models' performance. Hence, we propose a curriculum-style learning approach to minimizing the domain gap in urban scene semantic segmentation. The curriculum domain adaptation solves easy tasks first to infer necessary properties about the target domain; in particular, the first task is to learn global label distributions over images and local distributions over landmark superpixels. These are easy to estimate because images of urban scenes have strong idiosyncrasies (e.g., the size and spatial relations of buildings, streets, cars, etc.). We then train a segmentation network, while regularizing its predictions in the target domain to follow those inferred properties. In experiments, our method outperforms the baselines on two datasets and three backbone networks. We also report extensive ablation studies about our approach. Yang Zhang 0035, Philip David, Hassan Foroosh, Boqing Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | CAMOU: Learning Physical Vehicle Camouflages to Adversarially Attack Detectors in the Wild
Yang Zhang 0035, Hassan Foroosh, Philip David, Boqing Gong |
ICLR (Poster) | 3 |
| 2018 | FarSight: Long-Range Depth Estimation from Outdoor ImagesabstractThis paper introduces the problem of long-range monocular depth estimation for outdoor urban environments. Range sensors and traditional depth estimation algorithms (both stereo and single view) predict depth for distances of less than 100 meters in outdoor settings and 10 meters in indoor settings. The shortcomings of outdoor single view methods that use learning approaches are, to some extent, due to the lack of long-range ground truth training data, which in turn is due to limitations of range sensors. To circumvent this, we first propose a novel strategy for generating synthetic long-range ground truth depth data. We utilize Google Earth images to reconstruct large-scale 3D models of different cities with proper scale. The acquired repository of 3D models and associated RGB views along with their long-range depth renderings are used as training data for depth prediction. We then train two deep neural network models for long-range depth estimation: i) a Convolutional Neural Network (CNN) and ii) a Generative Adversarial Network (GAN). We found in our experiments that the GAN model predicts depth more accurately. We plan to open-source the database and the baseline models for public use. Md. Alimoor Reza, Jana Kosecka, Philip David |
IROS | 3 |
| 2017 | Curriculum Domain Adaptation for Semantic Segmentation of Urban ScenesabstractDuring the last half decade, convolutional neural networks (CNNs) have triumphed over semantic segmentation, which is a core task of various emerging industrial applications such as autonomous driving and medical imaging. However, to train CNNs requires a huge amount of data, which is difficult to collect and laborious to annotate. Recent advances in computer graphics make it possible to train CNN models on photo-realistic synthetic data with computer-generated annotations. Despite this, the domain mismatch between the real images and the synthetic data significantly decreases the models' performance. Hence we propose a curriculum-style learning approach to minimize the domain gap in semantic segmentation. The curriculum domain adaptation solves easy tasks first in order to infer some necessary properties about the target domain; in particular, the first task is to learn global label distributions over images and local distributions over landmark superpixels. These are easy to estimate because images of urban traffic scenes have strong idiosyncrasies (e.g., the size and spatial relations of buildings, streets, cars, etc.). We then train the segmentation network in such a way that the network predictions in the target domain follow those inferred properties. In experiments, our method significantly outperforms the baselines as well as the only known existing approach to the same problem. Yang Zhang 0035, Philip David, Boqing Gong |
ICCV | 2 |
| 2016 | Ray Saliency: Bottom-Up Visual Saliency for a Rotating and Zooming Camera
Garrett Warnell, Philip David, Rama Chellappa |
Int. J. Comput. Vis. | 2 |
| 2013 | Ascending stairway modeling from dense depth imagery for traversability analysisabstractLocalization and modeling of stairways by mobile robots can enable multi-floor exploration for those platforms capable of stair traversal. Existing approaches focus on either stairway detection or traversal, but do not address these problems in the context of path planning for the autonomous exploration of multi-floor buildings. We propose a system for detecting and modeling ascending stairways while performing simultaneous localization and mapping, such that the traversability of each stairway can be assessed by estimating its physical properties. The long-term objective of our approach is to enable exploration of multiple floors of a building by allowing stairways to be considered during path planning as traversable portals to new frontiers. We design a generative model of a stairway as a single object. We localize these models with respect to the map, and estimate the dimensions of the stairway as a whole, as well as its steps. With these estimates, a robot can determine if the stairway is traversable based on its climbing capabilities. Our system consists of two parts: a computationally efficient detector that leverages geometric cues from dense depth imagery to detect sets of ascending stairs, and a stairway modeler that uses multiple detections to infer the location and parameters of a stairway that is discovered during exploration. We demonstrate the performance of this system when deployed on several mobile platforms using a Microsoft Kinect sensor. Jeffrey A. Delmerico, David Baran, Philip David, Julian Ryde, Jason J. Corso |
ICRA | 3 |
| 2013 | Building facade detection, segmentation, and parameter estimation for mobile robot stereo vision
Jeffrey A. Delmerico, Philip David, Jason J. Corso |
Image Vis. Comput. | 2 |
| 2012 | Ascending stairway modeling: A first step toward autonomous multi-floor explorationabstractMany robotics platforms are capable of ascending stairways, but all existing approaches for autonomous stair climbing use stairway detection as a trigger for immediate traversal. In the broader context of autonomous exploration, the ability to travel between floors of a building should be compatible with path planning, such that the robot can traverse a stairway at a time that is appropriate to its navigation goals. No system yet presented is capable of both localizing stairways on a map and estimating their properties, functions that in combination would enable stairways to be considered as traversable terrain in a path planning algorithm. We propose a method for modeling stairways as objects and localizing them on a map, such that they can be subsequently traversed if they are of dimensions that the robotic platform is capable of climbing. Our system consists of two parts: a computationally efficient detector that leverages geometric cues from depth imagery to detect sets of ascending stairs, and a stairway modeler that uses multiple detections to infer the location and parameters of a stairway that is discovered during exploration. This video demonstrates the performance of the system in a number of real-world situations, modeling and localizing a variety of stairway types in both indoor and outdoor environments. Jeffrey A. Delmerico, Jason J. Corso, David Baran, Philip David, Julian Ryde |
IROS | 4 |
| 2011 | Orientation descriptors for localization in urban environmentsabstractAccurately determining the position and orientation of an observer (a vehicle or a human) in outdoor urban environments is an important and challenging problem. The standard approach is to use the Global Positioning System (GPS), but this system performs poorly near tall buildings where line of sight to a sufficient number of satellites cannot be obtained. Most previous vision-based approaches for localization register ground imagery to a previously generated ground-level model of the environment. Generating such a model can be difficult and time consuming, and is impractical in some environments. Instead, we propose to perform localization by registering a single omnidirectional ground image to a 2D urban terrain model that is easily generated from aerial imagery. We introduce a novel image descriptor that encodes the position and orientation of a camera relative to buildings in the environment. The descriptor is efficiently generated from edges and vanishing points in an omnidirectional image and is registered to descriptors previously generated for the terrain model. Rather than constructing a local CAD-like model of the environment, which is difficult in cluttered environments, our descriptor measures, at equally spaced intervals over the 360° field of view, the orientation of visible building facades projected onto the ground plane (i.e., the building footprints). We evaluate our approach on an urban data set with significant clutter and demonstrate an accuracy of about 1 m, which is an order of magnitude better than commercial GPS operating in open environments. Philip David, Sean Ho |
IROS | 1 |
| 2011 | Building facade detection, segmentation, and parameter estimation for mobile robot localization and guidanceabstractBuilding facade detection is an important problem in computer vision, with applications in mobile robotics and semantic scene understanding. In particular, mobile platform localization and guidance in urban environments can be enabled with an accurate segmentation of the various building facades in a scene. Toward that end, we present a system for segmenting and labeling an input image that for each pixel, seeks to answer the question ¿Is this pixel part of a building facade, and if so, which one?¿ The proposed method determines a set of candidate planes by sampling and clustering points from the image with Random Sample Consensus (RANSAC), using local normal estimates derived from Principal Component Analysis (PCA) to inform the planar model. The corresponding disparity map and a discriminative classification provide prior information for a two-layer Markov Random Field model. This MRF problem is solved via Graph Cuts to obtain a labeling of building facade pixels at the mid-level, and a segmentation of those pixels into particular planes at the high-level. The results indicate a strong improvement in the accuracy of the binary building detection problem over the discriminative classifier alone, and the planar surface estimates provide a good approximation to the ground truth planes. Jeffrey A. Delmerico, Philip David, Jason J. Corso |
IROS | 2 |
| 2005 | Object Recognition in High Clutter Images Using Line FeaturesabstractWe present an object recognition algorithm that uses model and image line features to locate complex objects in high clutter environments. Finding correspondences between model and image features is the main challenge in most object recognition systems. In our approach, corresponding line features are determined by a three-stage process. The first stage generates a large number of approximate pose hypotheses from correspondences of one or two lines in the model and image. Next, the pose hypotheses from the previous stage are quickly ranked by comparing local image neighborhoods to the corresponding local model neighborhoods. Fast nearest neighbor and range search algorithms are used to implement a distance measure that is unaffected by clutter and partial occlusion. The ranking of pose hypotheses is invariant to changes in image scale, orientation, and partially invariant to affine distortion. Finally, a robust pose estimation algorithm is applied for refinement and verification, starting from the few best approximate poses produced by the previous stages. Experiments on real images demonstrate robust recognition of partially occluded objects in very high clutter environments Philip David, Daniel DeMenthon |
ICCV | 1 |
| 2004 | SoftPOSIT: Simultaneous Pose and Correspondence Determination
Philip David, Daniel DeMenthon, Ramani Duraiswami, Hanan Samet |
Int. J. Comput. Vis. | 1 |
| 2003 | Simultaneous Pose and Correspondence Determination using Line FeatureabstractWe present a new robust line matching algorithm for solving the model-to-image registration problem. Given a model consisting of 3D lines and a cluttered perspective image of this model, the algorithm simultaneously estimates the pose of the model and the correspondences of model lines to image lines. The algorithm combines softassign for determining correspondences and POSIT for determining pose. Integrating these algorithms into a deterministic annealing procedure allows the correspondence and pose to evolve from initially uncertain values to a joint local optimum. This research extends to line features the SoftPOSIT algorithm proposed recently for point features. Lines detected in images are typically more stable than points and are less likely to be produced by clutter and noise, especially in man-made environments. Experiments on synthetic and real imagery with high levels of clutter, occlusion, and noise demonstrate the robustness of the algorithm. Philip David, Daniel DeMenthon, Ramani Duraiswami, Hanan Samet |
CVPR (2) | 1 |
| 2002 | SoftPOSIT: Simultaneous Pose and Correspondence Determination
Philip David, Daniel DeMenthon, Ramani Duraiswami, Hanan Samet |
ECCV (3) | 1 |