EDBT 2026 Demo / reviewers in the wild / expert
Avideh Zakhor
dblp:28/1974
· DBLP profile ↗
150ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0003-4770-6353ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 114 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 13 · 3 since 2021Computer networks · 13Systems, architecture and hardware · 11 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Swap Path Network for Robust Person Search Pre-trainingabstractIn person search, we detect and rank matches to a query person image within a set of gallery scenes. Most person search models make use of a feature extraction back-bone, followed by separate heads for detection and re-identification. While pretraining methods for vision back-bones are well-established, pretraining additional modules for the person search task has not been previously exam-ined. In this work, we present the first framework for end-to-end person search pretraining. Our framework splits person search into object-centric and query-centric methodolo-gies, and we show that the query-centric framing is robust to label noise, and trainable using only weakly-labeled person bounding boxes. Further, we provide a novel model dubbed Swap Path Net (SPNet) which implements both query-centric and object-centric training objectives, and can swap between the two while using the same weights. Using SPNet, we show that query-centric pretraining, followed by object-centric fine-tuning, achieves state-of-the-art results on the standard PRW and CUHK-SYSU person search benchmarks, with 96.4% mAP on CUHK-SYSU and 61.2% mAP on PRW. In addition, we show that our method is more effective, efficient, and robust for person search pretraining than recent backbone-only pretraining alter-natives. Lucas Jaffe, Avideh Zakhor |
WACV | 2 |
| 2024 | Versatile Locomotion Skills for Hexapod RobotsabstractHexapod robots are potentially suitable for carrying out tasks in cluttered environments since they are stable, compact, and light weight. They also have multi-joint legs and variable height bodies that make them good candidates for tasks such as stairs climbing and squeezing under objects in a typical home environment or an attic. Expanding on our previous work on joist climbing in attics, we train a legged hexapod equipped with a depth camera and visual inertial odometry (VIO) to perform three tasks: climbing stairs, avoiding obstacles, and squeezing under obstacles such as a table. Our policies are trained with simulation data only and can be deployed on low-cost hardware not requiring real-time joint state feedback. We train our model in a teacher-student model with 2 phases: In phase 1, we use reinforcement learning with access to privileged information such as height maps and joint feedback. In phase 2, we use supervised learning to distill the model into one with access to only onboard observations, consisting of egocentric depth images and robot pose captured by a tracking VIO camera. By manipulating available privileged information, constructing simulation terrains, and refining reward functions during phase 1 training, we are able to train the robots with skills that are robust in non-ideal physical environments. We demonstrate successful sim-to-real transfer and achieve high success rates across all three tasks in physical experiments. Tomson Qu, Dichen Li, Avideh Zakhor, Wenhao Yu 0003, Tingnan Zhang |
IROS | 3 |
| 2024 | Monocular Depth Estimation for Drone Obstacle Avoidance in Indoor EnvironmentsabstractAutonomous nano-quadcopters possess large potential for indoor use. Existing works on autonomous flight however rely on large amounts of compute, therefore resulting in heavy and bulky platforms that can only be safely deployed outdoors. We present a monocular depth estimation method for autonomous indoor obstacle avoidance and waypoint navigation of nano-quadcopters demonstrated on the Bitcraze Crazyflie 2.1 which weighs a mere 33g. Our depth estimation model has 1.56 million parameters and is 4 MB, which after quantization becomes 1 MB. We transmit the images via WiFi from the onboard grayscale camera on the Bitcraze to a laptop, which then runs the 1 MB quantized model to generate small-size depth maps. Subsequently, we run our navigation algorithms on a laptop and transmit high-level motion commands back to the drone. We demonstrate obstacle avoidance capability of this end-to-end system through real-world flights in a variety of indoor environments. Haokun Zheng, Sidhant Rajadnya, Avideh Zakhor |
IROS | 3 |
| 2023 | Close-Range Indoor Proximity Detection for COVID-19 Exposure Notifications Using Smartphone Magnetometer TracesabstractSmartphone apps for exposure notification and contact tracing have been shown to be effective in controlling the COVID-19 pandemic. However, there is evidence that the Bluetooth Low Energy approach used for proximity detection by existing apps can be error-prone in areas with high numbers of metallic objects. In this paper, we present a new method for detecting whether or not two smartphones are 2 or fewer meters apart, intended to augment BLE-based proximity detection methods. We design a set of binary machine learning classifiers that take as input pairs of 10-second-long segments of magnetometer traces. These classifiers identify pairs of trace segments for which the two recording devices were 2 or fewer meters apart for at least 75% of the segment duration. We introduce a simple method of compensating for different magnetometer biases in heterogeneous devices. We confirm that our classifiers can generalize well to both new buildings and new devices whose traces are not present in their training data, and characterize their overall accuracy for heterogeneous-device proximity detection to be between 93.1% and 96.2%. Zach Van Hyfte, Avideh Zakhor |
IPIN | 2 |
| 2023 | Perceptive Hexapod Legged Locomotion for Climbing Joist EnvironmentsabstractAttics are one of the largest sources of energy loss in residential homes, but they are uncomfortable and dangerous for human workers to conduct air sealing and insulation. Hexapod robots are potentially suitable for carrying out those tasks in tight attic spaces since they are stable, compact, and lightweight. For hexapods to succeed in these tasks, they must be able to navigate inside tight attic spaces of single-family residential homes in the U.S., which typically contain rows of approximately 6 or 8-inch tall joists placed 16 inches apart from each other. Climbing over such obstacles is challenging for autonomous robotics systems. In this work, we develop a perceptive walking model for legged hexapods that can traverse over terrain with random joist structures using egocentric vision. Our method can be used on low-cost hardware not requiring real-time joint state feedback. We train our model in a teacher-student fashion with 2 phases: In phase 1, we use reinforcement learning with access to privileged information such as local elevation maps and joint feedback. In phase 2, we use supervised learning to distill the model into one with access to only onboard observations, consisting of egocentric depth images and robot orientation captured by a tracking camera. We demonstrate zero-shot sim-to-real transfer on a Hiwonder[1] SpiderPi robot, equipped with a depth camera onboard, climbing over joist courses we construct to simulate the environment in the field. Our proposed method achieves nearly 100% success rate climbing over the test courses, significantly outperforming the model without perception and the controller provided by the manufacturer. Zixian Zang, Maxime Kawawa-Beaudan, Wenhao Yu 0003, Tingnan Zhang, Avideh Zakhor |
IROS | 5 |
| 2023 | RSF: Optimizing Rigid Scene Flow From 3D Point Clouds Without LabelsabstractWe present a method for optimizing object-level rigid 3D scene flow over two successive point clouds without any annotated labels in autonomous driving settings. Rather than using pointwise flow vectors, our approach represents scene flow as the composition a global ego-motion and a set of bounding boxes with their own rigid motions, exploiting the multi-body rigidity commonly present in dynamic scenes. We jointly optimize these parameters over a novel loss function based on the nearest neighbor distance using a differentiable bounding box formulation. Our approach achieves state-of-the-art accuracy on KITTI Scene Flow and nuScenes without requiring any annotations, outperforming even supervised methods. Additionally, we demonstrate the effectiveness of our approach on motion segmentation and ego-motion estimation. Lastly, we visualize our predictions and validate our loss function design with an ablation study. David Deng, Avideh Zakhor |
WACV | 2 |
| 2023 | Gallery Filter Network for Person SearchabstractIn person search, we aim to localize a query person from one scene in other gallery scenes. The cost of this search operation is dependent on the number of gallery scenes, making it beneficial to reduce the pool of likely scenes. We describe and demonstrate the Gallery Filter Network (GFN), a novel module which can efficiently discard gallery scenes from the search process, and benefit scoring for persons detected in remaining scenes. We show that the GFN is robust under a range of different conditions by testing on different retrieval sets, including cross-camera, occluded, and low-resolution scenarios. In addition, we develop the base SeqNeXt person search model, which improves and simplifies the original SeqNet model. We show that the SeqNeXt+GFN combination yields significant performance gains over other state-of-the-art methods on the standard PRW and CUHK-SYSU person search datasets. To aid experimentation for this and other models, we provide standardized tooling for the data processing and evaluation pipeline typically used for person search research. Lucas Jaffe, Avideh Zakhor |
WACV | 2 |
| 2022 | Temporal Axial Attention For Lidar-Based 3d Object Detection In Autonomous Drivingabstract3D object detection is a core problem of the perception systems of autonomous vehicles. Despite recent progress in the field, the temporal aspect of LiDAR data has not been fully explored in current state-of-the-art detectors. This work proposes a modified CenterPoint architecture that uses temporal axial attention to exploit the sequential nature of autonomous driving data for 3D object detection. The last ten LiDAR sweeps are split into three groups of frames, and the axial attention transformer block captures both spatial and temporal dependencies among the features extracted from each group. Our proposal is evaluated using the nuScenes dataset. With this novel approach, we obtain an average mAP improvement of 3.8 and 2.3 points over the original CenterPoint in the fine/coarse pillar settings, respectively. Manuel Carranza-García, José Cristóbal Riquelme Santos, Avideh Zakhor |
ICIP | 3 |
| 2022 | Multi-modal Semantic Inconsistency Detection in Social Media News Posts
Scott McCrae, Avideh Zakhor |
MMM (2) | 3 |
| 2022 | Multimodal Semantic Mismatch Detection in Social Media PostsabstractShort videos have become the most popular form of social media in recent years. In this work, we focus on the threat scenario where video, audio, and their text description are semantically mismatched to mislead the audience. We develop self-supervised methods to detect semantic mismatch across multiple modalities, namely video, audio and text. We use state-of-the-art language, video and audio models to extract dense features from each modality, and explore transformer architecture together with contrastive learning methods on a dataset of one million Twitter posts from 2021 to 2022. Our best-performing method benefits from the robustness of Noise-Contrastive loss and the context provided by fusing modalities together using a cross-transformer. It outperforms state-of-the-art by over 9% in accuracy. We further characterize the performance of our system on topic-specific datasets containing COVID-19 and Russia-Ukraine related tweets, and shows that it outperforms state-of-the-art by over 17% in accuracy. Seth Z. Zhao, Avideh Zakhor, John F. Canny |
MMSP | 4 |
| 2021 | Fast, Accurate Barcode Detection in Ultra High-Resolution ImagesabstractObject detection in Ultra High-Resolution (UHR) images has long been a challenging problem in computer vision due to the varying scales of the targeted objects. When it comes to barcode detection, resizing UHR input images to smaller sizes often leads to the loss of pertinent information, while processing them directly is highly in-efficient and computationally expensive. In this paper, we propose using semantic segmentation to achieve a fast and accurate detection of barcodes of various scales in UHR images. Our pipeline involves a modified Region Proposal Network (RPN) on images of size greater than $10k \times10k$ and a newly proposed Y-Net segmentation network, followed by a post-processing workflow for fitting a bounding box around each segmented barcode mask. The end-to-end system has a latency of 16 milliseconds, which is $2.5 \times$ faster than YOLOv4 and $5.9 \times$ faster than Mask R-CNN. In terms of accuracy, our method outperforms YOLOv4 and Mask R-CNN by a mAP of 5.5% and 47.1% respectively, on a synthetic dataset. We have made available the generated synthetic barcode dataset and its code at http://www.github.com/viplabB/SBD/. Jerome Quenum, Avideh Zakhor |
ICIP | 3 |
| 2021 | Immediate Proximity Detection Using Wi-Fi-Enabled SmartphonesabstractSmartphone apps for exposure notification and contact tracing have been shown to be effective in controlling the COVID-19 pandemic. However, Bluetooth Low Energy tokens similar to those broadcast by existing apps can still be picked up far away from the transmitting device. In this paper, we present a new class of methods for detecting whether or not two Wi-Fi–enabled devices are in immediate physical proximity, i.e. 2 or fewer meters apart, as established by the U.S. Centers for Disease Control and Prevention (CDC). Our goal is to enhance the accuracy of smartphone-based exposure notification and contact tracing systems. We present a set of binary machine learning classifiers that take as input pairs of Wi-Fi RSSI fingerprints. We empirically verify that a single classifier cannot generalize well to a range of different environments with vastly different numbers of detectable Wi-Fi Access Points (APs). However, specialized classifiers, tailored to situations where the number of detectable APs falls within a certain range, are able to detect immediate physical proximity significantly more accurately. As such, we design three classifiers for situations with low, medium, and high numbers of detectable APs. These classifiers distinguish between pairs of RSSI fingerprints recorded 2 or fewer meters apart and pairs recorded further apart but still in Bluetooth range. We characterize their balanced accuracy for this task to be between 66.8% and 77.8%. Zach Van Hyfte, Avideh Zakhor |
IPIN | 2 |
| 2020 | Temporal LiDAR Frame Prediction for Autonomous DrivingabstractAnticipating the future in a dynamic scene is critical for many fields such as autonomous driving and robotics. In this paper we propose a class of novel neural network architectures to predict future LiDAR frames given previous ones. Since the ground truth in this application is simply the next frame in the sequence, we can train our models in an self-supervised fashion. Our proposed architectures are based on FlowNet3D and Dynamic Graph CNN. We use Chamfer Distance (CD) and Earth Mover's Distance (EMD) as loss functions and evaluation metrics. We train and evaluate our models using the newly released nuScenes dataset, and characterize their performance and complexity with several baselines. Compared to directly using FlowNet3D, our proposed architectures achieve CD and EMD nearly an order of magnitude lower. In addition, we show that our predictions generate reasonable scene flow approximations without using any labelled supervision. David Deng, Avideh Zakhor |
3DV | 2 |
| 2020 | Indoor Query System for the Visually Impaired
Lizhi Yang, Ilian Herzi, Avideh Zakhor, Anup Hiremath, Sahm Bazargan, Robert Tames-Gadam |
ICCHP (1) | 3 |
| 2020 | 3d Object Detection For Autonomous Driving Using Temporal Lidar Dataabstract3D object detection is a fundamental problem in the space of autonomous driving, and pedestrians are some of the most important objects to detect. The recently introduced PointPillars architecture has been shown to be effective in object detection. It voxelizes 3D LiDAR point clouds to produce a 2D pseudo-image to be used for object detection. In this work, we modify PointPillars to become a recurrent network, using fewer LiDAR frames per forward pass. Specifically, as compared to the original PointPillars model which uses 10 LiDAR frames per forward pass, our recurrent model uses 3 frames and recurrent memory. With this modification, we observe an 8% increase in pedestrian detection and a slight decline in performance on vehicle detection in a coarsely voxelized setting. Furthermore, when given 3 frames of data as input to both models, our recurrent architecture outperforms PointPillars by 21% and 1% in pedestrian and vehicle detection, respectively. Scott McCrae, Avideh Zakhor |
ICIP | 2 |
| 2020 | Few Shot Learning For Point Cloud Data Using Model Agnostic Meta LearningabstractThe ability of deep neural networks to extract complex statistics and learn high level features from vast datasets is proven. Yet current deep learning approaches suffer from poor sample efficiency in stark contrast to human perception. Few shot learning algorithms such as matching networks or Model Agnostic Meta Learning (MAML) mitigate this problem, enabling fast learning with few examples. In this paper, we extend the MAML algorithm to point cloud data using a PointNet Architecture. We construct N × K-shot classification tasks from the ModelNet40 point cloud dataset to show that this method performs classification as well as supervised deep learning methods with the added benefit of being able to adapt after a single gradient step on a single N × K task. We empirically search for optimal values of N and K for few shot classification and show our method to achieve 90% meta test accuracy compared to traditional PointNet with 89.2%. We also adapt a meta-trained PointNet to a support set of 9, N = 3, K = 3, never before seen point clouds which are drawn from an entirely different dataset, ShapeNet. Once adapted the model achieves 7.1/9 classification accuracy on average across 100 query sets of the same classes with new, unique instances. This result far exceeds the supervised Stochastic Gradient Descent (SGD) training result of 3.1/9 accuracy on the query sets which is equivalent to a random baseline. Rishi Puri, Avideh Zakhor, Raul Puri |
ICIP | 2 |
| 2020 | Point Cloud Segmentation using RGB Drone ImageryabstractIn recent years, the ubiquity of drones equipped with RGB cameras has made aerial 3D model generation significantly more cost effective than traditional aerial LiDAR-based methods. Most existing aerial 3D point cloud segmentation approaches use geometric methods and are tailored to 3D LiDAR data. In this paper, we propose a pipeline for semantic segmentation of 3D point clouds obtained via photogrammetry from aerial RGB camera images. Our basic approach is to directly apply deep learning segmentation methods to the very RGB images used to create the point cloud itself, followed by back-projecting the pixel class in segmented images onto the 3D points. This is a particularly attractive solution, since deep learning methods for image segmentation are more mature and advanced as compared to 3D point cloud segmentation. Furthermore, GPU engines for 2D image convolutions are likely to result in higher processing speeds than could be achieved using 3D point cloud data. We demonstrate our segmentation approach on two RGB Drone image datasets captured in Alameda, California, and compare its performance with manually labelled ground truth data. We use F1 and Jaccard similarity coefficient scores to show that our methodology outperforms existing methods such as PointNet++ and commercially available packages such as Pix4D. Marc WuDunn, James Dunn 0003, Avideh Zakhor |
ICIP | 3 |
| 2019 | Duodepth: Static Gesture Recognition Via Dual Depth SensorsabstractStatic gesture recognition is an effective non-verbal communication channel between a user and their devices; however many modern methods are sensitive to the relative pose of the user's hands with respect to the capture device, as parts of the gesture can become occluded. We present two methodologies for gesture recognition via synchronized recording from two depth cameras to alleviate this occlusion problem. One is a more classic approach using iterative closest point registration to accurately fuse point clouds and a single PointNet architecture for classification, and the other is a dual Point-Net architecture for classification without registration. On a manually collected data-set of 20,100 point clouds we show a 39.2% reduction in misclassification for the fused point cloud method, and 53.4% for the dual PointNet, when compared to a standard single camera pipeline. Ilya Chugunov, Avideh Zakhor |
ICIP | 2 |
| 2017 | AtomMap: A probabilistic amorphous 3D map representation for robotics and surface reconstructionabstractWe present a new 3D probabilistic occupancy map representation for robotics applications by relaxing the commonly-assumed constraint that space must be perfectly tessellated. We replace the regular structure of 3D grids with an unstructured collection of non-overlapping, equally-sized spheres, which we call “atoms”. Abandoning the grid structure allows a more accurate representation of space directly tangent to surfaces, which facilitates a number of applications such as high fidelity surface reconstruction and surface-guided path planning. Maps composed of atoms can distinguish between free, occupied, and unknown space, support computationally efficient insertions and collision queries, provide free space planning guarantees, and achieve state-of-the-art memory efficiency over large volumes. This is achieved while simultaneously reducing quantization effects in the vicinity of surfaces and defining a useful implicit surface representation. David Fridovich-Keil, Erik Nelson, Avideh Zakhor |
ICRA | 3 |
| 2017 | Cooperative inchworm localization with a low cost teamabstractIn this paper we address the problem of multi-robot localization with a heterogeneous team of low-cost mobile robots. The team consists of a single centralized observer with an inertial measurement unit (IMU) and monocular camera, and multiple picket robots with only IMUs and Red Green Blue (RGB) light emitting diodes (LED). This team cooperatively navigates a visually featureless environment while localizing all robots. A combination of camera imagery captured by the observer and IMU measurements from the pickets and observer are fused to estimate motion of the team. A team movement strategy, referred to as inchworm, is formulated as follows: Pickets move ahead of the observer and then act as temporary landmarks for the observer to follow. This cooperative approach employs a single Extended Kalman Filter (EKF) to localize the entire heterogeneous multi-robot team, using a formulation of the measurement Jacobian to relate the pose of the observer to the poses of the pickets with respect to the global reference frame. An initial experiment with the inchworm strategy has shown localization within 0.14 m position error and 2.18° orientation error over a path-length of 5 meters in an environment with irregular ground, partial occlusions, and a ramp. This demonstrates improvement over a camera-only localization technique that was adapted to our team dynamic which produced 0.18m position error and 3.12° orientation error over the same dataset. In addition, we demonstrate improvement in localization accuracy with an increasing number of picket robots. Brian E. Nemsick, Austin Buchan, Anusha Nagabandi, Ronald S. Fearing, Avideh Zakhor |
ICRA | 5 |
| 2015 | Automatic Indoor 3D Surface Reconstruction with Segmented Building and Object ElementsabstractAutomatic generation of 3D indoor building models is important for applications in augmented and virtual reality, indoor navigation, and building simulation software. This paper presents a method to generate high-detail watertight models from laser range data taken by an ambulatory scanning device. Our approach can be used to segment the permanent structure of the building from the objects within the building. We use distinct techniques to mesh the building structure and the objects to efficiently represent large planar surfaces, such as walls and floors, while still preserving the fine detail of segmented objects, such as furniture or light fixtures. Our approach is scalable enough to be applied on large models composed of several dozen rooms, spanning over 14,000 square feet. We experimentally verify this method on several datasets from diverse building environments. Avideh Zakhor |
3DV | 2 |
| 2015 | Access Point Selection for Multi-Rate IEEE 802.11 Wireless LANsabstractAccess Point (AP) selection is important in WLANs as it affects the throughput of the joining station (STA). In this paper, we propose a class of AP selection algorithms to maximize the joining STA's expected throughput by considering interference at STAs, and transmit opportunities (TXOPs) at APs. Specifically, we collect a binary-valued local channel occupancy signal, called busy-idle (BI) signal, at each node and require the APs to periodically broadcast their BI signal and a quantity representing their TXOPs. This enables the joining STA to estimate throughput from candidate APs before selecting one. We use NS-2 simulations to demonstrate the effectiveness of our algorithms for saturated UDP and TCP downlink traffic, and compare them with received signal power (rxpwr) algorithm, load based algorithm, and Fukuda algorithm. For a random topology consisting of 24 APs and 60 STAs, our algorithms increase the joining STAs average throughput by as much as 42% and 24% compared to rxpwr for UDP and TCP respectively. In addition, the achieved average throughput is 90% and 93% of that obtained via the optimal selection. We also show that, in contrast to rxpwr, the throughput of proposed algorithms remains close to optimal with the increase in AP or STA density. Shicong Yang, Michael N. Krishnan, Avideh Zakhor |
GLOBECOM | 3 |
| 2015 | Sensor fusion for semantic segmentation of urban scenesabstractSemantic understanding of environments is an important problem in robotics in general and intelligent autonomous systems in particular. In this paper, we propose a semantic segmentation algorithm which effectively fuses information from images and 3D point clouds. The proposed method incorporates information from multiple scales in an intuitive and effective manner. A late-fusion architecture is proposed to maximally leverage the training data in each modality. Finally, a pairwise Conditional Random Field (CRF) is used as a post-processing step to enforce spatial consistency in the structured prediction. The proposed algorithm is evaluated on the publicly available KITTI dataset [1] [2], augmented with additional pixel and point-wise semantic labels for building, sky, road, vegetation, sidewalk, car, pedestrian, cyclist, sign/pole, and fence regions. A per-pixel accuracy of 89.3% and average class accuracy of 65.4% is achieved, well above current state-of-the-art [3]. Richard Zhang 0001, Stefan A. Candra, Kai Vetter, Avideh Zakhor |
ICRA | 4 |
| 2014 | Contention window adaptation using the busy-idle signal in 802.11 WLANsabstractIn a wireless local area network (LAN), packets can be lost for a variety of reasons, including collisions due to high traffic and channel errors due to poor channel conditions. In practice, however, nodes cannot easily differentiate between these types of loss. As a result, adaptations based on packet loss alone can result in significantly degraded performance. In 802.11 networks, wireless nodes avoid collisions via the Binary Exponential Backoff (BEB) protocol. This performs well for moderate numbers of nodes and low channel error rates, but is inefficient for large numbers of nodes, high channel error rates, or in the presence of hidden terminals. In this paper, we propose a contention window adaptation scheme in which nodes use information shared by the AP to optimize contention window sizes in a distributed fashion to improve network utility. We show via NS-2 simulations that our method can improve throughput by as much as 24% in the high node count scenario, 35% in the high channel error scenario, and 350% in the presence of hidden terminals. Michael N. Krishnan, Shicong Yang, Avideh Zakhor |
GLOBECOM | 3 |
| 2014 | Interactive shadow analysis for camera heading in outdoor imagesabstractImage geo-localization is an important problem with many applications such as augmented reality and navigation. The most common ways to geo-localize an image are to use its meta-data such as GPS or to match it against a geotagged database. When neither of those is available, it is still possible to apply shadow analysis to determine the camera heading for outdoor images. This could be useful pruning the search space in geo-localization applications, for example by removing roads with incompatible orientations from a database such as Open Street Map. In this paper, we develop a novel interactive method for deducing the global heading of a query image using the shadows in it. We start by constructing a model of the sun-earth system to determine all shadows possible at a given approximate latitude, and compare shadows within the query to those possible under the model to determine the range of possible headings. We demonstrate this on 54 query images with known ground truth, and show that in 52 cases the ground truth lies in the computed range. Matthew Clements, Avideh Zakhor |
ICIP | 2 |
| 2014 | Simultaneous fingerprinting and mapping for multimodal image and WiFi indoor positioningabstractIn this paper, we propose an end-to-end system which can be used to simultaneously generate (a) 3D models and associated 2D floor plans and (b) multiple sensor e.g. WiFi and imagery signature databases for the large scale indoor environments in a fast, automated, scalable way. We demonstrate ways of recovering the position of a user carrying a mobile device equipped with a camera and WiFi sensor in an indoor environment. The acquisition system consists of a man portable backpack of sensors carried by an operator inside buildings walking at normal speeds. The sensor suite consists of laser scanners, cameras and an IMU. Particle filtering algorithms are used to recover 2D and 3D path of the operator, a 3D point cloud, the 2D floor plan, and 3D models of the environment. The same walkthrough that produces 2D maps also generates multi-modal sensor databases, in our case WiFi and imagery. The resulting WiFi database is generated much more rapidly than existing systems due to continuous, rather than stop-and-go or crowd-sourced WiFi signature acquisition. We also use particle filtering algorithms in an Android application to combine inertial sensors on the mobile device, with 2D maps and WiFi and image sensor databases to localize the user. Experimental for the second floor of the electrical engineering building at UC Berkeley campus show that our system achieves an average localization error of under 2m. Plamen Levchev, Michael N. Krishnan, Chaoran Yu, Joseph Menke, Avideh Zakhor |
IPIN | 5 |
| 2014 | Large Area Cell Based Image LocalizationabstractWe present a memory scalable image localization system that uses distributed kd-trees created on overlapping geographic cells using a database of 10 million Google Street View images for an area of approximately 10,000 square kilometers in Taiwan. Given a collection of images over a region of interest (ROI), we generate a database by dynamically creating geographic cells that are optimized so that each cell contains roughly the same number of images. We then create kd-trees for each cell from SIFT features extracted from the images in that cell. When querying the system, we run traditional feature matching on each cell and pool the results for each cell to rerank with a geometric constraint. The key idea is the subdivisions of the ROI into overlapping geographic cells, allowing our system to scale to 10 million images and to efficiently utilize prior query location information when available. We evaluate our system on a test set of 29 geo-tagged images, not from Google Street View, taken throughout Taiwan with various resolutions, aspect ratios, and qualities. We also evaluate our system on a set of 97 images without geo-tag data. Andrew Zhai, Matthew Clements, Avideh Zakhor |
ISM | 3 |
| 2014 | Automatic identification of window regions on indoor point clouds using LiDAR and camerasabstractIn this paper, we propose an algorithm to automatically identify window regions on exterior facing facades of buildings using interior 3D point cloud resulting from an ambulatory backpack sensor system, outfitted with multiple LiDAR sensors and cameras. We develop a set of discriminative features for the task, namely visual brightness, infrared opaqueness, and an occlusion indicator, within a Markov Random Field (MRF) framework to provide structured prediction for window or glass regions. A preprocessing classifier is trained on the features to produce node potentials, and large margin parameter training is used to boost performance. Our algorithm has been trained on data taken at the 3rdfloor of Cory Hall at UC Berkeley, with a total façade area of 269.1 m2, and has been tested on walls taken on the 2ndfloor of Cory Hall, a Walgreens, and an office building in San Francisco, with a total exterior façade area of 454.6 m2. Window regions are successfully identified with 85.5% F1-score and 94.2% accuracy. Richard Zhang 0001, Avideh Zakhor |
WACV | 2 |
| 2013 | Watertight Planar Surface Meshing of Indoor Point-Clouds with Voxel Carvingabstract3D modeling of building architecture from point-cloud scans is a rapidly advancing field. These models are used in augmented reality, navigation, and energy simulation applications. State-of-the-art scanning produces accurate point-clouds of building interiors containing hundreds of millions of points. Current surface reconstruction techniques either do not preserve sharp features common in a man-made structures, do not guarantee water tightness, or are not constructed in a scalable manner. This paper presents an approach that generates watertight triangulated surfaces from input point-clouds, preserving the sharp features common in buildings. The input point-cloud is converted into a voxelized representation, utilizing a memory-efficient data structure. The triangulation is produced by analyzing planar regions within the model. These regions are represented with an efficient number of elements, while still preserving triangle quality. This approach can be applied to data of arbitrary size to result in detailed models. We apply this technique to several data sets of building interiors and analyze the accuracy of the resulting surfaces with respect to the input point-clouds. Avideh Zakhor |
3DV | 2 |
| 2013 | Reduced-complexity data acquisition system for image-based localization in indoor environmentsabstractImage-based localization has important commercial applications such as augmented reality and customer analytics. In prior work, we developed a three step pipeline for image-based localization of mobile devices in indoor environments. In the first step, we generate a 2.5D georeferenced image database using an ambulatory backpack-mounted system originally developed for 3D modeling of indoor environments. Specifically, we first create a dense 3D point cloud and polygonal model from the side laser scanner measurements of the backpack, and then use it to generate dense 2.5D database image depthmaps by raytracing the 3D model. In the second step, a query image is matched against the image database to retrieve the best-matching database image. In the final step, the pose of the query image is recovered with respect to the best-matching image. Since the pose recovery in step three only requires sparse depth information at certain SIFT feature keypoints in the database image, in this paper we improve upon our previous method by only calculating depth values at these keypoints, thereby reducing the required number of sensors in our data acquisition system. To do so, we use a modified version of the classic multi-camera 3D scene reconstruction algorithm, thereby eliminating the need for expensive geometry laser range scanners. Our experimental results in a shopping mall indicate that the proposed reduced complexity sparse depthmap approach is nearly as accurate as our previous dense depth map method. Jason Zhi Liang, Nicholas Corso, Avideh Zakhor |
IPIN | 4 |
| 2013 | Geometric calibration for a multi-camera-projector systemabstractIn this paper, we describe a calibration method for multi-camera-projector systems in which sensors face each other as well as share a common viewpoint. We use a translucent planar sheet framed in PVC piping as a calibration target which is placed at multiple positions and orientations within a scene. In each position, the target is captured by the cameras while it is being illuminated by a set of projected patterns from various projectors. The translucent sheet allows the projected patterns to be visible from both sides, allowing correspondences between devices that face each other. The set of correspondences generated between the devices using this target are input into a bundle adjustment framework to estimate calibration parameters. We demonstrate the effectiveness of this approach on a multiview structured light system made of three projectors and nine cameras. Ricardo R. Garcia, Avideh Zakhor |
WACV | 2 |
| 2013 | Single view pose estimation of mobile devices in urban environmentsabstractPose estimation of mobile devices is useful for a wide variety of applications, including augmented reality and geo-tagging. Even though most of today's cell phones are equipped with sensors such as GPS, accelerometers, and gyros, the pose estimated via these is often inaccurate, particularly in urban environments. In this paper, we describe an image based localization algorithm for estimating the pose of cell phones in urban environments. Our proposed approach solves for a homography transformation matrix between the cell phone image and a matching database image, constrained by knowledge of the change in orientation obtained from the cell phone gyro, and augmented with 3D information from the database to achieve an estimate of pose which improves upon readings from the GPS and compass. We characterize the performance of this approach for a dataset in Oakland, CA and show that for a query set of 92 images, our computed location (yaw) is within 10 meters (degrees) for 92% (96%) of queries as compared to 31% (26%) for the GPS (compass) on the cell phone. Aaron Hallquist, Avideh Zakhor |
WACV | 2 |
| 2012 | Planar 3D modeling of building interiors from point cloud dataabstractWe present an automatic system for planar 3D modeling of building interiors from point cloud data generated by range scanners. This is motivated by the observation that most building interiors may be modeled as a collection of planes representing ceilings, floors, walls and staircases. Our proposed system, which employs model-fitting and RANSAC, is capable of detecting large-scale architectural structures, such as ceilings and floors, as well as small-scale architectural structures, such as staircases. We experimentally validate our system on a number of challenging point clouds of real architectural scenes. Victor Sanchez, Avideh Zakhor |
ICIP | 2 |
| 2012 | Sharp geometry reconstruction of building facades using range dataabstractIn this paper we describe a method for detailed geometry reconstruction of building façades in an urban environment, given a 3D point-cloud of LiDAR range data. Our approach separates planar faces and interpolates their shape with Moving Least-Squares (MLS) sampling. A method is then proposed to reconstruct occluded areas of the building whereby gaps in the building surface are modeled with axis-aligned planes fit to the gap boundary vertices. This approach reconstructs unsampled areas of building surfaces under the assumption that buildings have 3D rectilinear, axis-aligned features. We demonstrate the effectiveness of our approach on a number of building façades. Avideh Zakhor |
ICIP | 2 |
| 2011 | A Method for Estimating Access Delay Distribution in IEEE 802.11 NetworksabstractThe volume of multimedia traffic over wireless networks has been steadily increasing over the past decade. Unlike web browsing applications, multimedia data needs to satisfy stringent delay requirements since late packets are as good as lost packets. In this paper, we present a framework for the nodes in 802.11 networks to estimate the distribution of uplink access delay in Distributed Coordination Function (DCF) MAC mechanism using locally available information. The access delay for a packet is defined as the time between the packet arriving at the head of line of MAC queue, and its ACK being received. In our proposed framework, each node periodically records channel occupancy information to estimate the distribution of access delay. We use NS-2 simulations to verify the accuracy of our proposed approach. Ehsan Haghani, Michael N. Krishnan, Avideh Zakhor |
GLOBECOM | 3 |
| 2011 | Packet Length Adaptation in WLANs with Hidden Nodes and Time-Varying ChannelsabstractIn a wireless local area network (LAN), packets can be lost due to a multitude of reasons. It is possible to reduce the probability of occurrence of some of these loss mechanisms by reducing packet length at the medium access control (MAC) layer. However, there is an inherent tradeoff in that shorter packets decrease efficiency with respect to overhead. In current packet length adaptation literature, simplified or incomplete packet loss models are used, neglecting channel fading or collisions due to hidden nodes. In this paper, we apply a more complete packet loss model and propose a local packet length adaptation algorithm whereby each node dynamically adjusts its packet length based on estimates of the probabilities of each significant type of packet loss. In our technique, the access point periodically broadcasts channel occupancy information which each node uses in conjunction with its own local observations in order to estimate current network conditions. These are used to estimate the derivative of throughput with respect to packet length at each node under the current network conditions and to adapt the packet lengths accordingly. We demonstrate throughput gains of up to 20% via NS-2 simulations. Michael N. Krishnan, Ehsan Haghani, Avideh Zakhor |
GLOBECOM | 3 |
| 2011 | Surface completion of shape and texture based on energy minimizationabstractIn this paper, we propose a novel surface completion method to generate plausible shapes and textures for missing regions of 3D models. The missing regions are filled in by minimizing two energy functions for shape and texture, which are both based on similarities between the missing region and the rest of the object; in doing so, we take into account the positive correlation between shape and texture. We demonstrate the effectiveness of the proposed method experimentally by applying it to two models. Norihiko Kawai, Avideh Zakhor, Tomokazu Sato, Naokazu Yokoya |
ICIP | 2 |
| 2011 | Fast approximation for geometric classification of LiDAR returnsabstractCurrent LiDAR classification methods are excessively slow to be used in real-time navigation systems, even though they are useful for human perception. These methods typically analyze curvature by applying Principal Component Analysis (PCA) to each point in a point cloud. For variable-density aerial LiDAR obtained by at a shallow angle with respect to the ground rather than in a top-down fashion, the variations in density pose special challenges in terms of choosing the appropriate PCA parameters. In this paper we use gridded approximate nearest neighbor searches for fast classification of geometric features in large LiDAR point clouds. The underlying algorithm exploits spatial hashes and the forgiving nature of PCA as a part of geometric classification. We show a factor of 10-20 speed up for both actual and simulated point clouds with little or no loss in classification performance. Our approach is applicable to both uniform and variable-density aerial LiDAR datasets. Xiaozhe Shi, Avideh Zakhor |
ICIP | 2 |
| 2011 | Location-based image retrieval for urban environmentsabstractImage based localization is an important problem with many applications. The basic idea is to match a user generated query image against a database of geo-tagged images with known 6 degrees of freedom poses. Once this retrieval problem is solved, it is possible to recover the pose of the query image. A challenging problem in image retrieval is performance degradation as the size of the image database grows. In this paper we describe an approach to large scale image retrieval for user localization in urban environment by taking advantage of coarse position estimates available, e.g. via cell tower triangulation, on many mobile devices today. The basic idea is to partition the large image database for a large region into a number of overlapping cells each with its own prebuilt search and retrieval structure. We demonstrate retrieval results over a ~12,000 image database covering a 1 km2area of downtown Berkeley. Jerry Zhang, Aaron Hallquist, Eric Liang, Avideh Zakhor |
ICIP | 4 |
| 2011 | Local estimation of collision probabilities in 802.11 WLANs: An experimental studyabstractCurrent 802.11 networks do not typically achieve the maximum potential throughput despite link adaptation and cross-layer optimization techniques designed to alleviate many causes of packet loss. A primary contributing factor is the difficulty in distinguishing between various causes of packet loss, including collisions caused by high network use, co-channel interference from neighboring networks, and errors due to poor channel conditions. In previous work, we used NS-2 simulations to show that estimating various components of loss probability such as direct collisions, staggered collisions, and physical layer errors, can be used to improve the throughput of 802.11 networks via link adaptation, carrier sense threshold adaptation, and MAC layer packet length adaptation. We have also proposed a method to estimate the various components of loss probability by comparing channel occupancy at a station with that of its access point. In this paper, we use Ath5k open source wireless card driver in an experimental testbed in order to experimentally verify the accuracy of our previously proposed approach to estimating collision probability. We show that our proposed methodology accurately estimates overall collision probability to within 5%. This experimental verification demonstrates the feasibility of our collision probability estimation approach and the resulting throughput gains in practice. Miklos Christine, Michael N. Krishnan, Ehsan Haghani, Avideh Zakhor |
WCNC | 4 |
| 2010 | Adaptive Carrier-Sensing for Throughput Improvement in IEEE 802.11 NetworksabstractAs a Carrier Sense Multiple Access (CSMA) network, the performance of IEEE 802.11 networks highly depends on the accuracy of the carrier sensing procedure. However, conventional carrier sensing approaches suffer from the well known hidden and exposed node problems, adversely affecting aggregate throughput of the IEEE 802.11 networks. In this paper, we propose a novel scheme through which each station can adaptively select its Carrier Sense Threshold (CST) in order to mitigate the hidden/exposed node problems. The basic idea behind our approach is for the Access Point (AP) to periodically transmit a Busy/Idle (BI) signal to all the stations. Individual stations then use the BI signal from the AP together with their own local BI signal in order to adjust their CST. We use NS-2 simulations to show that our approach can enhance the aggregate throughput by as much as 50%. Ehsan Haghani, Michael N. Krishnan, Avideh Zakhor |
GLOBECOM | 3 |
| 2010 | Indoor localization and visualization using a human-operated backpack systemabstractAutomated 3D modeling of building interiors is useful in applications such as virtual reality and entertainment. Using a human-operated backpack system equipped with 2D laser scanners and inertial measurement units (IMU), we develop scan matching based algorithms to localize the backpack in complex indoor environments such as a T-shaped corridor intersection, a staircase, and two indoor hallways from two separate floors connected by a staircase. When building 3D textured models, we find that the localization resulting from scan matching is not pixel accurate, resulting in misalignment between successive images used for texturing. To address this, we propose an image based pose estimation algorithm to refine the results from our scan matching based localization. Finally, we use the localization results within an image based renderer to enable virtual walkthroughs of indoor environments using imagery from cameras on the same backpack. Our renderer uses a three-step process to determine which image to display, and a RANSAC framework to determine homographies to mosaic neighboring images with common SIFT features. In addition, our renderer uses plane-fitted models of the 3D point cloud resulting from the laser scans to detect occlusions. We characterize the performance of our image based renderer on an unstructured set of 2709 images obtained during a five minute backpack data acquisition for a T-shaped corridor intersection. Timothy Liu, Matthew Carlberg, Jacky Chen, John Kua, Avideh Zakhor |
IPIN | 6 |
| 2010 | Throughput Improvement in 802.11 WLANs Using Collision Probability Estimates in Link AdaptationabstractThe 802.11 standard includes several modulation rates, each of which is optimal for a different channel condition. However, there are no simple and reliable methods for nodes to determine their current channel conditions. Existing link adaptation techniques use packet losses as an indication of poor channel conditions; however, when there is a significant probability of collision, this assumption fails, leading to degraded throughput. In this paper, we show that an estimate of the probability of collision can be used to improve link adaptation in 802.11 networks with hidden terminals, and significantly increase throughput by up to a factor of five. We demonstrate this through NS-2 simulations of a few link adaptation techniques including a new algorithm, called SNRg. Michael N. Krishnan, Avideh Zakhor |
WCNC | 2 |
| 2009 | Local Estimation of Probabilities of Direct and Staggered Collisions in 802.11 WLANsabstractCurrent 802.11 networks do not typically achieve the maximum potential throughput despite link adaptation and cross-layer optimization techniques designed to alleviate many causes of packet loss. A primary contributing factor is the difficulty in distinguishing between various causes of packet loss, including collisions caused by high network use, co-channel interference from neighboring networks, and errors due to poor channel conditions. In this paper, we propose a novel method for estimating various collision type probabilities locally at a given node of an 802.11 network. Our approach is based on combining locally observable quantities with information observed and broadcast by the access point (AP) in order to obtain partial spatial information about the network traffic. We provide a systematic assessment and definition of the different types of collision, and show how to approximate each of them using only local and AP information. Additionally, we show how to approximate the sensitivity of these probabilities to key related configuration parameters including carrier sense threshold and packet length. We verify our methods through NS-2 simulations, and characterize estimation accuracy of each of the considered collision types. Michael N. Krishnan, Sofie Pollin, Avideh Zakhor |
GLOBECOM | 3 |
| 2009 | Classifying urban landscape in aerial LiDAR using 3D shape analysisabstractThe classification of urban landscape in aerial lidar point clouds is useful in 3D modeling and object recognition applications in urban environments. In this paper, we introduce a multi-category classification system for identifying water, ground, roof, and trees in airborne lidar. The system is organized as a cascade of binary classifiers, each of which performs unsupervised region growing followed by supervised, segment-wise classification. Categories with the most discriminating features, such as water and ground, are identified first and are used as context for identifying more complex categories, such as trees. We use 3D shape analysis and region growing to identify ¿planar¿ and ¿scatter¿ regions that likely correspond to ground/roof and trees respectively. We demonstrate results on two urban datasets, the larger of which contains 200 million lidar returns over 7km2. We show that our ground, roof, and tree classifiers, when trained on one dataset, perform well on the other dataset. Matthew Carlberg, Peiran Gao, Avideh Zakhor |
ICIP | 4 |
| 2009 | 2D tree detection in large urban landscapes using aerial LiDAR dataabstractWe present a scalable approach to tree detection in large urban landscapes using aerial LiDAR data. Similar to our previous work in 2006, our current method consists of segmentation followed by classification. However, unlike our previous work, the current approach does not use color information or aerial imagery, and hence is more generally applicable. Also, our current approach has been successfully tested on two very large datasets, which are many orders of magnitude larger than the dataset used in 2006. Specifically, we use a North American dataset, containing 125 million LiDAR returns over 3 km2, and a European dataset, containing 200 million LiDAR returns over 7 km2. For both datasets, we report precision and recall rates of over 95%. Avideh Zakhor |
ICIP | 2 |
| 2009 | Image augmented laser scan matching for indoor dead reckoningabstractMost existing approaches to indoor localization focus on using either cameras or laser scanners as the primary sensor for pose estimation. In scan matching based localization, finding scan point correspondences across scans is challenging as individual scan points lack unique attributes. In camera based localization, one has to deal with images with few or no visual features as well as scale factor ambiguities to recover absolute distances. In this paper, we develop multimodal approaches for two indoor localization problems by fusing a camera and laser scanners in order to alleviate the drawbacks of each individual modality. For our first problem we recover 3 degrees of freedom (DoF) of a camera-laser rig on a rolling cart in a 2D plane, by using visual odometry to facilitate scan correspondence estimation. We demonstrate this approach to result in a 0.3% loop closure error for a 60 m loop around the interior corridor of a building. In our second problem, we recover 6 DoF of a human operator carrying a backpack system mounted with sensors in 3D, by merging rotation estimates from scan matching and translation estimates from visual odometry, resulting in a 1% loop closure error. Nikhil Santosh Naikal, John Kua, Avideh Zakhor |
IROS | 4 |
| 2009 | Adaptive packetization for error-prone transmission over 802.11 WLANs with hidden terminalsabstractCollision and fading are the two main sources of packet loss in wireless local area networks (WLANs) and as such, both are affected by the packetization at the medium access control (MAC) layer.While a larger packet is preferred to balance protocol header overhead, a shorter packet is less vulnerable to packet loss due to channel fading errors or staggered collisions in the presence of hidden terminals. Direct collisions due to backoff are not affected by packet size. Recently, Krishnan et. al. have developed a new technique for estimating probabilities of various components of packet loss, namely, direct and staggered collisions and fading. Motivated by this work, in this paper, we exploit ways in which packetization can be used to improve throughput performance of WLANs. We first show analytically that the effective throughput is a unimodal function of the packet size when considering both channel fading and staggered collisions. We then develop a measurement-based algorithm based on golden section search to arrive at an optimal packet size for MAC-layer transmissions. Our simulations demonstrate that packetization based on our search algorithm can greatly improve the effective throughput of sensing-limited nodes, and reduce video frame transfer delay in WLANs. Michael N. Krishnan, Avideh Zakhor |
MMSP | 3 |
| 2009 | Interference Aware Multipath Selection for Video Streaming in Wireless Ad Hoc NetworksabstractIn this paper, we propose a novel multipath selection framework for video streaming over wireless ad hoc networks. We propose a heuristic interference-aware multipath routing protocol based on the estimation of concurrent packet drop probability of two paths, taking into account interference between links. Through both simulations and actual experiments, we show that the performance of the proposed protocol is close to that of the optimal solution, and is better than that of other heuristic protocols. Wei Wei 0023, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Automatic registration of aerial imagery with untextured 3D LiDAR modelsabstractA fast 3D model reconstruction methodology is desirable in many applications such as urban planning, training, and simulations. In this paper, we develop an automated algorithm for texture mapping oblique aerial images onto a 3D model generated from airborne light detection and ranging (LiDAR) data. Our proposed system consists of two steps. In the first step, we combine vanishing points and global positioning system aided inertial system readings to roughly estimate the extrinsic parameters of a calibrated camera. In the second step, we refine the coarse estimate of the first step by applying a series of processing steps. Specifically, We extract 2D corners corresponding to orthogonal 3D structural corners as features from both images and the untextured 3D LiDAR model. The correspondence between an image and the 3D model is then performed using Hough transform and generalized M-estimator sample consensus. The resulting 2D corner matches are used in Lowepsilas algorithm to refine camera parameters obtained earlier. Our system achieves 91% correct pose recovery rate for 90 images over the downtown Berkeley area, and overall 61% accuracy rate for 358 images over the residential, downtown and campus portions of the city of Berkeley. Min Ding 0005, Kristian Lyngbaek, Avideh Zakhor |
CVPR | 3 |
| 2007 | Lossless Compression Algorithms for Post-OPC IC LayoutabstractAn important step in today's integrated circuit (IC) manufacturing is optical proximity correction (OPC). While OPC increases the fidelity of pattern transfer to the wafer, it also results in significant increase in IC layout file size. In this paper, we develop two techniques for compressing post-OPC layout data while remaining compliant with existing industry standard data formats such as OASIS and GDSII. The motivation for doing so is for the resulting compressed files to be viewed and edited by any industry standard CAD tools without a decoder. Our approach is to eliminate redundancies in the representation of the geometric data by finding repeating groups of polygons between multiple cells as well as within a cell. We refer to the former as "inter-cell sub-cell detection" and the later as "intra-cell sub-cell detection". Both problems are NP hard, and as such, we propose two sets of greedy algorithms to solve them. We show the results of our proposed inter-cell and intra-cell algorithms on actual 90nm, 130nm, and 180nm IC layouts. Allan Gu, Avideh Zakhor |
ICIP (2) | 2 |
| 2007 | Tree Detection in Urban Regions Using Aerial Lidar and Image DataabstractIn this letter, we present an approach to detecting trees in registered aerial image and range data obtained via lidar. The motivation for this problem comes from automated 3-D city modeling, in which such data are used to generate the models. Representing the trees in these models is problematic because the data are usually too sparsely sampled in tree regions to create an accurate 3-D model of the trees. Furthermore, including the tree data points interferes with the polygonization step of the building roof top models. Therefore, it is advantageous to detect and remove points that represent trees in both lidar and aerial imagery. In this letter, we propose a two-step method for tree detection consisting of segmentation followed by classification. The segmentation is done using a simple region-growing algorithm using weighted features from aerial image and lidar, such as height, texture map, height variation, and normal vector estimates. The weights for the features are determined using a learning method on random walks. The classification is done using the weighted support vector machines, allowing us to control the misclassification rate. The overall problem is formulated as a binary detection problem, and the results presented as receiver operating characteristic curves are shown to validate our approach John Secord, Avideh Zakhor |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2007 | Multiple Tree Video Multicast Over Wireless Ad Hoc NetworksabstractIn this paper, we propose multiple tree construction schemes and routing protocols for video streaming over wireless ad hoc networks. The basic idea is to split the video into multiple parts and send each part over a different tree, which are constructed to be disjoint with each other so as to increase robustness to loss and other transmission degradations. Specifically, we propose two novel multiple tree multicast protocols. Our first scheme constructs two disjoint multicast trees in a serial, but distributed fashion, and is referred to as serial multiple disjoint tree multicast routing protocol. It achieves reasonable tree connectivity while maintaining disjointness of two trees. In order to reduce routing overhead and construction delay, we further propose parallel multiple nearly-disjoint multicast trees protocol, which is also shown to achieve reasonable tree connectivity. Simulations show that resulting video quality for either scheme is significantly higher than that of single tree multicast, with similar routing overhead and forwarding efficiency Wei Wei 0023, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Tree Detection in Aerial Lidar and Image DataabstractIn this paper, we present an approach to detecting trees in registered aerial image and range data obtained via LiDAR. The motivation for this problem comes from automated 3D city modeling, in which such data is used to generate the models. Representing the trees in these models is problematic because the data is usually too sparsely sampled in tree regions to create an accurate 3-D model of the trees. Furthermore, including the tree data points interferes with the polygonization step of the building roof top models. Therefore, it is advantageous to detect and remove points that represent trees in both LiDAR and aerial imagery. In this paper we propose a two-step method for tree detection consisting of segmentation followed by classification. The segmentation is done using a simple region-growing algorithm using weighted features from aerial image and LiDAR, such as height, texture map, height variation, and normal vector estimates. The weights for the features are determined using a learning method on random walks. The classification is done using weighted support vector machines (SVM), allowing us to control the mis-classification rate. The overall problem is formulated as a binary detection problem, and receiver operating characteristic curves are shown to validate our approach. John Secord, Avideh Zakhor |
ICIP | 2 |
| 2006 | Path Selection for Multi-Path Streaming in Wireless Ad Hoc NetworksabstractIn this paper, we propose a novel multi-path selection framework for streaming over wireless ad hoc networks. Our approach is to approximately estimate the concurrent packet drop probability of two paths by taking into account the interference between different links, and to select the best path pair based on that estimation. We prove the optimal path selection problem to be NP-hard, and propose a heuristic solution, whose performance is shown to be close to that of the optimal solution, while significantly outperforming other heuristic protocols. Wei Wei 0023, Avideh Zakhor |
ICIP | 2 |
| 2006 | Multiple Tree Video Multicast Over Wireless Ad Hoc NetworksabstractIn this paper, we propose a multiple tree multicast streaming scheme for video applications over wireless ad hoc networks. Specifically, we propose a multiple tree construction protocol, which builds two nearly disjoint trees simultaneously in a distributed way. Simulation shows that video quality of our proposed scheme to be superior to that of single tree multicast, even though they have similar control overhead and forwarding efficiency. Avideh Zakhor, Wei Wei 0023 |
ICIP | 1 |
| 2006 | Flow Control Over Wireless Network and Application Layer ImplementationabstractAbstract — Flow control, including congestion control for data transmission, and rate control for multimedia streaming, is an important issue in information transmission in both wireline and wireless networks. Widely accepted flow control methods in wireline networks are TCP [1] for data, and TCP Friendly Rate Control (TFRC) [2] for multimedia. Kelly [3] [4] has laid down theoretical framework for TCP in wireline networks demonstrating its optimality, fairness, and stability. However, TCP and TFRC both assume that packet loss in wireline networks is primarily due to congestion, and as such, are not applicable to wireless networks in which the bulk of packet loss is due to errors at the physical layer. In this paper we first show flow control in the wireless networks can be formulated as the same concave optimization problem Kelly defined in the wireline networks. TCP and TFRC in the wireless networks pursue the optimal solution using inaccurate feedback. All existing approaches to this TCP/TFRC over wireless problem correct the inaccurate feedback by casting modifications to existing protocols, such as TCP, or infrastructure elements such as routers, thereby making them hard to deploy in practice. In this paper, we formulate the problem as another concave optimization problem with a different utility function, and propose a new class of solutions. Our approach is end-to-end, and achieves reasonable performance by adjusting the number of connections of a user according to a properly selected control law. The control law is based on only one bit of information, which can be reliably measured at the application layer. We show that the control system has a unique stable equilibrium that solves the concave optimization problem, implying scalability and optimality of the solution. We apply our results to design a practical rate control scheme for data transmission over wireless networks, and characterize its performance using NS-2 simulations and actual experiments over Verizon Wireless 1xRTT data network. Analysis and simulation results also indicate our scheme is applicable to both wireline and wireless scenarios. I. Minghua Chen 0001, Avideh Zakhor |
INFOCOM | 2 |
| 2006 | Lossless compression of VLSI layout image dataabstractWe present a novel lossless compression algorithm called Context Copy Combinatorial Code (C4), which integrates the advantages of two very disparate compression techniques: context-based modeling and Lempel-Ziv (LZ) style copying. While the algorithm can be applied to many lossless compression applications, such as document image compression, our primary target application has been lossless compression of integrated circuit layout image data. These images contain a heterogeneous mix of data: dense repetitive data better suited to LZ-style coding, and less dense structured data, better suited to context-based encoding. As part of C4, we have developed a novel binary entropy coding technique called combinatorial coding which is simultaneously as efficient as arithmetic coding, and as fast as Huffman coding. Compression results show C4 outperforms JBIG, ZIP, BZIP2, and two-dimensional LZ, and achieves lossless compression ratios greater than 22 for binary layout image data, and greater than 14 for gray-pixel image data. Vito Dai, Avideh Zakhor |
IEEE Trans. Image Process. | 2 |
| 2006 | Multiple TFRC Connections Based Rate Control for Wireless NetworksabstractRate control is an important issue in video streaming applications for both wired and wireless networks. A widely accepted rate control method in wired networks is equation based rate control , in which the TCP friendly rate is determined as a function of packet loss rate, round trip time and packet size. This approach, also known as TCP friendly rate control (TFRC), assumes that packet loss in wired networks is primarily due to congestion, and as such is not applicable to wireless networks in which the bulk of packet loss is due to error at the physical layer. In this paper, we propose multiple TFRC connections as an end-to-end rate control solution for wireless video streaming. We show that this approach not only avoids modifications to the network infrastructure or network protocol, but also results in full utilization of the wireless channel. NS-2 simulations, actual experiments over 1$times$RTT CDMA wireless data network, and and video streaming simulations using traces from the actual experiments, are carried out to validate, and characterize the performance of our proposed approach. Minghua Chen 0001, Avideh Zakhor |
IEEE Trans. Multim. | 2 |
| 2005 | Data Processing Algorithms for Generating Textured 3D Building Facade Meshes from Laser Scans and Camera Images
Christian Früh, Avideh Zakhor |
Int. J. Comput. Vis. | 3 |
| 2005 | Fast similarity search and clustering of video sequences on the world-wide-webabstractWe define similar video content as video sequences with almost identical content but possibly compressed at different qualities, reformatted to different sizes and frame-rates, undergone minor editing in either spatial or temporal domain, or summarized into keyframe sequences. Building a search engine to identify such similar content in the World-Wide Web requires: 1) robust video similarity measurements; 2) fast similarity search techniques on large databases; and 3) intuitive organization of search results. In a previous paper, we proposed a randomized technique called the video signature (ViSig) method for video similarity measurement. In this paper, we focus on the remaining two issues by proposing a feature extraction scheme for fast similarity search, and a clustering algorithm for identification of similar clusters. Similar to many other content-based methods, the ViSig method uses high-dimensional feature vectors to represent video. To warrant a fast response time for similarity searches on high dimensional vectors, we propose a novel nonlinear feature extraction scheme on arbitrary metric spaces that combines the triangle inequality with the classical Principal Component Analysis (PCA). We show experimentally that the proposed technique outperforms PCA, Fastmap, Triangle-Inequality Pruning, and Haar wavelet on signature data. To further improve retrieval performance, and provide better organization of similarity search results, we introduce a new graph-theoretical clustering algorithm on large databases of signatures. This algorithm treats all signatures as an abstract threshold graph, where the distance threshold is determined based on local data statistics. Similar clusters are then identified as highly connected regions in the graph. By measuring the retrieval performance against a ground-truth set, we show that our proposed algorithm outperforms simple thresholding, single-link and complete-link hierarchical clustering techniques. Sen-Ching S. Cheung, Avideh Zakhor |
IEEE Trans. Multim. | 2 |
| 2005 | Effective bandwidth based scheduling for streaming mediaabstractWe propose a class of rate-distortion optimized packet scheduling algorithms for streaming media by generating a number of nested substreams, with more important streams embedding less important ones in a progressive manner. Our goal is to determine the optimum substream to send at any moment in time, using feedback information from the receiver and statistical characteristics of the video. To do so, we model the streaming system as a queueing system, compute the run-time decoding failure probability of a group of picture in each substream based on effective bandwidth approach, and determine the optimum substream to be sent at that moment in time. We evaluate our scheduling scheme with various video traffic models featuring short-range dependency (SRD), long-range dependency (LRD), and/or multifractal properties. From experiments with real video data, we show that our proposed scheduling scheme outperforms the conventional sequential sending scheme. Sang H. Kang, Avideh Zakhor |
IEEE Trans. Multim. | 2 |
| 2005 | Receiver-driven bandwidth sharing for TCP and its application to video streamingabstractApplications using Transmission Control Protocol (TCP), such as web-browsers, ftp, and various peer-to-peer (P2P) programs, dominate most of the Internet traffic today. In many cases, users have bandwidth-limited last mile connections to the Internet which act as network bottlenecks. Users generally run multiple concurrent networking applications that compete for the scarce bandwidth resource. Standard TCP shares bottleneck link capacity according to connection round-trip time (RTT), and consequently may result in a bandwidth partition which does not necessarily coincide with the user's desires. In this work, we present a receiver-based bandwidth sharing system (BWSS) for allocating the capacity of last-hop access links according to user preferences. Our system does not require modifications to the TCP protocol, network infrastructure or sending hosts, making it easy to deploy. By breaking fairness between flows on the access link, the BWSS can limit the throughput fluctuations of high-priority applications. We utilize the BWSS to perform efficient video streaming over TCP to receivers with bandwidth-limited last mile connections. We demonstrate the effectiveness of our proposed system through Internet experiments. Puneet Mehra, Christophe De Vleeschouwer, Avideh Zakhor |
IEEE Trans. Multim. | 3 |
| 2004 | Multipath Unicast and Multicast Video Communication over Wireless Ad Hoc NetworksabstractIn this paper, we address the problem of real-time video communication over wireless ad hoc networks. For the unicast case, we propose a robust, multipath source routing protocol for both interactive and video on-demand applications. Simulations show that our proposed scheme enhances the quality of video applications as compared to the existing protocols. For the multicast case, we propose multiple tree multicast streaming as a way to provide robustness for video multicast applications. Specifically, we propose a distributed double disjoint tree multicast routing protocol called serial MDTMR, and characterize its performance via simulations. We show that serial MDTMR achieves reasonable tree connectivity while maintaining disjointness of two trees, and that it outperforms single tree multicast communication. Wei Wei 0023, Avideh Zakhor |
BROADNETS | 2 |
| 2004 | Transmission protocols for streaming video over wireless
Minghua Chen 0001, Avideh Zakhor |
ICIP | 2 |
| 2004 | Robust multipath source routing protocol (RMPSR) for video communication over wireless ad hoc networksabstractMultipath routing is effective in wireless ad hoc networks, since connectivity along multiple paths is less likely to be broken. We propose a multipath extension to dynamic source routing to support multipath video communication over wireless ad hoc networks. The proposed scheme is compared to others for interactive video applications. Simulations show effectiveness of our proposed scheme Wei Wei 0023, Avideh Zakhor |
ICME | 2 |
| 2004 | Rate Control for Streaming Video over WirelessabstractRate control is an important issue in video streaming applications for both wired and wireless networks. A widely accepted rate control method in wired networks is equation based rate control (Sally Floyd et al., Aug. 2000), in which the TCP friendly rate is determined as a function of packet loss rate, round trip time and packet size. This approach, also known as TFRC, assumes that packet loss in wired networks is primarily due to congestion, and as such is not applicable to wireless networks in which the bulk of packet loss is due to error at the physical layer. We propose multiple TFRC connections as an end-to-end rate control solution for wireless video streaming. We show that this approach not only avoids modifications to the network infrastructure or network protocol, hut also results in full utilization of the wireless channel. NS-2 simulations and experiments over 1/spl times/RTT CDMA wireless data network are carried out to validate, and characterize the performance of our proposed approach. Minghua Chen 0001, Avideh Zakhor |
INFOCOM | 2 |
| 2004 | An Automated Method for Large-Scale, Ground-Based City Model Acquisition
Christian Früh, Avideh Zakhor |
Int. J. Comput. Vis. | 2 |
| 2004 | Dictionary design for matching pursuit and application to motion-compensated video codingabstractWe present a new algorithm for matching pursuit (MP) dictionary design. This technique uses existing vector-quantization design techniques and an inner product-based distortion measure to learn functions from a set of training patterns. While this scheme can be applied to many MP applications, we focus on motion-compensated video coding. Given a set of training sequences, data are extracted from the high-energy packets of the motion-compensated frames. Dictionaries with different regions of support are trained, pruned, and finally evaluated on MPEG test sequences. We find that for high bit-rate QCIF sequences we can achieve improvements of up to 0.66 dB with respect to conventional MP with separable Gabor functions. Philippe Schmid-Saugeon, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Multiple sender distributed video streamingabstractWith the explosive growth of video applications over the Internet, many approaches have been proposed to stream video effectively over packet switched, best-effort networks. We propose a receiver-driven protocol for simultaneous video streaming from multiple senders to a single receiver in order to achieve higher throughput, and to increase tolerance to packet loss and delay due to network congestion. Our receiver-driven protocol employs a novel rate allocation algorithm (RAA) and a packet partition algorithm (PPA). The RAA, run at the receiver, determines the sending rate for each sender by taking into account available network bandwidth, channel characteristics, and a prespecified, fixed level of forward error correction, in such a way as to minimize the probability of packet loss. The PPA, run at the senders based on a set of parameters estimated by the receiver, ensures that every packet is sent by one and only one sender, and at the same time, minimizes the startup delay. Using both simulations and Internet experiments, we demonstrate the effectiveness of our protocol in reducing packet loss. Thinh P. Q. Nguyen, Avideh Zakhor |
IEEE Trans. Multim. | 2 |
| 2003 | Constructing 3D City Models by Merging Ground-Based and Airborne ViewsabstractIn this paper, we present a fast approach to automated generation of textured 3D city models with both high details at ground level, and complete coverage for bird's-eye view. A close-range facade model is acquired at the ground level by driving a vehicle equipped with laser scanners and a digital camera under normal traffic conditions on public roads; a far-range Digital Surface Map (DSM), containing complementary roof and terrain shape, is created from airborne laser scans, then triangulated, and finally texture mapped with aerial imagery. The facade models are first registered with respect to the DSM by using Monte-Carlo-Localization, and then merged with the DSM by removing redundant parts and filling gaps. The developed algorithms are evaluated on a data set acquired in downtown Berkeley. Christian Früh, Avideh Zakhor |
CVPR (2) | 2 |
| 2003 | Binary Combinatorial CodingabstractSummary form only given. A novel binary entropy code, called combinatorial coding (CC), is presented. The theoretical basis for CC has been described previously under the context of universal coding, enumerative coding, and minimum description length. The code described in these references works as follows: assume the source data are binary of length M, memoryless, and generated with an unknown parameter /spl theta/ (the probability that a "1" occurs). The compression efficiency, and encoding and decoding speed of CC against Huffman and arithmetic coding were tested. Over the entire test, CC achieved the compression efficiency of arithmetic coding, together with the coding speed of Huffman coding. Vito Dai, Avideh Zakhor |
DCC | 2 |
| 2003 | Fast similarity search on video signaturesabstractVideo signatures are compact representations of video sequences designed for efficient similarity measurement. In this paper, we propose a feature extraction technique to support fast similarity search on large databases of video signatures. Our proposed technique transforms the high dimensional video signatures into low dimensional vectors where similarity search can be efficiently performed. We exploit both the upper and lower bounds of the triangle inequalities in approximating the high-dimensional metric, and combine this approximation with the classical PCA to achieve the target dimension. Experimental results on a large set of Web video sequences show that our technique outperforms fastmap, Haar wavelet, PCA, and triangle-inequality pruning. Sen-Ching S. Cheung, Avideh Zakhor |
ICIP (2) | 2 |
| 2003 | Effective bandwidth based scheduling for streaming multimediaabstractWe propose a class of packet scheduling algorithms for streaming media. The importance level of a video packet is determined by its relative position within its group of pictures, taking into account the motion-texture discrimination and temporal scalability. We generate a number of nested substreams, with more important streams embedding less important ones in a progressive manner. We model the streaming system as a queueing system, compute the run-time decoding failure probability of a frame in each substream based on effective bandwidth, and determine the optimum substream to be sent at any moment in time. The data within optimum substream is sent based on earliest-deadline-first scheduling, until the next channel report arrives, at which time the optimum substream is recomputed. From experiments with real video data, we show that our proposed scheduling scheme outperforms the conventional sequential sending scheme. Sang H. Kang, Avideh Zakhor |
ICIP (3) | 2 |
| 2003 | Matching pursuits based multiple description video coding for lossy environmentsabstractMultiple description coding (MDC) is an error resilient source coding scheme that creates multiple descriptions of the source with the aim of providing an acceptable reconstruction quality when only one description is received, and improved quality as more descriptions become available. Recently, we developed a matching pursuit multiple description video coder (MP-MDVC) based on a three loop structure originally proposed by Reibman and her colleagues (MTDC). While the MP-MDVC outperforms the discrete cosine transform based MTDC. it is not optimized for lossy environments. In this paper, we extend MP-MDVC by considering the network loss characteristics when coding multiple descriptions. In particular, we propose a fast steepest descent algorithm for creating multiple descriptions that results in minimum expected distortion, given network outage probability, bandwidth constraints, and maximum allowable distortion for each description. Analytical and experimental results show that by taking network loss characteristics into account, our approach outperforms existing MP-MDVC techniques. Thinh P. Q. Nguyen, Avideh Zakhor |
ICIP (1) | 2 |
| 2003 | Path diversity and bandwidth allocation for multimedia streamingabstractThe recent advent of widely available broadband Internet access has resulted in an explosive growth of new video streaming applications and research into methods to efficiently support, such applications over today's Internet. Many approaches, including source and channel coding techniques, have been proposed to deal with the delay, loss, and time-varying characteristics of best-effort packet-switched networks. In this paper we present two class of techniques for such networks: the first set of techniques is designed for streaming to receivers with bandwidth-limited last mile connections to the Internet, while the second set explores techniques relying on path diversity to handle situations in which the path to the video source, and not the access link itself, causes degradation in the quality of the video stream. We show results of simulations and Internet experiments, demonstrating the effectiveness of the two techniques. Thinh P. Q. Nguyen, Puneet Mehra, Avideh Zakhor |
ICME | 3 |
| 2003 | Receiver-Driven Bandwidth Sharing for TCPabstractApplications using TCP, such as Web-browsers, ftp, and various P2P programs, dominate most of the Internet traffic today. In many cases the last-hop access links are bottlenecks due to their limited bandwidth capability with users running many simultaneous network applications. Standard TCP shares bottleneck link capacity according to connection round-trip time (RTT), and may result in a bandwidth partition which does not necessarily coincide with the user's desires. We present a receiver-based control system for allocating bandwidth among TCP flows according to user preferences. Our system does not require any changes to network infrastructure, and works with standard TCP senders. NS-2 simulations, as well as actual Internet experiments, show that our system achieves desired bandwidth allocation in a wide variety of scenarios including interfering cross-traffic. We also demonstrate the viability of our system in multimedia streaming applications over TCP. Puneet Mehra, Christophe De Vleeschouwer, Avideh Zakhor |
INFOCOM | 3 |
| 2003 | Automated reconstruction of building facades for virtual walk-thrusabstractIn this paper, we present a fast, automated approach to generating a highly detailed, textured 3D building facade model. This model is acquired at the ground level by driving a vehicle equipped with laser scanners and a digital camera under normal traffic conditions on public roads, and then processed offline. We evaluate our approach on a large data set acquired in downtown Berkeley. Christian Früh, Avideh Zakhor |
SIGGRAPH | 2 |
| 2003 | Efficient video similarity measurement with video signatureabstractThe proliferation of video content on the Web makes similarity detection an indispensable tool in Web data management, searching, and navigation. We propose a number of algorithms to efficiently measure video similarity. We define video as a set of frames, which are represented as high dimensional vectors in a feature space. Our goal is to measure ideal video similarity (IVS), defined as the percentage of clusters of similar frames shared between two video sequences. Since IVS is too complex to be deployed in large database applications, we approximate it with Voronoi video similarity (VVS), defined as the volume of the intersection between Voronoi cells of similar clusters. We propose a class of randomized algorithms to estimate VVS by first summarizing each video with a small set of its sampled frames, called the video signature (ViSig), and then calculating the distances between corresponding frames from the two ViSigs. By generating samples with a probability distribution that describes the video statistics, and ranking them based upon their likelihood of making an error in the estimation, we show analytically that ViSig can provide an unbiased estimate of IVS. Experimental results on a large dataset of Web video and a set of MPEG-7 test sequences with artificially generated similar versions are provided to demonstrate the retrieval performance of our proposed techniques. Sen-Ching S. Cheung, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | In-loop atom modulus quantization for matching pursuit and its application to video codingabstractThis paper provides a precise analytical study of the selection and modulus quantization of matching pursuit (MP) coefficients. We demonstrate that an optimal rate-distortion trade-off is achieved by selecting the atoms up to a quality-dependent threshold, and by defining the modulus quantizer in terms of that threshold. In doing so, we take into account quantization error re-injection resulting from inserting the modulus quantizer inside the MP atom computation loop. In-loop quantization not only improves coding performance, but also affects the optimal quantizer design for both uniform and nonuniform quantization. We measure the impact of our work in the context of video coding. For both uniform and nonuniform quantization, the precise understanding of the relation between atom selection and quantization results in significant improvements in terms of coding efficiency. At high bitrates, the proposed nonuniform quantization scheme results in 0.5 to 2 dB improvement over the previous method. Christophe De Vleeschouwer, Avideh Zakhor |
IEEE Trans. Image Process. | 2 |
| 2002 | Efficient video similarity measurement with video signatureabstractThe video signature method has previously been proposed as a technique to summarize video efficiently for visual similarity measurements (see Cheung, S.-C. and Zakhor, A., Proc. SPIE, vol.3964, p.34-6, 2000; ICIP2000, vol.1, p.85-9, 2000; ICIP2001, vol.1, p.649-52, 2001). We now develop the necessary theoretical framework to analyze this method. We define our target video similarity measure based on the fraction of similar clusters shared between two video sequences. This measure is too computationally complex to be deployed in database applications. By considering this measure geometrically on the image feature space, we find that it can be approximated by the volume of the intersection between Voronoi cells of similar clusters. In the video signature method, sampling is used to estimate this volume. By choosing an appropriate distribution to generate samples, and ranking the samples based upon their distances to the boundary between Voronoi cells, we demonstrate that our target measure can be well approximated by the video signature method. Experimental results on a large dataset of Web video and a set of MPEG-7 test sequences with artificially generated similar versions are used to demonstrate the retrieval performance of our proposed techniques. Sen-Ching S. Cheung, Avideh Zakhor |
ICIP (1) | 2 |
| 2002 | Protocols for distributed video streamingabstractWith the explosive growth of video applications over the packet switched networks, many approaches have been proposed to stream video effectively over packet switched, best-effort networks. In our previous work, we proposed a framework with a receiver driven protocol to coordinate simultaneous video streaming from multiple senders to a single receiver in order to achieve higher throughput, and to increase tolerance to packet loss and delay due to network congestion. The receiver-driven protocol employs two algorithms: the rate allocation and packet partition. The rate allocation algorithm determines the sending rate for each sender while the packet partition algorithm ensures no sender sends the same packets, and at the same time, minimizes the probability of late packets. We extend the rate allocation scheme to be used with forward error correction (FEC) in order to minimize the probability of packet loss in a bursty loss environment such as one due to network congestion. Using both simulations and actual Internet experiments, we demonstrate the effectiveness of our rate allocation scheme in reducing packet loss, and hence, achieving higher visual quality for the streamed video. Thinh P. Q. Nguyen, Avideh Zakhor |
ICIP (3) | 2 |
| 2002 | Atom modulus quantization for matching pursuit video codingabstractWe provide an analytical study of the selection and modulus quantization of matching pursuits (MP) coefficients. We demonstrate that an optimal rate-distortion trade-off is achieved by selecting the atoms up to a dead-zone threshold, and by defining the modulus quantizer in terms of that threshold. In doing so, we take into account quantization error re-injection resulting from inserting the modulus quantizer inside the MP atom computation loop. In-loop quantization affects the stepsize of the uniform quantizer, and results in a non-uniform optimal entropy constrained quantizer. Improvements larger than one dB are obtained for video coding. Christophe De Vleeschouwer, Avideh Zakhor |
ICIP (3) | 2 |
| 2002 | Cross layer techniques for adaptive video streaming over wireless networksabstractWireless multimedia delivery is becoming increasingly more important in today's networks. Unlike wired packet switched networks that suffer from congestion related loss and delay, the wireless networks have to deal with a time varying, error prone, physical channel that in many instances is also severely bandwidth constrained. As such, the solutions needed for wireless video streaming applications are fundamentally different from wired streaming. In this paper, we propose a set of end to end application layer techniques for adaptive video streaming over wireless networks. The adaptation is done both with respect to channel and data. Our approach combines the flexibility and programmability of application layer adaptations, with low delay and bandwidth efficiency of link layer techniques. Socket level simulations are presented to verify the effectiveness of our approach. Yufeng Shan, Avideh Zakhor |
ICME (1) | 2 |
| 2002 | Matching pursuit video coding .I. Dictionary approximationabstractWe have shown in previous works that overcomplete signal decomposition using matching pursuits is an efficient technique for coding motion-residual images in a hybrid video coder. Others have shown that alternate basis sets may improve the coding efficiency or reduce the encoder complexity. In this work, we introduce for the first time a design methodology which incorporates both coding efficiency and complexity in a systematic way. The key to the method is an algorithm which takes an arbitrary 2-D dictionary and generates approximations of the dictionary which have fast two-stage implementations according to the method of Redmill et al. (see Proc. IEEE Int. Conf. Image Processing, p.769-773, 1998). By varying the quality of the approximation, we can explore a systematic tradeoff between the coding efficiency and complexity of the resulting matching pursuit video encoder. As a practical result, we show that complexity reduction factors of up to 1000 are achievable with negligible coding efficiency losses of about 0.1-dB PSNR. Ralph Neff, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | Matching-pursuit video coding .II. Operational models for rate and distortionabstractFor pt. I see ibid., vol.12 , no.1, (2000).We introduce two models for predicting the rate and distortion of the matching-pursuit video codec. The first model is based on a pre-coding analysis pass using the full matching-pursuit dictionary. The second model is based on a reduced-complexity analysis pass. We evaluate these models for use within existing rate-distortion optimization techniques. Our prediction results suggest that the models have sufficient accuracy to be useful in this context, and that significant complexity reductions could be achieved compared to exact rate-distortion computation. Ralph Neff, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | Corrections to "matching pursuit video coding-part I: dictionary approximation"
Ralph Neff, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | Matching pursuits multiple description coding for wireless videoabstractMultiple description coding (MDC) is an error-resilient source coding scheme that creates multiple bitstreams of approximately equal importance. The reconstructed signal based on any single bitstream has an acceptable quality. However, a higher quality reconstruction can be achieved with a larger number of bitstreams. We develop a multiple description video coding scheme based on a three-loop structure. We modify the discrete cosine transform structure to the matching pursuits framework and evaluate performance gain using maximum-likelihood (ML) enhancement when both descriptions are available. We find that ML enhancement works best for low motion sequences and results in gains of up to 1.3 dB in terms of average PSNR. Rate distortion performance is characterized. Performance comparison is made between our MDC scheme and single description coding (SDC) schemes over lossy channels, including two state Markov channels and Rayleigh fading channels. We find that MDC outperforms SDC in bursty slowly varying environments. In the case of Rayleigh fading channels, interleaving helps SDC close the gap and even outperform MDC depending on the amount of interleaving performed, at the expense of additional delay. Xiaoyi Tang, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | 3D Model Generation for Cities Using Aerial Photographs and Ground Level Laser ScansabstractIn this paper we describe techniques for 3D textured model construction of urban areas using acquisition devices such as intensity cameras, as well as 2D laser scanner. Our experimental set up consists of a truck equipped with one camera and two fast, inexpensive 2D laser scanner, traveling on city streets under normal traffic conditions. The horizontal laser scans are used to determine the approximate component of motion along the movement of the acquisition vehicle. The vertical scanner is used to build 3D models of the facade of the buildings. To improve the accuracy of localization of the truck and hence our resulting 3D models of the city, two different methods are developed and compared: the first method employs a correlation technique and the second method is based on Markov Monte Carlo localization. Both techniques use digital road maps and aerial photographs in conjunction with laser scans. A fairly accurate textured, 3D model of downtown area has been acquired in a matter of few minutes, limited only by traffic conditions during the data acquisition phase. Christian Früh, Avideh Zakhor |
CVPR (2) | 2 |
| 2001 | Video similarity detection with video signature clusteringabstractThe proliferation of video content on the Web makes similarity detection an indispensable tool in Web data management, searching, and navigation. We have previously proposed a compact representation of video clips, called video signature, for retrieving similar video clips in large databases. In this paper, we propose a new signature clustering algorithm to further improve retrieval performance. The algorithm treats all the signatures as an abstract threshold graph, where the threshold is determined based on local data statistics. Similar clusters are identified as highly connected regions in the graph. This algorithm outperforms simple thresholding and hierarchical clustering techniques in identifying a set of manually-determined similar clusters from a dataset of 46,356 Web video clips. At 95% precision, our algorithm attains 85% recall while simple thresholding and complete-link hierarchical scheme attain 67% and 75% recall respectively. Applying our algorithm to the entire dataset, 6,900 similar clusters are identified, with an average cluster size of 2.81 video clips. The distribution of cluster sizes follows a power-law distribution, which has been shown to describe many Web phenomena. Sen-Ching S. Cheung, Avideh Zakhor |
ICIP (2) | 2 |
| 2001 | Learning dictionaries for matching pursuits based video codersabstractWe present a learning scheme for designing dictionaries of two-dimensional functions for matching pursuits (MP) based video coding. The motivation is to improve the performance of such codecs by adapting the structure of the dictionary functions to specific bit-rates or types of sequences. The scheme we propose is based on vector quantization (VQ), and uses an inner-product based distortion measure. The different processing steps, consisting of data extraction from the motion compensated error frames, training, pruning, and testing, are presented in detail. We find that for high bit-rate QCIF sequences we can achieve improvements of up to 0.66 dB. Philippe Schmid-Saugeon, Avideh Zakhor |
ICIP (3) | 2 |
| 2001 | Matching pursuits multiple description coding for wireless videoabstractMultiple description coding (MDC) is an error resilient source coding scheme that creates multiple bitstreams of approximately equal importance. We develop a 2 description video coding scheme based on the 3 loop structure originally studied in Reibman et al. (1999). We modify the discrete cosine transform structure to the matching pursuits framework and evaluate performance gain using maximum likelihood (ML) enhancement when both descriptions are available. We find that ML enhancement works best for low motion sequences. Performance comparison is made between our MDC scheme and single description coding (SDC) schemes over two-state Markov channels and Rayleigh fading channels. We find that MDC outperforms SDC in bursty slowly varying environments. In the case of Rayleigh fading channels, interleaving helps SDC close the gap and even outperform MDC depending on the amount of interleaving performed, at the expense of additional delay. Xiaoyi Tang, Avideh Zakhor |
ICIP (1) | 2 |
| 2001 | Constructing a Multivalued Representation for View Synthesis
Nelson L. Chang, Avideh Zakhor |
Int. J. Comput. Vis. | 2 |
| 2001 | Video multicast using layered FEC and scalable compressionabstractThe use of scalable video with layered multicast has been shown to be an effective method to achieve rate control in heterogeneous networks. We propose the use of layered forward error correction (FEC) as an error-control mechanism in a layered multicast framework. By organizing FEC into multiple layers, receivers can obtain different levels of protection commensurate with their respective channel conditions. Efficient network utilization is achieved as FEC streams are multicast, and only to receivers that need them. Furthermore, FEC is used without overall rate expansion by selectively dropping data layers to make room for FEC layers. Effects of bursty losses are amortized by staggering the FEC streams in time, giving rise to a tradeoff between delay and quality. For rate control at the receivers, we propose an equation-based approach that computes network usage as a function of measured network characteristics. We show that equation-based rate control achieves more fair bandwidth sharing amongst competing sessions as compared to existing multicast rate control schemes such as RLM and RLC. Fairness is achieved since competing sessions sharing a path will measure similar network characteristics. Simulations and actual MBONE experiments are performed using error-resilient, scalable video compression. We find that video quality is significantly improved at the same communication rate when layered FEC is used. Wai-tian Tan, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Efficient Video Similarity Measurement and SearchabstractWe consider the use of meta-data and/or video-domain methods to detect similar videos on the Web. Meta-data is extracted from the textual and hyperlink information associated with each video clip. In the video domain, we apply an efficient similarity detection algorithm called video signature. The idea is to form a signature for each clip by selecting a small number of its frames that are most similar to a set of random seed images. We then apply a statistical pruning algorithm to allow fast detection on very large databases. Using a small ground-truth set, we achieve 90% recall and 95% precision using only 8% of the total number of operations required without pruning. For a database of around 46,000 video clips crawled from the Web, the video signature technique significantly outperforms meta-data in precision and recall. We show that even better performance can be achieved by combining them together. Based on our measurements, each video clip in our database has, on average, 1.53 similar copies. Sen-Ching S. Cheung, Avideh Zakhor |
ICIP | 2 |
| 2000 | Dictionary Approximation for Matching Pursuit Video CodingabstractPreviously, we demonstrated an efficient video codec based on overcomplete signal decomposition using matching pursuits. Dictionary design is an important issue for this system, and others have shown alternate dictionaries which lead to either coding efficiency improvements or reduced encoder complexity. We introduce for the first time a design methodology which incorporates both coding efficiency and complexity in a systematic way. The key to our new method is an algorithm which takes an arbitrary 2-D dictionary and generates approximations of the dictionary which have fast 2-stage implementations. By varying the quality of the approximation, we can explore a systematic tradeoff between the coding efficiency and complexity of the matching pursuit video encoder. As a practical result, we show cases where complexity is reduced by a factor of 500 to 1000 in exchange for small coding efficiency losses of around 0.1 dB PSNR. Ralph Neff, Avideh Zakhor |
ICIP | 2 |
| 2000 | Performance Analysis of an H.263 Video Encoder for VIRAMabstractVIRAM (vector intelligent random access memory) is a vector architecture processor with embedded memory, designed for portable multimedia processing devices. Its vector processing capability results in high performance multimedia processing, while embedded DRAM technology provides high memory bandwidth with low energy consumption. We evaluate and compare the performance of VIRAM to digital signal processors (DSPs) and conventional SIMD (single instruction multiple data) media extensions in the context of video coding. In particular, we examine motion estimation (ME) and the discrete cosine transform (DCT) which have been shown to dominate typical video encoders such as H.263. We show that VIRAM outperforms other architectures by 4.6/spl times/ to 8.7/spl times/ in computing ME and by 1.2/spl times/ to 5.0/spl times/ in computing DCT. Thinh P. Q. Nguyen, Avideh Zakhor, Katherine A. Yelick |
ICIP | 2 |
| 2000 | Modulus quantization for matching-pursuit video codingabstractOvercomplete signal decomposition using matching pursuits has been shown to be an efficient technique for coding motion-residual images in a hybrid video coder. Unlike orthogonal decomposition, matching pursuit uses an in-the-loop modulus quantizer which must be specified before coding begins. This complicates the quantizer design, since the optimal quantizer depends on the statistics of the matching-pursuit coefficients which in turn depend on the in loop quantizer actually used. In this paper, we address the modulus quantizer design issue, specifically developing frame-adaptive quantization schemes for the matching-pursuit video coder. Adaptive dead-zone subtraction is shown to reduce the information content of the modulus source, and a uniform threshhold quantizer is shown to be optimal for the resulting source. Practical two-pass and one-pass algorithms are developed to jointly determine the quantizer parameters and the number of coded basis functions in order to minimize coding distortion for a given rate. The compromise one-pass scheme performs nearly as well as the full two-pass algorithm, but with the same complexity as a fixed-quantizer design. The adaptive schemes are shown to outperform the fixed quantizer used in earlier works, especially at high bit rates, where the gain is as high as 1.7 dB. Ralph Neff, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Bit allocation for joint source/channel coding of scalable videoabstractWe propose an efficient bit allocation algorithm for a joint source/channel video codec over noisy channels. The approach is to distribute the available source and channel coding bits among the subbands in such a way that the expected distortion is minimized. The constructed distortion curves bound the performance degradation should the channel be estimated incorrectly. The algorithm can be used in other similar distortion minimization problems with two constraints, such as power or complexity. Gene Cheung, Avideh Zakhor |
IEEE Trans. Image Process. | 2 |
| 1999 | A Multivalued Representation for View SynthesisabstractWe propose a new depth-based representation of 3-D scenes primarily for the problem of arbitrary view synthesis. The information contained in a given image sequence is compacted into a single multivalued array, where points are organized by occlusions rather than by coherent affine motions. This grouping facilitates an automatic process to determine the number of layers and helps to reduce the artifacts caused by occlusions in the scene. In addition, an iterative multiframe dynamic programming algorithm is described to produce piecewise smooth depth maps. A novel multiframe segmentation, tracking, and plane fitting algorithm is also proposed to handle the traditionally difficult low-contrast regions. Reconstructed views as well as arbitrary flyarounds of real scenes are presented to demonstrate the effectiveness of the approach. Nelson L. Chang, Avideh Zakhor |
ICIP (2) | 2 |
| 1999 | Rate Control for Layered Video Compression Using Matching PursuitsabstractThe use of matching pursuits (MP) to compress motion compensated residual video signals and its extension to layered coding have recently been demonstrated. In this paper we describe a multi-pass rate control scheme for SNR scalable encoding based on MP. The rate control algorithm enforces constant quality on each layer, while keeping the bit budget for each layer at a pre-specified target level. We formulate this as a zero finding problem, and solve it using Newton's method. Experimental results on 14 video Sequences are included, showing that layered video can be encoded at constant quality in about 5 encoding iterations per layer, while satisfying bit budget constraints with 1.5% tolerance. Eugene Miloslavsky, Avideh Zakhor |
ICIP (2) | 2 |
| 1999 | Adaptive Modulus Quantizer Design for Matching Pursuit Video CodingabstractOvercomplete signal decomposition using matching pursuits has been shown to be an efficient technique for coding motion residual images in a hybrid video coder. Unlike orthogonal decomposition where computation of the transform coefficients is decoupled from quantization, matching pursuit uses an in-loop quantizer which must be specified before coding begins. Optimal quantizer design thus depends on the computed matching pursuit coefficients, but these in turn depend on the chosen quantizer. To resolve this interdependency, we propose frame-adaptive quantization for matching pursuit based on adaptive dead-zone subtraction followed by uniform threshold quantization. Practical 2-pass and 1-pass algorithms are developed which jointly find the quantizer parameters and the number of coded basis functions which minimize coding distortion for a given rate. The compromise 1-pass scheme performs nearly as well as the full 2-pass algorithm, but with the same complexity as a fixed quantizer design. The adaptive schemes are shown to outperform the fixed quantizer used in earlier works, especially at high bit rates where the gain is up to 1.7 dB. Ralph Neff, Avideh Zakhor |
ICIP (2) | 2 |
| 1999 | Error Control for Video Multicast Using Hierarchical FecabstractBit-rate scalable video compression with layered multicast has been shown to be an effective method to achieve rate control in heterogeneous networks. We further propose the use of hierarchical FEC as an error control mechanism that allows receivers to individually trade-off latency for received video quality. The scheme is efficient since FEC packets are used to protect only the more important data layers and is multicast only to receivers that need them, thereby improving network utilization. Furthermore, there is no loss in error correcting capability by using hierarchical FEC when maximum distance separable codes are used. Actual MBONE experiments are performed to evaluate the performance of the proposed scheme. Wai-tian Tan, Avideh Zakhor |
ICIP (1) | 2 |
| 1999 | Video portals for the next century (panel session)
Nevenka Dimitrova, Rob Koenen, Hong Heather Yu, Avideh Zakhor, Francis Galliano, Charles A. Bouman |
ACM Multimedia (1) | 4 |
| 1999 | Video compression using matching pursuitsabstractThe use of matching pursuit (MP) to code video using overcomplete Gabor basis functions has recently been introduced. In this paper, we propose new functionalities such as SNR scalability and arbitrary shape coding for video coding based on matching pursuit. We improve the performance of the baseline algorithm presented earlier by proposing a new search and a new position coding technique. The resulting algorithm is compared to the earlier one and to DCT-based coding. Osama K. Al-Shaykh, Eugene Miloslavsky, Toshio Nomura, Ralph Neff, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 1999 | Content analysis of video using principal componentsabstractWe use principal component analysis (PCA) to reduce the dimensionality of features of video frames for the purpose of content description. This low-dimensional description makes practical the direct use of all the frames of a video sequence in later analysis. The PCA representation circumvents or eliminates several of the stumbling blocks in current analysis methods and makes new analyses feasible. We demonstrate this with two applications. The first accomplishes high-level scene description without shot detection and key-frame selection. The second uses the time sequences of motion data from every frame to classify sports sequences. Emile Sahouria, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Resolution enhancement of color video sequencesabstractWe propose a new multiframe algorithm to enhance the spatial resolution of frames in video sequences. Our technique specifically accounts for the possibility that motion estimation will be inaccurate and compensates for these inaccuracies. Experiments show that our multiframe enhancement algorithm yields perceptibly sharper enhanced images with significant signal-to-noise ratio (SNR) improvement over bilinear and cubic B-spline interpolation. Nimish R. Shah, Avideh Zakhor |
IEEE Trans. Image Process. | 2 |
| 1999 | Real-Time Internet Video Using Error Resilient Scalable Compression and TCP-Friendly Transport ProtocolabstractWe introduce a point to point real-time video transmission scheme over the Internet combining a low-delay TCP-friendly transport protocol in conjunction with a novel compression method that is error resilient and bandwidth-scalable. Compressed video is packetized into individually decodable packets of equal expected visual importance. Consequently, relatively constant video quality can be achieved at the receiver under lossy conditions. Furthermore, the packets can be truncated to instantaneously meet the time varying bandwidth imposed by a TCP-friendly transport protocol. As a result, adaptive flows that are friendly to other Internet traffic are produced. Actual Internet experiments together with simulations are used to evaluate the performance of the compression, transport, and the combined schemes. Wai-tian Tan, Avideh Zakhor |
IEEE Trans. Multim. | 2 |
| 1998 | Finite Sensor Effects for Estimating Structure-from-MotionabstractWe propose a computational framework for assessing uncertainty of structure estimates under perspective projection for a given camera configuration. Unlike probabilistic models, the approach takes into account finite sensor areas and relates the uncertainty in estimating a set of points to the volumes formed by the intersections of their associated frustums. Simulations demonstrate and compare the effectiveness of lateral and semicircular motion for structure recovery of arbitrarily-shaped objects. Nelson L. Chang, Avideh Zakhor |
ICIP (1) | 2 |
| 1998 | Decoder Complexity and Performance Comparison of Matching Pursuit and DCT-based MPEG-4 Video CodesabstractMatching pursuits is an overcomplete expansion technique which has been successfully applied to the problem of coding motion residual images in a hybrid video coder. In this paper, the coding efficiency and decoder complexity of the method are compared to that of the DCT-based MPEG-4 standard with and without post-processing. Without post-processing, matching pursuits is shown to have significantly better PSNR and visual quality and similar decoding complexity compared to the MPEG-4 DCT decoder. To achieve reasonable quality at low bit rates, the DCT-based scheme requires post-processing, while the patching pursuit scheme does not. We show that the MPEG-4 post-processing filters have a prohibitive cost, increasing decoder complexity by a factor of 3 to 8. Finally, we introduce an all-integer matching pursuit implementation. The performance is shown to be within 0.05 dB of the original floating point algorithm. Ralph Neff, Toshio Nomura, Avideh Zakhor |
ICIP (1) | 3 |
| 1998 | Content Analysis of Video using Principal Componets
Emile Sahouria, Avideh Zakhor |
ICIP (3) | 2 |
| 1998 | Internet Video using Error Resilient Scalable Compression and Cooperative Transport Protocol
Wai-tian Tan, Avideh Zakhor |
ICIP (3) | 2 |
| 1997 | Motion Indexing of VideoabstractA system has been developed to analyze and index surveillance videos based on the motions of objects in the scene. A segmentation and tracking system extracts trajectories from compressed video, which are represented in a multiresolution manner and stored in a database. Hand-drawn queries can be submitted to the system for imprecise searches. The system was tested on real video footage with numerous moving objects; the success rates for tracking and recall in most cases exceeded 70 percent, a promising result. Emile Sahouria, Avideh Zakhor |
ICIP (2) | 2 |
| 1997 | Disk-based storage for scalable videoabstractWe consider the placement of scalable video data on single and multiple disks for storage and real-time retrieval. For the single-disk case, we extend the principle of constant frame grouping from constant bit rate (CBR) to variable bit rate (VBR) scalable video data. When the number of admitted users exceeds the server capacity, the rate of data sent to each user is reduced to relieve the disk system overload, offering a graceful degradation in comparison with nonscalable data. We examine the qualities of video reconstructions obtained from a real disk video server and find the scalable video more visually appealing. In the VBR case, scalability is also used to improve interactivity by reducing the delay associated with using interactive functions in a predictive admission control environment. Finally, we consider the multiple disk scenario and prove that periodic interleaving results in lower system delay than striping in a video server using round-robin scheduling. We verify the results through detailed simulation of a four-disk array. Edward Y. Chang, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | Very low bit-rate video coding based on matching pursuitsabstractWe present a video compression algorithm which performs well on generic sequences at very low bit rates. This algorithm was the basis for a submission to the November 1995 MPEG-4 subjective tests. The main novelty of the algorithm is a matching-pursuit based motion residual coder. The method uses an inner-product search to decompose motion residual signals on an overcomplete dictionary of separable Gabor functions. This coding strategy allows residual bits to be concentrated in the areas where they are needed most, providing detailed reconstructions without block artifacts. Coding results from the MPEG-4 Class A compression sequences are presented and compared to H.263. We demonstrate that the matching pursuit system outperforms the H.263 standard in both peak signal-to-noise ratio (PSNR) and visual quality. Ralph Neff, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | View generation for three-dimensional scenes from video sequencesabstractThis paper focuses on the representation and view generation of three-dimensional (3-D) scenes. In contrast to existing methods that construct a full 3-D model or those that exploit geometric invariants, our representation consists of dense depth maps at several preselected viewpoints from an image sequence. Furthermore, instead of using multiple calibrated stationary cameras or range scanners, we derive our depth maps from image sequences captured by an uncalibrated camera with only approximately known motion. We propose an adaptive matching algorithm that assigns various confidence levels to different regions in the depth maps. Nonuniform bicubic spline interpolation is then used to fill in low confidence regions in the depth maps. Once the depth maps are computed at preselected viewpoints, the intensity and depth at these locations are used to reconstruct arbitrary views of the 3-D scene. Specifically, the depth maps are regarded as vertices of a deformable 2-D mesh, which are transformed in 3-D, projected to 2-D, and rendered to generate the desired view. Experimental results are presented to verify our approach. Nelson L. Chang, Avideh Zakhor |
IEEE Trans. Image Process. | 2 |
| 1996 | Deconvolution of electrograms to detect infarcted myocardiumabstractAblation to prevent cardiac arrhythmias requires interpretation of electrograms to locate the arrhythmogenic tissue. This study examined a novel signal processing technique employing deconvolution to calculate electrograms which best fit observed electrograms. We hypothesize that the power of difference between the calculated and the observed electrogram detects changes in morphology resulting from myocardial infarction. 380 electrograms were recorded from 10 dogs. Scintigraphic studies with Thallium-201 identified recording sites as normal or infarcted tissue. The power of the difference increased 65 percent for infarcted tissue as compared to normal tissue (p<0.0001). Receiver operating curves were plotted, and deconvolution detected 80% of the infarcted sites with a 20% false positive rate. Deconvolution's performance for detection of infarcted tissue was superior to previous metrics and shows promise for improving clinical electrogram interpretation. Willard S. Ellis, Susan J. Eisenberg, David M. Auslander, Michael W. Dae, Avideh Zakhor, Michael D. Lesh |
ICASSP | 5 |
| 1996 | Joint source/channel coding of scalable video over noisy channelsabstractWe propose an optimal bit allocation strategy for a joint source/channel video codec over noisy channel when the channel state is assumed to be known. Our approach is to partition source and channel coding bits in such a way that the expected distortion is minimized. The particular source coding algorithm we use is rate scalable and is based on 3D subband coding with multi-rate quantization. We show that using this strategy, transmission of video over very noisy channels still renders acceptable visual quality, and outperforms schemes that use equal error protection only. The flexibility of the algorithm also permits the bit allocation to be selected optimally when the channel state is in the form of a probability distribution instead of a deterministic state. Gene Cheung, Avideh Zakhor |
ICIP (3) | 2 |
| 1996 | Multiframe spatial resolution enhancement of color videoabstractThis paper focuses on the spatial resolution enhancement of a video sequence. In contrast to previous work with grayscale images and highly constrained motion, we present a technique for color video frames with general motion. The method consists of determining subpixel motion estimation between video frames and subsequently using these estimates along with the original low resolution frames to iteratively create a sequence of enhanced resolution frames. We present a novel motion estimation technique based on determining a set of candidate motion estimates per pixel. Experimental results for video sequences containing general motion are presented to verify our technique. Enhanced frames using this technique show significant improvement in both SNR and perceived visual quality over methods such as bilinear and cubic B-spline interpolation. Nimish R. Shah, Avideh Zakhor |
ICIP (1) | 2 |
| 1996 | A common framework for rate and distortion based scaling of highly scalable compressed videoabstractScalability refers to the ability to modify the resolution and/or bit rate associated with an already compressed data source in order to satisfy requirements which could not be foreseen at the time of compression. A number of researchers have already demonstrated the feasibility of efficient scalable image and video compression. The principle focus of this paper is to describe data structures for highly scalable compressed video, which are able to support simple, generic scaling approaches for both constant bit rate and constant distortion scaling criteria. Interactive video material presents particular challenges when the data stream is to be scaled to maintain an approximately constant level of distortion, rather than just a constant bit rate. Special attention is paid, therefore, to the development of generic, robust scaling algorithms for such applications. The data structures and scaling methodologies developed are particularly appealing for the distribution of highly scalable compressed video over heterogeneous media, because they simultaneously support both variable bit rate (VBR) and constant bit rate (CBR) services with a wide range of available service qualities, using only simple, generic mechanisms for scaling. The performance of the proposed scaling methodologies is experimentally investigated using a highly scalable video compression algorithm, which is able to achieve comparable compression performance to that of the inherently nonscalable MPEG-1 compression standard. David S. Taubman, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1995 | Matching Pursuit Video Coding at Very Low Bit RatesabstractMatching pursuits refers to a greedy algorithm which matches structures in a signal to a large dictionary of functions. In this paper, we present a matching-pursuit based video coding system which codes motion residual images using a large dictionary of Gabor functions. One feature of our system is that bits are assigned progressively to the highest-energy areas in the motion residual image. The large dictionary size is another advantage, since it allows structures in the motion residual to be represented using few significant coefficients. Experimental results compare the performance of the matching-pursuit system to a hybrid-DCT system at various bit rates between 6 and 128 kbit/s. Additional experiments show how the matching pursuit system performs if the Gabor dictionary is replaced by an 8/spl times/8 DCT dictionary. Ralph Neff, Avideh Zakhor |
Data Compression Conference | 2 |
| 1995 | Arbitrary view generation for three-dimensional scenes from uncalibrated video camerasabstractThis paper focuses on the representation and arbitrary view generation of three dimensional (3-D) scenes. In contrast to existing methods that construct a full 3-D model or those that exploit geometric invariants, our representation consists of dense depth maps at several preselected viewpoints from an image sequence. Furthermore, instead of using multiple calibrated stationary cameras or range data, we derive our depth maps from image sequences captured by an uncalibrated camera. We propose an adaptive matching algorithm which assigns various confidence levels to different regions. Nonuniform bicubic spline interpolation is then used to fill in low confidence regions in the depth maps. Once the depth maps are computed at preselected viewpoints, the intensity and depth at these locations are used to reconstruct arbitrary views of the 3-D scene. Experimental results are presented to verify our approach. Nelson L. Chang, Avideh Zakhor |
ICASSP | 2 |
| 1995 | Depth based recovery of human facial features from video sequencesabstractWe propose a way to locate facial features from a video sequence captured by a camcorder undergoing strong translational motion. Pairs of stereo images containing frontal views of the human subject are sampled from the video sequence. A multiresolution hierarchical matching algorithm finds point correspondences over a large disparity range. The task of locating facial features such as the eyes, nose and mouth is aided by depth information derived from the matching data. We present experimental results to verify our approach. G. Galicia, Avideh Zakhor |
ICIP | 2 |
| 1995 | An optimization approach for removing blocking effects in transform codingabstractOne drawback of the discrete cosine transform (DCT) is visible block boundaries due to coarse quantization of the coefficients. Most restoration techniques for the removing blocking effect are variations of low-pass filtering, and as such, result in unnecessary blurring. The authors propose a new approach for reducing the blocking effect which can be applied to conventional transform coding without introducing additional information or significant blurring. The method exploits the correlation between the intensity values of boundary pixels of two neighboring blocks. It is based on the theoretical and empirical observation that under mild assumptions, quantization of the DCT coefficients of two neighboring blocks increases the expected value of the mean squared difference of slope (MSDS) between the slope across two adjacent blocks, and the average between the boundary slopes of each of the two blocks. The amount of this increase is dependent upon the width of quantization intervals of the transform coefficients. Therefore, among all permissible inverse quantized coefficients, the set which reduces the expected value of this MSDS by an appropriate amount is most likely to decrease the blocking effect. To estimate the set of unquantized coefficients, the authors solve a constrained quadratic programming problem. The approach is based on the gradient projection method. It is shown that from a subjective viewpoint, the blocking effect is less noticeable in the author' processed images than in the ones using existing filtering techniques.> Shigenobu Minami, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1995 | Halftone to continuous-tone conversion of error-diffusion coded imagesabstractConsiders the problem of reconstructing a continuous-tone (contone) image from its halftoned version, where the halftoning process is done by error diffusion. The authors present an iterative nonlinear decoding algorithm for halftone-to-contone conversion and show simulation results that compare the performance of the algorithm to that of conventional linear low-pass filtering. They find that the new technique results in subjectively superior reconstruction. As there is a natural relationship between error diffusion and SigmaDelta modulation, the reconstruction algorithm can also be applied to the decoding problem for SigmaDelta modulators. Søren Hein, Avideh Zakhor |
IEEE Trans. Image Process. | 2 |
| 1994 | Rate and resolution scalable subband coding of videoabstractWe propose a full colour video compression strategy, based on 3-D subband coding with camera pan compensation, to generate a single, embedded, compressed bit stream supporting multiple decoder display formats and a wide, finely gradated range of bit rates. An experimental implementation of our algorithm produces a single bit stream, from which suitable subsets are extracted to be compatible with many useful decoder frame sizes and frame rates and to satisfy transmission bandwidth constraints ranging from several tens of kilo-bits per second to several mega-bits per second. The reconstructed video quality from any of these bit stream subsets is often found to exceed that obtained from an MPEG-1 implementation, operated with equivalent bit rate constraints, in both perceptual quality and mean squared error.> David S. Taubman, Avideh Zakhor |
ICASSP (5) | 2 |
| 1994 | Admissions Control and Data Placement for VBR Video ServersabstractIn this paper we compare techniques for storage and real-time retrieval of variable bit rate (VBR) video data for multiple simultaneous users. We compare two main classes of data placement techniques: constant time length (CTL) and constant data length (CDL). For each technique, we consider both deterministic and statistical admissions control policies and show that the statistical policies are more suitable for interactive applications. CDL-based data placement is shown to be able to achieve lower overload probabilities than CTL-based ones for a given user distribution at the expense of a much higher buffer requirement and higher delays. A "burst mode" technique for CDL is discussed that can reduce the delays at the expense of even higher buffer usage. We verify our theoretical overload probabilities for the statistical admissions control policies with experiments run on a discrete event simulator.> Edward Y. Chang, Avideh Zakhor |
ICIP (1) | 2 |
| 1994 | Highly Scalable, Low-Delay Video CompressionabstractWe propose a class of scalable video compression algorithms, within which compression performance may be exchanged for end-to-end delay. Each of the algorithms in this class produces a highly scalable bit stream, from which subsets may be extracted for compatibility with a wide range of display frame sizes, frame rates and bit rate constraints. Moreover, with modest end-to-end delay, the reconstructed video quality associated with any of these subsets is often superior to that obtained using an implementation of the inherently non-scalable MPEG-1 compression standard, operated with equivalent resolution and bit rate constraints. We describe scalable compressed data structures based on a layered substream abstraction, with simple, generic scaling operations, for both constant bit rate and constant distortion scaling criteria.> David S. Taubman, Avideh Zakhor |
ICIP (1) | 2 |
| 1994 | Analysis of Tones in the Double Loop SigmaDelta Modulator with Unstable Filter DynamicsabstractIn this paper we analyze the tone behavior of the double loop /spl Sigma//spl Delta/ modulator with unstable filter dynamics. We show that some unstable limit cycles have an attractor region in their neighborhood which results in tones in the spectrum corresponding to the fundamental or harmonics of these limit cycles. We develop the tools to determine whether or not an unstable limit cycle has an attractor in its neighborhood and obtain a method to approximate the boundaries of these attractors.> Mariam Motamed, Seth R. Sanders, Avideh Zakhor |
ISCAS | 3 |
| 1994 | Inverse Problem and Approximation of Fractal-like ImagesabstractIn this paper we consider an inverse problem and an approximation problem for fractal images. The inverse problem involves finding the Iterated Function System (IFS) parameters for a class of signals that are exactly generated via an IFS. We make use of the wavelet transform and of the image moments to solve the inverse problem. The approximation problem involves finding a fractal IFS generated image whose moments either match exactly or in a mean squared error sense a range of moments of the original image. The approximating measures are generated by an IFS model of a special form, and provide a general basis for the approximation of arbitrary images. Experimental results verifying our approach are presented.> Roberto Rinaldo, Avideh Zakhor |
ISCAS | 2 |
| 1994 | Inverse and approximation problem for two-dimensional fractal sets abstractThe geometry of fractals is rich enough that they have extensively been used to model natural phenomena and images. Iterated function systems (IFS) theory provides a convenient way to describe and classify deterministic fractals in the form of a recursive definition. As a result, it is conceivable to develop image representation schemes based on the IFS parameters that correspond to a given fractal image. In this paper, we consider two distinct problems: an inverse problem and an approximation problem. The inverse problem involves finding the IFS parameters of a signal that is exactly generated via an IFS. We make use of the wavelet transform and of the image moments to solve the inverse problem. The approximation problem involves finding a fractal IFS-generated image whose moments match, either exactly or in a mean squared error sense, a range of moments of the original image. The approximating measures are generated by an IFS model of a special form and provide a general basis for the approximation of arbitrary images. Experimental results verifying our approach will be presented. Roberto Rinaldo, Avideh Zakhor |
IEEE Trans. Image Process. | 2 |
| 1994 | Orientation adaptive subband coding of imagesabstractIn the subband coding of images, directionality of image features has thus far been exploited very little. The proposed subband coding scheme utilizes orientation of local image features to avoid the highly objectionable Gibbs-like phenomena observed at reconstructed image edges with conventional subband schemes at low bit rates, At comparable bit rates, the subjective image quality obtained by our orientation adaptive scheme is considerably enhanced over a conventional separable subband coding scheme, as well as other separable approaches such as the JPEG compression standard. David S. Taubman, Avideh Zakhor |
IEEE Trans. Image Process. | 2 |
| 1994 | Multirate 3-D subband coding of videoabstractWe propose a full color video compression strategy, based on 3-D subband coding with camera pan compensation, to generate a single embedded bit stream supporting multiple decoder display formats and a wide, finely gradated range of bit rates. An experimental implementation of our algorithm produces a single bit stream, from which suitable subsets are extracted to be compatible with many decoder frame sizes and frame rates and to satisfy transmission bandwidth constraints ranging from several tens of kilobits per second to several megabits per second. Reconstructed video quality from any of these bit stream subsets is often found to exceed that obtained from an MPEG-1 implementation, operated with equivalent bit rate constraints, in both perceptual quality and mean squared error. In addition, when restricted to 2-D, the algorithm produces some of the best results available in still image compression. David S. Taubman, Avideh Zakhor |
IEEE Trans. Image Process. | 2 |
| 1993 | Scalable video coding using 3-D subband velocity coding and multirate quantization
Edward Chang, Avideh Zakhor |
ICASSP (5) | 2 |
| 1993 | Halftone to continuous-tone conversion of error-diffusion coded images
Søren Hein, Avideh Zakhor |
ICASSP (5) | 2 |
| 1993 | Tones, Saturation, and SNR in Doubel Loop Sigma Delta Modulators
Mariam Motamed, Avideh Zakhor, Seth R. Sanders |
ISCAS | 2 |
| 1993 | Orientation adaptive subband coding of images
David S. Taubman, Avideh Zakhor |
ISCAS | 2 |
| 1993 | Edge-based 3-D camera motion estimation with application to video codingabstractTwo classes of algorithms for modeling camera motion in video sequences captured by a camera are proposed. The first class can be applied when there is no camera translation and the motion of the camera can be adequately modeled by zoom, pan, and rotation parameters. The second class is more general in that it can be applied when the camera is undergoing a translation motion, as well as a rotation and zoom and pan. This class uses seven parameters to describe the motion of the camera and requires the depth map to be known at the receiver. The salient feature of both algorithms is that the camera motion is estimated using binary matching of the edges in successive frames. The rate distortion characteristics of the algorithms are compared with that of the block matching algorithm and show that the former provide performance characteristics similar to those of the latter with reduced computational complexity. Avideh Zakhor, Francesco Lari |
IEEE Trans. Image Process. | 1 |
| 1993 | A new class of B/W halftoning algorithmsabstractA new class of dithering algorithms for black and white (B/W) images is presented. The basic idea behind the technique is to divide the image into small blocks and minimize the distortion between the original continuous-tone image and its low-pass-filtered halftone. This corresponds to a quadratic programming problem with linear constraints, which is solved via standard optimization techniques. Examples of B/W halftone images obtained by this technique are compared to halftones obtained via existing dithering algorithms. Avideh Zakhor, Farokh H. Eskafi |
IEEE Trans. Image Process. | 1 |
| 1992 | Reconstruction of oversampled band-limited signals from Sigma Delta encoded binary sequencesabstractThe authors consider the application of Sigma Delta modulators to analog-to-digital conversion. They have previously shown that for constant input signals, optimal nonlinear decoding can achieve large gains in signal-to-noise ratio (SNR) over linear decoding. A similar result is shown for band-limited input signals. The new nonlinear decoding algorithm is based on projections onto convex sets (POCS) and alternates between a band limitation and a time-domain operation to find a signal invariant under both. The band limitations can be based on singular value decomposition of a certain matrix. The time-domain operation results in a quadratic programming problem. Simulation results are shown for the SNR performance of a POCS-based decoder and a linear decoder for the single-loop Sigma Delta modulator and for a specific fourth-order interpolative modulator. Improvements in SNR of up to 20 dB can be achieved for the single-loop modulator depending on the oversampling ratio, and up to 10 dB for the fourth-order modulator.> Søren Hein, Avideh Zakhor |
ICASSP | 2 |
| 1992 | Inverse problem for two-dimensional fractal sets using the wavelet transform and the moment methodabstractFractal geometry has provided statistical and deterministic models for classes of signals and images that represent many natural phenomena and objects. Iterated function systems (IFS) theory provides a convenient way to describe and classify deterministic fractals in the form of a recursive definition. As a result, it is conceivable to develop image representation schemes based on the IFS parameters that correspond to a given fractal image. The authors propose the use of the wavelet transform and of the moment method for the solution of the inverse problem of recovering IFS parameters from fractal images. The redundancy of a fractal with respect to scale variation is mirrored by its wavelet decomposition, thus providing a method to estimate the scaling parameters for a class of IFSs modeling the image. Displacement parameters and probabilities are then found using the moment method. Experimental results verifying the approach are presented.> Roberto Rinaldo, Avideh Zakhor |
ICASSP | 2 |
| 1992 | A multi-start algorithm for signal adaptive subband systems (image coding)abstractA method for optimizing the filter coefficients of signal adaptive subband systems, in which the polyphase transfer matrix is paraunitary, using a mean-squared-error criterion is proposed. An efficient algorithm was developed for locating numerous local optima in the coefficient space, permitting a degree of confidence in the location of globally optimal or near-optimal solutions. Both separable systems and a particular class of nonseparable filter systems are studied. The application of the algorithm to a number of images is described.> David S. Taubman, Avideh Zakhor |
ICASSP | 2 |
| 1992 | New properties of sigma-delta modulators with DC inputsabstractNew properties of single- and double-loop sigma-delta modulators with constant inputs are derived by exploiting the inherent structure os the output sequences or codewords that the modulators are capable of producing. Upper bounds are derived on the number of N-bit codewords for the single- and double-loop modulators. Analytical lower bounds on the mean squared error (MSE) obtainable by any decoder, linear or nonlinear, in approximating the constant input are also derived. Optimal nonlinear decoders for constant inputs based on a table lookup approach which operates directly on the nonuniform quantization intervals are considered. Using simulations is is found that the optimal nonlinear decoders perform better than linear decoders, by about 3 and 20 dB for the single- and double-loop modulators, respectively. A cascade structure specifically for constant inputs is introduced, and its corresponding decoding algorithm is derived. It is shown that for a fixed latency, the MSE performance of the cascade structure is 12 dB superior, and its throughput is twice that of the conventional two-stage MASH modulator.> Søren Hein, Khalid Ibraham, Avideh Zakhor |
IEEE Trans. Commun. | 3 |
| 1992 | Neural net-based continuous phase modulation receiversabstractThe authors propose feedforward neural networks (NNs) as receivers for partial-response continuous-phase-modulation (CPM) systems. Their approach is to replace the entire receiver structure, excluding timing recovery, with a neural net unit whose inputs are time samples of the incoming baseband signals, and whose outputs are the decoded symbols. Simulation results for coherent and incoherent NN-based receivers are presented, and their performance is compared with that of the optimum maximum-likelihood receiver. The performance of NN-based receivers at large SNR is analyzed.> Gustavo de Veciana, Avideh Zakhor |
IEEE Trans. Commun. | 2 |
| 1992 | Iterative procedures for reduction of blocking effects in transform image codingabstractThe authors propose an iterative block reduction technique based on the theory of a projection onto convex sets. The idea is to impose a number of constraints on the coded image in such a way as to restore it to its original artifact-free form. One such constraint can be derived by exploiting the fact that the transform-coded image suffering from blocking effects contains high-frequency vertical and horizontal artifacts corresponding to vertical and horizontal discontinuities across boundaries of neighboring blocks. Another constraint has to be with the quantization intervals of the transform coefficients. Specifically, the decision levels associated with transform coefficient quantizers can be used as lower and upper bounds on transform coefficients, which in turn define boundaries of the convex set for projection. A few examples of the proposed approach are presented.> Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1991 | A new class of B/W and color halftoning algorithmsabstractA new class of dithering algorithms for black and white (B/W) and color images is presented. The basic idea behind the technique is to divide the image into small blocks and minimize the distortion between the original continuous tone image and its low-pass filtered halftone. This corresponds to a quadratic programming problem with linear constraints which is solved via the branch and bound algorithm. Examples of B/W and color dither images using the technique are shown and compared to halftones obtained via existing dithering algorithms.> Avideh Zakhor, Farokh H. Eskafi |
ICASSP | 1 |
| 1990 | Optimal binary image design based on the branch and bound algorithmabstractMathematical programming techniques for systematic determination of optimal binary masks for precompensation in incoherent optical systems are developed. The feasibility of applying combinatorial optimization techniques to binary mask design for optical lithography is demonstrated. The mask is optimized in such a way as to precompensate the distortions due to optical diffraction of the system. The problem is formulated as a binary linear programming problem and solved with the branch and bound algorithm. Two different formulations corresponding to two optimization criteria are proposed and evaluated. Variation of the optimal mask as a function of the optical system bandwidth is discussed. One- and two-dimensional examples involving bars and a corner are presented.> Avideh Zakhor |
ICASSP | 2 |
| 1990 | Stability of reconstruction of real and complex multidimensional signals from Fourier transform magnitudeabstractThe lower bound on the condition number of the problem of reconstruction for two classes of signals is derived. The lower bound for N*N images in which the norm of the Fourier transform magnitude (FTM) vector is dominated by DC and low-frequency components is shown to grow with N. The lower bound for images in which the elements of the FTM vector contribute more or less equally to its norm is 1. The problem of reconstruction from FTM and known phase in space domain is discussed. Since the introduction of sufficiently random space-domain phase results in flattened FTM distribution, the lower bound on the condition number of the reconstruction problem with respect to the FTM vector again becomes one. This is in agreement with experimental results, indicating that randomizing the phase in space domain improves the convergence rate of the Grechberg-Saxton algorithm.> Avideh Zakhor |
ICASSP | 1 |
| 1990 | Reconstruction of two-dimensional signals from level crossingsabstractRecent results indicate that reconstruction of two-dimensional signals from crossings of one level requires, in theory and practice, extreme accuracy in positions of the samples. The representation of signals with one-level crossings can be viewed as a tradeoff between bandwidth and dynamic range, in the sense that if the available bandwidth is sufficient to preserve the level crossings accurately, then the dynamic range requirements are significantly reduced. On the other hand, representation of signals by their samples at the Nyquist rate can be considered as requiring relatively small bandwidth and large dynamic range, because, at least in theory, amplitude information at prespecified points is needed, to infinite precision. An overview of existing results in zero crossing representation is presented, and a number of new results on sampling schemes for reconstruction from multiple-level threshold crossing are developed. The quantization characteristics of these sampling schemes appear to lie between those of Nyquist sampling and one-level crossing representations, thus bridging the gap between explicit Nyquist sampling and implicit one-level crossing sampling strategies.> Avideh Zakhor, Alan V. Oppenheim |
Proc. IEEE | 1 |
| 1988 | Sampling schemes for reconstruction of multidimensional signals from multiple level threshold crossingsabstractTwo intermediate sampling schemes are developed which bridge the gap between Nyquist sampling and one-level crossing representation by enabling signals to be recovered from multiple-level threshold crossings. To this end, semi-implicit and implicit sampling strategies and their corresponding reconstruction algorithms are derived. A preliminary investigation of the quantization characteristics of some of the proposed sampling and reconstruction schemes, is included.> Avideh Zakhor, Alan V. Oppenheim |
ICASSP | 1 |