VLDB 2026 Research / reviewers in the wild / expert
Anup Basu
dblp:48/1283
· DBLP profile ↗
166ranked-venue papers
20as first author
32since 2021 · last 2026
0000-0002-7695-4148ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 83 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 71 · 14 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 6 since 2021Human-computer interaction and ubiquitous computing · 19 · 3 first-authorSystems, architecture and hardware · 15 · 5 first-authorComputer networks · 5 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StyleFM: Frequency Manipulation Empowered by Recursive Attention on Diffusion Models for Arbitrary Style TransferabstractGiven the remarkable performance of diffusion models in image generation, recent research has been exploring their adaptation to style transfer. However, current diffusion-based approaches encounter persistent challenges, such as style distortions and the reliance on textual prompts for content preservation. To address these limitations, we introduce StyleFM, a novel training-free diffusion-based style transfer approach that incorporates optimization strategies into both the frequency and temporal domains. The proposed method provides two core innovations: (1) Tripartite Frequency Manipulation: To more precisely tailor frequency manipulation, StyleFM introduces a tripartite frequency design with a buffer band accounting for the overlap of content and style representations. In addition, StyleFM designs a frequency superposition editing method to achieve frequency enhancement. (2) Recursive Attention: StyleFM proposes the recursive attention strategy within the diffusion process, which facilitates the progressive and consistent injection of style information throughout the temporal process without reliance on text guidance. Experiments demonstrate that StyleFM outperforms state-of-the-art methods. It effectively preserves content fidelity while achieving sufficient style embedding. Yingnan Ma, Zhenye Liu, Anup Basu |
AAAI | 4 |
| 2026 | SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion Diarization
Farzaneh Jafari, Stefano Berretti, Anup Basu |
ICPR (12) | 3 |
| 2026 | Local Autoregression with Finite-Support Random Variables for Image Generation
Chenqiu Zhao, Anup Basu |
ICPR (3) | 2 |
| 2026 | Dual-branch non-negative matrix factorization guided by information decoupling for multi-view clustering
Mingxia Gong, Chengcai Leng, Zhao Pei, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Neurocomputing | 7 |
| 2025 | Block information strategy for multi-modal remote sensing image registration
Yameng Hong, Chengcai Leng, Beihua Liu, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Dual graph-regularized low-rank representation for hyperspectral image denoising
Chengcai Leng, Mingpei Tang, Zhao Pei, Jinye Peng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Orthogonal Diversity Nonnegative Matrix Factorization for multi-view clustering
Xinling Zhang, Chengcai Leng, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | MFEL-YOLO for small object detection in UAV aerial images
Ting Hou, Chengcai Leng, Zhao Pei, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Expert Syst. Appl. | 7 |
| 2025 | Multi-view data representation via adaptive label propagation nonnegative matrix factorization
Chengcai Leng, Jinye Peng 0001, Zhao Pei, Anup Basu |
Inf. Sci. | 5 |
| 2025 | Scale- and Shape-Aware Network With Prediction Decoupling for Building Fine-Grained Change DetectionabstractBuilding change detection (BCD) is a hot topic in geoscience and remote sensing (RS) with widespread applications. However, most existing BCD methods only focus on areas where changes have occurred, but ignore the change statuses. To address this problem, a building fine-grained change detection (BFCD) task is further explored in this work, which aims to judge the time-related “disappeared”, “appeared”, and “rebuilt” change types of buildings. Meanwhile, a scale- and shape-aware network (S2Net) with prediction decoupling is designed. Firstly, a prediction decoupling framework with dual decoders is built to ensure the prediction consistency with the temporal order of bi-temporal images. Secondly, considering the rebuilt type is the changes between building instances, which are often reflected in the scale and shape differences of the buildings. Thereby, a scale-aware module (ScAM) and a shape-aware module (ShAM) are designed. These two modules help extract the discriminative features of buildings with different scales and shapes for subsequent change detection (CD). In addition, two BCD datasets widely used, LEVIR-CD+ and WHU-CD, are relabeled in this work to support the study of BFCD. Experimental results show that S2Net achieves competitive performance, and its effectiveness is confirmed. The code and datasets will be publicly available at https://github.com/ptdoge/S2Net. Chengcai Leng, Xi Li 0001, Irene Cheng 0001, Anup Basu, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | LMoW: A Latent Random Variable Model for Unconditional Human Motion Generation
Justin Rozeboom, Hanran Song, Chenqiu Zhao, Anup Basu |
MMAsia | 5 |
| 2024 | Accelerating Inference of Networks in the Frequency Domain
Chenqiu Zhao, Guanfang Dong, Anup Basu |
MMAsia | 3 |
| 2024 | Incremental semi-supervised graph learning NMF with block-diagonal
Xue Lv, Chengcai Leng, Jinye Peng 0001, Zhao Pei, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | Feature matching based on Gaussian kernel convolution and minimum relative motion
Chengcai Leng, Huaiping Yan, Jinye Peng 0001, Zhao Pei, Anup Basu |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | Bayesian non-negative matrix factorization with Student's t-distribution for outlier removal and data clusteringabstractNon-negative Matrix Factorization (NMF) is an effective way to solve the redundancy of non-negative high-dimensional data. Most of the traditional probability-based NMF methods use Gaussian distribution to model the differences between the matrices before and after decomposition. However, the Gaussian distribution is strongly affected by outliers, and it may not fit all datasets accurately when there are no outliers in the data. In this article, we propose a novel Bayesian NMF with the Student’s t-distribution, i.e., TNMF. specifically, in order to reduce the impact of outliers on the algorithm, we use the Student’s t-distribution to fit the data points instead of the Gaussian distribution. In addition, it is possible to adjust the Degree of Freedom (DF) to make the Student’s t-distribution more flexible than the Gaussian distribution to fit data points when there are no outliers. Next, we combine the Automatic Relevance Determination (ARD) prior in our algorithm to simplify the model and allow for better performance of the algorithm. Finally, the article used 10 datasets to design two kinds of experiments, outlier removal and data clustering. The outlier removal results of this proposed algorithm are significantly better than the other methods, and it performs better in clustering compared to the other methods in the majority of cases. Ruixue Yuan, Chengcai Leng, Jinye Peng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Learning Temporal Distribution and Spatial Correlation Toward Universal Moving Object SegmentationabstractThe goal of moving object segmentation is separating moving objects from stationary backgrounds in videos. One major challenge in this problem is how to develop a universal model for videos from various natural scenes since previous methods are often effective only in specific scenes. In this paper, we propose a method called Learning Temporal Distribution and Spatial Correlation (LTS) that has the potential to be a general solution for universal moving object segmentation. In the proposed approach, the distribution from temporal pixels is first learned by our Defect Iterative Distribution Learning (DIDL) network for a scene-independent segmentation. Notably, the DIDL network incorporates the use of an improved product distribution layer that we have newly derived. Then, the Stochastic Bayesian Refinement (SBR) Network, which learns the spatial correlation, is proposed to improve the binary mask generated by the DIDL network. Benefiting from the scene independence of the temporal distribution and the accuracy improvement resulting from the spatial correlation, the proposed approach performs well for almost all videos from diverse and complex natural scenes with fixed parameters. Comprehensive experiments on standard datasets including LASIESTA, CDNet2014, BMC, SBMI2015 and 128 real world videos demonstrate the superiority of proposed approach compared to state-of-the-art methods with or without the use of deep learning networks. To the best of our knowledge, this work has high potential to be a general solution for moving object segmentation in real world environments. The code and real-world videos can be found on GitHub https://github.com/guanfangdong/LTS-UniverisalMOS. Guanfang Dong, Chenqiu Zhao, Xichen Pan, Anup Basu |
IEEE Trans. Image Process. | 4 |
| 2024 | Dual-Graph Global and Local Concept Factorization for Data ClusteringabstractConsidering a wide range of applications of nonnegative matrix factorization (NMF), many NMF and their variants have been developed. Since previous NMF methods cannot fully describe complex inner global and local manifold structures of the data space and extract complex structural information, we propose a novel NMF method called dual-graph global and local concept factorization (DGLCF). To properly describe the inner manifold structure, DGLCF introduces the global and local structures of the data manifold and the geometric structure of the feature manifold into CF. The global manifold structure makes the model more discriminative, while the two local regularization terms simultaneously preserve the inherent geometry of data and features. Finally, we analyze convergence and the iterative update rules of DGLCF. We illustrate clustering performance by comparing it with latest algorithms on four real-world datasets. Chengcai Leng, Irene Cheng 0001, Anup Basu, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | RAST: Restorable Arbitrary Style TransferabstractThe objective of arbitrary style transfer is to apply a given artistic or photo-realistic style to a target image. Although current methods have shown some success in transferring style, arbitrary style transfer still has several issues, including content leakage. Embedding an artistic style can result in unintended changes to the image content. This article proposes an iterative framework called Restorable Arbitrary Style Transfer (RAST) to effectively ensure content preservation and mitigate potential alterations to the content information. RAST can transmit both content and style information through multi-restorations and balance the content-style tradeoff in stylized images using the image restoration accuracy. To ensure RAST’s effectiveness, we introduce two novel loss functions: multi-restoration loss and style difference loss. We also propose a new quantitative evaluation method to assess content preservation and style embedding performance. Experimental results show that RAST outperforms state-of-the-art methods in generating stylized images that preserve content and embed style accurately. Yingnan Ma, Chenqiu Zhao, Bingran Huang, Anup Basu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Principal Component Approximation Network for Image CompressionabstractIn this work, we propose a novel principal component approximation network (PCANet) for image compression. The proposed network is based on the assumption that a set of images can be decomposed into several shared feature matrices, and an image can be reconstructed by the weighted sum of these matrices. The proposed PCANet is specifically devised to learn and approximate these feature matrices and weight vectors, which are used to encode images for compression. Unlike previous deep learning-based methods, a distinctive aspect of our approach is its consideration of network size in the bit-rate computation. Despite this inclusion, our proposed method yields promising results. Through extensive experiments conducted on standard datasets, we demonstrate the effectiveness of our approach in comparison to state-of-the-art techniques. To the best of our knowledge, this is the first machine learning approach that includes the size of networks during bitrate computation in image compression. Chenqiu Zhao, Anup Basu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | RAST: Restorable Arbitrary Style Transfer via Multi-restorationabstractArbitrary style transfer aims to reproduce the target image with the artistic or photo-realistic styles provided. Even though existing approaches can successfully transfer style information, arbitrary style transfer still faces many challenges, such as the content leak issue. Specifically, the embedding of artistic style can lead to content changes. In this paper, we solve the content leak problem from the perspective of image restoration. In particular, an iterative architecture is proposed to achieve the Restorable Arbitrary Style Transfer (RAST), which can realize transmission of both content and style information through multi-restorations. We control the content-style balance in stylized images by the accuracy of image restoration. In order to ensure effectiveness of the proposed RAST architecture, we design two novel loss functions: multi-restoration loss and style difference loss. In addition, we propose a new quantitative evaluation method to measure content preservation performance and style embedding performance. Comprehensive experiments comparing with state-of-the-art methods demonstrate that our proposed architecture can produce stylized images with superior performance on content preservation and style embedding. Yingnan Ma, Chenqiu Zhao, Anup Basu |
WACV | 3 |
| 2023 | β-divergence NMF with biorthogonal regularization for data representation
Ruixue Yuan, Chengcai Leng, Bing Li 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Robust dual-graph discriminative NMF for data classification
Chengcai Leng, Bing Li 0001, Licheng Jiao, Anup Basu |
Knowl. Based Syst. | 5 |
| 2023 | Hybrid Conv-ViT Network for Hyperspectral Image ClassificationabstractWith the success of ViT (Vision Transformer), Transformer is being increasingly used for hyperspectral image (HSI) classification given its ability to extract global context dependencies. However, existing methods based on transformers tend to classify HSI in the traditional patch-wise manner. Thus, these methods cannot obtain true global features because the inputs of the model are local patches. To solve these problems, a hybrid convolution and ViT network (HCVN) is proposed for HSI classification. HCVN realizes the classification task from the perspective of semantic segmentation, and its input is the entire HSI, which makes it possible to obtain truly meaningful global features. By improving the original ViT, an HCV module is proposed, which enhances the ability of local structure characterization while extracting global features. The HCVN hybrid convolution layer and HCV module realize the extraction and fusion of local and global features. Finally, the dual branch network architecture is used to integrate the spatial and spectral features. Extensive experiments on two datasets verify the effectiveness of the proposed method. Huaiping Yan, Erlei Zhang, Jun Wang 0078, Chengcai Leng, Anup Basu, Jinye Peng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Cosine Multilinear Principal Component Analysis for RecognitionabstractExisting two-dimensional principal component analysis methods can only handle second-order tensors (i.e., matrices). However, with the advancement of technology, tensors of order three and higher are gradually increasing. This brings new challenges to dimensionality reduction. Thus, a multilinear method called MPCA was proposed. Although MPCA can be applied to all tensors, using the square of the F-norm makes it very sensitive to outliers. Several two-dimensional methods, such as Angle 2DPCA, have good robustness but cannot be applied to all tensors. We extend the robust Angle 2DPCA method to a multilinear method and propose Cosine Multilinear Principal Component Analysis (CosMPCA) for tensor representation. Our CosMPCA method considers the relationship between the reconstruction error and projection scatter and selects the cosine metric. In addition, our method naturally uses the F-norm to reduce the impact of outliers. We introduce an iterative algorithm to solve CosMPCA. We provide detailed theoretical analysis in both the proposed method and the analysis of the algorithm. Experiments show that our method is robust to outliers and is suitable for tensors of any order. Chengcai Leng, Bing Li 0001, Anup Basu, Licheng Jiao |
IEEE Trans. Big Data | 4 |
| 2022 | Multi-step implicit Adams predictor-corrector network for fire detectionabstractAbstract Fire detection methods based on the Convolutional Neural Networks (CNN) have advantages of high accuracy, wide coverage and robustness, receiving significant attention from researchers. Among CNN‐based methods, ResNet has achieved better performance than other CNN frameworks in fire detection system, since it uses stacked residual blocks to enlarge the receptive field to overcome the vanishing gradient problem with residual learning. The merits of ResNet can be attributed to the similarity between ResNet and the single‐step explicit solver for Ordinary Differential Equations (ODEs), for example, the Euler method. Motivated by the theory of numerical ODE that a multi‐step implicit solver has higher accuracy than a single‐step explicit solver, the Multi‐step Implicit Adams predictor‐corrector (MIAPC) network for fire detection is proposed. The MIAPC method is first mapped to a corresponding predictor‐corrector Adams block which achieves higher accuracy than a single‐step explicit solver. Then, Adaptive Feature Fusion (AFF) and the Spatial Attention Layer (SAL) are utilized to extract hierarchical features from stacked predictor‐corrector Adams blocks, forming the corresponding Adams module. Finally, the 4 Adams modules which are made of 4, 6, 8, 10 predictor‐corrector Adams blocks and followed by AFF and SAL form the crucial ODE‐based approximation part in the proposed network. By adding a simple feature extraction and detection in front of and after the ODE‐based approximation part, the MIAPC network is built. Experiments demonstrate that the method achieves 87% accuracy in the challenging test dataset, outperforming existing methods by at least 6%. Besides, the 5.3M model size with inference speed of 4.7 frames/second in CPU and 65.7 frames/second in GPU enables the proposed method to be used in practical applications. Zhen Deng, Shuhao Hu, Shibai Yin, Yibin Wang 0001, Anup Basu, Irene Cheng 0001 |
IET Image Process. | 5 |
| 2022 | Max-Index Based Local Self-Similarity Descriptor for Robust Multi-Modal Image RegistrationabstractIn order to address problems, such as radiation and intensity differences in multi-modal images, this letter proposes a novel idea that integrates maximal indices into the construction of a local self-similarity (LSS) descriptor. The LSS vectors at the same angles but different radial intervals are added to construct the max-index similarity map (MISM) and form the proposed descriptor. This novel descriptor is named max-index-based local self-similarity (MLSS). The MLSS descriptor not only captures the shape similarity between images but is also robust to radiation distortions. Furthermore, a fast and robust algorithm is introduced based on the MLSS descriptor. Comprehensive analysis of accuracy, precision, and computational efficiency shows that the proposed method outperforms five other state-of-the-art methods with stable and better performance on nine pairs of multi-modal test images. Yameng Hong, Chengcai Leng, Xinyue Zhang 0013, Jinye Peng 0001, Licheng Jiao, Anup Basu |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Find Small Objects in UAV Images by Feature Mining and AttentionabstractWith the increasing popularity of Unmanned Aerial Vehicles (UAVs), the accuracy of detecting small objects in large-view images is also expected to increase. However, accurate small object detection is still a challenging problem. Currently, Image Pyramid Network, Feature Pyramid Network (FPN), rich training strategies and data augmentation are widely used to address this problem. To accurately detect small objects, the most important thing is to mine for more feature information. We propose Widened Residual Block (WRB) to break through the bottleneck of residual information gain to extract more feature information. The second is to emphasize or suppress features to prevent small objects from being overwhelmed by a broad background. We introduce an attention mechanism into PANet and propose Enhanced Attention PANet (EA-PANet), which consists of two parts: Context Attention Module (COAM) and Attention Enhancement Module (AEM). COAM outputs attention heatmaps with context, and AEM fuses features from the channel attention module (CAM) and COAM to avoid distraction from a vast background. In addition, we design a lightweight Decoupled Attention Head (DA-head) to dynamically compute important regions for specific tasks and achieve reliable predictions. Experiments show that our method outperforms state-of-the-art (SOTA) detectors. The source code for this work is available at https://github.com/liuxiaolei111/FindSmallObjects. Chengcai Leng, Xiaoming Niu, Zhao Pei, Irene Cheng 0001, Anup Basu |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Remote Sensing Image Registration Based on Local Affine Constraint With Circle DescriptorabstractMany methods have been developed to improve the performance of image registration. In this letter, we introduce a novel method based on a local affine constraint for remote sensing image registration, which can be widely used in image processing and pattern recognition. Our algorithm has three components. First, we exploit the scale invariant feature transform (SIFT) method to extract feature points and calculate the gradient magnitude to establish feature descriptors with a circular instead of square neighborhood. Second, an initial matching is implemented by the nearest neighbor distance ratio (NNDR) and the fast sample consensus (FSC) algorithm. Finally, fine registration is established using more correct matches obtained by the local affine transformation circular region search algorithm. Experimental results show that the proposed method achieves subpixel accuracy. In addition, both the correct matching rate and registration demonstrate the effectiveness and efficiency of our method. Chengcai Leng, Guo-Rong Cai, Zhao Pei, Naigong Yu, Anup Basu |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2022 | Universal Background Subtraction Based on Arithmetic Distribution Neural NetworkabstractWe propose a universal background subtraction framework based on the Arithmetic Distribution Neural Network (ADNN) for learning the distributions of temporal pixels. In our ADNN model, the arithmetic distribution operations are utilized to introduce the arithmetic distribution layers, including the product distribution layer and the sum distribution layer. Furthermore, in order to improve the accuracy of the proposed approach, an improved Bayesian refinement model based on neighboring information, with a GPU implementation, is incorporated. In the forward pass and backpropagation of the proposed arithmetic distribution layers, histograms are considered as probability density functions rather than matrices. Thus, the proposed approach is able to utilize the probability information of the histogram and achieve promising results with a very simple architecture compared to traditional convolutional neural networks. Evaluations using standard benchmarks demonstrate the superiority of the proposed approach compared to state-of-the-art traditional and deep learning methods. To the best of our knowledge, this is the first method to propose network layers based on arithmetic distribution operations for learning distributions during background subtraction. Chenqiu Zhao, Kangkang Hu, Anup Basu |
IEEE Trans. Image Process. | 3 |
| 2021 | A multi-scale attentive recurrent network for image dehazing
Yibin Wang 0001, Shibai Yin, Anup Basu |
Multim. Tools Appl. | 3 |
| 2021 | Total Variation Constrained Graph-Regularized Convex Non-Negative Matrix Factorization for Data RepresentationabstractWe propose a novel NMF algorithm, named Total Variation constrained Graph-regularized Convex Non-negative Matrix Factorization (TV-GCNMF), to incorporate total variation and graph Laplacian with convex NMF. In this model, the feature details of the data are preserved by a diffusion coefficient based on the gradient information. The graph regularization and convex constraints reveal the intrinsic geometry and structure information of the features; thereby, obtaining sparse and parts-based representations. Furthermore, we give the multiplicative update rules and prove convergence of the proposed algorithm. The results of clustering experiments on multiple datasets, under various noise conditions, show the effectiveness and robustness of the proposed method compared to state-of-the-art clustering methods and other related work. Chengcai Leng, Anup Basu |
IEEE Signal Process. Lett. | 4 |
| 2021 | Deep Variation Transformation Network for Foreground DetectionabstractIn existing literature, the distribution of pixel observations is analyzed with models designed for the video foreground detection task. However, it is possible that the background and foreground share similar observations, causing false detections. We propose a novel foreground detection method called Deep Variation Transformation Network (DVTN), focusing on analyzing the pixel variations instead of distributions. In particular, pixel variations are represented by a sequence of pixel observations, and DVTN is trained to transform the pixel variations into a new space, where the observations can be classified easily. Following this, the output of DVTN is utilized by a linear classifier to label pixels as foreground or background. As a result of the global analysis and the strong learning ability of DVTN, the proposed approach adaptively learns a good transformation from pixel variations to probabilities of labels to improve performance. Comprehensive experiments on several benchmark datasets demonstrate the superiority of our DVTN approach compared to both state-of-the-art deep learning and traditional methods, especially in scenes lacking texture and color information. Code is available at https://github.com/Zhangjunyin/DVTN. Yongxin Ge, Junyin Zhang, Xinyu Ren, Chenqiu Zhao, Anup Basu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Edge-guided CNN for Denoising Images from Portable Ultrasound DevicesabstractUltrasound is a non-invasive tool that is useful for medical diagnosis and treatment. To reduce long wait times and add convenience to patients, portable ultrasound scanning devices are becoming increasingly popular. These devices can be held in one hand, and are compatible with modern cell phones. However, the quality of ultrasound images captured from the portable scanners is relatively poor compared to standard ultrasound scanning systems in hospitals. To improve the quality of the ultrasound images obtained from portable ultrasound devices, we propose a new neural network architecture called Edge-guided Denoising Convolutional Neural Network (EDCNN), which can preserve significant edge information in ultrasound images when removing noise. We also study and compare the effectiveness of existing deep learning methods and classical filtering approaches in removing speckle noise in these images. Experimental results show that after applying the proposed EDCNN, various organs can be better recognized from ultrasound images. This approach is expected to lead to better accuracy in diagnostics in the future. Yingnan Ma, Anup Basu |
ICPR | 3 |
| 2020 | Multi-Scale Deep Pixel Distribution Learning for Concrete Crack DetectionabstractA number of methods including image processing technologies (IPTs) and deep learning methods, have been used to detect defects in civilian infrastructure. These methods have been introduced to extract features representing cracks in concrete surfaces. Inspired by recent advances of a pixel distribution learning method in background subtraction, we propose a novel multi-scale deep learning method (MS-DPDL) for concrete crack detection. The designed CNN network is trained on the dataset CRACK500 [1], [2] and tested on it for concrete segmentation. To show good transferability of our proposed model, it is later tested on the dataset Concrete Crack Images for Classification [3]. Several existing deep learning methods are used to compare the performance of the proposed MS-DPDL method. Results show that our method has good performance and can effectively find concrete cracks in practical situations. Xuanyi Wu, Jianfei Ma, Chenqiu Zhao, Anup Basu |
ICPR | 5 |
| 2020 | Image dehazing with uneven illumination prior by dense residual channel attention networkabstractExisting dehazing methods based on convolutional neural networks estimate the transmission map by treating channel‐wise features equally, which lacks flexibility in handling different types of haze information, leading to the poor representational ability of the network. Besides, the scene lights are predicted by an even illumination prior which does not work for a real situation. To solve these problems, the authors propose a dense residual channel attention network (DRCAN) for estimating the transmission map and use an image segmentation strategy to predict scene lights. Specifically, DRCAN is built based on the proposed dense residual block (DRB) and dense residual channel attention block (DRCAB). DRB extracts the hierarchical features with increasing receptive fields. DRCAB makes the network focus on the features containing heavy haze information. After the transmission map is estimated, fuzzy partition entropy combined with graph cuts is used to segment the transmission map into scene regions covered with varying scene lights. This strategy not only considers the fuzzy intensities of the low‐contrast transmission map but also takes spatial correlation into account. Finally, a clear image is obtained by the transmission map and varying scene lights. Extensive experiments demonstrate that our method is comparable to most of existing methods. Shibai Yin, Jin Xin, Yibin Wang 0001, Anup Basu |
IET Image Process. | 4 |
| 2020 | Dynamic Deep Pixel Distribution Learning for Background SubtractionabstractPrevious approaches to background subtraction usually approximate the distribution of pixels with artificial models. In this paper, we focus on automatically learning the distribution, using a novel background subtraction model named Dynamic Deep Pixel Distribution Learning (D-DPDL). In our D-DPDL model, a distribution descriptor named Random Permutation of Temporal Pixels (RPoTP) is dynamically generated as the input to a convolutional neural network for learning the statistical distribution, and a Bayesian refinement model is tailored to handle the random noise introduced by the random permutation. Because the temporal pixels are randomly permutated to guarantee that only statistical information is retained in RPoTP features, the network is forced to learn the pixel distribution. Moreover, since the noise is random, the Bayesian theorem is naturally selected to propose an empirical model as a compensation based on the similarity between pixels. Evaluations using standard benchmark demonstrates the superiority of the proposed approach compared with the state-of-the-art, including traditional methods as well as deep learning methods. Chenqiu Zhao, Anup Basu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Simplified Active Calibration
Mehdi Faraji, Anup Basu |
Image Vis. Comput. | 2 |
| 2018 | Adaptive Resolution Optimization and Tracklet Reliability Assessment for Efficient Multi-Object TrackingabstractRecent digital acquisition systems can acquire high-resolution videos, generating a large amount of dynamic data and leading to higher computational cost in online target tracking and learning, especially for complex scenes. We introduce an efficient and robust approach to improve the performance of multi-object online tracking and learning. Prior methods saved on computational cost by scaling down each video frame to a fixed smaller resolution, without considering the image features. Our algorithm computes the optimal image resolution adaptively by exploiting the correlation between an image's gray-value distribution and resolution. This dimensionality reduction step significantly improves the time performance in subsequent online tracking and learning, while preserving high tracking accuracy. Since a small detection error in one frame can cause cumulative error in the video sequence leading to incorrect labeling and tracking, we introduce a new tracklet reliability assessment metric to eliminate incorrect samples. Experimental results show that our approach can successfully track multiple objects in real time with both high precision and recall. Ruixing Yu, Irene Cheng 0001, Sweta Bedmutha, Anup Basu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Introduction to the Special Issue on Representation, Analysis, and Recognition of 3D HumansabstractNo abstract available. Stefano Berretti, Mohamed Daoudi, Pavan Turaga, Anup Basu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2018 | Representation, Analysis, and Recognition of 3D Humans: A SurveyabstractComputer Vision and Multimedia solutions are now offering an increasing number of applications ready for use by end users in everyday life. Many of these applications are centered for detection, representation, and analysis of face and body. Methods based on 2D images and videos are the most widespread, but there is a recent trend that successfully extends the study to 3D human data as acquired by a new generation of 3D acquisition devices. Based on these premises, in this survey, we provide an overview on the newly designed techniques that exploit 3D human data and also prospect the most promising current and future research directions. In particular, we first propose a taxonomy of the representation methods, distinguishing between spatial and temporal modeling of the data. Then, we focus on the analysis and recognition of 3D humans from 3D static and dynamic data, considering many applications for body and face. Stefano Berretti, Mohamed Daoudi, Pavan Turaga, Anup Basu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Facilitating player progression by implementing procedural music in videogamesabstractWhile the multi-million-dollar videogame industry sees constant improvements in visuals and processing power with every new processor, gaming console, or graphics card release, is there any avenue to pursue new innovations in this discipline? One answer lies in the introduction of procedural music in videogames as an innovative means of overcoming the repetitive nature of traditionally composed game music and to enhance the immersive response of game music to player actions. While most studies are preoccupied with the descriptive evaluation of how entertaining procedural music is in comparison to traditionally composed music, we pursue a novel study on the utility of implementing procedural music in videogames as a tool for facilitating player progression. To do so we employ objective, quantitative measures to gather results that can quantify the utility of implementing procedural videogame music, unlike other studies. We demonstrate that users playing a game with a procedural music model that actively instructs and assists players can complete game levels in a significantly more time efficient manner and are more likely to rate the procedural music model as having significantly contributed to the game's entertainment and engagement. Jayden Chan, Justin John Daza, William Kwan, Anup Basu |
SMC | 4 |
| 2017 | Multimodal interaction in augmented realityabstractWith the boost of computing power in mobile devices and availability of cloud APIs in recent years, mobile augmented reality (AR) applications have become increasingly embedded in people's everyday life. However, effective and intuitive interaction between virtual and real worlds in these AR applications is still an open question. In this paper, we make one step towards an answer by exploring the possibility of incorporating two input modalities, gesture and speech, for enhancing user experience in AR applications. Zhaorui Chen, Jinzhu Li, Yifan Hua, Rui Shen 0002, Anup Basu |
SMC | 5 |
| 2017 | Medical image compression based on region of interest using better portable graphics (BPG)abstractEveryday, an enormous number of medical images are produced by hospitals and medical imaging center for research, surgical and disease diagnostics. Therefore, compression is necessary for storing, managing and transferring these data to make storage manageable. Medical images have some parts which are more important called region of interest (ROI) with useful information for the diagnostic purpose that should be reconstructed with high quality during the image decompression process. In this paper, a state-of-the-art image compression format known as Better Portable Graphics (BPG), which is based on the High Efficiency Video Coding (HEVC), is used for medical image compression. In the proposed compression method, first the medical image is segmented into two parts: ROI and non-ROI regions. In the next step, lossless BPG compression algorithm is applied to the ROI areas, and lossy BPG is utilized for non-ROI regions. In the end, the resulting reconstructed images are combined to create a complete compressed image. The MRI scan dataset hosted by the University of Cyprus is used to evaluate the performance of the proposed compression method to demonstrate improvement between 10-25% in the compression rate compared to traditional image compression techniques used in the medical industry. David Yee, Sara Soltaninejad, Deborsi Hazarika, Gaylord Mbuyi, Rishi Barnwal, Anup Basu |
SMC | 6 |
| 2017 | Subjective and Objective Visual Quality Assessment of Textured 3D MeshesabstractObjective visual quality assessment of 3D models is a fundamental issue in computer graphics. Quality assessment metrics may allow a wide range of processes to be guided and evaluated, such as level of detail creation, compression, filtering, and so on. Most computer graphics assets are composed of geometric surfaces on which several texture images can be mapped to make the rendering more realistic. While some quality assessment metrics exist for geometric surfaces, almost no research has been conducted on the evaluation of texture-mapped 3D models. In this context, we present a new subjective study to evaluate the perceptual quality of textured meshes, based on a paired comparison protocol. We introduce both texture and geometry distortions on a set of 5 reference models to produce a database of 136 distorted models, evaluated using two rendering protocols. Based on analysis of the results, we propose two new metrics for visual quality assessment of textured mesh, as optimized linear combinations of accurate geometry and texture quality measurements. These proposed perceptual metrics outperform their counterparts in terms of correlation with human opinion. The database, along with the associated subjective scores, will be made publicly available online. Jinjiang Guo, Vincent Vidal 0002, Irene Cheng 0001, Anup Basu, Atilla Baskurt, Guillaume Lavoué |
ACM Trans. Appl. Percept. | 4 |
| 2016 | Highlighting objects of interest in an image by integrating saliency and depthabstractStereo images have been captured primarily for 3D reconstruction in the past. However, the depth information acquired from stereo can also be used along with saliency to highlight certain objects in a scene. This approach can be used to make still images more interesting to look at, and highlight objects of interest in the scene. We introduce this novel direction in this paper, and discuss the theoretical framework behind the approach. Even though we use depth from stereo in this work, our approach is applicable to depth data acquired from any sensor modality. Experimental results on both indoor and outdoor scenes demonstrate the benefits of our algorithm. Subhayan Mukherjee, Irene Cheng 0001, Anup Basu |
ICIP | 3 |
| 2016 | Robust Human Animation Skeleton Extraction Using Compatibility and Correctness ConstraintsabstractThe ability to automatically animate arbitrary 3D characters based on motion capture (MoCap) data has many applications in simulation, entertainment and multimedia transmission. However, defining trajectory key-points in human figures for animation without any manual intervention remains a challenging problem that makes complete automation difficult. To animate an articulated 3D character an animation skeleton needs to be extracted from, or be embedded into, a 3D model for deformation during animation. In conventional animation software, this process is mostly done manually by expert animators, which makes it a very tedious and time consuming step. The automatic rigging approaches proposed in the literature require a front facing model with neutral T-pose to accurately extract or embed an animation skeleton. We propose a fully automatic skeleton extraction approach based on optimization of constraints on human shape that can generate the animation skeleton, regardless of the model's orientation and position. Experimental results demonstrate the effectiveness of our approach. Robust skeleton extraction followed by efficient MoCap data compression can greatly improve the fidelity of 3D animated model transmission. Nasim Hajari, Irene Cheng 0001, Anup Basu |
ISM | 3 |
| 2016 | Spatio-Temporally Optimized Multi-sensor Motion FusionabstractThe latest advances in smart sensor technology, e.g., Leap Motion Sensor, has increased the precision in tracking fully articulated human hand and finger movements, without the need for placing electrical or optical markers. A remaining challenge is finger occlusion, which can affect tracking accuracy. In this paper, we introduce a spatio-temporal optimization technique for motion data generated from multiple sensors. We demonstrate that our algorithm can produce a fused stream of probabilistic optimal hand poses, by improving local spatial domain analysis and proposing a fast and effective flow analysis technique in the temporal domain, which computes how well the hand pose estimation in the current frame fits the movement flow within a time segment. By using an artificial hand to represent the hand pose ground truth at selected time steps, experimental results demonstrate that our spatio-temporal optimization algorithm increases the estimation accuracy by 6% compared to the reference method, achieving an overall accuracy of 91.29%. Our proposed method can be used offline or in real-time, and can benefit a wide range of applications, including surgical planning and training, where hand motion is the focus of performance efficiency and assessment. Xinyao Sun, Irene Cheng 0001, Anup Basu |
ISM | 3 |
| 2016 | Optimized per-joint compression of hand motion dataabstractMotion data is quickly expanding its application scope, following the recent advancements in smart sensing technology. In particular, it has been shown to be helpful for objective measurement and assessment of surgical dexterity among users at different levels of training. The goal is to allow trainees to evaluate their performance based on a reference set of hand movements. Similar to other multimedia data types, recording motion can produce a substantial amount of data, some of which are redundant for the application. Compression methods aim to optimize storage and transmission of motion capture (MoCap) data by taking advantage of temporal and spatial correlation. Hand motion data is a special sub-type of MoCap and is the focus of many applications, where hand movement evaluation is important. In this paper, we propose a lossy but visually indifferent, compression method that exploits redundancy found in hand motion data. Since individual joint movements have different impacts on the motion sequence, our technique is designed to minimize the overall distortion by providing a per-joint compression. We are able to demonstrate that our approach offers a quantitative gain for different compression ratios, while preserving visual quality. Antonio Carlos Furtado, Xinyao Sun, Anup Basu, Irene Cheng 0001 |
SMC | 3 |
| 2016 | Robust Lung Segmentation combining adaptive Concave Hulls with Active ContoursabstractLung segmentation is an important first step towards an automated CAD (Computer Aided Detection) system for a variety of medical applications. These applications range from lung nodule detection for identifying cancerous tumors to acinar shadow detection for identifying Tuberculosis. In our prior work we had used the Concave Hull algorithm for lung segmentation. However, our results showed over segmentation. In this work we introduce “Adaptive” concave hulls, combine it with Adaptive Median Filtering, and finally apply an Active Contour Model to make the results much more robust and eliminate the over segmentation and under segmentation problem. Our technique is especially useful for automated detection of Juxtapleural pulmonary nodules that are attached to the chest wall. Experimental results demonstrate the improvements achieved by our new algorithm. Sara Soltaninejad, Irene Cheng 0001, Anup Basu |
SMC | 3 |
| 2016 | A Multisensor Technique for Gesture Recognition Through Intelligent Skeletal Pose AnalysisabstractRecent advances in smart sensor technology and computer vision techniques have made the tracking of unmarked human hand and finger movements possible with high accuracy and at sampling rates of over 120 Hz. However, these new sensors also present challenges for real-time gesture recognition due to the frequent occlusion of fingers by other parts of the hand. We present a novel multisensor technique that improves the pose estimation accuracy during real-time computer vision gesture recognition. A classifier is trained offline, using a premeasured artificial hand, to learn which hand positions and orientations are likely to be associated with higher pose estimation error. During run-time, our algorithm uses the prebuilt classifier to select the best sensor-generated skeletal pose at each time step, which leads to a fused sequence of optimal poses over time. The artificial hand used to establish the ground truth is configured in a number of commonly used hand poses such as pinches and taps. Experimental results demonstrate that this new technique can reduce total pose estimation error by over 30% compared with using a single sensor, while still maintaining real-time performance. Our evaluations also demonstrate that our approach significantly outperforms many other alternative approaches such as weighted averaging of hand poses. An analysis of our classifier performance shows that the offline training time is insignificant, and our configuration achieves about 90.8% optimality for the dataset used. Our method effectively increases the robustness of touchless display interactions, especially in high-occlusion situations by analyzing skeletal poses from multiple views. Nathaniel Rossol, Irene Cheng 0001, Anup Basu |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2015 | Foveated High Efficiency Video Coding for Low Bit Rate TransmissionabstractThis work describes the design and subjective performance of Foveated High Efficiency Video Coding (FHEVC). Even though foveation has been widely used for various forms of compression since the early 1990s, we believe its use to improve HEVC is new. We consider the application of, possibly moving, foveated compression in this work and evaluate scenarios where it can be used to improve perceptual quality of videos under constrained transmission resources, e.g., bandwidth. A new method to reduce artifacts during remapping is also proposed. The preliminary implementation considers a single fovea only. Experiments summarizing user evaluations are presented to validate our implementation. Irene Cheng 0001, Masha Mohammadkhani, Anup Basu, Frédéric Dufaux |
ISM | 3 |
| 2015 | Normalized Gaussian Distance Graph Cuts for Image SegmentationabstractThis paper presents a novel, fast image segmentation method based on normalized Gaussian distance on nodes in conjunction with normalized graph cuts. We review the equivalence between kernel k-means and normalized cuts. Then we extend the framework of efficient spectral clustering and avoid choosing weights in the weighted graph cuts approach. Experiments on synthetic data sets and real-world images demonstrate that the proposed method is effective and accurate. Chengcai Leng, Wei Xu 0009, Irene Cheng 0001, Zhihui Xiong, Anup Basu |
ISM | 5 |
| 2015 | Perceptually motivated LSPIHT for motion capture data compression
Irene Cheng 0001, Amirhossein Firouzmanesh, Anup Basu |
Comput. Graph. | 3 |
| 2015 | Graph Matching Based on Stochastic PerturbationabstractThis paper presents a novel perspective on characterizing the spectral correspondence between the nodes of weighted graphs for image matching applications. The algorithm is based on the principal feature components obtained by stochastic perturbation of a graph. There are three areas of contributions in this paper. First, a stochastic normalized Laplacian matrix of a weighted graph is obtained by perturbing the matrix of a sensed graph model. Second, we obtain the eigenvectors based on an eigen-decomposition approach, where representative elements of each row of this matrix can be considered to be the feature components of a feature point. Third, correct correspondences are determined in a low-dimensional principal feature component space between the graphs. In order to further enhance image matching, we also exploit the random sample consensus algorithm, as a post-processing step, to eliminate mismatches in feature correspondences. The experiments on synthetic and real-world images demonstrate the effectiveness and accuracy of the proposed method. Chengcai Leng, Wei Xu 0009, Irene Cheng 0001, Anup Basu |
IEEE Trans. Image Process. | 4 |
| 2013 | Evaluation of 3D Model Segmentation Techniques Based on Animal Anatomyabstract3D model decomposition is a challenging and important problem in computer graphics. Several semantically based approaches have been proposed in the literature, however, due to the lack of proper evaluation criteria, comparison of these techniques is almost impossible. In this paper we suggest to use animal anatomy as the ground truth and compare the result of different segmentation techniques based on that. Differing from previous approaches which perform the evaluation based on ground truth databases created subjectively by human observers, we consider expert knowledge on anatomy of various animals. Based on this knowledge we specify the ground truth for different animals and compare alternative algorithms. Nasim Hajari, Irene Cheng 0001, Anup Basu, Guillaume Lavoué |
SMC | 3 |
| 2013 | Perceptual Quality Metrics for 3D Meshes: Towards an Optimal Multi-attribute Computational Modelabstract3D graphical data, commonly represented using triangular meshes, are deployed in a wide range of application processes including compression, filtering, watermarking, and simplification. These processes often introduce geometric distortions which affect the visual quality of the ultimate data visualization. In order to accurately evaluate perceptual impacts caused by the distortions, assessment metrics on 3D Mesh Visual Quality (MVQ) have been extensively discussed in the literature. Researchers recommended various metrics to predict the adverse effects that visual artifacts can have in applications. Most of these metrics are based on geometric attributes, conventional geometric distance, Laplacian coordinates, different types of curvature computation, and dihedral angles. We hypothesize that an optimal combination of multiple attributes associated with a 3D mesh surface can contribute to better perceptual prediction than single attributes used separately. In this paper, we use two user studies to validate our hypothesis. Our contributions are: (1) providing a detailed analysis of the most relevant geometric attributes for mesh quality assessment, and (2) introducing a new perceptual evaluation metric based on multiple attributes, with the optimal combination determined through machine learning techniques. Statistical quantitative analysis shows that our metric delivers better results than other state-of-the-art approaches. The proposed method is simple to implement and fast in execution. Moreover, our framework can easily be expanded to accommodate additional surface attributes. Guillaume Lavoué, Irene Cheng 0001, Anup Basu |
SMC | 3 |
| 2013 | Improved Robust Kernel Subspace for Object-Based Registration and Change DetectionabstractPostclassification comparison is an approach to detect changes in remote sensing images that have strongly inhomogeneous scenes. It is a challenging task to register pre- and postevent scenarios, because variform classifications may induce an inadequate number of homologous points to be used as tie points. In this letter, we show how the variform objects can be precisely registered using their robust kernel subspace. There are two primary contributions in our work. First, a robust kernel subspace analysis method is proposed to capture the common patterns of the variform objects. Second, a registration method based on the common patterns and their preimage are derived. The power of the proposed approach is demonstrated by two real applications: one for lake monitoring in the Jiayu region and the other for damage mapping of earthquake-induced barrier lake at Tangjiashan. The results show that the proposed method is effective in structural pattern analysis and object registration. Zheng Tian 0001, Mingtao Ding, Anup Basu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2013 | QoE-Based Multi-Exposure Fusion in Hierarchical Multivariate Gaussian CRFabstractMany state-of-the-art fusion methods, combining details in images taken under different exposures into one well-exposed image, can be found in the literature. However, insufficient study has been conducted to explore how perceptual factors can provide viewers better quality of experience on fused images. We propose two perceptual quality measures: perceived local contrast and color saturation, which are embedded in our novel hierarchical multivariate Gaussian conditional random field model, to illustrate improved performance for multi-exposure fusion. We show that our method generates images with better quality than existing methods for a variety of scenes. Rui Shen 0002, Irene Cheng 0001, Anup Basu |
IEEE Trans. Image Process. | 3 |
| 2012 | A general optimal pixel aspect ratio model for stereo-based 3D reconstruction and visualizationabstractWe propose a mathematical model for optimizing pixel aspect ratio for the best 3D estimation in both stereo-based 3D reconstruction and 3D viewing applications. We analyze the 3D reconstruction through a 3D display medium and reduce the whole process to a single stereo system so that a unified model can be applied to both direct reconstruction from stereo images and indirect reconstruction through stereo content presented on a 3D display. We use this unified model to extend our earlier work to a general solution for determining the optimal pixel aspect ratio for both applications. Unlike earlier work, the solution proposed here relates the optimal pixel aspect ratio to the device-specific parameters, rather than the stereo configuration parameters, which makes it more easily applicable in design and manufacture of stereo capture and display devices. In general, our mathematical model and subjective user studies suggest that, for a given total resolution, a finer horizontal discretization with a ratio of about 0.6 leads to a more accurate 3D reconstruction and a better 3D visual experience. Hossein Azari, Irene Cheng 0001, Anup Basu |
SMC | 3 |
| 2012 | Cross-selection kernel regression for super-resolution fusion of complementary panoramic imagesabstractComplementary catadioptric imaging technique was proposed to solve the problem of low and non-uniform resolution in omnidirectional imaging. To enhance this research, our paper focuses on how to generate a high-resolution panoramic image from the captured omnidirectional image. To avoid the interference between the inner and outer images while fusing the two complementary views, a cross-selection kernel regression method is proposed. First, in view of the complementarity of sampling resolution in the tangential and radial directions between the inner and the outer images respectively, the horizontal gradients in the expected panoramic image are estimated based on the scattered neighboring pixels mapped from the outer, while the vertical gradients are estimated using the inner image. Then, the size and shape of the regression kernel are adaptively steered based on the local gradients. Furthermore, the neighboring pixels in the next interpolation step of kernel regression are also selected based on the comparison between the horizontal and vertical gradients. In simulation and real-image experiments, the proposed method outperforms existing kernel regression methods and our previous wavelet-based fusion method in terms of both visual quality and objective evaluation. Lidong Chen, Anup Basu, Maojun Zhang, Wei Wang 0068 |
SMC | 2 |
| 2012 | Hand and face tracking under occlusion with anthropomorphic constraintsabstractWe propose a graphical model for a decentralized, simultaneous detection and tracking algorithm for efficient localization of hands from a sequence in color and range images. We deduce the location of key-points using a Bayesian framework. We use anthropomorphic constraints for modelling body part articulation. Furthermore, our algorithm reasons about occlusion and preserves data association to deal with ambiguities. Experimental results demonstrate that our system tracks face and hands more accurately in video, compared to prior research. Abhishek Sen, Irene Cheng 0001, Anup Basu |
SMC | 3 |
| 2012 | HW/SW co-design of an embedded omni-imaging systemabstractOmni-imaging can be used in many practical applications that need a wide field of view, therefore a real-time and high-definition embedded system design and implementation of omni-imaging is desired. In this study, we propose a hardware/software co-design method for the design and implementation of embedded omni-imaging systems. In order to achieve real-time and high-definition goals, we perform hardware/software partitioning based on the analysis of functional modules in a basic embedded omni-imaging system. In the experiments, the proposed hardware/software co-design omni-imaging system is implemented in a FPGA (Field Programmable Gate Array) plus DSP (Digital Signal Processor) system architecture. Results indicate that the omni-imaging speed achieved is 39fps with the imaging resolution set at 1024×768 for the original omni-image and 1280×288 for the unwarped image. Zhihui Xiong, Irene Cheng 0001, Maojun Zhang, Anup Basu |
SMC | 4 |
| 2012 | Optimal pixel aspect ratio for enhanced 3D TV visualization
Hossein Azari, Irene Cheng 0001, Kostas Daniilidis, Anup Basu |
Comput. Vis. Image Underst. | 4 |
| 2012 | Perceptually Coded Transmission of Arbitrary 3D Objects over Burst Packet Loss Channels Enhanced with a Generic JND FormulationabstractIn this work we propose a new approach to account for burst packet loss during transmission of 3D objects represented by texture and mesh over unreliable networks. Our strategy includes applying stripification on the 3D mesh following the valence-driven algorithm and distributing nearby vertices into different packets, combined with an interleaving technique that does not need texture or mesh packets to be re-transmitted. The perceptually-driven technique is able to successfully interpolate lost mesh features even under severe packet loss. The reconstructed mesh is further improved by applying our curvature-driven probabilistic strategy to safeguard visually significant structures on the 3D surface. Experimental results show that smoothness on the object surface is preserved even at 50% packet loss. At 75% packet loss, smoothness on the object surface deteriorates but the overall shape of an object is still preserved. We also define a Quality of Experience (QoE) metric to formulate the Just-Noticeable-Difference (JND) concept, to quantify the qualitative findings obtained from earlier subjective user studies, which provides flexibility to applications for reducing the transmission of visually redundant data. Irene Cheng 0001, Lihang Ying, Anup Basu |
IEEE J. Sel. Areas Commun. | 3 |
| 2012 | Choice of low resolution sample sets for efficient super-resolution signal reconstruction
Meghna Singh, Cheng Lu 0001, Anup Basu, Mrinal Mandal 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | Efficient omni-image unwarping using geometric symmetry
Zhihui Xiong, Irene Cheng 0001, Anup Basu, Wei Wang 0068, Wei Xu 0019, Maojun Zhang |
Mach. Vis. Appl. | 3 |
| 2011 | Optimized point splatting based on a fully-balanced hierarchical structureabstractWe propose an improved two-layer hierarchical point-based 3D model representation for interactive 3D rendering. A tree structure called multi-section tree is introduced that allows creating a fully-balanced hierarchy with a desired branching factor for any arbitrary number of points. The proximity problem, the challenge of keeping close-by points together in equi-partitioning process, is addressed by applying an initial bottom-up point-grouping process that divides the 3D model surface into small, nearly equal-size groups of neighboring points. Each group is represented by a small multi-section subtree but treated as a single point in the main top-down equi-splitting process. The top-down process creates the main balanced multi-section tree which holds the whole structure together. Results show that our two-step description leads to a very compact quantized representation with enhanced rendering quality. Hossein Azari, Irene Cheng 0001, Anup Basu |
ICME | 3 |
| 2011 | An in-place texture synthesis technique for memory constrained multimedia applicationsabstractDespite the rapid evolution of multimedia content, from 2D to 3D and to stereo on IMAX display, material and texture remain an indispensable component when rendering realistic and appealing graphics and animations. As the demand for high-definition displays increases, so is the necessity for high-resolution textures. Nevertheless, the available texture images very often have low resolution and are inadequate for high-quality rendering. Texture synthesis from examples offers an effective way not only for the creation of high resolution texture, but also useful in interpolating missing data resulted from unreliable transmission and overly compressed data. We present a new memory-efficient technique that facilitates high-quality texture synthesis. We compare the time performance and memory usage between our approach and the commonly used caching techniques. Experimental results show that our method can perform better, given limited memory resources. Alexey Badalov, Irene Cheng 0001, Anup Basu |
ICME | 4 |
| 2011 | Temporal-spatial face recognition using multi-atlas and Markov process modelabstractAlthough video-based face recognition algorithms can provide more information than image-based algorithms, their performance is affected by subjects' head poses, expressions, illumination and so on. In this paper, we present an effective video-based face recognition algorithm. Multi-atlas is employed to efficiently represent faces of individual persons under various conditions, such as different poses and expressions. The Markov process model is used to propagate the temporal information between adjacent video frames. The combination of multi-atlas and Markov model provides robust face recognition by taking both spatial and temporal information into account. The performance of our algorithm was evaluated on three standard test databases: the Honda/UCSD video database, the CMU Motion of Body database, and the multi-modal VidTIMIT database. Experimental results demonstrate that our video-based face recognition algorithm outperforms other methods on all three test databases. Gaopeng Gou, Rui Shen 0002, Yunhong Wang 0001, Anup Basu |
ICME | 4 |
| 2011 | Anatomy preserving 3D model decomposition based on robust skeleton-surface node correspondenceabstractIn this work, we present an effective anatomy preserving model decomposition technique. By extracting unit-width curve skeletons, which are robust to noise, and mapping skeleton branches to model surface nodes, our method accurately identifies the topology and geometry information of a 3D model, resulting in more semantically rich segmented components. Experiments on 2194 models from the Princeton Shape Benchmark and A Benchmark for 3D Mesh Segmentation demonstrate the advantage of the proposed technique. Our results preserve better anatomical structures compared to three commonly used model segmentation methods. Irene Cheng 0001, Anup Basu |
ICME | 3 |
| 2011 | Efficient video sequences alignment using unbiased bidirectional dynamic time warping
Cheng Lu 0001, Meghna Singh, Irene Cheng 0001, Anup Basu, Mrinal Mandal 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2011 | A 2-point algorithm for 3D reconstruction of horizontal lines from a single omni-directional image
Irene Cheng 0001, Zhihui Xiong, Anup Basu, Maojun Zhang |
Pattern Recognit. Lett. | 4 |
| 2011 | Generalized Random Walks for Fusion of Multi-Exposure ImagesabstractA single captured image of a real-world scene is usually insufficient to reveal all the details due to under- or over-exposed regions. To solve this problem, images of the same scene can be first captured under different exposure settings and then combined into a single image using image fusion techniques. In this paper, we propose a novel probabilistic model-based fusion technique for multi-exposure images. Unlike previous multi-exposure fusion methods, our method aims to achieve an optimal balance between two quality measures, i.e., local contrast and color consistency, while combining the scene details revealed under different exposures. A generalized random walks framework is proposed to calculate a globally optimal solution subject to the two quality measures by formulating the fusion problem as probability estimation. Experiments demonstrate that our algorithm generates high-quality images at low computational cost. Comparisons with a number of other techniques show that our method generates better results in most cases. Rui Shen 0002, Irene Cheng 0001, Jianbo Shi, Anup Basu |
IEEE Trans. Image Process. | 4 |
| 2011 | Perceptually Guided Fast Compression of 3-D Motion Capture DataabstractA time efficient compression technique, incorporating attention stimulating factors, for motion capture data is proposed. Compression ratios of 25:1 to 30:1 can be achieved with very little noticeable degradation in perceptual quality of animation. Experimental analysis shows that the proposed algorithm is much faster than comparable approaches using wavelets, thereby making our approach feasible for motion capture, transmission, and real-time synthesis on mobile devices, where processing power and memory capacity are limited. Amirhossein Firouzmanesh, Irene Cheng 0001, Anup Basu |
IEEE Trans. Multim. | 3 |
| 2010 | Fully automatic brain tumor segmentation using a normalized Gaussian Bayesian Classifier and 3D Fluid Vector FlowabstractBrain tumor segmentation from Magnetic Resonance Images (MRIs) is an important task to measure tumor responses to treatments. However, automatic segmentation is very challenging. This paper presents an automatic brain tumor segmentation method based on a Normalized Gaussian Bayesian classification and a new 3D Fluid Vector Flow (FVF) algorithm. In our method, a Normalized Gaussian Mixture Model (NGMM) is proposed and used to model the healthy brain tissues. Gaussian Bayesian Classifier is exploited to acquire a Gaussian Bayesian Brain Map (GBBM) from the test brain MRIs. GBBM is further processed to initialize the 3D FVF algorithm, which segments the brain tumor. This algorithm has two major contributions. First, we present a NGMM to model healthy brains. Second, we extend our 2D FVF algorithm to 3D space and use it for brain tumor segmentation. The proposed method is validated on a publicly available dataset. Irene Cheng 0001, Anup Basu |
ICIP | 3 |
| 2010 | Automatic segmentation of spinal cord mri using symmetric boundary tracingabstractWe develop an adaptive active contour tracing algorithm for extraction of spinal cord from MRI that is fully automatic, unlike existing approaches that need manually chosen seeds. We can accurately extract the target spinal cord and construct the volume of interest to provide visual guidance for strategic rehabilitation surgery planning. Dipti Prasad Mukherjee, Irene Cheng 0001, Nilanjan Ray, Vivian Mushahwar, R. Marc Lebel, Anup Basu |
IEEE Trans. Inf. Technol. Biomed. | 6 |
| 2009 | Distortion metric for robust 3D point cloud transmissionabstractThis paper discusses using a forward error correction (FEC) algorithm to protect the transmission of progressively compressed 3D point clouds against packets loss. We design a metric to evaluate each layer's quality contribution to the decoding result of the progressively compressed model. With this metric, we minimize the expected distortion when applying an Unequal Error Protection (UEP) strategy to allocate channel bits to different layers of the model. The performance of employing UEP and Equal Error Protection (EEP) are compared with respect to the expected distortion. Experimental results show that by incorporating our distortion estimation metric with UEP, the rendering quality of a reconstructed 3D model degrades more gracefully as the packet-loss rate increases. Feng Chen 0003, Irene Cheng 0001, Anup Basu |
ICME | 3 |
| 2009 | Interactive Graphics for Computer Adaptive TestingabstractAbstract Interactive graphics are commonly used in games and have been shown to be successful in attracting the general audience. Instead of computer games, animations, cartoons, and videos being used only for entertainment, there is now an interest in using interactive graphics for ‘innovative testing’. Rather than traditional pen‐and‐paper tests, audio, video and graphics are being conceived as alternative means for more effective testing in the future. In this paper, we review some examples of graphics item types for testing. As well, we outline how games can be used to interactively test concepts; discuss designing chemistry item types with interactive 3D graphics; suggest approaches for automatically adjusting difficulty level in interactive graphics based questions; and propose strategies for giving partial marks for incorrect answers. We study how to test different cognitive skills, such as music, using multimedia interfaces; and also evaluate the effectiveness of our model. Methods for estimating difficulty level of a mathematical item type using Item Response Theory (IRT) and a molecule construction item type using Graph Edit Distance are discussed. Evaluation of the graphics item types through extensive testing on some students is described. We also outline the application of using interactive graphics over cell phones. All of the graphics item types used in this paper are developed by members of our research group. Irene Cheng 0001, Anup Basu |
Comput. Graph. Forum | 2 |
| 2008 | Optimization of Symmetric Transfer Error for Sub-frame Video Synchronization
Meghna Singh, Irene Cheng 0001, Mrinal Mandal 0001, Anup Basu |
ECCV (2) | 4 |
| 2008 | An Algorithm for Automatic Difficulty Level Estimation of Multimedia Mathematical Test ItemsabstractThe use of multimedia in learning and testing has been gaining popularity over the last few years. In this paper, we describe a method for estimating difficulty level of a multimedia mathematical item type using Item Response Theory (IRT). Experimental results on determining the parameters of an IRT model through tests on a group of students are also presented. Linear regression equations are fitted based on test data collected on a group of students. Results show a high degree of correlation between observed difficulty levels and fitted values. Irene Cheng 0001, Rui Shen 0002, Anup Basu |
ICALT | 3 |
| 2008 | A confidence measure and iterative rank-based method for temporal registrationabstractIn this paper we develop a confidence measure that can determine if a given set of samples is suitable for inclusion in the reconstruction of a higher resolution dataset. The confidence measure is formulated as a weighted combination of two well defined objective functions. We discuss the scope of the confidence measure and the two key factors that affect it: (i) non-uniformity of the samples and (ii) error in temporal registration. We also present a greedy iterative rank-based method that uses the confidence measure for reconstruction from multiple sample sets. The proposed method is evaluated with real video, audio and MRI data. Meghna Singh, Mrinal Mandal 0001, Anup Basu |
ICASSP | 3 |
| 2008 | Stereo matching using random walksabstractThis paper presents a novel two-phase stereo matching algorithm using the random walks framework. At first, a set of reliable matching pixels is extracted with prior matrices defined on the penalties of different disparity configurations and Laplacian matrices defined on the neighbourhood information of pixels. Following this, using the reliable set as seeds, the disparities of unreliable regions are determined by solving a Dirichlet problem. The variance of illumination across different images is taken into account when building the prior matrices and the Laplacian matrices, which improves the accuracy of the resulting disparity maps. Even though random walks have been used in other applications, our work is the first application of random walks in stereo matching. The proposed algorithm demonstrates good performance using the Middlebury stereo datasets. Rui Shen 0002, Irene Cheng 0001, Xiaobo Li 0001, Anup Basu |
ICPR | 4 |
| 2008 | Human Activity Recognition Based on Silhouette DirectionalityabstractRecent advances in computer vision and pattern recognition have fueled numerous initiatives that aim to intelligently recognize human activities. In this paper, we propose an algorithm for nonintrusive human activity recognition. We use an adaptive background-foreground separation technique to extract motion information and generate silhouettes (foreground) from the input videos. We then derive directionality-based feature vectors (directional vectors) from the silhouette contours and use the distinct data distribution of directional vectors in a vector space for clustering and recognition. We also exploit the dynamic characteristic of human motion in order to smooth decisions over time and reduce errors in activity recognition. Our approach is monocular, tolerant to moderate view changes, and can be applied to both frontal and lateral views of most activities. Experiments with short and long video sequences show robust recognition under conditions of varying view angles, zoom depths, backgrounds, and frame rates. Meghna Singh, Anup Basu, Mrinal Mandal 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Contrast Enhancement from Multiple Panoramic ImagesabstractIn this work we discuss an efficient strategy for combining multiple panoramic scans of the same scene to create a single higher contrast image. Prior research papers mainly consider that precise exposure times for multiple images are known and create a new high dynamic range representation for each pixel. We simply consider multiple scans of the same scene where the exposure times are not known, and no special filters or image detectors are used in the image acquisition process. In this unrestricted scenario, the problem is how to select different parts of a scene from the "best" image among a collection of images in order to optimize the overall clarity or contrast of a single final image. Experimental results are shown, and steps to improve results using modified contrast measures for color images are described. The results are enhanced using morphological filtering and image blending. Irene Cheng 0001, Anup Basu |
ICCV | 2 |
| 2007 | Multimedia Adaptive Computer based Testing: An OverviewabstractInstead of computer games, animations, cartoons, and videos being used only for entertainment by kids, there is now an interest in using multimedia for "innovative testing." Rather than traditional paper-and-pencil tests, audio, video and graphics are being conceived as alternative means for more effective testing in the future. In this paper we review some examples of multimedia item types for testing. As well, we will outline research on testing the seven types of intelligence through multimedia; describe how games can be used to test physics concepts; discuss designing chemistry item types with interactive multimedia; consider architecture for supporting online multimedia testing; and suggest approaches for automatically determining difficulty level in interactive mathematical questions. Detailed description on various topics will be given in the other papers to follow in the special session. Anup Basu, Irene Cheng 0001, Mun Prasad, Gautam Rao |
ICME | 1 |
| 2007 | An Effective Multimedia Item Shell Design for Individualized EducationabstractThe advantages of creating multimedia item types and applying computer-based adaptive testing in education are: First, the capability to motivate learning by making the learners feel more engaged and interactive, as well as a better representation of concepts, which are not possible when using conventional multiple choice tests. Second, instead of following a curriculum designed for average students, an individual is given a customized curriculum suitable for his or her learning capability. However, the issue to address when achieving these goals is the enormous amount of item types required to transform the current multiple choice questions into multimedia formats, and the criteria used to determine the difficulty level of a multimedia question item. In this paper we propose a multimedia item shell design that not only reduces the number of item types required, but also computes difficulty level of an item automatically. The concept of question seed is also introduced to make content creation more cost-effective. Irene Cheng 0001, Anup Basu |
ICME | 2 |
| 2007 | Multimedia Games for Learning and Testing PhysicsabstractMany difficult concepts can often be best explained or understood through simple games. In order to motivate the understanding of movements and trajectories of projectiles, and other Physics concepts, we propose using educational games for learning and testing physics so that the students can feel more engaged and rewarding, and thus are motivated to acquire new knowledge. Our novel approach includes automatically generating the next game episode at a different difficulty level based on how the student performs in the current episode. This adaptive approach is associated with a computer scoring system which can evaluate the student's understanding of physics concepts and computational skills, as well as strategic planning. Saul D. Rodriguez, Irene Cheng 0001, Anup Basu |
ICME | 3 |
| 2007 | A note on 'A fully parallel 3D thinning algorithm and its applications'
Anup Basu |
Pattern Recognit. Lett. | 2 |
| 2007 | Perceptually Optimized 3-D Transmission Over Wireless NetworksabstractMany protocols optimized to transmissions over wireless networks have been proposed. However, one issue that has not been looked into is considering human perception in deciding a transmission strategy for three-dimensional (3D) objects. Several factors, such as the number of vertices and the resolution of texture, can affect the display quality of 3D objects. When the resources of a graphics system are not sufficient to render the ideal image, degradation is inevitable. It is therefore important to study how individual factors affect the overall quality, and how the degradation can be controlled given limited bandwidth resources and possibility of data loss. In this paper, the essential factors determining the display quality are reviewed. We provide an overview of our research on designing a 3D perceptual quality metric integrating two important ones, resolution of texture and resolution of mesh, that control the transmission bandwidth requirements. A review of robust mesh transmission considering packet loss is presented, followed by a discussion of the difference of existing literature with our problem and approach. We then suggest alternative strategies for packet transmission of both 3D texture and mesh. These strategies are then compared with respect to preserving 3D perceptual quality under packet loss Irene Cheng 0001, Anup Basu |
IEEE Trans. Multim. | 2 |
| 2007 | Event Dynamics Based Temporal RegistrationabstractTemporal registration is the establishment of correspondence between two (or more) temporal frames of video sequences, or 3-D volume data. In this paper, we propose to use event dynamics, a property that is inherent to an event and is thus common to all acquisitions of the event, for both global and local temporal registration of video sequences in order to generate high temporal resolution video. We compare our approach to a widely used linear interpolation based temporal registration algorithm and demonstrate that in the case of low temporal acquisition rate, a global event dynamics based approach, such as ours, has smaller temporal registration error. We also present a unique application of our work in solving 3-D (2D + time) high temporal resolution medical data visualization problem. Meghna Singh, Anup Basu, Mrinal Mandal 0001 |
IEEE Trans. Multim. | 2 |
| 2006 | Image Based Temporal Registration of MRI Data for Medical VisualizationabstractThe capability of creating video data from MRI has many advantages in visualization for medical practitioners, including (i) not subjecting patients to harmful radiations, (ii) being able to monitor patients at short inter-exam time intervals, and (iii) being able to capture 3D volume data. The quality and speed with which MRI data can be acquired, however, poses a challenge towards supporting good quality visualization. In this work we present results from our preliminary attempts at enhancing the temporal resolution of video captured via MRI. Our initial focus is on visualization of swallowing and associated problems that are broadly categorized as Dysphagia. We present a method to register data from multiple swallows to generate high temporal resolution MRI videos. 1. Meghna Singh, Richard Thompson, Anup Basu, Jana Rieger, Mrinal Mandal 0001 |
ICIP | 3 |
| 2006 | Packet Loss Modeling for Perceptually Optimized 3D TransmissionabstractTransmissions over wireless and other unreliable networks can lead to packet loss. An area that has received limited research attention is how to tailor multimedia information taking into account the way packets are lost. We provide a brief overview of our research on designing a 3D perceptual quality metric integrating two important factors, resolution of texture and resolution of mesh, which control transmission bandwidth. We then suggest alternative strategies for packet 3D transmission of both texture and mesh. These strategies are then compared with respect to preserving 3D perceptual quality under packet loss in ad hoc wireless networks. Experiments are conducted to study how the time between consecutive packet transmission and packet size affects loss in wireless channels. A preliminary model for estimating the optimal packet size is then proposed Irene Cheng 0001, Lihang Ying, Anup Basu |
ICME | 3 |
| 2006 | Improving Multimedia Innovative Item Types for Computer Based TestingabstractInstead of computer games, animations, cartoons, and videos being used only for entertainment by kids, there is now an interest in using some of these media for educational purposes as well. Along with content creation, multimedia has potential for use in "innovative testing". Rather than traditional paper-and-pencil tests, audio, video and graphics are being conceived as alternative means for more effective testing in the future [1,17,21,28,29,30,33,42,44,49,50]. For example, we would like to use animation and games to help in learning concepts; consider how image, graphics and audio tools can be used for innovative testing; and develop techniques for measuring the impact of multimedia in improving performance or arousing interest in students. In this paper we discuss some examples of multimedia item types for testing, followed by a strategy for adaptive testing using those item types. We also show how techniques for perceptual evaluations can be used to improve strategies for adaptive testing Irene Cheng 0001, Anup Basu |
ISM | 2 |
| 2006 | Traceroute-Based Fast Peer Selection without Offline DatabaseabstractThe extreme heterogeneity in the P2P Internet environment causes the connections to candidate peers to vary significantly. Thus, selecting good candidate peers is critical to P2P networking performance. While most research in the literature focus on selecting peers with low latency, high uploading bandwidth, and high serving stability by time-consuming end-to-end measurement, we address how to quickly narrow the scope of good candidate peers. In this paper, we proposed to pick out the close peers from the same location or with low round-trip time from each other, by observing that close peers share routers along the traceroute path to the same destination. Our method only maintains online peers's traceroute records, without any offline database. Experiments with real traceroute records from 352 different IP addresses around the world verify the efficiency of our method Lihang Ying, Anup Basu |
ISM | 2 |
| 2005 | Balanced incomplete designs for 3D perceptual quality estimationabstractMany factors, such as the number of vertices and the resolution of texture, can affect the display quality of 3D objects. When the resources of a graphics system are not sufficient to render the ideal image, degradation is inevitable. It is therefore important to study how individual factors affect the overall quality, and how the degradation can be controlled given limited resources. The essential factors determining the display quality are reviewed and a 3D perceptual estimation method is described. One of the major concerns in designing perceptual experiments is the large number of evaluations to be performed by judges, which results in fatigue and errors in judgement. To reduce judging fatigue and increase reliability of evaluations we propose using a statistical approach, following the design of experiments technique of balanced incomplete block design (BIBD). We develop models following BIBD for perceptual experiments and validate our model with experimental results. Even though the BIBD framework for perceptual evaluations is described in the context of 3D quality estimation, the approach can be used in other perceptual evaluation scenarios as well. Anup Basu, Irene Cheng 0001 |
ICIP (1) | 1 |
| 2005 | Visual gesture recognition for ground air traffic control using the Radon transformabstractHuman gesture recognition is an active topic of vision research which has applications in diverse fields such as collaborative virtual environments and robot teleoperation. We propose a novel method for the recognition of hand gestures, used by air marshals for steering aircraft on the runway, using the Radon transform. Various aspects of the algorithm, including acquisition, segmentation, labeling and recognition using the parametric Radon transform are addressed in this paper. A binary skeleton representation of the human subject is computed. The Radon transform is used to generate maxima corresponding to specific orientations of the skeletal representation. Feature vectors are extracted from the transform space by computing the normalized cumulative projections of the Radon transform on the angle axis. K-means clustering is then applied to recognize static gestures from the extracted features. This technique has the potential to provide information about the exact orientation of gesture segments and can find use in ground control of unmanned air vehicles. Experiments with image data corresponding to the various ground air traffic control gestures used in directing aircrafts, highlight the potential application of this approach. Meghna Singh, Mrinal Mandal 0001, Anup Basu |
IROS | 3 |
| 2005 | Panoramic stereo reconstruction using non-SVP optics
Mark Fiala, Anup Basu |
Comput. Vis. Image Underst. | 2 |
| 2005 | Gaussian and Laplacian of Gaussian weighting functions for robust feature based tracking
Meghna Singh, Mrinal Mandal 0001, Anup Basu |
Pattern Recognit. Lett. | 3 |
| 2005 | Quality metric for approximating subjective evaluation of 3-D objectsabstractMany factors, such as the number of vertices and the resolution of texture, can affect the display quality of three-dimensional (3-D) objects. When the resources of a graphics system are not sufficient to render the ideal image, degradation is inevitable. It is, therefore, important to study how individual factors will affect the overall quality, and how the degradation can be controlled given limited resources. In this paper, the essential factors determining the display quality are reviewed. We then integrate two important ones, resolution of texture and resolution of wireframe, and use them in our model as a perceptual metric. We assess this metric using statistical data collected from a 3-D quality evaluation experiment. The statistical model and the methodology to assess the display quality metric are discussed. A preliminary study of the reliability of the estimates is also described. The contribution of this paper lies in: 1) determining the relative importance of wireframe versus texture resolution in perceptual quality evaluation and 2) proposing an experimental strategy for verifying and fitting a quantitative model that estimates 3-D perceptual quality. The proposed quantitative method is found to fit closely to subjective ratings by human observers based on preliminary experimental results. Yixin Pan, Irene Cheng 0001, Anup Basu |
IEEE Trans. Multim. | 3 |
| 2005 | Image retrieval based on histogram of fractal parametersabstractImage indexing and retrieval techniques are important for efficient management of visual databases. These techniques are generally developed based on the associated compression techniques. In the fractal domain, luminance offset and contrast scaling parameter are typically used as the fractal indices. However, luminance offset and contrast scaling parameter are strongly correlated. In this paper, we prove that range block mean and contrast scaling parameters are independent. Based on this independence, we propose four statistical indices for efficient image retrieval. In addition, we propose an efficient hierarchical indexing strategy based on the de and ac component analysis. Experimental results on a database of 416 texture images, created by decomposing 26 images, indicate that the proposed indices significantly improve the retrieval rate, compared to other retrieval methods. Ming Hong Pi, Mrinal Mandal 0001, Anup Basu |
IEEE Trans. Multim. | 3 |
| 2004 | A comparison of non-orthogonal and orthogonal fractal decodingabstractThe model coefficients for the case of the non-orthogonal basis in Jacquin's mappings are contrast scaling and luminance offset. After orthogonalization, the model coefficients become range block mean and contrast scaling. These two fractal coding algorithms have the same encoding procedure except that luminance offset is replaced by range block mean, however, their decoding algorithms are different. In this paper, we prove that the orthogonal decoding algorithm converges faster than the non-orthogonal algorithm, while the two decoding algorithms produce the same iteration series if the initial image is the range-averaged image and the step size equals the side-length of the range block. Ming Hong Pi, Anup Basu, Mrinal Mandal 0001 |
ICIP | 2 |
| 2004 | Robust wireless transmission of regions of interest in jpeg2000abstractIn this paper, we present a technique to robustly transmit regions of interest in the JPEG2000 framework. The technique assumes a prioritized region-of-interest coding and optimally assigns channel protection to the coded data according to the importance of every packet in the final bit-stream. The mean energy of the transform coefficients contained in a packet and the distance of a packet from the region of interest determine the importance of every packet. The channel protection is achieved by means of a concatenation of a cyclic redundancy check outer coder and an inner rate-compatible convolutional coder. Simulation results performed over a Rayleigh fading channel show an improvement in the visual quality of the reconstructed images. Victor F. Sanchez, Mrinal Mandal 0001, Anup Basu |
ICIP | 3 |
| 2004 | Robot navigation using panoramic tracking
Mark Fiala, Anup Basu |
Pattern Recognit. | 2 |
| 2004 | Scalable edge enhancement with automatic optimization for digital radiographic images
Lijun Yin 0001, Anup Basu, Ja Kwei Chang |
Pattern Recognit. | 2 |
| 2004 | Prioritized region of interest coding in JPEG2000abstractA method is proposed to encode multiple regions of interest in the JPEG2000 image-coding framework. The algorithm is based on the rearrangement of packets in the code-stream to place the regions of interest before the background coefficients. In order to improve the quality of the reconstructed image, partial background information is included with the regions of interest. The method makes use of a Gaussian priority distribution to assign different priority levels to background and region of interest packets. The priority level is in turn used to determine how much background information should be included with the regions of interest. The proposed technique is fully compatible with the current JPEG2000 standard and allows transmission of different regions of interest with different priorities. Experimental results demonstrating the validity of the proposed approach are presented and compared with existing region of interest coding techniques. Victor F. Sanchez, Anup Basu, Mrinal Mandal 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | Image retrieval based on histogram of new fractal parametersabstractImage indexing and retrieval techniques are important for efficient management of visual databases. These techniques are generally developed based on the associated compression techniques. In the fractal domain, the luminance offset and contrast scaling parameters are typically used as the fractal index. We propose to use the range block mean and contrast scaling as the fractal index. The image retrieval is performed in two steps. First, a coarse search is performed using the histogram of the range block means. Subsequently, a fine search is performed using the 2D joint histogram of the range block mean and contrast scaling parameters. Experimental results on a database of 416 texture images indicate that the proposed indices significantly improve the retrieval rate, compared to other retrieval methods. Ming Hong Pi, Mrinal Mandal 0001, Anup Basu |
ICASSP (3) | 3 |
| 2003 | Perceptual quality metric for qualitative 3D scene evaluationabstractVarious factors affect the quality of 3D images, such as the number of vertices and the resolution of texture. In this report, we discuss the factors determining the quality of 3D images. Among these factors two important ones, most relevant for bandwidth constrained online applications, are combined to model a perceptual metric. We estimate this metric from the statistical data collected in a 3D quality evaluation experiment. The theory behind modeling a quality metric and the details of experiments performed for deriving the parameters of this metric are described. Yixin Pan, Irene Cheng 0001, Anup Basu |
ICIP (3) | 3 |
| 2003 | A new decoding algorithm based on range block mean and contrast scalingabstractA new fractal decoding algorithm based on range block mean and contrast scaling is presented. As a result, the decoding image is represented as the sum of the DC and AC components. We prove that new decoding algorithm converges faster than existing ones. Experiments show that the new decoding algorithm converges at least 2-3 times faster than existing methods. Ming Hong Pi, Anup Basu, Mrinal Mandal 0001 |
ICIP (2) | 2 |
| 2003 | Image retrieval based on 2-D histogram of fractal parametersabstractAn image can be characterized by its fractal parameters, and hence, the fractal parameters can be used as the image signature to retrieve the images. In this paper, based on the principle that fractal transform is completely determined by luminance offset and contrast scaling, we first propose histogram of luminance offset as a statistical index, and we further propose three composite indices by combining individual histograms to enhance retrieval rate and reduce computational complexity. Experimental results on a database of 416 texture images indicate that the proposed indices significantly improve the retrieval rate, compared to other retrieval methods. Ming Hong Pi, Mrinal Mandal 0001, Anup Basu |
ICME | 3 |
| 2003 | Recognizing facial expressions using active textures with wrinklesabstractThis paper explores the use of facial wrinkle textures for recognizing the facial expressions. Based on the observation of the wrinkles appearance and change along with performed expressions, we propose to extract the partial texture information in both the facial organ areas (e.g., eyes and mouth) and the facial wrinkle areas, and use the texture dissimilarity between the neutral expression and the active expression to extract the active texture for the expression representation. We present a novel method using multiple levels of detail to measure the active texture dissimilarity. The rate of change between levels is used as the rule for discriminating 6 types of universal expressions. The experiments on video sequences demonstrate the simplicity and efficiency of the proposed method for recognizing expressions with an 82.8% correct recognition rate. Lijun Yin 0001, Sergey Royt, Matt T. Yourst, Anup Basu |
ICME | 4 |
| 2003 | QoS based video delivery with foveation and bandwidth monitoring
Irene Cheng 0001, Anup Basu |
Pattern Recognit. Lett. | 2 |
| 2003 | Optimal adaptive bandwidth monitoring for QoS based retrievalabstractNetwork aware multimedia delivery applications are a class of applications that provide certain level of quality of service (QoS) guarantees to end users while not assuming underlying network resource reservations. These applications guarantee QoS parameters like media object transmission time limit by actively monitoring the available bandwidth of the network and adapting the object to a target size that can be transmitted within a given time limit. A critical problem is how to obtain an accurate enough estimation of available bandwidth while not wasting too much time in bandwidth testing. In this paper, we present an algorithm to determine optimal amount of bandwidth testing given a probabilistic confidence level for network-aware multimedia object retrieval applications. The model treats the bandwidth testing as sampling from an actual bandwidth population. It uses statistical estimation method to quantify the benefit of each new bandwidth-testing sample, which is used to determine the optimal amount of bandwidth testing by balancing the benefit with the cost of each sample. Our implementation and experiments shows the algorithm determines the optimal amount of bandwidth testing effectively with minimum computation overhead. Yinzhe Yu, Irene Cheng 0001, Anup Basu |
IEEE Trans. Multim. | 3 |
| 2002 | Color-based mouth shape tracking for synthesizing realistic facial expressionsabstractMouth shape analysis and synthesis play an important role for realistic facial expression generation. The conventional deformable template-based method requires a fairly accurate initial localization of the template because the energy minimization process only finds a local minimum. We present a new method for accurate mouth corner detection and mouth shape estimation using color information. The proposed method can deal with a variety of shapes for both open and closed positions of the mouth. Accurate mouth corner detection and mouth status determination (open and closed) increases the accuracy of template initialization. The extracted mouth shape parameters can be used for synthesizing realistic virtual facial expressions in model based coding. Experiments on video sequences demonstrate the advantage of the proposed algorithm. Lijun Yin 0001, Anup Basu |
ICIP (1) | 2 |
| 2002 | Synthesis-based scalable image enhancement for digital radiographyabstractThe ability to resolve fine picture detail is of paramount importance in a medical imaging system when it comes to view small tissues, bone structure and anatomy in X-ray images. In order to enhance diagnostic information and suppress irrelevant detail, we present a new digital X-ray imaging processing system with the property of scalability and adaptability. First, a new optimum contrast enhancement algorithm is proposed for display. The adaptive detection of the region-of-interest is developed. Second, a so called "scalable edge enhancement algorithm" is proposed to improve the image quality for showing subtle structures of the digital X-ray images. The advantage of the scheme is demonstrated by an experiments on 200 X-ray images, in which different parts of human body structures are captured. Lijun Yin 0001, Ja Kwei Chang, Anup Basu |
ICIP (2) | 3 |
| 2002 | Active Tracking and Cloning of Facial Expressions Using Spatio-Temporal InformationabstractThis paper presents a new method to analyze and synthesize facial expressions, in which a spatio-temporal gradient based method (i.e., optical flow) is exploited to estimate the movement of facial feature points. We proposed a method (called motion correlation) to improve the conventional block correlation method for obtaining motion vectors. The tracking of facial expressions under an active camera is addressed. With the motion vectors estimated, a facial expression can be cloned by adjusting the existing 3D facial model, or synthesized using different facial models. The experimental results demonstrate that the approach proposed is feasible for applications such as low bit rate video coding and face animation. Lijun Yin 0001, Anup Basu, Matt T. Yourst |
ICTAI | 2 |
| 2002 | Analysis of depth estimation error for cylindrical stereo imaging
Anup Basu, Hossein Sahabi |
Pattern Recognit. | 1 |
| 2002 | Hough transform for feature detection in panoramic images
Mark Fiala, Anup Basu |
Pattern Recognit. Lett. | 2 |
| 2001 | Nose shape estimation and tracking for model-based codingabstractFeature extraction on the face plays an important role in applications of model based coding and human face recognition. Traditionally, the eyes and mouth are considered to be the most significant features contributing to different facial expressions. However, detecting and tracking the nose shape is non-trivial, and plays an equally important role as eyes and mouth for model based coding, especially for analysis and synthesis of realistic facial expressions. A feature detection method on the facial organ areas is presented. Individual templates are designed for the nostril and nose-side. First, the feature regions are limited to certain areas by using two-stage region growing methods. Second, the pre-defined templates are applied to extract the shape of the nostril and nose-side. Finally, the extracted feature shapes are exploited to guide a facial model to complete an accurate adaptation. The advantage of the proposed scheme is demonstrated by experiments on real video sequences for low bit rate video coding. Lijun Yin 0001, Anup Basu |
ICASSP | 2 |
| 2001 | QoS based video delivery with foveationabstractSpatially varying sensing (foveation) was first used as a means for image compression in our past research. We extend previous work to address the advantages of foveation in improving the performance of MPEG compression over bandwidth limited channels, such as the Internet. Unlike other approaches to foveating MPEG which used multiresolution representations, we use continuously spatially varying resolution and demonstrate that this approach is indeed advantageous over others. Two parameters, scaling and distortion, are used to allow us to adapt MPEG video to various compression ratios. Experimental results are presented, and can be viewed on the Web, validating our approach. Irene Cheng 0001, Anup Basu |
ICIP (1) | 2 |
| 2001 | Generating Realistic Facial Expressions with Wrinkles for Model-Based Coding
Lijun Yin 0001, Anup Basu |
Comput. Vis. Image Underst. | 2 |
| 2001 | Synthesizing realistic facial animations using energy minimization for model-based coding
Lijun Yin 0001, Anup Basu, Stefan Bernögger, Axel Pinz |
Pattern Recognit. | 2 |
| 2001 | Improving image and video transmission quality over ATM with foveal prioritization and priority dithering
Kevin James Wiebe, Anup Basu |
Pattern Recognit. Lett. | 2 |
| 2000 | Texture decomposition and correlation thresholding for realistic low-bit rate model-based codingabstractFacial texture updating and compression are crucial in achieving realistic facial animation for low-bit rate coding. We present an efficient way to update and encode facial textures in model-based coding. Differing from traditional methods which update the entire facial texture, we propose a partial texture updating method for realistic facial expression synthesis with facial wrinkles. Experiments on video sequences demonstrate the advantage of the proposed algorithm in keeping transmission cost low while producing realistic expressions. Lijun Yin 0001, Anup Basu |
ICASSP | 2 |
| 2000 | 3D Estimation using Panoramic StereoabstractOmni-directional sensors are useful in obtaining a 360/spl deg/ field of view with a single lens camera. Fiala et al. (1996) described an approach to designing a stereo panoramic imaging system using a double-lobed mirror surface. This paper extends our earlier work by relating 3D points to the 2D projections obtained by our system. Applications of the system include surveillance and security, automatic and passive monitoring for peace-keeping tasks, and fast detection of corrosion and deposits in pipelines and industrial containers. Experimental results verifying the mathematical modeling are also provided. Jonathan Baldwin, Anup Basu |
ICPR | 2 |
| 2000 | Analysis of Cylindrical Stereo ImagingabstractConventional stereo imaging uses area CCDs for which depth perception error has been analyzed in our past research (Sahabi and Basu, 1996). In this work we analyze the depth perception error for a stereo system that uses two rotating linear CCD cameras to create cylindrical stereo images. Theoretical analysis shows certain advantages to cylindrical stereo imaging over conventional methods. Initial experimental results are presented to validate the theoretical results. Unlike normal area CCD cameras, rotating linear CCD cameras can capture very high resolution images thereby substantially reducing depth estimation error; more importantly, the error has certain directionally uniform characteristics. Anup Basu, Hossein Sahabi |
ICPR | 1 |
| 2000 | Hybrid video using motion estimationabstractOne of the major problems in low-bandwidth telerobotics applications is determining what visual data is sufficient to allow the operator to successfully perform a task. When the available bandwidth is extremely low, we must severely restrict the data being transmitted. For very-low-bandwidth applications, displaying certain types of features, such as edges in indoor navigation applications, allows for successful completion of the operator's task. As the available bandwidth increases, the operator can be presented with more information than just the edge features. Using simple overlay techniques to overlay a low-frame-rate video stream on a high-frame-rate edge feature stream provides the operator with more information about the environment. Registering the video data with the moving edge data poses a significant problem. In this paper, we propose to use simple motion estimation techniques based on the edge image stream to estimate the motion of the edge images and use this data to register the slower frame-rate video stream to the edge feature stream. Using these simple estimation techniques allows us to perform the video stream registration quickly as the edge feature data is being presented to the operator. Jonathan Baldwin, Anup Basu, Hong Zhang 0013 |
SMC | 2 |
| 1999 | Face Model Adaptation with Active TrackingabstractA system for face model adaptation combining active tracking is presented. Input from an active camera is used for MPEG4 model based coding. First, the background is compensated considering a moving camera (tilt or pan). Second, the talking face is segmented from the compensated background using fusion of frame differences. A morphological filter is then applied to make the system less sensitive to noise. Third, Hough transform and deformable template coupled with color information are exploited to detect the facial features, e.g., eyes, mouth. Fourth, a wireframe model is adapted to the extracted face by an extended dynamic mesh. The feasibility of the proposed system is demonstrated using several real active video sequences. Lijun Yin 0001, Anup Basu |
ICIP (4) | 2 |
| 1999 | Panoramic Video with Predictive Windows for Telepresence ApplicationsabstractWe describe the application of a predictive Kalman filter to the display of panoramic images. We discuss integrating a panoramic imaging system with prediction of viewing direction to create an effective telepresence system over low bandwidth links. Panoramic imaging using a reflective mirror surface offers an alternative to pan-tilt systems for obtaining a 360 degree field of view. Selecting a small window within a panoramic image allows a meaningful part of an image from a remote site to be seen at a higher refresh rate. Because of the delay in transmitting an image from a remote site, it is necessary to have additional image information available locally. This information can be used to simulate continuously flowing pictures with reduced apparent delay. Continuity in image viewing is achieved by predicting the next viewpoint of an operator and preemptively transmitting parts of an image. Experimental results are given to evaluate the proposed telepresence system. Jonathan Baldwin, Anup Basu, Hong Zhang 0013 |
ICRA | 2 |
| 1999 | Integrating active face tracking with model based coding
Lijun Yin 0001, Anup Basu |
Pattern Recognit. Lett. | 2 |
| 1998 | Eye tracking and animation for MPEG-4 codingabstractAccurate localization and tracking of facial features are crucial for developing high quality model-based coding (MPEG-4) systems. For teleconferencing applications at very low bit rates, it is necessary to track eye and lip movements accurately over time. These movements can be coded and transmitted to a remote site, where animation techniques can be used to synthesize facial movements on a model of a face. In this paper we describe and simple heuristics which are effective in improving the results of well-known facial feature detection and tracking algorithms. Animation models are also presented, along with experimental results to demonstrate the system being developed. We focus our discussion only on the detection, tracking and modeling of eye movements. Stefan Bernögger, Lijun Yin 0001, Anup Basu, Axel Pinz |
ICPR | 3 |
| 1998 | Predictive Windows for Delay Compensation in Telepresence ApplicationsabstractPredictive Kalman filters can be used to predict positions of a mouse when it is operated under some basic assumptions. This prediction can be used to estimate what portions of a larger image an operator wants to view. This paper discusses theory and experimentation being done at the University of Alberta using predictive Kalman filters to provide predictive windows for low bandwidth telepresence applications. We compare several state models used in prediction with each other and also with having no prediction, both with numerical measures, and on human subjects. We show that the constant velocity model provides the best prediction results. Jonathan Baldwin, Anup Basu, Hong Zhang 0013 |
ICRA | 2 |
| 1998 | Analysis and synthesis of facial expressions for MPEG-4 systemabstractThis paper presents a new method to analyze and synthesize facial expressions for model-based coding, which uses optical flow to estimate the movements of facial feature points and then obtains the motion vectors based on motion correlation and block correlation. With the motion vectors estimated, a facial expression can be synthesized by adjusting the existing 3-D facial model. Preliminary experimental results demonstrate that the approach proposed is feasible for very low bit rate transmission (e.g. MPEG-4 application). Lijun Yin 0001, Anup Basu |
SMC | 2 |
| 1998 | From 2d Surface Patches To 3d Reconstructed Models: Theory And Applications
Ashraf Elnagar, Anup Basu |
Pattern Recognit. | 2 |
| 1998 | Enhancing videoconferencing using spatially varying sensingabstractHuman vision can be characterized as a variable resolution system-the region around the fovea (point of attention) is observed with great detail, whereas the periphery is viewed in lesser detail. In this work, we show that spatially varying sensing (resembling the human eye) can indeed be useful in videoconferencing. A system that can incorporate multiple and moving foveae is described. There are various possible ways of implementing multiple foveae and combining information from them. Some of these alternative strategies are discussed and results are compared. We also show the advantage of using spatially varying sensing as a preprocessor to JPEG. The methods described here can be useful in designing teleconferencing systems and image databases. Anup Basu, Kevin James Wiebe |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 1997 | MPEG4 Face Modeling Using Fiducial PointsabstractWe present an approach for modeling a person's face for model-based coding (MPEG4) which will cater to very low bit-rate video-conferencing applications. This approach is performed entirely automatically. We utilize two views of a person's face, and the fiducial points defined on a generic facial model, to modify the generic model into a 3D individual model of a face. To deform the generated individual facial model for tracking the expression of a face over time, a new method named the layered force spreading method (LFSM) is proposed which makes the animation of facial expressions looks more natural. The feasibility of our approach is demonstrated using a real facial image. Lijun Yin 0001, Anup Basu |
ICIP (1) | 2 |
| 1997 | Modelling ecologically specialized biological visual systems
Kevin James Wiebe, Anup Basu |
Pattern Recognit. | 2 |
| 1997 | Active camera calibration using pan, tilt and rollabstractThree dimensional vision applications, such as robot vision, require modeling of the relationship between the two-dimensional images and the three-dimensional world. Camera calibration is a process which accurately models this relationship. The calibration procedure determines the geometric parameters of the camera, such as focal length and center of the image. Most of the existing calibration techniques use predefined patterns and a static camera. Recently, a novel calibration technique for computing the focal length and image center, which uses an active camera, has been developed. This technique does not require any predefined patterns or point-to-point correspondence between images-only a set of scenes with some stable edges. It was observed that the algorithms developed for the image center are sensitive to noise and hence unreliable in real situations. This report extends the techniques provided to develop a simpler, yet more robust method for computing the image center. Anup Basu, Kavita Ravi |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1996 | Optimal non-uniform discretization for stereo reconstructionabstractIn stereo vision the depth of a 3D point is estimated based on the position of its projections on the left and right images. The imaging sensors of cameras, such as charge coupled devices (CCD), consist of discrete pixels. The discretization of images generates uncertainty in estimation of the depth at each 3D point. In this paper, we study the discretization of stereo images that provides the best 3D perception to an end-user. The research will be used to create optimal designs for head-mounted or LCD-based stereo displays. Anup Basu, Hossein Sahabi |
ICPR | 1 |
| 1996 | Panoramic stereoabstractOmni-directional sensors are useful in obtaining a 360/spl deg/ field of view with a single lens camera. This work describes how sensors can be designed for obtaining stereo. Alternative approaches are presented and evaluated. A technique for calibrating panoramic stereo cameras using colour coding is described. As well, real-time video hardware has been designed and built to support panoramic stereo display. David Southwell, Anup Basu, Mark Fiala, Jereome Reyda |
ICPR | 2 |
| 1996 | Improving image and video transmission quality over ATM with foveal prioritization and priority ditheringabstractA fundamental drawback to increasingly popular ATM-based switching is the possibility of information loss with congestion. We demonstrate that with intelligent, fovea driven priority assignment of image data, we can reduce the negative impact of information loss over ATM networks. ATM standards allow a single bit to indicate high or low packet priority. To reduce the effect of this restriction we introduce the concept of priority dithering. Network multimedia multicast scenarios over heterogeneous link capacities where foveal prioritization would be of benefit are described. Network simulation results of this method are included which demonstrate the advantages of priority dithered foveal prioritization over traditional methods. Kevin James Wiebe, Anup Basu |
ICPR | 2 |
| 1996 | A conical mirror pipeline inspection systemabstractPresents a novel system that provides high quality imaging of the interior surface of pipelines. The imaging device uses a refined conical mirror micro-surface machined from aluminum to image a 360/spl deg/ strip of the pipe in a single frame. A continuous stream of such images is transmitted to the surface where it is stored on a standard video recorder. The image sequences are processed off-line, ultimately producing high definition imagery of areas of interest. We address the design issues of the conic profile of the mirror, and some aspects of the image reconstruction/registration process. David Southwell, Basil Vandegriend, Anup Basu |
ICRA | 3 |
| 1996 | Analysis of Error in Depth Perception with Vergence and Spatially Varying Sensing
Hossein Sahabi, Anup Basu |
Comput. Vis. Image Underst. | 2 |
| 1995 | Surface Integration for Inspection TasksabstractIn underwater environments it is often difficult to obtain a big/clear picture of a scene. For that reason, a system that can integrate small pieces of images (taken from close range) into a composite 3D surface, is developed here. The device, along with 3D position/orientation estimation equipment, can be used for inspection of hulls of ships anchored in a bay, or for examination of underwater pipes and tanks. Experimental results are presented which validate the algorithms developed. Anup Basu, Ashraf Elnagar, Mark Fiala |
ICRA | 1 |
| 1995 | Active Camera Calibration Using Pan, Tilt and RollabstractThree dimensional vision applications, such as robot vision, require modelling of the relationship between the 2D images and the 3D world. Camera calibration is a process which accurately models this relationship. The calibration procedure determines the geometric parameters of the camera, such as focal length and center of the image. Most of the existing calibration techniques use predefined patterns and a static camera. Recently, A. Basu (1993) developed a novel calibration technique for computing the focal length and image center which uses an active camera. This technique does not require any predefined patterns or point to point correspondence between images-only a set of scenes with some stable edges. It was observed that the algorithms developed for image center are sensitive to noise and hence unreliable in real situations. The article extends the techniques provided by Basu to develop a simpler, yet more robust method for computing the image center. Anup Basu, Kavita Ravi |
ICRA | 1 |
| 1995 | Robust Detection of Moving Objects by a Moving Observer on Planar SurfacesabstractWe introduce a technique for detecting moving objects from an image sequence obtained with a moving camera using the planarity constraint. To increase the robustness of this technique, false motion caused by inaccuracies in sensor readings is eliminated by use of a morphological filter. This involves two successive operations-erosion and dilation-performed on a motion compensated image. Experimental results with real images are presented. Applications to the compression of moving images are now being investigated. Ashraf Elnagar, Anup Basu |
ICRA | 2 |
| 1995 | An active technique for piecewise calibration of robot manipulatorsabstractRobot calibration is essential to improve the positioning accuracy of robot manipulators. A mathematical (kinematic) model is used to describe the geometric structure of a robot manipulator. Robot calibration procedure involves calculating and improving the values of this model's parameters. The robot calibration technique presented in this paper uses a vision system, to calibrate the kinematic model. For an n-linked robot manipulator, the procedure calibrates one link at a time, starting with the n/sup th/ link by making small movements. The resulting equations are linear, consequently, the algorithms are simple. Kavita Ravi, Anup Basu |
IROS (1) | 2 |
| 1995 | Motion detection using background constraints
Ashraf Elnagar, Anup Basu |
Pattern Recognit. | 2 |
| 1995 | Improving boundary detection using variable resolution masks
Anup Basu, Manoj K. Jain, Xiaobo Li 0001 |
Pattern Recognit. Lett. | 1 |
| 1995 | Alternative models for fish-eye lenses
Anup Basu, Sergio Licardie |
Pattern Recognit. Lett. | 1 |
| 1995 | Active calibration of cameras: theory and implementationabstractThe problem of calibrating a camera has been widely addressed in the past. Almost all techniques described in the literature use a known calibrating pattern and a static camera. We introduce a novel technique, based on an active camera, which does not need any predefined patterns. All that is required is a scene with some strong and stable edges. Two alternative algorithms are presented and analysed-the second method is shown to be more robust to noise and useful in practical situations. Using our methods, an active camera can automatically calibrate itself. Experimental results are shown, demonstrating the validity of the algorithms.> Anup Basu |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1994 | Videoconferencing using spatially varying sensing with multiple and moving foveaeabstractA new method for videoconferencing using the concept of spatially varying sensing is introduced. Various techniques are discussed for combining information obtained from multiple points-of-interest (foveae) in an image. A fovea can be used to follow a moving object of interest as well. A fast videoconferencing prototype for desktop computers is also described. Anup Basu, Kevin James Wiebe |
ICPR (3) | 1 |
| 1994 | Crack Detection Using Contact SensingabstractIn this paper, we describe a robotic system employing contact sensing to detect the presence of cracks in surfaces. Automated detection of cracks in surfaces has many practical applications. Tasks such as inspection of underground pipes carrying fluids and the inspection of ship hulls are examples of jobs that are difficult for human operators. Robots equipped with vision are ill-suited for such tasks because of low visibility and fluid disturbance. Contact sensing is shown to provide a viable alternative for surface inspection tasks in situations where images cannot be obtained. The mathematical modeling of the system, some simulations and experimental results are provided.> Raju Patil, Anup Basu, Hong Zhang 0013 |
ICRA | 2 |
| 1994 | Motion Tracking with an Active CameraabstractThis paper describes a method for real-time motion detection using an active camera mounted on a pan/tilt platform. Image mapping is used to align images of different viewpoints so that static camera motion detection can be applied. In the presence of camera position noise, the image mapping is inexact and compensation techniques fail. The use of morphological filtering of motion images is explored to desensitize the detection algorithm to inaccuracies in background compensation, Two motion detection techniques are examined, and experiments to verify the methods are presented. The system successfully extracts moving edges from dynamic images even when the pan/tilt angles between successive frames are as large as 3.> Don Murray, Anup Basu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1993 | Active calibration: alternative strategy and analysisabstractAlmost all camera calibration techniques use a known calibrating pattern and a static camera. Techniques based on an active camera, which does not need any predefined patterns, are introduced. All that is required is a scene with some strong and stable edges. Two algorithms are presented and analyzed. It is shown that one strategy performs much better in the presence of noise, and thus is preferable in practical situations. Experimental results are shown, demonstrating the validity of the algorithms.> Anup Basu |
CVPR | 1 |
| 1993 | Modeling fish-eye lensesabstractThe human visual system can be characterized as a variable-resolution system: foveal information is processed at very high spatial resolution whereas peripheral information is processed at low spatial resolution. Various transforms have been proposed to model spatially varying resolution. Unfortunately, special sensors need to be designed to acquire images according to existing transforms. In this work, two models of fish-eye transform are presented. The validity of the transformations is demonstrated by fitting the alternative models to a real fish-eye lens. Anup Basu, Sergio Licardie |
IROS | 1 |
| 1993 | Smooth and acceleration minimizing trajectories for mobile robotsabstractAn approach to generating smooth piecewise local trajectories for mobile robots is proposed. Given the configurations (position and direction) of two points, one searches for the trajectory that minimizes the integral of acceleration (tangential and normal). The resulting trajectory should not only be smooth but also safe in order to be applicable in real-life situations, so the authors investigate two different obstacle-avoidance constraints that satisfy the minimization problem. Unfortunately, in this case the problem becomes more complex and unsuitable for real-time implementations. Therefore, the authors introduce two simple solutions, based on the idea of polynomial fitting to generate safe trajectories, once a collision is detected with the original smooth trajectory. Simulation results for the different algorithms are presented. Ashraf Elnagar, Anup Basu |
IROS | 2 |
| 1993 | Active trackingabstractThis work describes a method for real-time motion detection using an active camera mounted on a pan/tilt platform. Image mapping is used to align images of different viewpoints so that static camera motion detection can be applied. In the presence of camera position noise, the image mapping is inexact and compensation techniques fail. The use of morphological filtering of motion images is explored to desensitize the detection of algorithm in inaccuracies in background compensation. Two motion detection techniques are examined, and experiments to verify the methods are presented. The system successfully extracts moving edges from dynamic images, even when the pan/tilt angles between successive frames are as large as 3/spl deg/. Don Murray, Anup Basu |
IROS | 2 |
| 1993 | Heuristics for local path planningabstractA heuristic technique for solving the problem of path planning based on local information for a mobile robot with acceleration constraints moving amidst a set of stationary obstacles is described. The concept of safety is used to design a planning strategy. A path based on local information that maximizes the product of safety and attraction towards the goal is chosen. The safety function depends on the acceleration bounds. The attraction towards the goal depends on the distance from the goal. Two additional heuristics are proposed to improve the efficiency of the search process, and to enhance the ability of the robot to avoid obstacles.> Ashraf Elnagar, Anup Basu |
IEEE Trans. Syst. Man Cybern. | 2 |
| 1992 | Heuristics for local path planningabstractThe authors describe a heuristic technique for solving the problem of path planning based on local information for a mobile robot with acceleration constraints moving amidst a set of stationary obstacles. The concept of safety is introduced to design a planning strategy. A path which maximizes the product of safety based on local information and attraction towards the goal is chosen. The safety function depends on the acceleration bounds. The attraction toward the goal depends on the distance from the goal. Two additional heuristics are proposed to improve the efficiency of the search process and to enhance the ability of the robot to avoid obstacles. Some simulation examples of the algorithm corresponding to different navigational environments are discussed.> Ashraf Elnagar, Anup Basu |
ICRA | 2 |
| 1992 | Motion Planning With Acceleration ConstraintabstractIn this paper we examine the problem of finding a local collision free path between two points in a two dimensional space. We consider a non-h.olonomic robot with constraints on normal acceleration. Solution, of this problem is important in order to prevent a robot moving on wheels from slipping (or skidding) while turning. The exact solution (known so far) to the problem of reachability is exponential. We analyze th,e problem of finding an approximate path locally with an a priori probability (or approximation measure). Our algorithm is based on a variable size discretization, which is computed locally depending on the given constraints, the size of the local free space, and the closeness of the approximation. Implemen,tation results demonstrating the validity of the method are presented. Anup Basu, Goksin Bakir, Hong Zhang 0013 |
IROS | 1 |
| 1992 | Optimal discretization for stereo reconstruction
Anup Basu |
Pattern Recognit. Lett. | 1 |
| 1991 | Variable-resolution character thinning
Xiaobo Li 0001, Anup Basu |
Pattern Recognit. Lett. | 2 |
| 1990 | Approximate constrained motion planningabstractThe problem of finding a collision-free path connecting two points (start and goal) in the presence of obstacles, with constraints on the curvature of the path, is examined. This problem of curvature-constrained motion planning arises when, for example, a vehicle with constraints on its steering mechanism needs to be maneuvered through obstacles. Though no lower bound on the difficulty of the problem in 2-D is known, exact algorithms given to date for the reachability questions are exponential. It is shown that a variation of the problem is NP-hard. Notably, however, the same variation to polynomially solvable motion planning problems does not make them intractable. In addition, it is proven that epsilon -approximations to this problem cannot exist unless the underlying decision problem is polynomially solvable. An algorithm which is expected to find a desired path, when one exists, with a required probability is presented. Results indicate that a variable-size discretization is necessary for the task, linking the required probability to the size of the discretization locally.> Anup Basu, Yiannis Aloimonos |
ICRA | 1 |
| 1987 | A Robust Algorithm for Determining the Translation of a Rigidly Moving Surface without Correspondence, for Robotics Applications
Anup Basu, Yiannis Aloimonos |
IJCAI | 1 |
| 1987 | Algorithms and hardware for efficient image smoothing
Anup Basu |
Comput. Vis. Graph. Image Process. | 1 |
| 1987 | Texture, contour, shape, and motion
Yiannis Aloimonos, Mike Swain, Paul B. Chou, Anup Basu |
Pattern Recognit. Lett. | 5 |