Ghassan Al-Regib

dblp:83/1655 · also Ghassan AlRegib, Ghassan Alregib · DBLP profile ↗
← Back
146ranked-venue papers
14as first author
27since 2021 · last 2026
0000-0001-6818-8001ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 116 · 12 first-author · 16 since 2021Artificial intelligence and machine learning · 14 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 8 since 2021Computer networks · 10 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Countering Multi-modal Representation Collapse through Rank-targeted Fusion
abstract
Multi-modal fusion methods often suffer from two types of representation collapse: feature collapse where individual dimensions lose their discriminative power (as measured by eigenspectra), and modality collapse where one dominant modality overwhelms the other. Applications like human action anticipation that require fusing multifarious sensor data are hindered by both feature and modality collapse. However, existing methods attempt to counter feature collapse and modality collapse separately. This is because there is no unifying framework that efficiently addresses feature and modality collapse in conjunction. In this paper, we posit the utility of effective rank as an informative measure that can be utilized to quantify and counter both the representation collapses. We propose Rank-enhancing Token Fuser, a theoretically grounded fusion framework that selectively blends less informative features from one modality with complementary features from another modality. We show that our method increases the effective rank of the fused representation. To address modality collapse, we evaluate modality combinations that mutually increase each others’ effective rank. We show that depth maintains representational balance when fused with RGB, avoiding modality collapse. We validate our method on action anticipation, where we present R3D, a depth-informed fusion framework. Extensive experiments on NTURGBD, UTKinect, and DARai demonstrate that our approach significantly outperforms prior state-of-the-art methods by up to 3.74%. Our code is available at: https://github.com/olivesgatech/R3D.
Seulgi Kim, Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan Al-Regib
WACV4
2025 Multi-Level and Multi-Modal Action Anticipation
abstract
Action anticipation, the task of predicting future actions from partially observed videos, is crucial for advancing intelligent systems. Unlike action recognition, which operates on fully observed videos, action anticipation must handle incomplete information. Hence, it requires temporal reasoning, and inherent uncertainty handling. While recent advances have been made, traditional methods often focus solely on visual modalities, neglecting the potential of integrating multiple sources of information. Drawing inspiration from human behavior, we introduce Multi-level and Multi-modal Action Anticipation (m&m-Ant), a novel multi-modal action anticipation approach that combines both visual and textual cues, while explicitly modeling hierarchical semantic information for more accurate predictions. To address the challenge of inaccurate coarse action labels, we propose a fine-grained label generator paired with a specialized temporal consistency loss function to optimize performance. Extensive experiments on widely used datasets, including Breakfast, 50 Salads, and DARai, demonstrate the effectiveness of our approach, achieving state-of-the-art results with an average anticipation accuracy improvement of 3.08% over existing methods. This work underscores the potential of multi-modal and hierarchical modeling in advancing action anticipation and establishes a new benchmark for future research in the field. Our code is available at: https://github.com/olivesgatech/mM-ant.
Seulgi Kim, Ghazal Kaviani, Mohit Prabhushankar, Ghassan Al-Regib
ICIP4
2025 HEX: Hierarchical Emergence Exploitation in Self-Supervised Algorithms
abstract
In this paper, we propose an algorithm that can be used on top of a wide variety of self-supervised (SSL) approaches to take advantage of hierarchical structures that emerge during training. SSL approaches typically work through some invariance term to ensure consistency between similar samples and a regularization term to prevent global dimensional collapse. Dimensional collapse refers to data representations spanning a lower-dimensional subspace. Recent work has demonstrated that the representation space of these algorithms gradually reflects a semantic hierarchical structure as training progresses. Data samples of the same hierarchical grouping tend to exhibit greater dimensional collapse locally compared to the dataset as a whole due to sharing features in common with each other. Ideally, SSL algorithms would take advantage of this hierarchical emergence to have an additional regularization term to account for this local dimensional collapse effect. However, the construction of existing SSL algorithms does not account for this property. To address this, we propose an adaptive algorithm that performs a weighted decomposition of the denominator of the InfoNCE loss into two terms: local hierarchical and global collapse regularization respectively. This decomposition is based on an adaptive threshold that gradually lowers to reflect the emerging hierarchical structure of the representation space throughout training. It is based on an analysis of the cosine similarity distribution of samples in a batch. We demonstrate that this hierarchical emergence exploitation (HEX) approach can be integrated across a wide variety of SSL algorithms. Empirically, we show performance improvements of up to 5.6% relative improvement over baseline SSL approaches on classification accuracy on Imagenet with 100 epochs of training.
Kiran Kokilepersaud, Seulgi Kim, Mohit Prabhushankar, Ghassan Al-Regib
WACV4
2025 A Gating Model for Bias Calibration in Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning (GZSL) aims at training a model that can generalize to unseen class data by only using auxiliary information. One of the main challenges in GZSL is a biased model prediction toward seen classes caused by overfitting on only available seen class data during training. To overcome this issue, we propose a two-stream autoencoder-based gating model for GZSL. Our gating model predicts whether the query data is from seen classes or unseen classes, and utilizes separate seen and unseen experts to predict the class independently from each other. This framework avoids comparing the biased prediction scores for seen classes with the prediction scores for unseen classes. In particular, we measure the distance between visual and attribute representations in the latent space and the cross-reconstruction space of the autoencoder. These distances are utilized as complementary features to characterize unseen classes at different levels of data abstraction. Also, the two-stream autoencoder works as a unified framework for the gating model and the unseen expert, which makes the proposed method computationally efficient. We validate our proposed method in four benchmark image recognition datasets. In comparison with other state-of-the-art methods, we achieve the best harmonic mean accuracy in SUN and AWA2, and the second best in CUB and AWA1. Furthermore, our base model requires at least 20% less number of model parameters than state-of-the-art methods relying on generative models.
Gukyeong Kwon, Ghassan Al-Regib
IEEE Trans. Image Process.2
2024 Targeting Negative Flips in Active Learning using Validation Sets
abstract
The performance of active learning algorithms can be improved in two ways. The often used and intuitive way is by reducing the overall error rate within the test set. The second way is to ensure that correct predictions are not forgotten when the training set is increased in between rounds. The former is measured by the accuracy of the model and the latter is captured in negative flips between rounds. Negative flips are samples that are correctly predicted when trained with the previous/smaller dataset and incorrectly predicted after additional samples are labeled. In this paper, we discuss improving the performance of active learning algorithms both in terms of prediction accuracy and negative flips. The first observation we make in this paper is that negative flips and overall error rates are decoupled and reducing one does not necessarily imply that the other is reduced. Our observation is important as current active learning algorithms do not consider negative flips directly and implicitly assume the opposite. The second observation is that performing targeted active learning on subsets of the unlabeled pool has a significant impact on the behavior of the active learning algorithm and influences both negative flips and prediction accuracy. We then develop ROSE - a plug-in algorithm that utilizes a small labeled validation set to restrict arbitrary active learning acquisition functions to negative flips within the unlabeled pool. We show that integrating a validation set results in a significant performance boost in terms of accuracy, negative flip rate reduction, or both.
Ryan Benkert, Mohit Prabhushankar, Ghassan Al-Regib
IEEE Big Data3
2024 Benchmarking Human and Automated Prompting in the Segment Anything Model
abstract
The remarkable capabilities of the Segment Anything Model (SAM) for tackling image segmentation tasks in an intuitive and interactive manner has sparked interest in the design of effective visual prompts. Such interest has led to the creation of automated point prompt selection strategies, typically motivated from a feature extraction perspective. However, there is still very little understanding of how appropriate these automated visual prompting strategies are, particularly when compared to humans, across diverse image domains. Additionally, the performance benefits of including such automated visual prompting strategies within the finetuning process of SAM also remains unexplored, as does the effect of interpretable factors like distance between the prompt points on segmentation performance. To bridge these gaps, we leverage a recently released visual prompting dataset, PointPrompt, and introduce a number of benchmarking tasks that provide an array of opportunities to improve the understanding of the way human prompts differ from automated ones and what underlying factors make for effective visual prompts. We demonstrate that the resulting segmentation scores obtained by humans are approximately 29% higher than those given by automated strategies and identify potential features that are indicative of prompting performance with R2scores over 0.5. Additionally, we demonstrate that performance when using automated methods can be improved by up to 68% via a finetuning approach. Overall, our experiments not only showcase the existing gap between human prompts and automated methods, but also highlight potential avenues through which this gap can be leveraged to improve effective visual prompt design. Further details along with the dataset links and codes are available at https://github.com/olivesgatech/PointPrompt
Jorge Quesada, Zoe Fowler, Mohammad Alotaibi, Mohit Prabhushankar, Ghassan Al-Regib
IEEE Big Data5
2024 Are Objective Explanatory Evaluation Metrics Trustworthy? An Adversarial Analysis
abstract
Explainable AI (XAI) has revolutionized the field of deep learning by empowering users to have more trust in neural network models. The field of XAI allows users to probe the inner workings of these algorithms to elucidate their decision-making processes. The rise in popularity of XAI has led to the advent of different strategies to produce explanations, all of which only occasionally agree. Thus several objective evaluation metrics have been devised to decide which of these modules give the best explanation for specific scenarios. The goal of the paper is twofold: (i) we employ the notions of necessity and sufficiency from causal literature to come up with a novel explanatory technique called SHifted Adversaries using Pixel Elimination(SHAPE) which satisfies all the theoretical and mathematical criteria of being a valid explanation, (ii) we show that SHAPE is, infact, an adversarial explanation that fools causal metrics that are employed to measure the robustness and reliability of popular importance based visual XAI methods. Our analysis shows that SHAPE outperforms popular explanatory techniques like GradCAM and GradCAM++ in these tests and is comparable to RISE, raising questions about the sanity of these metrics and the need for human involvement for an overall better evaluation.
Prithwijit Chowdhury, Mohit Prabhushankar, Ghassan Al-Regib, Mohamed Deriche 0001
ICIP3
2024 Taxes are All You Need: Integration Of Taxonomical Hierarchy Relationships Into the Contrastive Loss
abstract
In this work, we propose a novel supervised contrastive loss that enables the integration of taxonomic hierarchy information during the representation learning process. A supervised contrastive loss operates by enforcing that images with the same class label (positive samples) project closer to each other than images with differing class labels (negative samples). The advantage of this approach is that it directly penalizes the structure of the representation space itself. This enables greater flexibility with respect to encoding semantic concepts. However, the standard supervised contrastive loss only enforces semantic structure based on the downstream task (i.e. the class label). In reality, the class label is only one level of a hierarchy of different semantic relationships known as a taxonomy. For example, the class label is oftentimes the species of an animal, but between different classes there are higher order relationships such as all animals with wings being “birds”. We show that by explicitly accounting for these relationships with a weighting penalty in the contrastive loss we can out-perform the supervised contrastive loss. Additionally, we demonstrate the adaptability of the notion of a taxonomy by integrating our loss into medical and noise-based settings that show performance improvements by as much as 7%.
Kiran Kokilepersaud, Yavuz Yarici, Mohit Prabhushankar, Ghassan Al-Regib
ICIP4
2024 Intelligent Multi-View Test Time Augmentation
abstract
In this study, we introduce an intelligent Test Time Augmentation (TTA) algorithm designed to enhance the robustness and accuracy of image classification models against viewpoint variations. Unlike traditional TTA methods that indiscriminately apply augmentations, our approach intelligently selects optimal augmentations based on predictive uncertainty metrics. This selection is achieved via a two-stage process: the first stage identifies the optimal augmentation for each class by evaluating uncertainty levels, while the second stage implements an uncertainty threshold to determine when applying TTA would be advantageous. This methodological advancement ensures that augmentations contribute to classification more effectively than a uniform application across the dataset. Experimental validation across several datasets and neural network architectures validates our approach, yielding an average accuracy improvement of 1.73% over methods that use single-view images. This research underscores the potential of adaptive, uncertainty-aware TTA in improving the robustness of image classification in the presence of viewpoint variations, paving the way for further exploration into intelligent augmentation strategies. The code is available at: https://github.com/olivesgatech/Intelligent-Multi-View-TTA
Efe Ozturk, Mohit Prabhushankar, Ghassan Al-Regib
ICIP3
2024 Explaining Representation Learning With Perceptual Components
abstract
Self-supervised models create representation spaces that lack clear semantic meaning. This interpretability problem of representations makes traditional explainability methods ineffective in this context. In this paper, we introduce a novel method to analyze representation spaces using three key perceptual components: color, shape, and texture. We employ selective masking of these components to observe changes in representations, resulting in distinct importance maps for each. In scenarios, where labels are absent, these importance maps provide more intuitive explanations as they are integral to the human visual system. Our approach enhances the interpretability of the representation space, offering explanations that resonate with human visual perception. We analyze how different training objectives create distinct representation spaces using perceptual components. Additionally, we examine the representation of images across diverse image domains, providing insights into the role of these components in different contexts.
Yavuz Yarici, Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan Al-Regib
ICIP4
2024 Transitional Uncertainty with Layered Intermediate Predictions
abstract
In this paper, we discuss feature engineering for single-pass uncertainty estimation. For accurate uncertainty estimates, neural networks must extract differences in the feature space that quantify uncertainty. This could be achieved by current single-pass approaches that maintain feature distances between data points as they traverse the network. While initial results are promising, maintaining feature distances within the network representations frequently inhibits information compression and opposes the learning objective. We study this effect theoretically and empirically to arrive at a simple conclusion: preserving feature distances in the output is beneficial when the preserved features contribute to learning the label distribution and act in opposition otherwise. We then propose Transitional Uncertainty with Layered Intermediate Predictions (TULIP) as a simple approach to address the shortcomings of current single-pass estimators. Specifically, we implement feature preservation by extracting features from intermediate representations before information is collapsed by subsequent layers. We refer to the underlying preservation mechanism as transitional feature preservation. We show that TULIP matches or outperforms current single-pass methods on standard benchmarks and in practical settings where these methods are less reliable (imbalances, complex architectures, medical modalities).
Ryan Benkert, Mohit Prabhushankar, Ghassan Al-Regib
ICML3
2024 Effective Data Selection for Seismic Interpretation Through Disagreement
abstract
This article presents a discussion on data selection for deep learning in the field of seismic interpretation. To enable robust generalization, the training set must contain informative samples for the interpretation process. The selection of the training set from a target volume is a critical factor in determining the effectiveness of the deep learning algorithm for interpreting seismic volumes. This article proposes the inclusion of interpretation disagreement as a valuable and intuitive factor in the process of selecting training sets. The development of a novel data selection framework is inspired by established practices in seismic interpretation. The framework we have developed utilizes representation shifts to effectively model interpretation disagreement within neural networks. Additionally, it incorporates the disagreement measure to enhance attention toward geologically interesting regions throughout the data selection workflow. By combining this approach with active learning, a well-known machine learning paradigm for data selection, we arrive at a comprehensive and innovative framework for training set selection in seismic interpretation. In addition, we offer a specific implementation of our proposed framework, which we have named active transfer learning for attention sensitivity (ATLAS). This implementation serves as a means for data selection. In this study, we present the results of our comprehensive experiments, which clearly indicate thatATLASconsistently surpasses traditional active learning frameworks in the field of seismic interpretation. Our findings reveal thatATLASachieves improvements of up to 12% in mean intersection-over-union.
Ryan Benkert, Mohit Prabhushankar, Ghassan Al-Regib
IEEE Trans. Geosci. Remote. Sens.3
2024 Visual Attention-Guided Learning With Incomplete Labels for Seismic Fault Interpretation
abstract
Annotating geological faults on three dimensional seismic volumes is a laborious process. Typically, only a fraction of the actual faults are manually interpreted, leaving many others unlabeled. This is due to the way attention selectivity works to drive human perception. The human brain selectively focuses its attention to certain salient regions in a visual scene marked by prominent changes in color, contrast, and other low level signal cues. This bottom-up attention is further modulated by the individual’s goals, expectations, and constraints with respect to the task at hand, also called top-down attention. The fault annotations created by seismic interpreters reflect this cognitive process comprising of both bottom-up and top-down attentional mechanisms. 3D convolutional neural networks pretrained on synthetic seismic data for fault mapping can be finetuned on select seismic lines extracted and labeled on a real seismic volume of interest. Traditional finetuning approaches treat all pixels on labeled sections as the absolute ground-truth. This leads to the network incorrectly learning to predict regions of missing fault labels as negatives. We propose an attention-guided training framework that models and incorporates human visual attention to (1) condition the process of sampling training data and (2) modulate the loss value for each pixel. Through quantitative and qualitative evaluation of results on a real seismic volume from North Western Australia, we demonstrate that the proposed approach is able to predict both the annotated as well as the unlabeled faults significantly better compared to baseline approaches.
Ahmad Mustafa 0002, Reza Rastegar, Tim Brown, Gregory Nunes, Daniel Delilla, Ghassan Al-Regib
IEEE Trans. Geosci. Remote. Sens.6
2024 TrajPRed: Trajectory Prediction With Region-Based Relation Learning
abstract
Forecasting human trajectories in traffic scenes is critical for safety within mixed or fully autonomous systems. Human future trajectories are driven by two major stimuli, social interactions, and stochastic goals. Thus, reliable forecasting needs to capture these two stimuli. Edge-based relation modeling represents social interactions using pairwise correlations from precise individual states. Nevertheless, edge-based relations can be vulnerable under perturbations. To alleviate these issues, we propose a region-based relation learning paradigm that models social interactions via region-wise dynamics of joint states, i.e., the changes in the density of crowds. In particular, region-wise agent joint information is encoded within convolutional feature grids. Social relations are modeled by relating the temporal changes of local joint information from a global perspective. We show that region-based relations are less susceptible to perturbations. In order to account for the stochastic individual goals, we exploit a conditional variational autoencoder to realize multi-goal estimation and diverse future prediction. Specifically, we perform variational inference via the latent distribution, which is conditioned on the correlation between input states and associated target goals. Sampling from the latent distribution enables the framework to reliably capture the stochastic behavior in test data. We integrate multi-goal estimation and region-based relation learning to model the two stimuli, social interactions, and stochastic goals, in a prediction framework. We evaluate our framework on the ETH-UCY dataset and Stanford Drone Dataset (SDD). We show that diverse prediction benefits from region-based relation learning. The predicted intermediate location distributions better fit the ground truth when incorporating the relation module. Our framework outperforms the state-of-the-art models on SDD by$27.61\%$/$18.20\%$of ADE/FDE metrics. Our code is available at https://github.com/olivesgatech/TrajPRed
Ghassan Al-Regib, Armin Parchami, Kunjan Singh
IEEE Trans. Intell. Transp. Syst.2
2023 FOCAL: A Cost-Aware Video Dataset for Active Learning
abstract
In this paper, we introduce the FOCAL (Ford-OLIVES Collaboration on Active Learning) dataset which enables the study of the impact of annotation-cost within a video active learning setting. Annotation-cost refers to the time it takes an annotator to label and quality-assure a given video sequence. A practical motivation for active learning research is to minimize annotation-cost by selectively labeling informative samples that will maximize performance within a given budget constraint. However, previous work in video active learning lacks real-time annotation labels for accurately assessing cost minimization and instead operates under the assumption that annotation-cost scales linearly with the amount of data to annotate. This assumption does not take into account a variety of real-world confounding factors that contribute to a nonlinear cost such as the effect of an assistive labeling tool and the variety of interactions within a scene such as occluded objects, weather, and motion of objects. FOChL addresses this discrepancy by providing real annotation-cost labels for 126 video sequences across 69 unique city scenes with a variety of weather, lighting, and seasonal conditions. These videos have a wide range of interactions that are at the intersection of infrastructure-assisted autonomy and autonomous vehicle communities. We show through a statistical analysis of the FOChL dataset that cost is more correlated with a variety of factors beyond just the length of a video sequence. We also introduce a set of conformal active learning algorithms that take advantage of the sequential structure of video data in order to achieve a better trade-off between annotation-cost and performance while also reducing floating point operations (FLOPS) overhead by at least 77.67%. We show how these approaches better reflect how annotations on videos are done in practice through a sequence selection framework. We further demonstrate the advantage of these approaches by introducing two performance-cost metrics and show that the best conformal active learning method is cheaper than the best traditional active learning method by 113 hours. The code associated with this paper can be found at this link. The data can be downloaded at this location.
Kiran Kokilepersaud, Yash-Yee Logan, Ryan Benkert, Mohit Prabhushankar, Ghassan Al-Regib, Enrique Corona, Kunjan Singh, Mostafa Parchami
IEEE Big Data6
2023 Clinically Labeled Contrastive Learning for OCT Biomarker Classification
abstract
This article presents a novel positive and negative set selection strategy for contrastive learning of medical images based on labels that can be extracted from clinical data. In the medical field, there exists a variety of labels for data that serve different purposes at different stages of a diagnostic and treatment process. Clinical labels and biomarker labels are two examples. In general, clinical labels are easier to obtain in larger quantities because they are regularly collected during routine clinical care, while biomarker labels require expert analysis and interpretation to obtain. Within the field of ophthalmology, previous work has shown that clinical values exhibit correlations with biomarker structures that manifest within optical coherence tomography (OCT) scans. We exploit this relationship by using the clinical data as pseudo-labels for our data without biomarker labels in order to choose positive and negative instances for training a backbone network with a supervised contrastive loss. In this way, a backbone network learns a representation space that aligns with the clinical data distribution available. Afterwards, we fine-tune the network trained in this manner with the smaller amount of biomarker labeled data with a cross-entropy loss in order to classify these key indicators of disease directly from OCT scans. We also expand on this concept by proposing a method that uses a linear combination of clinical contrastive losses. We benchmark our methods against state of the art self-supervised methods in a novel setting with biomarkers of varying granularity. We show performance improvements by as much as 5% in total biomarker detection AUROC.
Kiran Kokilepersaud, Stephanie Trejo Corona, Mohit Prabhushankar, Ghassan Al-Regib, Charles C. Wykoff
IEEE J. Biomed. Health Informatics4
2022 Forgetful Active Learning with Switch Events: Efficient Sampling for Out-of-Distribution Data
abstract
This paper considers deep out-of-distribution active learning. In practice, fully trained neural networks interact randomly with out-of-distribution (OOD) inputs and map aberrant samples randomly within the model representation space. Since data representations are direct manifestations of the training distribution, the data selection process plays a crucial role in outlier robustness. For paradigms such as active learning, this is especially challenging since protocols must not only improve performance on the training distribution most effectively but further render a robust representation space. However, existing strategies directly base the data selection on the data representation of the unlabeled data which is random for OOD samples by definition. For this purpose, we introduce forgetful active learning with switch events (FALSE) - a novel active learning protocol for out-of-distribution active learning. Instead of defining sample importance on the data representation directly, we formulate "informativeness" with learning difficulty during training. Specifically, we approximate how often the network "forgets" unlabeled samples and query the most "forgotten" samples for annotation. We report up to 4.5% accuracy improvements in over 270 experiments, including four commonly used protocols, two OOD benchmarks, one in-distribution benchmark, and three different architectures.
Ryan Benkert, Mohit Prabhushankar, Ghassan Al-Regib
ICIP3
2022 Gradient-Based Severity Labeling for Biomarker Classification in Oct
abstract
In this paper, we propose a novel selection strategy for contrastive learning for medical images. On natural images, contrastive learning uses augmentations to select positive and negative pairs for the contrastive loss. However, in the medical domain, arbitrary augmentations have the potential to distort small localized regions that contain the biomarkers we are interested in detecting. A more intuitive approach is to select samples with similar disease severity characteristics, since these samples are more likely to have similar structures related to the progression of a disease. To enable this, we introduce a method that generates disease severity labels for unlabeled OCT scans on the basis of gradient responses from an anomaly detection algorithm. These labels are used to train a supervised contrastive learning setup to improve biomarker classification accuracy by as much as 6% above self-supervised baselines for key indicators of Diabetic Retinopathy.
Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan Al-Regib, Stephanie Trejo Corona, Charles C. Wykoff
ICIP3
2022 Patient Aware Active Learning for Fine-Grained OCT Classification
abstract
This paper considers making active learning more sensible from a medical perspective. In practice, a disease manifests itself in different forms across patient cohorts. Existing frameworks have primarily used mathematical constructs to engineer uncertainty or diversity-based methods for selecting the most informative samples. However, such algorithms do not present themselves naturally as usable by the medical community and healthcare providers. Thus, their deployment in clinical settings is very limited, if any. For this purpose, we propose a framework that incorporates clinical insights into the sample selection process of active learning that can be incorporated with existing algorithms. Our medically interpretable active learning framework captures diverse disease manifestations from patients to improve generalization performance of OCT classification. After comprehensive experiments, we report that incorporating patient insights within the active learning framework yields performance that matches or surpasses five commonly used paradigms on two architectures with a dataset having imbalanced patient distributions. Also, the framework integrates within existing medical practices and thus can be used by healthcare providers.
Yash-Yee Logan, Ryan Benkert, Ahmad Mustafa 0002, Ghassan Al-Regib
ICIP4
2022 Learning Trajectory-Conditioned Relations to Predict Pedestrian Crossing Behavior
abstract
In smart transportation, intelligent systems avoid potential collisions by predicting the intent of traffic agents, especially pedestrians. Pedestrian intent, defined as future action, e.g., start crossing, can be dependent on traffic surroundings. In this paper, we develop a framework to incorporate such dependency given observed pedestrian trajectory and scene frames. Our framework first encodes regional joint information between a pedestrian and surroundings over time into feature-map vectors. The global relation representations are then extracted from pairwise feature-map vectors to estimate intent with past trajectory condition. We evaluate our approach on two public datasets and compare against two state-of-the-art approaches. The experimental results demonstrate that our method helps to inform potential risks during crossing events with 0.04 improvement in F1-score on JAAD dataset and 0.01 improvement in recall on PIE dataset. Furthermore, we conduct ablation experiments to confirm the contribution of the relation extraction in our framework.
Ghassan Al-Regib, Armin Parchami, Kunjan Singh
ICIP2
2022 Introspective Learning : A Two-Stage approach for Inference in Neural Networks
abstract
In this paper, we advocate for two stages in a neural network's decision making process. The first is the existing feed-forward inference framework where patterns in given data are sensed and associated with previously learned patterns. The second stage is a slower reflection stage where we ask the network to reflect on its feed-forward decision by considering and evaluating all available choices. Together, we term the two stages as introspective learning. We use gradients of trained neural networks as a measurement of this reflection. A simple three-layered Multi Layer Perceptron is used as the second stage that predicts based on all extracted gradient features. We perceptually visualize the post-hoc explanations from both stages to provide a visual grounding to introspection. For the application of recognition, we show that an introspective network is 4% more robust and 42% less prone to calibration errors when generalizing to noisy data. We also illustrate the value of introspective networks in downstream tasks that require generalizability and calibration including active learning, out-of-distribution detection, and uncertainty estimation. Finally, we ground the proposed machine introspection to human introspection for the application of image quality assessment.
Mohit Prabhushankar, Ghassan Al-Regib
NeurIPS2
2022 OLIVES Dataset: Ophthalmic Labels for Investigating Visual Eye Semantics
abstract
Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans. Clinical practitioners use all available data modalities for diagnosing and treating eye diseases like Diabetic Retinopathy (DR) or Diabetic Macular Edema (DME). Enabling usage of machine learning algorithms within the ophthalmic medical domain requires research into the relationships and interactions between all relevant data over a treatment period. Existing datasets are limited in that they neither provide data nor consider the explicit relationship modeling between the data modalities. In this paper, we introduce the Ophthalmic Labels for Investigating Visual Eye Semantics (OLIVES) dataset that addresses the above limitation. This is the first OCT and near-IR fundus dataset that includes clinical labels, biomarker labels, disease labels, and time-series patient treatment information from associated clinical trials. The dataset consists of 1268 near-IR fundus images each with at least 49 OCT scans, and 16 biomarkers, along with 4 clinical labels and a disease diagnosis of DR or DME. In total, there are 96 eyes' data averaged over a period of at least two years with each eye treated for an average of 66 weeks and 7 injections. We benchmark the utility of OLIVES dataset for ophthalmic data as well as provide benchmarks and concrete research directions for core and emerging machine learning paradigms within medical image analysis.
Mohit Prabhushankar, Kiran Kokilepersaud, Yash-Yee Logan, Stephanie Trejo Corona, Ghassan Al-Regib, Charles C. Wykoff
NeurIPS5
2022 Example Forgetting: A Novel Approach to Explain and Interpret Deep Neural Networks in Seismic Interpretation
abstract
In recent years, deep neural networks have significantly impacted the seismic interpretation process. Due to the simple implementation and low interpretation costs, deep neural networks are an attractive component for the common interpretation pipeline. However, neural networks are frequently met with distrust due to their property of producing semantically incorrect outputs when exposed to sections the model was not trained on. We address this issue by explaining model behaviour and improving generalization properties through example forgetting: First, we introduce a method that effectively relates semantically malfunctioned predictions to their respectful positions within the neural network representation manifold. More concrete, our method tracks how models ”forget” seismic reflections during training and establishes a connection to the decision boundary proximity of the target class. Second, we use our analysis technique to identify frequently forgotten regions within the training volume and augment the training set with state-of-the-art style transfer techniques from computer vision. We show that our method improves the segmentation performance on underrepresented classes while significantly reducing the forgotten regions in the F3 volume in the Netherlands.
Ryan Benkert, Oluwaseun Joseph Aribido, Ghassan Al-Regib
IEEE Trans. Geosci. Remote. Sens.3
2021 Explaining Deep Models Through Forgettable Learning Dynamics
abstract
Even though deep neural networks have shown tremendous success in countless applications, explaining model behaviour or predictions is an open research problem. In this paper, we address this issue by employing a simple yet effective method by analysing the learning dynamics of deep neural networks in semantic segmentation tasks. Specifically, we visualize the learning behaviour during training by tracking how often samples are learned and forgotten in subsequent training epochs. This further allows us to derive important information about the proximity to the class decision boundary and identify regions that pose a particular challenge to the model. Inspired by this phenomenon, we present a novel segmentation method that actively uses this information to alter the data representation within the model by increasing the variety of difficult regions. Finally, we show that our method consistently reduces the amount of regions that are forgotten frequently. We further evaluate our method in light of the segmentation performance.
Ryan Benkert, Oluwaseun Joseph Aribido, Ghassan Al-Regib
ICIP3
2021 Open-Set Recognition With Gradient-Based Representations
abstract
Neural networks for image classification tasks assume that any given image during inference belongs to one of the training classes. This closed-set assumption is challenged in real-world applications where models may encounter inputs of unknown classes. Open-set recognition aims to solve this problem by rejecting unknown classes while classifying known classes correctly. In this paper, we propose to utilize gradient-based representations obtained from a known classifier to train an unknown detector with instances of known classes only. Gradients correspond to the amount of model updates required to properly represent a given sample, which we exploit to understand the model’s capability to characterize inputs with its learned features. Our approach can be utilized with any classifier trained in a supervised manner on known classes without the need to model the distribution of unknown samples explicitly. We show that our gradient-based approach outperforms state-of-the-art methods by up to 11.6% in open-set classification.
Jinsol Lee, Ghassan Al-Regib
ICIP2
2021 Man-Recon: Manifold Learning For Reconstruction With Deep Autoencoder For Smart Seismic Interpretation
abstract
Deep learning can extract rich data representations if provided sufficient quantities of labeled training data. For many tasks however, annotating data has significant costs in terms of time and money, owing to the high standards of subject matter expertise required, for example in medical and geophysical image interpretation tasks. Active learning can identify the most informative training examples for the interpreter to train, leading to higher efficiency. We propose an active learning method based on jointly learning representations for supervised and unsupervised tasks. The learned manifold structure is later utilized to identify informative training samples most dissimilar from the learned manifold from the error profiles on the unsupervised task. We verify the efficiency of the proposed method on a seismic facies segmentation dataset from the Netherlands F3 block survey, significantly outperforming contemporary methods to achieve the highest mean Intersection-Over-Union value of 0.773.
Ahmad Mustafa 0002, Ghassan Al-Regib
ICIP2
2021 Extracting Causal Visual Features For Limited Label Classification
abstract
Neural networks trained to classify images do so by identifying features that allow them to distinguish between classes. These sets of features are either causal or context dependent. Grad-CAM is a popular method of visualizing both sets of features. In this paper, we formalize this feature divide and provide a methodology to extract causal features from Grad-CAM. We do so by defining context features as those features that allow contrast between predicted class and any contrast class. We then apply a set theoretic approach to separate causal from contrast features for COVID-19 CT scans. We show that on average, the image regions with the proposed causal features require 15% less bits when encoded using Huffman encoding, compared to Grad-CAM, for an average increase of 3% classification accuracy, over Grad-CAM. Moreover, we validate the transfer-ability of causal features between networks and comment on the non-human interpretable causal nature of current networks.
Mohit Prabhushankar, Ghassan Al-Regib
ICIP2
2020 Action Segmentation With Joint Self-Supervised Temporal Domain Adaptation
abstract
Despite the recent progress of fully-supervised action segmentation techniques, the performance is still not fully satisfactory. One main challenge is the problem of spatiotemporal variations (e.g. different people may perform the same activity in various ways). Therefore, we exploit unlabeled videos to address this problem by reformulating the action segmentation task as a cross-domain problem with domain discrepancy caused by spatio-temporal variations. To reduce the discrepancy, we propose SelfSupervised Temporal Domain Adaptation (SSTDA), which contains two self-supervised auxiliary tasks (binary and sequential domain prediction) to jointly align cross-domain feature spaces embedded with local and global temporal dynamics, achieving better performance than other Domain Adaptation (DA) approaches. On three challenging benchmark datasets (GTEA, 50Salads, and Breakfast), SSTDA outperforms the current state-of-the-art method by large margins (e.g. for the F1@25 score, from 59.6% to 69.1% on Breakfast, from 73.4% to 81.5% on 50Salads, and from 83.6% to 89.1% on GTEA), and requires only 65% of the labeled training data for comparable performance, demonstrating the usefulness of adapting to unlabeled target videos across variations. The source code is available at https://github.com/cmhungsteve/SSTDA.
Min-Hung Chen, Baopu Li, Sid Ying-Ze Bao, Ghassan Al-Regib, Zsolt Kira
CVPR4
2020 Backpropagated Gradient Representations for Anomaly Detection
Gukyeong Kwon, Mohit Prabhushankar, Dogancan Temel, Ghassan Al-Regib
ECCV (21)4
2020 Learning to Generate Grounded Visual Captions Without Localization Supervision
Chih-Yao Ma, Yannis Kalantidis, Ghassan Al-Regib, Peter Vajda, Marcus Rohrbach, Zsolt Kira
ECCV (18)3
2020 Multiple Events Detection In Seismic Structures Using A Novel U-Net Variant
abstract
Seismic data interpretation is a fundamental process in the pipeline of identifying hydrocarbon structural traps such as salt domes and faults. This process is highly demanding and challenging in terms of expert-knowledge, time, and efforts. The interpretation process becomes even more challenging when it comes to identifying multiple seismic events taking place simultaneously. In recent years, the technology trend has been directed towards the automation of seismic interpretation using advanced computational techniques and in particular deep learning (DL) networks. In this paper, we present our DL solution for concurrent salt domes and faults identification with very promising preliminary results obtained through applications to real world seismic data. The proposed workflow leads to excellent detection results even with small size training datasets. Furthermore, the resulting probability maps can be extended to even a larger number of structure types. Precisions of the order of more than 96% were obtained with real data when three types of seismic structures are present concurrently.
Mustafa Alfarhan, Mohamed Deriche 0001, Ahmed Maalej, Ghassan Al-Regib, Hasan Al-Marzouqi
ICIP4
2020 Self-Supervised Annotation of Seismic Images Using Latent Space Factorization
abstract
Annotating seismic data is expensive, laborious and subjective due to the number of years required for seismic interpreters to attain proficiency in interpretation. In this paper, we develop a framework to automate annotating pixels of a seismic image to delineate geological structural elements given image-level labels assigned to each image. Our framework factorizes the latent space of a deep encoder-decoder network by projecting the latent space to learned sub-spaces. Using constraints in the pixel space, the seismic image is further factorized to reveal confidence values on pixels associated with the geological element of interest. Details of the annotated image are provided for analysis and qualitative comparison is made with similar frameworks.
Oluwaseun Joseph Aribido, Ghassan Al-Regib, Mohamed Deriche 0001
ICIP2
2020 Novelty Detection Through Model-Based Characterization of Neural Networks
abstract
In this paper, we propose a model-based characterization of neural networks to detect novel input types and conditions. Novelty detection is crucial to identify abnormal inputs that can significantly degrade the performance of machine learning algorithms. Majority of existing studies have focused on activation-based representations to detect abnormal inputs, which limits the characterization of abnormality from a data perspective. However, a model perspective can also be informative in terms of the novelties and abnormalities. To articulate the significance of the model perspective in novelty detection, we utilize back-propagated gradients. We conduct a comprehensive analysis to compare the representation capability of gradients with that of activation and show that the gradients outperform the activation in novel class and condition detection. We validate our approach using four image recognition datasets including MNIST, Fashion-MNIST, CIFAR10, and CURE-TSR. We achieve a significant improvement on all four datasets with an average AUROC of 0.953, 0.918, 0.582, and 0.746, respectively.
Gukyeong Kwon, Mohit Prabhushankar, Dogancan Temel, Ghassan Al-Regib
ICIP4
2020 Gradients as a Measure of Uncertainty in Neural Networks
abstract
Despite tremendous success of modern neural networks, they are known to be overconfident even when the model encounters inputs with unfamiliar conditions. Detecting such inputs is vital to preventing models from making naive predictions that may jeopardize real-world applications of neural networks. In this paper, we address the challenging problem of devising a simple yet effective measure of uncertainty in deep neural networks. Specifically, we propose to utilize backpropagated gradients to quantify the uncertainty of trained models. Gradients depict the required amount of change for a model to properly represent given inputs, thus providing a valuable insight into how familiar and certain the model is regarding the inputs. We demonstrate the effectiveness of gradients as a measure of model uncertainty in applications of detecting unfamiliar inputs, including out-of-distribution and corrupted samples. We show that our gradient-based method outperforms state-of-the-art methods by up to 4.8% of AUROC score in out-of-distribution detection and 35.7% in corrupted input detection.
Jinsol Lee, Ghassan Al-Regib
ICIP2
2020 On the Structures of Representation for the Robustness of Semantic Segmentation to Input Corruption
abstract
Semantic segmentation is a scene understanding task at the heart of safety-critical applications where robustness to corrupted inputs is essential. Implicit Background Estimation (IBE) has demonstrated to be a promising technique to improve the robustness to out-of-distribution inputs for semantic segmentation models for little to no cost. In this paper, we provide analysis comparing the structures learned as a result of optimization objectives that use Softmax, IBE, and Sigmoid in order to improve understanding their relationship to robustness. As a result of this analysis, we propose combining Sigmoid with IBE (SCrIBE) to improve robustness. Finally, we demonstrate that SCrIBE exhibits superior segmentation performance aggregated across all corruptions and severity levels with a mIOU of 42.1 compared to both IBE 40.3 and the Softmax Baseline 37.5.
Charles Lehman, Dogancan Temel, Ghassan Al-Regib
ICIP3
2020 Robustness And Overfitting Behavior Of Implicit Background Models
abstract
In this paper, we examine the overfitting behavior of image classification models modified with Implicit Background Estimation (SCrIBE), which transforms them into weakly supervised segmentation models that provide spatial domain visualizations without affecting performance. Using the segmentation masks, we derive an overfit detection criterion that does not require testing labels. In addition, we assess the change in model performance, calibration, and segmentation masks after applying data augmentations as overfitting reduction measures and testing on various types of distorted images.
Shirley Liu, Charles Lehman, Ghassan Al-Regib
ICIP3
2020 Contrastive Explanations In Neural Networks
abstract
Visual explanations are logical arguments based on visual features that justify the predictions made by neural networks. Current modes of visual explanations answer questions of the form `Why P.7'. These Why questions operate under broad contexts thereby providing answers that are irrelevant in some cases. We propose to constrain these Why questions based on some context Q so that our explanations answer contrastive questions of the form `Why P, rather than Q.7'. In this paper, we formalize the structure of contrastive visual explanations for neural networks. We define contrast based on neural networks and propose a methodology to extract defined contrasts. We then use the extracted contrasts as a plug-in on top of existing `Why P?' techniques, specifically Grad-CAM. We demonstrate their value in analyzing both networks and data in applications of large-scale recognition, fine-grained recognition, subsurface seismic analysis, and image quality assessment.
Mohit Prabhushankar, Gukyeong Kwon, Dogancan Temel, Ghassan Al-Regib
ICIP4
2020 S6: Semi-Supervised Self-Supervised Semantic Segmentation
abstract
Semi-supervised learning provides a means to leverage unlabeled data when labels are expensive to obtain. In this work, we propose a constrained framework that better learns from unlabeled data. The proposed algorithm adds a self-supervised task, image reconstruction, to the target segmentation task. The extra reconstruction task improves the model's geometric reasoning about different textures in an image. It happens that this improvement is transferable from reconstruction to segmentation since they both share some common parts of the architecture. Such extra task allows the trained model to have richer representations and better geometric understanding. Our results show that the proposed constrained framework achieves an improvement in mean Intersection over Union by 18% over unconstrained one, using only 2% of labeled examples. The performance gain and reduction in amounts of labeled data is crucial for applications in which obtaining labels is expensive and labor intensive such as Biomedical Imaging and Seismic Interpretation.
Moamen Soliman, Charles Lehman, Ghassan Al-Regib
ICIP3
2020 Implicit Saliency In Deep Neural Networks
abstract
In this paper, we show that existing recognition and localization deep architectures, that have not been exposed to eye tracking data or any saliency datasets, are capable of predicting the human visual saliency. We term this as implicit saliency in deep neural networks. We calculate this implicit saliency using expectancy-mismatch hypothesis in an unsupervised fashion. Our experiments show that extracting saliency in this fashion provides comparable performance when measured against the state-of-art supervised algorithms. Additionally, the robustness outperforms those algorithms when we add large noise to the input images. Also, we show that semantic features contribute more than low-level features for human visual saliency detection. Based on these properties and performances, our proposed method greatly lowers the threshold for saliency detection in terms of required data and bridges the gap between human visual saliency and model saliency.
Mohit Prabhushankar, Ghassan Al-Regib
ICIP3
2020 Action Segmentation with Mixed Temporal Domain Adaptation
abstract
The main progress for action segmentation comes from densely-annotated data for fully-supervised learning. Since manual annotation for frame-level actions is time-consuming and challenging, we propose to exploit auxiliary unlabeled videos, which are much easier to obtain, by shaping this problem as a domain adaptation (DA) problem. Although various DA techniques have been proposed in recent years, most of them have been developed only for the spatial direction. Therefore, we propose Mixed Temporal Domain Adaptation (MTDA) to jointly align frame-and video-level embedded feature spaces across domains, and further integrate with the domain attention mechanism to focus on aligning the frame-level features with higher domain discrepancy, leading to more effective domain adaptation. Finally, we evaluate our proposed methods on three challenging datasets (GTEA, 50Salads, and Breakfast), and validate that MTDA outperforms the current state-of-the-art methods on all three datasets by large margins (e.g. 6.4% gain on F1@50 and 6.8% gain on the edit score for GTEA).
Min-Hung Chen, Baopu Li, Sid Ying-Ze Bao, Ghassan Al-Regib
WACV4
2020 Texture classification using block intensity and gradient difference (BIGD) descriptor
Yuting Hu 0001, Zhen Wang 0007, Ghassan Al-Regib
Signal Process. Image Commun.3
2020 Relative Afferent Pupillary Defect Screening Through Transfer Learning
abstract
Abnormalities in pupillary light reflex can indicate optic nerve disorders that may lead to permanent visual loss if not diagnosed in an early stage. In this study, we focus on relative afferent pupillary defect (RAPD), which is based on the difference between the reactions of the eyes when they are exposed to light stimuli. Incumbent RAPD assessment methods are based on subjective practices that can lead to unreliable measurements. To eliminate subjectivity and obtain reliable measurements, we introduced an automated framework to detect RAPD. For validation, we conducted a clinical study with lab-on-a-headset, which can perform automated light reflex test. In addition to benchmarking handcrafted algorithms, we proposed a transfer learning-based approach that transformed a deep learning-based generic object recognition algorithm into a pupil detector. Based on the conducted experiments, proposed algorithm RAPDNet can achieve a sensitivity and a specificity of 90.6% over 64 test cases in a balanced set, which corresponds to an AUC of 0.929 in ROC analysis. According to our benchmark with three handcrafted algorithms and nine performance metrics, RAPDNet outperforms all other algorithms in every performance category.
Dogancan Temel, Melvin J. Mathew, Ghassan Al-Regib, Yousuf M. Khalifa
IEEE J. Biomed. Health Informatics3
2020 Traffic Sign Detection Under Challenging Conditions: A Deeper Look into Performance Variations and Spectral Characteristics
abstract
Traffic signs are critical for maintaining the safety and efficiency of our roads. Therefore, we need to carefully assess the capabilities and limitations of automated traffic sign detection systems. Existing traffic sign datasets are limited in terms of type and severity of challenging conditions. Metadata corresponding to these conditions are unavailable and it is not possible to investigate the effect of a single factor because of the simultaneous changes in numerous conditions. To overcome the shortcomings in existing datasets, we introduced the CURE-TSD-Real dataset, which is based on simulated challenging conditions that correspond to adversaries that can occur in real-world environments and systems. We test the performance of two benchmark algorithms and show that severe conditions can result in an average performance degradation of 29% in precision and 68% in recall. We investigate the effect of challenging conditions through spectral analysis and show that the challenging conditions can lead to distinct magnitude spectrum characteristics. Moreover, we show that mean magnitude spectrum of changes in video sequences under challenging conditions can be an indicator of detection performance. The CURE-TSD-Real dataset is available online at https://github.com/olivesgatech/CURE-TSD.
Dogancan Temel, Min-Hung Chen, Ghassan Al-Regib
IEEE Trans. Intell. Transp. Syst.3
2019 The Regretful Agent: Heuristic-Aided Navigation Through Progress Estimation
abstract
As deep learning continues to make progress for challenging perception tasks, there is increased interest in combining vision, language, and decision-making. Specifically, the Vision and Language Navigation (VLN) task involves navigating to a goal purely from language instructions and visual information without explicit knowledge of the goal. Recent successful approaches have made in-roads in achieving good success rates for this task but rely on beam search, which thoroughly explores a large number of trajectories and is unrealistic for applications such as robotics. In this paper, inspired by the intuition of viewing the problem as search on a navigation graph, we propose to use a progress monitor developed in prior work as a learnable heuristic for search. We then propose two modules incorporated into an end-to-end architecture: 1) A learned mechanism to perform backtracking, which decides whether to continue moving forward or roll back to a previous state (Regret Module) and 2) A mechanism to help the agent decide which direction to go next by showing directions that are visited and their associated progress estimate (Progress Marker). Combined, the proposed approach significantly outperforms current state-of-the-art methods using greedy action selection, with 5% absolute improvement on the test server in success rates, and more importantly 8% on success rates normalized by the path length.
Chih-Yao Ma, Zuxuan Wu, Ghassan Al-Regib, Caiming Xiong, Zsolt Kira
CVPR3
2019 Temporal Attentive Alignment for Large-Scale Video Domain Adaptation
abstract
Although various image-based domain adaptation (DA) techniques have been proposed in recent years, domain shift in videos is still not well-explored. Most previous works only evaluate performance on small-scale datasets which are saturated. Therefore, we first propose two large-scale video DA datasets with much larger domain discrepancy: UCF-HMDB_full and Kinetics-Gameplay. Second, we investigate different DA integration methods for videos, and show that simultaneously aligning and learning temporal dynamics achieves effective alignment even without sophisticated DA methods. Finally, we propose Temporal Attentive Adversarial Adaptation Network (TA3N), which explicitly attends to the temporal dynamics using domain discrepancy for more effective domain alignment, achieving state-of-the-art performance on four video DA datasets (e.g. 7.9% accuracy gain over “Source only” from 73.9% to 81.8% on “HMDB → UCF”, and 10.3% gain on “Kinetics → Gameplay”). The code and data are released at http://github.com/cmhungsteve/TA3N.
Min-Hung Chen, Zsolt Kira, Ghassan Al-Regib, Jaekwon Yoo, Ruxin Chen
ICCV3
2019 Multi-Level Texture Encoding and Representation (Multer) Based on Deep Neural Networks
abstract
In this paper, we propose a multi-level texture encoding and representation network (MuLTER) for texture-related applications. Based on a multi-level pooling architecture, the MuLTER network simultaneously leverages low-and high-level features to maintain both texture details and spatial information. Such a pooling architecture involves few extra parameters and keeps feature dimensions fixed despite of the changes of image sizes. In comparison with state-of-the-art texture descriptors, the MuLTER network yields higher recognition accuracy on typical texture datasets such as MINC-2500 and GTOS-mobile with a discriminative and compact representation. In addition, we analyze the impact of combining features from different levels, which supports our claim that the fusion of multi-level features efficiently enhances recognition performance. Our source code will be published on GitHub (https://github.com/olivesgatech).
Yuting Hu 0001, Zhiling Long, Ghassan Al-Regib
ICIP3
2019 Distorted Representation Space Characterization Through Backpropagated Gradients
abstract
In this paper, we utilize weight gradients from backpropagation to characterize the representation space learned by deep learning algorithms. We demonstrate the utility of such gradients in applications including perceptual image quality assessment and out-of-distribution classification. The applications are chosen to validate the effectiveness of gradients as features when the test image distribution is distorted from the train image distribution. In both applications, the proposed gradient based features outperform activation features. In image quality assessment, the proposed approach is compared with other state of the art approaches and is generally the top performing method on TID 2013 and MULTI-LIVE databases in terms of accuracy, consistency, linearity, and monotonic behavior. Finally, we analyze the effect of regularization on gradients using CURE-TSR dataset for out-of-distribution classification.
Gukyeong Kwon, Mohit Prabhushankar, Dogancan Temel, Ghassan Al-Regib
ICIP4
2019 Implicit Background Estimation For Semantic Segmentation
abstract
Scene understanding and semantic segmentation are at the core of many computer vision tasks, many of which, involve interacting with humans in potentially dangerous ways. It is therefore paramount that techniques for principled design of robust models be developed. In this paper, we provide analytic and empirical evidence that correcting potentially errant non-distinct mappings that result from the softmax function can result in improving robustness characteristics on a stateof-the-art semantic segmentation model with minimal impact to performance and minimal changes to the code base.
Charles Lehman, Dogancan Temel, Ghassan Al-Regib
ICIP3
2019 Object Recognition Under Multifarious Conditions: A Reliability Analysis and a Feature Similarity-Based Performance Estimation
abstract
In this paper, we investigate the reliability of online recognition platforms, Amazon Rekognition and Microsoft Azure, with respect to changes in background, acquisition device, and object orientation. We focus on platforms that are commonly used by the public to better understand their real-world performances. To assess the variation in recognition performance, we perform a controlled experiment by changing the acquisition conditions one at a time. We use three smartphones, one DSLR, and one webcam to capture side views and overhead views of objects in a living room, an office, and photo studio setups. Moreover, we introduce a framework to estimate the recognition performance with respect to backgrounds and orientations. In this framework, we utilize both handcrafted features based on color, texture, and shape characteristics and data-driven features obtained from deep neural networks. Experimental results show that deep learning- based image representations can estimate the recognition performance variation with a Spearman's rank-order correlation of 0.94 under multifarious acquisition conditions.
Dogancan Temel, Jinsol Lee, Ghassan Al-Regib
ICIP3
2019 Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan Al-Regib, Zsolt Kira, Richard Socher, Caiming Xiong
ICLR (Poster)4
2019 Texture retrieval using periodically extended and adaptive curvelets
Hasan Al-Marzouqi, Yuting Hu 0001, Ghassan Al-Regib
Signal Process. Image Commun.3
2019 TS-LSTM and temporal-inception: Exploiting spatiotemporal dynamics for activity recognition
Chih-Yao Ma, Min-Hung Chen, Zsolt Kira, Ghassan Al-Regib
Signal Process. Image Commun.4
2019 Perceptual image quality assessment through spectral analysis of error representations
Dogancan Temel, Ghassan Al-Regib
Signal Process. Image Commun.2
2018 Attend and Interact: Higher-Order Object Interactions for Video Understanding
abstract
Human actions often involve complex interactions across several inter-related objects in the scene. However, existing approaches to fine-grained video understanding or visual relationship detection often rely on single object representation or pairwise object relationships. Furthermore, learning interactions across multiple objects in hundreds of frames for video is computationally infeasible and performance may suffer since a large combinatorial space has to be modeled. In this paper, we propose to efficiently learn higher-order interactions between arbitrary subgroups of objects for fine-grained video understanding. We demonstrate that modeling object interactions significantly improves accuracy for both action recognition and video captioning, while saving more than 3-times the computation over traditional pairwise relationships. The proposed method is validated on two large-scale datasets: Kinetics and ActivityNet Captions. Our SINet and SINet-Caption achieve state-of-the-art performances on both datasets even though the videos are sampled at a maximum of 1 FPS. To the best of our knowledge, this is the first work modeling object interactions on open domain large-scale video datasets, and we additionally model higher-order object interactions which improves the performance with low computational costs.
Chih-Yao Ma, Asim Kadav, Iain Melvin, Zsolt Kira, Ghassan Al-Regib, Hans Peter Graf
CVPR5
2018 Fault Detection Using Attention Models Based on Visual Saliency
abstract
In this paper, we present an approach for detecting faults within seismic volumes using a saliency detection framework that employs a 3D-FFT local spectra and multi-dimensional plane projections. The projection scheme divides a 3D-FFT local spectrum into three distinct components, each depicting variations along different dimensions of the data. To detect seismic structures oriented at different angles and to capture directional features within 3D volume, we modify the center-surround model to incorporate directional comparisons around each voxel. The weighted combination of the obtained features then yields a saliency map. Experimental results on a real seismic dataset from the Great South Basin in New Zealand show the effectiveness of the proposed algorithm in the detection of complex fault networks, which are hardly conspicuous within original seismic volume. The subjective evaluation of the results show that the proposed method outperforms the state-of-the-art saliency algorithms and seismic attributes in detecting complex structures and holds a promising future in computer-aided extraction of other geologic features as well.
Muhammad Amir Shafiq, Zhiling Long, Haibin Di, Ghassan Al-Regib, Mohamed Deriche 0001
ICASSP4
2018 Semantically Interpretable and Controllable Filter Sets
abstract
In this paper, we generate and control semantically interpretable filters that are directly learned from natural images in an unsupervised fashion. Each semantic filter learns a visually interpretable local structure in conjunction with other filters. The significance of learning these interpretable filter sets is demonstrated on two contrasting applications. The first application is image recognition under progressive decolorization, in which recognition algorithms should be color-insensitive to achieve a robust performance. The second application is image quality assessment where objective methods should be sensitive to color degradations. In the proposed work, the sensitivity and lack thereof are controlled by weighing the semantic filters based on the local structures they represent. To validate the proposed approach, we utilize the CURE-TSR dataset for image recognition and the TID 2013 dataset for image quality assessment. We show that the proposed semantic filter set achieves state-of-the-art performances in both datasets while maintaining its robustness across progressive distortions.
Mohit Prabhushankar, Gukyeong Kwon, Dogancan Temel, Ghassan Al-Regib
ICIP4
2018 CURE-OR: Challenging Unreal and Real Environments for Object Recognition
abstract
In this paper, we introduce a large-scale, controlled, and multi-platform object recognition dataset denoted as Challenging Unreal and Real Environments for Object Recognition (CURE-OR). In this dataset, there are 1,000,000 images of 100 objects with varying size, color, and texture that are positioned in five different orientations and captured using five devices including a webcam, a DSLR, and three smartphone cameras in real-world (real) and studio (unreal) environments. The controlled challenging conditions include underexposure, overexposure, blur, contrast, dirty lens, image noise, resizing, and loss of color information. We utilize CURE-OR dataset to test recognition APIs - Amazon Rekognition and Microsoft Azure Computer Vision - and show that their performance significantly degrades under challenging conditions. Moreover, we investigate the relationship between object recognition and image quality and show that objective quality algorithms can estimate recognition performance under certain photometric challenging conditions. The dataset is publicly available at https://ghassanalregib.com/cure-or/ https://ghassanalregib.com/cure-or/.
Dogancan Temel, Jinsol Lee, Ghassan Al-Regib
ICMLA3
2018 A High-Speed, Real-Time Vision System for Texture Tracking and Thread Counting
abstract
In garment manufacturing, an automatic sewing machine is desirable to reduce cost. To accomplish this, a high-speed vision system is required to track fabric motions and recognize repetitive weave patterns with high accuracy, from a microperspective near a sewing zone. In this letter, we present an innovative framework for real-time texture tracking and weave pattern recognition. Our framework includes a module for motion estimation using blob detection and feature matching. It also includes a module for lattice detection to facilitate the weave pattern recognition. Our lattice-detection algorithm utilizes blob detection and template matching to assess pair-wise similarity in blobs' appearance. In addition, it extracts information of dominant orientations to obtain a global constraint in the topology. By incorporating both constraints in the appearance similarity and the global topology, the algorithm determines a lattice that characterizes the topological structure of the repetitive weave pattern, thus allowing for thread counting. In our experiments, the proposed thread-based texture tracking system is capable of tracking denim fabric with high accuracy (e.g., 0.03° rotation and 0.02 weave-thread translation errors) and high speed (3 frames per second), demonstrating its high potential for automatic real-time textile manufacturing.
Yuting Hu 0001, Zhiling Long, Ghassan Al-Regib
IEEE Signal Process. Lett.3
2018 Unsupervised Uncertainty Estimation Using Spatiotemporal Cues in Video Saliency Detection
abstract
In this paper, we address the problem of quantifying the reliability of computational saliency for videos, which can be used to improve saliency-based video processing algorithms and enable more reliable performance and objective risk assessment of saliency-based video processing applications. Our approach to quantify such reliability is twofold. First, we explore spatial correlations in both the saliency map and the eye-fixation map. Then, we learn the spatiotemporal correlations that define a reliable saliency map. We first study spatiotemporal eye-fixation data from the public CRCNS data set and investigate a common feature in human visual attention, which dictates a correlation in saliency between a pixel and its direct neighbors. Based on the study, we then develop an algorithm that estimates a pixel-wise uncertainty map that reflects our supposed confidence in the associated computational saliency map by relating a pixel's saliency to the saliency of its direct neighbors. To estimate such uncertainties, we measure the divergence of a pixel, in a saliency map, from its local neighborhood. In addition, we propose a systematic procedure to evaluate uncertainty estimation performance by explicitly computing uncertainty ground truth as a function of a given saliency map and eye fixations of human subjects. In our experiments, we explore multiple definitions of locality and neighborhoods in spatiotemporal video signals. In addition, we examine the relationship between the parameters of our proposed algorithm and the content of the videos. The proposed algorithm is unsupervised, making it more suitable for generalization to most natural videos. Also, it is computationally efficient and flexible for customization to specific video content. Experiments using three publicly available video data sets show that the proposed algorithm outperforms state-of-the-art uncertainty estimation methods with improvement in accuracy up to 63% and offers efficiency and flexibility that make it more useful in practical situations.
Tariq Alshawi, Zhiling Long, Ghassan Al-Regib
IEEE Trans. Image Process.3
2017 Scale selective extended local binary pattern for texture classification
abstract
In this paper, we propose a new texture descriptor, scale selective extended local binary pattern (SSELBP), to characterize texture images with scale variations. We first utilize multi-scale extended local binary patterns (ELBP) with rotation-invariant and uniform mappings to capture robust local micro- and macro-features. Then, we build a scale space using Gaussian filters and calculate the histogram of multi-scale ELBPs for the image at each scale. Finally, we select the maximum values from the corresponding bins of multi-scale ELBP histograms at different scales as scale-invariant features. A comprehensive evaluation on public texture databases (KTH-TIPS and UMD) shows that the proposed SSELBP has high accuracy comparable to state-of-the-art texture descriptors on gray-scale-, rotation-, and scale-invariant texture classification but uses only one-third of the feature dimension.
Yuting Hu 0001, Zhiling Long, Ghassan Al-Regib
ICASSP3
2017 Phase Congruency for image understanding with applications in computational seismic interpretation
abstract
Phase Congruency (PC) can highlight small discontinuities in images with varying illumination and contrast using the congruency of phase in Fourier components. PC can not only detect the subtle variations in the image intensity but can also highlight the anomalous values to develop a deeper understanding of the images content and context. In this paper, we propose a new method based on PC for computational seismic interpretation with an application to subsurface structures delineation within migrated seismic volumes. We show the effectiveness of the proposed method as compared to the edge- and texture-based methods for salt domes boundary detection. The subjective and objective evaluation of the experimental results on the real seismic dataset from the North Sea, F3 block show that the proposed method is not only computationally very efficient but also outperforms the state of the art methods for salt dome delineation.
Muhammad Amir Shafiq, Yazeed Alaudah, Ghassan Al-Regib, Mohamed Deriche 0001
ICASSP3
2017 Generating adaptive and robust filter sets using an unsupervised learning framework
abstract
In this paper, we introduce an adaptive unsupervised learning framework, which utilizes natural images to train filter sets. The applicability of these filter sets is demonstrated by evaluating their performance in two contrasting applications - image quality assessment and texture retrieval. While assessing image quality, the filters need to capture perceptual differences based on dissimilarities between a reference image and its distorted version. In texture retrieval, the filters need to assess similarity between texture images to retrieve closest matching textures. Based on experiments, we show that the filter responses span a set in which a monotonicity-based metric can measure both the perceptual dissimilarity of natural images and the similarity of texture images. In addition, we corrupt the images in the test set and demonstrate that the proposed method leads to robust and reliable retrieval performance compared to existing methods.
Mohit Prabhushankar, Dogancan Temel, Ghassan Al-Regib
ICIP3
2017 Saliency detection for seismic applications using multi-dimensional spectral projections and directional comparisons
abstract
In this paper, we propose a novel approach for saliency detection for seismic applications using 3D-FFT local spectra and multi-dimensional plane projections. We develop a projection scheme by dividing a 3D-FFT local spectrum of a data volume into three distinct components, each depicting changes along a different dimension of the data. The saliency detection results obtained using each projected component are then combined to yield a saliency map. To accommodate the directional nature of seismic data, in this work, we modify the center-surround model, proven to be biologically plausible for visual attention, to incorporate directional comparisons around each voxel in a 3D volume. Experimental results on real seismic dataset from the F3 block in Netherlands offshore in the North Sea prove that the proposed algorithm is effective, efficient, and scalable. Furthermore, a subjective comparison of the results shows that it outperforms the state-of-the-art methods for saliency detection.
Muhammad Amir Shafiq, Zhiling Long, Tariq Alshawi, Ghassan Al-Regib
ICIP4
2017 Power of tempospatially unified spectral density for perceptual video quality assessment
abstract
We propose a perceptual video quality assessment (PVQA) metric for distorted videos by analyzing the power spectral density (PSD) of a group of pictures. This is an estimation approach that relies on the changes in video dynamic calculated in the frequency domain and are primarily caused by distortion. We obtain a feature map by processing a 3D PSD tensor obtained from a set of distorted frames. This is a full-reference tempospatial approach that considers both temporal and spatial PSD characteristics. This makes it ubiquitously suitable for videos with varying motion patterns and spatial contents. Our technique does not make any assumptions on the coding conditions, streaming conditions or distortion. This approach is also computationally inexpensive which makes it feasible for real-time and practical implementations. We validate our proposed metric by testing it on a variety of distorted sequences from PVQA databases. The results show that our metric estimates the perceptual quality at the sequence level accurately. We report the correlation coefficients with the differential mean opinion scores (DMOS) reported in the databases. The results show high and competitive correlations compared with the state of the art techniques.
Mohammed A. Aabed, Gukyeong Kwon, Ghassan Al-Regib
ICME3
2017 Curvelet transform with learning-based tiling
Hasan Al-Marzouqi, Ghassan Al-Regib
Signal Process. Image Commun.2
2016 SalSi: A new seismic attribute for salt dome detection
abstract
In this paper, we propose a saliency-based attribute, SalSi, to detect salt dome bodies within seismic volumes. SalSi is based on the saliency theory and modeling of the human vision system (HVS). In this work, we aim to highlight the parts of the seismic volume that receive highest attention from the human interpreter, and based on the salient features of a seismic image, we detect the salt domes. Experimental results show the effectiveness of SalSi on the real seismic dataset acquired from the North Sea, F3 block. Subjectively, we have used the ground truth and the output of different salt dome delineation algorithms to validate the results of SalSi. For the objective evaluation of results, we have used the receiver operating characteristics (ROC) curves and area under the curves (AUC) to demonstrate SalSi is a promising and an effective attribute for seismic interpretation.
Muhammad Amir Shafiq, Tariq Alshawi, Zhiling Long, Ghassan Al-Regib
ICASSP4
2016 Tensor-based subspace learning for tracking salt-dome boundaries constrained by seismic attributes
abstract
We propose a method to delineate salt-dome structures by tracking manually labeled boundaries through seismic volumes. We first extract texture features from boundary regions using the tensor-based subspace learning method. Then, we utilize one seismic attribute, the gradient of texture (GoT), as a constraint on the tracking process. Using texture features and GoT maps, we can identify tracked points and optimally connect them to synthesize the boundaries. The proposed method is evaluated using real-world seismic data and experimental results show that it outperforms the state of the art in accuracy, robustness, and computational efficiency.
Zhen Wang 0007, Zhiling Long, Ghassan Al-Regib
ICASSP3
2016 Weakly-supervised labeling of seismic volumes using reference exemplars
abstract
Localizing seismic structures that can form traps for hydrocarbon reservoirs within large seismic volumes is a very challenging task. Due to the lack of accurately labeled data, we propose a weakly-supervised model for labeling seismic volumes using only a few labeled exemplars. Using six manually-labeled patches, we are able to extract patches that contain instances of similar geophysical structures. Features based on the effective singular values of curvelet coefficients are then used to train a classifier that can label an entire seismic volume with relatively high accuracy. Experiments on reference seismic sections in the Netherlands North Sea seismic dataset result in 73.8% mean pixel accuracy and 75.5% mean class accuracy, with an average labeling time of 5.2 seconds per section. These results are promising considering the nature of seismic images, and the lack of accurate edges between different geological structures.
Yazeed Alaudah, Ghassan Al-Regib
ICIP2
2016 Completed local derivative pattern for rotation invariant texture classification
abstract
In this paper, we propose a new texture descriptor, completed local derivative pattern (CLDP). In contrast to completed local binary pattern (CLBP), which involves only local differences at each scale, CLDP encodes the directional variation of the local differences of two scales as a complementary component to local patterns in CLBP. The new component in CLDP, with regarded as the directional derivative pattern, reflects the directional smoothness of local textures without increasing computation complexity. Experimental results on the Outex database show that CLDP, as a uni-scale pattern, outperforms uni-scale state-of-the-art texture descriptors on texture classification and has comparable performance with multi-scale texture descriptors.
Yuting Hu 0001, Zhiling Long, Ghassan Al-Regib
ICIP3
2016 ReSIFT: Reliability-weighted sift-based image quality assessment
abstract
This paper presents a full-reference image quality estimator based on SIFT descriptor matching over reliability-weighted feature maps. Reliability assignment includes a smoothing operation, a transformation to perceptual color domain, a local normalization stage, and a spectral residual computation with global normalization. The proposed method ReSIFT is tested on the LIVE and the LIVE Multiply Distorted databases and compared with 11 state-of-the-art full-reference quality estimators. In terms of the Pearson and the Spearman correlation, ReSIFT is the best performing quality estimator in the overall databases. Moreover, ReSIFT is the best performing quality estimator in at least one distortion group in compression, noise, and blur category.
Dogancan Temel, Ghassan Al-Regib
ICIP2
2016 Understanding spatial correlation in eye-fixation maps for visual attention in videos
abstract
In this paper, we present an analysis of recorded eye-fixation data from human subjects viewing video sequences. The purpose is to better understand visual attention for videos. Utilizing the eye-fixation data provided in the CRCNS (Collaborative Research in Computational Neuroscience) dataset, this paper focuses on the relation between the saliency of a pixel and that of its direct neighbors, without making any assumption about the structure of the eye-fixation maps. By employing some basic concepts from information theory, the analysis shows substantial correlation between the saliency of a pixel and the saliency of its neighborhood. The analysis also provides insights into the structure and dynamics of the eye-fixation maps, which can be very useful in understanding video saliency and its applications.
Tariq Alshawi, Zhiling Long, Ghassan Al-Regib
ICME3
2016 BLeSS: Bio-inspired low-level spatiochromatic similarity assisted image quality assessment
abstract
This paper proposes a biologically-inspired low-level spatiochromatic-model-based similarity method (BLeSS) to assist full-reference image-quality estimators that originally oversimplify color perception processes. More specifically, the spatiochromatic model is based on spatial frequency, spatial orientation, and surround contrast effects. The assistant similarity method is used to complement image-quality estimators based on phase congruency, gradient magnitude, and spectral residual. The effectiveness of BLeSS is validated using FSIM, FSIMc and SR-SIM methods on LIVE, Multiply Distorted LIVE, and TID 2013 databases. In terms of Spearman correlation, BLeSS enhances the performance of all quality estimators in color-based degradations and the enhancement is at 100% for both feature- and spectral residualbased similarity methods. Moreover, BleSS significantly enhances the performance of SR-SIM and FSIM in the full TID 2013 database.
Dogancan Temel, Ghassan Al-Regib
ICME2
2016 Perceptual video quality assessment: Spatiotemporal pooling strategies for different distortions and visual maps
abstract
In this paper, we investigate the challenge of distortion map feature selection and spatiotemporal pooling in perceptual video quality assessment (PVQA). We analyze three distortion maps representing different visual features spatially and temporally: squared error, local pixel-level SSIM, and absolute difference of optical flow magnitudes. We examine the performance of each of these maps with different spatial and temporal pooling strategies across three databases. We identify the most effective statistical pooling strategies spatially and temporally with respect to PVQA. We also show the most significant spatial and temporal features correlated with perception for every distortion/feature map. Our results show that varying the pooling strategy and distortion maps yields a significant improvement in perceptual quality estimation. We also deduce insights from our results to better understand the sensitivity of human vision to distortions. We aim for these findings to provide perceptual cues and guidelines to researchers during metric design, perceptual feature selection, HVS modeling and pooling selection/optimization. We further show that the same distortions across databases can yield different results in terms of PVQA evaluation and verification.
Mohammed A. Aabed, Ghassan Al-Regib
MMSP2
2016 Content-adaptive non-parametric texture similarity measure
abstract
In this paper, we introduce a non-parametric texture similarity measure based on the singular value decomposition of the curvelet coefficients followed by a content-based truncation of the singular values. This measure focuses on images with repeating structures and directional content such as those found in natural texture images. Such textural content is critical for image perception and its similarity plays a vital role in various computer vision applications. In this paper, we evaluate the effectiveness of the proposed measure using a retrieval experiment. The proposed measure outperforms the state-of-the-art texture similarity metrics on CUReT and PerTex texture databases, respectively.
Motaz Alfarraj, Yazeed Alaudah, Ghassan Al-Regib
MMSP3
2016 Boosting in image quality assessment
abstract
In this paper, we analyze the effect of boosting in image quality assessment through multi-method fusion. On the contrary of existing studies that propose a single quality estimator, we investigate the generalizability of multi-method fusion as a framework. In addition to support vector machines that are commonly used in the multi-method fusion studies, we propose using neural networks in the boosting. To span different types of image quality assessment algorithms, we use quality estimators based on fidelity, perceptually-extended fidelity, structural similarity, spectral similarity, color, and learning. In the experiments, we perform k-fold cross validation using the LIVE, the multiply distorted LIVE, and the TID 2013 databases and the performance of image quality assessment algorithms are measured via accuracy-, linearity-, and ranking-based metrics. Based on the experiments, we show that boosting methods generally improve the performance of image quality assessment and the level of improvement depends on the type of the boosting algorithm. Our experimental results also indicate that boosting the worst performing quality estimator with two or more methods lead to statistically significant performance enhancements independent of the boosting technique and neural network-based boosting outperforms support vector machine-based boosting when two or more methods are fused.
Dogancan Temel, Ghassan Al-Regib
MMSP2
2016 CSV: Image quality assessment based on color, structure, and visual system
Dogancan Temel, Ghassan Al-Regib
Signal Process. Image Commun.2
2016 UNIQUE: Unsupervised Image Quality Estimation
abstract
In this letter, we estimate perceived image quality using sparse representations obtained from generic image databases through an unsupervised learning approach. A color space transformation, a mean subtraction, and a whitening operation are used to enhance descriptiveness of images by reducing spatial redundancy; a linear decoder is used to obtain sparse representations; and a thresholding stage is used to formulate suppression mechanisms in a visual system. A linear decoder is trained with 7 GB worth of data, which corresponds to 100 000 8 × 8 image patches randomly obtained from nearly 1000 images in the ImageNet 2013 database. A patch-wise training approach is preferred to maintain local information. The proposed quality estimator UNIQUE is tested on the LIVE, the Multiply Distorted LIVE, and the TID 2013 databases and compared with 13 quality estimators. Experimental results show that UNIQUE is generally a top performing quality estimator in terms of accuracy, consistency, linearity, and monotonic behavior.
Dogancan Temel, Mohit Prabhushankar, Ghassan Al-Regib
IEEE Signal Process. Lett.3
2016 Air-Writing Recognition - Part I: Modeling and Recognition of Characters, Words, and Connecting Motions
abstract
Air-writing refers to writing of linguistic characters or words in a free space by hand or finger movements. Air-writing differs from conventional handwriting; the latter contains the pen-up-pen-down motion, while the former lacks such a delimited sequence of writing events. We address air-writing recognition problems in a pair of companion papers. In Part I, recognition of characters or words is accomplished based on six-degree-of-freedom hand motion data. We address air-writing on two levels: motion characters and motion words. Isolated air-writing characters can be recognized similar to motion gestures although with increased sophistication and variability. For motion word recognition in which letters are connected and superimposed in the same virtual box in space, we build statistical models for words by concatenating clustered ligature models and individual letter models. A hidden Markov model is used for air-writing modeling and recognition. We show that motion data along dimensions beyond a 2-D trajectory can be beneficially discriminative for air-writing recognition. We investigate the relative effectiveness of various feature dimensions of optical and inertial tracking signals and report the attainable recognition performance correspondingly. The proposed system achieves a word error rate of 0.8% for word-based recognition and 1.9% for letter-based recognition. We also subjectively and objectively evaluate the effectiveness of air-writing and compare it with text input using a virtual keyboard. The words per minute of air-writing and virtual keyboard are 5.43 and 8.42, respectively.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
IEEE Trans. Hum. Mach. Syst.2
2016 Air-Writing Recognition - Part II: Detection and Recognition of Writing Activity in Continuous Stream of Motion Data
abstract
Air-writing refers to writing of characters or words in the free space by hand or finger movements. We address air-writing recognition problems in two companion papers. Part 2 addresses detecting and recognizing air-writing activities that are embedded in a continuous motion trajectory without delimitation. Detection of intended writing activities among superfluous finger movements unrelated to letters or words presents a challenge that needs to be treated separately from the traditional problem of pattern recognition. We first present a dataset that contains a mixture of writing and nonwriting finger motions in each recording. The LEAP from Leap Motion is used for marker-free and glove-free finger tracking. We propose a window-based approach that automatically detects and extracts the air-writing event in a continuous stream of motion data, containing stray finger movements unrelated to writing. Consecutive writing events are converted into a writing segment. The recognition performance is further evaluated based on the detected writing segment. Our main contribution is to build an air-writing system encompassing both detection and recognition stages and to give insights into how the detected writing segments affect the recognition result. With leave-one-out cross validation, the proposed system achieves an overall segment error rate of 1.15% for word-based recognition and 9.84% for letter-based recognition.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
IEEE Trans. Hum. Mach. Syst.2
2015 Reduced-reference perceptual quality assessment for video streaming
abstract
We propose a perceptual video quality monitoring metric for streaming applications using the optical flow features. This approach is a reduced-reference pixel-based and relies only on the deviation of the optical flow of the corrupted frames. This techniques compares an optical flow descriptor from the corrupted frame against the descriptor obtained from the anchor frame. This approach is suitable for videos with complex motion patterns. Our technique does not make any assumptions on the coding conditions, network loss patterns or error concealment techniques. We validate our proposed metric by testing it on a variety of distorted sequences from the LIVE mobile database. Our results show that our metric estimates the perceptual quality at the sequence level accurately. We report the correlation coefficients with the differential mean opinion scores (DMOS) reported in the database. The results show Spearman's correlations of 0.88 for the tested sequences.
Mohammed A. Aabed, Ghassan Al-Regib
ICIP2
2015 A curvelet-based distance measure for seismic images
abstract
We introduce a new curvelet-based distance measure for post-migration seismic data. The measure exploits the highly directional content of seismic images. It calculates the sum of the squared chord distances between the histograms of the curvelet coefficients of two images, over all orientations and scales. In a retrieval task consisting of 918 retrieval instances, the proposed method successfully retrieves all images. This wasn't achieved by other methods in the literature. Furthermore, in comparison to the state-of-the-art, the proposed measure required 86% less computation time.
Yazeed Alaudah, Ghassan Al-Regib
ICIP2
2015 PerSIM: Multi-resolution image quality assessment in the perceptually uniform color domain
abstract
An average observer perceives the world in color instead of black and white. Moreover, the visual system focuses on structures and segments instead of individual pixels. Based on these observations, we propose a full reference objective image quality metric modeling visual system characteristics and chroma similarity in the perceptually uniform color domain (Lab). Laplacian of Gaussian features are obtained in the L channel to model the retinal ganglion cells in human visual system and color similarity is calculated over the a and b channels. In the proposed perceptual similarity index (PerSIM), a multi-resolution approach is followed to mimic the hierarchical nature of human visual system. LIVE and TID2013 databases are used in the validation and PerSIM outperforms all the compared metrics in the overall databases in terms of ranking, monotonic behavior and linearity.
Dogancan Temel, Ghassan Al-Regib
ICIP2
2015 Tensor-based subspace learning for tracking salt-dome boundaries
abstract
The exploration of petroleum reservoirs has a close relationship with the identification of salt domes. To efficiently interpret salt-dome structures, in this paper, we propose a method that tracks salt-dome boundaries through seismic volumes using a tensor-based subspace learning algorithm. We build texture tensors by classifying image patches acquired along the boundary regions of seismic sections and contrast maps. With features extracted from the subspaces of texture tensors, we can identify tracked points in neighboring sections and label salt-dome boundaries by optimally connecting these points. Experimental results show that the proposed method outperforms the state-of-the-art salt-dome detection method by employing texture information and tensor-based analysis.
Zhen Wang 0007, Zhiling Long, Ghassan Al-Regib
ICIP3
2014 No-reference quality assessment of HEVC videos in loss-prone networks
abstract
In this paper, we propose a no-reference quality assessment measure for high efficiency video coding (HEVC). We analyze the impact of network losses on HEVC videos and the resulting error propagation. We estimate channel-induced distortion in the video assuming we have access to the decoded video only without access to the bitstream or the decoder. Our model does not make any assumptions on the coding conditions, network loss patterns or error concealment techniques. The proposed approach relies only on the temporal variations of the power spectrum across the decoded frames. We validate our proposed quality measure by testing it on a variety of HEVC coded videos subject to network losses. Our simulation results show that the proposed model accurately captures channel-induced distortions. For the test videos, the correlation coefficients between the proposed measure and the full-reference SSIM values range between 0.70 and 0.80.
Mohammed A. Aabed, Ghassan Al-Regib
ICASSP2
2014 Fault detection in seismic datasets using hough transform
abstract
In this paper, we present a semi-automatic algorithm to detect faults in seismic datasets using Hough transform. As a multistage approach, our method first highlights the likely fault points from the discontinuity map of one seismic section. Hough transform is then applied to detect faults features. Considering geological constraints of faults, false features are removed using a double-threshold method. Then, we get an initial fault line by connecting the remaining faults features. In the last stage, by incorporating the discontinuity information from step one, we tweak the initial fault line to obtain more accurate and reliable results. Our experimental results show that our method can delineate the fault lines in seismic sections more accurately than a state-of-the-art method.
Zhen Wang 0007, Ghassan Al-Regib
ICASSP2
2014 A comparative study of computational aesthetics
abstract
Objective metrics model image quality by quantifying image degradations or estimating perceived image quality. However, image quality metrics do not model what makes an image more appealing or beautiful. In order to quantify the aesthetics of an image, we need to take it one step further and model the perception of aesthetics. In this paper, we examine computational aesthetics models that use hand-crafted, generic and hybrid descriptors. We show that generic descriptors can perform as well as state of the art hand-crafted aesthetics models that use global features. However, neither generic nor hand-crafted features is sufficient to model aesthetics when we only use global features without considering spatial composition or distribution. We also follow a visual dictionary approach similar to state of the art methods and show that it performs poorly without the spatial pyramid step.
Dogancan Temel, Ghassan Al-Regib
ICIP2
2014 Automatic fault tracking across seismic volumes via tracking vectors
abstract
The identification of reservoir regions has a close relationship with the detection of faults in seismic volumes. However, only relying on human intervention, most fault detection algorithms are inefficient. In this paper, we present a new technique that automatically tracks faults across a 3D seismic volume. To implement automation, we propose a two-way fault line projection based on estimated tracking vectors. In the tracking process, projected fault lines are integrated into a synthesized line as the tracked fault line, through an optimization process with local geological constraints. The tracking algorithm is evaluated using real-world seismic data sets with promising results. The proposed method provides comparable accuracy to the detection of faults explicitly in every seismic section, and it also reduces computational complexity.
Zhen Wang 0007, Zhiling Long, Ghassan Al-Regib, Asjad Amin, Mohamed Deriche 0001
ICIP3
2013 Searching for the optimal curvelet tiling
abstract
Curvelets were recently introduced as a popular extension of wavelets. In the curvelet domain the input image is represented by sets of coefficients representing signal energy in different scales and angular directions. In this paper an algorithm that searches for optimal tilings for use with the curvelet transform is introduced. We consider two adaptations: scale locations, and the number of angular divisions per scale. A search algorithm that searches for the optimal tiling with respect to denoising performance is introduced. Results show significant improvement over original curvelet tilings. Tiling results were also tested with a seismic compressed sensing recovery problem. A similar performance advantage is reported.
Hasan Al-Marzouqi, Ghassan Al-Regib
ICASSP2
2013 Joint Framework for Motion Validity and Estimation Using Block Overlap
abstract
This paper presents a block-overlap-based validity metric for use as a measure of motion vector (MV) validity and to improve the quality of the motion field. In contrast to other validity metrics in the literature, the proposed metric is not sensitive to image features and does not require the use of neighboring MVs or manual thresholds. Using a hybrid de-interlacer, it is shown that the proposed metric outperforms other block-based validity metrics in the literature. To help regularize the ill-posed nature of motion estimation, the proposed validity metric is also used as a regularizer in an energy minimization framework to determine the optimal MV. Experimental results show that the proposed energy minimization framework outperforms several existing motion estimation methods in the literature in terms of MV and interpolation quality. For interpolation quality, our algorithm outperforms all other block-based methods as well as several complex optical flow methods. In addition, it is one of the fastest implementations at the time of this writing.
Michael Santoro, Ghassan Al-Regib, Yücel Altunbasak
IEEE Trans. Image Process.2
2013 Cooperative Delivery Techniques to Support Video-on-Demand Service in IPTV Networks
abstract
In this paper, we study the use of peer-assisted server-based cooperative transmission strategies in IPTV networks for the delivery of on-demand services to end users. The proposed techniques aim to support the resource efficient delivery of on-demand content to end users in a timely manner. Within the proposed framework, a cooperative transmission strategy suggests that users who have access to the requested content cooperatively transmit to the targeted set of users. In doing so, we can minimize the servicing requirements at the server side and improve the scalability performance in the network. We conducted extensive simulations and showed that significant performance improvements can be achieved with the proposed delivery techniques to enable efficient access to an ever growing on-demand content. We also showed the robustness of the proposed techniques in regard to variations observed in network state.
Aytac Azgin, Ghassan Al-Regib, Yücel Altunbasak
IEEE Trans. Multim.2
2013 Feature Processing and Modeling for 6D Motion Gesture Recognition
abstract
A 6D motion gesture is represented by a 3D spatial trajectory and augmented by another three dimensions of orientation. Using different tracking technologies, the motion can be tracked explicitly with the position and orientation or implicitly with the acceleration and angular speed. In this work, we address the problem of motion gesture recognition for command-and-control applications. Our main contribution is to investigate the relative effectiveness of various feature dimensions for motion gesture recognition in both user-dependent and user-independent cases. We introduce a statistical feature-based classifier as the baseline and propose an HMM-based recognizer, which offers more flexibility in feature selection and achieves better performance in recognition accuracy than the baseline system. Our motion gesture database which contains both explicit and implicit motion information allows us to compare the recognition performance of different tracking signals on a common ground. This study also gives an insight into the attainable recognition rate with different tracking devices, which is valuable for the system designer to choose the proper tracking technology.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
IEEE Trans. Multim.2
2012 Cooperative on-demand delivery for IPTV networks
abstract
In this paper, we investigate the use of cooperative transmission strategies to support timely and efficient delivery of on-demand content to end users. Within the proposed framework, a cooperative transmission strategy suggests that users who have access to the requested content cooperatively transmit to the targeted user(s) to minimize the servicing requirements at the server side and to improve the scalability performance in the network. Through extensive simulation-based studies, we show that significant performance improvements can be achieved with the proposed framework to enable efficient access to an ever growing on-demand content.
Aytac Azgin, Ghassan Al-Regib, Yücel Altunbasak
GLOBECOM2
2012 6D motion gesture recognition using spatio-temporal features
abstract
Depending on the tracking technology in use, a 6D motion gesture can be tracked and represented explicitly by the position and orientation or implicitly by the acceleration and angular speed. In this work, we first present the reasoning for the definition and recognition of motion gestures. Five basic feature vectors are then derived from the 6D motion data. Our main contribution is to investigate the relative effectiveness of various feature dimensions for motion gesture recognition in both user dependent and user independent cases. We also propose a feature normalization procedure and prove its effectiveness in achieving “scale” invariance especially in the user independent case. Our study gives an insight into the attainable recognition rate with different tracking devices.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
ICASSP2
2012 Block-overlap-based validity metric for hybrid de-interlacing
abstract
This paper presents a block-overlap-based validity metric in the context of hybrid de-interlacing. The proposed metric uses the degree of motion-compensated block overlap as a measure of motion vector (MV) validity. In contrast to other validity methods in the literature, the proposed method is not sensitive to image features and does not require the use of neighboring MVs or manual thresholds. To perform hybrid de-interlacing, the confidence value for each MV is used to weigh the contribution from two de-interlacing methods. Experimental results show that the de-interlaced images produced with the proposed method are of significant higher quality than those produced by existing validity methods.
Michael Santoro, Ghassan Al-Regib, Yücel Altunbasak
ICIP2
2012 Misalignment correction for depth estimation using stereoscopic 3-D cameras
abstract
This paper presents a misalignment correction method for reducing depth errors that result from camera shift. The proposed method uses a real-time motion estimation approach to correct for alignment errors between stereo cameras. Unlike existing methods in the literature, the natural disparity between stereo views is incorporated into a constrained motion estimation framework. The proposed method is shown to accurately estimate synthetic misalignments due to translation, rotation, scaling, and perspective transformation. In addition, real images taken from a stereo camera rig confirm that the proposed method is capable of significantly reducing misalignments due to camera yaw, pitch, and roll.
Michael Santoro, Ghassan Al-Regib, Yücel Altunbasak
MMSP2
2012 Motion estimation using block overlap minimization
abstract
This paper presents a block-overlap-based method for handling the ill-posed nature of motion estimation. The proposed method uses the volume of motion-compensated block overlap as an additional term in minimizing the overall energy. By reducing the amount of block overlap, the proposed method results in a significant improvement in the quality of the motion field. Experimental results show that the proposed method outperforms several existing methods in the literature in terms of motion vector (MV) and interpolation quality. In terms of interpolation quality, our algorithm outperforms all other block-based methods as well as several complex optical flow methods. In addition, it is the fastest non-GPU implementation at the time of this writing.
Michael Santoro, Ghassan Al-Regib, Yücel Altunbasak
MMSP2
2012 6DMG: a new 6D motion gesture database
abstract
Motion-based control is gaining popularity, and motion gestures form a complementary modality in human-computer interactions. To achieve more robust user-independent motion gesture recognition in a manner analogous to automatic speech recognition, we need a deeper understanding of the motions in gesture, which arouses the need for a 6D motion gesture database. In this work, we present a database that contains comprehensive motion data, including the position, orientation, acceleration, and angular speed, for a set of common motion gestures performed by different users. We hope this motion gesture database can be a useful platform for researchers and developers to build their recognition algorithms as well as a common test bench for performance comparisons.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
MMSys2
2012 MIQM: A Multicamera Image Quality Measure
abstract
Although several subjective and objective quality assessment methods have been proposed in the literature for images and videos from single cameras, no comparable effort has been devoted to the quality assessment of multicamera images. With the increasing popularity of multiview applications, quality assessment of multicamera images and videos is becoming fundamental to the development of these applications. Image quality is affected by several factors, such as camera configuration, number of cameras, and the calibration process. In order to develop an objective metric specifically designed for multicamera systems, we identified and quantified two types of visual distortions in multicamera images: photometric distortions and geometric distortions. The relative distortion between individual camera scenes is a major factor in determining the overall perceived quality. In this paper, we show that such distortions can be translated into luminance, contrast, spatial motion, and edge-based structure components. We propose three different indices that can quantify these components. We provide examples to demonstrate the correlation among these components and the corresponding indices. Then, we combine these indices into one multicamera image quality measure (MIQM). Results and comparisons with other measures, such as peak signal-to noise ratio, mean structural similarity, and visual information fidelity show that MIQM outperforms other measures in capturing the perceptual fidelity of multicamera images. Finally, we verify the results against subjective evaluation.
Mashhour Solh, Ghassan Al-Regib
IEEE Trans. Image Process.2
2011 Trajectory triangulation: 3D motion reconstruction with ℓ1 optimization
abstract
In this paper, we first explain the formulation of the trajectory triangulation: 3D reconstruction of a moving point from a series of 2D projections. The system has to be overconstrained to be solved by least squares techniques. We take advantage of the sparseness of real-world motions in the transformed domain, and borrow the concept of compressive sampling to reformulate the problem with ℓ1optimization so that it is possible to reconstruct the trajectory even in an underconstrained system. Thus, fewer measurements are needed to reconstruct a 3D trajectory of even larger bandwidth coverage. We conduct experiments on both synthetic and real-world motion data to verify our proposed method, and compare the reconstruction results based on ℓ1and ℓ2optimization.
Mingyu Chen 0003, Ghassan Al-Regib, Biing-Hwang Juang
ICASSP2
2011 Improved DCT coefficient distribution modeling for H.264-like video coders based on block classification
abstract
Through extensive experimentation with a large set of video sequences, we show that modeling the statistical distribution of the transform coefficients in H.264-like video coders can be improved significantly in terms of accuracy by classifying the video source into multiple classes and modeling each class with a different statistical distribution. In this paper, we present a simple yet effective classification method and best practical models for each class and show that it is possible to improve the statistical modeling significantly without a significant complexity increase. We propose a two-class based approach in which one class is composed of very low detail blocks, and the other class is composed of high texture blocks and blocks with edges. Our two-class based statistical modeling reduces the approximation error up to 70% over the existing single-class modeling approaches for majority of the video sequences experimented. Furthermore, this approach also fits very well with the context of rate control with human visual system considerations, in which distortion in low detail regions of an image is more noticeable than in high detailed regions. In this work, we consider modeling the transform coefficients all lumped together.
Nejat Kamaci, Ghassan Al-Regib
ICIP2
2011 A no-reference quality measure for DIBR-based 3D videos
abstract
In this paper we present a no-reference objective quality measure for stereoscopic 3D videos generated by depth image-based rendering (DIBR). At first we will derive an ideal depth estimate for each pixel value. The ideal depth estimate will then be used to calculate three distortion measures: temporal outliers (TO), temporal inconsistencies (TI), and spatial outliers (SO). The combination of the three measures constitute the proposed no-reference measure. Its performance is verified using subjective DMO S scores and compared to the full reference version of the proposed algorithm. The subjective results show that the proposed measure highly correlates with subjective scores and is close in performance to the full reference version of the measure.
Mashhour Solh, Ghassan Al-Regib
ICME2
2011 3VQM: A vision-based quality measure for DIBR-based 3D videos
abstract
In this paper, we present a new method for objectively evaluating the quality of stereoscopic 3D videos generated by depth-image-based rendering (DIBR). First we show how to derive an ideal depth estimate at each pixel value that would constitute a distortion-free rendered video. The ideal depth estimate will then be used to derive three distortion measures to objectify the visual discomfort in the stereoscopic videos. The three measures are temporal outliers (TO), temporal inconsistencies (TI), and spatial outliers (SO). The combination of the three measures will constitute a vision-based quality measure for 3D DIBR-based videos, 3VQM. Finally, 3VQM will be presented and verified against a fully conducted subjective evaluation. The results show that our proposed measure is significantly accurate, coherent and consistent with the subjective scores.
Mashhour Solh, Ghassan Al-Regib, Judit Martinez Bauza
ICME2
2011 The Mosaic Camera: Streaming, Coding and Compositing Experiments
abstract
The HP Fan Camera is a panoramic mosaic king camera that is a composite of 24-imager array system. Streaming the captured video is a challenging problem due to several factors such as the large bandwidth requirements, the limited capabilities of the client's machines, and our desire to provide independent viewing controls for users. In the process of developing an optimal rate controller for the HP Fan Camera we developed a client-server framework for multi-camera streaming and performed a set of experiments using various bandwidth allocation schemes. From our preliminary research, we found that sending individual streams of the cameras over the network provides more interactivity to the end users and requires less bandwidth in case the behavior of the end users is aggressive in scene selection. In this paper we present this framework and share the results of our conducted experiments.
Mashhour Solh, Ghassan Al-Regib
ISM2
2010 Hierarchical Hole-Filling(HHF): Depth image based rendering without depth map filtering for 3D-TV
abstract
In this paper we propose a new approach for disocclusion removal in depth image-based rendering (DIBR) for 3D-TV. The new approach, Hierarchical Hole-Filling (HHF), eliminates the need for any preprocessing of the depth map. HHF uses a pyramid like approach to estimate the hole pixels from lower resolution estimates of the 3D wrapped image. The lower resolution estimates involves a pseudo zero canceling plus Gaussian filtering of the wrapped image. Then starting backwards from the lowest resolution hole-free estimate in the pyramid, we interpolate and use the pixel values to fill in the hole in the higher up resolution image. The procedure is repeated until the estimated image is hole-free. Experimental results show that HHF yields virtual images that are free of any geometric distortions, which is not the case in other algorithms that preprocess the depth map. Experiments has also shown that unlike previous DIBR techniques, HHF is not sensitive to depth maps with high percentage of bad matching pixels.
Mashhour Solh, Ghassan Al-Regib
MMSP2
2009 Context-adaptive hybrid variable length coding in H.264/AVC
abstract
It is well observed that the nonzero transform coefficients for block-based hybrid video coding is more clustered in the low frequency region while more scattered in the high frequency region. Based on this observation, hybrid variable length coding (HVLC) was recently proposed, which divides each transform block into low frequency (LF) region and high frequency (HF) region by a pre-defined breakpoint and encodes the clustered low frequency coefficients with new coding schemes, such as two dimensional position and one-dimensional amplitude coding (2DP1DA) or three-dimensional position and amplitude coding (3DPA). However, the transition between clustered and scattered nonzero coefficients varies with different blocks and sequences. In this paper, we propose a context-adaptive hybrid variable length coding (CAHVLC) scheme, where multiple code tables are used and the code table is adaptively chosen based on the derivable context information from coded neighboring blocks or from the coded portion of the currently being coded block. The experimental results show that CAHVLC achieves about 6.5% ∼ 8% bit rate reduction for a wide range of quantization parameters (QP) compared with CAVLC in H.264.
Ghassan Al-Regib, Dihong Tian, Wen H. Chen, Pi Sheng Chang
ICIP2
2008 Joint position and amplitude coding in hybrid variable length coding for video compression
abstract
Hybrid variable length coding (HVLC) was recently proposed as a novel entropy coding scheme for block-based image and video compression, which divides each transform block into low frequency (LF) region and high frequency (HF) region and codes them differently. To efficiently code LF region, a two-dimensional position and one-dimensional amplitude coding scheme (2DP1DA) was also proposed, which jointly codes the 2D position information, i.e., run of consecutive zero-valued coefficients and run of consecutive nonzero coefficients. To further explore the potential of HVLC concept, we propose a new scheme for coding LF region, which codes the 2D position and amplitude information of each nonzero cluster jointly with manageable complexity. The experimental results show that compared with CAVLC in H.264, about 3.5% bit rate reduction is achieved by the proposed method for a wide range of quantization parameters (QP).
Ghassan Al-Regib, Dihong Tian, Pi Sheng Chang, Wen H. Chen
ICASSP2
2008 Maximizing Network Lifetime for Estimation in Multi-Hop Wireless Sensor Networks
abstract
In this paper, we consider distributed estimation in energy-limited wireless sensor networks from lifetime-distortion perspective, where the goal is to maximize the network lifetime for a given distortion requirement. To take into account both local quantization and multi-hop transmission, which are essential to save transmission energy and thus prolong the network lifetime, the network lifetime maximization problem is formulated as a nonlinear programming (NLP) problem, where there are three factors needed to be optimized jointly: (i) source coding at each sensor, (ii) source throughput of each sensor, and (iii) multi-hop routing path. Furthermore, we show that this NLP problem can be decoupled without loss of optimality and reformulated as a linear programming (LP) problem. The proposed algorithm is optimal and the simulation results show that a significant gain is achieved by the proposed algorithm compared with heuristic methods.
Ghassan Al-Regib
ICCCN2
2008 3-D Position and Amplitude VLC Coding in H.264/AVC
abstract
Hybrid variable length coding (HVLC) was recently proposed as a novel entropy coding scheme for block-based image and video compression, which divides each transform block into low frequency (LF) region and high frequency (HF) region and codes them differently. To take advantage of the clustered nature of the nonzero coefficients in the LF region, a two-dimensional position and one-dimensional amplitude coding scheme (2DP1DA) was also proposed, which codes the 2D position information, i.e., run of consecutive zero-valued coefficients and run of consecutive nonzero coefficients, jointly by a 2D code table. To further improve the coding efficiency, joint position and amplitude coding is desirable. In this paper, we propose a three-dimensional position and amplitude coding scheme (3DPA) to jointly code the position and amplitude information for each nonzero cluster by a 3D code table. The experimental results show that HVLC with 3DPA for coding LF region achieves about 3.2%~3.7% bit rate reduction for a wide range of quantization parameters (QP) compared with CAVLC in H.264. Moreover, the 3DPA scheme is a compact way to realize the joint position and amplitude coding and it is very suitable for context-adaptive HVLC design.
Ghassan Al-Regib, Dihong Tian, Pi Sheng Chang, Wen H. Chen
ICCCN2
2008 Arm Movement Prediction Using Neural Networks
abstract
Whether interacting with a collaborative virtual environment, or CVE, locally or one networked across the Internet, any delay in the system can lead to a reduced sense of immersion. Input sensor delay and network delay are two common problems in CVE design that can be overcome with the application of prediction algorithms to the system. The purpose of this experiment was to assess the quality of feed forward back propagation neural networks in predicting natural avatar arm movement typically used in a CVE. In addition the experiment attempts to find the bounds for precise neural network prediction. The results show many different combinations of back propagation neural network topologies are capable of predicting up to 400 ms of human arm movements relatively accurately.
Fred Stakem, Ghassan Al-Regib
ICCCN2
2008 Three-dimensional position and amplitude VLC coding in H.264/AVC
abstract
Hybrid variable length coding (HVLC) was recently proposed as a novel entropy coding scheme for block-based image and video compression, which divides each transform block into low frequency (LF) region and high frequency (HF) region and codes them differently. To take advantage of the clustered nature of the nonzero coefficients in the LF region, a two-dimensional position and one- dimensional amplitude coding scheme (2DP1DA) was also proposed, which codes the 2D position information, i.e., run of consecutive zero-valued coefficients and run of consecutive nonzero coefficients, jointly by a 2D code table. To further improve the coding efficiency, joint position and amplitude coding is desirable. In this paper, we propose a three-dimensional position and amplitude coding scheme (3DPA) to jointly code the position and amplitude information for each nonzero cluster by a 3D code table. The experimental results show that HVLC with 3DPA for coding LF region achieves about 3.2%~3.7% bit rate reduction for a wide range of quantization parameters (QP) compared with CAVLC in H.264. Moreover, the 3DPA scheme is a compact way to realize the joint position and amplitude coding and it is very suitable for context-adaptive HVLC design.
Ghassan Al-Regib, Dihong Tian, Pi Sheng Chang, Wen H. Chen
ICIP2
2008 Technology and Tools to Enhance Distributed Engineering Education
abstract
While the ongoing information technology (IT) revolution is providing us with tremendous educational opportunities, educators and IT researchers face numerous obstacles and pedagogical questions. The Georgia Institute of Technology (Georgia Tech) has a long history in engineering education research and has developed and designed many different tools for instructor authoring, course capturing, indexing, and retrieval. Special attention has been applied to the design and deployment of distributed learning environments. This paper describes the environment and challenges facing Georgia Tech as it expands its campus worldwide while maintaining integration of faculty and students across these campuses. The focus is primarily on the problems of synchronous delivery to multiple sites with a description of how technology is currently being deployed in Georgia Tech's distance-learning (DL) classrooms, as well as technology that is under development for the DL classroom of the future.
Ghassan Al-Regib, Monson H. Hayes III, Elliot Moore II, Douglas B. Williams
Proc. IEEE1
2008 Function-Based Network Lifetime for Estimation in Wireless Sensor Networks
abstract
Network lifetime is a critical concern in the design of wireless sensor networks. Though many different definitions of network lifetime have been used in the literature, we introduce a function-based network lifetime definition, which focuses on whether the network can perform a given task instead of whether any individual sensor is dead. Then we derive an upper bound of function-based network lifetime for estimation, and we maximize it by introducing a concept of equivalent unit-resource mean square error (MSE) function. The proposed algorithm is optimal and the simulation results show that a significant gain is achieved by the proposed algorithm compared with heuristic methods.
Ghassan Al-Regib
IEEE Signal Process. Lett.2
2008 BaTex3: Bit Allocation for Progressive Transmission of Textured 3-D Models
abstract
The efficient progressive transmission of a textured 3D model, which has a texture image mapped on a geometric mesh, requires careful organization of the bit stream to quickly present high-resolution visualization on the user's screen. Bit rates for the mesh and the texture need to be properly balanced so that the decoded model has the best visual fidelity during transmission. This problem is addressed in the paper, and a bit allocation framework is proposed. In particular, the relative importance of the mesh geometry and the texture image on visual fidelity is estimated using a fast quality measure (FQM). Optimal bit distribution between the mesh and the texture is then computed under the criterion of maximizing the quality measured by FQM. Both the quality measure and the bit allocation algorithm are designed with low computation complexity. Empirical studies show that not only the bit allocation framework maximizes the receiving quality of the textured 3D model, but also it does not require sending additional information to indicate bit boundaries between the mesh and the texture in the multiplexed bit stream, which makes the streaming application more robust to bit errors that may occur randomly during transmission.
Dihong Tian, Ghassan Al-Regib
IEEE Trans. Circuits Syst. Video Technol.2
2007 Hybrid Variable Length Coding for Image and Video Compression
abstract
The conventional run-level variable length coding (RL-VLC), commonly adopted in block-based image and video compression to code quantized transform coefficients, is not efficient in coding consecutive nonzero coefficients. To overcome the deficiency, hybrid variable length coding (HVLC) is proposed in this paper. HVLC takes advantage of the clustered nature of nonzero transform coefficients in the low-frequency (LF) region and the scattered nature of nonzero transform coefficients in the high-frequency (HF) region by employing two types of VLC schemes. A novel two-dimensional position and one-dimensional amplitude coding scheme is proposed to code the LF coefficients while RL-VLC or an equivalent VLC scheme is retained to code the HF coefficients. Experimental results show that HVLC greatly favors the coding of high-resolution, high-complexity scenes, while it preserves low computational complexity.
Dihong Tian, W. Homer Chen, Pi Sheng Chang, Ghassan Al-Regib, Russell M. Mersereau
ICASSP (1)4
2007 Joint Source and Channel Coding for 3-D Scene Databases Using Vector Quantization and Embedded Parity Objects
abstract
Three-dimensional graphic scenes contain various mesh objects in one geometric space where different objects have potentially unequal importance regarding display. This paper proposes an object-oriented system for efficiently coding and streaming 3-D scene databases in lossy and rate-constrained environments. Vector quantization (VQ) is exploited to code 3-D scene databases into multiresolution hierarchies. For the best distortion-rate performance, adaptive quantization precisions are allocated to different objects and different layers of each object based on a weighted distortion model. Upon transmission, scalably coded objects are delivered in respective packet sequences to preserve their manipulation independency. For packet loss resilience, a plurality of FEC codes are generated as "parity objects" parallel to graphic objects, which protect the graphic objects concurrently and also preferentially in regard to their unequal decoding importance. A rate-distortion optimization framework is then developed, which performs rate allocation between graphic objects and parity objects and generates the parity data properly. We show that, by treating graphic objects jointly and preferentially in source and channel coding while preserving their independencies in transport, the proposed system reduces the receiving distortion of the 3-D database significantly compared to conventional methods.
Dihong Tian, Ghassan Al-Regib
IEEE Trans. Image Process.3
2007 Multistreaming of 3-D Scenes With Optimized Transmission and Rendering Scalability
abstract
Three-dimensional (3D) graphic scenes require considerable network bandwidth to be transmitted and computing power to be rendered on a user's terminal. Toward high-quality display in real time, we propose a sender-driven mechanism for streaming 3D scenes in a resource-constrained environment. In doing so, objects are encoded into multiresolutions to provide transmission and rendering scalability, and a weighted distortion metric is developed to measure the quality of a scene rendered with multiresolution objects, modeling objects' unequal importance regarding display. To preserve the manipulation independency of multiple objects in data delivery while provide preferential treatment for different objects as well as different layers of each object, transmission of the objects is performed over multiple streams in a partially sequenced and partially reliable fashion. A rate-distortion optimization framework is developed, which determines an optimal level of reliability for every chunk of data in each stream, taking into account the rendering importance of the object, the distortion-rate performance of the data chunks, and the statistics of the network link. Compared with heuristical methods, simulation results show that the proposed framework maximizes the display quality of the scene while minimizing the amount of data that needs to be processed by the client's rendering engine
Dihong Tian, Ghassan Al-Regib
IEEE Trans. Multim.2
2006 Latency-Minimized Delivery for Multi-Resolution Mesh Geometry
abstract
Three-dimensional (3D) meshes are used intensively in distributed graphics applications where model data is transmitted on demand to users’ terminals and rendered for interactive manipulation. This paper presents a transmission system for such applications with an objective of minimizing the latency between the user input and the response. In particular, we represent 3D models by multiple resolutions to allow fast and scalable rendering, and provide them with unequal error protection and/or retransmission when sending the data over a lossy link. The transmission policies are determined adaptive to environment variables and in linear computation time. Simulation results show that the proposed transmission system achieves 20-30% reduction on delivering latency compared to the state-of-the-art approach in the literature.
Ghassan Al-Regib, Dihong Tian
ICASSP (5)1
2006 PODs: Partially Ordered Delivery for 3D Scenes in Resource-Constrained Environments
abstract
Three-dimensional (3D) graphic scenes require considerable network bandwidth to be transmitted and computing power to be rendered on users’ terminals. Toward high-quality display in real time, we propose a sender-driven mechanism for streaming 3D scenes in a resource-constrained environment. By pre-processing the database, objects in the scene are properly weighted upon their rendering importance, and their resolutions are selected accordingly to reduce the bit rate. Partially ordered delivery is then performed using decoding independencies between the objects. Simulation results show the efficacy of the proposed mechanism. For a test benchmark, for example, the proposed algorithm outperforms the comparing heuristic by 4 dB under a 100-KB bit rate.
Dihong Tian, Ghassan Al-Regib
ICASSP (5)2
2006 Rate-Constrained Distributed Estimation in Wireless Sensor Networks
abstract
We consider the distributed parameter estimation in wireless sensor networks where a total bit rate constraint is imposed. There is a tradeoff between the number of active sensors and the quantization bit rate for each active sensor to minimize the estimation mean square error (MSE). We first present an optimal distributed estimation algorithm for homogeneous sensor networks and introduce a concept of the equivalent 1-bit MSE function. Then, we propose a quasi-optimal distributed estimation algorithm for heterogeneous sensor networks, which is also based on the equivalent 1-bit MSE function, and the upper bound of the estimation MSE of the proposed algorithm is also addressed. Furthermore, a theoretical lower bound of the estimation MSE under the total bit rate constraint is stated and it is shown that our proposed algorithm is quasi-optimal within a factor 2.2872 of the theoretical lower bound. Simulation results also show that significant reduction in estimation MSE is achieved by the proposed algorithm when compared to other uniform methods.
Ghassan Al-Regib
ICCCN2
2006 Multi-Streaming of Visual Scenes with Scalable Partial Reliability
abstract
Three-dimensional (3D) visual scenes with pluralities of graphic objects require considerable network bandwidth to be transmitted and computing power to be rendered on a user's terminal. This paper proposes an application and transport cross-layer solution for efficiently streaming 3D scenes toward high-quality display. In so doing, 3D objects are scalably coded and are transmitted over multiple streams in a partially sequenced and partially reliable fashion. A rate-distortion optimization framework is developed to determine the optimal level of reliability for every chunk of data in each stream. Compared to heuristical methods, simulation results show that the proposed framework maximizes the display quality of the scene while minimizing the amount of data that needs to be processed by the client's rendering engine.
Ghassan Al-Regib, Dihong Tian
ICIP1
2006 Adaptive Multi-Resolution Coding for 3D Scenes using Vector Quantization
abstract
This paper proposes a multi-resolution coding system for 3D graphic scenes using vector quantization (VQ). For improved compression efficiency, the proposed system generates the VQ codebook for a given 3D scene using mesh objects contained in the scene database instead of separate training models. Furthermore, objects are weighted according to their unequal display importance resulting from interactions, which is incorporated into both codebook training and rate-distortion optimization. Empirical results are presented to verify the efficiency of the proposed coding system.
Dihong Tian, Ghassan Al-Regib
ICIP3
2006 Design of a Transmission Protocol for a CVE
abstract
Virtual reality is a useful tool to train individuals for various different situations or scenarios. A single person utilizing a virtual environment can greatly enhance their skill set, but the greatest gain in skill is obtained when multiple users realistically interact in the virtual environment. To accomplish this, a network protocol must be designed to efficiently communicate vital information to participants in the virtual environment. The purpose of this paper is to document the design of a transmission protocol for a collaborative virtual environment (CVE). In addition, the network protocol was tested with a network emulator to understand how well the protocol functioned under various network conditions. When the jitter level was below 50 ms the transmission protocol was able to reproduce the smooth avatar movement even at transmission intervals as high as 100 ms. However as the jitter increased the transmission interval had to be kept as low as 10 ms to reproduce realistic human movement.
Fred Stakem, Ghassan Al-Regib, Biing-Hwang Juang, Mourad Bouzit
ICIP2
2006 Parity-Object Embedded Streaming for Synthetic Graphics
abstract
This paper proposes an object-oriented approach for streaming synthetic graphic contents over a packet-erasure channel. Graphic objects such as geometric meshes and texture images are coded independently into scalable resolutions and are transmitted in respective sequences to preserve their manipulation independency during packet delivery. For error resilience, FEC codes are generated as "parity objects" parallel to graphic objects. Using parity objects, multiple graphic objects are protected concurrently while also preferentially, in regard to their estimated importance on the display quality. Both the optimization framework and operational algorithm are developed for the generation of parity objects, and simulation results are presented to show the efficacy of the proposed mechanism.
Dihong Tian, Ghassan Al-Regib
ICIP2
2006 Mobility Support for Cooperative Wireless Networks
abstract
In a cooperative network, by combining multiple copies of the same packet that are received either simultaneously (i.e., same time-slot) or non-simultaneously (i.e., different time-slot), the aggregate power level of the cooperatively transmitting set, required to achieve reliable reception at the receiver node, can be considerably lowered. This approach is especially advantageous for networks, where the number of nodes in the network, the network density, or the potential for wireless broadcasting advantage is high. In such networks, required transmit energy can be considerably reduced by routing packets over cooperative paths. Since the success of cooperation strictly depends on the selection of the appropriate power levels for each cooperating node, cooperative routing becomes difficult in networks where the nodes are mobile. In such networks, we need to select paths that can maximize the energy savings averaged over the connection time rather than the ones that only maximize savings at the current time. For this reason, we need to design routing protocols that takes into account the mobility characteristics within the network. In this paper, we propose a cooperative routing protocol that can improve the reliability level in the network against mobility variations.
Aytac Azgin, Yücel Altunbasak, Ghassan Al-Regib
VTC Fall3
2006 On-demand transmission of 3D models over lossy networks
Dihong Tian, Ghassan Al-Regib
Signal Process. Image Commun.2
2006 Vector Quantization in Multiresolution Mesh Compression
abstract
Irregularly sampled triangular meshes contain vertices in a geometric space with arbitrary connectivity degrees. From a compression point of view, this letter presents a multiresolution analysis for irregular meshes using progressive vector quantization (VQ). We show that VQ exploits the joint distribution of coordinates in the geometric space, thus improving the compression efficiency. Based on the presented compression algorithm, we further propose an equalization algorithm that properly balances between connectivity downsampling and geometry quantization and achieves the optimal rate-distortion performance under constrained bit rates.
Dihong Tian, Ghassan Al-Regib
IEEE Signal Process. Lett.3
2005 Cooperative MAC and routing protocols for wireless ad hoc networks
abstract
Cooperative diversity techniques exploit the spatial characteristics of the network to create transmit-diversity, in which the same information can be forwarded through multiple paths towards a single destination or a set of destination nodes. In this paper, we study the integration of cooperative diversity into wireless routing protocols by developing distributed cooperative MAC (C-MAC) and routing protocols. The proposed protocols employ efficient relay selection-coordination and power allocation techniques to maximize the cooperation benefits in the network. Simulation results show that the energy-saving performance of the minimum-energy routing protocols can be significantly improved when they are implemented together with the proposed C-MAC protocol (%50). We also show that the performance of the C-MAC protocol can be further enhanced when the initial path is selected using the cooperation characteristics of the network (%11 more energy-savings compared to the previous case, i.e., C-MAC with minimum-energy routing)
Aytac Azgin, Yücel Altunbasak, Ghassan Al-Regib
GLOBECOM3
2005 Parallel distributed detection for wireless sensor networks: performance analysis and design
abstract
Parallel distributed detection for wireless sensor networks is studied in this paper. The network consists of a set of local sensors and a fusion center. Each local sensor makes a binary (single-bit) or M-ary (multi-bit) decision and passes it to the fusion center where a final decision is made. The links between the local sensors and the fusion center are subject to fading and additive noise resulting in corruption of the transmitted decisions. We analyze the performance of the decision fusion based on likelihood ratio tests and derive false alarm and detection probabilities. Based on the theoretical probability expressions, we design optimal decision rules for the local sensors and the fusion center. Finally, we illustrate the performance of the parallel fusion by numerical examples
Israfil Bahceci, Ghassan Al-Regib, Yücel Altunbasak
GLOBECOM2
2005 Progressive streaming of textured 3D models over bandwidth-limited channels
abstract
The bitstream of a progressively encoded textured model consists of multiple refinement layers. Decoding each layer produces a model with a simplified mesh and a resolution-reduced texture. We have proposed a quality measure that captures the visual fidelity of the multi-resolution textured models (Tian, D.H. and AlRegib, G., Proc. ACM Multimedia 2004, p.684-91, 2004). Based on that quality measure, we consider the problem of streaming progressively encoded textured models over a bandwidth-limited channel. We develop a bit-allocation algorithm that optimally packetizes the source bits in every transmitted data unit, such that the perceptual quality of the model displayed on the client's screen is maximized. Experimental results confirm the effectiveness of the proposed bit-allocation algorithm.
Dihong Tian, Ghassan Al-Regib
ICASSP (2)2
2005 Serial distributed detection for wireless sensor networks
abstract
Serial distributed detection for wireless sensor networks is studied in this paper. Unlike the traditional serial distributed detection schemes where it is assumed that the sensor decision at one stage is known exactly to the subsequent sensor node, we assume that the links between the consecutive sensor nodes are subject to fading and additive noise resulting in corruption of the transmitted decisions. Incorporating the effect of fading in the detection process, a decision fusion rule based on likelihood ratio tests is derived. Optimality of likelihood ratio test is also investigated. By numerical examples, we illustrate the performance of the proposed serial detection scheme and compare it with that of the parallel distributed detection
Israfil Bahceci, Ghassan Al-Regib, Yücel Altunbasak
ISIT2
2005 Bit allocation for joint source and channel coding of progressively compressed 3-D models
abstract
This work presents a joint source and channel coding method for transmission of progressively compressed three-dimensional (3-D) models where the bit budget is allocated optimally. The proposed system uses the Compressed Progressive Mesh (CPM) algorithm to produce a hierarchical bitstream representing different levels of detail (LODs). In addition to optimizing the transmitted bitstream with respect to the channel characteristics, we also choose the optimum combination of quantization and tessellation to maximize the expected decoded model quality. That is, given a 3-D mesh, a total bit budget (B) and a channel packet-loss rate (P/sub LR/), we determine the optimum combination of: 1) the geometry coordinates quantizer step size (l-bit); 2) the number of transmitted batches (L); 3) the total number of channel coding bits (C); and 4) the distribution of these channel coding bits among the transmitted levels (C/sub L/). Experimental results show that, with our unequal error protection approach that uses the proposed bit-allocation algorithm, the decoded model quality degrades more gracefully (compared to either no error protection or equal error protection methods) as the packet-loss rate increases.
Ghassan Al-Regib, Yücel Altunbasak, Russell M. Mersereau
IEEE Trans. Circuits Syst. Video Technol.1
2005 An unequal error protection method for progressively transmitted 3D models
abstract
In this paper, we present a packet-loss resilient system for the transmission of progressively compressed three-dimensional (3D) models. It is based on a joint source and channel coding approach that trades off geometry precision for increased error resiliency to optimize the decoded model quality on the client side. We derive a theoretical framework for the overall system by which the channel packet loss behavior and the channel bandwidth can be directly related to the decoded model quality at the receiver. First, the 3D model is progressively compressed into a base mesh and a number of refinement layers. Then, we assign optimal forward error correction code rates to protect these layers according to their importance to the decoded model quality. Experimental results show that with the proposed unequal error protection approach, the decoded model quality degrades more gracefully (compared to either no error protection or equal error protection methods) as the packet-loss rate increases.
Ghassan Al-Regib, Yücel Altunbasak, Jarek Rossignac
IEEE Trans. Multim.1
2005 3TP: an application-Layer protocol for streaming 3-D models
abstract
This paper addresses the problem of streaming progressively compressed three-dimensional (3-D) models over lossy networks. Out of all encoded packets that can be transmitted, we intelligently choose a subset of packets to be transmitted using transport control protocol in order to meet a distortion constraint, while transmitting the remaining packets using user datagram protocol to minimize the end-to-end delay. We call this new application-layer protocol 3-D models transport protocol. We show the effectiveness of this protocol both experimentally and theoretically. We compare the performance of the proposed protocol with systems that do not optimize transmission according to the content of the encoded bitstream. When the maximum distortion is 30, measured using the Hausdorff distance, we achieve savings in delay time ranging from 39% to 68% for packet-loss rates between 1% and 19%.
Ghassan Al-Regib, Yücel Altunbasak
IEEE Trans. Multim.1
2005 Error-resilient transmission of 3D models
abstract
In this article, we propose an error-resilient transmission method for progressively compressed 3D models. The proposed method is scalable with respect to both channel bandwidth and channel packet-loss rate. We jointly design source and channel coders using a statistical measure that (i) calculates the number of both source and channel coding bits, and (ii) distributes the channel coding bits among the transmitted refinement levels in order to maximize the expected decoded model quality. In order to keep the total number of bits before and after applying error protection the same, we transmit fewer triangles in the latter case to accommodate the channel coding bits. When the proposed method is used to transmit a typical model over a channel with a 10% packet-loss rate, the distortion (measured using the Hausdorff distance between the original and the decoded models) is reduced by 50% compared to the case when no error protection is applied.
Ghassan Al-Regib, Yücel Altunbasak, Jarek Rossignac
ACM Trans. Graph.1
2004 A low-complexity video encoder with decoder motion estimator
abstract
We investigate the use of video compression methods that require a simple encoder and a complex decoder for applications such as video surveillance, smart spaces, and sensor networks. In our proposed method, the encoder tries to identify the locations whose content cannot be predicted at the decoder, and codes such areas at higher fidelity. Typically, high-motion macro-blocks represent such significant state regions. A shape-adaptive (SA) SPIHT (set partitioning in hierarchical trees) encoder is then used to code these regions efficiently. On the decoder side, we perform motion extrapolation using the previously decoded frames to construct an estimate for the current frame. This estimate then serves as the side information at the decoder. Receiving the SA-SPIHT coded and the motion extrapolated frame, the decoder fuses the information in these two frames to produce a frame that is of higher quality than the component images. Experimental results show that the proposed codec outperforms H.264 intra mode by 3 dB.
Sibel Yaman, Ghassan Al-Regib
ICASSP (3)2
2004 Packetized media streaming over multiple wireless channels
abstract
This paper addresses the problem of streaming packetized media through a proxy server to a mobile client over multiple wireless channels. We follow a rate-distortion optimized streaming approach. A real-time sender-driven streaming framework is employed at the proxy server to maximize the playback quality at the client. The experimental results show that the proposed algorithm improves the video quality by 1.3-3dB as compared to the conventional system which simply transmits the video data in the order they are displayed.
Dihong Tian, Yen-Chi Lee, Ghassan Al-Regib, Yücel Altunbasak
ICC3
2004 FQM: a fast quality measure for efficient transmission of textured 3D models
abstract
In this paper, we propose an efficient transmission method to stream textured 3D models. We develop a bit-allocation algorithm that distributes the bit budget between the geometry and the mapped texture to maximize the quality of the model displayed on the client's screen. Both the geometry and the texture are progressively and independently compressed. The resolutions for the geometry and the texture are selected to maximize the quality for a given bitrate. We further propose a novel and fast quality measure (FQM) to quantify the perceptual fidelity of the simplified model. Experimental results demonstrate the effectiveness of the proposed bit-allocation algorithm using FQM. For example, when the bit budget is 10KB, the quality of the Zebra model is improved by 15% using the proposed method compared to distributing the bit budget equally between the geometry and the texture.
Dihong Tian, Ghassan Al-Regib
ACM Multimedia2
2004 Optimal packet scheduling for wireless video streaming with error-prone feedback
abstract
In wireless video transmission, burst packet errors generally produce more catastrophic results than equal number of isolated errors. To miniimize the playback distortion it is crucial for the sender to know the packet errors at the receiver and then optimally schedule next transmissions. Unfortunately, in practice, feedback errors result in inaccurate observations of the receiving status. In this paper, we develop an optimal scheduling framework to minimize the expected distortion by first estimating the receiving status. Then, we jointly consider the source and channel characteristics and optimally choose the packets to transmit. The optimal transmission strategy is computed through a partially observable Markov decision process. The experimental results show that the proposed framework improves the average peak signal-to-noise ratio (PSNR) by 0.6-1.3 dB upon using a traditional system without packet scheduling. Moreover, we show that the proposed method smoothes out the bursty distortion periods and results in less fluctuating PSNR values.
Dihong Tian, Xiaohuan Li 0003, Ghassan Al-Regib, Yücel Altunbasak, Joel R. Jackson
WCNC3
2003 3TP: an application-layer protocol for streaming 3-D graphics
abstract
This paper addresses the problem of streaming progressively compressed 3-D models over lossy networks. Out of all encoded packets that can be transmitted, we intelligently choose a subset of packets to be transmitted using TCP in order to meet a distortion constraint, while transmitting the remaining packets using UDP to minimize the end-to-end delay. We call this new application-layer protocol 3TP (3-D models transport protocol). Experimental results show savings of 42% and 68% in delay time at packet-loss rates of 6% and 19%, respectively, compared to systems that do not optimize transmission according to the encoded bitstream content.
Ghassan Al-Regib, Yücel Altunbasak
ICME1
2003 Hierarchical motion estimation with content-based meshes
abstract
Two-dimensional mesh-based models provide a good alternative to motion estimation and compensation. The estimation of the best node-point motion vectors is a challenging task. To this effect, Nakaya and Harashima (1994) proposed a hexagonal matching procedure. Toklu et al. improved the hexagonal search algorithm in terms of both motion-estimation accuracy and computational complexity by employing a hierarchy of regular meshes. Recognizing the limitations of regular meshes, Van Beek et al. (1997) extended Toklu et al.'s (1996) work by utilizing content-based meshes. Here, we provide an alternative hierarchical motion-estimation method with content-based meshes where hierarchical representations are employed for both the images and the irregular meshes in order to provide further improvements in computational complexity as well as motion accuracy. The comparison results are provided with real video sequences.
Ghassan Al-Regib, Yücel Altunbasak, Russell M. Mersereau
IEEE Trans. Circuits Syst. Video Technol.1
2002 An unequal error protection method for progressively compressed 3-D meshes
abstract
In this paper, we present a packet-loss resilient 3-D graphics transmission system that is scalable with respect to both channel bandwidth and channel error characteristics. The algorithm trades off source coding efficiency for increased bit-stream error resilience to optimize the decoded mesh quality on the client side. It uses the Compressed Progressive Mesh (CPM) algorithm to generate a hierarchical bit-stream representing different levels of details (LODs). We assign optimal forward error correction (FEC) code rates to protect different parts of the bit-stream differently. These optimal FEC code rates are determined theoretically via a distortion function that accounts for: the channel packet loss rate, the nature of the encoded 3-D mesh and the error protection bit-budget. We present experimental results, which show that with our unequal error protection (UEP) optimal approach, the decoded mesh quality degrades more gracefully (compared to either no error protection (NEP) or equal error protection (EEP) methods) as the packet loss rate increases.
Ghassan Al-Regib, Yücel Altunbasak, Jarek Rossignac
ICASSP1
2002 A joint source and channel coding approach for progressively compressed 3-D mesh transmission
abstract
In this paper, we present an unequal error protection method for packet-loss resilient transmission of progressively compressed 3D meshes. The proposed method is based on a source and channel coding approach where we set up a theoretical framework for the overall system by which the channel packet loss behavior and the channel bandwidth can be directly related to the decoded mesh quality at the receiver. In particular, we develop a statistical distortion measure and optimize it to compute the best combination of (i) the number of triangles to transmit, (ii) the total number of channel coding bits, and (iii) the distribution of these error-protection bits among the transmitted layers in order to maximize the expected decoded mesh quality at the receiver. The proposed method differs from the earlier approaches in two major aspects: (i) determination of the number of channel coding bits (C) and (ii) the approach of reducing the source rate in order to accommodate for channel coding bits. When the proposed method is used to transmit a typical 3D mesh over a channel with a 10% packet loss rate, the distortion (measured using the Hausdorff distance between the original and the decoded overly sampled meshes) is reduced by 50% compared to the case when no error protection is applied.
Ghassan Al-Regib, Yücel Altunbasak, Jarek Rossignac
ICIP (2)1
2002 Protocol for streaming compressed 3-D animations over lossy channels
abstract
We propose a protocol for efficient streaming of 3-D animations over lossy channels. In order to improve the expected quality on the client's side, we first transmit a crude model of the 3-D mesh, of its texture, and of its immediate evolution. We endow this data with significant error-protection against transmission error. Then we transmit a series of upgrades that refine the accuracy of the model and/or of the animation. These are encoded with lower levels of error protection that is proportional to their impact on the quality of the upgraded animation. We propose the following types of upgrade chunks: selection of a subset of vertices, adjustments of the accuracy of the positions of the selected vertices, connectivity refinements, motion adjustments of the selected subset of vertices, adjustments of the accuracy of the texture coordinates of the selected vertices, and upgrade of the quality of the texture. Finally, given the allocated bit-budget, we determine the optimal number of both source and channel coding bits assigned for each chunk to maximize the animation quality on the client's side. The optimization takes into account both the bandwidth and the error characteristics of the channel.
Ghassan Al-Regib, Yücel Altunbasak, Jarek Rossignac, Russell M. Mersereau
ICME (1)1
2002 An Unequal Error Protection Method for Packet Loss Resilient 3-D Mesh Transmission
abstract
A packet-loss resilient, bandwidth-scalable 3D graphics streaming system is proposed. It uses the compressed progressive mesh (CPM) algorithm (see Pajarola, R. and Rossignac, J., IEEE Trans. on Visualization and Computer Graphics, vol.6, no.1, p.79-93, 2000) to generate a hierarchical bit-stream to represent different levels of details (LODs). We assign forward error correction (FEC) codes to each layer in the encoded bit-stream according to its importance. To this end, we develop a new distortion measure that quantifies the distortion in the reconstructed mesh when part of the bit-stream is lost. Experimental results with a collection of 3D meshes have been conducted to demonstrate the efficacy of our solution. They indicate that with our proposed unequal error protection (UEP), the decoded mesh quality degrades more gracefully, compared to either no error protection (NEP) or equal error protection (EEP) methods, as the packet loss rate increases.
Ghassan Al-Regib, Yücel Altunbasak
INFOCOM1
2001 2-D motion estimation with hierarchical content based meshes
abstract
Two-dimensional mesh-based models provide a good alternative to motion estimation and compensation. The estimation of the best motion vectors at the node-point motion vectors is a challenging task. To this effect, Nakaya et al. (1994), proposed a hexagonal matching procedure. Toklu et al. (1996) improved the hexagonal search algorithm in terms of both motion estimation accuracy and computational complexity by employing a hierarchy of regular meshes. Recognizing the limitations of regular meshes, Van Beek et al. (1999), extended Toklu's work by utilizing content-based meshes. Here, we provide an alternative hierarchical motion estimation method with content-based meshes where hierarchical representations are employed for images as well as meshes in order to provide further improvements in computational complexity as well as motion accuracy.
Ghassan Al-Regib, Yücel Altunbasak
ICASSP1
1999 Improved Selective Encryption Techniques for Secure Transmission of Mpeg Video Bit-Streams
abstract
This paper presents three new selective encryption techniques for secure transmission of MPEG-I video bit-streams. These techniques maintain higher security levels than previously proposed selective encryption techniques while maintaining reasonable processing times. In the first of these methods, the encryption is applied to the data associated with every n/sup th/ I-macroblock. In the second method, the encryption is applied to the headers of all the predicted macroblocks as well as to the data associated with every n/sup th/ I-macroblock. In the third method, encryption is applied to every n/sup th/ I-macroblock as well as the header of every n/sup th/ predicted macroblock. The last method, with n=2, is found to be the most efficient of the three proposed methods. This method achieves a 60-82% reduction in the processing time over "total" encryption, and simulation results show that the encrypted/decoded video is fully disguised.
Adnan M. Alattar, Ghassan Al-Regib, Saud Al-Semari
ICIP (4)2