VLDB 2026 Research / reviewers in the wild / expert
Andreas Spanias
dblp:38/4969 · also Andreas S. Spanias
· DBLP profile ↗
163ranked-venue papers
10as first author
19since 2021 · last 2025
0000-0003-0306-9348ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 100 · 6 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 22 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 21 · 7 since 2021Systems, architecture and hardware · 11 · 1 since 2021Computer networks · 7Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Implementation of Quantum Machine Learning on Educational Data
Sofía Ramos-Pulido, Neil Hernández-Gress, Glen S. Uehara, Andreas Spanias, Hector G. Ceballos |
ICAART (3) | 4 |
| 2025 | Quantum-Enhanced Cancer Detection for Histopathologic ImagesabstractThis paper investigates quantum image processing for histopathological cancer detection, focusing on enhanced feature extraction and improved classification. We propose a hybrid quantum-classical framework in which quantum preprocessing circuits, specifically, Quantum Hadamard Edge Detection (QHED) and the Quantum Fourier Transform (QFT), extract subtle spatial and spectral features from hematoxylin and eosin-stained whole-slide images. These quantum-enhanced features are then classified using deep neural networks such as ResNet-50, VGG-16, and UNet. Experiments on the PatchCamelyon dataset demonstrate that quantum-enhanced preprocessing increases sensitivity in detecting cancerous regions, while integrating hybrid preprocessing with an SVM classifier further improves classification accuracy. Although quantum noise remains a challenge, our results highlight the promise of quantum image pre-processing and feature extraction for advancing histopathological image classification. Nandika Goyal, Glen S. Uehara, Andreas Spanias |
ICIP | 3 |
| 2025 | Image Analysis-Synthesis Using the Quantum Fourier TransformabstractThis paper introduces a hybrid two-dimensional Quantum Fourier Transform (2D hybrid QFT) method for image analysis and synthesis. The QFT and Inverse QFT (IQFT) functions enable analysis-synthesis implementation using frequency spectrum-based selection approaches for compression. We study the resolution and precision of the QFT by comparing its outputs against classical two-dimensional Fast Fourier Transform (FFT) techniques using pixel-level Signal to Noise Ratio (SNR) as a metric. Formative and summative assessments have been deployed as a laboratory exercise in two National Science Foundation (NSF) workforce development programs. Preliminary evaluation results are presented in this paper. Tanay Kamlesh Patel, Danielle Knutson, Glen S. Uehara, Frank Marfai, Carly Jazwin, Jean S. Larson, Andreas Spanias |
ISCAS | 7 |
| 2024 | WIP: Building a Research Experience for Undergraduates in Quantum Machine LearningabstractThis work in progress research-to-practice study describes the development of a new undergraduate research training site on Quantum Machine Learning (QML), hosted at Arizona State University, a large Hispanic-Serving Institution. The objectives of this project are to a) recruit and prepare students from diverse pathways to increase representation of those traditionally underrepresented in QML research, b) increase awareness of career opportunities in the QML field, c) engage students in theoretical and experimental quantum information processing and machine learning (ML), d) motivate students to continue QML research into graduate school, and e) provide professional development training including presenting to stakeholders, developing publications/patents, and building an awareness on social implications, ethics, and privacy. The project adopts an integrative theory, application, and hands-on training approach by immersing undergraduate students in ML algorithm and quantum computing studies with hands-on quantum circuit design tasks. Participants are embedded in research labs, guided by graduate students and faculty mentors on quantum computing research studies. The program is evaluated by both the Center for Evaluating the Research Pipeline (CERP) and an independent evaluator. Formative and summative assessments include pre- and post-surveys, a mid-point check-in survey, and a document review of program deliverables. Findings are described in a final evaluation report. This paper describes the importance of introducing QML research at the undergraduate level, methods for recruiting a diverse group of participants, program format, research projects, and preliminary program evaluation results. Jean S. Larson, Deep Pujara, David F. Ramirez, Leslie Miller, Tanay Kamlesh Patel, Niraj Anil Babar, Andreas Spanias |
FIE | 7 |
| 2023 | Introducing Quantum Computing in a Sophomore Signals and Systems CourseabstractThis Innovative Practice Work in Progress Paper describes the development and assessment of a web-based simulation lab exercise introducing basic quantum computing concepts in a sophomore signals and systems course. Specifically, students make the connection between quantum computing and signals and systems theory through a comparative study of using Quantum Fourier Transforms and Fast Fourier Transforms for a speech analysis-synthesis application. In addition, quantum noise models are introduced and simulated to show their effect on computation performance. Statistics from pre/post quizzes show that there is significant knowledge improvement by completing the lab exercise. Aradhita Sharma, Glen S. Uehara, Leslie Miller, Deep Pujara, Wendy M. Barnard, Jean S. Larson, Andreas Spanias |
FIE | 8 |
| 2023 | Signal Analysis-Synthesis Using the Quantum Fourier TransformabstractThis paper presents the development of Quantum Fourier transform (QFT) education tools in the object-oriented Java-DSP (J-DSP) simulation environment. More specifically, QFT and Inverse QFT (IQFT) user-friendly J-DSP functions are developed to expose undergraduate students to quantum computing. These functions provide opportunities to examine QFT resolution, precision (qubits), and the effects of quantum measurement noise. In our study, we also describe a laboratory exercise on QFT-based speech analysis-synthesis which has been deployed in our senior-level DSP class and in our NSF workforce development programs. The software and the laboratory exercise are evaluated using formative and summative assessments. Aradhita Sharma, Glen S. Uehara, Vivek Sivaraman Narayanaswamy, Leslie Miller, Andreas Spanias |
ICASSP | 5 |
| 2023 | AN L2-Normalized Spatial Attention Network for Accurate and Fast Classification of Brain Tumors in 2D T1-Weighted CE-MRI ImagesabstractWe propose an accurate and fast classification network for classification of brain tumors in MRI images that outperforms all lightweight methods investigated in terms of accuracy. We test our model on a challenging 2D T1-weighted CE-MRI dataset containing three types of brain tumors: Meningioma, Glioma and Pituitary. We introduce an l2-normalized spatial attention mechanism that acts as a regularizer against overfitting during training. We compare our results against the state-of-the-art on this dataset and show that by integrating l2-normalized spatial attention into a baseline network we achieve a performance gain of 1.79 percentage points. Even better accuracy can be attained by combining our model in an ensemble with the pretrained VGG16 at the expense of execution speed. Our code is publicly available at https://github.com/juliadietlmeier/MRI_image_classification. Grace Billingsley, Julia Dietlmeier, Vivek Sivaraman Narayanaswamy, Andreas Spanias, Noel E. O'Connor |
ICIP | 4 |
| 2023 | Software-Defined Imaging: A SurveyabstractHuge advancements have been made over the years in terms of modern image-sensing hardware and visual computing algorithms (e.g., computer vision, image processing, and computational photography). However, to this day, there still exists a current gap between the hardware and software design in an imaging system, which silos one research domain from another. Bridging this gap is the key to unlocking new visual computing capabilities for end applications in commercial photography, industrial inspection, and robotics. In this survey, we explore existing works in the literature that can be leveraged to replace conventional hardware components in an imaging system with software for enhanced reconfigurability. As a result, the user can program the image sensor in a way best suited to the end application. We refer to this as software-defined imaging (SDI), where image sensor behavior can be altered by the system software depending on the user’s needs. The scope of our survey covers imaging systems for single-image capture, multi-image, and burst photography, as well as video. We review works related to the sensor primitives, image signal processor (ISP) pipeline, computer architecture, and operating system elements of the SDI stack. Finally, we outline the infrastructure and resources for SDI systems, and we also discuss possible future research directions for the field. Suren Jayasuriya, Odrika Iqbal, Venkatesh Kodukula, Victor Isaac Torres Muro, Robert LiKamWa, Andreas Spanias |
Proc. IEEE | 6 |
| 2022 | Undergraduate Research and Education in Quantum Machine LearningabstractThis Work-In-Progress paper describes a program in quantum machine learning launched in the academic year of 2021-22. The program engaged undergraduate students from STEM areas with faculty and industry mentors. Because of the COVID-19 conditions, this undergraduate engagement was offered in a virtual format. In 2022, some face-to-face meetings with presentations were also held. The program included: a) training in machine learning with quantum simulators, b) weekly presentations, and c) semester end presentations. The assessment of the program included surveys, interviews, and presentation observations. Challenges and opportunities from virtual engagement were also part of the assessment. Glen S. Uehara, Jean S. Larson, Wendy M. Barnard, Michael Esposito, Filippo Posta, Maxwell Yarter, Aradhita Sharma, Niki Kyriacou, Matthew Dobson, Andreas Spanias |
FIE | 10 |
| 2022 | Predicting the Generalization Gap in Deep Models using AnchoringabstractWe address the problem of predicting the generalization gap of deep neural networks under large, natural, and synthetic distribution shifts between source and target domains. This is crucial in understanding how models behave in uncontrollable ‘in-the-wild’ scenarios, but existing techniques fail when target domain becomes very different from the source. Accurately capturing the relationship and distance between the source and target domains is critical for a reliable post-hoc estimation of generalization. In this paper, we propose a novel strategy for directly predicting accuracy on unseen target data with the help of anchoring and pre-text encoding in predictive models. Anchoring has been shown previously to perform effectively in characterizing domain shifts, which we exploit for predicting the generalization gap. Our experiments on the PACS dataset along with synthetic ablations indicate that our approach produces well calibrated accuracy estimates outperforming existing baselines. Vivek Sivaraman Narayanaswamy, Rushil Anirudh, Irene Kim, Yamen Mubarka, Andreas Spanias, Jayaraman J. Thiagarajan |
ICASSP | 5 |
| 2022 | Improved StyleGAN-v2 based Inversion for Out-of-Distribution ImagesabstractInverting an image onto the latent space of pre-trained generators, e.g., StyleGAN-v2, has emerged as a popular strategy to leverage strong image priors for ill-posed restoration. Several studies have showed that this approach is effective at inverting images similar to the data used for training. However, with out-of-distribution (OOD) data that the generator has not been exposed to, existing inversion techniques produce sub-optimal results. In this paper, we propose SPHInX (StyleGAN with Projection Heads for Inverting X), an approach for accurately embedding OOD images onto the StyleGAN latent space. SPHInX optimizes a style projection head using a novel training strategy that imposes a vicinal regularization in the StyleGAN latent space. To further enhance OOD inversion, SPHInX can additionally optimize a content projection head and noise variables in every layer. Our empirical studies on a suite of OOD data show that, in addition to producing higher quality reconstructions over the state-of-the-art inversion techniques, SPHInX is effective for ill-posed restoration tasks while offering semantic editing capabilities. Rakshith Subramanyam, Vivek Sivaraman Narayanaswamy, Mark Naufel, Andreas Spanias, Jayaraman J. Thiagarajan |
ICML | 4 |
| 2021 | Uncertainty-Matching Graph Neural Networks to Defend Against Poisoning AttacksabstractGraph Neural Networks (GNNs), a generalization of neural networks to graph-structured data, are often implemented using message passes between entities of a graph. While GNNs are effective for node classification, link prediction and graph classification, they are vulnerable to adversarial attacks, i.e., a small perturbation to the structure can lead to a non-trivial performance degradation. In this work, we propose Uncertainty Matching GNN (UM-GNN), that is aimed at improving the robustness of GNN models, particularly against poisoning attacks to the graph structure, by leveraging epistemic uncertainties from the message passing framework. More specifically, we propose to build a surrogate predictor that does not directly access the graph structure, but systematically extracts reliable knowledge from a standard GNN through a novel uncertainty-matching strategy. Interestingly, this uncoupling makes UM-GNN immune to evasion attacks by design, and achieves significantly improved robustness against poisoning attacks. Using empirical studies with standard benchmarks and a suite of global and target attacks, we demonstrate the effectiveness of UM-GNN, when compared to existing baselines including the state-of-the-art robust GCN. Uday Shankar Shanthamallu, Jayaraman J. Thiagarajan, Andreas Spanias |
AAAI | 3 |
| 2021 | Accurate and Robust Feature Importance Estimation under Distribution ShiftsabstractWith increasing reliance on the outcomes of black-box models in critical applications, post-hoc explainability tools that do not require access to the model internals are often used to enable humans understand and trust these models. In particular, we focus on the class of methods that can reveal the influence of input features on the predicted outputs. Despite their wide-spread adoption, existing methods are known to suffer from one or more of the following challenges: computational complexities, large uncertainties and most importantly, inability to handle real-world domain shifts. In this paper, we propose PRoFILE (Producing Robust Feature Importances using Loss Estimates), a novel feature importance estimation method that addresses all these challenges. Through the use of a loss estimator jointly trained with the predictive model and a causal objective, PRoFILE can accurately estimate the feature importance scores even under complex distribution shifts, without any additional re-training. To this end, we also develop learning strategies for training the loss estimator, namely contrastive and dropout calibration, and find that it can effectively detect distribution shifts. Using empirical studies on several benchmark image and non-image data, we show significant improvements over state-of-the-art approaches, both in terms of fidelity and robustness. Jayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Rushil Anirudh, Peer-Timo Bremer, Andreas Spanias |
AAAI | 5 |
| 2021 | Research Experiences for Teachers in Machine LearningabstractMachine learning and Artificial Intelligence (AI) are national priority areas for research, education and workforce development. This work in progress paper describes a Research Experiences for Teachers program in sensors and machine learning launched in the summer of 2020. Motivated by national AI workforce needs, we designed a program that engaged high school teachers from STEM fields in machine learning research. In 2020, the program focused on AI algorithms for solar energy systems. Because of the COVID-19 conditions, the research experience was virtual and ran with a smaller teacher group than originally planned. The program included development of training content, algorithm and software training, research in solar energy monitoring, development of research reports and lesson plans, research presentations, and assessment. The assessment of the program included surveys, interviews, presentation observations, and follow-up in high school content delivery. Kristen Jaskie, Jean S. Larson, Milton Johnson, Kathy Turner, Megan A. O'Donnell, Jennifer Blain Christen, Sunil Rao, Andreas Spanias |
FIE | 8 |
| 2021 | Experiences with Web-based Signal Analysis Laboratories and Online Training during the COVID-19 PeriodabstractThis work in progress paper describes our efforts and challenges in delivering undergraduate and graduate courses during COVID-19 conditions. More specifically, we focus on the adaptation and delivery of digital signal analysis laboratories for all the remote learners during the pandemic conditions. Methods for online labs and workforce training have been developed and deployed on a virtual basis. These labs and simulation environments have been deployed in signals and systems and DSP classes as well as in workforce development programs such as the REU and RET. The assessment of these efforts included evaluation forms and interviews. Challenges and opportunities from virtual delivery of content and labs were also part of the assessment. Vivek Sivaraman Narayanaswamy, Photini Spanias, Sunil Rao, Andreas Spanias |
FIE | 4 |
| 2021 | Using Deep Image Priors to Generate Counterfactual ExplanationsabstractThrough the use of carefully tailored convolutional neural network architectures, a deep image prior (DIP) can be used to obtain pre-images from latent representation encodings. Though DIP inversion has been known to be superior to conventional regularized inversion strategies such as total variation, such an over-parameterized generator is able to effectively reconstruct even images that are not in the original data distribution. This limitation makes it challenging to utilize such priors for tasks such as counterfactual reasoning, wherein the goal is to generate small, interpretable changes to an image that systematically leads to changes in the model prediction. To this end, we propose a novel regularization strategy based on an auxiliary loss estimator jointly trained with the predictor, which efficiently guides the prior to re-cover natural pre-images. Our empirical studies with a real-world ISIC skin lesion detection problem clearly evidence the effectiveness of the proposed approach in synthesizing meaningful counterfactuals. In comparison, we find that the standard DIP inversion often proposes visually imperceptible perturbations to irrelevant parts of the image, thus providing no additional insights into the model behavior. Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Andreas Spanias |
ICASSP | 3 |
| 2021 | On the Design of Deep Priors for Unsupervised Audio RestorationabstractUnsupervised deep learning methods for solving audio restoration problems extensively rely on carefully tailored neural architectures that carry strong inductive biases for defining priors in the time or spectral domain. In this context, lot of recent success has been achieved with sophisticated convolutional network constructions that recover audio signals in the spectral domain. However, in practice, audio priors require careful engineering of the convolutional kernels to be effective at solving ill-posed restoration tasks, while also being easy to train. To this end, in this paper, we propose a new U-Net based prior that does not impact either the network complexity or convergence behavior of existing convolutional architectures, yet leads to significantly improved restoration. In particular, we advocate the use of carefully designed dilation schedules and dense connections in the U-Net architecture to obtain powerful audio priors. Using empirical studies on standard benchmarks and a variety of ill-posed restoration tasks, such as audio denoising, in-painting and source separation, we demonstrate that our proposed approach consistently outperforms widely adopted audio prior architectures. Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Andreas Spanias |
Interspeech | 3 |
| 2021 | Designing Counterfactual Generators using Deep Model InversionabstractExplanation techniques that synthesize small, interpretable changes to a given image while producing desired changes in the model prediction have become popular for introspecting black-box models. Commonly referred to as counterfactuals, the synthesized explanations are required to contain discernible changes (for easy interpretability) while also being realistic (consistency to the data manifold). In this paper, we focus on the case where we have access only to the trained deep classifier and not the actual training data. While the problem of inverting deep models to synthesize images from the training distribution has been explored, our goal is to develop a deep inversion approach to generate counterfactual explanations for a given query image. Despite their effectiveness in conditional image synthesis, we show that existing deep inversion methods are insufficient for producing meaningful counterfactuals. We propose DISC (Deep Inversion for Synthesizing Counterfactuals) that improves upon deep inversion by utilizing (a) stronger image priors, (b) incorporating a novel manifold consistency objective and (c) adopting a progressive optimization strategy. We find that, in addition to producing visually meaningful explanations, the counterfactuals from DISC are effective at learning classifier decision boundaries and are robust to unknown test-time corruptions. Jayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Deepta Rajan, Jason Liang, Akshay Chaudhari, Andreas Spanias |
NeurIPS | 6 |
| 2021 | Coverage-Based Designs Improve Sample Mining and Hyperparameter Optimization
Gowtham Muniraju, Bhavya Kailkhura, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Cihan Tepedelenlioglu, Andreas Spanias |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | A Regularized Attention Mechanism for Graph Attention NetworksabstractMachine learning models that can exploit the inherent structure in data have gained prominence. In particular, there is a surge in deep learning solutions for graph-structured data, due to its wide-spread applicability in several fields. Graph attention networks (GAT), a recent addition to the broad class of feature learning models in graphs, utilizes the attention mechanism to efficiently learn continuous vector representations for semi-supervised learning problems. In this paper, we perform a detailed analysis of GAT models, and present interesting insights into their behavior. In particular, we show that the models are vulnerable to heterogeneous rogue nodes and hence propose novel regularization strategies to improve the robustness of GAT models. Using benchmark datasets, we demonstrate performance improvements on semi-supervised learning, using the proposed robust variant of GAT. Uday Shankar Shanthamallu, Jayaraman J. Thiagarajan, Andreas Spanias |
ICASSP | 3 |
| 2020 | Design and FPGA Implementation of an Adaptive video Subsampling Algorithm for Energy-Efficient Single Object TrackingabstractImage sensors with programmable region-of-interest (ROI) readout are a new sensing technology important for energyefficient embedded computer vision. In particular, ROIs can subsample the number of pixels being readout while performing single object tracking in a video. In this paper, we develop adaptive sampling algorithms which perform joint object tracking and predictive video subsampling. We utilize an object detection consisting of either mean shift tracking or a neural network, coupled with a Kalman filter for prediction. We show that our algorithms achieve mean average precision of 0.70 or higher on a dataset of 20 videos in software. Further, we implement hardware acceleration of mean shift tracking with Kalman filter adaptive subsampling on an FPGA. Hardware results show a 23 × improvement in clock cycles and latency as compared to baseline methods and achieves 38FPS real-time performance. This research points to a new domain of hardware-software co-design for adaptive video subsampling in embedded computer vision. Odrika Iqbal, Saquib Siddiqui, Joshua Martin, Sameeksha Katoch, Andreas Spanias, Daniel W. Bliss, Suren Jayasuriya |
ICIP | 5 |
| 2020 | Unsupervised Audio Source Separation Using Generative PriorsabstractState-of-the-art under-determined audio source separation systems rely on supervised end-end training of carefully tailored neural network architectures operating either in the time or the spectral domain. However, these methods are severely challenged in terms of requiring access to expensive source level labeled data and being specific to a given set of sources and the mixing process, which demands complete re-training when those assumptions change. This strongly emphasizes the need for unsupervised methods that can leverage the recent advances in data-driven modeling, and compensate for the lack of labeled data through meaningful priors. To this end, we propose a novel approach for audio source separation based on generative priors trained on individual sources. Through the use of projected gradient descent optimization, our approach simultaneously searches in the source-specific latent spaces to effectively recover the constituent sources. Though the generative priors can be defined in the time domain directly, e.g. WaveGAN, we find that using spectral domain loss functions for our optimization leads to good-quality source estimates. Our empirical studies on standard spoken digit and instrument datasets clearly demonstrate the effectiveness of our approach over classical as well as state-of-the-art unsupervised baselines. Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Rushil Anirudh, Andreas Spanias |
INTERSPEECH | 4 |
| 2020 | Consensus Based Distributed Spectral Radius EstimationabstractA consensus based distributed algorithm to compute the spectral radius of a network is proposed. The spectral radius of the graph is the largest eigenvalue of the adjacency matrix, and is a useful characterization of the network graph. Conventionally, centralized methods are used to compute the spectral radius, which involves eigenvalue decomposition of the adjacency matrix of the underlying graph. Our distributed algorithm uses a simple update rule to reach consensus on the spectral radius, using only local communications. We consider time-varying graphs to model packet loss and imperfect transmissions, and provide the convergence characteristics of our algorithm, for both static and time-varying graphs. We prove that the convergence error is a function of principal eigenvector of adjacency matrix of the graph and reduces as O(1/t), where t is the number of iterations. The algorithm works for any connected graph structure. Simulation results supporting the theory are also presented. Gowtham Muniraju, Cihan Tepedelenlioglu, Andreas Spanias |
IEEE Signal Process. Lett. | 3 |
| 2020 | GrAMME: Semisupervised Learning Using Multilayered Graph Attention ModelsabstractModern data analysis pipelines are becoming increasingly complex due to the presence of multiview information sources. While graphs are effective in modeling complex relationships, in many scenarios, a single graph is rarely sufficient to succinctly represent all interactions, and hence, multilayered graphs have become popular. Though this leads to richer representations, extending solutions from the single-graph case is not straightforward. Consequently, there is a strong need for novel solutions to solve classical problems, such as node classification, in the multilayered case. In this article, we consider the problem of semisupervised learning with multilayered graphs. Though deep network embeddings, e.g., DeepWalk, are widely adopted for community discovery, we argue that feature learning with random node attributes, using graph neural networks, can be more effective. To this end, we propose to use attention models for effective feature learning and develop two novel architectures, GrAMME-SG and GrAMME-Fusion, that exploit the interlayer dependences for building multilayered graph embeddings. Using empirical studies on several benchmark data sets, we evaluate the proposed approaches and demonstrate significant performance improvements in comparison with the state-of-the-art network embedding strategies. The results also show that using simple random features is an effective choice, even in cases where explicit node attributes are not available. Uday Shankar Shanthamallu, Jayaraman J. Thiagarajan, Huan Song, Andreas Spanias |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Introducing machine learning concepts using hands-on Android-based exercisesabstractIn this innovative practice work-in-progress paper, we discuss novel methods to teach machine learning concepts to undergraduate students. Teaching machine learning involves introducing students to complex concepts in statistics, linear algebra, and optimization. In order for students to better grasp concepts in machine learning, we provide them with hands-on exercises. These types of immersive experiences will expose students to the different stages of the practical uses of machine learning. The data collection apparatus is based on applications (apps) developed for the Android platform. Due to the accessible nature of the app and the exercises based on the app, this approach is useful for students across all majors.We provide the students with three different sets of activities, the first of which will introduce the basics of machine learning with specially designed artificial datasets. The second and third activities involve data collection, modeling, training, and testing, as applied to machine learning algorithms. The second activity will involve collecting touch/swipe data on mobile devices from students as they use a touch logger app. The third activity uses the Reflections app to collect cross-correlation data from rooms with different purposes. These hands-on activities guide the students through every step of the machine learning process. Student learning is assessed for each activity by holding workshops for undergraduate students. A workshop with the first activity outlining the basics of machine learning was given in the fall of 2018 and significant student learning was demonstrated. Workshops for the second and third activities are planned for the fall semester of 2019. Results from these workshops will be presented at the conference. Blaine Ayotte, Justin Au-Yeung, Mahesh K. Banavar, Dana M. Barry, Gowtham Muniraju, Sunil Rao, Andreas Spanias, Cihan Tepedelenlioglu |
FIE | 7 |
| 2019 | An REU Experience in Machine Learning and Computational CamerasabstractIn this work in progress paper, we describe an REU summer experience on imaging sensors that involved a female junior level Electrical Engineering student, a graduate student advisor, and three faculty. A research plan was designed to embed the student in a sensor and machine learning research with specific emphasis on energy-efficient cameras. The motivation for submitting this paper is the unique planning and the quality of the overall student experience which resulted in continuous engagement of the REU student with the faculty after the REU summer program completed. The program resulted in a major presentation at an international event, an NSF I/UCRC poster presentation, a research conference submission which is remarkable for an undergraduate student, and finally a new research direction for the graduate mentor and faculty. This paper describes successful strategies for research engagement for undergraduates in state-of-the-art research fields which yield positive outcomes for all participants, and is grounded in contemporary educational methodology and theoretical frameworks. Divya Mohan, Sameeksha Katoch, Suren Jayasuriya, Pavan Turaga, Andreas Spanias |
FIE | 5 |
| 2019 | Introducing Machine Learning in a Sophomore Signals and Systems CourseabstractThis Innovative Practice Work in Progress Paper describes the experience and assessment of introducing machine learning concepts in a sophomore signals and systems course. Advanced machine learning concepts are typically covered in graduate level courses. However, as machine learning applications become more and more ubiquitous in our daily lives, it is important to expose students to machine learning concepts early at the undergraduate level. Signals and Systems I is a sophomore level course in the Electrical Engineering online bachelor degree curriculum. As the first course in signals and systems, it focuses on the basic concepts including signal transformation, linear time-invariant systems, Fourier series, Fourier transforms, Laplace and Z transforms. The course was taught using lecture videos and reading materials. MATLAB labs were also incorporated to introduce students to practical applications. Feedback from students showed their preference for more real world applications. This paper describes how a web-based simulation lab exercise was introduced to expose students to machine learning concepts. Specifically, students made the connection between machine learning and signals and systems concepts through a speech recognition application. In particular, students applied spectral analysis and identified voice features through pole/zero representation. To evaluate the effectiveness of the exercise, statistics from pre/post quizzes as well as student comments from a survey is analyzed. Abhinav Dixit, Andreas Spanias, Sunil Rao |
FIE | 3 |
| 2019 | Graph Filtering with Multiple Shift MatricesabstractWe propose a novel graph filtering method for semi-supervised classification that adopts multiple graph shift matrices to obtain more flexibility in dealing with misleading features. The resulting optimization problem is solved with a computationally efficient alternating minimization approach. In simulation experiments, we implement both conventional and our proposed graph filters as semi-supervised classifiers on real and synthetic datasets to demonstrate advantages of our algorithms in terms of classification performance. Cihan Tepedelenlioglu, Andreas Spanias |
ICASSP | 3 |
| 2019 | Designing an Effective Metric Learning Pipeline for Speaker DiarizationabstractState-of-the-art speaker diarization systems utilize knowledge from external data, in the form of a pre-trained distance metric, to effectively determine relative speaker identities to unseen data. However, much of recent focus has been on choosing the appropriate feature extractor, ranging from pre-trained i-vectors to representations learned via different sequence modeling architectures (e.g. 1D-CNNs, LSTMs, attention models), while adopting off-the-shelf metric learning solutions. In this paper, we argue that, regardless of the feature extractor, it is crucial to carefully design a metric learning pipeline, namely the loss function, the sampling strategy and the discriminative margin parameter, for building robust diarization systems. Furthermore, we propose to adopt a fine-grained validation process to obtain a comprehensive evaluation of the generalization power of metric learning pipelines. To this end, we measure diarization performance across different language speakers, and variations in the number of speakers in a recording. Using empirical studies, we provide interesting insights into the effectiveness of different design choices and make recommendations. Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Huan Song, Andreas Spanias |
ICASSP | 4 |
| 2019 | Distributed Bayesian Estimation with Low-rank Data: Application to Solar Array ProcessingabstractIn this paper, we present a distributed array processing algorithm to analyze the power output of solar photo-voltaic (PV) installations, leveraging the low-rank structure inherent in the data to estimate possible faults. Our multi-agent algorithm requires near-neighbor communications only and is also capable of jointly estimating the common low rank cloud profile and local shading of panels. To illustrate the workings of our algorithm, we perform experiments to detect shading faults in solar PV installations within a single ZIP code. Additionally, we also derive a Bayesian lower bound on the shading parameter's mean squared estimation error. The results are promising and show that we can successfully estimate the fraction of partial shading in solar installations that can usually go unnoticed. Raksha Ramakrishna, Anna Scaglione, Andreas Spanias, Cihan Tepedelenlioglu |
ICASSP | 3 |
| 2019 | Introducing Machine Learning in Undergraduate DSP ClassesabstractMachine Learning (ML) and Artificial Intelligence (AI) algorithms are enabling several modern smart products and devices. Furthermore, several initiatives such as smart cities and autonomous vehicles utilize AI and ML computational engines. The current and emerging applications and the growing industrial interest in AI necessitate introducing ML algorithms at the undergraduate level. In this paper, we describe a series of activities to introduce ML in undergraduate digital signal processing (DSP) classes. These activities include a computational comparative study of ML algorithms for spoken digit recognition using spectral representations of speech. We choose spectral representations and features for speech as those concepts associate with the core topics in DSP such as FFT and autoregressive spectra. Our primary objective is to introduce undergraduate DSP students to feature extraction and classification using appropriate signal analysis and ML tools. An online module on ML along with a computer exercise are developed and assigned as a semester project in the DSP class. The exercise is developed in Python and also on the online JDSP HTML5 environments. An assessment study of the modules and computer exercises are also part of this effort. Uday Shankar Shanthamallu, Sunil Rao, Abhinav Dixit, Vivek Sivaraman Narayanaswamy, Andreas Spanias |
ICASSP | 6 |
| 2019 | Spatially-Varying Sharpness Map Estimation Based on the Quotient of Spectral BandsabstractNatural images suffer from defocus blur due to the presence of objects at different depths from the camera. Automatic estimation of spatially-varying sharpness has several applications including depth estimation, image quality assessment, information retrieval, image restoration among others. In this paper, we propose a sharpness metric based on the quotient of high- to low-frequency bands of the log-spectrum of the image gradients. Using the proposed sharpness metric, we obtain a descriptive dense sharpness map. We also propose a simple yet effective method to segment out-of-focus regions using a global threshold which is defined using weak textured regions present in the input image. Results over two publicly available databases show that the proposed method provides competitive performance when compared with state-of-the-art methods. Juan Andrade, Pavan Turaga, Andreas Spanias |
ICIP | 3 |
| 2018 | Attend and Diagnose: Clinical Time Series Analysis Using Attention ModelsabstractWith widespread adoption of electronic health records, there is an increased emphasis for predictive models that can effectively deal with clinical time-series data. Powered by Recurrent Neural Network (RNN) architectures with Long Short-Term Memory (LSTM) units, deep neural networks have achieved state-of-the-art results in several clinical prediction tasks. Despite the success of RNN, its sequential nature prohibits parallelized computing, thus making it inefficient particularly when processing long sequences. Recently, architectures which are based solely on attention mechanisms have shown remarkable success in transduction tasks in NLP, while being computationally superior. In this paper, for the first time, we utilize attention models for clinical time-series modeling, thereby dispensing recurrence entirely. We develop the SAnD (Simply Attend and Diagnose) architecture, which employs a masked, self-attention mechanism, and uses positional encoding and dense interpolation strategies for incorporating temporal order. Furthermore, we develop a multi-task variant of SAnD to jointly infer models with multiple diagnosis tasks. Using the recent MIMIC-III benchmark datasets, we demonstrate that the proposed approach achieves state-of-the-art performance in all tasks, outperforming LSTM models and classical baselines with hand-engineered features. Huan Song, Deepta Rajan, Jayaraman J. Thiagarajan, Andreas Spanias |
AAAI | 4 |
| 2018 | Online Machine Learning Experiments in HTML5abstractThis work in progress paper describes software that enables online machine learning experiments in an undergraduate DSP course. This software operates in HTML5 and embeds several digital signal processing functions. The software can process natural signals such as speech and can extract various features, for machine learning applications. For example in the case of speech processing, LPC coefficients and formant frequencies can be computed. In this paper, we present speech processing, feature extraction and clustering of features using the K-means machine learning algorithm. The primary objective is to provide a machine learning experience to undergraduate students. The functions and simulations described provide a user-friendly visualization of phoneme recognition tasks. These tasks make use of the Levinson-Durbin linear prediction and the K-means machine learning algorithms. The exercise was assigned as a class project in our undergraduate DSP class. The description of the exercise along with assessment results is described. Abhinav Dixit, Uday Shankar Shanthamallu, Andreas Spanias, Visar Berisha, Mahesh K. Banavar |
FIE | 3 |
| 2018 | A Stem Reu Site on the Integrated Design of Sensor Devices and Signal Processing AlgorithmsabstractArizona State University (ASU) established an NSF Research Experiences for Undergraduates (REU) site to embed students in research projects related to integrated sensor and signal processing systems. The program includes both sensor hardware and algorithm/software design for a variety of applications including health monitoring. The site was funded in February 2017 and the Co-PIs recruited nine students from different universities and community colleges to spend the summer of 2017 in research laboratories at ASU. The program included structured training with modules in sensor design, signal processing, and machine learning. Cross-cutting training included research ethics, IEEE manuscript development, and building presentation skills. Nine undergraduate research projects were launched and the program went through an assessment by an independent evaluator. This paper describes the REU activities, modules, training, projects, and their assessment. Andreas Spanias, Jennifer Blain Christen |
ICASSP | 1 |
| 2018 | Fast Non-Linear Methods for Dynamic Texture PredictionabstractThis paper aims to develop a fast dynamic-texture prediction method, using tools from non-linear dynamical modeling, and fast approaches for approximate regression. We consider dynamic textures to be described by patch-level non-linear processes, thus requiring tools such as delay-embedding to uncover a phase-space where dynamical evolution can be more easily modeled. After mapping the observed time-series from a dynamic texture video to its recovered phase-space, a time-efficient approximate prediction method is presented which utilizes locality-sensitive hashing approaches to predict possible phase-space vectors, given the current phase-space vector. Our experiments show the favorable performance of the proposed approach, both in terms of prediction fidelity, and computational time. The proposed algorithm is applied to shading prediction in utility scale solar arrays. Sameeksha Katoch, Pavan Turaga, Andreas Spanias, Cihan Tepedelenlioglu |
ICIP | 3 |
| 2018 | Triplet Network with Attention for Speaker DiarizationabstractIn automatic speech processing systems, speaker diarization is a crucial front-end component to separate segments from different speakers.Inspired by the recent success of deep neural networks (DNNs) in semantic inferencing, triplet loss-based architectures have been successfully used for this problem.However, existing work utilizes conventional i-vectors as the input representation and builds simple fully connected networks for metric learning, thus not fully leveraging the modeling power of DNN architectures.This paper investigates the importance of learning effective representations from the sequences directly in metric learning pipelines for speaker diarization.More specifically, we propose to employ attention models to learn embeddings and the metric jointly in an end-to-end fashion.Experiments are conducted on the CALLHOME conversational speech corpus.The diarization results demonstrate that, besides providing a unified model, the proposed approach achieves improved performance when compared against existing approaches. Huan Song, Megan M. Willi, Jayaraman J. Thiagarajan, Visar Berisha, Andreas Spanias |
INTERSPEECH | 5 |
| 2018 | Optimizing Kernel Machines Using Deep LearningabstractBuilding highly nonlinear and nonparametric models is central to several state-of-the-art machine learning systems. Kernel methods form an important class of techniques that induce a reproducing kernel Hilbert space (RKHS) for inferring non-linear models through the construction of similarity functions from data. These methods are particularly preferred in cases where the training data sizes are limited and when prior knowledge of the data similarities is available. Despite their usefulness, they are limited by the computational complexity and their inability to support end-to-end learning with a task-specific objective. On the other hand, deep neural networks have become the de facto solution for end-to-end inference in several learning paradigms. In this paper, we explore the idea of using deep architectures to perform kernel machine optimization, for both computational efficiency and end-to-end inferencing. To this end, we develop the deep kernel machine optimization framework, that creates an ensemble of dense embeddings using Nyström kernel approximations and utilizes deep learning to generate task-specific representations through the fusion of the embeddings. Intuitively, the filters of the network are trained to fuse information from an ensemble of linear subspaces in the RKHS. Furthermore, we introduce the kernel dropout regularization to enable improved training convergence. Finally, we extend this framework to the multiple kernel case, by coupling a global fusion layer with pretrained deep kernel machines for each of the constituent kernels. Using case studies with limited training data, and lack of explicit feature sources, we demonstrate the effectiveness of our framework over conventional model inferencing techniques. Huan Song, Jayaraman J. Thiagarajan, Prasanna Sattigeri, Andreas Spanias |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Signal processing and machine learning concepts using the reflections echolocation appabstractThis paper describes the use of a space usage determination algorithm for teaching signal processing and machine learning concepts to undergraduate electrical engineering and computer science students. An Android device transmits a high-frequency signal in an unknown space. The device determines the reflective properties of this unknown space by analyzing the received signal. Based on the features extracted from this signal, the app measures distances and determines how the space can be utilized for various application such as libraries, conference rooms, or laboratories. The application and related algorithms use concepts such cross-correlation, feature extraction, learning/training algorithms, and discrimination/decision making. These concepts are typically covered in undergraduate classes such as Digital Signal Processing, Control Systems, and Probability and Statistics; and graduate-level classes such as Pattern Recognition and Detection and Estimation Theory. The app is used to create compelling demonstrations and immersive exercises to teach basic concepts related to signal processing and machine learning. Undergraduate student hands-on workshops and outreach activities are planned to evaluate the effectiveness of this approach. Assessment results will be presented at the conference. Mahesh K. Banavar, Houchao Gan, Benjamin Robistow, Andreas Spanias |
FIE | 4 |
| 2017 | Development of signal processing online labs using HTML5 and mobile platformsabstractSeveral web-based signal processing simulation packages for education have been developed in a Java environment. Although this environment has provided convenience and accessibility using standard browser technology, it has recently become vulnerable to cyber-attacks and is no longer compatible with secure browsers. In this paper, we describe our efforts to transform our award-winning J-DSP online laboratory by rebuilding it on an HTML5 framework. Along with a new simulation environment, we have redesigned the interface to enable several new functionalities and an entirely new educational experience. These new features include functions that enable real-time interfaces with sensor boards and mobile phones. The Web 4.0 HTML5 technology departs from older Java interfaces and provides an interactive graphical user interface (GUI) enabling seamless connectivity and both software and hardware experiences for students in DSP classes. Abhinav Dixit, Sameeksha Katoch, Photini Spanias, Mahesh K. Banavar, Huan Song, Andreas Spanias |
FIE | 6 |
| 2017 | Reflections: An eModule for echolocation educationabstractAn Android-based eModule app has been designed and developed for science, technology, engineering, and mathematics (STEM) education. The eModule consists of: (1) an Android demonstration of echolocation; (2) a set of notes describing the functionality of the app, the basics of echolocation, and its application to advanced signal processing systems such as RADAR, LIDAR, and SONAR; (3) quizzes to test the concepts introduced by the demonstration and the notes; and (4) companion videos. The eModule is, therefore, a holistic teaching and learning app that can be used across various grade levels including K-12, undergraduate signals and systems, and graduate DSP education.The app, “Reflections”, provides students a means to determine distances to objects while allowing them the ability to manipulate signal envelopes, signal shapes, signal types, and frequency constraints. The intuitive graphical user interface, combined with notes, videos and quizzes, creates a rich educational environment to help educate users with the fundamental concepts of signals, systems, and digital signal processing. In addition to its role in STEM education, the app has potential use in low-visibility environments and for spatial acoustic analysis. Preliminary assessments strongly support the effectiveness of the eModule as an education tool signals and systems and DSP classes. Benjamin Robistow, Robert Newman, Thomas H. DePue, Mahesh K. Banavar, Dana M. Barry, Paul Curtis, Andreas Spanias |
ICASSP | 7 |
| 2017 | A deep learning approach to multiple kernel fusionabstractKernel fusion is a popular and effective approach for combining multiple features that characterize different aspects of data. Traditional approaches for Multiple Kernel Learning (MKL) attempt to learn the parameters for combining the kernels through sophisticated optimization procedures. In this paper, we propose an alternative approach that creates dense embeddings for data using the kernel similarities and adopts a deep neural network architecture for fusing the embeddings. In order to improve the effectiveness of this network, we introduce the kernel dropout regularization strategy coupled with the use of an expanded set of composition kernels. Experiment results on a real-world activity recognition dataset show that the proposed architecture is effective in fusing kernels and achieves state-of-the-art performance. Huan Song, Jayaraman J. Thiagarajan, Prasanna Sattigeri, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
ICASSP | 5 |
| 2016 | An Android app for spatial acoustic analysis as a learning toolabstractAn Android app has been developed to assist in the education of individuals in a science, technology, engineering, and mathematics (STEM) course of study. The Android Reflection Application provides students a means to determine distances to objects while allowing them the ability to manipulate signal envelopes, signal shapes, signal types, and frequency constraints. The convenient and intuitive graphical user interface immerses the user into a richly educational environment allowing for the solidification of fundamental concepts regarding digital signal processing (DSP). In addition to the educational benefits, this application is also being applied to spatial acoustic analysis and assistance in low-visibility. This feature will allow users to determine the best use for a given space whether it is a quiet study room or a room better suited for conference meetings. The effectiveness of this application has not yet been formally tested but suggests a positive result. Thomas H. DePue, Benjamin Robistow, Robert Newman, Kevin Mack, Mahesh K. Banavar, Dana M. Barry, Paul Curtis, Andreas Spanias, Whitni Watkins |
FIE | 9 |
| 2016 | Development of course modules for multidisciplinary STEM educationabstractTraditional STEM education models in electrical engineering and computer science rely on structured classes, laboratories, and textbooks to transfer key concepts. Even though this process meets most of the ABET objectives, it does not respond well to current workforce needs that require widely accessible programs that will provide a large pool of graduates with STEM backgrounds, analytical and programming skills, critical thinking, and leadership abilities. In this work in progress paper, we describe our efforts to motivate students to pursue studies in STEM areas. We accomplish this by creating and disseminating modules that demonstrate how math and engineering theory enable modern applications such as those embedded in wireless devices. Andreas Spanias, Mahesh K. Banavar, Henry Braun, Photini Spanias, Yongpeng Zhang |
FIE | 1 |
| 2016 | Distributed Estimation of the Degree Distribution in Wireless Sensor NetworksabstractA distributed consensus algorithm for estimating the degree distribution of a graph is proposed. The proposed algorithm is based on average consensus and in-network empirical mass function estimation. It is fully distributed in the sense that each node in the network only needs to know its own degree, and nodes do not need to be labeled. The algorithm works for any connected graph structure in the presence of communication noise. The performance of the algorithm is analyzed. A discussion on how the properties of the graph degree distribution can be exploited for post-processing after consensus is reached is given. Simulation results corroborating the theory are also provided. Sai Zhang 0002, Jongmin Lee 0003, Cihan Tepedelenlioglu, Andreas Spanias |
GLOBECOM | 4 |
| 2016 | Consensus inference on mobile phone sensors for activity recognitionabstractThe pervasive use of wearable sensors in activity and health monitoring presents a huge potential for building novel data analysis and prediction frameworks. In particular, approaches that can harness data from a diverse set of low-cost sensors for recognition are needed. Many of the existing approaches rely heavily on elaborate feature engineering to build robust recognition systems, and their performance is often limited by the inaccuracies in the data. In this paper, we develop a novel two-stage recognition system that enables a systematic fusion of complementary information from multiple sensors in a linear graph embedding setting, while employing an ensemble classifier phase that leverages the discriminative power of different feature extraction strategies. Experimental results on a challenging dataset show that our framework greatly improves the recognition performance when compared to using any single sensor. Huan Song, Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias, Pavan Turaga |
ICASSP | 4 |
| 2016 | Empirically-estimable multi-class classification boundsabstractIn this paper, we extend previously developed non-parametric bounds on the Bayes risk in binary classification problems to multi-class problems. In comparison with the well-known Bhattacharyya bound which is typically calculated by employing parametric assumptions, the bounds proposed in this paper are directly estimable from data, provably tighter, and more robust to different types of data. We verify the tightness and validity of this bound using an illustrative synthetic example, and further demonstrate its value by incorporating it into a feature selection algorithm which we apply to the real-world problem of distinguishing between different neuro-motor disorders based on sentence-level speech data. Alan Wisler, Visar Berisha, Dennis Wei, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
ICASSP | 5 |
| 2016 | Auto-context modeling using multiple Kernel learningabstractIn complex visual recognition systems, feature fusion has become crucial to discriminate between a large number of classes. In particular, fusing high-level context information with image appearance models can be effective in object/scene recognition. To this end, we develop an auto-context modeling approach under the RKHS (Reproducing Kernel Hilbert Space) setting, wherein a series of supervised learners are used to approximate the context model. By posing the problem of fusing the context and appearance models using multiple kernel learning, we develop a computationally tractable solution to this challenging problem. Furthermore, we propose to use the marginal probabilities from a kernel SVM classifier to construct the auto-context kernel. In addition to providing better regularization to the learning problem, our approach leads to improved recognition performance in comparison to using only the image features. Huan Song, Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
ICIP | 4 |
| 2015 | A new signal processing course for digital cultureabstractSignal processing algorithms, software, and hardware are being used in several fields including non-engineering areas such as arts and media. Students in these fields and particularly in the new Digital Culture major at Arizona State University (ASU) use signal processing tools in several of their projects and artistic endeavors. Yet the blind use of these DSP tools in other disciplines, without understanding their properties has been a long-standing problem. In fact, the broader issue is the disconnect between engineers that develop tools and artists that use them to design the next generation digital art applications. In that context, ASU has formed the Arts Media and Engineering (AME) School and more recently, the multidisciplinary undergraduate Digital Culture degree granting program. In order to provide formal training in signal processing to students that are non-Electrical Engineering majors, we piloted a new course titled Signal Processing for Digital Culture. This course, which is being offered online, teaches non-majors some of the basics of signal processing and covers several applications. The only prerequisite to the course is general sophomore calculus. This new online course contains several topics and is focused on an approach that teaches concepts by connecting theory to compelling applications. Future plans include introducing this course at Clarkson University as a Knowledge Area course open to students from all majors. Andreas Spanias, Paul Curtis, Photini Spanias, Mahesh K. Banavar |
FIE | 1 |
| 2015 | Audio modeling and loudness estimation with IJDSP mobile simulationsabstractAudio signal modeling and simulation is important in several coding, noise removal, and recognition applications. This paper focuses on implementing models for loudness estimation and their use in estimating parameters on iOS mobile devices (iPhones and iPads). We briefly address estimating excitation patterns and loudness through auditory models. These loudness estimation and other algorithms were implemented in the award winning educational iOS app iJDSP for performing DSP simulations on mobile devices. The modules were introduced to graduate students in the general signal processing area, to evaluate their effectiveness as teaching tools. The evaluation process involved giving the students a pre-quiz, guiding them through hands-on activities on the iOS app, and finally, a post-quiz. Assessments results were positive with noticeable improvement of student understanding of topics such as spectrograms and linear predictive coding. Girish Kalyanasundaram, Mahesh K. Banavar, Andreas Spanias |
ICASSP | 3 |
| 2015 | Removing data with noisy responses in regression analysisabstractIn regression analysis, outliers in the data can induce a bias in the learned function, resulting in larger errors. In this paper we derive an empirically estimable bound on the regression error based on a Euclidean minimum spanning tree generated from the data. Using this bound as motivation, we propose an iterative approach to remove data with noisy responses from the training set. We evaluate the performance of the algorithm on experiments with real-world pathological speech (speech from individuals with neurogenic disorders). Comparative results show that removing noisy examples during training using the proposed approach yields better predictive performance on out-of- sample data. Alan Wisler, Visar Berisha, Karthikeyan Natesan Ramamurthy, Andreas Spanias, Julie M. Liss |
ICASSP | 4 |
| 2015 | Nonlinear diffusion adaptation with bounded transmission over distributed networksabstractThis paper introduces diffusion adaptation strategies over distributed networks with nonlinear transmissions, motivated by the necessity for bounded transmit power. Local information is exchanged in real-time with neighboring nodes in order to estimate a common parameter vector via constrained nonlinear transmissions, using an adaptive learning algorithm. We propose nonlinear diffusion strategies for such an adaptive estimation. We will study convergence properties of the proposed algorithm in the mean and the mean-square sense. Simulations support the performance analysis and show that the proposed algorithm performs close to the linear case with the added advantage of power savings. Jongmin Lee 0003, Cihan Tepedelenlioglu, Mahesh K. Banavar, Andreas Spanias |
ICC | 4 |
| 2015 | Learning Stable Multilevel Dictionaries for Sparse RepresentationsabstractSparse representations using learned dictionaries are being increasingly used with success in several data processing and machine learning applications. The increasing need for learning sparse models in large-scale applications motivates the development of efficient, robust, and provably good dictionary learning algorithms. Algorithmic stability and generalizability are desirable characteristics for dictionary learning algorithms that aim to build global dictionaries, which can efficiently model any test data similar to the training samples. In this paper, we propose an algorithm to learn dictionaries for sparse representations from large scale data, and prove that the proposed learning algorithm is stable and generalizable asymptotically. The algorithm employs a 1-D subspace clustering procedure, the K-hyperline clustering, to learn a hierarchical dictionary with multiple levels. We also propose an information-theoretic scheme to estimate the number of atoms needed in each level of learning and develop an ensemble approach to learn robust dictionaries. Using the proposed dictionaries, the sparse code for novel test data can be computed using a low-complexity pursuit procedure. We demonstrate the stability and generalization characteristics of the proposed algorithm using simulations. We also evaluate the utility of the multilevel dictionaries in compressed recovery and subspace learning applications. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Embedding Android signal processing apps in a high school math class - An RET projectabstractThe objective of this project is to develop and design mobile content for introducing engineering technology to high school students. More specifically, we intend to work on a sequence of modules that will establish connections between high school mathematics and physics to modern technologies associated with smart phones, iPods and other high-tech products. The participants of the project will use the previously developed AJDSP (for Android devices) and iJDSP (for iPhones and iPads) apps to facilitate this process. Additionally, modules have been developed that have been embedded in math classes. Anticipated benefits of the project include creating positive attitudes towards STEM areas that will help recruit high school students and minorities in engineering, math and science fields. After an initial pilot study and assessments at CDS High School, these activities will be disseminated to other high schools. In order to obtain feedback from high school students and teachers, we will hold workshops and collect assessment results. These results will also provide us assessments about the effectiveness of the project, and allow us to make modifications to the project as necessary. The project is part of TUES Phase 3 and I/UCRC RET activities. Mahesh K. Banavar, Deepta Rajan, Andrew Strom, Photini Spanias, Xue Zhang 0002, Henry Braun, Andreas Spanias |
FIE | 7 |
| 2014 | Signals and systems demonstrations for undergraduates using Android-based localizationabstractThis project aims to contribute to education research using mobile apps to demonstrate how signals and systems concepts are used in sensor network localization. An educational demonstration of sensor localization on mobile devices is described with the mobile devices acting as nodes of a sensor network. In our approach, we are developing an app that uses a modified version of time-difference of arrival (TDOA) using audio signals and commercial Android devices. At a high level, the app can be used to illustrate how triangulation can be used to localize devices, which is a concept that is used in GPS. Students are exposed to signal processing concepts such as correlation and the fast Fourier transform (FFT), and their utility in sensor localization. The use of FFT-based convolution for computing correlations is also demonstrated. Since the app has been developed for Android devices, it can be made widely available and has the added benefit of appealing to students that are eager to use educational apps on their mobile phones. Assessments will be performed to evaluate the effectiveness of the app in education. The work engages NSF REU and REV students who are involved in developing and testing the app. Paul Curtis, Mahesh K. Banavar, Xue Zhang 0002, Andreas Spanias, Vitor Weber |
FIE | 4 |
| 2014 | Modeling pathological speech perception from data with similarity labelsabstractThe current state of the art in judging pathological speech intelligibility is subjective assessment performed by trained speech pathologists (SLP). These tests, however, are inconsistent, costly and, oftentimes suffer from poor intra- and inter-judge reliability. As such, consistent, reliable, and perceptually-relevant objective evaluations of pathological speech are critical. Here, we propose a data-driven approach to this problem. We propose new cost functions for examining data from a series of experiments, whereby we ask certified SLPs to rate pathological speech along the perceptual dimensions that contribute to decreased intelligibility. We consider qualitative feedback from SLPs in the form of comparisons similar to statements "Is Speaker A's rhythm more similar to Speaker B or Speaker C?" Data of this form is common in behavioral research, but is different from the traditional data structures expected in supervised (data matrix + class labels) or unsupervised (data matrix) machine learning. The proposed method identifies relevant acoustic features that correlate with the ordinal data collected during the experiment. Using these features, we show that we are able to develop objective measures of the speech signal degradation that correlate well with SLP responses. Visar Berisha, Julie M. Liss, Steven Sandoval, Rene Utianski, Andreas Spanias |
ICASSP | 5 |
| 2014 | Direct tracking from compressive imagers: A proof of conceptabstractThe compressive sensing paradigm holds promise for more cost-effective imaging outside of the visible range, particularly in infrared wavelengths. However, the process of reconstructing compressively sensed images remains computationally expensive. The proof-of-concept tracker described here uses a particle filter with a likelihood update based on a “smashed filter” which estimates correlation directly, avoiding the reconstruction step. This approach leads to increased noise in correlation estimates, but by implementing the track-before-detect concept in the particle filter, tracker convergence may still be achieved with reasonable sensing rates. The tracker has been successfully tested on sequences of moving cars in the PETS2000 dataset. Henry Braun, Pavan Turaga, Andreas Spanias |
ICASSP | 3 |
| 2014 | Automatic image annotation using inverse maps from semantic embeddingsabstractHuman annotation in large scale image databases is time-consuming and error-prone. Since it is very hard to mine image databases using just visual features or textual descriptors, it is common to transform the image features into a semantically meaningful space. In this paper, we propose to perform image annotation in a semantic space inferred based on sparse representations. By constructing a semantic embedding for the visual features, that is constrained to be close to the tag embedding, we show that a robust inverse map can be used to predict the tags. Experiments using standard datasets show the effectiveness of the proposed approach in automatic image annotation when compared to existing methods. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri, Peer-Timo Bremer, Andreas Spanias |
ICIP | 5 |
| 2014 | A multi-modal approach to emotion recognition using undirected topic modelsabstractA multi-modal framework for emotion recognition using bag-of-words features and undirected, replicated softmax topic models is proposed here. Topic models ignore the temporal information between features, allowing them to capture the complex structure without a brute-force collection of statistics. Experiments are performed over face, speech and language features extracted from the USC IEMOCAP database. Performance on facial features yields an unweighted average recall of 60.71%, a relative improvement of 8.89% over state-of-the-art approaches. A comparable performance is achieved when considering only speech (57.39%) or a fusion of speech and face information (66.05%). Individually, each source is shown to be strong at recognizing either sadness (speech) or happiness (face) or neutral (language) emotions, while, a multi-modal fusion retains these properties and improves the accuracy to 68.92%. Implementation time for each source and their combination is provided. Results show that a turn of 1 second duration can be classified in approximately 666.65ms, thus making this method highly amenable for real-time implementation. Mohit Shah, Chaitali Chakrabarti, Andreas Spanias |
ISCAS | 3 |
| 2014 | Domain invariant speech features using a new divergence measureabstractExisting speech classification algorithms often perform well when evaluated on training and test data drawn from the same distribution. In practice, however, these distributions are not always the same. In these circumstances, the performance of trained models will likely decrease. In this paper, we discuss an underutilized divergence measure and derive an estimable upper bound on the test error rate that depends on the error rate on the training data and the distance between training and test distributions. Using this bound as motivation, we develop a feature learning algorithm that aims to identify invariant speech features that generalize well to data similar to, but different from, the training set. Comparative results confirm the efficacy of the algorithm on a set of cross-domain speech classification tasks. Alan Wisler, Visar Berisha, Julie M. Liss, Andreas Spanias |
SLT | 4 |
| 2014 | Multiple Kernel Sparse Representations for Supervised and Unsupervised LearningabstractIn complex visual recognition tasks, it is typical to adopt multiple descriptors, which describe different aspects of the images, for obtaining an improved recognition performance. Descriptors that have diverse forms can be fused into a unified feature space in a principled manner using kernel methods. Sparse models that generalize well to the test data can be learned in the unified kernel space, and appropriate constraints can be incorporated for application in supervised and unsupervised learning. In this paper, we propose to perform sparse coding and dictionary learning in the multiple kernel space, where the weights of the ensemble kernel are tuned based on graph-embedding principles such that class discrimination is maximized. In our proposed algorithm, dictionaries are inferred using multiple levels of 1D subspace clustering in the kernel space, and the sparse codes are obtained using a simple levelwise pursuit scheme. Empirical results for object recognition and image clustering show that our algorithm outperforms existing sparse coding based approaches, and compares favorably to other state-of-the-art methods. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
IEEE Trans. Image Process. | 3 |
| 2013 | Health monitoring laboratories by interfacing physiological sensors to mobile android devicesabstractThe recent sensing capabilities of mobile devices along with their interactivity and popularity in the student community can be used to create a unique learning environment in engineering education. Android Java-DSP (AJDSP) is a mobile educational application that interfaces with sensors and enables simulation and visualization of signal processing concepts. In this paper, we present the work done towards building non-invasive physiological signal monitoring tools in AJDSP through hardware interfaces to both external sensors and on-board device sensors. Examples of laboratory exercises that can be introduced in classes are presented. The proposed software tools can be used to provide intuitive understanding in wireless sensing and feature extraction to demonstrate the application of DSP to health monitoring systems. The effectiveness of the software modules in enhancing student understanding is demonstrated with the help of preliminary assessments. Deepta Rajan, Andreas Spanias, Suhas Ranganath, Mahesh K. Banavar, Photini Spanias |
FIE | 2 |
| 2013 | Java tools for teaching OFDM principles in undergraduate coursesabstractIn this paper, we describe a new set of software functions and associated exercises that can be used for teaching orthogonal frequency division multiplexing (OFDM) concepts in undergraduate DSP and communications courses. These tools can be used to simulate, visualize, and analyze the performance and behavior of OFDM systems by considering different input signals and communication channels. OFDM is a compelling paradigm because of its utility in WiFi and LTE. It is also a good demonstration of the use of the FFT in a communication system. We have developed the proposed set of functions as a part of the Java-DSP (J-DSP) visual programming environment. The functions can be used in undergraduate DSP and communications courses, in order to demonstrate to students, the application of DSP concepts in a communication system, as well as concepts such as FIR filter design, properties of the DFT matrix, random signals, and circular effects. Sai Zhang 0002, Mahesh K. Banavar, Andreas Spanias, Cihan Tepedelenlioglu, Xue Zhang 0002 |
FIE | 3 |
| 2013 | A heterogeneous dictionary model for representation and recognition of human actionsabstractIn this paper, we consider low-dimensional and sparse representation models for human actions, that are consistent with how actions evolve in high-dimensional feature spaces. We first show that human actions can be well approximated by piecewise linear structures in the feature space. Based on this, we propose a new dictionary model that considers each atom in the dictionary to be an affine subspace defined by a point and a corresponding line. When compared to centered clustering approaches such as K-means, we show that the proposed dictionary is a better generative model for human actions. Furthermore, we demonstrate the utility of this model in efficient representation and recognition of human activities that are not available in the training set. Rushil Anirudh, Karthikeyan Natesan Ramamurthy, Jayaraman J. Thiagarajan, Pavan Turaga, Andreas Spanias |
ICASSP | 5 |
| 2013 | Selecting disorder-specific features for speech pathology fingerprintingabstractThe general aim of this work is to learn a unique statistical signature for the state of a particular speech pathology. We pose this as a speaker identification problem for dysarthric individuals. To that end, we propose a novel algorithm for feature selection that aims to minimize the effects of speaker-specific features (e.g., fundamental frequency) and maximize the effects of pathology-specific features (e.g., vocal tract distortions and speech rhythm). We derive a cost function for optimizing feature selection that simultaneously trades off between these two competing criteria. Furthermore, we develop an efficient algorithm that optimizes this cost function and test the algorithm on a set of 34 dysarthric and 13 healthy speakers. Results show that the proposed method yields a set of features related to the speech disorder and not an individual's speaking style. When compared to other feature-selection algorithms, the proposed approach results in an improvement in a disorder fingerprinting task by selecting features that are specific to the disorder. Visar Berisha, Steven Sandoval, Rene Utianski, Julie M. Liss, Andreas Spanias |
ICASSP | 5 |
| 2013 | Optical flow for compressive sensing video reconstructionabstractAlthough considerable effort has been devoted to the problem of reconstructing compressively sensed video, no existing algorithm achieves results comparable to commonly available video compression methods such as H.264. One possible avenue for improving compressively sensed video reconstruction is the use of optical flow information. Current efforts reported in the literature have not fully utilized optical flow information, instead focusing on limited cases such as stationary backgrounds with sparse foreground motion. In this paper, a reconstruction method is presented which fully utilizes optical flow information to increase the quality of reconstruction. The special cases of known image motion and constant global image motion are presented, and the performance of the algorithm on existing datasets is evaluated. Henry Braun, Pavan Turaga, Cihan Tepedelenlioglu, Andreas Spanias |
ICASSP | 4 |
| 2013 | Sinusoidal component selection based on partial loudness criteriaabstractSinusoidal models are widely used in parametric speech and audio coding schemes. A common requirement in these applications is to select only a subset of components that provide the greatest perceptual benefit particularly at low bitrates. Usually, perceptual sinusoidal component selection algorithms make use of greedy algorithms that are computationally expensive. In this paper, we present a new algorithm that selects sinusoidal components based on the partial loudness model proposed by Moore & Glasberg. We compare the performance of the proposed algorithm in terms of perceptual benefit and computational complexity to other existing sinusoidal selection algorithms. Harish Krishnamoorthi, Andreas Spanias |
ICASSP | 2 |
| 2013 | Boosted dictionaries for image restoration based on sparse representationsabstractSparse representations using learned dictionaries have been successful in several image processing applications. However, using a single dictionary model in inverse problems may lead to instability in estimation. In this paper, we propose to perform image restoration using an ensemble of weak dictionaries that incorporate prior knowledge about the form of linear corruption. The dictionary learned in each round of the training procedure is optimized for the training examples having high reconstruction error in the previous round. The weak dictionaries are either obtained using a weighted K-Means or an example-selection approach. The final restored data is computed as a convex combination of data restored in individual rounds. Results with compressed recovery of standard images show that the proposed dictionaries result in a better performance compared to using a single dictionary obtained with a traditional alternating minimization approach. Karthikeyan Natesan Ramamurthy, Jayaraman J. Thiagarajan, Andreas Spanias, Prasanna Sattigeri |
ICASSP | 3 |
| 2013 | A speech emotion recognition framework based on latent Dirichlet allocation: Algorithm and FPGA implementationabstractIn this paper, we present a speech-based emotion recognition framework based on a latent Dirichlet allocation model. This method assumes that incoming speech frames are conditionally independent and exchangeable. While this leads to a loss of temporal structure, it is able to capture significant statistical information between frames. In contrast, a hidden Markov model-based approach captures the temporal structure in speech. Using the German emotional speech database EMO-DB for evaluation, we achieve an average classification accuracy of 80.7% compared to 73% for hidden Markov models. This improvement is achieved at the cost of a slight increase in computational complexity. We map the proposed algorithm onto an FPGA platform and show that emotions in a speech utterance of duration 1.5s can be identified in 1.8ms, while utilizing 70% of the resources. This further demonstrates the suitability of our approach for real-time applications on hand-held devices. Mohit Shah, Lifeng Miao, Chaitali Chakrabarti, Andreas Spanias |
ICASSP | 4 |
| 2013 | CRLB for the localization error in the presence of fadingabstractLocalization accuracy is crucial in sensor networks. A wireless sensor network (WSN) with M anchors and one node is considered in this paper. The estimation is based on time of arrival (TOA) in the presence of fading channels. The Cramer-Rao lower bound (CRLB) for localization error in the presence of fading is derived under different scenarios. Firstly, fading coefficients are considered as unknown random parameters with a prior distribution. The ML estimator for this case is also derived. If the distribution of fading is unknown to the estimator then the modified CRLB (MCRLB) is applied and shown to be equal to the CRLB in the absence of fading. This is used to conclude that fading always deteriorates the CRLB in localization. It is shown that there is a loss of about 5dB in CRLB due to Rayleigh fading. Xue Zhang 0002, Cihan Tepedelenlioglu, Mahesh K. Banavar, Andreas Spanias |
ICASSP | 4 |
| 2012 | Automated tumor segmentation using kernel sparse representationsabstractIn this paper, we describe a pixel based approach for automated segmentation of tumor components from MR images. Sparse coding with data-adapted dictionaries has been successfully employed in several image recovery and vision problems. Since it is trivial to obtain sparse codes for pixel values, we propose to consider their non-linear similarities to perform kernel sparse coding in a high dimensional feature space. We develop the kernel K-lines clustering procedure for inferring kernel dictionaries and use the kernel sparse codes to determine if a pixel belongs to a tumorous region. By incorporating spatial locality information of the pixels, contiguous tumor regions can be efficiently identified. A low complexity segmentation approach, which allows the user to initialize the tumor region, is also presented. Results show that both of the proposed approaches lead to accurate tumor identification with a low false positive rate, when compared to manual segmentation by an expert. Jayaraman J. Thiagarajan, Deepta Rajan, Karthikeyan Natesan Ramamurthy, David H. Frakes, Andreas Spanias |
BIBE | 5 |
| 2012 | Workshop: Interactive education tools for earth systems and sustainability applicationsabstractEarth system signals include indicators of climate change. In this workshop, the participants will use the Java-DSP/Earth Systems Edition in order to analyze and understand the components and drivers of climate change in the twentieth century. The session will be interactive and will be useful to researchers, practitioners and instructors with interests in Earth systems signal analysis. People with interests in general STEM related areas will also find this workshop useful as an important interdisciplinary application of signal processing. Linda Hinnov, Andreas Spanias, Karthikeyan Natesan Ramamurthy, Girish Kalyanasundaram |
FIE | 2 |
| 2012 | Work in progress: Performing signal analysis laboratories using Android devicesabstractIn this paper, we present a graphical-programming application to support signal processing education on the Android operating system. This application features a simulation environment and a palette of DSP functions, which will allow students to perform laboratories using Android smartphones and tablets. In order to demonstrate the application of the software in a classroom setting, a number of laboratories which incorporate the proposed functionalities have been developed. A set of assessments designed to evaluate the effectiveness of the software is also presented. Suhas Ranganath, Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Mahesh K. Banavar, Andreas Spanias |
FIE | 6 |
| 2012 | Signal processing for fault detection in photovoltaic arraysabstractPhotovoltaics (PV) is an important and rapidly growing area of research. With the advent of power system monitoring and communication technology collectively known as the “smart grid,” an opportunity exists to apply signal processing techniques to monitoring and control of PV arrays. In this paper a monitoring system which provides real-time measurements of each PV module's voltage and current is considered. A fault detection algorithm formulated as a clustering problem and addressed using the robust minimum covariance determinant (MCD) estimator is described; its performance on simulated instances of arc and ground faults is evaluated. The algorithm is found to perform well on many types of faults commonly occurring in PV arrays. Henry Braun, Santoshi T. Buddha, Venkatachalam Krishnan, Andreas Spanias, Cihan Tepedelenlioglu, Ted Yeider, Toru Takehara |
ICASSP | 4 |
| 2012 | Interactive DSP laboratories on mobile phones and tabletsabstractThe use of mobile devices and tablets in engineering education has been gaining lot of interest, due to its interactive capabilities and its ability to stimulate student interest. On the other hand, this technology can also enable instructors to broaden the scope of their curriculum and increase student participation. In this paper, we describe an interactive application to perform signal processing simulations on iOS devices such as the iPhone and the iPad. Furthermore, we describe two laboratory exercises to introduce continuous/discrete convolution and filter design. The exercises and the proposed application will be evaluated by students of an undergraduate DSP course at Arizona State University during Fall 2011. Finally, we describe the planned assessment methodology which will enable us to provide prescriptive recommendations for using i-JDSP in DSP courses. Jinru Liu, Jayaraman J. Thiagarajan, Xue Zhang 0002, Suhas Ranganath, Mahesh K. Banavar, Andreas Spanias |
ICASSP | 7 |
| 2012 | Supervised local sparse coding of sub-image features for image retrievalabstractThe success of sparse representations in image modeling and recovery has motivated its use in computer vision applications. Image retrieval and classification tasks require extracting features that discriminate different image classes. State-of-the-art object recognition methods based on sparse coding use spatial pyramid features obtained from dense descriptors. In this paper, we develop a feature extraction method that uses multiple global/local features extracted from large overlapping regions of an image, which we refer to as sub-images. We propose a procedure for dictionary design and supervised local sparse coding of sub-image heterogeneous features. We perform image retrieval on the Microsoft Research Cambridge image dataset and show that the proposed features outperform the spatial pyramid features obtained using dense descriptors. Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri, Andreas Spanias |
ICIP | 4 |
| 2012 | On the Effectiveness of Multiple Antennas in Distributed Detection over Fading MACsabstractA distributed detection problem over fading Gaussian multiple-access channels is considered. Sensors observe a phenomenon and transmit their observations to a fusion center using the amplify and forward scheme. The fusion center has multiple antennas with different channel models considered between the sensors and the fusion center, and different cases of channel state information are assumed at the sensors. The performance is evaluated in terms of the error exponent for each of these cases, where the effect of multiple antennas at the fusion center is studied. When there is channel information at the sensors, the gain in error exponent due to having multiple antennas at the fusion center is shown to be limited to a factor of 8/π for Rayleigh fading channels between the sensors and the fusion center, and independent of the number of antennas at the fusion center. Simple practical schemes and numerical methods using semidefinite relaxation techniques are presented that utilize the limited possible gains available. Simulations are used to establish the accuracy of the results. Mahesh K. Banavar, Anthony D. Smith, Cihan Tepedelenlioglu, Andreas Spanias |
IEEE Trans. Wirel. Commun. | 4 |
| 2011 | Work in progress: The J-DSP/ESE software for analyzing Earth systems signalsabstractJava-DSP (J-DSP) is a free online Java applet that has been extensively used in signal processing education and research. We present the functionalities of J-DSP Earth Systems Edition (J-DSP/ESE) that uses the basic architecture of J-DSP, but has functions tailor-made for Earth systems signals. No text-based programming is required, so that users can focus on understanding signal processing concepts. Here, we describe the functionalities in the current version of J-DSP/ESE. A coherency analysis of Earth time series is presented. In order to overcome the inherent limitations of J-DSP/ESE in terms of memory and computations, a standalone Java application is proposed. This will greatly enhance the functionalities of the existing J-DSP/ESE applet. The standalone application will be platform independent and available for free. These additional functionalities of the application make it suitable for use in research as well as education. Linda Hinnov, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
FIE | 3 |
| 2011 | Work in progress - Interactive signal-processing labs and simulations on iOS devicesabstractHandheld devices are increasingly finding more applications in STEM education. In this paper, we present the design of an interactive signal processing simulation software operating on both the iPhone OS (iOS) and Android platforms. This object-oriented application is called i-JDSP and is conceptually based on the award-winning Java-DSP (J-DSP) simulation environment. The i-JDSP app offers a user-friendly visual programming interface and provides users with a compelling multi-touch programming experience. It supports basic signal processing simulation functions such as the FFT, filtering, frequency response, pole-zero plots, and sound recording and playback. Initial assessments have been promising and we believe that this new attractive smartphone interface will make signal processing education among undergraduate students more appealing. Jinru Liu, Andreas Spanias, Mahesh K. Banavar, Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Xue Zhang 0002 |
FIE | 2 |
| 2011 | Improved sparse coding using manifold projectionsabstractSparse representations using predefined and learned dictionaries have widespread applications in signal and image processing. Sparse approximation techniques can be used to recover data from its low dimensional corrupted observations, based on the knowledge that the data is sparsely representable using a known dictionary. In this paper, we propose a method to improve data recovery by ensuring that the data recovered using sparse approximation is close its manifold. This is achieved by performing regularization using examples from the data manifold. This technique is particularly useful when the observations are highly reduced in dimensions when compared to the data and corrupted with high noise. Using an example application of image inpainting, we demonstrate that the proposed algorithm achieves a reduction in reconstruction error in comparison to using only sparse coding with predefined and learned dictionaries, when the percentage of missing pixels is high. Karthikeyan Natesan Ramamurthy, Jayaraman J. Thiagarajan, Andreas Spanias |
ICIP | 3 |
| 2011 | Error probability-based optimal training for linearly decoded orthogonal space-time block coded wireless systemsabstractAn optimal training strategy is devised for the linearly decoded orthogonal space–time block coded (OSTBC) wireless systems in quasi-static fading channel, based on the performance analysis using pairwise error probability (PEP) and symbol error probability (SEP). The PEP/SEP analyses allow us to find a generic expression for the performance improvement due to optimal training compared to the conventional case for OSTBC system equipped with any number of transmit and receive antennas and any linear modulation scheme. It is observed that the linear processing in the receiver, the most attractive feature of OSTBC, although destroys the orthogonality in the presence of channel estimation error, does not reduce diversity, but causes performance penalty as a loss of signal-to-noise ratio (LoSNR) due to training. This loss is quantified analytically and minimised by optimal allocation of power between training and data symbols. The performance of optimal power allocation improves with the higher number of space–time blocks in a frame. Furthermore, the LoSNR depends only on the OSTBC and is independent of any modulation scheme and the full rate Alamouti and other high rate OSTBCs suffer more in terms of performance due to training compared to the lower rate OSTBC. Khawza I. Ahmed, Cihan Tepedelenlioglu, Andreas Spanias, Mohammad N. Patwary, Hongnian Yu |
IET Commun. | 3 |
| 2011 | Optimality and stability of the K-hyperline clustering algorithm
Jayaraman J. Thiagarajan, Karthikeyan Natesan Ramamurthy, Andreas Spanias |
Pattern Recognit. Lett. | 3 |
| 2011 | On the Asymptotic Efficiency of Distributed Estimation Systems With Constant Modulus Signals Over Multiple-Access ChannelsabstractA distributed estimation problem is considered with multiple-access channels between sensors and a fusion center. The sensors phase-modulate their noisy observations before transmitting them to the fusion center, where a signal parameter is estimated. The asymptotic efficiency of this estimator is then determined by using two inequalities that relate the Fisher information and the characteristic function. A necessary and sufficient condition for equality is found for the first time in the literature. The loss in efficiency of the distributed estimation scheme relative to the centralized approach is quantified for different sensing noise distributions. It is shown that this distributed estimation system does not incur an efficiency loss if and only if the sensing noise distribution is Gaussian. Cihan Tepedelenlioglu, Mahesh K. Banavar, Andreas Spanias |
IEEE Trans. Inf. Theory | 3 |
| 2010 | Distributed detection over fading macs with multiple antennas at the fusion centerabstractWe consider a distributed detection problem over fading multiple-access channels. Sensors observe a phenomenon and transmit their observations to a fusion center using the amplify-and-forward scheme. The fusion center has multiple antennas and uses the transmissions from the sensors to run a detection algorithm. The channels are Ricean fading, and the sensors have no channel information. The performance is evaluated in terms of error exponent and compared with the AWGN channels case. The benefit of having multiple antennas at the fusion center is also quantified. Mahesh K. Banavar, Anthony D. Smith, Cihan Tepedelenlioglu, Andreas Spanias |
ICASSP | 4 |
| 2010 | An individually tunable perfect reconstruction filterbank for banded waveguide synthesisabstractIn gaming and other multimedia applications, there is an increasing need to simulate and transform the sounds of physical interactions with real-world objects. When these sounds are based upon source recordings, this motivates the need for flexible, yet ecologically valid resynthesis. With modal analysis as a front end, banded waveguide resynthesis enables flexible modifications to dynamic spectral content and varied forms of interaction. Unfortunately, traditional banded waveguide models using biquad filters can have problems with undesired transient behavior. To this end, we introduce a new perfect reconstruction filterbank that allows for each filter to be individually tuned to a desired mode frequency and precisely cancel the other selected frequencies. Comparisons with traditional biquad models amply demonstrate the improved transient response and overall fidelity of our model. In addition, examples of resynthesis, transformation, and alternative interaction with real-world recordings are given. Alex Fink, Harvey D. Thornburg, Andreas Spanias |
ICASSP | 3 |
| 2010 | Least-squares based feature extraction and sensor fusion for explosive detectionabstractThe effective and reliable detection of explosive compounds in complex environments is an important problem in many environment and security-related applications. This paper develops an explosive detection approach based on multi-modal sensing and sensor data fusion. A least-squares feature extraction technique is designed to isolate explosive signatures in data collected using electrochemical and polymer nanojunction sensors. The information obtained from the two sensors is then efficiently combined using a Bayesian decision fusion scheme. Results are presented for the detection of the explosive compound TNT showing the merit of the proposed approach. Narayan Kovvali, Chad Prior, Karel Cizek, Michal Galik, Alvaro Diaz, Erica Forzani, Avi Cagan, Nongjian Tao, Douglas Cochran, Andreas Spanias, Ray Tsui |
ICASSP | 11 |
| 2010 | An auditory-domain based speech enhancement algorithmabstractTypically, speech enhancement algorithms minimize a suitable error criterion in the spectral or time domain. Although the error criterions have included perceptual properties such as masking thresholds, non-uniform frequency resolution and sensitivity of the auditory system, these are only done heuristically and the error criterion does not explicitly include an auditory model in their formulation. In this paper, we propose an auditory-domain based speech enhancement algorithm that minimizes the distortion between the auditory representation of the estimated and desired signal. Simulation results indicate that the proposed algorithm performs effectively under different noise conditions and also results in a lower average loudness error. Harish Krishnamoorthi, Andreas Spanias, Visar Berisha, Homin Kwon, Harvey D. Thornburg |
ICASSP | 2 |
| 2010 | Combining semantic, social, and acoustic similarity for retrieval of environmental soundsabstractRecent work in audio information retrieval has demonstrated the effectiveness of combining semantic information, such as descriptive, tags with acoustic content. However, these methods largely ignore the possibility of tag queries that do not yet exist in the database and the possibility of similar terms. In this work, we propose a network structure integrating similarity between semantic tags, content-based similarity between environmental audio recordings, and the collective sound descriptions provided by a user community. We then demonstrate the effectiveness of our approach by comparing the use of existing similarity measures for incorporating new vocabulary into an audio annotation and retrieval system. Brandon Mechtley, Gordon Wichern, Harvey D. Thornburg, Andreas Spanias |
ICASSP | 4 |
| 2010 | Time-frequency based biological sequence queryingabstractWe investigate the use of time-frequency (TF) methods to query biological sequences in search of regions of similarity or critical relationships among the sequences. Existing querying approaches are insensitive to repeats, especially in low-complexity regions, and do not provide much support for efficiently querying sub-sequences with inserts and deletes (or gaps). Our approach uses highly-localized basis functions and multiple transformations in the TF plane to map characters in a sequence as well as different properties of a sub-sequence, such as its position in the sequence or number of gaps between sub-sequences. We analyze gapped query-based alignment methods using transformations in the TF plane while demonstrating the method's possible operation in real-time without pre-processing. The algorithm's performance is compared to the widely-accepted BLAST alignment approach, and a significance improvement is observed for queries with repetitive segments. Lakshminarayan Ravichandran, Antonia Papandreou-Suppappola, Andreas Spanias, Zoé Lacroix, Christophe Legendre |
ICASSP | 3 |
| 2010 | Automatic audio tagging using covariate shift adaptationabstractAutomatically annotating or tagging unlabeled audio files has several applications, such as database organization and recommender systems. We are interested in the case where the system is trained using clean high-quality audio files, but most of the files that need to be automatically tagged during the test phase are heavily compressed and noisy, for instance if they were captured on a mobile device. In this situation we assume the audio files follow a covariate shift model in the acoustic feature space, i.e., the feature distributions are different in the training and test phases, but the conditional distribution of labels given features remains unchanged. Our method uses a specially designed audio similarity measure as input to a set of weighted logistic regressors, which attempt to alleviate the influence of covariate shift. Results on a freely available database of sound files contributed and labeled by non-expert users, demonstrate effective automatic tagging performance. Gordon Wichern, Makoto Yamada, Harvey D. Thornburg, Masashi Sugiyama, Andreas Spanias |
ICASSP | 5 |
| 2010 | Enhanced direction of arrival estimation via reassigned space-time-frequency methodsabstractTwo new space-time-frequency direction of arrival estimation algorithms are presented that decrease the estimation variance under specific conditions of narrow differential direction of arrival when multiple signals impinge upon the sensor array. The first algorithm, called Wide-Lane, provides a modest extension of existing space-time-frequency (STF) techniques by integrating a wide path along the signal's instantaneous frequency track through the time-frequency (TF) plane. The second algorithm, called Reassigned STF, utilizes the TF reassignment method to increase the signal's localization within the TF plane prior to estimating the direction of arrival with MUSIC. The performance variation of space-time MUSIC to a variety of signal structures is also presented. Steven R. Miller, Andreas Spanias, Antonia Papandreou-Suppappola, Robert W. Santucci |
ISCAS | 2 |
| 2010 | Segmentation, Indexing, and Retrieval for Environmental and Natural SoundsabstractWe propose a method for characterizing sound activity in fixed spaces through segmentation, indexing, and retrieval of continuous audio recordings. Regardingsegmentation, we present a dynamic Bayesian network (DBN) that jointly infers onsets and end times of the most prominent sound events in the space, along with an extension of the algorithm for covering large spaces with distributed microphone arrays. Each segmented sound event isindexedwith a hidden Markov model (HMM) that models the distribution of example-based queries that a user would employ toretrievethe event (or similar events). In order to increase the efficiency of the retrieval search, we recursively apply a modified spectral clustering algorithm to group similar sound events based on the distance between their corresponding HMMs. We then conduct a formal user study to obtain the relevancy decisions necessary for evaluation of our retrieval algorithm on both automatically and manually segmented sound clips. Furthermore, our segmentation and retrieval algorithms are shown to be effective in both quiet indoor and noisy outdoor recording conditions. Gordon Wichern, Jiachen Xue, Harvey D. Thornburg, Brandon Mechtley, Andreas Spanias |
IEEE Trans. Speech Audio Process. | 5 |
| 2009 | Acquiring and Classifying Signals from Nanopores and Ion-Channels
Bharatan Konnanath, Prasanna Sattigeri, Trupthi Mathew, Andreas Spanias, Shalini Prasad, Michael Goryll, Trevor J. Thornton, Peter Knee |
ICANN (2) | 4 |
| 2009 | Nonlinear acoustic echo control using an accelerometerabstractA typical echo canceller uses an FIR adaptive filter to estimate the impulse response of the near-end echo path. However nonlinear echo components due to the loudspeaker and driving amplifier are not captured by such an arrangement. We propose a novel scheme for nonlinear acoustic echo cancellation where an accelerometer sensor captures loudspeaker vibration and this sensor signal is used along with the error signal to adapt the adaptive filter coefficients. The accelerometer sensor is used in lieu of explicit nonlinear modeling in the echo cancelling loop and hence it is compatible with typical NLMS and RLS echo cancellation algorithms. Experiments with an accelerometer mounted on the magnet of a hands-free kit show an echo reduction improvement by as much as 15 dB in situations of nonlinear loudspeaker distortion. Seth B. Suppappola, Andreas Spanias |
ICASSP | 3 |
| 2009 | Low-complexity sinusoidal component selection using loudness patternsabstractSinusoidal modeling of audio at low-bit rates involves selecting a limited number of parameters according to a quantitative or perceptual criterion. Most perceptual sinusoidal component selection strategies are computationally intensive and not suitable for real-time applications. In this paper, a computationally efficient sinusoidal selection algorithm based on a novel hybrid loudness estimation scheme is presented. The hybrid scheme first estimates efficiently the loudness of a multi-tone signal from the loudness patterns of its constituent sinusoidal components. Then it refines this estimate by performing a full evaluation of loudness but only in select critical bands. Experimental results show that the proposed technique maintains a low perceptual sinusoidal synthesis error at a much lower computational complexity. Harish Krishnamoorthi, Visar Berisha, Andreas Spanias, Homin Kwon |
ICASSP | 3 |
| 2009 | Energy-constrained discriminant analysisabstractDimensionality reduction algorithms have become an indispensable tool for working with high-dimensional data in classification. Linear discriminant analysis (LDA) is a popular analysis technique used to project high-dimensional data into a lower-dimensional space while maximizing class separability. Although this technique is widely used in many applications, it suffers from overfitting when the number of training examples is on the same order as the dimension of the original data space. When overfitting occurs, the direction of the LDA solution can be dominated by low-energy noise and therefore the solution becomes non-robust to unseen data. In this paper, we propose a novel algorithm, energy-constrained discriminant analysis (ECDA), that overcomes the limitations of LDA by finding lower dimensional projections that maximize inter-class separability, while also preserving signal energy. Our results show that the proposed technique results in higher classification rates when compared to comparable methods. The results are given in terms of SAR image classification, however the algorithm is broadly applicable and can be generalized to any classification problem. Scott Philips, Visar Berisha, Andreas Spanias |
ICASSP | 3 |
| 2009 | Multi-channel audio segmentation for continuous observation and archival of large spacesabstractIn most real-world situations, a single microphone is insufficient for the characterization of an entire auditory scene. This often occurs in places such as office environments which consist of several interconnected spaces that are at least partially acoustically isolated from one another. To this end, we extend our previous work on segmentation of natural sounds to perform scene characterization using a sparse array of microphones, strategically placed to ensure that all parts of the environment are within range of at least one microphone. By accounting for which microphones are active for a given sound event, we perform a multi-channel segmentation that captures sound events occurring in any part of the space. The segmentation is inferred from a custom dynamic Bayesian network (DBN) that models how event boundaries influence changes in audio features. Example recordings illustrate the utility of our approach in a noisy office environment. Gordon Wichern, Harvey D. Thornburg, Andreas Spanias |
ICASSP | 3 |
| 2009 | Fast image registration with non-stationary Gauss-Markov random field templatesabstractNon-stationary Gauss-Markov random fields are required in modeling images with complex patterns. In this paper, we propose a framework for registering images to a non-stationary Gauss-Markov random field template in an M×M lattice, with a complexity of order M2log M, considering only global translations. We simplify the likelihood computation by expressing it as a scalar product and we estimate the maximal likelihood translation using 2-D FFTs. We demonstrate the utility of this framework by applying it to image registration in a wavelet-domain template learning application. Results reveal that significant complexity reduction is achieved in image registration compared to straightforward registration in the wavelet domain. Karthikeyan Natesan Ramamurthy, Jayaraman J. Thiagarajan, Andreas Spanias |
ICIP | 3 |
| 2009 | A Sensor Network for Real-time Acoustic Scene AnalysisabstractAcoustic scene analysis can be used to extract relevant information in applications such as homeland security, surveillance and environmental monitoring. Wireless sensor networks have been of particular interest in monitoring acoustic scenes. Sensors embedded in such a network typically operate under several constraints such as low power and limited bandwidth. In this paper, we consider resource-efficient acoustic sensing tasks that extract and transmit relevant information to a central station where information assessment can be conducted. We propose a series of acoustic scene analysis tasks that are performed in a hierarchical manner. Hierarchical tasks include sound and speech discrimination, estimation of the number of speakers from the acquired sound, gender and emotional state, and ultimately voice monitoring and key word spotting. We apply support vector machine and Gaussian mixture model algorithms on sound features. A real-time implementation is accomplished using crossbow motes interfaced with a TI DSP board. A series of experiments are presented to characterize the performance of the algorithms under different conditions. Homin Kwon, Harish Krishnamoorthi, Visar Berisha, Andreas Spanias |
ISCAS | 4 |
| 2009 | Perceptually-Motivated All-Pole ModelingabstractA new speech analysis-synthesis approach that is based on a perceptually-motivated all-pole (PMAP) modeling is described. The main idea is to directly estimate the perceptually relevant pole locations using an auditory excitation pattern-matching method. The all-pole model is synthesized using the perceptual poles and produces improved spectral fitting. We show that the prediction residual obtained from the PMAP analysis has lower perceptual loudness relative to that of the conventional LP. Venkatraman Atti, Andreas Spanias |
IEEE Signal Process. Lett. | 2 |
| 2009 | A Frequency/Detector Pruning Approach for Loudness EstimationabstractIn this letter, we propose a frequency and detector pruning approach for reducing the computational complexity associated with loudness estimation. The frequency pruning approach exploits the principles of psychoacoustics such that the total neural activity is preserved. The detector pruning approach evaluates the excitation/loudness patterns at nonuniform sample locations and employs signal interpolation techniques to obtain their corresponding high resolution estimates. Comparative results with the Moore and Glasberg loudness estimation process reveal that the proposed pruning approach for loudness estimation performs consistently well for different types of audio signals with a significant reduction in the computational complexity. Harish Krishnamoorthi, Andreas Spanias, Visar Berisha |
IEEE Signal Process. Lett. | 2 |
| 2008 | Ultrasound imaging media layer texture analysis of the carotid arteryabstractThe intima-media thickness (IMT) of the common carotid artery (CCA) is widely used as an early indicator of cardiovascular disease (CVD). It was proposed but not thoroughly investigated that the media layer (ML), its composition and texture, may be indicative for identifying the risk of stroke and differentiating between patients of high and low risk. In this study we investigate the usefulness of texture analysis of the ML of the CCA. The study was performed on 100 longitudinal ultrasound images acquired from asymptomatic subjects at risk of atherosclerosis. The images were separated into three different age groups, namely below 50, 50 to 60, and above 60 years old. A total of 61 different texture features were extracted from the intima-media complex (IMC), ML and the intima layer (IL). The IMC and ML were segmented manually by a neurovascular expert and automatically by a snakes segmentation system. It was shown that texture features extracted from the IL, ML and IMC are significantly different (mean, gray scale median (GSM), standard deviation, contrast, difference variance, periodicity) and that some of them can be associated with the increase (difference variance, entropy) or decrease (GSM) of patientpsilas age. It was also shown that the GSM of the ML falls linearly with increasing ML thickness (MLT) and with increasing age. Further research on more subjects is required for estimating other features that may provide information for patients at risk of stroke. Philipos C. Loizou, Marios Pantziaris, Andrew Nicolaides, Andreas Spanias, Marios S. Pattichis, Constantinos S. Pattichis |
BIBE | 4 |
| 2008 | Performance of distributed estimation over multiple access fading channels with partial feedbackabstractWe consider a wireless sensor network for distributed estimation over Rayleigh fading channels. The sensors transmit their observations over fading channels to a fusion center, where a source parameter is estimated. Since the sensor transmissions add incoherently over a multiple access channel, we consider partial channel knowledge at the sensors to improve performance. We calculate the variance of the estimate when the channel phase is quantized uniformly and fed back to the sensors. We show that as few as 3 bits of feedback is sufficient for a loss in performance of about 5%. We also show that the performance is robust in the presence of feedback errors. Mahesh K. Banavar, Cihan Tepedelenlioglu, Andreas Spanias |
ICASSP | 3 |
| 2008 | A low-complexity loudness estimation algorithmabstractAudio processing applications such as rate determination, bandwidth extension, compression, and noise reduction make use of loudness metrics. Most loudness estimation algorithms are computationally expensive and often not suitable for real time applications. In this paper, we present a low-complexity loudness estimation algorithm applicable to both steady and time-varying sounds. The model computes an estimate of the excitation pattern by simultaneously pruning the frequency components and detector locations. Comparative results indicate that the proposed algorithm performs consistently well for different types of audio signals at a reduced complexity. Harish Krishnamoorthi, Visar Berisha, Andreas Spanias |
ICASSP | 3 |
| 2008 | Fast query by example of environmental sounds via robust and efficient cluster-based indexingabstractThere has been much recent progress in the technical infrastructure necessary to continuously characterize and archive all sounds, or more precisely auditory streams, that occur within a given space or human life. Efficient and intuitive access, however, remains a considerable challenge. In specifically musical domains, i.e., melody retrieval, query-by-example (QBE) has found considerable success in accessing music that matches a specific query. We propose an extension of the QBE paradigm to the broad class of natural and environmental sounds, which occur frequently in continuous recordings. We explore several cluster-based indexing approaches, namely non-negative matrix factorization (NMF) and spectral clustering to efficiently organize and quickly retrieve archived audio using the QBE paradigm. Experiments on a test database compare the performance of the different clustering algorithms in terms of recall, precision, and computational complexity. Initial results indicate significant improvements over both exhaustive search schemes and traditional K- means clustering, and excellent overall performance in the example-based retrieval of environmental sounds. Jiachen Xue, Gordon Wichern, Harvey D. Thornburg, Andreas Spanias |
ICASSP | 4 |
| 2008 | Gradient projection-based channel equalization under sustained fading
Venkatraman Atti, Andreas Spanias, Kostas Tsakalis, Constantinos Panayiotou, Leonidas D. Iasemidis, Visar Berisha |
Signal Process. | 2 |
| 2007 | A Scalable Bandwidth Extension AlgorithmabstractMost modern bandwidth extension techniques predict the high- frequency band based on features extracted from the lower band. While this works for some frames, problems arise when the correlation between the low and the high band is insufficient. In these situations, additional high-band information must be sent to the decoder. In this paper, we propose a scalable speech coding method based on the principles of bandwidth extension. The rate selection is based on explicit psychoacoustic criteria, while the bandwidth extension is performed using a constrained MMSE estimation technique. Objective and subjective evaluations indicate that the proposed system performs at a lower average bit rate when compared to other similar algorithms while improving speech quality. Visar Berisha, Andreas Spanias |
ICASSP (4) | 2 |
| 2007 | Sparse Manifold Learning with Applications to SAR Image ClassificationabstractNonlinear data-driven dimensionality reduction techniques have recently gained popularity due to the emergence of high dimensional data sets. The algorithmic complexity and storage requirements of these techniques, however, can make them prohibitive in resource-limited applications. It is therefore beneficial to reduce the number of exemplar samples required for performing an out-of-sample extension to a test point. In this paper, we propose a novel method for selecting a minimal set of exemplars and performing the out-of-sample extension. In the case of two-class target recognition with synthetic aperture radar (SAR) data, we compare the efficacy of the proposed approach with other approaches for selecting a subset of the available training samples. We show that the proposed algorithm outperforms the existing methods by providing low-dimensional embeddings that maintain interclass separability using fewer retained exemplars. Visar Berisha, Nitesh Shah, Donald E. Waagen, Harry Schmitt, Salvatore Bellofiore, Andreas Spanias, Douglas Cochran |
ICASSP (3) | 6 |
| 2007 | Analysis of the Matrix Processing (MxP) ArchitectureabstractIn this paper, we describe and evaluate a new DSP architecture called the matrix processing (MxP™). MxP exploits data parallelism using hardwired matrix operations and instructions. The MxP matrix computation capabilities are optimized for multidimensional transforms and filtering. A detailed analysis of the MxP performance, in terms of precision, execution time, and instruction memory requirements, is presented. We focus particularly on matrix and vector manipulations of the type embedded in video processing standards, e.g., filtering, DCT, and motion estimation. We present comparison tables of the performance of the MxP with other widely used DSPs such as the TI TMS320c55x™, and the TI TMS320c64x™. Lakshmi Girija, Andreas Spanias |
ICASSP (2) | 2 |
| 2007 | Tracking the Path Shape Qualities of Human MotionabstractWe propose a probabilistic generative model for extracting intended path shape qualities of an object moving under human control in real time. At each instant, we decide whether the object is moving in a straight, curved, or random path, or whether it has stopped moving. Our model incorporates sensor noise as well as human imperfections in the intended motion. As well as tracking the object's position, velocity, and motion direction, we compute the posterior probability of each shape quality hypothesis given all sensed-data in the horizon [t - N + 1,t]; the hypothesis maximizing this posterior is taken as the decision. The posterior is computed using the unscented Kalman filter (UKF), as our model is inherently nonlinear. The path-shape quality tracking is successfully embedded in a hybrid physical-digital interface where the position of an illuminated ball, sensed by a low-cost video camera array, triggers multimodal feedback in a mediated learning environment. We show successful results on a variety of real-world motion paths where the participant is given only verbal descriptions of how to move. Our generative model is further validated by user studies involving a simple color-based interaction, where participants discover shape quality controls as they interact. Kai Tu, Harvey D. Thornburg, Matthew Fulmer, Andreas Spanias |
ICASSP (2) | 4 |
| 2007 | Analysis and Mitigation of Doppler Rate Effect in a Multipath ChannelabstractEstablishing robust communications in the presence of Doppler rate is a requirement for several waveforms such as SATCOM and HF. In this paper, we motivate the need to model Doppler rate by analyzing the kinematics of motion in the 2-D plane. We formulate the time-varying, multipath channel output as a composite chirp signal with a convolutive mix. We propose a data-aided method to transform the channel output into a composite signal whose chirp parameters can be estimated by existing techniques. We present a method to decouple chirps from the channel output and estimate channel coefficients. The performance of the channel estimator is characterized in terms of mean-squared-error. The sensitivity of the algorithm to errors in chirp estimates is studied as a function of symbol rate. Ghassan Maalouli, Andreas Spanias |
ICC | 2 |
| 2007 | An Adaptive Low Rank Algorithm for Semispherical Antenna ArraysabstractAdaptive algorithms have been applied in beam-forming applications with adaptive linear and planar arrays. In this paper, we are concerned with adaptive beamforming algorithms for 3D semispherical arrays. We present a low rank adaptive algorithm, which we call the 3D low rank recursive least squares approximate power iteration. The algorithm uses a high performance subspace tracker integrated with a low rank RLS beamformer. The algorithm is computationally efficient compared to the RLS beamformer with better steady state performance and, in certain situations, equally fast convergence. The algorithm is evaluated in terms of computational complexity and adaptation speed with comparative results to the RLS beamformer. Jeffrey A. Foutz, Andreas Spanias |
ISCAS | 2 |
| 2007 | Dual-Mode Wideband Speech CompressionabstractMany bandwidth extension techniques attempt to predict the high-band frequencies based on features extracted from the lower band. Recent work suggests that such methods are limiting because the correlation between the low band and the high band is insufficient for adequate representation. As a result, additional high-band information must be sent to the decoder. In this paper, we propose a dual mode wideband speech coding algorithm based on the principles of bandwidth extension. The principal contributions include a mode selection algorithm based on greedy algorithm that maximizes the loudness criteria, and a bandwidth extension algorithm based on a constrained MMSE estimator. Results reveal that the proposed system improves the quality of narrowband speech while performing at a lower bit rate. Visar Berisha, Andreas Spanias |
MMSP | 2 |
| 2006 | Real-Time Collaborative Monitoring in Wireless Sensor NetworksabstractIn recent years, wireless sensor networks (WSN) have shown success in distributed real-time signal processing systems. In collaborative signal processing environments, each sensor is responsible for extracting pertinent information from the surrounding environment and transmitting it to other sensors and/or to the main processing station. Often times, the sensors operate under a number of constraints, such as limited processing power and low bandwidth. In this paper we propose a collaborative signal processing framework that is implemented in an acoustic monitoring scenario. A low-complexity voice activity detector and a gender classifier are implemented on the Crossbow sensor motes. A series of experiments are presented that characterize the performance of the algorithms under varying SNR conditions and in different environments. Visar Berisha, Homin Kwon, Andreas Spanias |
ICASSP (3) | 3 |
| 2006 | Equalization of a Time-Varying Channel in the Presence of Doppler RateabstractIn this paper, we consider the problem of equalizing a time-varying channel in the presence of Doppler rate. Blind estimation of a time-varying channel in the presence of Doppler frequency alone based on a complex-exponential, basis expansion model (CE-BEM) has been studied by several researchers in the field. Recently, we proposed a fractionally spaced channel model that does not require basis expansion and possesses a structure that models Doppler rate. We presented a data-aided, LS method for channel estimation. In this work we analyze the performance of the channel estimator using exact and perturbed Doppler estimates. We present a symbol-by-symbol based, decision-feedback equalizer for the time-varying channel and assess its performance in terms of BER Ghassan Maalouli, Andreas Spanias |
ICASSP (3) | 2 |
| 2006 | Optimal Training and Performance Analysis of Space-Time-Frequency Coded OFDMabstractWe consider space-time-frequency (STF) coded MIMO-OFDM systems with imperfect channel estimates due to channel estimation error. We address the problem of optimal training design for acquiring the channel state information (CSI) and performance analysis using pairwise error probability (PEP) in the presence of channel estimation error. The optimal power allocation factor between the training and the information-bearing data is determined using the novel training-based PEP for MIMO-OFDM. Furthermore, the loss of coding gain due to channel estimation error is quantified to demonstrate performance benefits of power allocation optimization in MIMO-OFDM systems between the training and information symbols. Khawza I. Ahmed, Cihan Tepedelenlioglu, Andreas Spanias |
ICC | 3 |
| 2006 | Real-time acoustic monitoring using wireless sensor motesabstractWireless sensor networks (WSN) have recently gained popularity in distributed monitoring and surveillance applications. The objective of these devices is to extract pertinent information under several constrains such as low computational capabilities, limited arithmetic precision, and the need to conserve power. One of the most revealing environmental cues is audio. In this paper, we propose a voice activity detector and a simple gender classifier for use in a distributed acoustic sensing system. This algorithm makes use of low-complexity audio features and a pre-trained regression tree to classify incoming speech by gender. The algorithm is implemented real-time on the Crossbow sensor motes and a series of results are given that characterize the algorithm performance and complexity. Challenges in this real-time implementation include designing the algorithm and software architecture such that the signal processing is appropriately distributed between the sensor mote and the base station. At the base station, a data fusion algorithm considers a linear combination of individual mote decisions to form a final decision. Visar Berisha, Homin Kwon, Andreas Spanias |
ISCAS | 3 |
| 2006 | A transform-domain G-PrOBE algorithmabstractWe propose a computationally attractive transform-domain gradient-with-projection-on-bounding-ellipsoid (G-PrOBE) algorithm. The main idea in the proposed frequency-domain G-PrOBE algorithm is that the system parameters are computed by projecting the gradient estimate onto a set that is updated using a priori information from the instantaneous error magnitude. Experimental results demonstrate that the transform-domain G-PrOBE algorithm is more robust towards burst errors and tends to exhibit smaller parameter variance relative to gradient-type algorithms especially under low signal excitation A. Natarajan, Venkatraman Atti, Andreas Spanias, Kostas Tsakalis, Leonidas D. Iasemidis |
ISCAS | 3 |
| 2006 | Bandwidth Extension of Audio Based on Partial Loudness CriteriaabstractMost modern speech coders operate on a limited bandwidth. This tends to decrease the naturalness of the synthesized audio and often also affects the intelligibility of certain sounds. While a few wideband speech coders have been standardized, implementing them in existing systems would require significant changes to the infrastructure. One solution is to use bandwidth extension techniques that predict the high-frequency band based on low-band features. Problems arise however when the correlation between the low and the high band is insufficient for an adequate representation of the wideband signal. In this paper, we propose a novel source-filter bandwidth extension algorithm that makes use of psychoacoustic concepts to determine the perceptual benefits that a particular audio frame gains from a more exact representation of the high band. Preliminary results indicate that the proposed system performs at a lower average bit rate when compared to other similar algorithms without compromising the audio quality Visar Berisha, Andreas Spanias |
MMSP | 2 |
| 2005 | PEP-based optimal training for MIMO systems in wireless channelsabstractIn this paper, PEP (pairwise error probability) is derived for MIMO systems in the presence of channel estimation error, where the channel is estimated using training insertion and linear channel estimation. This training-based PEP expression is utilized to design the training sequence and for optimal power allocation between the training and data symbols. The loss of coding gain due to the training is quantified at high SNR. It is shown that both the proposed PEP-based and the existing capacity-based optimal training approach of B. Hassibi and B. Hochwald (IEEE Trans. Inform. Theory, vol. 49, no. 4, pp. 951-963, 2003) reveal the same power allocation for constant modulus (CM) constellations. The performance of our optimal power allocation scheme falls between the nonoptimized equal power allocation and perfectly known channel cases, and converges to the latter for larger data transmission intervals. Numerical and simulation results are presented to justify the usefulness of our analytical findings. Khawza I. Ahmed, Cihan Tepedelenlioglu, Andreas Spanias |
ICASSP (3) | 3 |
| 2005 | Speech Analysis by Estimating Perceptually Relevant Pole LocationsabstractAn approach for estimating the perceptually-relevant pole locations is described. These "perceptual poles" are determined by using an auditory excitation pattern-matching method. The estimated perceptual poles are then used to construct a perceptually-motivated all-pole (PMAP) filter for use in speech analysis/synthesis. The proposed PMAP approach is compared against some of the existing perceptually-based linear prediction (LP) methods, i.e., the perceptual LP and the warped LP. The PMAP approach compares well against the perceptual LP and the warped LP in terms of speech reconstruction quality and estimation of the formant frequencies. Venkatraman Atti, Andreas Spanias |
ICASSP (1) | 2 |
| 2005 | Interactive Java modules for the MPEG-1 psychoacoustic model [audio coding teaching applications]abstractThis paper presents a collection of interactive Java modules for the purpose of introducing undergraduate DSP students to perceptual audio coding principles. This effort is part of a combined research and curriculum program funded by NSF that aims towards exposing undergraduate students to advanced concepts and research in signal processing. A computer laboratory with several supporting exercises and Java functions has been developed for use in our undergraduate DSP course. This exercise along with the accompanying Java software was assigned and assessed in the Summer of 2004 and will be reassessed in the Fall of 2004. Results of this assessment along with student comments are presented at the end of the paper. Andreas Spanias, Venkatraman Atti, Visar Berisha |
ICASSP (5) | 2 |
| 2005 | Enhancing the Quality of Coded Audio Using Perceptual CriteriaabstractCode excited linear predictive (CELP) coding standards often fail to properly represent non-speech signals because they are inherently optimized for speech. Most modern CELP coders include provisions for the inclusion of indirect perceptual criteria to counteract this problem; however no direct psychoacoustic models are employed. In this paper, we present a pre- and postprocessor for the vocoder that makes use of the MPEG-1 psychoacoustic model 1 in order to enhance the quality of the coded audio. A novel frequency-domain technique is proposed that attempts to shape the residual of a vocoder such that it falls below psychoacoustic thresholds Visar Berisha, Andreas Spanias |
MMSP | 2 |
| 2005 | Perceptual segmentation and component selection for sinusoidal representations of audioabstractThis paper presents two fundamental enhancements in a hybrid audio signal model consisting of sinusoidal, transient, and noise (STN) components. The first enhancement involves a novel application of a perceptual metric for optimal time segmentation for the analysis of transients. In particular, Moore and Glasberg's model of partial loudness is modified for use with general signals and then integrated into a novel time segmentation scheme. The second, and perhaps more significant STN enhancement is concerned with a new methodology for ranking and selection of the most perceptually relevant sinusoids. A systematic procedure is developed for the selection of a compact set of sinusoids and comparative results are given to demonstrate the merit of this method. Edward M. Painter, Andreas Spanias |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | On the performance of optimal training-based OFDM with channel estimation errorabstractWe analyze the performance of an uncoded orthogonal frequency-division multiplexing (OFDM) system in a quasistatic fading channel where optimal training is being employed for the acquisition of channel state information (CSI). A bit-error probability (BEP) expression is obtained for the linear minimum mean-square error (LMMSE) based channel estimator where the mean-square error (MSE) of the channel estimator is minimized by equispaced, equipowered training with the number of training tones equal to the channel length. Optimal power allocation is obtained by minimizing the BEP expressions. The performances of BPSK, QPSK and 16-QAM schemes are considered. It is shown that while equal power training performs around 3 dB worse than the known channel case, our optimal power channel estimator performance varies between these two, depending on the number of subcarrier to channel length ratio. Numerical examples are presented to further clarify and confirm the analytical findings. Khawza I. Ahmed, Cihan Tepedelenlioglu, Andreas Spanias |
GLOBECOM | 3 |
| 2004 | Effect of channel estimation on pair-wise error probability in OFDMabstractPair-wise error probability (PEP) is analyzed in the presence of channel estimation error (CEE) for orthogonal frequency division multiplexing (OFDM) in a quasi-static Rayleigh fading channel. Subcarriers in OFDM are grouped in an equi-spaced manner with the number of subcarriers in a group equal to the number of channel taps. One group is dedicated for training and the rest are for data transmission. A linear minimum mean square error (LMMSE) based channel estimation for coherent detection is considered in this paper. This enables the signal dependent CEE to be uncorrelated to the data. Since ML decoding employed with perfect CSI becomes suboptimal in the presence of CEE, an optimal ML decoding is derived. It is observed from our training-based PEP expression that the CEE does not reduce the diversity order but contributes to a loss of coding gain. Moreover, a code design criterion is established. To reduce the loss of coding gain, an optimal training scheme is developed, based on the PEP expression. Loss of performance due to an imperfect channel estimate is quantified in terms of bounds on bit error probability (BEP) for high SNR. Our analytical findings are corroborated by simulation examples. Khawza I. Ahmed, Cihan Tepedelenlioglu, Andreas Spanias |
ICASSP (4) | 3 |
| 2004 | Web-based experiments for introducing speech recognition basics in a DSP courseabstractIn this paper, we describe Web-based educational software tools tailored to expose students in an undergraduate DSP course to the basics of hidden Markov model (HMM)-based speech recognition. In particular, we developed Java software that enables on-line computer laboratories on the essential preprocessing, HMM, and Viterbi algorithms as used in a basic speech recognition task. The software is complemented by streaming lectures, a set of on-line demonstrations with animation, and exercises that take the student through HMM training and recognition. The software is being made available to students in the Fall of 2003 and we expect to present assessment results at the conference. Venkatraman Atti, Andreas Spanias |
ICASSP (5) | 2 |
| 2004 | On-line signal processing using J-DSPabstractThe Java-DSP (J-DSP) on-line laboratory software has been developed from the ground up at Arizona State University to support the computer lab portion of the senior-level DSP course. J-DSP provides capabilities for web-based DSP simulations that can be run using a PC with a Java-enabled browser. J-DSP is accompanied by exercises that actively engage students in several concepts including Z-transforms, filter design, spectral analysis, and random signal processing. Tools and on-line instruments for assessment of the J-DSP software and the associated lab exercises have been developed and described in this letter. Statistical analysis of the pre/post assessment data revealed that the effect of student involvement and the student learning has been enhanced by using J-DSP. Andreas Spanias, Venkatraman Atti, Antonia Papandreou-Suppappola, Khawza I. Ahmed, Moushumi Zaman, Thrassos Thrasyvoulou |
IEEE Signal Process. Lett. | 1 |
| 2004 | Measuring the direction and the strength of coupling in nonlinear Systems-a modeling approach in the State spaceabstractWe present a novel signal processing methodology to determine the direction and the strength of coupling between coupled nonlinear systems. The methodology is based on multivariate local linear prediction in the reconstructed state spaces of the observed variables from each multivariable nonlinear system. Application of the method is illustrated with systems of coupled Rossler and Lorenz oscillators in various coupling configurations. The obtained results are compared with ones produced by the use of the directed transfer function, a model-based method in the time domain. Through a surrogate analysis, it is shown that the proposed method is more reliable than the directed transfer function in identifying the direction and strength of the involved interactions. Balaji Veeramani, K. Narayanan, Awadhesh Prasad, Leonidas D. Iasemidis, Andreas Spanias, Kostas Tsakalis |
IEEE Signal Process. Lett. | 5 |
| 2003 | On-line laboratories for image and two-dimensional signal processing using 2D J-DSPabstractJava Digital Signal Processing (J-DSP) is a Java-based object-oriented programming environment that was developed at Arizona State University (ASU) for use in undergraduate- and graduate-level engineering classes. J-DSP is written as a platform-independent Java applet that resides on the Web and is thereby accessible by all students using a Web browser. We describe an innovative software extension to J-DSP, called 2D J-DSP, to accommodate on-line laboratories for two-dimensional digital signal processing. Two-dimensional DSP capabilities in J-DSP include: 2D signal generation; 2D FIR filter design and implementation; 2D transforms. Image processing capabilities include image restoration and enhancement. In order to illustrate 2D concepts graphically, contour (2D) and perspective (3D) plots have been incorporated in 2D J-DSP. Online laboratory exercises have been developed in the aforementioned areas for use in the graduate-level multidimensional signal processing and image processing courses at ASU, and are posted on the Website (http://jdsp.asu.edu). Statistical and qualitative evaluations that assess the learning experiences of the students that use 2D J-DSP are also presented. Muhammad Yasin, Lina J. Karam, Andreas Spanias |
ICASSP (3) | 3 |
| 2003 | Fast adaptive algorithms using eigenspace projections
N. Gopalan Nair, Andreas Spanias |
Signal Process. | 2 |
| 2002 | Partial band interference excision for GPS using frequency-domain exponentsabstractGlobal Positioning System (GPS) receivers for military applications are highly sensitive to hostile jamming. This paper describes a partial band frequency-domain interference excision technique for GPS receivers. This has been demonstrated to be quite effective in field tests. Although the results of the field tests are classified [6] and not presented here, the algorithm is described in the paper along with a series of simulation results that demonstrate the effectiveness of this technique. Brad Badke, Andreas Spanias |
ICASSP | 2 |
| 2002 | A simulation tool for introducing MPEG - audio (MP3) concepts in a DSP courseabstractThis paper presents a simulation tool1for introducing perceptual audio coding concepts in senior undergraduate and graduate DSP courses. The tool consists of a user-friendly graphical interface along with a complete MA TLAB realization of all aspects of the audio MPEG-1 Layer 3 (MP3) algorithm. The tool is accompanied by a series of computer experiments and exercises that can be used to provide hands-on training to class participants. The tool may also be used by instructors in a class setting to demonstrate key signal processing concepts associated with the processing of high-fidelity audio. The MATLAB MP3 tool has been used in Arizona State University undergraduate DSP courses as well as in a graduate course on speech and audio coding and in a continuing education short course. Evaluation of the tool is being performed by an educational software assessment specialist. Ram Rangachar, Andreas Spanias |
ICASSP | 2 |
| 2002 | Adaptive modified covariance algorithms for spectral analysis
Kyriakos Kitsios, Andreas Spanias, Bruno D. Welfert |
Signal Process. | 2 |
| 2001 | Perceptual segmentation and component selection in compact sinusoidal representations of audioabstractThis paper presents two fundamental enhancements in a hybrid audio signal model consisting of sinusoidal, transient, and noise (STN) components. The first enhancement involves a novel application of a perceptual metric for optimal time segmentation for the analysis of transients. In particular, Moore and Glasberg's model of partial loudness is modified for use with general signals and then integrated into a, novel time segmentation scheme. The second and perhaps more significant STN enhancement is concerned with a new methodology for ranking and selection of the most perceptually relevant sinusoids. Edward M. Painter, Andreas Spanias |
ICASSP | 2 |
| 2001 | Development of new functions and scripting capabilities in JavaA-DSP for easy creation and seamless integration of animated DSP simulations in Web coursesabstractArizona State University (ASU) has developed an on-line DSP laboratory that is based on an object-oriented Java tool called Java Digital Signal Processing (J-DSP). J-DSP is currently being used in a senior-level DSP course at ASU. J-DSP has a rich suite of signal processing functions that facilitate interactive on-line simulations of modern statistical signal and spectral analysis algorithms, filter design tools, QMF banks, and speech analysis. We present a series of significant functionality extensions of J-DSP enabling on-line laboratories in the systems-related areas of communications, image processing, and control. The extensions in communications are presented in some detail. In addition to these important functionality extensions, we present enhancements in the infrastructure of J-DSP that provide embedded scripting capabilities. Scripts enable easy creation and seamless integration of interactive animations in DSP Web content. The latter is very important for instructors creating their own DSP-related Web courses as it provides for user-friendly development of interactive visualization modules through the J-DSP applets. Andreas Spanias, Fikre Bizuneh |
ICASSP | 1 |
| 2001 | Perceptual component selection in sinusoidal coding of audioabstractThis paper presents a new method for selecting the sinusoids in a hybrid audio signal model consisting of sinusoidal, transient, and noise (STN) components. The method relies on a new methodology for ranking and selection of the most perceptually relevant sinusoids. We call this method excitation similarity weighting (ESW) ranking and selection procedure. Whereas current coders tend to choose maximum signal-to-mask ratio components the ESW methodology seeks to maximize the matching between the excitation patterns evoked by the coded and original signals on a short-time basis. It is shown that the proposed method provides a perceptual improvement. Edward M. Painter, Andreas Spanias |
MMSP | 2 |
| 2001 | Low bit-rate speech coding based on an improved sinusoidal model
Sassan Ahmadi, Andreas Spanias |
Speech Commun. | 2 |
| 2000 | Development and evaluation of a Web-based signal and speech processing laboratory for distance learningabstractWe describe an Internet-based signal processing laboratory that provides hands-on learning experiences in distributed learning environments. The laboratory is based on an object-oriented Java/sup TM/ tool called Java Digital Signal Processing (J-DSP). J-DSP has been developed at Arizona State University (ASU) and is being used for a virtual laboratory in a senior-level DSP course. J-DSP is written as a platform-independent Java applet that resides on the Web and is thereby accessible by all students through the use of a Web browser. J-DSP has a rich suite of signal processing functions that facilitate interactive on-line simulations of modern statistical signal and spectral analysis algorithms, filter design tools, QMF banks, and state-of-the-art vocoders. J-DSP is accompanied by administrative software tools for secure Internet-based lab-report submission and evaluation including servlets for maintaining Web-based grade books. A series of J-DSP laboratory exercises has been developed and delivered using the ASU distance learning facilities. Student evaluations as well as assessments by experts have been compiled and preliminary results are quite encouraging. Andreas Spanias, Susan Urban, Argyris Constantinou, Maya Tampi, Axel Clausen, Xiaopeng Zhang 0002, Jeffrey A. Foutz, Georgios Stylianou |
ICASSP | 1 |
| 2000 | Minimum-variance phase prediction and frame interpolation algorithms for low bit rate sinusoidal speech codingabstractA number of improved algorithms for phase prediction and frame interpolation in the context of sinusoidal speech coding are presented. A minimum-variance sinusoidal phase estimation scheme is proposed. It is shown that reasonably accurate estimates for short-time sinusoidal phases corresponding to voiced frames can be obtained. In addition, improved algorithms for interpolation of sine wave parameters are presented which result in further reduction in bit rate while preserving the subjective equality of the reproduced speech at low bit rates. The performance of the proposed algorithms were evaluated on a large speech database and the results of statistical analysis are provided. The proposed algorithms were successfully integrated into a 2.4 kbps sinusoidal coder, where speech of good quality intelligibility, and naturalness was obtained. Sassan Ahmadi, Andreas Spanias |
ISCAS | 2 |
| 2000 | Perceptual coding of digital audioabstractDuring the last decade, CD-quality digital audio has essentially replaced analog audio. Emerging digital audio applications for network, wireless, and multimedia computing systems face a series of constraints such as reduced channel bandwidth, limited storage capacity, and low cost. These new applications have created a demand for high-quality digital audio delivery at low bit rates. In response to this need, considerable research has been devoted to the development of algorithms for perceptually transparent coding of high-fidelity (CD-quality) digital audio. As a result, many algorithms have been proposed, and several have now become international and/or commercial product standards. This paper reviews algorithms for perceptually transparent coding of CD-quality digital audio, including both research and standardization activities. This paper is organized as follows. First, psychoacoustic principles are described, with the MPEG psychoacoustic signal analysis model 1 discussed in some detail. Next, filter bank design issues and algorithms are addressed, with a particular emphasis placed on the modified discrete cosine transform, a perfect reconstruction cosine-modulated filter bank that has become of central importance in perceptual audio coding. Then, we review methodologies that achieve perceptually transparent coding of FM- and CD-quality audio signals, including algorithms that manipulate transform components, subband signal decompositions, sinusoidal signal components, and linear prediction parameters, as well as hybrid algorithms that make use of more than one signal model. These discussions concentrate on architectures and applications of those techniques that utilize psychoacoustic models to exploit efficiently masking characteristics of the human receiver. Several algorithms that have become international and/or commercial standards receive in-depth treatment, including the ISO/IEC MPEG family (-1, -2, -4), the Lucent Technologies PAC/EPAC/MPAC, the Dolby AC-2/AC-3, and the Sony ATRAC/SDDS algorithms. Then, we describe subjective evaluation methodologies in some detail, including the ITU-R BS.1116 recommendation on subjective measurements of small impairments. This paper concludes with a discussion of future research directions. Edward M. Painter, Andreas Spanias |
Proc. IEEE | 2 |
| 2000 | An improved approach to robust speech recognition using minimum error classification
Min-Tau Lin, Andreas Spanias, Philipos C. Loizou |
Speech Commun. | 2 |
| 1999 | An interactive tool for bit error rate analysis of speech coding algorithmsabstractA GUI-based software tool that provides a framework for evaluation of different speech coding algorithms is presented. The tool is designed to measure the susceptibility of speech coding algorithms to errors added on the encoded bit-stream during transmission. In particular, the errors can be added individually to the parameters that comprise the encoded bit-stream. This enables a designer of a speech codec to evaluate its performance under adverse or impaired channel conditions. The tool is universally applicable to different speech coding algorithms, by means of a user-defined bit-stream definition file. In fact, the tool has been used in the past to evaluate a number of standardized speech coding algorithms. This paper describes the features of the software. In addition, the paper presents sample results generated by this tool during a study of the ETSI GSM Enhanced Full Rate algorithm. Hiren C. Bhagatwala, Edward M. Painter, Andreas Spanias |
ICASSP | 3 |
| 1999 | Cepstrum-based pitch detection using a new statistical V/UV classification algorithmabstractAn improved cepstrum-based voicing detection and pitch determination algorithm is presented. Voicing decisions are made using a multifeature voiced/unvoiced classification algorithm based on statistical analysis of cepstral peak, zero-crossing rate, and energy of short-time segments of the speech signal. Pitch frequency information is extracted by a modified cepstrum-based method and then carefully refined using pitch tracking, correction, and smoothing algorithms. Performance analysis on a large database indicates considerable improvement relative to the conventional cepstrum method. The proposed algorithm is also shown to be robust to additive noise. Sassan Ahmadi, Andreas Spanias |
IEEE Trans. Speech Audio Process. | 2 |
| 1999 | Improved speech recognition using a subspace projection approachabstractTwo class separability criteria based on the divergence measure are proposed to improve speech recognition performance. The average and weighted average divergence measures are used as criteria for finding a transformation matrix which maps the original features into a more discriminative subspace. Results are presented for a highly confusable task. Philipos C. Loizou, Andreas Spanias |
IEEE Trans. Speech Audio Process. | 2 |
| 1998 | A Java signal analysis tool for signal processing experimentsabstractA program to simulate discrete time linear systems is presented. The software is written as a Java applet and can be accessed on the Internet. Object oriented programming allows the user to construct and simulate a variety of systems. The program uses a graphical user interface which is easy to learn and it provides a visualization of the system and signal flow. The software is currently being used at Arizona State University to support an online software laboratory for a senior level digital signal processing (DSP) course. This paper presents our experiences gained by using the program in a class setting and gives examples of possible laboratory problems. Axel Clausen, Andreas Spanias, Anand Xavier, Maya Tampi |
ICASSP | 2 |
| 1998 | A new phase model for sinusoidal transform coding of speechabstractA phase modeling algorithm for sinusoidal analysis-synthesis of speech is presented, where short-time sinusoidal phases are approximated using a combination of linear prediction, spectral sampling, delay compensation, and phase correction techniques. The algorithm is different to phase compensation methods proposed for source-system LPC in that it has been tailored to sinusoidal representation of speech. Performance analysis on a large speech data base reveals an improvement in temporal and spectral signal matching, as well as in the subjective quality of reconstructed speech. The method can be applied to enhance phase matching in low bit rate sinusoidal coders, where underlying sine wave amplitudes are extracted from an all-pole model. Preliminary subjective results are presented for a 2.4 kb/s sinusoidal coder. Sassan Ahmadi, Andreas Spanias |
IEEE Trans. Speech Audio Process. | 2 |
| 1997 | A new sinusoidal phase modeling algorithmabstractA new phase modeling algorithm for sinusoidal analysis and synthesis of speech signals is presented. Short-time sinusoidal phases are efficiently approximated by incorporating linear prediction, spectral sampling, delay compensation, and phase correction techniques. The algorithm is different than phase compensation methods proposed for multi-pulse LPC in that it has been tailored to sinusoidal transform coding of speech signals. Performance analysis on a large speech database indicates considerable improvement in temporal and spectral matching between the original and reconstructed signals as compared to other sinusoidal phase models as well as improved subjective quality of the reproduced speech. Sassan Ahmadi, Andreas Spanias |
ICASSP | 2 |
| 1997 | HMM-based speech enhancement using harmonic modelingabstractThis paper describes a technique for reduction of non-stationary noise in electronic voice communication systems. Removal of noise is needed in many such systems, particularly those deployed in harsh mobile or otherwise dynamic acoustic environments. The proposed method employs state-based statistical models of both speech and noise, and is thus capable of tracking variations in noise during sustained speech. This work extends the hidden Markov model (HMM) based minimum mean square error (MMSE) estimator to incorporate a ternary voicing state, and applies it to a harmonic representation of voiced speech. Noise reduction during voiced sounds is thereby improved. Performance is evaluated using speech and noise from standard databases. The extended algorithm is demonstrated to improve speech quality as measured by informal preference tests and objective measures, to preserve speech intelligibility as measured by informal diagnostic rhyme tests, and to improve the performance of a low bit-rate speech coder and a speech recognition system when used as a pre-processor. Michael E. Deisher, Andreas Spanias |
ICASSP | 2 |
| 1996 | A MATLAB software tool for the introduction of speech coding fundamentals in a DSP courseabstractA MATLAB educational simulation program was developed to explore and understand standardized speech coding algorithms such as the FS-1015 LPC-10e, the FS-1016 CELP, the ETSI GSM, the IS-54 VSELP, the G.721 ADPCM, and the G.728 LD-CELP algorithms. The simulation software provides an interactive environment that allows users to display time- and frequency-domain representations of input and reconstructed speech on an IBM PC compatible equipped with a standard sound card. Scores associated with subjective and objective performance measures are computed using simulation tools. This tool is used in undergraduate and graduate DSP courses to introduce speech coding fundamentals. Edward M. Painter, Andreas Spanias |
ICASSP | 2 |
| 1996 | High-performance alphabet recognitionabstractAlphabet recognition is needed in many applications for retrieving information associated with the spelling of a name, such as telephone numbers, addresses, etc. This is a difficult recognition task due to the acoustic similarities existing between letters in the alphabet (e.g., the E-set letters). This paper presents the development of a high-performance alphabet recognizer that has been evaluated on studio quality as well as on telephone-bandwidth speech. Unlike previously proposed systems, the alphabet recognizer presented is based on context-dependent phoneme hidden Markov models (HMMs), which have been found to outperform whole-word models by as much as 8%. The proposed recognizer incorporates a series of new approaches to tackle the problems associated with the confusions occurring between the stop consonants in the E-set and the confusions between the nasals (i.e., letters M and N). First, a new feature representation is proposed for improved stop consonant discrimination, and second, two subspace approaches are proposed for improved nasal discrimination. The subspace approach was found to yield a 45% error-rate reduction in nasal discrimination. Various other techniques are also proposed, yielding a 97.3% speaker-independent performance on alphabet recognition and 95% speaker-independent performance on E-set recognition, A telephone alphabet recognizer was also developed using context-dependent HMMs. When tested on the recognition of 300 last names (which are contained in a list of 50,000 common last names) spelled by 300 speakers, the recognizer achieved 91.7% correct letter recognition with 1.1% letter insertions. Philipos C. Loizou, Andreas Spanias |
IEEE Trans. Speech Audio Process. | 2 |
| 1994 | High Sample Rate Architectures for Block Adaptive FiltersabstractIn this paper we propose a variety of architectures for implementing block adaptive filters in the time-domain. These filters are based on a block implementation of the least mean squares (BLMS) algorithm. First, we present an architecture which directly maps the BLMS algorithm into an array of processors. Next, we describe an architecture where the weight vector is updated without explicitly computing the filter error. Third, we describe an architecture which exploits the redundant computations of overlapping windows. All the architectures have a significantly smaller sample period compared to frequency domain implementations. Moreover, the sample periods can be reduced even further by applying relaxed look-ahead techniques.> Srikanth Karkada, Chaitali Chakrabarti, Andreas Spanias |
ISCAS | 3 |
| 1994 | Context-Dependent Modeling in Alphabet RecognitionabstractAlphabet recognition is known to be a difficult task due to the acoustic similarities among different letters, especially letters in the E-set. Recognition systems based on whole-word Hidden Markov Models (HMM) perform poorly on this task due to the inability of the models to capture fine phonetic details, especially details occurring within segments of short duration. Letters B and D, for example, differ mainly in the 10-20 msec segment prior to vowel onset. In this paper, we use context-dependent phoneme-based HMMs to capture the fine phonetic detail that is required to discriminate such a confusable vocabulary. Our results reveal that context-dependent modeling gives about 9% improvement on speaker-independent performance over whole-word modeling, and an 18% improvement on the E-set. Furthermore, using an improved spectral representation of the stop consonants in the E-set, an additional 6% improvement in the E-set can be achieved. Our best speaker-independent E-set performance over 15 speakers is 90.3%, with overall alphabet recognition of 94.1%.> Philipos C. Loizou, Andreas Spanias |
ISCAS | 2 |
| 1994 | Speech coding: a tutorial reviewabstractThe past decade has witnessed substantial progress towards the application of low-rate speech coders to civilian and military communications as well as computer-related voice applications. Central to this progress has been the development of new speech coders capable of producing high-quality speech at low data rates. Most of these coders incorporate mechanisms to: represent the spectral properties of speech, provide for speech waveform matching, and "optimize" the coder's performance for the human ear. A number of these coders have already been adopted in national and international cellular telephony standards. The objective of this paper is to provide a tutorial overview of speech coding methodologies with emphasis on those algorithms that are part of the recent low-rate standards for cellular communications. Although the emphasis is on the new low-rate coders, we attempt to provide a comprehensive survey by covering some of the traditional methodologies as well. We feel that this approach will not only point out key references but will also provide valuable background to the beginner. The paper starts with a historical perspective and continues with a brief discussion on the speech properties and performance measures. We then proceed with descriptions of waveform coders, sinusoidal transform coders, linear predictive vocoders, and analysis-by-synthesis linear predictive coders. Finally, we present concluding remarks followed by a discussion of opportunities for future research.> Andreas Spanias |
Proc. IEEE | 1 |
| 1993 | Speech enhancement using the bispectrum
Ralph Fulchiero, Andreas Spanias |
ICASSP (4) | 2 |
| 1993 | Speech processing using higher order statistics
Ines Jebali Gdoura, Philipos C. Loizou, Andreas Spanias |
ISCAS | 3 |
| 1992 | Block time and frequency domain modified covariance algorithmsabstractA block modified covariance algorithm (BMCA) is proposed for autoregressive (AR) parametric spectral estimation. The BMCA can be implemented either in the time or in the frequency domain with the latter being more efficient in high-order AR estimation. Although the BMCA operates on a block of data it allows for both block and sequential (sample-by-sample) updates. The performance of the algorithm is evaluated in terms of computational complexity and convergence characteristics. Results are given for narrowband and broadband processes.> Andreas Spanias, Gim Lim, Philipos C. Loizou, Michael E. Deisher |
ICASSP | 1 |
| 1991 | Real-time implementation of a frequency-domain adaptive filter on a fixed-point signal processorabstractThe authors study some problems associated with the implementation of a frequency-domain adaptive algorithm on a fixed-point signal processor. In particular, they propose methods to improve the convergence speed and reduce the computational complexity of a constrained frequency-domain algorithm that uses a time-varying step size. In addition, they study the effects of finite word length and fixed-point arithmetic. Improvements are realized by adopting a novel data reusing scheme and by applying running and pruned FFTs (fast Fourier transforms). Results are given using synthetic data as well as data from noise cancellation experiments.> Michael E. Deisher, Andreas Spanias |
ICASSP | 2 |
| 1991 | Vector quantization of transform components for speech coding at 1200 bpsabstractA transform-based system is described for vector quantizing speech at and below 1200 b/s. A fixed-dimension codebook is used to quantize low-resolution harmonic components of speech. To reduce the errors of the low-resolution process, transform components, within the baseband and adjacent to the pitch-harmonic components, are also encoded. Speech produced at 1200 b/s using this method was found to be intelligible, particularly for female speakers. This is because the number of harmonic components required for female (high-pitch) speakers is generally smaller than the number required for male (low-pitch) speakers.> Philipos C. Loizou, Andreas Spanias |
ICASSP | 2 |
| 1991 | A hybrid transform method for analysis/synthesis of speech
Andreas Spanias |
Signal Process. | 1 |
| 1991 | Transform methods for seismic data compressionabstractThe authors consider the development and evaluation of transform coding algorithms for the storage of seismic signals. Transform coding algorithms are developed using the discrete Fourier transform (DFT), the discrete cosine transform (DCT), the Walsh-Hadamard transform (WHT), and the Karhunen-Loeve transform (KLT). These are evaluated and compared to a linear predictive coding algorithm for data rates ranging from 150 to 550 bit/s. The results reveal that sinusoidal transforms are well-suited for robust, low-rate seismic signal representation. In particular, it is shown that a DCT coding scheme reproduces faithfully the seismic waveform at approximately one-third of the original rate.> Andreas Spanias, Stefan B. Jonsson, Samuel D. Stearns |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 1988 | A fast frequency-domain adaptive algorithmabstractA time-varying convergence factor mu is proposed for the frequency-domain LMS (least-mean-square) adaptive algorithm, which results in the frequency-domain optimal block algorithm (FOBA). The FOBA is the frequency-domain implementation of the recently proposed optimum block algorithm (OBA). The FOBA results in computational savings in comparison to the OBA and in performance enhancement relative to the frequency-domain LMS algorithm.> Wasfy B. Mikhael, Andreas Spanias |
Proc. IEEE | 2 |
| 1985 | A two stage approach for adaptive prediction of ARMA processesabstractA two-stage predictor is proposed. It is capable of predicting ARMA processes accurately with a reduced number of parameters. This is achieved by cascading a classical linear predictor, where the constraint on the predictor order is relaxed, and a pole-zero recursive like structure. Adaptive algorithms and sample results are given to demonstrate the excellent performance of the proposed technique. Wasfy B. Mikhael, Andreas Spanias, Frank H. Wu |
ICASSP | 2 |