EDBT 2026 Demo / reviewers in the wild / expert
Jack W. Stokes
dblp:24/6478
· DBLP profile ↗
31ranked-venue papers
7as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 3 since 2021Security and privacy · 6 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Heterogeneous Graph Neural Network on Semantic TreeabstractThe recent past has seen an increasing interest in Heterogeneous Graph Neural Networks (HGNNs), since many real-world graphs are heterogeneous in nature, from citation graphs to email graphs. However, existing methods ignore a tree hierarchy among metapaths, naturally constituted by different node types and relation types. In this paper, we present HetTree, a novel HGNN that models both the graph structure and heterogeneous aspects in a scalable and effective manner. Specifically, HetTree builds a semantic tree data structure to capture the hierarchy among metapaths. To effectively encode the semantic tree, HetTree uses a novel subtree attention mechanism to emphasize metapaths that are more helpful in encoding parent-child relationships. Moreover, HetTree proposes carefully matching pre-computed features and labels correspondingly, constituting a complete metapath representation. Our evaluation of HetTree on a variety of real-world datasets demonstrates that it outperforms all existing baselines on open benchmarks and efficiently scales to large real-world graphs with millions of nodes and edges. Mingyu Guan, Jack W. Stokes, Qinlong Luo, Fuchen Liu, Purvanshi Mehta, Elnaz Nouri, Taesoo Kim |
AAAI | 2 |
| 2024 | Interpretable User Satisfaction Estimation for Conversational Systems with Large Language ModelsabstractYing-Chun Lin, Jennifer Neville, Jack Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, Deepak Gupta, Sujay Kumar Jauhar, Xia Song, Georg Buscher, Saurabh Tiwary, Brent Hecht, Jaime Teevan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ying-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang 0001, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Sujay Kumar Jauhar, Georg Buscher, Saurabh Tiwary, Brent J. Hecht, Jaime Teevan |
ACL (1) | 3 |
| 2021 | Detection Of Malicious DNS and Web Servers using Graph-Based ApproachesabstractThe DNS hijacking attack represents a significant threat to users. In this type of attack, a malicious DNS server redirects a victim domain to an attacker-controlled web server. Existing defenses are not scalable and have not been widely deployed. In this work, we propose both unsupervised and semi-supervised defenses based on the available knowledge of the defender. Specifically, our unsupervised defense is a graph-based detection approach employing a new variant of the community detection algorithm. When the IP addresses of several compromised DNS servers are available, we also propose a semi-supervised defense for the detection of compromised or malicious web servers which host the web content. We evaluate our defenses on a real-world attack. The experimental results show that our defenses can successfully identify these malicious web servers and/or DNS server IPs. Moreover, we find that a deep learning-based algorithm, i.e., node2vec, outperforms one which employs belief propagation. Jinyuan Jia 0001, Jack W. Stokes |
ICASSP | 4 |
| 2021 | AMP: authentication of media via provenanceabstractAdvances in graphics and machine learning have led to the general availability of easy-to-use tools for modifying and synthesizing media. The proliferation of these tools threatens to cast doubt on the veracity of all media. One approach to thwarting the flow of fake media is to detect modified or synthesized media through machine learning methods. While detection may help in the short term, we believe that it is destined to fail as the quality of fake media generation continues to improve. Soon, neither humans nor algorithms will be able to reliably distinguish fake versus real content. Thus, pipelines for assuring the source and integrity of media will be required---and increasingly relied upon. We present AMP, a system that ensures the authentication of media via certifying provenance. AMP creates one or more publisher-signed manifests for a media instance uploaded by a content provider. These manifests are stored in a database allowing fast lookup from applications such as browsers. For reference, the manifests are also registered and signed by a permissioned ledger, implemented using the Confidential Consortium Framework (CCF). CCF employs both software and hardware techniques to ensure the integrity and transparency of all registered manifests. AMP, through its use of CCF, enables a consortium of media providers to govern the service while making all its operations auditable. The authenticity of the media can be communicated to the user via visual elements in the browser, indicating that an AMP manifest has been successfully located and verified. Paul England, Henrique S. Malvar, Eric Horvitz, Jack W. Stokes, Cédric Fournet, Rebecca Burke-Aguero, Amaury Chamayou, Sylvan Clebsch, Manuel Costa, John Deutscher, Shabnam Erfani, Matt Gaylor, Andrew Jenks, Kevin Kane, Elissa M. Redmiles, Alex Shamis, Isha Sharma, John C. Simmons, Sam Wenker, Anika Zaman |
MMSys | 4 |
| 2021 | Living-Off-The-Land Command Detection Using Active LearningabstractIn recent years, enterprises have been targeted by advanced adversaries who leverage creative ways to infiltrate their systems and move laterally to gain access to critical data. One increasingly common evasive method is to hide the malicious activity behind a benign program by using tools that are already installed on user computers. These programs are usually part of the operating system distribution or another user-installed binary, therefore this type of attack is called “Living-Off-The-Land”. Detecting these attacks is challenging, as adversaries may not create malicious files on the victim computers and anti-virus scans fail to detect them. Talha Ongun, Jack W. Stokes, Jonathan Bar Or, Ke Tian, Farid Tajaddodianfar, Joshua Neil, Christian Seifert, Alina Oprea, John C. Platt |
RAID | 2 |
| 2021 | Designing Media Provenance Indicators to Combat Fake MediaabstractWith the growth of technology that produces misinformation, there is a growing need to help users identify emerging types of fake media such as edited images and manipulated videos. In this work, we conduct a mixed-methods investigation into how we can provide provenance indicators to assist users in detecting newer forms of fake media. Specifically, we interview users regarding their experiences with different misinformation modes (text, image, video) to inform the design and content of indicators for previously unexplored media, especially fake videos. We find that media provenance – the source of the information – is a key heuristic used to evaluate all forms of fake media, and a heuristic that can be addressed by emerging technology. Thus, we subsequently design and investigate the use of provenance indicators to help users identify fake videos. We conduct a participatory design study to develop and design provenance indicators and evaluate participant-designed indicators via both expert evaluations and quantitative surveys (n=1,456) with end-users. Our results provide concrete design guidelines for the emerging issue of fake media. Our findings also raise concerns regarding users’ tendency to overgeneralize indicators used to assist users in identifying misinformation, suggesting the need for further research on warning design in the ongoing fight against misinformation. Imani N. S. Munyaka, Jack W. Stokes, Elissa M. Redmiles |
RAID | 2 |
| 2020 | Actor Critic Deep Reinforcement Learning for Neural Malware ControlabstractIn addition to using signatures, antimalware products also detect malicious attacks by evaluating unknown files in an emulated environment, i.e. sandbox, prior to execution on a computer's native operating system. During emulation, a file cannot be scanned indefinitely, and antimalware engines often set the number of instructions to be executed based on a set of heuristics. These heuristics only make the decision of when to halt emulation using partial information leading to the execution of the file for either too many or too few instructions. Also this method is vulnerable if the attackers learn this set of heuristics. Recent research uses a deep reinforcement learning (DRL) model employing a Deep Q-Network (DQN) to learn when to halt the emulation of a file. In this paper, we propose a new DRL-based system which instead employs a modified actor critic (AC) framework for the emulation halting task. This AC model dynamically predicts the best time to halt the file's execution based on a sequence of system API calls. Compared to the earlier models, the new model is capable of handling adversarial attacks by simulating their behaviors using the critic model. The new AC model demonstrates much better performance than both the DQN model and antimalware engine's heuristics. In terms of execution speed (evaluated by the halting decision), the new model halts the execution of unknown files by up to 2.5% earlier than the DQN model and 93.6% earlier than the heuristics. For the task of detecting malicious files, the proposed AC model increases the true positive rate by 9.9% from 69.5% to 76.4% at a false positive rate of 1% compared to the DQN model, and by 83.4% from 41.2% to 76.4% at a false positive rate of 1% compared to a recently proposed LSTM model. Yu Wang 0091, Jack W. Stokes, Mady Marinescu |
AAAI | 2 |
| 2020 | Privacy-Preserving Phishing Web Page Classification Via Fully Homomorphic EncryptionabstractThis work introduces a fast and lightweight homomorphic-encryption pipeline that enables privacy-preserving machine learning for phishing web page recognition. The primary goals are to use visual features to train an accurate model and to implement an inference pipeline with practical runtime and communication costs. To do so, we deploy a variety of techniques that cover deep learning and optical character recognition to extract salient visual features, and optimize the inner mechanisms of state-of-the-art homomorphic encryption schemes to reduce the encryption-related costs. Our presented system is able to achieve over 90% on the visual classification task, while using less than 250 KB of communication bandwidth and around 0.7 seconds of computation time. We hope our work not only demonstrates a private visual phishing detection pipeline, but also outlines techniques to practically utilize homomorphic encryption in a variety of machine learning tasks. Edward J. Chou, Arun Gururajan, Kim Laine, Nitin Kumar Goel, Anna Bertiger, Jack W. Stokes |
ICASSP | 6 |
| 2020 | Detection of Malicious Vbscript Using Static and Dynamic Analysis with Recurrent Deep LearningabstractAttackers have used malicious VBScripts as an important computer infection vector. In this study, we explore a system that employs both static and dynamic analysis to detect malicious VBScripts. For the static analysis, we investigate two deep recurrent models, LaMP (LSTM and Max Pooling) and CPoLS (Convoluted Partitioning of Long Sequences), which process a VBScript as a byte sequence. Lower layers capture the sequential nature of these byte sequences while higher layers classify the resulting embedding as malicious or benign. Our models are trained in an end-to-end fashion allowing discriminative training even for the sequential processing layers. Dynamic analysis allows us to investigate obfuscated VBScripts an additional files which may be dropped during execution. Evaluating these models on a large corpus of 240,504 VBScript files indicates that the best performing LaMP model has a 69.3% true positive rate (TPR) at a false positive rate (FPR) of 1.0%. Similarly, the best CPoLS model has a TPR of 67.9% at an FPR of 1.0%. Our system is general in nature and can be applied to other scripting languages (e.g., JavaScript) as well. Jack W. Stokes, Rakshit Agrawal, Geoff McDonald |
ICASSP | 1 |
| 2020 | Texception: A Character/Word-Level Deep Learning Model for Phishing URL DetectionabstractPhishing is the starting point for many cyberattacks that threaten the confidentiality, availability and integrity of enterprises' and consumers' data. The URL of a web page that hosts the attack provides a rich source of information to determine the maliciousness of the web server. In this work, we propose a novel deep learning architecture, Texception, that takes a URL as input and predicts whether it belongs to a phishing attack. Architecturally, Texception uses both character-level and word-level information from the incoming URL and does not depend on manually crafted features or feature engineering. This makes it different from classical approaches. In addition, Texception benefits from multiple parallel convolutional layers and can grow deeper or wider. We show that this flexibility enables Texception to generalize better for new URLs. Our results on production data show that Texception is able to significantly outperform a traditional text classification method by increasing the true positive rate by 126.7% at an extremely low false positive rate (0.01%) which is crucial for our model's healthy operation at internet scale. Farid Tajaddodianfar, Jack W. Stokes, Arun Gururajan |
ICASSP | 2 |
| 2020 | PrivateEye: Scalable and Privacy-Preserving Compromise Detection in the Cloud
Behnaz Arzani, Selim Ciraci, Stefan Saroiu, Alec Wolman, Jack W. Stokes, Geoff Outhred, Lechao Diwu |
NSDI | 5 |
| 2019 | Attention in Recurrent Neural Networks for Ransomware DetectionabstractRansomware, as a specialized form of malicious software, has recently emerged as a major threat in computer security. With an ability to lock out user access to their content, recent ransomware attacks have caused severe impact at an individual and organizational level. While research in malware detection can be adapted directly for ransomware, specific structural properties of ransomware can further improve the quality of detection. In this paper, we adapt the deep learning methods used in malware detection for detecting ransomware from emulation sequences. We present specialized recurrent neural networks for capturing local event patterns in ransomware sequences using the concept of attention mechanisms. We demonstrate the performance of enhanced LSTM models on a sequence dataset derived by the emulation of ransomware executables targeting the Windows environment. Rakshit Agrawal, Jack W. Stokes, Karthik Selvaraj, Mady Marinescu |
ICASSP | 2 |
| 2019 | Detecting Cyber Attacks Using Anomaly Detection with Explanations and Expert FeedbackabstractDetecting cyber attacks in large computer networks is crucial for many organizations. To that purpose, different types of detectors capture the important signals resembling a security attack from individual computers and bring that to the attention of a security analyst. Unfortunately, the analyst sometimes has no indications about why the particular computer was identified as being "under attack". In addition, the analyst may have no method to provide feedback to the detector if the computer was actually identified for some benign reason. In this paper, we use a state-of-the-art anomaly detector called an Isolation Forest [1] for attack detection and generate explanations about why the detector identified certain computers as anomalous. These explanations allow the analyst to direct their investigation in order to save time. We then take the feedback from the analyst in the form of true and false positives and update the anomaly detector to capture signals that align better with the given feedback. Our experiments on actual network data show that the explanations give more insight into the detections, and the analyst's feedback increases the attack detection rate. Md Amran Siddiqui, Jack W. Stokes, Christian Seifert, Evan Argyle, Robert McCann, Joshua Neil, Justin Carroll |
ICASSP | 2 |
| 2018 | Neural Sequential Malware Detection with ParametersabstractSequential models which analyze system API calls have shown promise for detecting unknown malware. Athiwaratkun and Stokes recently proposed a two-stage model which uses a long short-term memory (LSTM) model for learning a set of features which are then input to a second classifier. Kolosnjaji et al., first use a convolutional neural network followed by an LSTM to predict unknown malware. However, neither of these models consider the parameters which are input to the system API calls. These input parameters offer significant information regarding malicious intent. In this paper, we extend Athiwaratkun's model to include each system API call's two most input parameters. We then show that the proposed model dominates these previously proposed models in terms of the receiver operating characteristic (ROC) curve. Rakshit Agrawal, Jack W. Stokes, Mady Marinescu, Karthik Selvaraj |
ICASSP | 2 |
| 2017 | Malware classification with LSTM and GRU language models and a character-level CNNabstractMalicious software, or malware, continues to be a problem for computer users, corporations, and governments. Previous research [1] has explored training file-based, malware classifiers using a two-stage approach. In the first stage, a malware language model is used to learn the feature representation which is then input to a second stage malware classifier. In Pascanu et al. [1], the language model is either a standard recurrent neural network (RNN) or an echo state network (ESN). In this work, we propose several new malware classification architectures which include a long short-term memory (LSTM) language model and a gated recurrent unit (GRU) language model. We also propose using an attention mechanism similar to [12] from the machine translation literature, in addition to temporal max pooling used in [1], as an alternative way to construct the file representation from neural features. Finally, we propose a new single-stage malware classifier based on a character-level convolutional neural network (CNN). Results show that the LSTM with temporal max pooling and logistic regression offers a 31.3% improvement in the true positive rate compared to the best system in [1] at a false positive rate of 1%. Ben Athiwaratkun, Jack W. Stokes |
ICASSP | 2 |
| 2016 | MtNet: A Multi-Task Neural Network for Dynamic Malware Classification
Wenyi Huang, Jack W. Stokes |
DIMVA | 2 |
| 2016 | Asking for a second opinion: Re-querying of noisy multi-class labelsabstractIn this paper, we propose a new maximum margin-based, active learning algorithm for identifying incorrectly labeled training data. The algorithm combines a round-robin approach for investigating each class with a simple, yet effective ranking metric called maximum negative margin (MNM). Samples are given to an expert for re-evaluation to determine if they are indeed mislabeled. We also propose using five active learning metrics, including uncertainty sampling with margin sampling (USMS) and minimum margin, for the noisy label task which have previously been used in the standard active learning setting for identifying new samples to label. USMS is very competitive with maximum negative margin. In addition, we consider other information theoretic objective criteria for this new task including uncertainty sampling with entropy, query-by-committee with voting entropy, and K-nearest neighbor with voting entropy, but these consistently perform worse than MNM and USMS. The MNM noisy label active learning algorithm can be useful in several different scenarios including data cleansing as a preprocessing step before training and identifying mislabeled examples in the test set. Jack W. Stokes, Ashish Kapoor, Debajyoti Ray |
ICASSP | 1 |
| 2015 | Malware classification with recurrent networksabstractAttackers often create systems that automatically rewrite and reorder their malware to avoid detection. Typical machine learning approaches, which learn a classifier based on a handcrafted feature vector, are not sufficiently robust to such reorderings. We propose a different approach, which, similar to natural language modeling, learns the language of malware spoken through the executed instructions and extracts robust, time domain features. Echo state networks (ESNs) and recurrent neural networks (RNNs) are used for the projection stage that extracts the features. These models are trained in an unsupervised fashion. A standard classifier uses these features to detect malicious files. We explore a few variants of ESNs and RNNs for the projection stage, including Max-Pooling and Half-Frame models which we propose. The best performing hybrid model uses an ESN for the recurrent model, Max-Pooling for non-linear sampling, and logistic regression for the final classification. Compared to the standard trigram of events model, it improves the true positive rate by 98.3% at a false positive rate of 0.1%. Razvan Pascanu, Jack W. Stokes, Hermineh Sanossian, Mady Marinescu, Anil Thomas |
ICASSP | 2 |
| 2013 | Detecting malicious landing pages in Malware Distribution NetworksabstractDrive-by download attacks attempt to compromise a victim's computer through browser vulnerabilities. Often they are launched from Malware Distribution Networks (MDNs) consisting of landing pages to attract traffic, intermediate redirection servers, and exploit servers which attempt the compromise. In this paper, we present a novel approach to discovering the landing pages that lead to drive-by downloads. Starting from partial knowledge of a given collection of MDNs we identify the malicious content on their landing pages using multiclass feature selection. We then query the webpage cache of a commercial search engine to identify landing pages containing the same or similar content. In this way we are able to identify previously unknown landing pages belonging to already identified MDNs, which allows us to expand our understanding of the MDN. We explore using both a rule-based and classifier approach to identifying potentially malicious landing pages. We build both systems and independently verify using a high-interaction honeypot that the newly identified landing pages indeed attempt drive-by downloads. For the rule-based system 57% of the landing pages predicted as malicious are confirmed, and this success rate remains constant in two large trials spaced five months apart. This extends the known footprint of the MDNs studied by 17%. The classifier-based system is less successful, and we explore possible reasons. Gang Wang 0011, Jack W. Stokes, Cormac Herley, David Felstead |
DSN | 2 |
| 2013 | Large-scale malware classification using random projections and neural networksabstractAutomatically generated malware is a significant problem for computer users. Analysts are able to manually investigate a small number of unknown files, but the best large-scale defense for detecting malware is automated malware classification. Malware classifiers often use sparse binary features, and the number of potential features can be on the order of tens or hundreds of millions. Feature selection reduces the number of features to a manageable number for training simpler algorithms such as logistic regression, but this number is still too large for more complex algorithms such as neural networks. To overcome this problem, we used random projections to further reduce the dimensionality of the original input space. Using this architecture, we train several very large-scale neural network systems with over 2.6 million labeled samples thereby achieving classification results with a two-class error rate of 0.49% for a single neural network and 0.42% for an ensemble of neural networks. George E. Dahl, Jack W. Stokes, Li Deng 0001, Dong Yu 0001 |
ICASSP | 2 |
| 2013 | Robust scareware image detectionabstractIn this paper, we propose an image-based detection method to identify web-based scareware attacks that is robust to evasion techniques. We evaluate the method on a large-scale data set that resulted in an equal error rate of 0.018%. Conceptually, false positives may occur when a visual element, such as a red shield, is embedded in a benign page. We suggest including additional orthogonal features or employing graders to mitigate this risk. A novel visualization technique is presented demonstrating the acquired classifier knowledge on a classified screenshot. Christian Seifert, Jack W. Stokes, Christina Colcernian, John C. Platt, Long Lu |
ICASSP | 2 |
| 2012 | Using File Relationships in Malware Classification
Nikos Karampatziakis, Jack W. Stokes, Anil Thomas, Mady Marinescu |
DIMVA | 2 |
| 2012 | Scalable Telemetry Classification for Automated Malware Detection
Jack W. Stokes, John C. Platt, Helen J. Wang, Joe Faulhaber, Jonathan Keller, Mady Marinescu, Anil Thomas, Marius Gheorghescu |
ESORICS | 1 |
| 2011 | ARROW: GenerAting SignatuRes to Detect DRive-By DOWnloadsabstractA drive-by download attack occurs when a user visits a webpage which attempts to automatically download malware without the user's consent. Attackers sometimes use a malware distribution network (MDN) to manage a large number of malicious webpages, exploits, and malware executables. In this paper, we provide a new method to determine these MDNs from the secondary URLs and redirect chains recorded by a high-interaction client honeypot. In addition, we propose a novel drive-by download detection method. Instead of depending on the malicious content used by previous methods, our algorithm first identifies and then leverages the URLs of the MDN's central servers, where a central server is a common server shared by a large percentage of the drive-by download attacks in the same MDN. A set of regular expression-based signatures are then generated based on the URLs of each central server. This method allows additional malicious webpages to be identified which launched but failed to execute a successful drive-by download attack. The new drive-by detection system named ARROW has been implemented, and we provide a large-scale evaluation on the output of a production drive-by detection system. The experimental results demonstrate the effectiveness of our method, where the detection coverage has been boosted by 96% with an extremely low false positive rate. Junjie Zhang 0004, Christian Seifert, Jack W. Stokes, Wenke Lee |
WWW | 3 |
| 2008 | Nonlinear residual acoustic echo suppression for high levels of harmonic distortionabstractLinear adaptive filters are often used for acoustic echo cancellation (AEC) but sometimes fail to perform well in notebook computers and inexpensive telephony devices. Low-quality speakers and poorly-designed enclosures that produce vibrations often generate harmonic distortion, and this nonlinear effect degrades the performance of linear AEC algorithms considerably. In this work, we present a new AEC architecture that consists of a linear, subband adaptive AEC filter followed a nonlinear residual echo suppression (RES) stage specifically designed to address harmonic distortion. In addition to suppressing the residual echo in the primary subband, the proposed model also suppresses the residual echo in a window of bands surrounding the higher order harmonics. Results show considerable improvement over other proposed algorithms, and the new algorithm has much lower implementation costs compared to nonlinear AEC models based on Volterra filters and a previously proposed, nonlinear residual echo suppression algorithm. Diego A. Bendersky, Jack W. Stokes, Henrique S. Malvar |
ICASSP | 2 |
| 2007 | Normalized Double-Talk Detection Based on Microphone and AEC Error Cross-CorrelationabstractIn this paper, we present two different double-talk detection schemes for Acoustic Echo Cancellation (AEC). First, we present a novel normalized detection statistic based on the cross-correlation coefficient between the microphone signal and the cancellation error. The decision statistic is designed in such a way that it meets the needs of an optimal double-talk detector. We also show that the proposed detection statistic converges to the recently proposed normalized cross-correlation based double-talk detector, the best known cross-correlation based detector. Next, we present a new hybrid double-talk detection scheme based on a cross-correlation coefficient and two signal detectors. The hybrid algorithm not only detects double-talk but also detects and tracks any echo-path variations efficiently. We compare our results with other cross-correlation based double-talk detectors to show their effectiveness. Asif Iqbal Mohammad, Jack W. Stokes, Steven L. Grant |
ICME | 2 |
| 2006 | Robust Rls with Round Robin Regularization Including Application to Stereo Acoustic Echo CancellationabstractThis paper introduces a new algorithm for implementing subband, adaptive filtering using recursive least squares (RLS) with round robin regularization. We show that modern microprocessors with SEMD (single instruction, multiple data) instructions can now implement RLS for practical problems thereby avoiding the numerical stability issues associated with fast RLS (FRLS). The desired signal may be multichannel as in the stereo, acoustic echo cancellation (AEC) problem where the separate channels of the playback signals are often highly correlated. In this case, the recursive computation of the inverse correlation matrix in RLS will diverge. To avoid this problem, we extend adaptive subband RLS to include round robin regularization. The new, regularized RLS (RRLS) algorithm has been implemented in real-time on a personal computer (PC) for the stereo AEC problem and performs well in typical PC scenarios Jack W. Stokes, John C. Platt |
ICASSP (3) | 1 |
| 2006 | Acoustic Echo Cancelation for High Noise EnvironmentsabstractAcoustic echo cancellation (AEC) is highly imperative for enhanced communication in noisy environments such as a car or a conference room. In this work, we present a dual-structured AEC architecture that improves both the convergence time and misadjustment of a conventional adaptive sub-band AEC algorithm in high noise environments. In this architecture, one part performs smooth adaptation while the other part performs fast adaptation; a convergence detector is implemented to facilitate switching between the fast and smooth adaptations. We propose the momentum normalized least mean square (MNLMS) algorithm for smooth adaptation and we implement the NLMS algorithm for fast adaptation. The current architecture provides up to 3-4 dB echo reduction improvement over a conventional adaptive subband AEC algorithm and it helps minimize near-end distortion and artifacts in the post-processed AEC output Amit Chhetri, Jack W. Stokes, Dinei A. F. Florêncio |
ICME | 2 |
| 2006 | Speaker Identification using a Microphone Array and a Joint HMM with Speech Spectrum and Angle of ArrivalabstractIn this paper, we present a speaker identification algorithm for a microphone array based on a first-order joint hidden Markov model (HMM) where the observations correspond to the angle of arrival of the speech and the speech spectrum. The goal of the research is to investigate whether including angle of arrival information improves the speaker identification error rates compared to an algorithm based on the speech spectrum only. The spectral model consists of a Gaussian mixture model (GMM) using multiple discriminant analysis (MDA) coefficients and the angle model includes a separate histogram for each participant. The convergence time of the joint HMM is improved by estimating the GMM for each of the meeting participants prior to the start of the meeting and initializing each participant's spectral GMM in the joint HMM to the pretrained parameter values. The performance of the algorithm is analyzed from data collected during live meetings recorded using an eight element, circular microphone array. For meetings where the participants are stationary, the results show significant improvement over a single channel speaker ID algorithms based on spectrum only Jack W. Stokes, John C. Platt, Sumit Basu |
ICME | 1 |
| 2004 | Acoustic echo cancellation with arbitrary playback sampling rateabstractThis paper introduces a new architecture for implementing subband acoustic echo cancellation (AEC) with arbitrary playback sampling rate. Typically, in AEC algorithms for audio or videoconferencing, the sampling rates for the signals played through the speakers and captured from the microphones are identical. For speech recognition while playing CD-quality music and Internet gaming with voice chat, the playback sampling rate is usually higher than the capture rate. A direct solution is to apply a sampling rate converter to the playback signal before feeding it to the AEC, but that is complicated if many sampling frequencies must be supported. We propose a more efficient solution for subband AEC: we perform the sampling rate conversion as a frequency-domain interpolation that matches the transform lengths of the playback and capture signals. Results show that the new AEC architecture has a small computational cost and only a minimal reduction in echo attenuation. Jack W. Stokes, Henrique S. Malvar |
ICASSP (4) | 1 |
| 2001 | Performance analysis of DS/CDMA systems with shadowing and flat fading
Jack W. Stokes, James A. Ritcey |
Signal Process. | 1 |