VLDB 2026 Research / reviewers in the wild / expert
Tuka Al Hanai
dblp:207/8063 · also Tuka Alhanai, Tuka Waddah AlHanai
· DBLP profile ↗
17ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-1591-3908ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural AdjustmentsabstractLarge Language Models (LLMs) have shown remarkable performance across various tasks, yet significant disparities remain for non-English languages, and especially native African languages. This paper addresses these disparities by creating approximately 1 million human-translated words of new benchmark data in 8 low-resource African languages, covering a population of over 160 million speakers of: Amharic, Bambara, Igbo, Sepedi (Northern Sotho), Shona, Sesotho (Southern Sotho), Setswana, and Tsonga. Our benchmarks are translations of Winogrande and three sections of MMLU: college medicine, clinical knowledge, and virology. Using the translated benchmarks, we report previously unknown performance gaps between state-of-the-art (SOTA) LLMs in English and African languages. Finally, using results from over 400 fine-tuned models, we explore several methods to reduce the LLM performance gap, including high-quality dataset fine-tuning (using an LLM-as-an-Annotator), cross-lingual transfer, and cultural appropriateness adjustments. Key findings include average mono-lingual improvements of 5.6% with fine-tuning (with 5.4% average mono-lingual improvements when using high-quality data over low-quality data), 2.9% average gains from cross-lingual transfer, and a 3.0% out-of-the-box performance boost on culturally appropriate questions. The publicly available benchmarks, translations, and code from this study support further research and development aimed at creating more inclusive and effective language technologies. Tuka Al Hanai, Adam Kasumovic, Mohammad M. Ghassemi, Aven Zitzelberger, Jessica M. Lundin, Guillaume Chabot-Couture |
AAAI | 1 |
| 2025 | Distribution-Free Uncertainty Quantification in Mechanical Ventilation Treatment: A Conformal Deep Q-Learning FrameworkabstractMechanical Ventilation (MV) is a critical life-support intervention in intensive care units (ICUs). However, optimal ventilator settings are challenging to determine because of the complexity of balancing patient-specific physiological needs with the risks of adverse outcomes that impact morbidity, mortality, and healthcare costs. This study introduces ConformalDQN, a novel distribution-free conformal deep Q-learning approach for optimizing mechanical ventilation in intensive care units. By integrating conformal prediction with deep reinforcement learning, our method provides reliable uncertainty quantification, addressing the challenges of Q-value overestimation and out-of-distribution actions in offline settings. We trained and evaluated our model using ICU patient records from the MIMIC-IV database. ConformalDQN extends the Double DQN architecture with a conformal predictor and employs a composite loss function that balances Q-learning with well-calibrated probability estimation. This enables uncertainty-aware action selection, allowing the model to avoid potentially harmful actions in unfamiliar states and handle distribution shifts by being more conservative in out-of-distribution scenarios. Evaluation against baseline models, including physician policies, policy constraint methods, and behavior cloning, demonstrates that ConformalDQN consistently makes recommendations within clinically safe and relevant ranges, outperforming other methods by increasing the 90-day survival rate. Notably, our approach provides an interpretable measure of confidence in its decisions, which is crucial for clinical adoption and potential human-in-the-loop implementations. Niloufar Eghbali, Tuka Al Hanai, Mohammad M. Ghassemi |
AAAI | 2 |
| 2025 | An LSTM Feature Imitation Network for Hand Movement Recognition from sEMG SignalsabstractSurface Electromyography (sEMG) is a non-invasive signal that is used in the recognition of hand movement patterns, the diagnosis of diseases, and the robust control of prostheses. Despite the remarkable success of recent end-to-end Deep Learning approaches, they are still limited by the need for large amounts of labeled data. To alleviate the requirement for big data, we propose utilizing a feature-imitating network (FIN) for closed-form temporal feature learning over a 300ms signal window on Ninapro DB2, and applying it to the task of 17 hand movement recognition. We implement a lightweight LSTM-FIN network to imitate four standard temporal features (entropy, root mean square, variance, simple square integral). We observed that the LSTM-FIN network can achieve up to 99% R2 accuracy in feature reconstruction and 80% accuracy in hand movement recognition. Our results also showed that the model can be robustly applied for both within- and cross-subject movement recognition, as well as simulated low-latency environments. Overall, our work demonstrates the potential of the FIN modeling paradigm in data-scarce scenarios for sEMG signal processing. Chuheng Wu, Seyed Farokh Atashzar, Mohammad M. Ghassemi, Tuka Al Hanai |
ICASSP | 4 |
| 2025 | GLoG-CSUnet: Enhancing Vision Transformers with Adaptable Radiomic Features for Medical Image SegmentationabstractVision Transformers (ViTs) have shown promise in medical image semantic segmentation (MISS) by capturing longrange correlations. However, ViTs often struggle to model local spatial information effectively, which is essential for accurately segmenting fine anatomical details, particularly when applied to small datasets without extensive pre-training. We introduce Gabor and Laplacian of Gaussian Convolutional Swin Network (GLoG-CSUnet), a novel architecture enhancing Transformer-based models by incorporating learnable radiomic features. This approach integrates dynamically adaptive Gabor and Laplacian of Gaussian (LoG) filters to capture texture, edge, and boundary information, enhancing the feature representation processed by the Transformer model. Our method uniquely combines the longrange dependency modeling of Transformers with the texture analysis capabilities of Gabor and LoG features. Evaluated on the Synapse multi-organ and ACDC cardiac segmentation datasets, GLoG-CSUnet demonstrates significant improvements over stateof-the-art models, achieving a 1.14% increase in Dice score for Synapse and 0.99% for ACDC, with minimal computational overhead (only 15 and 30 additional parameters, respectively). GLoG-CSUnet’s flexible design allows integration with various base models, offering a promising approach for incorporating radiomics-inspired feature extraction in Transformer architectures for medical image analysis. The code implementation is available on GitHub at: https://github.com/HAAIL/GLoGCSUnet. Niloufar Eghbali, Hassan Bagher-Ebadian, Tuka Al Hanai, Mohammad M. Ghassemi |
ICASSP | 3 |
| 2025 | Investigating the Temporal Association of Biomedical Research on Small Business Funding: A Bibliometric and Data Analytic ApproachabstractThe relationship between scientific innovation in biomedical sciences and its impact on industrial activities is a complex and dynamic process. This article investigates the relationship between science and industrial innovation, focusing on how the historical impact and content of scientific paper abstracts are associated with future funding and innovation grant application content for small businesses. The research incorporates bibliometric analyses along with small business innovation research (SBIR) data to yield a holistic view of the science-industry interface. We quantify the temporal effects and impact latency of scientific advancements on industrial activity across 10873 topics and take into account their taxonomic relationships, spanning from 2010 to 2021. We find that the impact of scientific advances on industrial projects across different thematic depths consistently exhibitedp-values less than 0.05, underscoring the significant predictive power of contemporary scientific activities on future industrial projects. Further, we demonstrate that the semantic contents of scientific paper abstracts within a topic are associated with future industrial project description text embeddings. The frequency analysis reveals that various scientific activities significantly inform future industrial project funding across varying depths of MeSH topic categorization, highlighting the significant role of science in steering industrial innovation. This study demonstrates that the impact of scientific research on industrial innovation extends beyond the mere volume of scientific output, but is greatly influenced by its impact, the broader themes it advances, and the meaningful narratives it presents. Reza Khanmohammadi, Simerjot Kaur, Charese Smiley, Tuka Al Hanai, Ivan Brugere, Armineh Nourbakhsh, Mohammad M. Ghassemi |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2022 | Feature Imitating NetworksabstractWe introduce a novel approach to neural learning: the Feature-Imitating-Network (FIN). A FIN is a neural network with weights that are initialized to reliably approximate one or more closed-form statistical features, such as Shannon’s entropy. In this paper, we demonstrate that FINs (and FIN ensembles) provide best-in-class performance for a variety of downstream signal processing and inference tasks, while using less data and requiring less fine-tuning compared to other networks of similar (or even greater) representational power. We conclude that FINs can help bridge the gap between domain experts and machine learning practitioners by enabling re-searchers to harness insights from feature-engineering to enhance the performance of contemporary representation learning approaches. Sari Saba-Sadiya, Tuka Al Hanai, Mohammad M. Ghassemi |
ICASSP | 2 |
| 2022 | Speak: A Toolkit Using Amazon Mechanical Turk to Collect and Validate Speech Audio RecordingsabstractWe present Speak, a toolkit that allows researchers to crowdsource speech audio recordings using Amazon Mechanical Turk (MTurk). Speak allows MTurk workers to submit speech recordings in response to a task prompt and stimulus (e.g. image, text excerpt, audio file) defined by researchers, a functionality that is not natively offered by MTurk at the time of writing this paper. Importantly, the toolkit employs numerous measures to ensure that speech recordings collected are of adequate quality, in order to avoid accepting unusable data and prevent abuse/fraud. Speak has demonstrated utility, having collected over 600,000 recordings to date. The toolkit is open-source and available for download. Christopher Song, David F. Harwath, Tuka Al Hanai, James R. Glass |
LREC | 3 |
| 2021 | Modeling Simultaneous Preferences for Age, Gender, Race, and Professional Profiles in Government-Expense Spending: A Conjoint AnalysisabstractBias can have devastating outcomes on everyday life, and may manifest in subtle preferences for particular attributes (age, gender, ethnicity, profession). Understanding bias is complex, but first requires identifying the variety and interplay of individual preferences. In this study, we deployed a sociotechnical, web-based human-subject experiment to quantify individual preferences in the context of selecting an advisor to successfully pitch a government-expense. We utilized conjoint analysis to rank the preferences of 722 U.S. based subjects, and observed that their ideal advisor was White, middle-aged, and of either a government or STEM-related profession (0.68 AUROC, p < 0.05). The results motivate the simultaneous measurement of preferences as a strategy to offset preferences that may yield negative consequences (e.g. prejudice, disenfranchisement) in contexts where social interests are being represented. Lujain Ibrahim, Mohammad M. Ghassemi, Tuka Al Hanai |
HCOMP | 3 |
| 2020 | EEG Channel Interpolation Using Deep Encoder-decoder NetworksabstractElectrode “pop” artifacts originate from the spontaneous loss of connectivity between a surface and an electrode. Electroencephalography (EEG) uses a dense array of electrodes, hence popped segments are among the most pervasive type of artifact seen during the collection of EEG data. In many cases, the continuity of EEG data is critical for downstream applications (e.g. brain machine interface) and requires that popped segments be accurately interpolated. In this paper we frame the interpolation problem as a self-learning task using a deep encoder-decoder network. We compare our approach against contemporary interpolation methods on a publicly available EEG data set. Our approach exhibited a minimum of ~15% improvement over contemporary approaches when tested on subjects and tasks not used during model training. We demonstrate how our model's performance can be enhanced further on novel subjects and tasks using transfer learning. All code and data associated with this study is open-source to enable ease of extension and practical use. To our knowledge, this work is the first solution to the EEG interpolation problem that uses deep learning. Sari Saba-Sadiya, Tuka Al Hanai, Taosheng Liu, Mohammad M. Ghassemi |
BIBM | 2 |
| 2018 | Detecting Depression with Audio/Text Sequence Modeling of Interviews
Tuka Al Hanai, Mohammad M. Ghassemi, James R. Glass |
INTERSPEECH | 1 |
| 2017 | Predicting Latent Narrative Mood Using Audio and Physiologic DataabstractInferring the latent emotive content of a narrative requires consideration of para-linguistic cues (e.g. pitch), linguistic content (e.g. vocabulary) and the physiological state of the narrator (e.g. heart-rate). In this study we utilized a combination of auditory, text, and physiological signals to predict the mood (happy or sad) of 31 narrations from subjects engaged in personal story-telling. We extracted 386 audio and 222 physiological features (using the Samsung Simband) from the data. A subset of 4 audio, 1 text, and 5 physiologic features were identified using Sequential Forward Selection (SFS) for inclusion in a Neural Network (NN). These features included subject movement, cardiovascular activity, energy in speech, probability of voicing, and linguistic sentiment (i.e. negative or positive). We explored the effects of introducing our selected features at various layers of the NN and found that the location of these features in the network topology had a significant impact on model performance. To ensure the real-time utility of the model, classification was performed over 5 second intervals. We evaluated our model’s performance using leave-one-subject-out crossvalidation and compared the performance to 20 baseline models and a NN with all features included in the input layer. Tuka Al Hanai, Mohammad M. Ghassemi |
AAAI | 1 |
| 2017 | Spoken language biomarkers for detecting cognitive impairmentabstractIn this study we developed an automated system that evaluates speech and language features from audio recordings of neuropsychological examinations of 92 subjects in the Framingham Heart Study. A total of 265 features were used in an elastic-net regularized binomial logistic regression model to classify the presence of cognitive impairment, and to select the most predictive features. We compared performance with a demographic model from 6,258 subjects in the greater study cohort (0.79 AUC), and found that a system that incorporated both audio and text features performed the best (0.92 AUC), with a True Positive Rate of 29% (at 0% False Positive Rate) and a good model fit (Hosmer-Lemeshow test > 0.05). We also found that decreasing pitch and jitter, shorter segments of speech, and responses phrased as questions were positively associated with cognitive impairment. Tuka Al Hanai, Rhoda Au, James R. Glass |
ASRU | 1 |
| 2017 | An open-source tool for the transcription of paper-spreadsheet data: Code and supplemental materials available online: Https: //github.com/deskool/images to spreadsheetsabstractClinical researchers, historians, educators and field researchers alike still regularly capture data on paper spreadsheets. In the case of health care and education, data will often contain sensitive personal information, further complicating the process of transcribing paper-based archives into digital form. In this work, we describe a tool that utilizes machine learning and crowd intelligence to automatically transcribe images of paper-based spreadsheets into electronic form while protecting sensitive personal information. Our solution consists of four high-level stages: (1) the extraction of cell-level images from the spreadsheet grid, (2) machine recognition of digits within the cells, (3) human transcription of cell contents that the machine was uncertain of and (4) feedback of human transcription results to the machine to improve future classification performance. We test the algorithm on a novel data-set of 135 heterogeneous clinical flow-sheet images collected from the Massachusetts General Hospital (MGH), 2 hand-drawn spreadsheets, one chalk-board drawing, and one printed table. we demonstrate that our algorithm provides a generalized solution for spreadsheet transcription that maintains privacy, is up to 10 times faster and twice as cost effective than existing alternatives. Our work is valuable both as a tool and as a starting point for the development of better algorithms. Mohammad M. Ghassemi, Willow Jarvis, Tuka Al Hanai, Emery N. Brown, Roger G. Mark, M. Brandon Westover |
IEEE BigData | 3 |
| 2017 | QMDIS: QCRI-MIT Advanced Dialect Identification System
Sameer Khurana, Maryam Najafian, Ahmed Ali 0002, Tuka Al Hanai, Yonatan Belinkov, James R. Glass |
INTERSPEECH | 4 |
| 2016 | Development of the MIT ASR system for the 2016 Arabic Multi-genre Broadcast ChallengeabstractThe Arabic language, with over 300 million speakers, has significant diversity and breadth. This proves challenging when building an automated system to understand what is said. This paper describes an Arabic Automatic Speech Recognition system developed on a 1,200 hour speech corpus that was made available for the 2016 Arabic Multi-genre Broadcast (MGB) Challenge. A range of Deep Neural Network (DNN) topologies were modeled including; Feed-forward, Convolutional, Time-Delay, Recurrent Long Short-Term Memory (LSTM), Highway LSTM (H-LSTM), and Grid LSTM (GLSTM). The best performance came from a sequence discriminatively trained G-LSTM neural network. The best overall Word Error Rate (WER) was 18.3% (p <; 0:001) on the development set, after combining hypotheses of 3 and 5 layer sequence discriminatively trained G-LSTM models that had been rescored with a 4-gram language model. Tuka Al Hanai, Wei-Ning Hsu, James R. Glass |
SLT | 1 |
| 2014 | Recent advances in ASR applied to an Arabic transcription system for Al-JazeeraabstractThis paper describes a detailed comparison of several state-of-the-art speech recognition techniques applied to a limited Ara-bic broadcast news dataset. The different approaches were all trained on 50 hours of transcribed audio from the Al-Jazeera news channel. The best results were obtained using i-vector-based speaker adaptation in a training scenario using the Min-imum Phone Error (MPE) criteria combined with sequential Deep Neural Network (DNN) training. We report results for two different types of test data: broadcast news reports, with a best word error rate (WER) of 17.86%, and a broadcast conver-sations with a best WER of 29.85%. The overall WER on this test set is 25.6%. Index Terms: Arabic, ASR system, Kaldi 1. Patrick Cardinal, Ahmed Ali 0002, Najim Dehak, Yu Zhang 0033, Tuka Al Hanai, James R. Glass, Stephan Vogel |
INTERSPEECH | 5 |
| 2014 | Lexical modeling for Arabic ASR: a systematic approachabstractArabic has an ambiguous mapping between words and pronunciations, making it a deep orthographic system. This ambiguity can be resolved through diacritics, which if displayed, would compose 30% of characters in a text. We investigate the different dimensions of lexical modeling, covering diacritics, pronunciation rules, and acoustic based pronunciation modeling. We show the impact of explicitly modeling the different classes of diacritics (short vowels, geminates, nunnations). We further show that a phonetic lexicon, derived by applying simple pronunciation rules to diacritized words, offers the best gains in ASR performance. Finally, deriving pronunciations from acoustics, yields improvements, beyond a canonical lexicon. Index Terms: automatic speech recognition, Arabic, diacritics, pronunciation rules, language model, lexical model, joint sequence model, pronunciation mixture model. Tuka Al Hanai, James R. Glass |
INTERSPEECH | 1 |