Bang Tran

dblp:257/2795 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Exploring Safer Image Sharing: A Vision-Language Approach to Privacy Risk Detection and Protection
abstract
The widespread sharing of images on the Internet poses significant privacy and security concerns, including unauthorized distribution, identity theft, and metadata exploitation. These concerns are exacerbated by the rise of AI-powered platforms capable of analyzing the semantic content of images to extract sensitive information about individuals, locations, and activities, often without consent. Consequently, even seemingly harmless photos can become sources of privacy leakage in today’s interconnected digital ecosystem. In this work, we present IPPA (Image Privacy Protection Assistant), a general-purpose system designed to help identify and mitigate privacy leakage in images prior to sharing with on-line platforms; the system also allows users to reconstruct protected content when needed. Specifically, the IPPA comprises three main components: (1) an image content analysis module that interprets image semantics at the natural language level to extract sensitive keywords and identify potential privacy-related objects; (2) a privacy classification module that confirms privacy content using a fine-tuned image recognition model; and (3) a privacy protection module that leverages a watermark-based embedding technique to obfuscate or eliminate identified objects. Our experiments show that IPPA effectively detects and removes privacy objects of the input image. The semantic similarity between a set of keywords before and after removing privacy objects reduces from 1.0 to 0.32, indicating minimal residual sensitivity. Furthermore, the watermark-based technique employed in the privacy protection module achieves an average bit-accuracy of 93.05% in reconstructing the removed privacy content.
Bang Tran
MASS1
2024 CCPA: cloud-based, self-learning modules for consensus pathway analysis using GO, KEGG and Reactome
abstract
This manuscript describes the development of a resource module that is part of a learning platform named 'NIGMS Sandbox for Cloud-based Learning' (https://github.com/NIGMS/NIGMS-Sandbox). The module delivers learning materials on Cloud-based Consensus Pathway Analysis in an interactive format that uses appropriate cloud resources for data access and analyses. Pathway analysis is important because it allows us to gain insights into biological mechanisms underlying conditions. But the availability of many pathway analysis methods, the requirement of coding skills, and the focus of current tools on only a few species all make it very difficult for biomedical researchers to self-learn and perform pathway analysis efficiently. Furthermore, there is a lack of tools that allow researchers to compare analysis results obtained from different experiments and different analysis methods to find consensus results. To address these challenges, we have designed a cloud-based, self-learning module that provides consensus results among established, state-of-the-art pathway analysis techniques to provide students and researchers with necessary training and example materials. The training module consists of five Jupyter Notebooks that provide complete tutorials for the following tasks: (i) process expression data, (ii) perform differential analysis, visualize and compare the results obtained from four differential analysis methods (limma, t-test, edgeR, DESeq2), (iii) process three pathway databases (GO, KEGG and Reactome), (iv) perform pathway analysis using eight methods (ORA, CAMERA, KS test, Wilcoxon test, FGSEA, GSA, SAFE and PADOG) and (v) combine results of multiple analyses. We also provide examples, source code, explanations and instructional videos for trainees to complete each Jupyter Notebook. The module supports the analysis for many model (e.g. human, mouse, fruit fly, zebra fish) and non-model species. The module is publicly available at https://github.com/NIGMS/Consensus-Pathway-Analysis-in-the-Cloud. This manuscript describes the development of a resource module that is part of a learning platform named ``NIGMS Sandbox for Cloud-based Learning'' https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses.
Van-Dung Pham, Hung Nguyen 0005, Bang Tran, Juli Petereit, Tin Chi Nguyen
Briefings Bioinform.4
2023 Early Detection of Cognitive Decline Using Voice Assistant Commands
abstract
Early detection of Alzheimer's Disease and Related Dementias (ADRD) is critical in treating the progression of the disease. Previous studies have shown that ADRD can be detected and classified using machine learning models trained on samples of spontaneous speech. We propose using Voice-Assistant Systems (VAS), e.g., Amazon Alexa, to monitor and collect data from at-risk adults, and we show that this data can be used to achieve functional accuracy in classifying their cognitive status. In this paper, we develop multiple unique feature sets from VAS data that can be used in the training of machine learning models. We then perform multi-class classification, binary classification, and regression using these features on our dataset of older adults with three varying stages of cognitive decline interacting with VAS. Our results show that the VAS data can be used to classify Dementia (DM), Mild Cognitive Impairment (MCI), and Healthy Control (HC) participants with an accuracy up to 74.7%, and classify between HC and MCI with accuracy up to 62.8%.
Eli Kurtz, Youxiang Zhu, Tiffany M. Driesse, Bang Tran, John A. Batsis, Robert M. Roth, Xiaohui Liang 0002
ICASSP4
2023 Exploiting Relevance of Speech to Sleepiness Detection via Attention Mechanism
abstract
Excessive sleepiness in critical tasks and jobs can lead to adverse outcomes, such as work accidents and car crashes. Detecting and monitoring sleepiness levels can prevent these adverse events from happening. In this paper, we propose an attention-based sleepiness detection method using HuBERT embeddings and eGeMAPS features of human speech. Specifically, we propose an attention-based convolutional neural network (CNN) model that achieves accurate 82.57 % sleepiness detection using HuBERT embeddings plus age and gender as inputs. We also show that the embedded attention layers significantly improve the detection accuracy in different cases of inputs. We then explore the attention weights from the attention layers and observe that the long and semantically-different responses from “Picture description”, “Microphone test”, and “Free speech” tasks are more relevant to sleepiness detection when the model is trained with HuBERT only; the short and semantically-similar responses from “Sustained phonation” and “Diadochokinetic” tasks are more relevant when trained with HuBERT plus age and gender. The attention mechanism enables our model to take all responses as one input, simplifying the data pre-processing and identifying the relevant speech responses to sleepiness detection.
Bang Tran, Youxiang Zhu, James W. Schwoebel, Xiaohui Liang 0002
ICC1
2023 VPASS: Voice Privacy Assistant System for Monitoring In-home Voice Commands
abstract
Voice assistant systems (VAS), such as Google Assistant or Amazon Alexa, provide convenient means for users to interact verbally with online services. VAS is particularly important for users with severe health conditions or motor skills impairment. At the same time, voice commands may contain highly-sensitive information about individuals. Therefore, sharing such data with service providers must be done in a carefully controlled and transparent manner in order to prevent privacy breaches. One important challenge is identifying which voice commands contain sensitive information. Different individuals are likely to have distinct interpretations of what is sensitive and what must be kept private, depending on gender, age, cultural background, etc. Furthermore, even for the same individual, the context in which a command is issued can result in significantly different sensitivity perceptions. We introduce a framework named VPASS that supports the management of personalized privacy requirements for VAS systems. Specifically, we propose mechanisms to quantify two key aspects: the amount of information disclosure and the level of privacy sensitivity that each voice command has. Our mechanisms employ deep transfer learning techniques for processing voice commands and can accurately detect privacy-sensitive commands based on an individual’s prior history of VAS interaction. Finally, VPASS generates monthly reports or immediate privacy alerts based on the privacy policies pre-defined by users.
Bang Tran, Sai Harshavardhan Reddy Kona, Xiaohui Liang 0002, Gabriel Ghinita, Caroline Summerour, John A. Batsis
PST1
2022 Speech Tasks Relevant to Sleepiness Determined With Deep Transfer Learning
abstract
Excessive sleepiness in attention-critical contexts can lead to adverse events, such as car crashes. Detecting and monitoring sleepiness can help prevent these adverse events from happening. In this paper, we use the Voiceome dataset to extract speech from 1,828 participants to develop a deep transfer learning model using Hidden-Unit BERT (HuBERT) speech representations to detect sleepiness from individuals. Speech is an under-utilized source of data in sleep detection, but as speech collection is easy, cost-effective, and non-invasive, it provides a promising resource for sleepiness detection. Two complementary techniques were conducted in order to seek converging evidence regarding the importance of individual speech tasks. Our first technique, masking, evaluated task importance by combining all speech tasks, masking selected responses in the speech, and observing systematic changes in model accuracy. Our second technique, separate training, compared the accuracy of multiple models, each of which used the same architecture, but was trained on a different subset of speech tasks. Our evaluation shows that the best-performing model utilizes the memory recall task and categorical naming task from the Boston Naming Test, which achieved an accuracy of 80.07% (F1-score of 0.85) and 81.13% (F1-score of 0.89), respectively.
Bang Tran, Youxiang Zhu, Xiaohui Liang 0002, James W. Schwoebel, Lindsay A. Warrenburg
ICASSP1
2022 Towards Interpretability of Speech Pause in Dementia Detection Using Adversarial Learning
abstract
Speech pause is an effective biomarker in dementia detection. Recent deep learning models have exploited speech pauses to achieve highly accurate dementia detection, but have not exploited the interpretability of speech pauses, i.e., what and how positions and lengths of speech pauses affect the result of dementia detection. In this paper, we will study the positions and lengths of dementia-sensitive pauses using adversarial learning approaches. Specifically, we first utilize an adversarial attack approach by adding the perturbation to the speech pauses of the testing samples, aiming to reduce the confidence levels of the detection model. Then, we apply an adversarial training approach to evaluate the impact of the perturbation in training samples on the detection model. We examine the interpretability from the perspectives of model accuracy, pause context, and pause length. We found that some pauses are more sensitive to dementia than other pauses from the model's perspective, e.g., speech pauses near to the verb "is". Increasing lengths of sensitive pauses or adding sensitive pauses leads the model inference to Alzheimer's Disease (AD), while decreasing the lengths of sensitive pauses or deleting sensitive pauses leads to non-AD.
Youxiang Zhu, Bang Tran, Xiaohui Liang 0002, John A. Batsis, Robert M. Roth
ICASSP2
2021 Exploiting Physical Presence Sensing to Secure Voice Assistant Systems
abstract
Voice Assistant System (VAS) provides a convenient way for users to interact with smart-home devices via a voice interface. However, it raises unique security issues, including voice replay and injection attacks, where attackers remotely and maliciously control the smart-home devices via a voice interface. In this paper, we consider a typical smart-home scenario in which a VAS device and a compromised speaker device are placed in close physical proximity. The attacker can remotely play malicious voice commands through the speaker device to manipulate the VAS device for malicious purposes. We propose a defense system on the VAS device to secure the VAS device against both voice replay and injection attacks, without any additional devices and without any extra user effort. Specifically, our system aims to collect voice data and wireless data continuously from the VAS device and then extracts the Mel-Cepstral Frequency Coefficients (MFCC) features from voice and wireless data. We consider that both voice and wireless data are affected by the same present users' physical activities, and the correlation can be used to detect the attacks. Finally, our system applies a deep learning model that learns from previous time-series data and analyzes real-time data to infer whether the real-time voice command is generated from a user or the speaker device. We have tested our system in certain real-world smart-home scenarios. Our experiments showed that the proposed system has a probability between 76.4% to 89.1% to successfully detect the voice replay and injection attacks in the considered scenarios.
Bang Tran, Shenhui Pan, Xiaohui Liang 0002, Honggang Zhang 0003
ICC1
2021 A comprehensive survey of regulatory network inference methods using single cell RNA sequencing data
abstract
Gene regulatory network is a complicated set of interactions between genetic materials, which dictates how cells develop in living organisms and react to their surrounding environment. Robust comprehension of these interactions would help explain how cells function as well as predict their reactions to external factors. This knowledge can benefit both developmental biology and clinical research such as drug development or epidemiology research. Recently, the rapid advance of single-cell sequencing technologies, which pushed the limit of transcriptomic profiling to the individual cell level, opens up an entirely new area for regulatory network research. To exploit this new abundant source of data and take advantage of data in single-cell resolution, a number of computational methods have been proposed to uncover the interactions hidden by the averaging process in standard bulk sequencing. In this article, we review 15 such network inference methods developed for single-cell data. We discuss their underlying assumptions, inference techniques, usability, and pros and cons. In an extensive analysis using simulation, we also assess the methods' performance, sensitivity to dropout and time complexity. The main objective of this survey is to assist not only life scientists in selecting suitable methods for their data and analysis purposes but also computational scientists in developing new methods by highlighting outstanding challenges in the field that remain to be addressed in the future development.
Hung Nguyen 0005, Bang Tran, Bahadir Pehlivan, Tin Chi Nguyen
Briefings Bioinform.3
2021 Exploiting peer-to-peer communications for query privacy preservation in voice assistant systems
Bang Tran, Xiaohui Liang 0002
Peer-to-Peer Netw. Appl.1