Siyi Tang

dblp:184/7801 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-7504-5885ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Evidence-Augmented Generative Explanation for Health Rumor Detection
Siyi Tang, Zhong Qian 0001, Peifeng Li 0001, Qiaoming Zhu
ICIC (23)1
2023 Semi-Supervised Learning for Sparsely-Labeled Sequential Data: Application to Healthcare Video Processing
abstract
Labeled data is a critical resource for training and evaluating machine learning models. However, many real-life datasets are only partially labeled. We propose a semi-supervised machine learning training strategy to improve event detection performance on sequential data, such as video recordings, when only sparse labels are available, such as event start times without their corresponding end times. Our method uses noisy guesses of the events’ end times to train event detection models. Depending on how conservative these guesses are, mislabeled samples may be introduced into the training set. We further propose a mathematical model for explaining and estimating the evolution of the classification performance for increasingly noisier end time estimates. We show that neural networks can improve their detection performance by leveraging more training data with less conservative approximations despite the higher proportion of incorrect labels. We adapt sequential versions of CIFAR-10 and MNIST, and use the Berkeley MHAD and HMBD51 video datasets to empirically evaluate our method, and find that our risk-tolerant strategy outperforms conservative estimates by 3.5 points of mean average precision for CIFAR, 30 points for MNIST, 3 points for MHAD, and 14 points for HMBD51. Then, we leverage the proposed training strategy to tackle a real-life application: processing continuous video recordings of epilepsy patients, and show that our method outperforms baseline labeling methods by 17 points of average precision, and reaches a classification performance similar to that of fully supervised models. We share part of the code for this article at the following repository: fpgdubost/CIFAR-10-Sparsely-Labeled-Sequential-Data.
Florian Dubost, Erin Hong, Siyi Tang, Nandita Bhaskhar, Christopher Lee-Messer, Daniel L. Rubin
WACV3
2023 Predicting 30-Day All-Cause Hospital Readmission Using Multimodal Spatiotemporal Graph Neural Networks
abstract
Reduction in 30-day readmission rate is an important quality factor for hospitals as it can reduce the overall cost of care and improve patient post-discharge outcomes. While deep-learning-based studies have shown promising empirical results, several limitations exist in prior models for hospital readmission prediction, such as: (a) only patients with certain conditions are considered, (b) do not leverage data temporality, (c) individual admissions are assumed independent of each other, which ignores patient similarity, (d) limited to single modality or single center data. In this study, we propose a multimodal, spatiotemporal graph neural network (MM-STGNN) for prediction of 30-day all-cause hospital readmission, which fuses in-patient multimodal, longitudinal data and models patient similarity using a graph. Using longitudinal chest radiographs and electronic health records from two independent centers, we show that MM-STGNN achieved an area under the receiver operating characteristic curve (AUROC) of 0.79 on both datasets. Furthermore, MM-STGNN significantly outperformed the current clinical reference standard, LACE+ (AUROC = 0.61), on the internal dataset. For subset populations of patients with heart disease, our model significantly outperformed baselines, such as gradient-boosting and Long Short-Term Memory models (e.g., AUROC improved by 3.7 points in patients with heart disease). Qualitative interpretability analysis indicated that while patients' primary diagnoses were not explicitly used to train the model, features crucial for model prediction may reflect patients' diagnoses. Our model could be utilized as an additional clinical decision aid during discharge disposition and triaging high-risk patients for closer post-discharge follow-up for potential preventive measures.
Siyi Tang, Amara Tariq, Jared Dunnmon, Praneetha Elugunti, Daniel L. Rubin, Bhavik N. Patel, Imon Banerjee
IEEE J. Biomed. Health Informatics1
2022 Graph-based Fusion Modeling and Explanation for Disease Trajectory Prediction
Amara Tariq, Siyi Tang, Hifza Sakhi, Leo A. Celi, Janice M. Newsome, Daniel L. Rubin, Hari Trivedi, Judy Gichoya, Bhavik N. Patel, Imon Banerjee
AMIA2
2022 Self-Supervised Graph Neural Networks for Improved Electroencephalographic Seizure Analysis
Siyi Tang, Jared Dunnmon, Khaled Saab 0002, Qianying Huang, Florian Dubost, Daniel L. Rubin, Christopher Lee-Messer
ICLR1
2021 Tracking the Evolution of COVID-19 via Temporal Comorbidity Analysis from Multi-Modal Data
Sutanay Choudhury, Khushbu Agarwal, Colby Ham, Pritam Mukherjee, Siyi Tang, Sindhu Tipirneni, Veysel Kocaman, Suzanne Tamang, Robert Rallo, Chandan K. Reddy
AMIA5
2021 Comparison of segmentation-free and segmentation-dependent computer-aided diagnosis of breast masses on a public mammography dataset
abstract
PURPOSE: To compare machine learning methods for classifying mass lesions on mammography images that use predefined image features computed over lesion segmentations to those that leverage segmentation-free representation learning on a standard, public evaluation dataset. METHODS: We apply several classification algorithms to the public Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM), in which each image contains a mass lesion. Segmentation-free representation learning techniques for classifying lesions as benign or malignant include both a Bag-of-Visual-Words (BoVW) method and a Convolutional Neural Network (CNN). We compare classification performance of these techniques to that obtained using two different segmentation-dependent approaches from the literature that rely on specific combinations of end classifiers (e.g. linear discriminant analysis, neural networks) and predefined features computed over the lesion segmentation (e.g. spiculation measure, morphological characteristics, intensity metrics). RESULTS: values of 0.73 for a segmentation-free BoVW method, 0.86 for a segmentation-free CNN method, 0.75 for a segmentation-dependent linear discriminant analysis of Rubber-Band Straightening Transform features, and 0.58 for a hybrid rule-based neural network classification using a small number of hand-designed features. CONCLUSIONS: We find that malignancy classification performance on the CBIS-DDSM dataset using segmentation-free BoVW features is comparable to that of the best segmentation-dependent methods we study, but also observe that a common segmentation-free CNN model substantially and significantly outperforms each of these (p < 0.05). These results reinforce recent findings suggesting that representation learning techniques such as BoVW and CNNs are advantageous for mammogram analysis because they do not require lesion segmentation, the quality and specific characteristics of which can vary substantially across datasets. We further observe that segmentation-dependent methods achieve performance levels on CBIS-DDSM inferior to those achieved on the original evaluation datasets reported in the literature. Each of these findings reinforces the need for standardization of datasets, segmentation techniques, and model implementations in performance assessments of automated classifiers for medical imaging.
Rebecca Sawyer Lee, Jared Dunnmon, Ann He, Siyi Tang, Christopher Ré, Daniel L. Rubin
J. Biomed. Informatics4
2020 Emotion Detection in Online Social Networks: A Multilabel Learning Approach
abstract
Emotion detection in online social networks (OSNs) can benefit kinds of applications, such as personalized advertisement services, recommendation systems, etc. Conventionally, emotion analysis mainly focuses on the sentence level polarity prediction or single emotion label classification, however, ignoring the fact that emotions might coexist from users' perspective. To this end, in this work, we address the multiple emotions detection in OSNs from user-level view, and formulate this problem as a multilabel learning problem. First, we discover emotion labels correlations, social correlations, and temporal correlations from an annotated Twitter data set. Second, based on the above observations, we adopt a factor graph-based emotion recognition model to incorporate emotion labels correlations, social correlations, and temporal correlations into a general framework, and detect the multiple emotions based on the multilabel learning approach. Performance evaluation demonstrates that the factor graph-based emotion detection model can outperform the existing baselines.
Xiao Zhang 0015, Haochao Ying, Feng Li 0002, Siyi Tang, Sanglu Lu
IEEE Internet Things J.5
2016 Pose-Invariant Object Recognition for Event-Based Vision with Slow-ELM
Rohan Ghosh, Siyi Tang, Mahdi Rasouli, Nitish V. Thakor, Sunil L. Kukreja
ICANN (2)2