Chunlei Tang

dblp:211/4125 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
10since 2021 · last 2023
0000-0002-6460-0246ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2023 Automatic ICD Coding Based on Segmented ClinicalBERT with Hierarchical Tree Structure Learning
Beichen Kang, Xiaosu Wang, Yun Xiong, Yao Zhang 0009, Chaofan Zhou, Yangyong Zhu, Jiawei Zhang 0001, Chunlei Tang
DASFAA (4)8
2023 MulEA: Multi-type Entity Alignment of Heterogeneous Medical Knowledge Graphs
Mingxia Wang, Yun Xiong, Jingwen Yue, Yao Zhang 0009, Chunlei Tang
DASFAA (2)6
2023 Solving combinatorial optimization problems over graphs with BERT-Based Deep Reinforcement Learning
Qi Wang 0044, Kenneth H. Lai, Chunlei Tang
Inf. Sci.3
2023 Learning from undercoded clinical records for automated International Classification of Diseases (ICD) coding
abstract
OBJECTIVES: To develop an unbiased objective for learning automatic coding algorithms from clinical records annotated with only partial relevant International Classification of Diseases codes, as annotation noise in undercoded clinical records used as training data can mislead the learning process of deep neural networks. MATERIALS AND METHODS: We use Medical Information Mart for Intensive Care III as our dataset. We employ positive-unlabeled learning to achieve unbiased loss estimation, which is free of misleading training signal. We then utilize reweighting mechanism to compensate for the imbalance between positive and negative samples. To further close the performance gap caused by poor quality annotation, we integrate the supervision provided by the automatic annotation tool Medical Concept Annotation Toolkit which can ease the heavy burden of manual validation. RESULTS: Our benchmarking results show that positive-unlabeled learning with reweighting outperforms competitive baseline methods over a range of missing label ratios. Integrating supervision provided by annotation tool further boosted the performance. DISCUSSION: Considering the annotation noise and severe imbalance, unbiased loss estimation and reweighting mechanism are both important for learning from undercoded clinical records. Unbiased loss requires the estimation of false negative ratios and estimation through trained models is practical and competitive. CONCLUSIONS: The combination of positive-unlabeled learning with reweighting and supervision provided by the annotation tool is a promising solution to learn from undercoded clinical records.
Yun Xiong, Dan Shi 0006, Yifei Lin, Lifang He 0001, Yao Zhang 0009, Joseph M. Plasek, Li Zhou 0007, David W. Bates, Chunlei Tang
J. Am. Medical Informatics Assoc.10
2023 Mastering construction heuristics with self-play deep reinforcement learning
Qi Wang 0044, Chunlei Tang
Neural Comput. Appl.3
2022 Generative Adversarial Imitation Learning to Search in Branch-and-Bound Algorithms
Qi Wang 0044, Suzanne V. Blackley, Chunlei Tang
DASFAA (2)3
2021 Addressing Disparities in Diabetes Using Temporal Fairness Models
Joseph M. Plasek, Chunlei Tang, Yun Xiong, Yangyong Zhu, Yanming He, Patricia C. Dykes, David W. Bates, Li Zhou 0007
AMIA2
2021 Radiology Report Generation for Rare Diseases via Few-shot Transformer
abstract
Reliable automatic radiology report generation is highly desired to reduce the labor-intensive and error-prone workload for healthcare workers. While some multi-modal learning models have been proposed to study on this task, few of them paid attention to the radiology report generation for rare diseases, except for RareGen which solved this problem by enhancing the semantic representations of rare diseases. However, there still exist several open problems to be addressed. The first lies in the low proportion of disease regions in an image, making the visual information redundant or irrelevant to rare diseases to be encoded. The second lies in that correlations modeled in the encoding stage may not be effectively decoded in the decoding stage due to the multi-modal representation. To address these two issues, we propose a few-shot Transformer radiology report generation model, namely TransGen, for rare diseases. It integrates the advantages of Transformer with two key modules assembled. Specifically, in the encoding stage, a Semantic-aware Visual Learning (SVL) module is introduced to capture the regions of rare diseases. Following that, in the decoding stage, a Memory Augmented Semantic Enhancement (MASE) module is proposed to enhance intermediate representations. It could make full use of the semantic information contained in the historical-generated sentences to benefit report generation involving rare diseases. Extensive experiments have been conducted on two public datasets of IU X-Ray and MIMIC-CXR to demonstrate the effectiveness of our proposed model.
Xing Jia, Yun Xiong, Jiawei Zhang 0001, Yao Zhang 0009, Suzanne V. Blackley, Yangyong Zhu, Chunlei Tang
BIBM7
2021 Deep reinforcement learning for transportation network combinatorial optimization: A survey
Qi Wang 0044, Chunlei Tang
Knowl. Based Syst.2
2021 Estimating Time to Progression of Chronic Obstructive Pulmonary Disease With Tolerance
abstract
We defined tolerance range as the distance of observing similar disease conditions or functional status from the upper to the lower boundaries of a specified time interval. A tolerance range was identified for linear regression and support vector machines to optimize the improvement rate (defined as IR) on accuracy in predicting mortality risk in patients with chronic obstructive pulmonary disease using clinical notes. The corpus includes pulmonary, cardiology, and radiology reports of 15,500 patients who died between 2011 and 2017. Their performance was compared against a long short-term memory recurrent neural network. The results demonstrate an overall improvement by those basic machine learning approaches after considering an optimal tolerance range: the average IR of linear regression was 90.1% and the maximum IR of support vector machines was 66.2%. There was a similitude between the time segments produced by our tolerance algorithms and those produced by the long short-term memory.
Chunlei Tang, Joseph M. Plasek, Meihan Wan, Min-Jeoung Kang, Sevan M. Dulgarian, Yun Xiong, David W. Bates, Li Zhou 0007
IEEE J. Biomed. Health Informatics1
2020 Following data as it crosses borders during the COVID-19 pandemic
abstract
Data change the game in terms of how we respond to pandemics. Global data on disease trajectories and the effectiveness and economic impact of different social distancing measures are essential to facilitate effective local responses to pandemics. COVID-19 data flowing across geographic borders are extremely useful to public health professionals for many purposes such as accelerating the pharmaceutical development pipeline, and for making vital decisions about intensive care unit rooms, where to build temporary hospitals, or where to boost supplies of personal protection equipment, ventilators, or diagnostic tests. Sharing data enables quicker dissemination and validation of pharmaceutical innovations, as well as improved knowledge of what prevention and mitigation measures work. Even if physical borders around the globe are closed, it is crucial that data continues to transparently flow across borders to enable a data economy to thrive, which will promote global public health through global cooperation and solidarity.
Joseph M. Plasek, Chunlei Tang, Yangyong Zhu, Yajun Huang, David W. Bates
J. Am. Medical Informatics Assoc.2
2019 Data Reconstruction Based on Temporal Expressions in Clinical Notes
abstract
Learning representations of clinical notes poses challenges in handling complex content that necessitates preprocessing steps to make the data more suitable for data mining. An important issue, addressed here, is that of temporal expressions, where cues indicate the time when clinical events occur. We present a three-step data reconstruction algorithm for transforming similar clinical entities (e.g., symptoms, complications) into sequential data through unsupervised annotation of temporal expressions. First, the data reconstruction algorithm detects if an expression has temporal intent. Second, it decomposes and rewrites the expression into non-temporal sub-expression and temporal constraints. Finally, it clusters similar non-temporal sub-expressions by using unsupervised sentence embedding under the modified K-medoids paradigm. We experimented with our proposed algorithm on clinical notes associated with chronic obstructive pulmonary disease (COPD). Visualizing reconstruction results of cardiology reports for a longitudinal cohort of patients with COPD demonstrated that this algorithm is feasible.
Chunlei Tang, Joseph M. Plasek, Yun Xiong, Min-Jeoung Kang, Patricia C. Dykes, David W. Bates, Li Zhou 0007
BIBM2
2018 RegionAl: an Optimized Regional Classifier to Predict Mortality in Chronic Obstructive Pulmonary Disease Patients
Chunlei Tang, Joseph M. Plasek, Yun Xiong, Li Zhou 0007, David W. Bates
AMIA2
2018 A Deep Learning Approach to Handling Temporal Variation in Chronic Obstructive Pulmonary Disease Progression
Chunlei Tang, Joseph M. Plasek, Yun Xiong, David W. Bates, Li Zhou 0007
BIBM1
2018 Predicting Disease-related Associations by Heterogeneous Network Embedding
Yun Xiong, Lu Ruan 0002, Mengjie Guo, Chunlei Tang, Xiangnan Kong, Yangyong Zhu, Wei Wang 0010
BIBM4
2017 Developing a regional classifier to track patient needs in medical literature using spiral timelines on a geographical map
abstract
Research clues can be expressed as coherent chains of keywords grouped by theme. Capturing clues to research from the vast and expanding medical literature is valuable. Yet, it is difficult to automatically create clear visualizations of research clues despite the presence of many competing summarization tools. In this paper, we propose a linear classifier based on a spiral, which we call a regional classifier. The study emphasizes the development of visualization methods and the process of finding a specific research clue to track patient needs reported in medical literature. When timelines are combined with a spiral geographical map, they show a geometric shape that helps to reveal the clues from different spatial viewpoints and periodical constraints. Our evaluation showed that the regional classifier produces better visual effects than support vector machine classifiers. It covers important concepts of each theme and is able to represent the relationships among papers in a way that captures continuous developments and changes in key themes.
Chunlei Tang, Kenneth H. Lai, Yuxuan She, Yun Xiong, Li Zhou 0007
BIBM1