Moxian Song

dblp:198/2984 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
12since 2021 · last 2024
0000-0002-3847-6384ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Time pattern reconstruction for classification of irregularly sampled time series
Hongyan Li 0002, Moxian Song, Derun Cai, Baofeng Zhang, Shenda Hong
Pattern Recognit.3
2024 A Ranking-Based Cross-Entropy Loss for Early Classification of Time Series
abstract
Early classification tasks aim to classify time series before observing full data. It is critical in time-sensitive applications such as early sepsis diagnosis in the intensive care unit (ICU). Early diagnosis can provide more opportunities for doctors to rescue lives. However, there are two conflicting goals in the early classification task-accuracy and earliness. Most existing methods try to find a balance between them by weighing one goal against the other. But we argue that a powerful early classifier should always make highly accurate predictions at any moment. The main obstacle is that the key features suitable for classification are not obvious in the early stage, resulting in the excessive overlap of time series distributions in different time stages. The indistinguishable distributions make it difficult for classifiers to recognize. To solve this problem, this article proposes a novel ranking-based cross-entropy (RCE) loss to jointly learn the feature of classes and the order of earliness from time series data. In this way, RCE can help classifier to generate probability distributions of time series in different stages with more distinguishable boundary. Thus, the classification accuracy at each time step is finally improved. Besides, for the applicability of the method, we also accelerate the training process by focusing the learning process on high-ranking samples. Experiments on three real-world datasets show that our method can perform classification more accurately than all baselines at all moments.
Hongyan Li 0002, Moxian Song, Shenda Hong
IEEE Trans. Neural Networks Learn. Syst.3
2023 SPL-LDP: a label distribution propagation method for semi-supervised partial label learning
Moxian Song, Derun Cai, Shenda Hong, Hongyan Li 0002
Appl. Intell.1
2023 Adaptive model training strategy for continuous classification of time series
Hongyan Li 0002, Moxian Song, Derun Cai, Baofeng Zhang, Shenda Hong
Appl. Intell.3
2022 Deep Ordinal Neural Network for Length of Stay Estimation in the Intensive Care Units
abstract
Length of Stay (LoS) estimation is important for efficient healthcare resource management. Since the distribution of LoS is highly skewed, some previous works frame the LoS estimation as a multi-class classification problem by dividing the range of LoS into buckets. However, they ignore the ordinal relationship between labels. The distribution of bucketed LoS, with a heavy head and a heavy tail, is still imbalanced since the long tail is grouped into the last bucket. This paper proposes a Deep Ordinal neural network for Length of stay Estimation in the intensive care units (DOSE). DOSE can exploit the ordinal relationship and mitigate the skewness. The ordinal classification problem is decomposed into a series of binary classification sub-problems by using multiple binary classifiers. To maintain consistency among binary classifiers, the monotonicity constraint penalty is proposed. The number of samples whose labels are higher or lower than a given threshold is at the same level due to the heavy head and tail of the distribution. Therefore, the training data of each binary classifier are balanced. Experiments are conducted on the real-world healthcare dataset. DOSE outperforms all baseline methods in all metrics. The distribution of the prediction of DOSE is more aligned with the ground truth.
Derun Cai, Moxian Song, Baofeng Zhang, Shenda Hong, Hongyan Li 0002
CIKM2
2022 Confidence-Guided Learning Process for Continuous Classification of Time Series
abstract
In the real world, the class of a time series is usually labeled at the final time, but many applications require to classify time series at every time point. e.g. the outcome of a critical patient is only determined at the end, but he should be diagnosed at all times for timely treatment. Thus, we propose a new concept: Continuous Classification of Time Series (CCTS). It requires the model to learn data in different time stages. But the time series evolves dynamically, leading to different data distributions. When a model learns multi-distribution, it always forgets or overfits. We suggest that meaningful learning scheduling is potential due to an interesting observation: Measured by confidence, the process of model learning multiple distributions is similar to the process of human learning multiple knowledge. Thus, we propose a novel Confidence-guided method for CCTS (C3TS). It can imitate the alternating human confidence described by the Dunning-Kruger Effect. We define the objective-confidence to arrange data, and the self-confidence to control the learning duration. Experiments on four real-world datasets show that C3TS is more accurate than all baselines for CCTS.
Moxian Song, Derun Cai, Baofeng Zhang, Shenda Hong, Hongyan Li 0002
CIKM2
2022 Hypergraph Structure Learning for Hypergraph Neural Networks
abstract
Hypergraphs are natural and expressive modeling tools to encode high-order relationships among entities. Several variations of Hypergraph Neural Networks (HGNNs) are proposed to learn the node representations and complex relationships in the hypergraphs. Most current approaches assume that the input hypergraph structure accurately depicts the relations in the hypergraphs. However, the input hypergraph structure inevitably contains noise, task-irrelevant information, or false-negative connections. Treating the input hypergraph structure as ground-truth information unavoidably leads to sub-optimal performance. In this paper, we propose a Hypergraph Structure Learning (HSL) framework, which optimizes the hypergraph structure and the HGNNs simultaneously in an end-to-end way. HSL learns an informative and concise hypergraph structure that is optimized for downstream tasks. To efficiently learn the hypergraph structure, HSL adopts a two-stage sampling process: hyperedge sampling for pruning redundant hyperedges and incident node sampling for pruning irrelevant incident nodes and discovering potential implicit connections. The consistency between the optimized structure and the original structure is maintained by the intra-hyperedge contrastive learning module. The sampling processes are jointly optimized with HGNNs towards the objective of the downstream tasks. Experiments conducted on 7 datasets show shat HSL outperforms the state-of-the-art baselines while adaptively sparsifying hypergraph structures.
Derun Cai, Moxian Song, Baofeng Zhang, Shenda Hong, Hongyan Li 0002
IJCAI2
2022 Hypergraph Contrastive Learning for Electronic Health Records
abstract
Electronic Health Records (EHR) is the repository of patients' involved medical codes in the hospital, including diagnosis codes, medication codes, procedure codes, lab codes, and so on. EHR inherently contains various kinds of relationships such as the code-code, the patient-patient, and the patient-code relationship. Recent research shows that graph representation learning can be an effective tool for capturing complex relationships. However, none of the existing methods considered high-order interactions between patients and medical codes or considered the three relationships together. In this paper, we propose Hypergraph Contrastive Learning (HCL), to jointly learn patient embeddings and code embeddings from the combination of the above three relationships. HCL first constructs a hypergraph from the EHR data. Then, the medical code graph and the patient graph are constructed based on the hypergraph. Empowered with hypergraph attention network, Transformer, and graph attention network, HCL learns representations from three graphs respectively. Next, contrastive learning is applied to aggregate information from these graphs. Finally, the learned representations can support downstream tasks in supervised learning settings and self-supervised learning settings. Experiments are conducted on eICU and MIMIC-III datasets with mortality prediction and readmission prediction tasks. Results show that our method outperforms almost all compared methods on all evaluation metrics and HCL can learn patient representations from medical codes even without labeled data.
Derun Cai, Moxian Song, Baofeng Zhang, Shenda Hong, Hongyan Li 0002
SDM3
2022 GRP-FED: Addressing Client Imbalance in Federated Learning via Global-Regularized Personalization
abstract
Since data is presented long-tailed in reality, it is challenging for Federated Learning (FL) to train across decentralized clients as practical applications. We present Global-Regularized Personalization (GRP-FED) to tackle the data imbalanced issue by considering a single global model and multiple local models for each client. With adaptive aggregation, the global model treats multiple clients fairly and mitigates the global long-tailed issue. Each local model is learned from the local data and aligns with its distribution for customization. To prevent the local model from just overfitting, GRP-FED applies an adversarial discriminator to regularize between the learned global-local features. Extensive results show that our GRP-FED improves under both global and local scenarios on real-world MIT-BIH and synthesis CIFAR-10 datasets, achieving comparable performance and addressing client imbalance.
Yen-hsiu Chou, Shenda Hong, Derun Cai, Moxian Song, Hongyan Li 0002
SDM5
2022 Dlsa: Semi-supervised partial label learning via dependence-maximized label set assignment
Moxian Song, Hongyan Li 0002, Derun Cai, Shenda Hong
Inf. Sci.1
2022 Classifying vaguely labeled data based on evidential fusion
Moxian Song, Derun Cai, Shenda Hong, Hongyan Li 0002
Inf. Sci.1
2021 TE-ESN: Time Encoding Echo State Network for Prediction Based on Irregularly Sampled Time Series Data
abstract
Prediction based on Irregularly Sampled Time Series (ISTS) is of wide concern in real-world applications. For more accurate prediction, methods had better grasp more data characteristics. Different from ordinary time series, ISTS is characterized by irregular time intervals of intra-series and different sampling rates of inter-series. However, existing methods have suboptimal predictions due to artificially introducing new dependencies in a time series and biasedly learning relations among time series when modeling these two characteristics. In this work, we propose a novel Time Encoding (TE) mechanism. TE can embed the time information as time vectors in the complex domain. It has the properties of absolute distance and relative distance under different sampling rates, which helps to represent two irregularities. Meanwhile, we create a new model named Time Encoding Echo State Network (TE-ESN). It is the first ESNs-based model that can process ISTS data. Besides, TE-ESN incorporates long short-term memories and series fusion to grasp horizontal and vertical relations. Experiments on one chaos system and three real-world datasets show that TE-ESN performs better than all baselines and has better reservoir property.
Shenda Hong, Moxian Song, Yen-hsiu Chou, Yongyue Sun, Derun Cai, Hongyan Li 0002
IJCAI3
2020 Knowledge-shot learning: An interpretable deep model for classifying imbalanced electrocardiography data
Yen-hsiu Chou, Shenda Hong, Junyuan Shang, Moxian Song, Hongyan Li 0002
Neurocomputing5
2017 A New Interval Numbers Power Average Operator in Multiple Attribute Decision Making
abstract
How to fuse uncertain information in multiple attribute decision making (MADM) efficiently is still an open issue. The power average operation is an effective tool to aggregate interval data. However, existing methods to aggregate interval numbers based on power average operator are relatively complicated. In this paper, a simple and effective support function of interval data is proposed. Then, a novel interval number power average operation operator is presented. Finally, a practical MADM problem is used to show the efficiency of the developed method.
Moxian Song, Wen Jiang 0002, Chunhe Xie
Int. J. Intell. Syst.1