VLDB 2026 Research / reviewers in the wild / expert
Hongyan Li 0002
dblp:62/5909-2
· DBLP profile ↗
57ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0001-7174-2851ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 37 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 28 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Knowledge to Causality: Self-supervised Representation Learning for Granger Causal Discovery in Groups of Time Series
Bo Liu 0113, Di Dai, Hongyan Li 0002, Shenda Hong |
DASFAA (4) | 3 |
| 2026 | LLM-GC: Advancing Granger Causal Discovery from Time Series with Multimodel Language ModelingabstractRecent advances in neural Granger causal methods have shown promise in modeling temporal nonlinear dependencies. However, existing approaches remain confined to raw time-series data, inherently lacking contextual semantics and tending to overfit, which undermines their real-world applicability. To address these challenges, we propose LLM-GC, a novel LLM-empowered multimodal Granger causality discovery framework that enriches unimodal temporal dynamics with semantic priors and world knowledge distilled from large language models (LLMs). LLM-GC leverages dual-modality encoding to capture and align temporal and contextual dynamics by Cross-Modal Dual Retrieval while avoiding causal entanglement across modalities. To extract multimodal causal features, we introduce a causality-aware self-attention mechanism by simply inverting the conventional self-attention structure, enabling a shared causality augmenter to effectively highlight consistent causal patterns across modalities. LLM-GC is the first to bridge LLMs and Granger causality, and experiments on synthetic and real-world benchmark datasets demonstrate that LLM-GC outperforms existing state-of-the-art methods in Granger causal discovery. Bo Liu 0113, Hongyan Li 0002, Shenda Hong |
WSDM | 2 |
| 2025 | Preacher: Paper-to-Video Agentic System
Ling Yang 0006, Hao Luo 0004, Fan Wang 0019, Hongyan Li 0002, Mengdi Wang 0001 |
ICCV | 5 |
| 2025 | DiffuGC: Diffusion Model Can Help Discover Granger Causality from Interventional Time SeriesabstractDiscovering Granger causality from time series data is fundamental to understanding dynamic systems, yet most existing methods struggle with unknown intervention targets or causal structures in real-world scenarios. In this paper, we propose DiffuGC, a novel diffusion-based framework that unifies observational and interventional causal discovery through a generative denoising process. By introducing diffusive interventions, which apply progressive interventions without any prior knowledge, DiffuGC amplifies causal signals while preserving structural information. Furthermore, we introduce a denoising NoiFormer with adaptive attention to both short- and long-term causal dependencies, which disentangles trend and seasonal components to enable accurate reconstruction of causal structures from interventional data. To the best of our knowledge, we are the first to integrate diffusion models with interventional Granger causal discovery. Extensive experiments on synthetic, quasi-real, and real-world benchmarks demonstrate that DiffuGC consistently outperforms state-of-the-art baselines in both observational and interventional data. Moreover, we introduce an intriguing notion, Causality Acceleration, characterized by the early emergence of informative causal patterns within the diffusion path, which may open up promising directions for future research on efficient and adaptive causal discovery. Bo Liu 0113, Hongyan Li 0002, Shenda Hong |
ICDM | 2 |
| 2025 | Reading Your Heart: Learning ECG Words and Sentences via Pre-training ECG Language ModelabstractElectrocardiogram (ECG) is essential for the clinical diagnosis of arrhythmias and other heart diseases, but deep learning methods based on ECG often face limitations due to the need for high-quality annotations. Although previous ECG self-supervised learning (eSSL) methods have made significant progress in representation learning from unannotated ECG data, they typically treat ECG signals as ordinary time-series data, segmenting the signals using fixed-size and fixed-step time windows, which often ignore the form and rhythm characteristics and latent semantic relationships in ECG signals. In this work, we introduce a novel perspective on ECG signals, treating heartbeats as words and rhythms as sentences. Based on this perspective, we first designed the QRS-Tokenizer, which generates semantically meaningful ECG sentences from the raw ECG signals. Building on these, we then propose HeartLang, a novel self-supervised learning framework for ECG language processing, learning general representations at form and rhythm levels. Additionally, we construct the largest heartbeat-based ECG vocabulary to date, which will further advance the development of ECG language processing. We evaluated HeartLang across six public ECG datasets, where it demonstrated robust competitiveness against other eSSL methods. Our data and code are publicly available at https://github.com/PKUDigitalHealth/HeartLang. Jiarui Jin, Hongyan Li 0002, Shenda Hong |
ICLR | 3 |
| 2024 | TEST: Text Prototype Aligned Embedding to Activate LLM's Ability for Time SeriesabstractThis work summarizes two ways to accomplish Time-Series (TS) tasks in today's Large Language Model (LLM) context: LLM-for-TS (model-centric) designs and trains a fundamental large model, or fine-tunes a pre-trained LLM for TS data; TS-for-LLM (data-centric) converts TS into a model-friendly representation to enable the pre-trained LLM to handle TS data. Given the lack of data, limited resources, semantic context requirements, and so on, this work focuses on TS-for-LLM, where we aim to activate LLM's ability for TS data by designing a TS embedding method suitable for LLM. The proposed method is named TEST. It first tokenizes TS, builds an encoder to embed TS via instance-wise, feature-wise, and text-prototype-aligned contrast, where the TS embedding space is aligned to LLM’s embedding layer space, then creates soft prompts to make LLM more open to that embeddings, and finally implements TS tasks using the frozen LLM. We also demonstrate the feasibility of TS-for-LLM through theory and experiments. Experiments are carried out on TS classification, forecasting, and representation tasks using eight frozen LLMs with various structures and sizes. The results show that the pre-trained LLM with TEST strategy can achieve better or comparable performance than today's SOTA TS models, and offers benefits for few-shot and generalization. By treating LLM as the pattern machine, TEST can endow LLM's ability to process TS data without compromising language ability. We hope that this study will serve as a foundation for future work to support TS+LLM progress. Hongyan Li 0002, Yaliang Li, Shenda Hong |
ICLR | 2 |
| 2024 | Retrieval-Augmented Diffusion Models for Time Series ForecastingabstractWhile time series diffusion models have received considerable focus from many recent works, the performance of existing models remains highly unstable. Factors limiting time series diffusion models include insufficient time series datasets and the absence of guidance. To address these limitations, we propose a Retrieval-Augmented Time series Diffusion model (RATD). The framework of RATD consists of two parts: an embedding-based retrieval process and a reference-guided diffusion model. In the first part, RATD retrieves the time series that are most relevant to historical time series from the database as references. The references are utilized to guide the denoising process in the second part. Our approach allows leveraging meaningful samples within the database to aid in sampling, thus maximizing the utilization of datasets. Meanwhile, this reference-guided mechanism also compensates for the deficiencies of existing time series diffusion models in terms of guidance. Experiments and visualizations on multiple datasets demonstrate the effectiveness of our approach, particularly in complicated prediction tasks. Our code is available at https://github.com/stanliu96/RATD Ling Yang 0006, Hongyan Li 0002, Shenda Hong |
NeurIPS | 3 |
| 2024 | Synthesis of Standard 12-Lead ECG from Single-Lead ECG Using Shifted Diffusion Models
Hongyan Li 0002, Shenda Hong |
ECML/PKDD (9) | 2 |
| 2024 | Time pattern reconstruction for classification of irregularly sampled time series
Hongyan Li 0002, Moxian Song, Derun Cai, Baofeng Zhang, Shenda Hong |
Pattern Recognit. | 2 |
| 2024 | A Ranking-Based Cross-Entropy Loss for Early Classification of Time SeriesabstractEarly classification tasks aim to classify time series before observing full data. It is critical in time-sensitive applications such as early sepsis diagnosis in the intensive care unit (ICU). Early diagnosis can provide more opportunities for doctors to rescue lives. However, there are two conflicting goals in the early classification task-accuracy and earliness. Most existing methods try to find a balance between them by weighing one goal against the other. But we argue that a powerful early classifier should always make highly accurate predictions at any moment. The main obstacle is that the key features suitable for classification are not obvious in the early stage, resulting in the excessive overlap of time series distributions in different time stages. The indistinguishable distributions make it difficult for classifiers to recognize. To solve this problem, this article proposes a novel ranking-based cross-entropy (RCE) loss to jointly learn the feature of classes and the order of earliness from time series data. In this way, RCE can help classifier to generate probability distributions of time series in different stages with more distinguishable boundary. Thus, the classification accuracy at each time step is finally improved. Besides, for the applicability of the method, we also accelerate the training process by focusing the learning process on high-ranking samples. Experiments on three real-world datasets show that our method can perform classification more accurately than all baselines at all moments. Hongyan Li 0002, Moxian Song, Shenda Hong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | SPL-LDP: a label distribution propagation method for semi-supervised partial label learning
Moxian Song, Derun Cai, Shenda Hong, Hongyan Li 0002 |
Appl. Intell. | 5 |
| 2023 | Adaptive model training strategy for continuous classification of time series
Hongyan Li 0002, Moxian Song, Derun Cai, Baofeng Zhang, Shenda Hong |
Appl. Intell. | 2 |
| 2022 | Deep Ordinal Neural Network for Length of Stay Estimation in the Intensive Care UnitsabstractLength of Stay (LoS) estimation is important for efficient healthcare resource management. Since the distribution of LoS is highly skewed, some previous works frame the LoS estimation as a multi-class classification problem by dividing the range of LoS into buckets. However, they ignore the ordinal relationship between labels. The distribution of bucketed LoS, with a heavy head and a heavy tail, is still imbalanced since the long tail is grouped into the last bucket. This paper proposes a Deep Ordinal neural network for Length of stay Estimation in the intensive care units (DOSE). DOSE can exploit the ordinal relationship and mitigate the skewness. The ordinal classification problem is decomposed into a series of binary classification sub-problems by using multiple binary classifiers. To maintain consistency among binary classifiers, the monotonicity constraint penalty is proposed. The number of samples whose labels are higher or lower than a given threshold is at the same level due to the heavy head and tail of the distribution. Therefore, the training data of each binary classifier are balanced. Experiments are conducted on the real-world healthcare dataset. DOSE outperforms all baseline methods in all metrics. The distribution of the prediction of DOSE is more aligned with the ground truth. Derun Cai, Moxian Song, Baofeng Zhang, Shenda Hong, Hongyan Li 0002 |
CIKM | 6 |
| 2022 | Confidence-Guided Learning Process for Continuous Classification of Time SeriesabstractIn the real world, the class of a time series is usually labeled at the final time, but many applications require to classify time series at every time point. e.g. the outcome of a critical patient is only determined at the end, but he should be diagnosed at all times for timely treatment. Thus, we propose a new concept: Continuous Classification of Time Series (CCTS). It requires the model to learn data in different time stages. But the time series evolves dynamically, leading to different data distributions. When a model learns multi-distribution, it always forgets or overfits. We suggest that meaningful learning scheduling is potential due to an interesting observation: Measured by confidence, the process of model learning multiple distributions is similar to the process of human learning multiple knowledge. Thus, we propose a novel Confidence-guided method for CCTS (C3TS). It can imitate the alternating human confidence described by the Dunning-Kruger Effect. We define the objective-confidence to arrange data, and the self-confidence to control the learning duration. Experiments on four real-world datasets show that C3TS is more accurate than all baselines for CCTS. Moxian Song, Derun Cai, Baofeng Zhang, Shenda Hong, Hongyan Li 0002 |
CIKM | 6 |
| 2022 | Hypergraph Structure Learning for Hypergraph Neural NetworksabstractHypergraphs are natural and expressive modeling tools to encode high-order relationships among entities. Several variations of Hypergraph Neural Networks (HGNNs) are proposed to learn the node representations and complex relationships in the hypergraphs. Most current approaches assume that the input hypergraph structure accurately depicts the relations in the hypergraphs. However, the input hypergraph structure inevitably contains noise, task-irrelevant information, or false-negative connections. Treating the input hypergraph structure as ground-truth information unavoidably leads to sub-optimal performance. In this paper, we propose a Hypergraph Structure Learning (HSL) framework, which optimizes the hypergraph structure and the HGNNs simultaneously in an end-to-end way. HSL learns an informative and concise hypergraph structure that is optimized for downstream tasks. To efficiently learn the hypergraph structure, HSL adopts a two-stage sampling process: hyperedge sampling for pruning redundant hyperedges and incident node sampling for pruning irrelevant incident nodes and discovering potential implicit connections. The consistency between the optimized structure and the original structure is maintained by the intra-hyperedge contrastive learning module. The sampling processes are jointly optimized with HGNNs towards the objective of the downstream tasks. Experiments conducted on 7 datasets show shat HSL outperforms the state-of-the-art baselines while adaptively sparsifying hypergraph structures. Derun Cai, Moxian Song, Baofeng Zhang, Shenda Hong, Hongyan Li 0002 |
IJCAI | 6 |
| 2022 | Hypergraph Contrastive Learning for Electronic Health RecordsabstractElectronic Health Records (EHR) is the repository of patients' involved medical codes in the hospital, including diagnosis codes, medication codes, procedure codes, lab codes, and so on. EHR inherently contains various kinds of relationships such as the code-code, the patient-patient, and the patient-code relationship. Recent research shows that graph representation learning can be an effective tool for capturing complex relationships. However, none of the existing methods considered high-order interactions between patients and medical codes or considered the three relationships together. In this paper, we propose Hypergraph Contrastive Learning (HCL), to jointly learn patient embeddings and code embeddings from the combination of the above three relationships. HCL first constructs a hypergraph from the EHR data. Then, the medical code graph and the patient graph are constructed based on the hypergraph. Empowered with hypergraph attention network, Transformer, and graph attention network, HCL learns representations from three graphs respectively. Next, contrastive learning is applied to aggregate information from these graphs. Finally, the learned representations can support downstream tasks in supervised learning settings and self-supervised learning settings. Experiments are conducted on eICU and MIMIC-III datasets with mortality prediction and readmission prediction tasks. Results show that our method outperforms almost all compared methods on all evaluation metrics and HCL can learn patient representations from medical codes even without labeled data. Derun Cai, Moxian Song, Baofeng Zhang, Shenda Hong, Hongyan Li 0002 |
SDM | 6 |
| 2022 | GRP-FED: Addressing Client Imbalance in Federated Learning via Global-Regularized PersonalizationabstractSince data is presented long-tailed in reality, it is challenging for Federated Learning (FL) to train across decentralized clients as practical applications. We present Global-Regularized Personalization (GRP-FED) to tackle the data imbalanced issue by considering a single global model and multiple local models for each client. With adaptive aggregation, the global model treats multiple clients fairly and mitigates the global long-tailed issue. Each local model is learned from the local data and aligns with its distribution for customization. To prevent the local model from just overfitting, GRP-FED applies an adversarial discriminator to regularize between the learned global-local features. Extensive results show that our GRP-FED improves under both global and local scenarios on real-world MIT-BIH and synthesis CIFAR-10 datasets, achieving comparable performance and addressing client imbalance. Yen-hsiu Chou, Shenda Hong, Derun Cai, Moxian Song, Hongyan Li 0002 |
SDM | 6 |
| 2022 | Dlsa: Semi-supervised partial label learning via dependence-maximized label set assignment
Moxian Song, Hongyan Li 0002, Derun Cai, Shenda Hong |
Inf. Sci. | 2 |
| 2022 | Classifying vaguely labeled data based on evidential fusion
Moxian Song, Derun Cai, Shenda Hong, Hongyan Li 0002 |
Inf. Sci. | 5 |
| 2021 | TE-ESN: Time Encoding Echo State Network for Prediction Based on Irregularly Sampled Time Series DataabstractPrediction based on Irregularly Sampled Time Series (ISTS) is of wide concern in real-world applications. For more accurate prediction, methods had better grasp more data characteristics. Different from ordinary time series, ISTS is characterized by irregular time intervals of intra-series and different sampling rates of inter-series. However, existing methods have suboptimal predictions due to artificially introducing new dependencies in a time series and biasedly learning relations among time series when modeling these two characteristics. In this work, we propose a novel Time Encoding (TE) mechanism. TE can embed the time information as time vectors in the complex domain. It has the properties of absolute distance and relative distance under different sampling rates, which helps to represent two irregularities. Meanwhile, we create a new model named Time Encoding Echo State Network (TE-ESN). It is the first ESNs-based model that can process ISTS data. Besides, TE-ESN incorporates long short-term memories and series fusion to grasp horizontal and vertical relations. Experiments on one chaos system and three real-world datasets show that TE-ESN performs better than all baselines and has better reservoir property. Shenda Hong, Moxian Song, Yen-hsiu Chou, Yongyue Sun, Derun Cai, Hongyan Li 0002 |
IJCAI | 7 |
| 2020 | Knowledge-shot learning: An interpretable deep model for classifying imbalanced electrocardiography data
Yen-hsiu Chou, Shenda Hong, Junyuan Shang, Moxian Song, Hongyan Li 0002 |
Neurocomputing | 6 |
| 2020 | Semantics-aware influence maximization in social networks
Yipeng Chen, Qiang Qu 0001, Yuanxiang Ying, Hongyan Li 0002, Jialie Shen 0001 |
Inf. Sci. | 4 |
| 2019 | GAMENet: Graph Augmented MEmory Networks for Recommending Medication CombinationabstractRecent progress in deep learning is revolutionizing the healthcare domain including providing solutions to medication recommendations, especially recommending medication combination for patients with complex health conditions. Existing approaches either do not customize based on patient health history, or ignore existing knowledge on drug-drug interactions (DDI) that might lead to adverse outcomes. To fill this gap, we propose the Graph Augmented Memory Networks (GAMENet), which integrates the drug-drug interactions knowledge graph by a memory module implemented as a graph convolutional networks, and models longitudinal patient records as the query. It is trained end-to-end to provide safe and personalized recommendation of medication combination. We demonstrate the effectiveness and safety of GAMENet by comparing with several state-of-the-art methods on real EHR data. GAMENet outperformed all baselines in all effectiveness measures, and also achieved 3.60% DDI rate reduction from existing EHR data. Junyuan Shang, Cao Xiao, Tengfei Ma 0001, Hongyan Li 0002, Jimeng Sun 0001 |
AAAI | 4 |
| 2019 | RDPD: Rich Data Helps Poor Data via ImitationabstractIn many situations, we need to build and deploy separate models in related environments with different data qualities. For example, an environment with strong observation equipments (e.g., intensive care units) often provides high-quality multi-modal data, which are acquired from multiple sensory devices and have rich-feature representations. On the other hand, an environment with poor observation equipment (e.g., at home) only provides low-quality, uni-modal data with poor-feature representations. To deploy a competitive model in a poor-data environment without requiring direct access to multi-modal data acquired from a rich-data environment, this paper develops and presents a knowledge distillation (KD) method (RDPD) to enhance a predictive model trained on poor data using knowledge distilled from a high-complexity model trained on rich, private data. We evaluated RDPD on three real-world datasets and shown that its distilled model consistently outperformed all baselines across all datasets, especially achieving the greatest performance improvement over a model trained only on low-quality data by 24.56% on PR-AUC and 12.21% on ROC-AUC, and over that of a state-of-the-art KD model by 5.91% on PR-AUC and 4.44% on ROC-AUC. Shenda Hong, Cao Xiao, Trong Nghia Hoang, Tengfei Ma 0001, Hongyan Li 0002, Jimeng Sun 0001 |
IJCAI | 5 |
| 2019 | MINA: Multilevel Knowledge-Guided Attention for Modeling Electrocardiography SignalsabstractElectrocardiography (ECG) signals are commonly used to diagnose various cardiac abnormalities. Recently, deep learning models showed initial success on modeling ECG data, however they are mostly black-box, thus lack interpretability needed for clinical usage. In this work, we propose MultIlevel kNowledge-guided Attention networks (MINA) that predict heart diseases from ECG signals with intuitive explanation aligned with medical knowledge. By extracting multilevel (beat-, rhythm- and frequency-level) domain knowledge features separately, MINA combines the medical knowledge and ECG data via a multilevel attention model, making the learned models highly interpretable. Our experiments showed MINA achieved PR-AUC 0.9436 (outperforming the best baseline by 5.51%) in real world ECG dataset. Finally, MINA also demonstrated robust performance and strong interpretability against signal distortion and noise contamination. Shenda Hong, Cao Xiao, Tengfei Ma 0001, Hongyan Li 0002, Jimeng Sun 0001 |
IJCAI | 4 |
| 2019 | K-margin-based Residual-Convolution-Recurrent Neural Network for Atrial Fibrillation DetectionabstractAtrial Fibrillation (AF) is an abnormal heart rhythm which can trigger cardiac arrest and sudden death. Nevertheless, its interpretation is mostly done by medical experts due to high error rates of computerized interpretation. One study found that only about 66% of AF were correctly recognized from noisy ECGs. This is in part due to insufficient training data, class skewness, as well as semantical ambiguities caused by noisy segments in an ECG record. In this paper, we propose a K-margin-based Residual-Convolution-Recurrent neural network (K-margin-based RCR-net) for AF detection from noisy ECGs. In detail, a skewness-driven dynamic augmentation method is employed to handle the problems of data inadequacy and class imbalance. A novel RCR-net is proposed to automatically extract both long-term rhythm-level and local heartbeat-level characters. Finally, we present a K-margin-based diagnosis model to automatically focus on the most important parts of an ECG record and handle noise by naturally exploiting expected consistency among the segments associated for each record. The experimental results demonstrate that the proposed method with 0.8125 F1NAOP score outperforms all state-of-the-art deep learning methods for AF detection task by 6.8%. Shenda Hong, Junyuan Shang, Hongyan Li 0002, Junqing Xie |
IJCAI | 6 |
| 2018 | Negative-Aware Influence Maximization on Social NetworksabstractHow to minimize the impact of negative users within the maximal set of influenced users? The Influenced Maximization (IM) is important for various applications. However, few studies consider the negative impact of some of the influenced users.We propose a negative-aware influence maximization problem by considering users' negative impact. A novel algorithm is proposed to solve the problem. Experiments on real-world datasets show the proposed algorithm can achieve 70% improvement on average in expected influence compared with rivals. Yipeng Chen, Hongyan Li 0002, Qiang Qu 0001 |
AAAI | 2 |
| 2018 | Generative Adversarial Network for Abstractive Text SummarizationabstractIn this paper, we propose an adversarial process for abstractive text summarization, in which we simultaneously train a generative model G and a discriminative model D. In particular, we build the generator G as an agent of reinforcement learning, which takes the raw text as input and predicts the abstractive summarization. We also build a discriminator which attempts to distinguish the generated summary from the ground truth summary. Extensive experiments demonstrate that our model achieves competitive ROUGE scores with the state-of-the-art methods on CNN/Daily Mail dataset. Qualitatively, we show that our model is able to generate more abstractive, readable and diverse summaries. Linqing Liu, Min Yang 0007, Qiang Qu 0001, Jia Zhu 0003, Hongyan Li 0002 |
AAAI | 6 |
| 2018 | Knowledge Guided Multi-instance Multi-label Learning via Neural Networks in Medicines PredictionabstractPredicting medicines for patients with co-morbidity has long been recognized as a hard task due to complex dependencies between diseases and medicines. Efforts have been made recently to build high-order dependency between diseases and medicines by extracting knowledge from electronic health records (EHR). But current works failed to utilize additional knowledge and ignored the data skewness problem which lead to sub-optimal combination of medicines. In this paper, we formulate the medicines prediction task in multi-instance multi-label learning framework considering the multi-diagnoses as input instances and multi-medicines as output labels. We propose a knowledge-guided multi-instance multi-label networks called \mname where two types of additional knowledge are incorporated into a RNN encoder-decoder model. The utilization of structural knowledge like clinical ontology provides a way to learn better representation called tree embedding by utilizing the ancestors’ information. Contextual knowledge is a global summarization of input instances which is informative for personal prediction. Experiments are conducted on a real world clinical dataset which showed the necessity to combine both contextual and structural knowledge and the \mname performs better than baselines up to 4+% in terms of Jaccard similarity score. Junyuan Shang, Shenda Hong, Hongyan Li 0002 |
ACML | 5 |
| 2017 | Assessing Death Risk of Patients with Cardiovascular Disease from Long-Term Electrocardiogram Streams Summarization
Shenda Hong, Hongyan Li 0002 |
PAKDD (1) | 4 |
| 2016 | FVBM: A Filter-Verification-Based Method for Finding Top-k Closeness Centrality on Dynamic Social Networks
Yiyong Lin 0003, Yuanxiang Ying, Shenda Hong, Hongyan Li 0002 |
APWeb (2) | 5 |
| 2016 | Real-Time Anomaly Detection over ECG Data Stream Based on Component Spectrum
Shenda Hong, Hongyan Li 0002 |
APWeb (2) | 4 |
| 2016 | Online Learning for Accurate Real-Time Map Matching
Biwei Liang, Tengjiao Wang 0003, Shun Li 0001, Wei Chen 0021, Hongyan Li 0002, Kai Lei |
PAKDD (2) | 5 |
| 2016 | Inferring Social Roles of Mobile Users Based on Communication Behaviors
Yipeng Chen, Hongyan Li 0002, Gaoshan Miao |
WAIM (1) | 2 |
| 2016 | Detecting Data-model-oriented Anomalies in Parallel Business Process
Ning Yin, Hongyan Li 0002, Lilue Fan |
WAIM (2) | 3 |
| 2015 | An Adaptive Skew Handling Join Algorithm for Large-scale Data Analysis
Tengjiao Wang 0003, Shun Li 0001, Hongyan Li 0002, Kai Lei |
WAIM | 5 |
| 2014 | A Segment-Wise Method for Pseudo Periodic Time Series Prediction
Ning Yin, Shenda Hong, Hongyan Li 0002 |
ADMA | 4 |
| 2014 | An Adaptive Skew Insensitive Join Algorithm for Large Scale Data Analytics
Wenjing Liao, Tengjiao Wang 0003, Hongyan Li 0002, Dongqing Yang, Kai Lei |
APWeb | 3 |
| 2014 | BF-Matrix: A Secondary Index for the Cloud Storage
Hongyan Li 0002, Yue Wang 0014, Tengjiao Wang 0003, Dongqing Yang |
WAIM | 2 |
| 2014 | Finding Vacant Taxis Using Large Scale GPS Traces
Hongyan Li 0002, Shenda Hong, Yiyong Lin 0003, Nana Fan, Gaoyan Ou, Tengjiao Wang 0003, Lilue Fan |
WAIM | 2 |
| 2013 | Logistic Regression Bias Correction for Large Scale Data with Rare Events
Hongyan Li 0002, Hanchen Su, Gaoyan Ou, Tengjiao Wang 0003 |
ADMA (2) | 2 |
| 2013 | An effective method to analyze variations of high-dimensional patterns over medical streamsabstractIn medical field, patterns over time-varied data streams usually imply high domain value. The variations of patterns can often be very complex and hard to evaluate. Traditional methods usually take each pattern as a whole to analyze data stream variations or only focus on one type of variation; however, few works have achieved a widely applicable resolution. This paper considers the feature of sub parts for data stream patterns and studies their variations and relationships from the perspective of multiple dimensions, to explore a comprehensive understanding for the variation history and effectively support different types of queries to help analyze the variations. This paper first decomposes patterns into different dimensions and then evaluates the variations of each dimension. After that, a data cube called VS-Cube is used to find out the variations of a single dimension as well as the relationships between different dimensions within a certain pattern. At last, a case study on disease MI over medical stream is given to demonstrate the effectiveness and efficiency of our proposed methods. Hongyan Li 0002, Feifei Li 0003, Lilue Fan |
BIBM | 2 |
| 2012 | PCG: An Efficient Method for Composite Pattern Matching over Data Streams
Cheng Ju, Hongyan Li 0002, Feifei Li 0003 |
ADMA | 2 |
| 2012 | VS-Cube: Analyzing Variations of Multi-dimensional Patterns over Data Streams
Hongyan Li 0002, Feifei Li 0003, Gaoshan Miao |
ADMA | 2 |
| 2011 | Efficient Topological OLAP on Information Networks
Qiang Qu 0001, Feida Zhu 0001, Xifeng Yan, Jiawei Han 0001, Philip S. Yu, Hongyan Li 0002 |
DASFAA (1) | 6 |
| 2010 | A General Multi-relational Classification Approach Using Feature Generation and Selection
Miao Zou, Tengjiao Wang 0003, Hongyan Li 0002, Dongqing Yang |
ADMA (2) | 3 |
| 2010 | A Heuristic Method for Unstructured Pattern Management over Data StreamsabstractPattern management is an important task in data stream mining and has attracted increasing attention recently. Variations of data stream patterns typically imply some fundamental changes of underlying objects and possess significant domain meanings. Many database applications require investigating the history information to get the knowledge about the evolving process of data streams. However, in most circumstances, the data stream patterns are unstructured: limited memory space cannot record all the patterns discovered online, no training sets or predefined models are available, and large numbers of noises bring another non-trivial challenge. This paper presents our research effort in online pattern management over such streams. A novel algorithm is proposed to detect stream changes, organize meaningful patterns and distinguish useful variations from noises. It extracts new trends from unstructured data heuristically, and involves a special parameter to identify whether the current event should be treated as significant. Several experiments are performed and the results prove this new method feasible and efficient. Gaoshan Miao, Hongyan Li 0002, Tengjiao Wang 0003 |
APWeb | 2 |
| 2008 | PEDS-VM: A Variation Management Prototype for Pattern Evolving Data StreamsabstractWidely applied in many domains, data stream processing has attracted more and more attention in database and sensor network communities. In this demo, we present our system PEDS-VM, a real-time surveillance system for managing pattern variation over evolving medical streams. PEDS-VM utilizes an effective strategy and a state-based window framework to extract the evolving patterns. After extraction, PEDS-VM uses a novel storage structure called PGG to efficiently record these incremental patterns. Moreover, several important application scenarios of PEDS-VM are also discussed in this demonstration. Xinbiao Zhou, Gaoshan Miao, Hongyan Li 0002, Lv-an Tang |
WAIM | 3 |
| 2008 | PGG: An Online Pattern Based Approach for Stream Variation Management
Lu-An Tang, Bin Cui 0001, Hongyan Li 0002, Gaoshan Miao, Dongqing Yang, Xinbiao Zhou |
J. Comput. Sci. Technol. | 3 |
| 2007 | Effective variation management for pseudo periodical streamsabstractMany database applications require the analysis and processing of data streams. In such systems, huge amounts of data arrive rapidly and their values change over time. The variations on streams typically imply some fundamental changes of the underlying objects and possess significant domain meanings. In some data streams, successive events seem to recur in a certain time interval, but the data indeed evolves with tiny differences as time elapses. This feature is called pseudo periodicity, which poses a non-trivial challenge to variation management in data streams. This paper presents our research effort in online variation management over such streams, and the idea can be applied to the problem domain of medical applications, such as patient vital signal monitoring. We propose a new method named Pattern Growth Graph (PGG) to detect and manage variations over pseudo periodical streams. PGG adopts the wave-pattern to capture the major information of data evolution and represent them compactly. With the help of wave-pattern matching algorithm, PGG detects the stream variations in a single pass over the stream data. PGG only stores the different segments of the pattern for incoming stream, and hence it can substantially compress the data without losing important information. The statistical information of PGG helps to distinguish meaningful data changes from noise and to reconstruct the stream with acceptable accuracy. Extensive experiments on real datasets containing millions of data items demonstrate the feasibility and effectiveness of the proposed scheme. Lv-an Tang, Bin Cui 0001, Hongyan Li 0002, Gaoshan Miao, Dongqing Yang, Xinbiao Zhou |
SIGMOD Conference | 3 |
| 2006 | DSEC: A Data Stream Engine Based Clinical Information System
Hongyan Li 0002, Zijing Hu, Jianlong Gao, Shiwei Tang, Xinbiao Zhou |
APWeb | 2 |
| 2006 | WISE: A Prototype for Ontology Driven Development of Web Information Systems
Lv-an Tang, Hongyan Li 0002, Baojun Qiu, Meimei Li, Dongqing Yang, Shiwei Tang |
APWeb | 2 |
| 2006 | DOPA: A Data-Driven and Ontology-Based Method for Ad Hoc Process Awareness in Web Information Systems
Meimei Li, Hongyan Li 0002, Lv-an Tang, Baojun Qiu |
WISE | 2 |
| 2005 | PODWIS: A Personalized Tool for Ontology Development in Domain Specific Web Information System
Lv-an Tang, Hongyan Li 0002, Zhiyong Pan, Shaohua Tan, Baojun Qiu, Shiwei Tang |
APWeb | 2 |
| 2005 | Understanding User Operations on Web Page in WISE
Hongyan Li 0002, Ming Xue, Shiwei Tang, Dongqing Yang |
WAIM | 1 |
| 2005 | An Ontology Based Approach to Construct Behaviors in Web Information Systems
Lv-an Tang, Hongyan Li 0002, Zhiyong Pan, Dongqing Yang, Meimei Li, Shiwei Tang, Ying Ying |
WAIM | 2 |
| 2001 | An XML Based Electronic Medical Record Integration System
Hongyan Li 0002, Shiwei Tang, Dongqing Yang |
WAIM | 1 |