EDBT 2026 Demo / reviewers in the wild / expert
Fan Lin
dblp:68/2682
· DBLP profile ↗
58ranked-venue papers
11as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 10 since 2021Systems, architecture and hardware · 9 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeepIndicator: A Decision-Aware Multi-factor Rebalancing Framework via Reinforcement Learning
Jingan Chen, Jingqi Gao, Lifan Chen, Fan Lin |
ICIC (7) | 7 |
| 2026 | TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMsabstractLarge language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess LLMs’ proficiency in following instructions on diverse real-world tasks. We construct a hierarchical task tree encompassing seven major areas covering over 200 categories and over 800 tasks, which covers diverse capabilities such as question answering, reasoning, multi-turn dialogue, and text generation, to evaluate LLMs in a comprehensive and in-depth manner. We also design detailed evaluation standards and processes to facilitate consistent, unbiased judgments from human evaluators. A test set of over 3,000 instances is released, spanning different difficulty levels and knowledge domains. Our work provides a standardized methodology to evaluate human alignment in LLMs for both English and Chinese. We also analyze the feasibility of automating parts of evaluation with a strong LLM (GPT-4). Our framework supports a thorough assessment of LLMs as they are integrated into real-world applications. We have made publicly available the task tree, TencentLLMEval dataset, and evaluation methodology which have been demonstrated as effective in assessing the performance of Tencent Hunyuan LLMs. By doing so, we aim to facilitate the benchmarking of advances in the development of safe and human-aligned LLMs. Shuyi Xie, Wenlin Yao, Yong Dai 0001, Zishan Xu, Fan Lin, Donglin Zhou, Lifeng Jin, Xinhua Feng, Pengzhi Wei, Zhichao Hu, Dong Yu 0001, Zhengyou Zhang |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2026 | Unsupervised Feature Selection-Driven Active Learning for Semi-Supervised Automatic ECG AnalysisabstractAutomatic analysis methods of electrocardiograms (ECGs) usually required large-scale annotated training data, but the annotation process is extremely time-consuming. While semi-supervised learning can leverage unlabeled data, its performance depends heavily on the quality of the initial labeled subset. Active learning has been used to identify the most informative samples for annotation, but conventional approaches face three critical limitations: (1) dependency on manual intervention for iterative query design, (2) prohibitive computational costs during sample selection, and (3) limited compatibility with semi-supervised learning frameworks. To address these limitations, we proposed an Unsupervised Active Feature-selective Semi-Supervised Learning (UAFSSL) framework for ECG analysis, including an unsupervised feature selection-based active learning module and a semi-supervised learning module. UAFSSL captures latent data distributions via unsupervised feature extraction, selects diverse and representative samples using pseudo-label clustering, and integrates seamlessly with semi-supervised learning to eliminate human intervention. We validated our algorithm on an ECG waveform segmentation task and an atrial fibrillation detection task. In the waveform segmentation task, our method improved the F1-score for P-wave delineation by 2.4% compared to random sampling, using only 5% of labeled samples. For the atrial fibrillation detection task, we evaluated our method on both the AFDB and a 24-hour dataset collected from 500 atrial fibrillation patients. Using only 200 labeled samples for model training, our method achieved AUC improvements of 2.5% and 2.2% over random sampling in five-fold cross-validation. This is the first study to integrate unsupervised active learning with semi-supervised learning for automatic ECG analysis, offering a robust, automated solution to reduce annotation costs while enhancing clinical applicability. Xiao Li 0036, Songyang An, Yizhe Huang, Fan Lin, Peng Zhang 0106 |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Progressive Two-Stage ECG Signal Generation: From Structural Skeletons to High-Fidelity Detail RefinementabstractAutomatic analysis of electrocardiogram signals has been widely applied in the intelligent detection of cardiovascular diseases, however, class imbalance caused by rare disease samples limits model performance. In this paper, we propose a progressive generation strategy for synthesizing ECG signals, consisting of two stages: from structural skeletons to high-fidelity detail refinement. Specifically, we first train a generator to produce structural representations of ECG signals in the form of square-wave encodings, which capture coarse patterns. Then, an attention-based refinement module is introduced to fuse the detailed features from real ECG signals with the coarse features of the square-wave representation. This design ensures stable training and high-quality signal generation. Using the Resting ECG Segmentation Dataset, we synthesize AFIB, AT, and AF signals. Our method outperforms SOTA models with the lowest ED (12.815) and KLD (0.325) in feature space, and achieves an F1-score of 0.672-4% higher than the best baseline (0.632). Guangqi Chen, Fan Lin, Wenxuan Dai, Siyi Fang |
BIBM | 2 |
| 2025 | RemoteChess: Enhancing Older Adults' Social Connectedness via Designing a Virtual Reality Chinese Chess (Xiangqi) CommunityabstractThe decline of social connectedness caused by distance and physical limitations severely affects older adults' well-being and mental health. While virtual reality (VR) is promising for older adults to socialize remotely, existing social VR designs primarily focus on verbal communication (e.g., reminiscent, chat). Actively engaging in shared activities is also an important aspect of social connection. We designed RemoteChess, which constructs a social community and a culturally relevant activity (i.e., Chinese chess) for older adults to play while engaging in social interaction. We conducted a user study with groups of older adults interacting with each other through RemoteChess. Our findings indicate that RemoteChess enhanced participants' social connectedness by offering familiar environments, culturally relevant social catalysts, and asymmetric interactions. We further discussed design guidelines for designing culturally relevant social activities in VR to promote social connectedness for older adults. Qianjie Wei, Xiaoying Wei, Yiqi Liang, Fan Lin, Nuonan Si, Mingming Fan 0001 |
CHI | 4 |
| 2025 | Adaptive Bidirectional State Space Model for High-frequency Portfolio ManagementabstractState space models (SSMs) have recently shown great potential on long-range sequence modeling tasks. Benefiting from SSMs' low spatio-temporal overhead and powerful modeling capabilities, utilizing them for high-frequency portfolio management is an appealing research direction. However, representing financial data is challenging for SSMs due to: 1) the non-stationary nature of financial markets and 2) the requirement of asset correlations for financial understanding. In this paper, under a deep reinforcement learning (DRL) paradigm for high-frequency portfolio management, we propose a novel Adaptive Bidirectional State Space Model (ABSSM) to tackle the above challenges. Specifically, in order to cope with changing market conditions, we design an adaptive linear time-varying structure, which precisely captures domain shifts in temporal patterns through an input-dependent state transition matrix, thereby seizing fleeting arbitrage opportunities. Furthermore, we enhance this framework by constructing a bidirectional state space layer, which extracts asset correlations by compressing the global context. To the best of our knowledge, this is the first work that solves the high-frequency portfolio management problem by devising a specialized state space model in the DRL framework. Through extensive experiments on real-world data from the U.S., China, and cryptocurrency markets, we show that our proposed ABSSM significantly outperforms state-of-the-art benchmark methods in balancing profits and risks. Hanpeng Jiang, Ruibo Xiong, Yongrong Wu, Jingan Chen, Lifan Chen, Fan Lin |
CIKM | 8 |
| 2025 | FD-RLPO: Feature Domain-based Reinforcement Learning Framework for Portfolio OptimizationabstractPortfolio optimization is a critical issue in finance. In the past decade, reinforcement learning has advanced con-siderably in this area owing to its outstanding ability to solve complex sequential decision-making problems. Nevertheless, ex-isting methods still have great room for improvement in fully considering the dynamic fluctuations of intra-domain and inter-domain features in financial markets. To address this issue, we propose FD-RLPO, a Feature Domain-based Reinforcement Learning Framework for Portfolio Optimization, which inte-grates Prediction and Relation Modules while leveraging rein-forcement learning agents for adaptive strategy adjustment. In particular, the Prediction Module utilizes an inversion mechanism within a Transformer architecture to effectively extract sequential fluctuations in intra-domain features. Meanwhile, the Relation Module employs a self-supervised approach to capture inter-domain relational distribution patterns, mitigating the impact of data noise on reconstructed asset relationships. Ultimately, the Decision Module constructs a comprehensive representation of the market feature domain by synthesizing the outputs of the preceding modules and leveraging the Actor-Critic Network to achieve portfolio optimization. The simultaneous perception of both intra-domain and inter-domain features enables FD-RLPO to dynamically adjust to the volatility of financial markets and increase returns. Experiments on the CSI-300 and NASDAQ-lOO datasets demonstrate that FD-RLPO outperforms previous state-of-the-art methods in key metrics, including the Annualized Rate of Return (ARR) and Annualized Sharpe Ratio (ASR). Ablation studies further validate the effectiveness of each component. Jingyuan Feng, Qiyue Wu, Fan Lin |
CSCWD | 3 |
| 2025 | PSNet: A Multi-Period and Multi-Scale Model for Long-term Time Series ForecastingabstractLong-term Time Series Forecasting (LTSF) aims to predict time series data over extended future horizons. In recent years, multi-scale mixing and multi-period analysis have gained significant traction in LTSF models. However, existing approaches often fail to integrate multi-scale mixing with multi-period analysis effectively. Additionally, while the emerging cross-period sparse forecasting technique has shown promise in modeling time series periodicity, it struggles with data characterized by multiple and long periods. To address these problems, this paper introduces PSNet, a novel model designed to tackle the challenges posed by complex periodic patterns and long-term dependencies in real-world time series data. PSNet integrates multi-period mixing and multi-scale mixing into the cross-period sparse forecasting framework, enabling the decoupling of multiple periods and the effective capture of both high-frequency and low-frequency patterns in time series. Specifically, the model incorporates two key components: the intra-period multi-scale mixing module, which enhances the modeling of periodicity within individual periods, and the multi-period mixing module, which facilitates the fusion of diverse periodic patterns. Extensive experiments demonstrate the superiority of PSNet, achieving state-of-the-art performance in LTSF tasks and showcasing its ability to handle complex temporal dynamics effectively. Zehang Chen, Licun Dai, Fan Lin |
IJCNN | 3 |
| 2025 | Enhancing Sequential Recommendations with Sequence Diffusion Models for Long-Tail UsersabstractIn sequential recommendation systems (SRS), the prediction of users' next interests is fundamentally challenged by the long-tail user problem and data sparsity. Tail users often experience suboptimal performance in existing SRS models, attributable to their limited interactions and shorter interaction sequences compared to head users. To bridge this performance gap without degrading the experience of head users, we propose a method SDMLTU that generates diverse and high-quality sequences for tail users using a conditional diffusion model. This method first encodes item sequences with a pre-trained model and then employs a diffusion model to create pseudo sequences that augment the data for tail users. To ensure the quality of the generated sequences, a filtering mechanism is designed to select the most beneficial sequences. This enriched dataset is used to fine-tune the pre-trained sequence model, aiming to enhance the recommendation performance for tail users while maintaining the high standards for head users. Through extensive experiments, we demonstrate the effectiveness of our method in improving the recommendation quality for tail users. Qiyue Wu, Jingyuan Feng, Fan Lin |
IJCNN | 3 |
| 2025 | A Comprehensive Survey on Deep Learning Techniques in Educational Data MiningabstractAbstract Educational Data Mining (EDM) has emerged as a vital field of research, which harnesses the power of computational techniques to analyze educational data. With the increasing complexity and diversity of academic data, Deep Learning techniques have shown significant advantages in addressing the challenges associated with analyzing and modeling this data. Existing studies are scattered across various domains, making it challenging to gain a comprehensive understanding of how Deep Learning techniques can transform educational practices. This survey aims to systematically review the state-of-the-art in EDM with Deep Learning. We begin by providing a brief introduction to EDM and Deep Learning, highlighting their relevance in the context of modern education. Next, we present a detailed review of Deep Learning techniques applied in four typical educational scenarios, including knowledge tracing, student behavior detection, performance prediction, and personalized recommendation. Furthermore, a comprehensive overview of public datasets and processing tools for EDM is provided. Finally, we point out emerging trends and future directions, aiming to guide researchers and practitioners in advancing the field of EDM. Yuanguo Lin, Wei Xia 0001, Fan Lin, Zongyue Wang, Yong Liu 0020 |
Data Sci. Eng. | 4 |
| 2025 | Y-Net-ECG: A Multi-Lead informed and interpretable architecture for ECG segmentation across diverse rhythms
Peng Zhang 0106, Xiaoli Feng, Kaibiao Huang, Yinuo Zhao, Zuoming Fu, Zhigang Ye, Tao Wang 0138, Xiaoyun Yang, Fan Lin, Qiang Li 0018 |
Expert Syst. Appl. | 14 |
| 2025 | MDSGCN: Predicting Multiple Types of Mutation-Drug Association Through Signed Graph Convolution NetworkabstractCancer constitutes a significant global public health challenge, resulting in millions of fatalities annually. Gene mutations are pivotal in initiating and advancing cancer, disrupting regular cellular growth and differentiation mechanisms, thereby fostering tumor development. Consequently, comprehending the intricacies of gene mutations and their interplay with pharmaceuticals is imperative for cancer prevention, diagnosis, and treatment. Despite drug therapy being a cornerstone in cancer treatment, prognosticating and assessing multiple types of mutation-drug association remains a laborious and costly work. To address this problem, we develop a deep learning model grounded in signed graph convolution network (MDSGCN) to predict multiple types of mutation-drug association. We establish mutation-drug association as a signed bipartite network, comprising mutation nodes, drug nodes and two edge types including sensitive or resistant of mutations in drugs. MDSGCN extracts the subgraphs from the mutation-drug pairs in signed bipartite network, utilizing a label algorithm to learn subgraph structural features. Furthermore, MDSGCN integrates biological features (i.e., mutation-mutation similarity and drug-drug similarity) as the auxiliary information with the subgraph structural features to construct the prediction model. Experimental results demonstrate that our model consistently outperforms the state-of-the-art methods. The case study shows that MDSGCN can discover novel mutation-drug association and the association type. Haisong Feng, Fan Lin |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | Trend-Heuristic Reinforcement Learning Framework for News-Oriented Stock Portfolio ManagementabstractRecent studies have shown that reinforcement learning (RL) methods have brought significant performance gains for stock portfolio management (PM) because they effectively utilize historical price information and directly generate portfolio weights. We found, however, that there is still great room for improvement in how to fully consider the impact of financial news and stock trends on PM while avoiding model instability caused by feature disparity and fragile convergence due to RL itself. Addressing this allows us to develop more profitable and robust PM strategies. To this end, we propose TrendTrader, a novel RL framework for news-oriented stock portfolio management with trend heuristics. Specifically, TrendTrader utilizes a large language model (LLM) to obtain sentiment scores of stock news and generates heterogeneous contexts based on global sentiment embedding, which enhances model stability in processing multimodal features. Further, TrendTrader incorporates trend heuristics into both the network architecture and the reward function, while combining supervised learning and reinforcement learning in the form of incremental training to facilitate policy network convergence. Extensive experiments in the U.S. market and the China market verify the state-of-the-art performance of TrendTrader. Zhennan Chen, Hanpeng Jiang, Yuanguo Lin, Fan Lin |
ICASSP | 5 |
| 2024 | Asformer: Learning From Adjacent ScaleabstractLong-term series forecasting is crucial in real-world applications. Existing works have leveraged self-attention mechanisms and multi-scale temporal data to learn complex dependencies and patterns in long-term series prediction. However, most prediction models directly aggregate together multi-scale data resulting in the inability to appropriately capture key implicit information at each time step and lack of flexibility. To alleviate this problem, in this paper, we propose a novel Transformer-based network, dubbed ASformer, which can capture the spatio-temporal dependence of time series from adjacent scale features while obtaining global information. Specifically, we design a Scale Bonder that empowers ASformer with progressive capacities in capturing the correlations between data of two adjacent scales. We achieve the state-of-the-art performances with a 20.18% relative improvement on three benchmarks covering two practical applications: energy and disease. Our code will be made publicly available at: https://github.com/Hanpengjiang/ASformer. Hanpeng Jiang, Zhennan Chen, Fan Lin |
ICASSP | 4 |
| 2024 | IDGen: Item Discrimination Induced Prompt Generation for LLM EvaluationabstractAs Large Language Models (LLMs) become more capable of handling increasingly complex tasks, the evaluation set must keep pace with these advancements to ensure it remains sufficiently discriminative. Item Discrimination (ID) theory, which is widely used in educational assessment, measures the ability of individual test items to differentiate between high and low performers. Inspired by this theory, we propose an ID-induced prompt synthesis framework for evaluating LLMs so that the evaluation set continually updates and refines according to model abilities.
Our data synthesis framework prioritizes both breadth and specificity. It can generate prompts that comprehensively evaluate the capabilities of LLMs while revealing meaningful performance differences between models, allowing for effective discrimination of their relative strengths and weaknesses across various tasks and domains.
To produce high-quality data, we incorporate a self-correct mechanism into our generalization framework and develop two models to predict prompt discrimination and difficulty score to facilitate our data synthesis framework, contributing valuable tools to evaluation data synthesis research. We apply our generated data to evaluate five SOTA models. Our data achieves an average score of 51.92, accompanied by a variance of 10.06. By contrast, previous works (i.e., SELF-INSTRUCT and WizardLM) obtain an average score exceeding 67, with a variance below 3.2.
The results demonstrate that the data generated by our framework is more challenging and discriminative compared to previous works.
We will release a dataset of over 3,000 carefully crafted prompts to facilitate evaluation research of LLMs. Fan Lin, Shuyi Xie, Yong Dai 0001, Wenlin Yao, Tianjiao Lang, Yu Zhang 0004 |
NeurIPS | 1 |
| 2024 | Co-learning-assisted progressive dense fusion network for cardiovascular disease detection using ECG and PCG signals
Haobo Zhang 0003, Peng Zhang 0106, Fan Lin, Lianying Chao, Zhiwei Wang 0002, Qiang Li 0018 |
Expert Syst. Appl. | 3 |
| 2024 | Knowledge-aware reasoning with self-supervised reinforcement learning for explainable recommendation in MOOCs
Yuanguo Lin, Wei Zhang 0252, Fan Lin, Wenhua Zeng, Xiuze Zhou |
Neural Comput. Appl. | 3 |
| 2024 | A Survey on Reinforcement Learning for Recommender SystemsabstractRecommender systems have been widely applied in different real-life scenarios to help us find useful information. In particular, reinforcement learning (RL)-based recommender systems have become an emerging research topic in recent years, owing to the interactive nature and autonomous learning ability. Empirical results show that RL-based recommendation methods often surpass supervised learning methods. Nevertheless, there are various challenges in applying RL in recommender systems. To understand the challenges and relevant solutions, there should be a reference for researchers and practitioners working on RL-based recommender systems. To this end, we first provide a thorough overview, comparisons, and summarization of RL approaches applied in four typical recommendation scenarios, including interactive recommendation, conversational recommendation, sequential recommendation, and explainable recommendation. Furthermore, we systematically analyze the challenges and relevant solutions on the basis of existing literature. Finally, under discussion for open issues of RL and its limitations of recommender systems, we highlight some potential research directions in this field. Yuanguo Lin, Yong Liu 0020, Fan Lin, Lixin Zou, Wenhua Zeng, Huanhuan Chen 0001, Chunyan Miao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Diffusion Model for Camouflaged Object DetectionabstractCamouflaged object detection is a challenging task that aims to identify objects that are highly similar to their background. Due to the powerful noise-to-image denoising capability of denoising diffusion models, in this paper, we propose a diffusion-based framework for camouflaged object detection, termed diffCOD, a new framework that considers the camouflaged object segmentation task as a denoising diffusion process from noisy masks to object masks. Specifically, the object mask diffuses from the ground-truth masks to a random distribution, and the designed model learns to reverse this noising process. To strengthen the denoising learning, the input image prior is encoded and integrated into the denoising diffusion model to guide the diffusion process. Furthermore, we design an injection attention module (IAM) to interact conditional semantic features extracted from the image with the diffusion noise embedding via the cross-attention mechanism to enhance denoising learning. Extensive experiments on four widely used COD benchmark datasets demonstrate that the proposed method achieves favorable performance compared to the existing 11 state-of-the-art methods, especially in the detailed texture segmentation of camouflaged objects. Our code will be made publicly available at: https://github.com/ZNan-Chen/diffCOD. Zhennan Chen, Rongrong Gao, Tian-Zhu Xiang, Fan Lin |
ECAI | 4 |
| 2023 | mTrader: A Multi-Scale Signal Optimization Deep Reinforcement Learning Framework for Financial Trading (S)abstractIt is universally acknowledged that the financial trading is a thorny issue in time-series scenarios.On the one hand, due to the great randomness and instability in financial markets, the existing machine learning methods are inadequate for modeling high-frequency financial data.On the other hand, it remains a challenge to identify the validity of transaction actions to avoid high fees.To address these issues, we propose a novel trading framework, namely mTrader, to offer suitable trading strategies automatically.We creatively design a multiscale signal matrix to describe the temporal trends of markets.On this basis, Vector Quantized Variational AutoEncoder (VQ-VAE) was introduced to capture discrete latent variables.In addition, an offline Action Optimizer (AO) based on Proximal Policy Optimization (PPO) could help filter out sub-optimal trading action.Extensive experiments have shown that our model achieves state-of-the-art performance on many popular stock markets. Zhennan Chen, Lingyue Wei, Shibo Feng, Fan Lin |
SEKE | 6 |
| 2023 | Multi-Domain Feature Representation, Multi-Dimensional Feature Interaction for Person-Job FitabstractPerson-job fit aims to use the algorithms to match jobseekers with job postings to overcome information overload on online recruitment platforms.Traditional matching algorithms are not ideal in the feature representation and interaction of resumes and job postings.To this end, we propose a person-job fit model PJFFRFI based on multi-domain feature representation and multi-dimensional feature interaction, which comprehensively considers the features of various domains and learns feature correlation vectors in different dimensions.Specifically, we first divide the features in resumes and job postings into seven domains, and design different representation methods according to the data type.Then we propose a feature enhancement module (FEM) based on multi-head self-attention to learn the feature correlation vectors in resumes and job postings.Moreover, we propose a feature interaction module (FIM) to facilitate feature interaction both inside and outside the domain.Extensive experiments on a real-world dataset demonstrate that the proposed method significantly surpasses the state-of-the-art methods. Jianwen Ding, Zhennan Chen, Hanpeng Jiang, Fan Lin |
SEKE | 5 |
| 2023 | TransIntegrator: capture nearly full protein-coding transcript variants via integrating Illumina and PacBio transcriptomesabstractGenes have the ability to produce transcript variants that perform specific cellular functions. However, accurately detecting all transcript variants remains a long-standing challenge, especially when working with poorly annotated genomes or without a known genome. To address this issue, we have developed a new computational method, TransIntegrator, which enables transcriptome-wide detection of novel transcript variants. For this, we determined 10 Illumina sequencing transcriptomes and a PacBio full-length transcriptome for consecutive embryo development stages of amphioxus, a species of great evolutionary importance. Based on the transcriptomes, we employed TransIntegrator to create a comprehensive transcript variant library, namely iTranscriptome. The resulting iTrancriptome contained 91 915 distinct transcript variants, with an average of 2.4 variants per gene. This substantially improved current amphioxus genome annotation by expanding the number of genes from 21 954 to 38 777. Further analysis manifested that the gene expansion was largely ascribed to integration of multiple Illumina datasets instead of involving the PacBio data. Moreover, we demonstrated an example application of TransIntegrator, via generating iTrancriptome, in aiding accurate transcriptome assembly, which significantly outperformed other hybrid methods such as IDP-denovo and Trinity. For user convenience, we have deposited the source codes of TransIntegrator on GitHub as well as a conda package in Anaconda. In summary, this study proposes an affordable but efficient method for reliable transcriptomic research in most species. Yangmei Qin, Hao Chen 0113, Mindong Zhong, Te An, Linshan Chen, Yiquan Wang, Fan Lin, Zhi-Liang Ji |
Briefings Bioinform. | 9 |
| 2023 | DistVAE: Distributed Variational Autoencoder for sequential recommendation
Li Li 0122, Jianbing Xiahou, Fan Lin, Songzhi Su |
Knowl. Based Syst. | 3 |
| 2023 | Self-label correction for image classification with noisy labels
Yu Zhang 0004, Fan Lin, Siya Mi, Yali Bian |
Pattern Anal. Appl. | 2 |
| 2022 | Kernel-based Hybrid Interpretable Transformer for High-frequency Stock Movement PredictionabstractIt is universally acknowledged that the prediction of the stock movement is a thorny issue of the time-series prediction tasks. Often high-frequency stock behavior in the financial market is a prerequisite for investors to make effective investment strategies. Besides, the existing solution to the traditional stock movement prediction is generally converted into a classification task, which affects the practicability of the model. In this paper, we propose a Kernel-based Hybrid Interpretable Transformer(KHIT) model, which combines with a novel loss function to cope with the prediction task of non-stationary stock markets. Specifically, inspired by the Donchian price channels theory, we propose an adaptive renormalization kernel function that converts the original binary classification task (Rise or Fall) into the time-series prediction. Furthermore, we design a multi-order differential sequence loss function to identify the movement of high-frequency stock in future multi-scale periods. Finally, based on the Information Bottleneck (IB) theory and Transformer structure, we introduce the tensor decomposition and interpretable attention mechanism for improving the discrimination to the multi-types factors and the interpretability of our model. To be the best of our knowledge, it is the first work to achieve the high-frequency stock movement prediction task rather than classification. Experimental results show our proposed model outperforms several competitive methods in stock movement prediction on two non-stationary datasets. Fan Lin, Yuanguo Lin, Zhennan Chen, Huanyu You, Shibo Feng |
ICDM | 1 |
| 2022 | GCNCPR-ACPs: a novel graph convolution network method for ACPs predictionabstractBACKGROUND: Anticancer peptide (ACP) inhibits and kills tumor cells. Research on ACP is of great significance for the development of new drugs, and the prediction of ACPs and non-ACPs is the new hotspot. RESULTS: We propose a new machine learning-based method named GCNCPR-ACPs (a Graph Convolutional Neural Network Method based on collapse pooling and residual network to predict the ACPs), which automatically and accurately predicts ACPs using residual graph convolution networks, differentiable graph pooling, and features extracted using peptide sequence information extraction. The GCNCPR-ACPs method can effectively capture different levels of node attributes for amino acid node representation learning, GCNCPR-ACPs uses node2vec and one-hot embedding methods to extract initial amino acid features for ACP prediction. CONCLUSIONS: Experimental results of ten-fold cross-validation and independent validation based on different metrics showed that GCNCPR-ACPs significantly outperformed state-of-the-art methods. Specifically, the evaluation indicators of Matthews Correlation Coefficient (MCC) and AUC of our predicator were 69.5% and 90%, respectively, which were 4.3% and 2% higher than those of the other predictors, respectively, in ten-fold cross-validation. And in the independent test, the scores of MCC and SP were 69.6% and 93.9%, respectively, which were 37.6% and 5.5% higher than those of the other predictors, respectively. The overall results showed that the GCNCPR-ACPs method proposed in the current paper can effectively predict ACPs. Xiujin Wu, Wenhua Zeng, Fan Lin |
BMC Bioinform. | 3 |
| 2022 | MBRep: Motif-based representation learning in heterogeneous networks
Fan Lin, Beizhan Wang, Chunyan Li 0002 |
Expert Syst. Appl. | 2 |
| 2022 | Federated low-rank tensor projections for sequential recommendation
Li Li 0122, Fan Lin, Jianbing Xiahou, Yuanguo Lin, Yong Liu 0020 |
Knowl. Based Syst. | 2 |
| 2022 | Hierarchical reinforcement learning with dynamic recurrent mechanism for course recommendation
Yuanguo Lin, Fan Lin, Wenhua Zeng, Jianbing Xiahou, Li Li 0122, Yong Liu 0020, Chunyan Miao |
Knowl. Based Syst. | 2 |
| 2022 | Relation-aware dynamic attributed graph attention network for stocks recommendation
Shibo Feng, Yu Zuo, Fan Lin, Jianbing Xiahou |
Pattern Recognit. | 5 |
| 2022 | Semi-Supervised Learning for Automatic Atrial Fibrillation Detection in 24-Hour Holter MonitoringabstractParoxysmal atrial fibrillation (AF) is generally diagnosed by long-term dynamic electrocardiogram (ECG) monitoring. Identifying AF episodes from long-term ECG data can place a heavy burden on clinicians. Many machine-learning-based automatic AF detection methods have been proposed to solve this issue. However, these methods require numerous annotated data to train the model, and the annotation of AF in long-term ECG is extremely time-consuming. Reducing the demand for labeled data can effectively improve the clinical practicability of automatic AF detection methods. In this study, we developed a novel semi-supervised learning method that generated modified low-entropy labels of unlabeled samples for training a deep learning model to automatically detect paroxysmal AF in 24 h Holter monitoring data. Our method employed a 1D CNN-LSTM neural network with RR intervals as input and used few labeled training data with numerous unlabeled data for training the neural network. This method was evaluated using a 24 h Holter monitoring dataset collected from 1000 paroxysmal AF patients. Using labeled samples from only 10 patients for model training, our method achieved a sensitivity of 97.8%, specificity of 97.9%, and accuracy of 97.9% in five-fold cross-validation. Compared to the supervised learning method with complete labeled samples, the detection accuracy of our method was only 0.5% lower, while the workload of data annotation was significantly reduced by more than 98%. In general, this is the first study to apply semi-supervised learning techniques for automatic AF detection using ECG. Our method can effectively reduce the demand for AF data annotations and can improve the clinical practicability of automatic AF detection. Peng Zhang 0106, Fan Lin, Xiaoyun Yang, Qiang Li 0018 |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Temporal Convolution Network Based on Attention for Intelligent Anomaly Detection of Wind Turbine Blades
Jianwen Ding, Fan Lin, Shengbo Lv |
ICA3PP (1) | 2 |
| 2021 | NeuRank: learning to rank with neural networks for drug-target interaction predictionabstractBACKGROUND: Experimental verification of a drug discovery process is expensive and time-consuming. Therefore, recently, the demand to more efficiently and effectively identify drug-target interactions (DTIs) has intensified. RESULTS: We treat the prediction of DTIs as a ranking problem and propose a neural network architecture, NeuRank, to address it. Also, we assume that similar drug compounds are likely to interact with similar target proteins. Thus, in our model, we add drug and target similarities, which are very effective at improving the prediction of DTIs. Then, we develop NeuRank from a point-wise to a pair-wise, and further to list-wise model. CONCLUSION: Finally, results from extensive experiments on five public data sets (DrugBank, Enzymes, Ion Channels, G-Protein-Coupled Receptors, and Nuclear Receptors) show that, in identifying DTIs, our models achieve better performance than other state-of-the-art methods. Xiujin Wu, Wenhua Zeng, Fan Lin, Xiuze Zhou |
BMC Bioinform. | 3 |
| 2021 | Adaptive course recommendation in MOOCs
Yuanguo Lin, Shibo Feng, Fan Lin, Wenhua Zeng, Yong Liu 0020 |
Knowl. Based Syst. | 3 |
| 2021 | A Deep Segmentation Network of Multi-Scale Feature Fusion Based on Attention Mechanism for IVOCT Lumen ContourabstractRecently, coronary heart disease has attracted more and more attention, where segmentation and analysis for vascular lumen contour are helpful for treatment. And intravascular optical coherence tomography (IVOCT) images are used to display lumen shapes in clinic. Thus, an automatic segmentation method for IVOCT lumen contour is necessary to reduce the doctors' workload while ensuring diagnostic accuracy. In this paper, we proposed a deep residual segmentation network of multi-scale feature fusion based on attention mechanism (RSM-Network, Residual Squeezed Multi-Scale Network) to segment the lumen contour in IVOCT images. Firstly, three different data augmentation methods including mirror level turnover, rotation and vertical flip are considered to expand the training set. Then in the proposed RSM-Network, U-Net is contained as the main body, considering its characteristic of accepting input images with any sizes. Meanwhile, the combination of residual network and attention mechanism is applied to improve the ability of global feature extraction and solve the vanishing gradient problem. Moreover, the pyramid feature extraction structure is introduced to enhance the learning ability for multi-scale features. Finally, in order to increase the matching degree between the actual output and expected output, the cross entropy loss function is also used. A series of metrics are presented to evaluate the performance of our proposed network and the experimental results demonstrate that the proposed RSM-Network can learn the contour details better, contributing to strong robustness and accuracy for IVOCT lumen contour segmentation. Chenxi Huang 0001, Yisha Lan, Gaowei Xu, Xiaojun Zhai, Jipeng Wu, Fan Lin, Nianyin Zeng, Qingqi Hong, E. Y. K. Ng, Yonghong Peng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | DeepBAN: A Temporal Convolution-Based Communication Framework for Dynamic WBANsabstractWireless body area network (WBAN) has become a promising technology, which can be widely applied in health monitoring, and so on. However, the performance of a practical WBAN may severely suffer from the degradation caused by dynamic nature of wireless channels with the movements of human body. Traditional communication frameworks cannot catch up with the channel variation of dynamic WBANs, which may severely degrade the performance, so an accurate channel prediction model is necessary for developing an efficient transmission strategy. In this paper, we propose a DeepBAN communication framework for dynamic WBANs. In our proposed framework, a temporal convolution network (TCN) based deep learning approach is adopted for channel prediction, the computationally intensive task of which is processed by mobile edge computing (MEC), to reduce the response time. Given the predicted channel conditions, we propose a joint power control, time-slot allocation, and relay selection algorithm to maximize the energy efficiency of the system, taking into account the transmission reliability and end-to-end latency requirements. We evaluate the performance of DeepBAN, and the results show that it can achieve energy-efficient, reliable, and low-latency data transmission in dynamic WBANs, which can improve the system energy efficiency by 15% compared with the stochastic scheduling scheme. Kunqian Liu, Feng Ke, Rong Yu 0001, Fan Lin, Yueqian Wu, Derrick Wing Kwan Ng |
IEEE Trans. Commun. | 5 |
| 2021 | Self-Ensembling Co-Training Framework for Semi-Supervised COVID-19 CT SegmentationabstractThe coronavirus disease 2019 (COVID-19) has become a severe worldwide health emergency and is spreading at a rapid rate. Segmentation of COVID lesions from computed tomography (CT) scans is of great importance for supervising disease progression and further clinical treatment. As labeling COVID-19 CT scans is labor-intensive and time-consuming, it is essential to develop a segmentation method based on limited labeled data to conduct this task. In this paper, we propose a self-ensembled co-training framework, which is trained by limited labeled data and large-scale unlabeled data, to automatically extract COVID lesions from CT scans. Specifically, to enrich the diversity of unsupervised information, we build a co-training framework consisting of two collaborative models, in which the two models teach each other during training by using their respective predicted pseudo-labels of unlabeled data. Moreover, to alleviate the adverse impacts of noisy pseudo-labels for each model, we propose a self-ensembling strategy to perform consistency regularization for the up-to-date predictions of unlabeled data, in which the predictions of unlabeled data are gradually ensembled via moving average at the end of every training epoch. We evaluate our framework on a COVID-19 dataset containing 103 CT scans. Experimental results show that our proposed method achieves better performance in the case of only 4 labeled CT scans compared to the state-of-the-art semi-supervised segmentation networks. Caizi Li, Qi Dou 0001, Fan Lin, Kebao Zhang, Zuxin Feng, Weixin Si, Xuesong Deng, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | A dynamic priority strategy for IoV data scheduling towards key data
Chenxi Huang 0001, Gaowei Xu, Wen Zhou 0005, Yongqiang Cheng 0001, Yonghong Peng, Kaijian Xia, Fan Lin |
J. Supercomput. | 10 |
| 2020 | Detection of Abnormal Behavior Based on the Scene of Anti-photographing
Wei Zhang 0252, Fan Lin |
ICIC (1) | 2 |
| 2019 | Sparse online collaborative filtering with dynamic regularization
Kangkang Li 0001, Xiuze Zhou, Fan Lin, Wenhua Zeng, Beizhan Wang, Gil Alterovitz |
Inf. Sci. | 3 |
| 2018 | Accurate geometry modeling of vasculatures using implicit fitting with 2D radial basis functions
Qingqi Hong, Qingde Li, Beizhan Wang, Kunhong Liu 0001, Fan Lin, Juncong Lin, Zhihong Zhang 0001, Ming Zeng 0008 |
Comput. Aided Geom. Des. | 5 |
| 2018 | Chinese Character CAPTCHA Recognition and performance estimation via deep neural network
Dazhen Lin, Fan Lin, Yanping Lv, Feipeng Cai, Donglin Cao |
Neurocomputing | 2 |
| 2018 | Multi-datacenter cloud storage service selection strategy based on AHP and backward cloud generator model
Jianbing Xiahou, Fan Lin, Qihua Huang, Wenhua Zeng |
Neural Comput. Appl. | 2 |
| 2017 | An artificial neural network approach for screening test escapesabstractIn this paper we investigate the application of an artificial neural network (ANN) for screening test escapes. Specifically, we propose to train an autoencoder, an ANN, in an unsupervised way to fit the good chip population, i.e. using good chips only as the training set. The autoencoder is designed with both its input and output layers representing a set of features that characterize the test data of the chips under test, where we use the Euclidean distance between the values in the input and output layers as the cost function for training. Based on the trained autoencoder, if the test measurement of a query chip has an abnormally large value for the cost function, the chip is likely to be a test escape because it does not fit the characteristics of the good chip population captured by the model. We demonstrate that an autoencoder-based classification could achieve a higher detection rate for test escapes and a significant reduction in runtime and memory usage, compared with an SVM applied on the same features and some additional proximity features generated from multiple nonlinear transformations. Fan Lin, Kwang-Ting Cheng |
ASP-DAC | 1 |
| 2017 | A MOEA/D-based multi-objective optimization algorithm for remote medical
Shufu Lin, Fan Lin, Haishan Chen, Wenhua Zeng |
Neurocomputing | 2 |
| 2017 | Multi-kernel learning for multivariate performance measures optimization
Fan Lin, Jingbin Wang, Jianbing Xiahou, Nancy McDonald |
Neural Comput. Appl. | 1 |
| 2017 | Cloud computing system risk estimation and service selection approach based on cloud focus theory
Fan Lin, Wenhua Zeng, Lvqing Yang, Shufu Lin, Jiasong Zeng |
Neural Comput. Appl. | 1 |
| 2017 | Multiobjective Evolutionary Algorithm Based on Nondominated Sorting and Bidirectional Local Search for Big DataabstractThe improved differential evolutionary algorithm (EA) discussed in this paper is used to solve high-dimensional big data. Specifically, the algorithm improves population diversity by expanding the searching scope of the population, prevents premature deaths of the population through wider and more specific searches, and aims to solve the high-dimensional issue. To achieve this improvement goal, the paper suggests a multilayer hierarchical architecture on the basis of the above-mentioned heuristic mechanism. In each layer of the hierarchical architecture in the dynamic subpopulation, individuals who are more suitable for isolated evolution can better coexist with the original main population. We propose a new multiobjective optimization algorithm based on nondominated sorting and bidirectional local search (NSBLS). The algorithm takes the local beam search as the main body. NSBLS outputs the nondominated solution set through a continuous iterative search when the iteration termination condition is satisfied. It is worthy to note that the iteration of NSBLS is similar to the generation of the EA; therefore, this paper uses generation to represent the iterations. An algorithm introduces a new distribution maintaining strategy based on the sampling theory to combine with the fast nondominated sorting algorithm in order to select a new population into the next iteration. NSBLS will compare with three classical algorithms: NSGA-II, MOEA/D-DE, and MODEA through a series of bi-objective test problems. The proposed nondominated sorting and local search is able to find a better spread of solutions and better convergence to the true Pareto-optimal front compared to the other four algorithms. The outstanding performance of the proposed technology was proven in well-known benchmark problems. Fan Lin, Jiasong Zeng, Jianbing Xiahou, Beizhan Wang, Wenhua Zeng, Haibin Lv |
IEEE Trans. Ind. Informatics | 1 |
| 2016 | TCM clinic records data mining approaches based on weighted-LDA and multi-relationship LDA model
Fan Lin, Jianbing Xiahou, Zhuxiang Xu |
Multim. Tools Appl. | 1 |
| 2015 | Pairwise Proximity-Based Features for Test Escape ScreeningabstractTest escapes are chips that pass the chip-level test program but fail system-level test or in the field. It is known that statistical analysis based on chip production test data could identify abnormalities for screening test escapes. It has also been shown that from the chip test data, we can generate revealing features for statistical analysis by comparing the measurement data to different references such as the measurement mean of a wafer, the spatial pattern of a wafer, and the measurements of neighboring chips. Given these existing features as the base features, this paper proposes a new class of transformations which could generate additional informative features based on pairwise proximities between chips on the same wafer. Specifically, we apply multiple distance functions in a feature space composed of the base features and calculate the corresponding pairwise proximities between each pair of chips. Each of the resulting proximities could potentially embed some unique information that reveals the abnormalities of some test escapes. Then we convert the proximities into Euclidean vector spaces using constant shift embedding (CSE), which preserves the cluster structure through the conversion, so that traditional outlier analysis algorithms such as local outlier factor (LOF) can be applied. The LOF value and the first dimension in each embedded space are used as additional features for each sample. These new features, jointly analyzed with the base features, provide more revealing information about test escapes and thus further improve the test escape detection rate in our experiment based on production test data. Fan Lin, Chun-Kai Hsu, Alberto Giovanni Busetto, Kwang-Ting Cheng |
ICCAD | 1 |
| 2015 | AdaTest: An efficient statistical test framework for test escape screeningabstractStatistical analyses based on production test data can help identify test escapes, which are chips that pass the test program but fail later at system-level test or in field. Such analyses do not require extra physical measurements and can be referred to as statistical tests. For designing effective statistical tests, this paper investigates the use of a learning framework based on Adaptive Boosting, which has demonstrated great success in real-time face and object recognition. The framework is composed of a cascade of AdaBoost classifiers, each of which uses a small set of most relevant features that are automatically selected in the training phase, to identify a subset of test escapes. This framework therefore generates only the features that are most relevant for classification and significantly reduces the runtime and memory usage for statistical tests during test application. We also propose a new feature set to characterize the chips under test and demonstrate that including the new feature set as input to the proposed feature selection framework could reveal more test escapes. Fan Lin, Chun-Kai Hsu, Kwang-Ting Cheng |
ITC | 1 |
| 2014 | Learning from Production Test Data: Correlation Exploration and Feature EngineeringabstractThe huge amount of test data of a modern chip produced during manufacturing test could be mined for valuable information about the device under test (DUT), far more than the pass/fail information of each test item. Exploring the hidden correlations and patterns in the test data allows better understanding of the DUT and could therefore lead to test cost reduction or test quality improvement. There are several known types of correlations embedded in the test data: spatial correlations, inter-test-item correlations, and temporal correlations, each of which may involve a large number of data dimensions. Deriving and selecting the most relevant features for a specific application is critical for designing an effective and efficient mining solution. This paper provides an overview of recent research efforts on correlation exploration and development of a framework of feature engineering for learning from production test data. Fan Lin, Chun-Kai Hsu, Kwang-Ting Cheng |
ATS | 1 |
| 2014 | Joint Virtual Probe: Joint exploration of multiple test items' spatial patterns for efficient silicon characterization and test predictionabstractVirtual Probe (VP), proposed for characterization of spatial variations and for test time reduction, can effectively reconstruct the spatial pattern of a test item for an entire wafer using measurement values from only a small fraction of dies on the wafer. However, VP calculates the spatial signature of each test item separately, one item at a time, resulting in very long runtime for complex chips which often require hundreds, or even thousands, of test items in production. In this paper, we propose a new method, named Joint Virtual Probe (JVP), which can jointly derive spatial patterns of multiple test items. By simultaneously handling a large group of test items, JVP significantly reduces the overall runtime. And the prediction accuracy can also be improved because of JVP's implicit use of inter-test-item correlations in predicting spatial patterns. The experimental results on two industrial products, with 277 and 985 parametric test items in the production test programs respectively, demonstrate that, JVP achieves an average speedup of ~ 170X and ~ 50X over VP in the pre-test analysis and the test application phases respectively, as well as a slightly higher prediction accuracy than VP. Shuangyue Zhang, Fan Lin, Chun-Kai Hsu, Kwang-Ting Cheng |
DATE | 2 |
| 2014 | Feature engineering with canonical analysis for effective statistical tests screening test escapesabstractIt is known that statistical analysis of test data can help screen potential test escapes without additional physical measurements. Based on analysis of production test data, this paper focuses on feature engineering for statistical tests to screen test escapes. The features are engineered in two aspects: development of effective features and transformation of features into a different space in which the inherent difference between the test escapes and the normal population can be compacted into a small number of features. In feature development, we generate two sets of features to characterize a chip based on the amounts of the chip's test measurements deviated from the measurement means and their amounts deviated from the spatial patterns among dies on the same wafer. In feature transformation, the features are projected into the canonical space, in which the separation between the test escapes and the good chips are encapsulated into the first few dimensions. We show that each set of features reveals a unique set of test escapes, and the transformation of features can result in significant runtime reduction while keeping a comparable differentiating power as that in the original features. Therefore, both sets of features should be utilized and the canonical transformation should be applied when developing statistical tests for test escape reduction. Fan Lin, Chun-Kai Hsu, Kwang-Ting Cheng |
ITC | 1 |
| 2013 | Test data analytics - Exploring spatial and test-item correlations in production test dataabstractThe discovery of patterns and correlations hidden in the test data could help reduce test time and cost. In this paper, we propose a methodology and supporting statistical regression tools that can exploit and utilize both spatial and inter-test-item correlations in the test data for test time and cost reduction. We first describe a statistical regression method, called group lasso, which can identify inter-test-item correlations from test data. After learning such correlations, some test items can be identified for removal from the test program without compromising test quality. An extended version of this method, weighted group lasso, allows taking into account the distinct test time/cost of each individual test item in the formulation as a weighted optimization problem. As a result, its solution would favor more costly test items for removal from the test program. We further integrate weighted group lasso with another statistical regression technique, virtual probe, which can learn spatial correlations of test data across a wafer. The integrated method could then utilize both spatial and inter-test-item correlations to maximize the number of test items whose values can be predicted without measurement. Experimental results of a high-volume industrial device show that utilizing both spatial and inter-test-item correlations can help reduce test time by up to 55%. Chun-Kai Hsu, Fan Lin, Kwang-Ting Cheng, Wangyang Zhang, Xin Li 0001, John M. Carulli Jr., Kenneth M. Butler |
ITC | 2 |
| 2003 | Ubiquitous media agents: a framework for managing personally accumulated multimedia files
Wenyin Liu, Zheng Chen 0001, Fan Lin, HongJiang Zhang, Wei-Ying Ma |
Multim. Syst. | 3 |
| 2002 | User Intention Modelling in Web Applications Using Data Mining
Zheng Chen 0001, Fan Lin, Huan Liu 0001, Wei-Ying Ma, Wenyin Liu |
World Wide Web | 2 |
| 2001 | Ubiquitous media agents for managing personal multimedia filesabstractA novel idea of ubiquitous media agents is presented. Media agents are intelligent systems that are able to automatically collect and build personalized semantic indices of multimedia data on behalf of the user whenever and wherever he/she accesses/uses these multimedia data. The sources of these semantic descriptions are the textual context of the same documents that contain these multimedia data. The URLs of these multimedia data are indexed using these textual features. When the user wants to use these multimedia data once again, the media agents can also help the user find relevant multimedia data and provide proper suggestions based on the semantic indices. The media agents can also learn form the user's interaction records to refine the semantic indices and to model the user intentions and preferences. In our experiments, the media agents are effective in gathering relevant semantics for media objects and learning to provide precise suggestions when the user wants to re-use relevant media objects again. Wenyin Liu, Zheng Chen 0001, Fan Lin, Rui Yang 0003, Mingjing Li, HongJiang Zhang |
ACM Multimedia | 3 |