EDBT 2026 Demo / reviewers in the wild / expert
Fan Yang 0010
dblp:29/3081-10
· DBLP profile ↗
38ranked-venue papers
6as first author
22since 2021 · last 2025
0000-0002-3523-1138ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Systems, architecture and hardware · 7 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Simultaneously local and global contrastive learning of graph representations
Binsheng Hong, Zhaori Guo, Shunzhi Zhu, Kaibiao Lin, Fan Yang 0010 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Health Insurance Fraud Detection via Multiview Heterogeneous Information Networks With Augmented Graph Structure LearningabstractWith the continuous development of the health insurance system, the problem of health insurance fraud is becoming more and more prominent. Health insurance fraud not only harms health insurance organizations financially but also causes serious disruptions to the normal provision of healthcare services. Current fraud detection methods usually rely on a single data source or specific features, making it difficult to comprehensively capture the complex patterns and intrinsic associations of fraudulent behavior. Therefore, the aim of this study is to efficiently integrate feature information from different perspectives of health insurance data using multiperspective representation learning. Health insurance fraud detection via multiview heterogeneous information networks with augmented graph structure learning (MHINAGSL) is proposed. Initially, multiple heterogeneous graphs are constructed based on health insurance datasets, encompassing diffusion, topology, feature, and semantic graphs. To enhance these foundational perspectives, a confidence fusion optimizer is employed. Additionally, a multichannel semantic graph convolution with shared parameters is introduced to effectively integrate diverse meta-path semantic graphs. Subsequently, an adaptive attention mechanism assigns weights to different views, facilitating the comprehensive fusion of information from multiple perspectives and enabling a more accurate characterization of the richness of health insurance data. Ultimately, the combination of multiple loss functions is employed for end-to-end model training, aiming to fully leverage the complementary advantages of multiview data, thereby enhancing the robustness and generalization ability of fraud detection models. We fully apply the proposed method to real health insurance datasets and provide in-depth validation of its comprehensive utility in real-world scenarios. Extensive experimental results fully confirm that our approach achieves significant results in improving accuracy and recall. Particularly noteworthy is that in the anomaly detection experiments, our method achieved an F1 score as high as 0.92, further proving its excellent performance. This also demonstrates the practical application and research value of the approach based on multiview heterogeneous information network structure learning in the field of health insurance fraud detection. Binsheng Hong, Ping Lu 0012, Kaibiao Lin, Fan Yang 0010 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | Causal-Driven Skill Prerequisite Structure DiscoveryabstractKnowing a prerequisite structure among skills in a subject domain effectively enables several educational applications, including intelligent tutoring systems and curriculum planning. Traditionally, educators or domain experts use intuition to determine the skills' prerequisite relationships, which is time-consuming and prone to fall into the trap of blind spots. In this paper, we focus on inferring the prerequisite structure given access to students' performance on exercises in a subject. Nevertheless, it is challenging since students' mastery of skills can not be directly observed, but can only be estimated, i.e., its latency in nature. To tackle this problem, we propose a causal-driven skill prerequisite structure discovery (CSPS) method in a two-stage learning framework. In the first stage, we learn the skills' correlation relationships presented in the covariance matrix from the student performance data while, through the predicted covariance matrix in the second stage, we consider a heuristic method based on conditional independence tests and standardized partial variance to discover the prerequisite structure. We demonstrate the performance of the new approach with both simulated and real-world data. The experimental results show the effectiveness of the proposed model for identifying the skills' prerequisite structure. Shenbao Yu, Yifeng Zeng, Fan Yang 0010, Yinghui Pan |
AAAI | 3 |
| 2024 | Temporal Coverage Optimization for Epoch-based Vehicular Crowdsensing RecruitmentabstractThe proliferation of vehicles equipped with onboard sensors has positioned vehicular crowdsensing (VCS) as a promising approach for urban sensing and monitoring. These mobile sensors offer a unique opportunity to capture comprehensive urban information. Consequently, the selection of appropriate participants is critical for the success of vehicular crowdsensing campaigns. The effectiveness of such campaigns is typically assessed based on the spatial and temporal coverage of the road network. However, existing research predominantly focuses on spatial coverage, often overlooking the temporal aspects of data collection. In this study, we bridge this gap by introducing a temporal coverage framework that accounts for both sensing frequency and duration. We formulate an optimization problem aimed at minimizing the overall recruitment cost, adhering to the principle of Epoch-based Participant Recruitment (EPR) and utilizing the deterministic trajectory information of vehicles. Our approach is validated through simulation experiments using a real-world dataset of vehicle trajectories, demonstrating superior performance compared to existing participant recruitment strategies. Ziying Pan, Yongxuan Lai, Yi Fan 0001, Fan Yang 0010 |
ICPADS | 6 |
| 2024 | Revisiting the loss functions in sequential recommendation
Fangyu Li 0003, Shenbao Yu, Fan Yang 0010 |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | Adaptive Neighbor Graph Aggregated Graph Attention Network for Heterogeneous Graph EmbeddingabstractGraph attention network can generate effective feature embedding by specifying different weights to different nodes. The key of the research on heterogeneous graph embedding is the way to combine its rich structural information with semantic relations to aggregate the neighborhood information. Most of the existing heterogeneous graph representation learning methods guide the selection of neighbors by defining various meta-paths on heterogeneous graphs. However, these models only consider the information contained in the nodes under different paths and ignore the potential semantic relationships of nodes in different neighbor graph structures, which leads to the underutilization of graph structure information. In this article, we propose a novel adaptive framework named Neighbor Graph Aggregated Graph Attention Network (NGGAN) to fully exploit graph topological details in heterogeneous graph, and aggregates their information to obtain an effective embedding. The key idea is to use different levels of sampling methods to define neighborhood, and use neighbor graphs to represent the complex structural interaction between nodes. In this way, the high-order relationship between nodes and the latent semantics of neighbor graphs can be fully explored. Afterward, a hierarchical attention mechanism is applied to adaptively learn the importance of different objects, including node information, path information, and neighbor graph information. Multiple downstream tasks are performed on four real-world heterogeneous graph datasets, and the experimental results demonstrate the effectiveness of NGGAN. Kaibiao Lin, Jinpo Chen, Ruicong Chen, Fan Yang 0010, Lin Min, Ping Lu 0012 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | SNMCF: A Scalable Non-Negative Matrix Co-Factorization for Student Cognitive ModelingabstractStudent cognitive modeling plays an important role in the rapid development of educational data mining research. It aims to discover students' proficiency in knowledge concepts as well as to predict students' performance in conducting exercises. Studies in the past few years have been mainly centered around two types of techniques: cognitive diagnosis models and data mining approaches. Cognitive diagnosis models focus on students' cognitive states and assess their knowledge concept proficiency through handcrafted features. The subjective features may trigger cascading errors in the students' performance prediction. On the other hand, data mining techniques, e.g., matrix factorization methods, achieve high prediction accuracy by directly modeling the students' exercising process. It lacks measuring the students' knowledge concept proficiency. To address the dilemma of the aforementioned methods, in this paper, we propose a scalable non-negative matrix co-factorization (SNMCF) model by jointly modeling the students' knowledge states and their exercising process. SNMCF can achieve high accuracy in predicting students' exercise performance while modeling their states of knowledge concepts in a given domain. We conduct extensive experiments on several real-world datasets, including large sparse ones, and demonstrate the effectiveness of our new approach in terms of prediction accuracy, cognitive diagnostic ability, and scalability Shenbao Yu, Yifeng Zeng, Yinghui Pan, Fan Yang 0010 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Traffic Demand Prediction Based on Multi-dimensional Graph Convolutional NetworkabstractTraffic demand prediction is of great importance for traffic management, yet it is also a challenging problem since the traffic data are usually with complex spatial-temporal dependencies and nonlinear relationships. In this paper, to better characterize and utilize the spatial and temporal features, we propose a Multi-dimensional Graph Convolutional Network (M-GCN) to capture dynamic spatial-temporal dependence of traffic data. M-GCN captures the explicit spatio-temporal dependencies by establishing a spatial adjacency graph and a temporal adjacency graph, and captures the hidden spatiotemporal dependencies by establishing a spatial adaptive graph and a temporal adaptive graph. We leverage graph convolutional networks for complex spatial and temporal dependencies modeling and design a gated fusion module to obtain the interactive spatial-temporal dependence. Besides, an attention mechanism is applied to alleviate the error propagation problem in long-term traffic prediction. Experimental results on two real-world traffic datasets demonstrate the superiority of M-GCN. The proposed M-GCN outperforms baseline methods by up to 4.6% improvement in MAE on the dataset TaxiNYC, and the training time and predicting time of $\mathrm{M}-\mathrm{GCN}$ are reduced by nearly half. Peiying Zeng, Liying Jiang, Yongxuan Lai, Fan Yang 0010 |
IEEE Big Data | 4 |
| 2023 | A Feedback-Enhanced Two-Stage Framework for judicial machine reading comprehension
Zhiqiang Lin 0003, Fan Yang 0010, Xuyang Wu 0006, Jinsong Su |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Online computation offloading with double reinforcement learning algorithm in mobile edge computing
Linbo Liao, Yongxuan Lai, Fan Yang 0010, Wenhua Zeng |
J. Parallel Distributed Comput. | 3 |
| 2022 | Hierarchical Reinforcement Learning-Based Mobility-Aware Content Caching and Delivery Policy for Vehicle Networks
Yongxuan Lai, Fan Yang 0010 |
ICA3PP | 3 |
| 2022 | Hierarchical Multi-Modal Fusion on Dynamic Heterogeneous Graph for Health Insurance Fraud DetectionabstractIn this paper, we devote to aggregating longitudinal and multi-modal information from heterogeneous neighbors to obtain an accurate node embedding on the dynamic heterogeneous graph. Recently, Heterogeneous Graph Neural Networks (GNNs) have attracted extensive attention in fraud detection. However, when faced with longitudinal and multi-modal data such as health insurance records, existing GNN-based fraud detectors always discard the multi-modal information (e.g., medication and treatment) of heterogeneous neighbors and ignore the inconsistent claimer behavior in the longitudinal records. To fully utilize the information, we represent the records in the form of a dynamic heterogeneous graph, and propose Hierarchical Multi-modal Fusion Graph Neural Network (HMF-GNN) which learns not only topological information, but also embeddings of longitudinal and multi-modal entities to improve the performance of fraud detection. Experimental results on two real-world health insurance datasets demonstrate that HMF-GNN outperforms state-of-the-art graph embedding methods and GNN-based fraud detectors. Fan Yang 0010, Kaibiao Lin, Yongxuan Lai |
ICME | 2 |
| 2022 | Exploiting Spatial Attention and Contextual Information for Document Image Segmentation
Yuman Sang, Yifeng Zeng, Ruiying Liu, Fan Yang 0010, Zhangrui Yao, Yinghui Pan |
PAKDD (3) | 4 |
| 2022 | DeepMPM: a mortality risk prediction model using longitudinal EHR dataabstractBACKGROUND: Accurate precision approaches have far not been developed for modeling mortality risk in intensive care unit (ICU) patients. Conventional mortality risk prediction methods can hardly extract the information in longitudinal electronic medical records (EHRs) effectively, since they simply aggregate the heterogeneous variables in EHRs, ignoring the complex relationship and interactions between variables and the time dependence in longitudinal records. Recently deep learning approaches have been widely used in modeling longitudinal EHR data. However, most existing deep learning-based risk prediction approaches only use the information of a single disease, neglecting the interactions between multiple diseases and different conditions. RESULTS: In this paper, we address this unmet need by leveraging disease and treatment information in EHRs to develop a mortality risk prediction model based on deep learning (DeepMPM). DeepMPM utilizes a two-level attention mechanism, i.e. visit-level and variable-level attention, to derive the representation of patient risk status from patient's multiple longitudinal medical records. Benefiting from using EHR of patients with multiple diseases and different conditions, DeepMPM can achieve state-of-the-art performances in mortality risk prediction. CONCLUSIONS: Experiment results on MIMIC III database demonstrates that with the disease and treatment information DeepMPM can achieve a good performance in terms of Area Under ROC Curve (0.85). Moreover, DeepMPM can successfully model the complex interactions between diseases to achieve better representation learning of disease and treatment than other deep learning approaches, so as to improve the accuracy of mortality prediction. A case study also shows that DeepMPM offers the potential to provide users with insights into feature correlation in data as well as model behavior for each prediction. Fan Yang 0010, Yongxuan Lai, Quan Zou 0001 |
BMC Bioinform. | 1 |
| 2022 | Optimized Large-Scale Road Sensing Through Crowdsourced VehiclesabstractModern vehicles are gradually becoming powerful mobile sensing, communication, computing and storage platforms, which bring about the concept of vehicular urban sensing that leverages sensor nodes as an effective and affordable solution for large-scale and fine-grained sensing. And there is a trend to combine the vehicular sensing and outsourcing technologies to solve the large-scale urban road sensing problem. However, how to select appropriate participated vehicles, how to least interrupt the original routes of vehicles, and how to actively maximize the benefit of the sensing remain challenging problems. In this paper, we introduce a Crowdsourced Vehicular Sensing (CVS) framework based on more realistic assumptions of the vehicular sensing, which consists of three steps: vehicle recruitment, candidate path calculation, and path computing. We define amaximal weighted sensing paths(MWSP) problem, which is NP-Complete, in vehicular crowdsensing scenario and use heuristic methods to speed up the solving process for large-scale crowdsensing in urban road networks. The MWSP problem is formulated as a maximal satisfiability (MaxSAT) problem, and aleast-interruptedurban sensing strategy is adopted. So trips are least disrupted when conducting the sensing tasks, which would increase the drivers’ willingness to participate in urban sensing. Experiments based on real-world road-network and historical origin-destination datasets verify the effectiveness of the proposed method. The results show that the proposed algorithm outperforms other solutions and it can solve the vehicular crowdsensing problem effectively and efficiently. Yongxuan Lai, Duojian Mai, Yi Fan 0001, Fan Yang 0010 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Data-Driven Flexible Vehicle Scheduling and Route OptimizationabstractThe flexible transit service reflects a trend of demand on the flexibility and convenience in urban public transport systems, within which the vehicle scheduling and passenger insertion are two challenging issues. Especially, finding the optimal solution for a flexible transit system can be viewed as an extension of the traveling salesman problem which is NP-complete. Yet most of the existing research mainly focuses on one aspect, i.e. route planning, stop selection or vehicle scheduling, where a combined integration and optimization of the whole system is largely neglected. In this paper, we propose a data-driven flexible transit system that integrates the origin-destination insertion algorithm and the milp-based (mixed-integer linear programming) scheduling scheme. Specifically, stops are mined from the historical datasets and some stops act as$backbone$stops that should be visited by the vehicles; and a heuristic backbone-based origin-destination insertion algorithm is proposed to schedule the routing path of vehicles, where the time loss caused by the optimal insertion positions is calculated for the vehicles to decide whether to accept the requests or not when constructing a path for the flexible routes. Moreover, a vehicle scheduling model based on milp is proposed to minimise the gap between the passenger flow and available seats. The proposed flexible transit systems are simulated in real-world taxi datasets, and experimental results show that the proposed flexible transit system can effectively increase the delivery ratio and decrease the passengers’ waiting time compared with existing methods. Yongxuan Lai, Fan Yang 0010, Ge Meng, Wei Lu 0015 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Utility-Based Matching of Vehicles and Hybrid Requests on Rider Demand Responsive SystemsabstractIn rider demand responsive systems riders submit requests to demand transit services and the incoming requests and vehicles are matched by the system. This demand-responsive transport problem is viewed by existing research as a kind of spatial matching problem between vehicles and riders. However, existing schemes mainly focus on maximizing the number of matching pairs. It neglects other factors like the length of pickup trajectories and riders’ waiting time. And the matching is not revocable, which loses the chance for further optimizations. Moreover, there is still not much work on handling and matching the appointment-based requests that play a key role in the demand-responsive transport market. In this article, we propose an algorithm called BMCF (Bipartite Minimal-Cost Flow) to solve the taxi-rider matching problem with appointment-based rider requests on a time-dependent road network. Unlike existing solutions that map the problem into the max flow problem, we transform the optimal matching to aminimal cost flow problemwhich aims to maximize a utility value and could be solved efficiently. Riders and vehicles are modeled as vertices in a bipartite graph, and factors like the length of pickup trajectories, riders’ waiting time, etc. are abstracted as utilities and denoted as weights of the edges which are taken into account during the matching process. To the best of our knowledge, the proposed scheme is the first to efficiently integrate and process both the real-time and appointment-based requests within the same framework, and the assignments between vehicles and requests could be dynamically adjusted and revoked to maximise the overall utility. Experimental results show that the proposed algorithm can effectively increase the appointment-based matching ratio and decrease the riders’ waiting time (> 10.4%) and vacant vehicles’ picking up time (> 60.9%) compared with existing schemes, while at the cost of acceptable increase on the running time. Yongxuan Lai, Shipeng Yang, Anshu Xiong, Fan Yang 0010, Lei Li 0003, Xiaofang Zhou 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Shortest Path Distance Prediction Based on CatBoost
Liying Jiang, Yongxuan Lai, Wenhua Zeng, Fan Yang 0010, Yi Fan 0001 |
WISA | 5 |
| 2021 | Efficient Estimation of Time-Dependent Shortest Paths Based on Shortcuts
Linbo Liao, Shipeng Yang, Yongxuan Lai, Wenhua Zeng, Fan Yang 0010, Min Jiang 0005 |
ICA3PP (2) | 5 |
| 2021 | A Semi-supervised Defect Detection Method Based on Image Inpainting
Huibin Cao, Yongxuan Lai, Fan Yang 0010 |
PRICAI (3) | 4 |
| 2021 | Dynamic Transit Flow Graph Prediction in Spatial-Temporal Network
Liying Jiang, Yongxuan Lai, Wenhua Zeng, Fan Yang 0010, Yi Fan 0001, Qisheng Liao |
WISE (1) | 5 |
| 2021 | An Adaptive Robust Semi-Supervised Clustering Framework Using Weighted Consensus of Random $k$k-Means EnsembleabstractSemi-supervised cluster ensemble usually introduces a small amount of supervision in the first stage of cluster ensemble, i.e., ensemble generation, by performing many runs of semi-supervised clustering algorithms. However, it is neither efficient in terms of computational complexity, nor flexible in a dynamic learning environment where limited supervision changes over time. In this article we propose a new framework which generates base partitions in an unsupervised manner and attributes different weights to each cluster of the base partitions. The weighting scheme considers both the internal validation measures of clustering and the degrees of satisfaction of pairwise constraints. A weighted co-association matrix based consensus approach is then applied to achieve a final partition. To handle high-dimensional data, we generate base partitions using k-means with both random sampling and random subspace techniques. The new framework retains a high accuracy, and is efficient since it avoids performing semi-supervised clustering in ensemble generation and the complexity of the weighting scheme is independent of the number of instances in a dynamic environment. It is more adaptive than the traditional approach because it does not require rerunning semi-supervised clustering algorithms when the limited supervision changes. Empirical results on 12 datasets demonstrate that it is also more robust to noisy constraints. Yongxuan Lai, Songyao He, Fan Yang 0010, Qifeng Zhou, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Bus Travel-Time Prediction Based on Deep Spatio-Temporal Model
Yongxuan Lai, Liying Jiang, Fan Yang 0010 |
WISE (1) | 4 |
| 2019 | Grouped Correlational Generative Adversarial Networks for Discrete Electronic Health RecordsabstractUsing Generative Adversarial Networks (GANs) to generate synthetic Electronic Health Records (EHR) has attracted increasing attention. However, in existing approaches, the events in EHRs are treated as separate variables which are indiscriminately entered into the model, without taking into account the meaning and grouping of them. Besides, the efficacy of treatment is often neglected. In this paper, we first embed the efficacy information into the disease diagnosis, and then propose Grouped Correlational GAN (GcGAN) to explicitly learn inherent correlations between different groups of variables. We also introduce a dense connection to strengthen the generator capacity in GcGAN. Experimental results on real-world data demonstrate that the generated data from GcGAN are able to simulate real-world data in terms of distribution statistics. The results on multi-label treatment recommendation tasks show that GcGAN can boost the performances by augmenting the training dataset with the generated data and outperforms state-of-the-art approaches. It can also automatically distinguish between disease-specific drugs and adjuvant drugs, which enhances the model interpretability. Fan Yang 0010, Zhongping Yu, Yunfan Liang, Xiaolu Gan, Kaibiao Lin, Quan Zou 0001, Yifeng Zeng |
BIBM | 1 |
| 2019 | Privacy-aware query processing in vehicular ad-hoc networks
Yongxuan Lai, Fan Yang 0010, Wei Lu 0001 |
Ad Hoc Networks | 3 |
| 2019 | Efficient data request answering in vehicular Ad-hoc networks based on fog nodes and filters
Yongxuan Lai, Hailin Lin, Fan Yang 0010, Tian Wang 0001 |
Future Gener. Comput. Syst. | 3 |
| 2019 | CASQ: Adaptive and cloud-assisted query processing in vehicular sensor networks
Yongxuan Lai, Fan Yang 0010, Tian Wang 0001, Kuanching Li |
Future Gener. Comput. Syst. | 3 |
| 2017 | Leaf node-level ensemble pruning approaches based on node-sample correlation for random forestabstractAs a state-of-the-art ensemble method, random forest which exhibits a good ability to predict and generalize on various dataset is often composed of a large number of trees. Redundancy of ensemble and connotative decision rules result in expensive operational costs as well as difficulties in comprehension. In this paper, novel leaf node-level pruning methods for random forest are proposed. Each leaf node extracted from a random forest model is regarded as a singe classifier or classification rule, and is then evaluated for pruning. Different from traditional tree-level pruning and rule pruning methods, the idea is to evaluate and extract rules according to node-sample correlation rather than eliminate trees from the ensemble or integrate rules themselves. Experiment results show that the proposed methods can efficiently reduce the size of rule set, and the resulted rules based ensemble achieves better interpretability without significant loss of accuracy. Qifeng Zhou, Fan Yang 0010 |
IECON | 3 |
| 2017 | Taxi Route Recommendation Based on Urban Traffic Coulomb's Law
Zheng Lyu, Yongxuan Lai, Kuanching Li, Fan Yang 0010, Minghong Liao, Xing Gao 0004 |
WISE (1) | 4 |
| 2017 | Copy-move forgery detection based on hybrid features
Fan Yang 0010, Wei Lu 0001, Jian Weng 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2017 | Cluster ensemble selection with constraints
Fan Yang 0010, Tao Li 0001, Qifeng Zhou |
Neurocomputing | 1 |
| 2017 | Keypoint-based copy-move detection scheme by adopting MSCRs and improved feature matching
Fan Yang 0010, Wei Lu 0001, Wei Sun 0007 |
Multim. Tools Appl. | 2 |
| 2015 | Two approaches for novelty detection using random forest
Qifeng Zhou, Yongpeng Ning, Fan Yang 0010, Tao Li 0001 |
Expert Syst. Appl. | 4 |
| 2014 | Modeling and broadening temporal user interest in personalized news recommendation
Lei Li 0001, Li Zheng 0001, Fan Yang 0010, Tao Li 0001 |
Expert Syst. Appl. | 3 |
| 2014 | Exploring the diversity in cluster ensemble generation: Random sampling and random projection
Fan Yang 0010, Qianmu Li, Tao Li 0001 |
Expert Syst. Appl. | 1 |
| 2013 | Using the Maximum Between-Class Variance for Automatic Gridding of cDNA Microarray ImagesabstractGridding is the first and most important step to separate the spots into distinct areas in microarray image analysis. Human intervention is necessary for most gridding methods, even if some so-called fully automatic approaches also need preset parameters. The applicability of these methods is limited in certain domains and will cause variations in the gene expression results. In addition, improper gridding, which is influenced by both the misalignment and high noise level, will affect the high throughput analysis. In this paper, we have presented a fully automatic gridding technique to break through the limitation of traditional mathematical morphology gridding methods. First, a preprocessing algorithm was applied for noise reduction. Subsequently, the optimal threshold was gained by using the improved Otsu method to actually locate each spot. In order to diminish the error, the original gridding result was optimized according to the heuristic techniques by estimating the distribution of the spots. Intensive experiments on six different data sets indicate that our method is superior to the traditional morphology one and is robust in the presence of noise. More importantly, the algorithm involved in our method is simple. Furthermore, human intervention and parameters presetting are unnecessary when the algorithm is applied in different types of microarray images. Fan Yang 0010, Qian Zhang 0005, Qifeng Zhou, Linkai Luo |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | Identify content quality in online social networksabstractThe flooding of low-quality user generated contents (UGC) in online social network (OSN) has been a threat to web knowledge management systems. Recently several domain-specific systems have been developed addressing this problem, for example, predict correct answer in QA community; recognise reliable comment in products review forums etc. Major drawback of most research efforts is the lack of a general framework applicable to all OSNs. In this study, the authors start by analysing the effects of distinguishing features on UGC quality in different types of OSNs. Extensive statistical analysis leads to the discovery of existence of diverse patterns of human information sharing activity in dissimilar OSNs. This discovery is employed as prior knowledge in the classification framework, which decompose the original highly imbalanced problem into several balanced sub-problems. Ensemble classifiers are adopted in samples from clusters generated by incompact features. Experiments show the proposed framework is both effective and efficient for several OSNs.Contributions of this study are two-fold: (i) model posting activity in different types of OSNs; (ii) propose novel classification framework to identify UGC quality. Chen Lin 0001, Zhenhua Huang 0001, Fan Yang 0010, Quan Zou 0001 |
IET Commun. | 3 |
| 2012 | Margin optimization based pruning for random forest
Fan Yang 0010, Wei-hang Lu, Linkai Luo, Tao Li 0001 |
Neurocomputing | 1 |