VLDB 2026 Research / reviewers in the wild / expert
Bo Jin 0001
dblp:74/3468-1
· DBLP profile ↗
51ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0002-4094-7499ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 7 first-author · 15 since 2021Databases, data management, data science and information retrieval · 22 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SiLP: Enhancing Non-Dominant Language Capabilities with a Selective Bidirectional Language Projection FrameworkabstractCurrent large language models (LLMs) often exhibit performance imbalances between dominant languages (e.g., English) and nondominant ones due to the skewed distribution of pretraining data.A common strategy to address this issue is to enhance cross-lingual alignment, thereby facilitating non-dominant language processing.However, existing methods typically rely on additional training objectives or language-specific parameters, which increase training complexity and cost.In this work, we propose a selective bidirectional language projection framework that enables efficient multilingual alignment and language shift using the intrinsic parameters.Specifically, we first identify the layers most sensitive to language projection between non-dominant and dominant languages through neuron activation analysis.We then perform sequential language projection within the selected layers by mapping non-dominant representations into the dominant language space and reverting them before generation.The bidirectional projection benefits the subsequent instruction tuning in non-dominant languages.Experiments on seven benchmarks demonstrate that our method remarkably enhances the performance of nondominant languages.Further analyses indicate that our method learns better internal representations and exhibits strong generalization capabilities. Junpeng Liu 0002, Jiuyi Li, Bo Jin 0001, Degen Huang, Hui Xiong 0001 |
ACL (1) | 4 |
| 2026 | Hierarchical prototype-guided representation learning for robust graph classification
Liang Zhang 0031, Kongyu Chen, Bo Jin 0001, Xiaopeng Wei |
Inf. Sci. | 3 |
| 2026 | Multi-Condition Latent Diffusion Network for Semantic-aware Knowledge Graph Completion
Liang Zhang 0031, Bo Jin 0001, Xiaopeng Wei |
Knowl. Based Syst. | 3 |
| 2025 | Transformer-Based Multi-Agent Reinforcement Learning Method With Credit-Oriented Strategy DifferentiationabstractThe problem of Multi-Agent Reinforcement Learning (MARL) shows a high level of both complexity in the environment and coordination between agents. In order to scale the algorithm to large-scale agent scenarios, neural networks designed for MARL are typically implemented with parameter sharing. These characteristics result in the challenges of partial observability, credit assignment and strategy homogenization. In this paper, a Transformer-Based Multi-Agent Reinforcement Learning Method With Credit-Oriented Strategy Differentiation (TMRC) is presented to address each of these challenges. First, we design a Temporal-Spatial Encoding module and an Attention-Based Value Decomposition module based on the Transformer architecture. The former leverages both temporal and spatial observation information, compensating for the missing environmental perspectives due to partial observability. The latter is designed to identify each agent’s individual contribution in complex interactions, effectively optimizing the credit assignment process. Then, we propose a Credit-Oriented Strategy Differentiation module that differentiates the entity representations of each agent based on their current task differences, allowing agents to have distinct real-time strategies, effectively mitigating the issue of strategy homogenization. We evaluate the proposed method on the SMAC benchmark. It demonstrates better final performance, faster convergence, and greater stability compared to other comparative methods. Additionally, a series of experiments are conducted to validate the effectiveness of the proposed modules. Our code is available at https://github.com/Hkxuan/TMRC.git. Kaixuan Huang, Bo Jin 0001, Haiyin Piao, Ziqi Wei 0001 |
IROS | 2 |
| 2025 | MGTNSyn: Molecular structure-aware graph transformer network with relational attention for drug synergy prediction
Yunjiong Liu, Peiliang Zhang, Chao Che, Bo Jin 0001 |
Expert Syst. Appl. | 5 |
| 2025 | MGHSTCKW: Predicting miRNA-drug sensitivity association using hypergraph sparse transformer and hypergraph-induced contrastive learning based on meta-path
Dong Ouyang, Bo Jin 0001, Pingtao Duan, Xiongfeng Zhu, Shaocan Fan, Rui Miao 0002, Ning Ai |
Expert Syst. Appl. | 2 |
| 2025 | Self-Supervised Disentangled Representation Learning for Time Series Anomaly DetectionabstractAnomaly detection is a fundamental component of intelligent monitoring in the Internet of Things (IoT), where accuracy, efficiency, and interpretability are critical requirements. However, existing methods often overlook the unique characteristics of IoT signals such as seasonality, trends, and irregular residual components, as well as the complex interactions among them. This oversight can lead to anomaly masking, increased false positives, and reduced interpretability in anomaly identification. Motivated by the effectiveness of disentangled representation learning, we propose TRAdetector, a novel disentangled reconstruction-based framework for IoT signals anomaly detection. TRAdetector explicitly models recurrent and consistent patterns, as well as irregular variations in the latent space by leveraging variational inference strategies, thereby enhancing probabilistic guidance in learning both regular and irregular temporal representations. A sparse coding strategy is incorporated within the latent space of the residual component to directly model inconsistent temporal fluctuations. Finally, a multihead cross-attention mechanism and a gated, decomposition-aware reconstruction strategy are designed to effectively model the complex interactions among different components. Extensive experiments show that our model achieves state-of-the-art performance on multiple benchmark datasets in terms of accuracy, efficiency, and interpretability. Liang Zhang 0031, Jianping Zhu 0002, Guangjie Han, Bo Jin 0001, Pengfei Wang 0013, Xiaopeng Wei |
IEEE Internet Things J. | 4 |
| 2024 | Explainable Origin-Destination Crowd Flow Interpolation via Variational Multi-Modal Recurrent Graph Auto-EncoderabstractOrigin-destination (OD) crowd flow, if more accurately inferred at a fine-grained level, has the potential to enhance the efficacy of various urban applications. While in practice for mining OD crowd flow with effect, the problem of spatially interpolating OD crowd flow occurs since the ineluctable missing values. This problem is further complicated by the inherent scarcity and noise nature of OD crowd flow data. In this paper, we propose an uncertainty-aware interpolative and explainable framework, namely UApex, for realizing reliable and trustworthy OD crowd flow interpolation. Specifically, we first design a Variational Multi-modal Recurrent Graph Auto-Encoder (VMR-GAE) for uncertainty-aware OD crowd flow interpolation. A key idea here is to formulate the problem as semi-supervised learning on directed graphs. Next, to mitigate the data scarcity, we incorporate a distribution alignment mechanism that can introduce supplementary modals into variational inference. Then, a dedicated decoder with a Poisson prior is proposed for OD crowd flow interpolation. Moreover, to make VMR-GAE more trustworthy, we develop an efficient and uncertainty-aware explainer that can provide explanations from the spatiotemporal topology perspective via the Shapley value. Extensive experiments on two real-world datasets validate that VMR-GAE outperforms the state-of-the-art baselines. Also, an exploratory empirical study shows that the proposed explainer can generate meaningful spatiotemporal explanations. Xinjiang Lu, Jingjing Gu, Bo Jin 0001 |
AAAI | 5 |
| 2024 | Adaptive Meta-Learning Probabilistic Inference Framework for Long Sequence PredictionabstractLong sequence prediction has broad and significant application value in fields such as finance, wind power, and weather. However, the complex long-term dependencies of long sequence data and the potential domain shift problems limit the effectiveness of traditional models in practical scenarios. To this end, we propose an Adaptive Meta-Learning Probabilistic Inference Framework (AMPIF) based on sequence decomposition, which can effectively enhance the long sequence prediction ability of various basic models. Specifically, first, we decouple complex sequences into seasonal and trend components through a frequency domain decomposition module. Then, we design an adaptive meta-learning task construction strategy, which divides the seasonal and trend components into different tasks through a clustering-matching approach. Finally, we design a dual-stream amortized network (ST-DAN) to capture shared information between seasonal-trend tasks and use the support set to generate task-specific parameters for rapid generalization learning on the query set. We conducted extensive experiments on six datasets, including wind power and finance scenarios, and the results show that our method significantly outperforms baseline methods in prediction accuracy, interpretability, and algorithm stability and can effectively enhance the long sequence prediction capabilities of base models. The source code is publicly available at https://github.com/Zhu-JP/AMPIF. Jianping Zhu 0002, Bo Jin 0001 |
AAAI | 6 |
| 2024 | FreAML: A Frequency-Domain Adaptive Meta-Learning Framework for EEG-Based Emotion RecognitionabstractEmotion recognition technology, especially methods based on electroencephalogram (EEG) signals, plays a critical role in revealing deep human emotions and enhancing human-computer interaction experiences. However, the high-frequency characteristics of EEG signals lead to significant individual differences, causing domain distribution shift issues. To address this, we propose a frequency-domain adaptive meta-learning probabilistic inference framework, named FreAML, leveraging the advantages of frequency-domain signals in preserving comprehensive views and learning global dependencies. Specifically, we first introduce an adaptive meta-task construction strategy based on a clustering-matching pattern. Then, through a dual-stream amortization network, we learn task-specific frequency-domain real and imaginary parameters from the support set. Finally, using these task-specific parameters, we achieve rapid generalization learning of the query set through a frequency-domain linear mapping. Extensive experiments on two public datasets validate the effectiveness and superiority of the proposed method. Lei Wang 0196, Jianping Zhu 0002, Le Du, Bo Jin 0001, Xiaopeng Wei |
BIBM | 4 |
| 2024 | Multi-view Time-frequency Contrastive Learning for Emotion Recognition
Lei Wang 0196, Jianping Zhu 0002, Bo Jin 0001, Xiaopeng Wei |
CogSci | 3 |
| 2024 | NCH-DDA: Neighborhood contrastive learning heterogeneous network for drug-disease association predictionabstractExploring new therapeutic diseases for existing drugs plays an essential role in reducing drug development costs. However, existing methods for predicting drug–disease association (DDA) lack fusion to multi-neighborhood information, which limits their ability to generalize and forces them to rely on prior knowledge. To this end, we propose a novel DDA model called the Neighborhood Contrastive Learning Heterogeneous Networks (NCH-DDA). NCH-DDA uses both single-neighborhood and multi-neighborhood feature extraction modules to extract important features of drugs and diseases in parallel from multiple potential spaces, such as heterogeneous networks and similarity networks. NCH-DDA fuses single-neighborhood and multi-neighborhood features using contrastive learning to enhance information interaction in different neighborhood spaces, ultimately obtaining universal domain features of drugs and diseases. NCH-DDA uses a combination of predictive loss and triplet loss to reduce dependence on prior knowledge. In different partition schemes of multiple datasets, NCH-DDA achieved the best performance in predicting DDA, outperforming several current state-of-the-art methods. Moreover, NCH-DDA demonstrated better performance in experiments on data sparsity and drug repositioning for Alzheimer’s disease, indicating its greater potential in DDA prediction with sparse omics data and drug repositioning applications. Peiliang Zhang, Chao Che, Bo Jin 0001, Jingling Yuan, Yongjun Zhu 0001 |
Expert Syst. Appl. | 3 |
| 2023 | Adaptive Bayesian Meta-Learning for EEG Signal ClassificationabstractAccurate classification of electroencephalogram (EEG) signals is crucial for brain activity understanding. However, EEG signals are characterized by data heterogeneity and label scarcity, which present a challenging low-data learning regime when building machine learning models. Existing methods tend to suffer from overfitting problem. To this end, we propose an adaptive Bayesian meta-learning framework for instance-specific learning and inference in EEG classification tasks. Specifically, first, a query set-driven dynamic parameter-based support set selection strategy is designed to adaptively fit the query set when constructing a meta-training task. Second, we employ an amortized variational inference network to generate task-specific adapted parameters given the support set, thereby achieving rapid model adaption for the inference of the data in the query set. Especially, a time- and frequency-aware representation learning encoder is leveraged to extract more task-relevant information guided by information bottleneck principle from time and frequency views, respectively, alleviating the low signal-to-noise ratio issue. Extensive experimental results on three public datasets demonstrate the superior effectiveness of our method. Jianping Zhu 0002, Liang Zhang 0031, Bo Jin 0001, Xiaopeng Wei |
BIBM | 4 |
| 2023 | Towards Long-Term Time-Series Forecasting: Feature, Pattern, and DistributionabstractLong-term time-series forecasting (LTTF) has become a pressing demand in many applications, such as wind power supply planning. Transformer models have been adopted to deliver high prediction capacity because of the high computational self-attention mechanism. Though one could lower the complexity of Transformers by inducing the sparsity in point-wise self-attentions for LTTF, the limited information utilization prohibits the model from exploring the complex dependencies comprehensively. To this end, we propose an efficient Transformer-based model, named Conformer, which differentiates itself from existing methods for LTTF in three aspects: (i) an encoder-decoder architecture incorporating a linear complexity without sacrificing information utilization is proposed on top of sliding-window attention and Stationary and Instant Recurrent Network (SIRN); (ii) a module derived from the normalizing flow is devised to further improve the information utilization by inferring the outputs with the latent variables in SIRN directly; (iii) the inter-series correlation and temporal dynamics in time-series data are modeled explicitly to fuel the downstream self-attention mechanism. Extensive experiments on seven real-world datasets demonstrate that Conformer outperforms the state-of-the-art methods on LTTF and generates reliable prediction results with uncertainty quantification. Xinjiang Lu, Haoyi Xiong, Jiantao Su, Bo Jin 0001, Dejing Dou |
ICDE | 6 |
| 2023 | A Global View-Guided Autoregressive Residual Network for Irregular Time Series Classification
Jianping Zhu 0002, Haocheng Tang, Liang Zhang 0031, Bo Jin 0001, Xiaopeng Wei |
PAKDD (4) | 4 |
| 2023 | CariesFG: A fine-grained RGB image classification framework with attention mechanism for dental caries
Hao Jiang 0052, Peiliang Zhang, Chao Che, Bo Jin 0001, Yongjun Zhu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | IEA-GNN: Anchor-aware graph neural network fused with information entropy for node classification and link prediction
Peiliang Zhang, Jiatao Chen, Chao Che, Liang Zhang 0031, Bo Jin 0001, Yongjun Zhu 0001 |
Inf. Sci. | 5 |
| 2023 | Predicting Drug-Target Interaction Via Self-Supervised LearningabstractRecent advances in graph representation learning provide new opportunities for computational drug-target interaction (DTI) prediction. However, it still suffers from deficiencies of dependence on manual labels and vulnerability to attacks. Inspired by the success of self-supervised learning (SSL) algorithms, which can leverage input data itself as supervision,we propose SupDTI, a SSL-enhanced drug-target interaction prediction framework based on a heterogeneous network (i.e., drug-protein, drug-drug, and protein-protein interaction network; drug-disease, drug-side-effect, and protein-disease association network; drug-structure and protein-sequence similarity network). Specifically, SupDTI is an end-to-end learning framework consisting of five components. First, localized and globalized graph convolutions are designed to capture the nodes' information from both local and global perspectives, respectively. Then, we develop a variational autoencoder to constrain the nodes' representation to have desired statistical characteristics. Finally, a unified self-supervised learning strategy is leveraged to enhance the nodes' representation, namely, a contrastive learning module is employed to enable the nodes' representation to fit the graph-level representation, followed by a generative learning module which further maximizes the node-level agreement across the global and local views by learning the probabilistic connectivity distribution of the original heterogeneous network. Experimental results show that our model can achieve better prediction performance than state-of-the-art methods. Jiatao Chen, Liang Zhang 0031, Ke Cheng 0003, Bo Jin 0001, Xinjiang Lu, Chao Che |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Diagnostic Sparse Connectivity Networks With Regularization TemplateabstractDynamic systems are often monitored with multivariate time series where each dimension represents a local component measured through a (virtual) sensor. Performing accurate diagnostic for dynamic systems while simultaneously taking into account their similarities/distinctions, is a non-trivial task. To this end, we develop an adaptive regularization approach to learning sparse connectivity structures in complex dynamic systems. The learned connectivity networks shed lights on the structural compositions of the system and hence can serve as highly informative inputs for various machine learning tasks such as classification. In particular, we focus on high-dimensional and semi-supervised learning scenarios and present a joint learning approach to recover system-wise connectivity patterns by adaptively constructing a shared, sparsity-inducing regularization template across all systems. The shared template can be physically interpreted and used as a modeling template for analyzing new systems. Moreover, our approach has the flexibility to incorporate supervising information such as must-links and cannot-links for constructing regularization templates. Overall, our approach, named sparse adaptive regularization (SAR), can extract structure-related connectivity features efficiently and effectively, and result in significant improvements for machine learning tasks in dynamic systems. We benchmark our approach against the state-of-the-art methods with real-world data. Our results demonstrate the superiority of our approach. Chuanren Liu, Kai Zhang 0001, Keli Xiao, Bo Jin 0001, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Bi-graph attention network for aspect category sentiment classification
Yongxue Shan, Chao Che, Xiaopeng Wei, Yongjun Zhu 0001, Bo Jin 0001 |
Knowl. Based Syst. | 6 |
| 2022 | Domain-Aware Word Segmentation for Chinese Language: A Document-Level Context-Aware ModelabstractWord segmentation is an essential and challenging task in natural language processing, especially for the Chinese language due to its high linguistic complexity. Existing methods for Chinese word segmentation, including statistical machine learning methods and neural network methods, usually have good performance in specific knowledge domains. Given the increasing importance of interdisciplinary and cross-domain studies, one of the challenges in cross-domain word segmentation is to handle the out-of-vocabulary (OOV) words. Existing methods show unsatisfactory performance to meet the practical standard. To this end, we propose a document-level context-aware model that can automatically perceive and identify OOV words from different domains. Our method jointly implements a word-based and a character-based model and then processes the results with a newly proposed reconstruction model. We evaluate the new method by designing and conducting comprehensive experiments on two real-world datasets (e.g., news from different domains). The results demonstrate the superiority of our method over the state-of-the-art models in handling texts from different domains. Importantly, when doing the word segmentation under the cross-domain scenario, our proposed method can improve the performance of OOV words recognition. Keli Xiao, Fengran Mo, Bo Jin 0001, Zhuang Liu 0001, Degen Huang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Prediction of Treatment Medicines With Dual Adaptive Sequential NetworksabstractPredicting treatment medicines is a key task in many intelligent healthcare systems. Prediction of treatment medicines can assist doctors in making informed prescription decisions for patients according to their Electronic Health Records (EHRs). However, predicting treatment medicines is a challenging task due to the following reasons: (1) heterogeneous nature of EHR data that typically includes laboratory results, treatment records, disease conditions, and demographic information; (2) complex correlations among EHR sequences, including inter-correlations between sequences and temporal intra-correlations within each sequence; (3) temporal dynamics of these correlations changing with disease progression. In this paper, we predict treatment medicines for patients with dual adaptive sequential networks (DASNet). Specifically, DASNet is designed with three components. First, a decomposed adaptive long short-term memory network (DA-LSTM) is designed to capture the intra- and inter-correlations in multiple heterogeneous temporal sequences. Then, we develop an attentive meta learning network (AT-MetaNet) to learn dynamic weight parameters for DA-LSTM, thus enabling it to model various correlation structures. Finally, we employ an attentive fusion network (AT-FuNet) to incorporate historical information and collectively fuse representation embeddings of heterogeneous data to predict treatment medicines. Our results on the public MIMIC-III dataset covering 11 medical conditions demonstrate that the proposed end-to-end model can achieve the state-of-the-art prediction performance while providing clinically useful insights. Liang Zhang 0031, Leilei Sun, Bo Jin 0001, Chuanren Liu, Ruiyun Yu, Xiaopeng Wei |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | CFFNN: Cross Feature Fusion Neural Network for Collaborative FilteringabstractNumerous state-of-the-art recommendation frameworks employ deep neural networks in Collaborative Filtering (CF). In this paper, we propose a cross feature fusion neural network (CFFNN) for the enhancement of CF. Existing studies overlook either user preferences for various item features or the relationship between item features and user features. To solve this problem, we construct a cross feature fusion network to enable the fusion of user features and item features as well as a self-attention network to determine users’ preferences for items. Specifically, we design a feature extraction layer with multiple MLP (Multilayer Perceptrons) modules to extract both user features and item features. Then, we introduce a cross feature fusion mechanism for an accurate determination of the relationship between different user-item interactions. The features of users and items are crossly embedded and then fed into a prediction network. The attention mechanism enables the model to focus on more effective features. The effectiveness of CFFNN model is demonstrated through extensive experiments on four real-world datasets. The experimental results indicate that CFFNN significantly outperforms the existing state-of-the-art models, with a relative improvement of 3.0 to 12.1 percent on hit ratio (HR) and normalized discounted cumulative gain (NDCG) compared with the baselines. Ruiyun Yu, Dezhi Ye, Biyun Zhang, Ann Move Oguti, Jie Li 0008, Bo Jin 0001, Fadi J. Kurdahi |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2021 | Aspect-Level Sentiment Classification of Chinese Patient Comments Based on Pre-trained Sentiment EmbeddingabstractWith the development of information technology, online health care service platforms have collected a large amount of patient comment information. Through fine-grained sentiment classification of this information, we can provide references for patients to seek medical treatment and help doctors understand their work. Therefore, we proposed a model that integrated pre-trained emotional information and semantic information at the word and character level for aspect-level sentiment classification. Specifically, we employed adversarial learning for training sentiment word embeddings and the two-layer bidirectional long short-term memory network to extract the sentiment embedding of the entire sentence in a specific aspect. We also combined the pre-trained sentiment feature vector with the structured semantic information by linear weighting and the multi-head self-attention mechanism, enabling the model to pay more attention to the information most relevant to a given aspect category. We performed experiments on the Chinese patient comments data set constructed by our research team and the proposed model outperformed the state-of-the-art methods, which proved the effectiveness of the proposed model for aspect-level sentiment classification of Chinese patient comments. Yongxue Shan, Zhaoqian Zhong, Chao Che, Bo Jin 0001, Xiaopeng Wei |
BIBM | 4 |
| 2021 | A Multi-view Confidence-calibrated Framework for Fair and Stable Graph Representation LearningabstractGraph Neural Networks (GNNs) are prone to adversarial attacks and discriminatory biases. The cutting-edge studies usually adopt a perturbation-invariant consistency regularization strategy without considering the inherent prediction uncertainties, which can lead to unsatisfactory overconfidence for incorrect prediction under intent graph topology or node features attacks. Besides, operating on the complete graph structure is biased towards global level graph noise and brings severe computational issues. In this work, we develop a multi-view confidence-calibrated framework, called MCCNIFTY, for unified fair and stable graph representation learning. At its core is a multi-view uncertainty-aware node embedding learning module derived from evidential theory, including an intra-view evidence calibration, an inter-view evidence fusion, and an uncertainty-aware message passing process in a GNN architecture, which simultaneously optimizes for counterfactual fairness and stability at the sub-graph level. Experimental results on three real-world datasets demonstrate that our method is capable of adequately capturing inherent uncertainties while improving the fairness and stability via subgraph-induced multiview confidence calibration. Xu Zhang 0026, Liang Zhang 0031, Bo Jin 0001, Xinjiang Lu |
ICDM | 3 |
| 2021 | EduHawkes: A Neural Hawkes Process Approach for Online Study Behavior ModelingabstractThe COVID-19 pandemic forces schools to move teaching online and stimulates the development of online tutoring platforms.Although online tutoring platforms provide students the access to learning materials and tools anytime and anywhere, the quality of studies is impeded by the fact that students learn by watching videos, which lacks interactions between teachers and students.Such dilemma prevents us from respectively understanding and improving the online learning patterns and efficiency of students.To achieve this goal, we need to solve three challenges: (1) How can we quantify the study quality of online learning?(2) How can we design an appropriate data structure to describe online study behaviors?(3) How can we model the online study behaviors to better mine online study patterns?To address the challenges, we first propose a new measurement to quantify the online study quality from the perspective of study engagement.We then define a study behavior sequence to describe online study behaviors.The study behavior at each timestamp is an event of a video lecture watching behavior type, such as, watching, dragging forward and dragging backward.Moreover, we develop a neural hawkes process framework (namely EduHawkes ) for online study behavior modeling.The EduHawkes is a novel hierarchical encode-decode architecture with simultaneously optimizing the study behavior prediction task (event-level) and the study quality prediction task (course-level).In the experiments, we apply EduHawkes to the applications of study quality prediction and flippant student identification in order to demonstrate the improved performances of our proposed method on modeling online study behaviors. Lu Jiang 0007, Pengyang Wang, Ke Cheng 0003, Kunpeng Liu 0001, Minghao Yin, Bo Jin 0001, Yanjie Fu |
SDM | 6 |
| 2021 | Completely blind image quality assessment via contourlet energy statisticsabstractAbstract An aim of completely blind image quality assessment (BIQA) is to develop algorithms which can grade image quality without any prior knowledge of the images. Here, a new contourlet energy statistics based completely on blind opinion‐unaware BIQA (OU‐BIQA) method is proposed, which can predict the perceptual severity of a range of image distortion types without requiring any prior knowledge. According to the energy distribution of the contourlet sub‐bands of natural images in log‐domain, the lower‐scale sub‐band energy can be predicted by the corresponding higher‐scale sub‐band energies of distorted images. A quality model is then constructed by quantifying the difference between predicted energy and realistic energy. Meanwhile, an effective method for adjusting and compensating an undesired distortion is integrated into the quality model. Experimental results show that the proposed new method outperforms state‐of‐the‐art OU‐BIQA models on relevant portions of TID2013 database, and is competitive on the LIVE IQA database. Moreover, the proposed model is very fast, suggesting a real‐time solution to high‐performance BIQA. Tuxin Guan, Yuhui Zheng, Bo Jin 0001, Xiaojun Wu 0001, Alan C. Bovik |
IET Image Process. | 4 |
| 2021 | Bi-DAINet: Bi-Directional Discard-Accept-Integrate Network for salient object detection
Cuili Yao, Lin Feng 0001, Yuqiu Kong, Bo Jin 0001, Leheng Li |
Neurocomputing | 4 |
| 2021 | MeSIN: Multilevel selective and interactive network for medication recommendation
Liang Zhang 0031, Mao You, Xueqing Tian, Bo Jin 0001, Xiaopeng Wei |
Knowl. Based Syst. | 5 |
| 2020 | Exploring Multi-level Mutual Information for Drug-target Interaction PredictionabstractRecent advances in graph representation learning provide new opportunities for computational drug-target interaction (DTI) prediction. Inspired by the emerging graph mutual information-based algorithms, we propose MMIDTI, a multi-level mutual information-aware DTI prediction framework based on a heterogeneous network (i.e., drug-protein, drug-drug and protein-protein interaction network; drug-disease, drug-side-effect, and protein-disease association network; drug-structure and protein-sequence similarity network). More specifically, MMIDTI leverages an encoder-decoder framework that can learn the type-aware and meta-path augmented node representations by following a contrastive learning paradigm. The encoder part is a Graph Convolutional Network (GCN) and the decoder is an inner product of the learned representations to recover the original heterogeneous network. Meanwhile, MMIDTI exploits two levels of mutual information: (1) maximizing local mutual information, to obtain node representations that capture the global information content of the entire heterogeneous graph. (2) maximizing the global mutual information, to constrain the node representation to have desired statistical characteristics. Experimental results show that our model can achieve better prediction performance than state-of-the-art methods. Jiatao Chen, Liang Zhang 0031, Ke Cheng 0003, Bo Jin 0001, Xinjiang Lu, Chao Che |
BIBM | 4 |
| 2020 | Partial Relationship Aware Influence Diffusion via a Multi-channel Encoding Scheme for Social RecommendationabstractSocial recommendation tasks exploit social connections to enhance recommendation performance. To fully utilize each user's first-order and high-order neighborhood preferences, recent approaches incorporate influence diffusion process for better user preference modeling. Despite the superior performance of these models, they either neglect the latent individual interests hidden in the user-item interactions or rely on computationally expensive graph attention models to uncover the item-induced sub-relations, which essentially determine the influence propagation passages. Considering the sparse substructures are derived from original social network, we name them as partial relationships between users. We argue such relationships can be directly modeled such that both personal interests and shared interests can propagate along a few channels (or dimensions) of latent users' embeddings. To this end, we propose a partial relationship aware influence diffusion structure via a computationally efficient multi-channel encoding scheme. Specifically, the encoding scheme first simplifies graph attention operation based on a channel-wise sparsity assumption, and then adds an InfluenceNorm function to maintain such sparsity. Moreover, ChannelNorm is designed to alleviate the oversmoothing problem in graph neural network models. Extensive experiments on two benchmark datasets show that our method is comparable to state-of-the-art graph attention-based social recommendation models while capturing user interests according to partial relationships more efficiently. Bo Jin 0001, Ke Cheng 0003, Liang Zhang 0031, Yanjie Fu, Minghao Yin, Lu Jiang 0007 |
CIKM | 1 |
| 2020 | Fast Sparse Connectivity Network Adaption via Meta-LearningabstractPartial correlation-based connectivity networks can describe the direct connectivity between features while avoiding spurious effects, and hence they can be implemented in diagnosing complex dynamic multivariate systems. However, existing studies mainly focus on single systems that are ill-equipped for incremental learning. Moreover, related methods estimate temporal connectivity network by imposing only sparse regularization without integrating pattern priors (e.g., inter-system shared pattern and intra-system intrinsic pattern), which have been proven effective in limiting noise interference. To this end, we develop an adaptive connectivity estimation model that incorporates prior patterns, namely Sparse Adaptive Meta-Learning Connectivity Network (SAMCN). Specifically, our model extends ideas of the gradient-based meta-learning to capture inter-system shared prior information by generating fast adaptive initialization parameters for the connectivity matrix. Then, a sparse variational autoencoder is proposed to generate a weight matrix for sparse regularization penalty in reweighted LASSO, which helps extract intra-system intrinsic patterns (local manifold structure). Experimental results on both synthetic data and real-world datasets demonstrate that our method is capable of adequately capturing the aforementioned pattern priors. Further, experiments from corresponding classification tasks validate the strength of the prior pattern-aware features connectivity network in resulting in better classification performance. Bo Jin 0001, Ke Cheng 0003, Liang Zhang 0031, Keli Xiao, Xinjiang Lu, Xiaopeng Wei |
ICDM | 1 |
| 2020 | RAHM: Relation augmented hierarchical multi-task learning framework for reasonable medication stocking
Yakun Mao, Liang Zhang 0031, Bo Jin 0001, Keli Xiao, Xiaopeng Wei, Jun Yan 0010 |
J. Biomed. Informatics | 4 |
| 2020 | Multi-view reconstructive preserving embedding for dimension reduction
Huibing Wang, Lin Feng 0001, Adong Kong, Bo Jin 0001 |
Soft Comput. | 4 |
| 2020 | Unified Generative Adversarial Networks for Multiple-Choice Oriented Machine ComprehensionabstractIn this article, we address the multiple-choice machine comprehension (MC) problem in natural language processing. Existing approaches for MC are usually designed for general cases; however, we specially develop a novel method for solving the multiple-choice MC problem. We take the inspiration generative adversarial networks (GANs) and first propose an adversarial framework for multiple-choice oriented MC, named McGAN . Specifically, our approach is designed as a GAN-based method that unifies both generative and discriminative MC models. Working together, the generative model focuses on predicting relevant answer given a passage (text) and a question; the discriminative model focuses on predicting their relevancy given an answer-passage-question set. Based on the competition via adversarial training in a minimize-maximize game, the proposed method takes advantages from both models. To evaluate the performance, we test our McGAN model on three well-known datasets for multiple-choice MC. Our results show that McGAN can achieve a significant increase in accuracy compared to existing models based on all three datasets, and it consistently outperforms all tested baselines, including state-of-the-art techniques. Zhuang Liu 0001, Keli Xiao, Bo Jin 0001, Degen Huang, Yunxia Zhang |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2019 | A Parallel Simulated Annealing Enhancement of the Optimal-Matching Heuristic for RidesharingabstractIn this paper, we develop an efficient parallel heuristic method to solve the global optimization problem associated with the ridesharing system. Based on the carefully formalized problem and objective function, we fully utilize the heuristic characteristics of the algorithm for handling the real-life constraints in ridesharing. Following the principles of simulated annealing, our method is adaptive in handling the matching and route optimization tasks. We develop an efficient parallel scheme with simulated annealing, named PCSA, for solving the global optimization problem for ridesharing. Our algorithm is capable of efficiently addressing the potential of ridesharing by exploiting the mobility information of the ride requests. Based on extensive experiments on large real-world data, we validate the performance of our parallel heuristic algorithm. Our results confirm the effectiveness and efficiency of the proposed method and its superiority over all other benchmarks. Zeyang Ye, Keli Xiao, Bo Jin 0001 |
ICDM | 4 |
| 2019 | Unsupervised EEG feature extraction based on echo state network
Leilei Sun, Bo Jin 0001, Jianing Tong, Chuanren Liu, Hui Xiong 0001 |
Inf. Sci. | 2 |
| 2019 | Learning a Distance Metric by Balancing KL-Divergence for Imbalanced DatasetsabstractIn many real-world domains, datasets with imbalanced class distributions occur frequently, which may confuse various machine learning tasks. Among all these tasks, learning classifiers from imbalanced datasets is an important topic. To perform this task well, it is crucial to train a distance metric which can accurately measure similarities between samples from imbalanced datasets. Unfortunately, existing distance metric methods, such as large margin nearest neighbor, information-theoretic metric learning, etc., care more about distances between samples and fail to take imbalanced class distributions into consideration. Traditional distance metrics have natural tendencies to favor the majority classes, which can more easily satisfy their objective function. Those important minority classes are always neglected during the construction process of distance metrics, which severely affects the decision system of most classifiers. Therefore, how to learn an appropriate distance metric which can deal with imbalanced datasets is of vital importance, but challenging. In order to solve this problem, this paper proposes a novel distance metric learning method named distance metric by balancing KL-divergence (DMBK). DMBK defines normalized divergences using KL-divergence to describe distinctions between different classes. Then it combines geometric mean with normalized divergences and separates samples from different classes simultaneously. This procedure separates all classes in a balanced way and avoids inaccurate similarities incurred by imbalanced class distributions. Various experiments on imbalanced datasets have verified the excellent performance of our novel method. Lin Feng 0001, Huibing Wang, Bo Jin 0001, Haohao Li, Mingliang Xue |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2018 | Dr. Right!: Embedding-Based Adaptively-Weighted Mixture Multi-classification Model for Finding Right Doctors with Healthcare Experience DataabstractFinding a right doctor with suitable expertise that meets one's health needs is important yet challenging. In this paper, we study the problem of finding high-rated doctors for a specific disease using imbalanced and heterogeneous healthcare experience rating data. We develop a data analytical framework, namely Dr. Right!, which incorporates the so-called network-textual embeddings, together with data-imbalance-aware mixture multi-classification models to rate doctors per specific disease. First, Dr. Right! collects the comments and rating records from patients for doctors on specific diseases from an online hospital and constructs a doctor-patient-disease network, where every edge weight is a pairwise average rating (experience score) among doctors, patients, and diseases. Then, Dr. Right! learns the embeddings of patient experiences from textual comments using the Word2Vec, as well as the embeddings of doctors and diseases from the doctor-patient-disease network via the Node2Vec. The two types of embeddings are fused to represent a doctor-patient pair. With the embedding representations of doctor-patient pairs, Dr. Right! learns an adaptively-weighted mixture multi-classification model to map a doctor-disease pair to an experience rating score, while addressing the challenges of data imbalance and group heterogeneity. Finally, extensive experimental results demonstrate the enhanced performances of Dr. Right! for predicting the disease-specific experience scores of doctors. Yanjie Fu, Haoyi Xiong, Bo Jin 0001, Shuli Hu, Minghao Yin |
ICDM | 4 |
| 2018 | CADEN: A Context-Aware Deep Embedding Network for Financial Opinions MiningabstractFollowing the recent advances of artificial intelligence, financial text mining has gained new potential to benefit theoretical research with practice impacts. An essential research question for financial text mining is how to accurately identify the actual financial opinions (e.g., bullish or bearish) behind words in plain text. Traditional methods mainly consider this task as a text classification problem with solutions based on machine learning algorithms. However, most of them rely heavily on the hand-crafted features extracted from the text. Indeed, a critical issue along this line is that the latent global and local contexts of the financial opinions usually cannot be fully captured. To this end, we propose a context-aware deep embedding network for financial text mining, named CADEN, by jointly encoding the global and local contextual information. Especially, we capture and include an attitude-aware user embedding to enhance the performance of our model. We validate our method with extensive experiments based on a real-world dataset and several state-of-the-art baselines for investor sentiment recognition. Our results show a consistently superior performance of our approach for identifying the financial opinions from texts of different formats. Liang Zhang 0031, Keli Xiao, Hengshu Zhu, Chuanren Liu, Jingyuan Yang 0001, Bo Jin 0001 |
ICDM | 6 |
| 2018 | A Treatment Engine by Predicting Next-Period PrescriptionsabstractRecent years have witnessed an opportunity for improving healthcare efficiency and quality by mining Electronic Medical Records (EMRs). This paper is aimed at developing a treatment engine, which learns from historical EMR data and provides a patient with next-period prescriptions based on disease conditions, laboratory results, and treatment records of the patient. Importantly, the engine takes consideration of both treatment records and physical examination sequences which are not only heterogeneous and temporal in nature but also often with different record frequencies and lengths. Moreover, the engine also combines static information (e.g., demographics) with the temporal sequences to provide personalized treatment prescriptions to patients. In this regard, a novel Long Short-Term Memory (LSTM) learning framework is proposed to model inter-correlations of different types of medical sequences by connections between hidden neurons. With this framework, we develop three multifaceted LSTM models: Fully Connected Heterogeneous LSTM, Partially Connected Heterogeneous LSTM, and Decomposed Heterogeneous LSTM. The experiments are conducted on two datasets: one is the public MIMIC-III ICU data, and the other comes from several Chinese hospitals. Experimental results reveal the effectiveness of the framework and the three models. The work is deemed important and meaningful for both academia and practitioners in the realm of medical treatment and prediction, as well as in other fields of applications where intelligent decision support becomes pervasive. Bo Jin 0001, Leilei Sun, Chuanren Liu, Jianing Tong |
KDD | 1 |
| 2018 | Particle classification optimization-based BP network for telecommunication customer churn prediction
Ruiyun Yu, Xuanmiao An, Bo Jin 0001, Ann Move Oguti, Yonghe Liu |
Neural Comput. Appl. | 3 |
| 2017 | Multitask Dyadic Prediction and Its Application in Prediction of Adverse Drug-Drug InteractionabstractAdverse drug-drug interactions (DDIs) remain a leading cause of morbidity and mortality around the world. Identifying potential DDIs during the drug design process is critical in guiding targeted clinical drug safety testing. Although detection of adverse DDIs is conducted during Phase IV clinical trials, there are still a large number of new DDIs founded by accidents after the drugs were put on market. With the arrival of big data era, more and more pharmaceutical research and development data are becoming available, which provides an invaluable resource for digging insights that can potentially be leveraged in early prediction of DDIs. Many computational approaches have been proposed in recent years for DDI prediction. However, most of them focused on binary prediction (with or without DDI), despite the fact that each DDI is associated with a different type. Predicting the actual DDI type will help us better understand the DDI mechanism and identify proper ways to prevent it. In this paper, we formulate the DDI type prediction problem as a multitask dyadic regression problem, where the prediction of each specific DDI type is treated as a task. Compared with conventional matrix completion approaches which can only impute the missing entries in the DDI matrix, our approach can directly regress those dyadic relationships (DDIs) and thus can be extend to new drugs more easily. We developed an effective proximal gradient method to solve the problem. Evaluation on real world datasets is presented to demonstrate the effectiveness of the proposed approach. Bo Jin 0001, Cao Xiao, Ping Zhang 0016, Xiaopeng Wei, Fei Wang 0001 |
AAAI | 1 |
| 2017 | Exploring risk factors and predicting UPDRS score based on Parkinson's speech signalsabstractThe unified Parkinson's disease rating scale (UPDRS) is the most widely employed scale for tracking Parkinson's disease (PD) symptom progression. However, conventional way to achieve UPDRS, mainly based on the physical examinations of clinic patients performed by the trained medical staffs, involves the disadvantages of inconvenience and high medical expense. Hence, in this study, we try to explore some risk factors and accurately predict the UPDRS for PD, using the speech signals of PD patients published on UCI machine-learning archive. More specifically, inspired by the idea of ensemble learning, we firstly construct a framework of ensemble feature selection (EFS) to select a suitable subset of features among numerous speech signals. Subsequently, a personalized predictive model, trained by adopting information from similar patients, is developed to be customized for an individual PD patient. Finally, we employ the personalized predictive model to predict UPDRS score combined with various classical regression algorithms. Compared to conventional models, our study has a potential to capture more relevant risk factors and produces more accurate UPDRS score for individual patient. Experimental results on real-world dataset from UCI machine-learning archive show that our personalized predictive model gets a promising performance. Jianxin Zhang 0001, Qiang Zhang 0008, Bo Jin 0001, Xiaopeng Wei |
Healthcom | 4 |
| 2017 | An RNN Architecture with Dynamic Temporal Matching for Personalized Predictions of Parkinson's DiseaseabstractParkinson's disease (PD) is a chronic disease that develops over years and varies dramatically in its clinical manifestations. A preferred strategy to resolve this heterogeneity and thus enable better prognosis and targeted therapies is to segment out more homogeneous patient sub-populations. However, it is challenging to evaluate the clinical similarities among patients because of the longitudinality and temporality of their records. To address this issue, we propose a deep model that directly learns patient similarity from longitudinal and multi-modal patient records with an Recurrent Neural Network (RNN) architecture, which learns the similarity between two longitudinal patient record sequences through dynamically matching temporal patterns in patient sequences. Evaluations on real world patient records demonstrate the promising utility and efficacy of the proposed architecture in personalized predictions. Chao Che, Cao Xiao, Jian Liang 0002, Bo Jin 0001, Jiayu Zho, Fei Wang 0001 |
SDM | 4 |
| 2016 | Minimizing Legal Exposure of High-Tech Companies through Collaborative Filtering MethodsabstractPatent litigation not only covers legal and technical issues, it is also a key consideration for managers of high-technology (high-tech) companies when making strategic decisions. Patent litigation influences the market value of high-tech companies. However, this raises unique challenges. To this end, in this paper, we develop a novel recommendation framework to solve the problem of litigation risk prediction. We will introduce a specific type of patent-related litigation, that is, Section 337 investigations, which prohibit all acts of unfair competition, or any unfair trade practices, when exporting products to the United States. To build this recommendation framework, we collect and exploit a large amount of published information related to almost all Section 337 investigation cases. This study has two aims: (1) to predict the litigation risk in a specific industry category for high-tech companies and (2) to predict the litigation risk from competitors for high-tech companies. These aims can be achieved by mining historical investigation cases and related patents. Specifically, we propose two methods to meet the needs of both aims: a proximal slope one predictor and a time-aware predictor. Several factors are considered in the proposed methods, including the litigation risk if a company wants to enter a new market and the risk that a potential competitor would file a lawsuit against the new entrant. Comparative experiments using real-world data demonstrate that the proposed methods outperform several baselines with a significant margin. Bo Jin 0001, Chao Che, Kuifei Yu, Li Guo 0008, Cuili Yao, Ruiyun Yu, Qiang Zhang 0008 |
KDD | 1 |
| 2015 | Efficient Methods for Multi-label Classification
Chonglin Sun, Chunting Zhou, Bo Jin 0001, Francis C. M. Lau 0001 |
PAKDD (1) | 3 |
| 2014 | Technology Prospecting for High Tech Companies through Patent MiningabstractTechnology prospecting is a process to evaluate the potential business values of high tech companies from the technology perspective. In this paper, we provide a new view-angle to understand technology prospecting by studying the evolving distributions of technologies in the companies. Specifically, we first exploit topic models to learn technological context in the form of probabilistic distributions of assignees and locations from large-scale patent documents. Then, we develop a matching solution to measure the relationships between patent topics and the description documents of technology terms. In this way, we can obtain the distribution of technologies for each company. In addition, we are able to assess the technology prospecting of a company by a designed indicator, which allows to compare the levels of discrepancies between the emerging technology distributions available as Garner Hype Cycles and the distribution of technologies of the company. Finally, experimental results on real-world patent data show the effectiveness of our approach for technology prospecting. Bo Jin 0001, Yong Ge 0001, Hengshu Zhu, Li Guo 0008, Hui Xiong 0001 |
ICDM | 1 |
| 2013 | UT-Tree: Efficient mining of high utility itemsets from data streamsabstractHigh utility itemsets mining is a hot topic in data stream mining. It is essential that the mining algorithm should be efficient in both time and space for data stream is continuous and unbounded. To the best of our knowledge, the existing algorithms Lin Feng 0001, Bo Jin 0001 |
Intell. Data Anal. | 3 |
| 2013 | Maximal Similarity Embedding
Lin Feng 0001, Shenglan Liu 0001, Zhen Yu Wu, Bo Jin 0001 |
Neurocomputing | 4 |
| 2007 | Chinese Patent Mining Based on Sememe Statistics and Key-Phrase Extraction
Bo Jin 0001, Hongfei Teng, Yanjun Shi, Fuzheng Qu |
ADMA | 1 |