VLDB 2026 Research / reviewers in the wild / expert
Qiming Fu 0001
dblp:162/2675-1 · also Qi-ming Fu 0001
· DBLP profile ↗
39ranked-venue papers
5as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VRQA: Context-adaptive view routing for long-document question-answer generationabstractLong-document question–answer (QA) generation often employs large-scale generation with strong filtering to reduce off-topic deviations. However, such strategies tend to concentrate generation on a few safe entry points, leading to uneven coverage and similar question expression. Conversely, expanding the scope of questioning may weaken topical consistency and the verifiability of supporting evidence. To address these challenges, we propose VRQA, a long-document QA generation framework with context-adaptive view routing under hierarchical anchor constraints. It selects better question entry points for different chunks and generates QA pairs with more dispersed semantic focuses. VRQA first constructs a semantic representation from chunk-level evidence and key points, then encodes view names and their descriptions into view prototypes, and obtains view-aware contextual features through conditional modulation. Subsequently, it updates the view routing policy online through reinforcement learning, where the reward is constructed using anchor consistency and QA answer quality scores from an evaluator. This enables VRQA to adaptively select the optimal view set and coordinate with submodular selection to achieve a dynamic balance between quality and diversity. To implement this framework, VRQA adopts a three-stage coupled training schedule, consisting of router warm-up, view-conditioned representation learning, and online updating, through which the router is progressively refined by evaluator feedback. Experiments on three vertical domains from the Elsevier OA CC-BY dataset show that VRQA achieves state-of-the-art performance, improving quality by 11.46% on AS-Chk and semantic diversity by 7.96% on VS over the strongest baseline, and thus offering a better trade-off between quality and diversity for long-document QA generation. Mengting Huang, Hongjie Wu, Fuyuan Hu, Lanhui Liu, Qiming Fu 0001 |
Expert Syst. Appl. | 7 |
| 2025 | LinFa-Q: Accurate Q-learning with linear function approximation
Zhechao Wang, Qiming Fu 0001, Quan Liu 0004, You Lu 0004, Hongjie Wu, Fuyuan Hu |
Neurocomputing | 2 |
| 2025 | GTD3-NET: A deep reinforcement learning-based routing optimization algorithm for wireless networks
You Lu 0004, Lanhui Liu, Qiming Fu 0001 |
Peer Peer Netw. Appl. | 5 |
| 2025 | TrGPCR: GPCR-Ligand Binding Affinity Prediction Based on Dynamic Deep Transfer LearningabstractPredicting G protein-coupled receptor (GPCR) -ligand binding affinity plays a crucial role in drug development. However, determining GPCR-ligand binding affinities is time-consuming and resource-intensive. Although many studies used data-driven methods to predict binding affinity, most of these methods required protein 3D structure, which was often unknown. Moreover, part of these studies only considered the sequence characteristics of the protein, ignoring the secondary structure of the protein. The number of known GPCR for affinity prediction is only a few thousand, which is insufficient for deep learning training. Therefore, this study aimed to propose a deep transfer learning method called TrGPCR, which used dynamic transfer learning to solve the problem of insufficient GPCR data. We used the Binding Database (BindingDB) as the source domain and the GLASS (GPCR-Ligand Association) database as the target domain. We also introduced protein secondary structures, called pockets, as features to predict binding affinities. Compared with DeepDTA, our model improved by 5.2% on RMSE (root mean square error) and 4.5% on MAE (mean squared error). Yaoyao Lu, Tengsheng Jiang, Qiming Fu 0001, Zhiming Cui 0002, Hongjie Wu |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Reinforced Metapath Optimization in Heterogeneous Information Networks for Drug-Target Interaction PredictionabstractGraph neural networks offer an effective avenue for predicting drug-target interactions. In this domain, researchers have found that constructing heterogeneous information networks based on metapaths using diverse biological datasets enhances prediction performance. However, the performance of such methods is closely tied to the selection of metapaths and the compatibility between metapath subgraphs and graph neural networks. Most existing approaches still rely on fixed strategies for selecting metapaths and often fail to fully exploit node information along the metapaths, limiting the improvement in model performance. This paper introduces a novel method for predicting drug-target interactions by optimizing metapaths in heterogeneous information networks. On one hand, the method formulates the metapath optimization problem as a Markov decision process, using the enhancement of downstream network performance as a reward signal. Through iterative training of a reinforcement learning agent, a high-quality set of metapaths is learned. On the other hand, to fully leverage node information along the metapaths, the paper constructs subgraphs based on nodes along the metapaths. Different depths of subgraphs are processed using different graph convolutional neural network. The proposed method is validated using standard heterogeneous biological benchmark datasets. Experimental results on standard datasets show significant advantages over traditional methods. Ben Xu, Qiming Fu 0001, You Lu 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | LoGo-GR: A Local to Global Graphical Reasoning Framework for Extracting Structured Information From Biomedical LiteratureabstractIn the biomedical literature, entities are often distributed within multiple sentences and exhibit complex interactions. As the volume of literature has increased dramatically, it has become impractical to manually extract and maintain biomedical knowledge, which would entail enormous costs. Fortunately, document-level relation extraction can capture associations between entities from complex text, helping researchers efficiently mine structured knowledge from the vast medical literature. However, how to effectively synthesize rich global information from context and accurately capture local dependencies between entities is still a great challenge. In this paper, we propose a Local to Global Graphical Reasoning framework (LoGo-GR) based on a novel Biased Graph Attention mechanism (B-GAT). It learns global context feature and information of local relation path dependencies from mention-level interaction graph and entity-level path graph respectively, and collaborates with global and local reasoning to capture complex interactions between entities from document-level text. In particular, B-GAT integrates structural dependencies into the standard graph attention mechanism (GAT) as attention biases to adaptively guide information aggregation in graphical reasoning. We evaluate our method on three publicly biomedical document-level datasets: Drug-Mutation Interaction (DV), Chemical-induced Disease (CDR), and Gene-Disease Association (GDA). LoGo-GR has advanced and stable performance compared to other state-of-the-art methods (it achieves state-of-the-art performance with 96.14%-97.39% F1 on DV dataset, advanced performance with 68.89% F1 and 84.22% F1 on CDR and GDA datasets, respectively). In addition, LoGo-GR also shows advanced performance on general-domain document-level relation extraction dataset, DocRED, which proves that it is an effective and robust document-level relation extraction framework. Xueyang Zhou, Qiming Fu 0001, Youbing Xia, You Lu 0004 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Temporal-difference emphasis learning with regularized correction for off-policy evaluation and control
Jiaqing Cao, Quan Liu 0004, Qiming Fu 0001 |
Appl. Intell. | 4 |
| 2023 | MAML2: meta reinforcement learning via meta-learning for task categories
Qiming Fu 0001, Zhechao Wang, Nengwei Fang |
Frontiers Comput. Sci. | 1 |
| 2023 | Reinforcement Learning in Few-Shot Scenarios: A Survey
Zhechao Wang, Qiming Fu 0001, You Lu 0004, Hongjie Wu |
J. Grid Comput. | 2 |
| 2023 | Extracting biomedical relation from cross-sentence text using syntactic dependency graph attention network
Xueyang Zhou, Qiming Fu 0001, Lanhui Liu, You Lu 0004, Hongjie Wu |
J. Biomed. Informatics | 2 |
| 2023 | Generalized gradient emphasis learning for off-policy evaluation and control with function approximation
Jiaqing Cao, Quan Liu 0004, Qiming Fu 0001 |
Neural Comput. Appl. | 4 |
| 2022 | Study on Path Planning of Multi-storey Parking Lot Based on Combined Loss Function
Zhongtian Hu, Yuli Wang, Qiming Fu 0001, Weizhong Lu, Hongjie Wu |
ICIC (3) | 5 |
| 2022 | AGCN: Augmented Graph Convolutional Network for Lifelong Multi-Label Image RecognitionabstractThe Lifelong Multi-Label (LML) image recognition builds an online class-incremental classifier in a sequential multilabel image recognition data stream. However, training on the data with different Partial Labels may result in more serious Catastrophic Forgetting in old classes. To solve the problem, the study proposes an Augmented Graph Convolutional Network (AGCN)to build an Augmented Correlation Matrix (ACM) across the sequential partial-label tasks and sustain the catastrophic forgetting. First, in ACM, the intra-task relations derive from the hard label statistics, while the inter-task relations further leverage the soft labels from a stored expert network. Then, based on the ACM, AGCN captures label dependencies with dynamic augmented structure and yields effective class representations. Our method is evaluated on two multi-label image benchmarks and the results show that the proposed method is effective for LML image recognition. Kaile Du, Fan Lyu, Fuyuan Hu, Wei Feng 0005, Fenglei Xu, Qiming Fu 0001 |
ICME | 7 |
| 2022 | Energy-efficient control of thermal comfort in multi-zone residential HVAC via reinforcement learningabstractEnergy efficient control of thermal comfort has been already an important part of residential heating, ventilation, and air conditioning (HVAC) systems. However, the optimisation of energy saving control for thermal comfort is not an easy task due to the complex dynamics of HVAC systems, the dynamics of thermal comfort and the trade-off between energy saving and thermal comfort. To solve the above problem, we propose a deep reinforcement learning-based thermal comfort control method in multi-zone residential HVAC. In this paper, firstly we design a SVR-DNN model, consisting of Support Vector Regression and a Deep Neural Network to predict thermal comfort value. Then, we apply Deep Deterministic Policy Gradient (DDPG) based on the output of the SVR-DNN model to achieve an optimal HVAC thermal comfort control strategy. This method can minimise energy consumption while satisfying occupants' thermal comfort. The experimental results show that our method can improve thermal comfort prediction performance by 20.5% compared with DNN; compared with deep Q-network (DQN), energy consumption and thermal comfort violation can be reduced by 3.52% and 64.37% respectively. Zhengkai Ding, Qiming Fu 0001, Hong-Jie Wu, You Lu 0004, Fuyuan Hu |
Connect. Sci. | 2 |
| 2022 | G Protein-Coupled Receptor Interaction Prediction Based on Deep Transfer LearningabstractG protein-coupled receptors (GPCRs) account for about 40% to 50% of drug targets. Many human diseases are related to G protein coupled receptors. Accurate prediction of GPCR interaction is not only essential to understand its structural role, but also helps design more effective drugs. At present, the prediction of GPCR interaction mainly uses machine learning methods. Machine learning methods generally require a large number of independent and identically distributed samples to achieve good results. However, the number of available GPCR samples that have been marked is scarce. Transfer learning has a strong advantage in dealing with such small sample problems. Therefore, this paper proposes a transfer learning method based on sample similarity, using XGBoost as a weak classifier and using the TrAdaBoost algorithm based on JS divergence for data weight initialization to transfer samples to construct a data set. After that, the deep neural network based on the attention mechanism is used for model training. The existing GPCR is used for prediction. In short-distance contact prediction, the accuracy of our method is 0.26 higher than similar methods. Tengsheng Jiang, Yuhui Chen, Zhongtian Hu, Weizhong Lu, Qiming Fu 0001, Yijie Ding, Haiou Li, Hongjie Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | A Reinforcement Learning-Based Model for Human MicroRNA-Disease Association Prediction
Linqian Cui, You Lu 0004, Qiming Fu 0001, Yijie Ding, Hongjie Wu |
ICIC (3) | 3 |
| 2021 | DNA-Binding Protein Prediction Based on Deep Learning Feature Fusion
Tengsheng Jiang, Weizhong Lu, Qiming Fu 0001, Haiou Li, Hongjie Wu |
ICIC (3) | 4 |
| 2021 | Multi-zone Residential HVAC Control with Satisfying Occupants' Thermal Comfort Requirements and Saving Energy via Reinforcement Learning
Zhengkai Ding, Qiming Fu 0001, Hongjie Wu, You Lu 0004, Fuyuan Hu |
PDCAT | 2 |
| 2021 | Research on RNA secondary structure predicting via bidirectional recurrent neural networkabstractBACKGROUND: RNA secondary structure prediction is an important research content in the field of biological information. Predicting RNA secondary structure with pseudoknots has been proved to be an NP-hard problem. Traditional machine learning methods can not effectively apply protein sequence information with different sequence lengths to the prediction process due to the constraint of the self model when predicting the RNA secondary structure. In addition, there is a large difference between the number of paired bases and the number of unpaired bases in the RNA sequences, which means the problem of positive and negative sample imbalance is easy to make the model fall into a local optimum. To solve the above problems, this paper proposes a variable-length dynamic bidirectional Gated Recurrent Unit(VLDB GRU) model. The model can accept sequences with different lengths through the introduction of flag vector. The model can also make full use of the base information before and after the predicted base and can avoid losing part of the information due to truncation. Introducing a weight vector to predict the RNA training set by dynamically adjusting each base loss function solves the problem of balanced sample imbalance. RESULTS: The algorithm proposed in this paper is compared with the existing algorithms on five representative subsets of the data set RNA STRAND. The experimental results show that the accuracy and Matthews correlation coefficient of the method are improved by 4.7% and 11.4%, respectively. CONCLUSIONS: The flag vector introduced allows the model to effectively use the information before and after the protein sequence; the introduced weight vector solves the problem of unbalanced sample balance. Compared with other algorithms, the LVDB GRU algorithm proposed in this paper has the best detection results. Weizhong Lu, Hongjie Wu, Yijie Ding, Zhengwei Song, Yu Zhang 0027, Qiming Fu 0001, Haiou Li |
BMC Bioinform. | 7 |
| 2021 | Gradient temporal-difference learning for off-policy evaluation using emphatic weightings
Jiaqing Cao, Quan Liu 0004, Fei Zhu 0003, Qiming Fu 0001 |
Inf. Sci. | 4 |
| 2021 | Empirical Potential Energy Function Toward ab Initio Folding G Protein-Coupled ReceptorsabstractApproximately 40-50 percent of all drugs targets are G protein-coupled receptors (GPCRs). Three-dimensional structure of GPCRs is important to probe their biophysical and biochemical functions and their pharmaceutical applications. Lacking reliable and high quality free function is one of the ugent problems of computational predicting the three-dimensional structure in this community. We proposed a GPCR-specified energy function composed of four novel empirical potential energy terms: a two-dimensional contact energy force field, knowledge-based helix pair connection distance energy term, knowledge-based helix pair angle restraint energy term and a disulfide bond energy term. To validate the energy function, we employed an ab initio GPCR three-dimensional structure predictor to test if the energy function improved the accuracy of prediction. We evaluated 28 solved GPCRs and found that 21(75 percent) targets were correctly folded (TM-score>0.5). Also, the average TM-score using the energy function was 0.54, which was improved 134 percent than the TM-score 0.23 for MODELLER energy function and 170 percent than the TM-score 0.20 for Rosetta membrane energy function. The results confirmed that our empirical potential energy function toward ab initio folding is competitive to state-of-the-art solutions for structural prediction of GPCRs. Hongjie Wu, Huajing Ling, Qiming Fu 0001, Weizhong Lu, Yijie Ding, Min Jiang 0009, Haiou Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Behavior Prediction for Unmanned Driving Based on Dual Fusions of Feature and DecisionabstractBehavioral decision systems may suffer from poor performance due to the failure in capturing the vibrations of environmental information. To better capture such vibrations and then make more accurate predictions, a parallel deep neural network based on dual fusions including feature and decision is proposed, called DFFD-Net. DFFD-NET is composed of two parts, the feature fusion network and the driving data network. The feature fusion model adopts two different operations, deconvolution and linear weighting, to fuse local features and global features, respectively. Deconvolution is applied between the convolutional layers, while linear weighting is operated among the outputs of SPP and LSTM. To further improve the accuracy of the prediction, the decisions generated from both networks are further weighed to get the final decision. Experimentally, DFFD-NET is implemented in the benchmarks BDDV and TORCS, and the results show that the final performance is benefited from both feature fusion and decision fusion. From the comparison, DFFD-NET can get state-of-the-art results on both perplexity and precision by only using the images captured from the front-facing camera as well as a few sensing data. Shengrong Gong, Kaijian Xia, Yuchen Fu, Qiming Fu 0001, Hongsheng Yin 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2020 | Prediction of Membrane Protein Interaction Based on Deep Residual Learning
Tengsheng Jiang, Hongjie Wu, Yuhui Chen, Haiou Li, Jin Qiu, Weizhong Lu, Qiming Fu 0001 |
ICIC (2) | 7 |
| 2020 | A Building Energy Consumption Prediction Method Based on Integration of a Deep Neural Network and Transfer Reinforcement LearningabstractWith respect to the problem of the low accuracy of traditional building energy prediction methods, this paper proposes a novel prediction method for building energy consumption, which is based on the seamless integration of the deep neural network and transfer reinforcement learning (DNN-TRL). The method introduces a stack denoising autoencoder to extract the deep features of the building energy consumption, and shares the hidden layer structure to transfer the common information between different building energy consumption problems. The output of the DNN model is used as the input of the Sarsa algorithm to improve the prediction performance of the target building energy consumption. To verify the performance of the DNN-TRL algorithm, based on the data recorded by American Power Balti Gas and Electric Power Company, and compared with Sarsa, ADE-BPNN, and BP-Adaboost algorithms, the experimental results show that the DNN-TRL algorithm can effectively improve the prediction accuracy of the building energy consumption. Qiming Fu 0001, QingSong Liu, Hongjie Wu, Baochuan Fu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2019 | Knowledge Based Helix Angle and Residue Distance Restraint Free Energy Terms of GPCRs
Huajing Ling, Hongjie Wu, Jiayan Han, Jiwen Ding, Weizhong Lu, Qiming Fu 0001 |
ICIC (2) | 6 |
| 2019 | SR-GAN: Semantic Rectifying Generative Adversarial Network for Zero-shot LearningabstractThe existing Zero-Shot learning (ZSL) methods may suffer from the vague class attributes that are highly overlapped for different classes. Unlike these methods that ignore the discrimination among classes, in this paper, we propose to classify unseen image by rectifying the semantic space guided by the visual space. First, we pre-train a Semantic Rectifying Network (SRN) to rectify semantic space with a semantic loss and a rectifying loss. Then, a Semantic Rectifying Generative Adversarial Network (SR-GAN) is built to generate plausible visual feature of unseen class from both semantic feature and rectified semantic feature. To guarantee the effectiveness of rectified semantic features and synthetic visual features, a pre-reconstruction and a post reconstruction networks are proposed, which keep the consistency between visual feature and semantic feature. Experimental results demonstrate that our approach significantly outperforms the state-of-the-arts on four benchmark datasets. Zihan Ye, Fan Lyu, Qiming Fu 0001, Jinchang Ren, Fuyuan Hu |
ICME | 4 |
| 2019 | Predicting RNA secondary structure via adaptive deep recurrent neural networks with energy-based filterabstractBACKGROUND: RNA secondary structure prediction is an important issue in structural bioinformatics, and RNA pseudoknotted secondary structure prediction represents an NP-hard problem. Recently, many different machine-learning methods, Markov models, and neural networks have been employed for this problem, with encouraging results regarding their predictive accuracy; however, their performances are usually limited by the requirements of the learning model and over-fitting, which requires use of a fixed number of training features. Because most natural biological sequences have variable lengths, the sequences have to be truncated before the features are employed by the learning model, which not only leads to the loss of information but also destroys biological-sequence integrity. RESULTS: To address this problem, we propose an adaptive sequence length based on deep-learning model and integrate an energy-based filter to remove the over-fitting base pairs. CONCLUSIONS: Comparative experiments conducted on an authoritative dataset RNA STRAND (RNA secondary STRucture and statistical Analysis Database) revealed a 12% higher accuracy relative to three currently used methods. Weizhong Lu, Hongjie Wu, Hongmei Huang, Qiming Fu 0001, Haiou Li |
BMC Bioinform. | 5 |
| 2019 | Ranking near-native candidate protein structures via random forest classificationabstractBACKGROUND: In ab initio protein-structure predictions, a large set of structural decoys are often generated, with the requirement to select best five or three candidates from the decoys. The clustered central structures with the most number of neighbors are frequently regarded as the near-native protein structures with the lowest free energy; however, limitations in clustering methods and three-dimensional structural-distance assessments make identifying exact order of the best five or three near-native candidate structures difficult. RESULTS: To address this issue, we propose a method that re-ranks the candidate structures via random forest classification using intra- and inter-cluster features from the results of the clustering. Comparative analysis indicated that our method was better able to identify the order of the candidate structures as comparing with current methods SPICKR, Calibur, and Durandal. The results confirmed that the identification of the first model were closer to the native structure in 12 of 43 cases versus four for SPICKER, and the same as the native structure in up to 27 of 43 cases versus 14 for Calibur and up to eight of 43 cases versus two for Durandal. CONCLUSIONS: In this study, we presented an improved method based on random forest classification to transform the problem of re-ranking the candidate structures by an binary classification. Our results indicate that this method is a powerful method for the problem and the effect of this method is better than other methods. Hongjie Wu, Hongmei Huang, Weizhong Lu, Qiming Fu 0001, Yijie Ding, Haiou Li |
BMC Bioinform. | 4 |
| 2019 | Research on predicting 2D-HP protein folding using reinforcement learning with full state spaceabstractBACKGROUND: Protein structure prediction has always been an important issue in bioinformatics. Prediction of the two-dimensional structure of proteins based on the hydrophobic polarity model is a typical non-deterministic polynomial hard problem. Currently reported hydrophobic polarity model optimization methods, greedy method, brute-force method, and genetic algorithm usually cannot converge robustly to the lowest energy conformations. Reinforcement learning with the advantages of continuous Markov optimal decision-making and maximizing global cumulative return is especially suitable for solving global optimization problems of biological sequences. RESULTS: In this study, we proposed a novel hydrophobic polarity model optimization method derived from reinforcement learning which structured the full state space, and designed an energy-based reward function and a rigid overlap detection rule. To validate the performance, sixteen sequences were selected from the classical data set. The results indicated that reinforcement learning with full states successfully converged to the lowest energy conformations against all sequences, while the reinforcement learning with partial states folded 50% sequences to the lowest energy conformations. Reinforcement learning with full states hits the lowest energy on an average 5 times, which is 40 and 100% higher than the three and zero hit by the greedy algorithm and reinforcement learning with partial states respectively in the last 100 episodes. CONCLUSIONS: Our results indicate that reinforcement learning with full states is a powerful method for predicting two-dimensional hydrophobic-polarity protein structure. It has obvious competitive advantages compared with greedy algorithm and reinforcement learning with partial states. Hongjie Wu, Qiming Fu 0001, Weizhong Lu, Haiou Li |
BMC Bioinform. | 3 |
| 2019 | Efficient reinforcement learning in continuous state and action spaces with Dyna and policy approximation
Quan Liu 0004, Zongzhang Zhang, Qiming Fu 0001 |
Frontiers Comput. Sci. | 4 |
| 2019 | Variational Bayesian Exploration-Based Active Sarsa AlgorithmabstractWe proposed an improved variational Bayesian exploration-based active Sarsa (VBE-ASAR) algorithm, which tries to balance the exploration and exploitation dilemma, and speeds up the convergence rate. First, in the learning process, variational Bayesian method is adopted to measure the information gain, which is used as an exploration factor to construct an internal reward function for heuristic exploration. In addition, before the learning process, in order to improve the exploration performance, transfer learning is used to initialize the value function, where Bisimulation metric is introduced to measure the distance between two states from the source MDP and the target MDP, respectively. Finally, we apply the proposed algorithm to the cliff walking problem, and compare with the Sarsa algorithm, the Q-Learning algorithm, the VFT-Sarsa algorithm and the Bayesian Sarsa (BS) algorithm. Experimental results show that the VBE-ASAR algorithm has a faster learning rate. Qiming Fu 0001, Zhengxia Yang, You Lu 0004, Hongjie Wu, Fuyuan Hu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2018 | Prediction of Indoor PM2.5 Index Using Genetic Neural Network Model
Hongjie Wu, Weisheng Liu, Qiming Fu 0001, Baochuan Fu, Dadong Dai |
ICIC (1) | 5 |
| 2018 | Optimizing GPCR Two-Dimensional Topology from Contact Map
Hongjie Wu, Dadong Dai, Huaxiang Shen, Weizhong Lu, Qiming Fu 0001 |
ICIC (3) | 6 |
| 2018 | RNA Secondary Structure Prediction Based on Long Short-Term Memory Model
Hongjie Wu, Weizhong Lu, Hongmei Huang, Qiming Fu 0001 |
ICIC (1) | 6 |
| 2018 | Optimizing HP Model Using Reinforcement Learning
Hongjie Wu, Qiming Fu 0001 |
ICIC (2) | 3 |
| 2018 | Single Trajectory Learning: Exploration Versus ExploitationabstractIn reinforcement learning (RL), the exploration/exploitation (E/E) dilemma is a very crucial issue, which can be described as searching between the exploration of the environment to find more profitable actions, and the exploitation of the best empirical actions for the current state. We focus on the single trajectory RL problem where an agent is interacting with a partially unknown MDP over single trajectories, and try to deal with the E/E in this setting. Given the reward function, we try to find a good E/E strategy to address the MDPs under some MDP distribution. This is achieved by selecting the best strategy in mean over a potential MDP distribution from a large set of candidate strategies, which is done by exploiting single trajectories drawn from plenty of MDPs. In this paper, we mainly make the following contributions: (1) We discuss the strategy-selector algorithm based on formula set and polynomial function. (2) We provide the theoretical and experimental regret analysis of the learned strategy under an given MDP distribution. (3) We compare these methods with the “state-of-the-art” Bayesian RL method experimentally. Qiming Fu 0001, Quan Liu 0004, Hongjie Wu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2016 | Reasoning and predicting POMDP planning complexity via covering numbers
Zongzhang Zhang, Qiming Fu 0001, Quan Liu 0004 |
Frontiers Comput. Sci. | 2 |
| 2015 | A Bayesian Sarsa Learning Algorithm with Bandit-Based Method
Shuhua You, Quan Liu 0004, Qiming Fu 0001, Fei Zhu 0003 |
ICONIP (1) | 3 |
| 2013 | The second order temporal difference error for Sarsa(λ)abstractTraditional reinforcement learning algorithms, such as Q-learning, Q(λ), Sarsa, and Sarsa(λ), update the action value function using temporal difference (TD) error, which is computed by the last action value function. From the perspective of the TD error, and with respect to the problems of low efficiency and slow convergence of the traditional Sarsa(λ) algorithm, this paper defines the nthorder TD Error, applies it in the traditional Sarsa(λ) algorithm, and develops a fast Sarsa(λ) algorithm based on the 2ndorder TD Error. The algorithm adjusts the Q value with the second-order TD Error and broadcasts the TD Error into the whole state-action space, which speeds up the convergence of the algorithm. This paper also analyzes the convergence rate, and under the condition of one-step update, the results show that the number of iteration depends primarily on γ, ε. Finally, using the proposed algorithm on the traditional reinforcement learning problems, the results show that the algorithm has both a faster convergence rate and better convergence performance. Qiming Fu 0001, Quan Liu 0004, Guixin Chen |
ADPRL | 1 |