VLDB 2026 Research / reviewers in the wild / expert
Jinfang Cai
dblp:196/2471
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2021
0000-0002-8517-6253ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Using Knowledge Concept Aggregation towards Accurate Cognitive DiagnosisabstractCognitive diagnosis is a crucial task in the field of educational measurement and psychology, which is aimed to mine and analyze the level of knowledge for a student in his or her learning process periodically. While a number of approaches and tools have been developed to diagnose the learning states of students, they do not fully learn the relationship between students, exercises and knowledge concepts in the learning system, or do not consider the traits that it is easier to complete diagnosis when focusing on a small part of knowledge concepts rather than all knowledge concepts. To address these limitations, we develop CDGK, a model based artificial neural network to deal with cognitive diagnosis. Our method not only captures non-linear interactions between exercise features, student scores, and their mastery on each knowledge concept, but also performs an aggregation of the knowledge concepts via converting them into graph structure, and only considering the leaf node in the knowledge concept tree, which can reduce the dimension of the model without accuracy loss. In our evaluation on two real-world datasets, CDGK outperforms the state-of-the-art related approaches in terms of accuracy, reasonableness and interpretability. Xinping Wang, Caidie Huang, Jinfang Cai, Liangyu Chen 0001 |
CIKM | 3 |
| 2021 | Using Surrounding Text of Formula towards More Accurate Mathematical Information RetrievalabstractFormula retrieval is an important research topic in Mathematical Information Retrieval (MIR).Most studies have focused on comparing formulae to determine the similarity between mathematical documents.However, two similar formulae may appear in completely different knowledge domains and have different meanings.Based on N-ary Tree-based Formula Embedding Model (NTFEM), we introduce a new hybrid retrieval model combining formula with its surrounding text for more accurate retrieval.Using keywords extraction technology, we extract keywords from text around the formula which can supplement the semantic information of formula.Then we get the representation vectors of keywords by FastText N-gram embedding model, and the representation vectors of formulae by NTFEM.Finally, documents are first sorted according to the similarity of keywords, and then the ranking results are optimized by formula similarity.Experimental results show that the accuracy of top-10 results is at least 20% higher than that of NTFEM and can be 50% in some specific topics. Cheng Chen 0015, Yuqi Shen, Jinfang Cai, Liangyu Chen 0001 |
SEKE | 4 |
| 2021 | A Hybrid Model Combining Formulae with Keywords for Mathematical Information RetrievalabstractFormula retrieval is an important research topic in Mathematical Information Retrieval (MIR). Most studies have focused on formula comparison to determine the similarity between mathematical documents. However, two similar formulae may appear in entirely different knowledge domains and have different meanings. Based on N-ary Tree-based Formula Embedding Model (NTFEM, our previous work in [Y. Dai, L. Chen, and Z. Zhang, An N-ary tree-based model for similarity evaluation on mathematical formulae, in Proc. 2020 IEEE Int. Conf. Systems, Man, and Cybernetics, 2020, pp. 2578–2584.], we introduce a new hybrid retrieval model, NTFEM-K, which combines formulae with their surrounding keywords for more accurate retrieval. By using keywords extraction technology, we extract keywords from context, which can supplement the semantic information of the formula. Then, we get the vector representations of keywords by FastText N-gram embedding model and the vector representations of formulae by NTFEM. Finally, documents are sorted according to the similarity between keywords, and then the ranking results are optimized by formula similarity. For performance evaluation, NTFEM-K is not only compared with NTFEM but also hybrid retrieval models combining formulae with long text and hybrid retrieval models combining formulae with their keywords using other keyword extraction algorithms. Experimental results show that the accuracy of top-10 results of NTFEM-K is at least 20% higher than that of NTFEM and can be 50% in some specific topics. Yuqi Shen, Cheng Chen 0015, Jinfang Cai, Liangyu Chen 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 4 |