EDBT 2026 Demo / reviewers in the wild / expert
Jinkun Chen
dblp:23/11539
· DBLP profile ↗
33ranked-venue papers
11as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unified Minimax Optimization Framework for Propensity Score Estimation in Debiased RecommendationabstractRecommendation systems commonly face selection bias from missing-not-at-random (MNAR) collected data. To address this bias, propensity-based methods such as inverse propensity scoring (IPS) and doubly robust (DR) estimators are widely used. In addition, many methods extend the vanilla IPS and DR to further control the bias, variance, propensity mis-calibration, and imbalance, but they only optimize some of the above metrics, limiting the debiasing performance. In this paper, we first empirically find that controlling one metric cannot guarantee the control of other important metrics, then we reveal a fundamental structural commonality among the above four important metrics, and propose a Unified Propensity Optimization (UPO) framework that optimizes all metrics simultaneously by a minimax optimization algorithm. Theoretically, we demonstrate that minimizing the UPO loss effectively controls all metrics, ensuring their simultaneous improvements without incurring additional bias, and achieving reduced variance compared to naively adding up multiple control losses in penalty terms. Empirically, experiments on a semi-synthetic dataset and three real-world datasets validate UPO’s effectiveness, demonstrating superior performance compared to state-of-the-art methods with minor computational overhead. We fully open-source our code. Chunyuan Zheng 0001, Jinkun Chen, Shufeng Zhang |
AAAI | 3 |
| 2026 | Attribute-coverage-oriented cognitive learning approach: Fuzzy concept-based perspective for knowledge discovery
Yidong Lin, Taoju Liang, Jinjin Li 0001, Jinkun Chen |
Fuzzy Sets Syst. | 4 |
| 2026 | Optimal scale reduction for generalized multi-scale decision tables based on conditional entropy
Jinkun Chen, Ling Wei |
Int. J. Approx. Reason. | 2 |
| 2026 | Robust feature selection with adaptive k-nearest-neighbor rough sets
Zhou-Ming Ma, Jiaping Li, Jinkun Chen, Guoping Lin, Meiling Zheng |
Pattern Recognit. | 4 |
| 2026 | TPFS: A three-phase heuristic feature selection algorithm for large-scale sample datasets
Haoran Su, Jinkun Chen |
Pattern Recognit. | 2 |
| 2025 | Addressing Correlated Latent Exogenous Variables in Debiased Recommender SystemsabstractRecommendation systems (RS) aim to provide personalized content, but they face a challenge in unbiased learning due to selection bias, where users only interact with items they prefer. This bias leads to a distorted representation of user preferences, which hinders the accuracy and fairness of recommendations. To address the issue, various methods such as error imputation based, inverse propensity scoring, and doubly robust techniques have been developed. Despite the progress, from the structural causal model perspective, previous debiasing methods in RS assume the independence of the exogenous variables. In this paper, we release this assumption and propose a learning algorithm based on likelihood maximization to learn a prediction model. We first discuss the correlation and difference between unmeasured confounding and our scenario, then we propose a unified method that effectively handles latent exogenous variables. Specifically, our method models the data generation process with latent exogenous variables under mild normality assumptions. We then develop a Monte Carlo algorithm to numerically estimate the likelihood function. Extensive experiments on synthetic datasets and three real-world datasets demonstrate the effectiveness of our proposed method. The code is at https://github.com/WallaceSUI/kdd25-background-variable. Shuqiang Zhang, Yuchao Zhang 0001, Jinkun Chen, Haochen Sui |
KDD (2) | 3 |
| 2025 | Feature selection considering synergy between features based on soft neighborhood rough sets
Lubin Chen, Jinkun Chen, Yaojin Lin |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | CAB-KWS : Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
Weinan Dai, Yifeng Jiang 0007, Yuanjing Liu, Jinkun Chen, Xin Sun 0034, Jinglei Tao |
ICPR (3) | 4 |
| 2024 | Long-form evaluation of model editingabstractDomenic Rosati, Robie Gonzales, Jinkun Chen, Xuemin Yu, Yahya Kayani, Frank Rudzicz, Hassan Sajjad. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Domenic Rosati, Robie Gonzales, Jinkun Chen, Xuemin Yu, Melis Erkan, Yahya Kayani, Satya Deepika Chavatapalli, Frank Rudzicz, Hassan Sajjad 0001 |
NAACL-HLT | 3 |
| 2024 | Label Distribution Learning Based on Horizontal and Vertical Mining of Label CorrelationsabstractLabel distribution learning (LDL) is a novel approach that outputs labels with varying degrees of description. To enhance the performance of LDL algorithms, researchers have developed different algorithms with mining label correlations globally, locally, and both globally and locally. However, existing LDL algorithms for mining local label correlations roughly assume that samples within a cluster share same label correlations, which may not be applicable to all samples. Moreover, existing LDL algorithms apply global and local label correlations to the same parameter matrix, which cannot fully exploit their respective advantages. To address these issues, a novel LDL method based on horizontal and vertical mining of label correlations (LDL-HVLC) is proposed in this paper. The method first encodes a unique local influence vector for each sample through the label distribution of its neighbor samples. Then, this vector is extended as additional features to assist in predicting unknown instances, and a penalty term is designed to correct wrong local influence vector (horizontal mining). Finally, to capture both local and global correlations of label, a new regularization term is constructed to constrain the global label correlations on the output results (vertical mining). Extensive experiments on real datasets demonstrate that the proposed method effectively solves the label distribution problem and outperforms the current state-of-the-art methods. Yaojin Lin, Yulin Li 0002, Chenxi Wang 0002, Lei Guo 0020, Jinkun Chen |
IEEE Trans. Big Data | 5 |
| 2023 | Semantic-gap-oriented feature selection in hierarchical classification learning
Yaojin Lin, Chenxi Wang 0002, Lei Guo 0020, Jinkun Chen |
Inf. Sci. | 5 |
| 2022 | Bring dialogue-context into RNN-T for streaming ASR
Junfeng Hou, Jinkun Chen, Yufeng Tang, Jun Zhang 0066, Zejun Ma 0001 |
INTERSPEECH | 2 |
| 2022 | A Spectral Feature Selection Approach With Kernelized Fuzzy Rough SetsabstractFeature evaluation is an important issue in constructing a feature selection algorithm in kernelized fuzzy rough sets, which has been proven to be an effective approach to deal with nonlinear classification tasks and uncertainty in learning problems. However, the feature evaluation function developed with kernelized fuzzy rough sets cannot better reflect the affinity relationship of samples and is time-consuming. To overcome these drawbacks, in this article, the problem of feature selection with kernelized fuzzy rough sets is studied based on the spectral graph theory. First, the within-class and between-class sample similarity matrices by using kernelized fuzzy approximation operators are constructed. Two operators, which can capture the affinity relationship of samples, are then introduced based on the sample similarity matrices. The proposed operator can be regarded as the sum of the weighted kernelized fuzzy approximation operators. Second, based on the ratio criterion, a feature evaluation function and its corresponding feature selection algorithm FRKF are presented, which can effectively evaluate the importance of features. Third, to illustrate the performance of the proposed algorithm, extensive experiments have been carried out to compare FRKF and other well-known feature selection methods, including the feature ranking methods and feature subset selection methods on various classification tasks. The experimental results on real-world datasets demonstrate that FRKF achieves the high performances in terms of the robustness, efficiency, and effectiveness. Jinkun Chen, Yaojin Lin, Ju-Sheng Mi, Shaozi Li, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2021 | HMM-Free Encoder Pre-Training for Streaming RNN TransducerabstractThis work describes an encoder pre-training procedure using frame-wise label to improve the training of streaming recurrent neural network transducer (RNN-T) model.Streaming RNN-T trained from scratch usually performs worse than nonstreaming RNN-T.Although it is common to address this issue through pre-training components of RNN-T with other criteria or frame-wise alignment guidance, the alignment is not easily available in end-to-end manner.In this work, frame-wise alignment, used to pre-train streaming RNN-T's encoder, is generated without using a HMM-based system.Therefore an allneural framework equipping HMM-free encoder pre-training is constructed.This is achieved by expanding the spikes of CTC model to their left/right blank frames, and two expanding strategies are proposed.To our best knowledge, this is the first work to simulate HMM-based frame-wise label using CTC model for pre-training.Experiments conducted on LibriSpeech and MLS English tasks show the proposed pre-training procedure, compared with random initialization, reduces the WER by relatively 5%∼11% and the emission latency by 60 ms.Besides, the method is lexicon-free, so it is friendly to new languages without manually designed lexicon. Jingyu Sun, Yufeng Tang, Junfeng Hou, Jinkun Chen, Jun Zhang 0066, Zejun Ma 0001 |
Interspeech | 5 |
| 2021 | Transitive Halifax: An Activity-Based Search Engine for Bus RoutesabstractTransitive Halifax is an activity-oriented mobility service that allows users to search for bus routes toward places where they can perform their desired activities. The service is based on the observation that individuals often go to a place to conduct an activity. Simultaneously, the activity is often not strictly related to a single place since one may go shopping or eating in different locations. Transitive Halifax has a web interface that helps the user find the most relevant bus routes and bus stops candidates that they could use to go to places where they can perform their intended activity. The system implements a search engine that ranks the bus stops candidates according to the user's preferences and desired activities. Jinkun Chen, Vinicius Monteiro de Lira, Fernando Vieira Paulovich, Amílcar Soares Júnior 0001 |
MDM | 1 |
| 2021 | Kernelized fuzzy rough sets based online streaming feature selection for large-scale hierarchical classification
Shengxing Bai, Yaojin Lin, Jinkun Chen, Chenxi Wang 0002 |
Appl. Intell. | 4 |
| 2021 | Causality-based online streaming feature selectionabstractAbstract Online streaming feature selection, as a well‐known and effective preprocessing approach in machine learning, is an eternal topic. Amount of online streaming feature selection algorithms have achieved a great deal of success in classification and prediction tasks. However, most of these existing algorithms only concentrate on the relevance between features and labels, and neglect the causal relationships between them. Discovering the potential causal relationships between features and labels, that is, the Markov blanket (MB) of class label, which can build a more interpretable and robust classification model. In this paper, we put forward a causality‐based online streaming feature selection algorithm with neighborhood conditional mutual information. First, we apply neighborhood symmetrical uncertainty to discover a candidate Markov blanket (CMB) with causal information. Then, neighborhood conditional mutual information instead of conditional independence test is used to delete the false positives in CMB, which can significantly alleviate the computational cost. Moreover, we utilize the updated CMB to choose the true spouses, which may be mistakenly deleted during the process of removing false positives, and then acquire an optimal MB as the online selected feature subset. Finally, causality‐based online streaming feature selection with neighborhood conditional mutual information is compared with four well‐established online streaming feature selection methods on 13 real‐world datasets. Experiment results show that the proposed algorithm outperforms these online streaming feature selection algorithms. Longzhu Li, Yaojin Lin, Hong Zhao 0002, Jinkun Chen, Shaozi Li |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | A new fuzzy multi-attribute group decision-making method with generalized maximal consistent block and its application in emergency management
Ju-Sheng Mi, Jinkun Chen, Wen Liu 0009 |
Knowl. Based Syst. | 3 |
| 2020 | A graph approach for fuzzy-rough feature selection
Jinkun Chen, Ju-Sheng Mi, Yaojin Lin |
Fuzzy Sets Syst. | 1 |
| 2020 | On-the-Fly Data Loader and Utterance-Level Aggregation for Speaker and Language RecognitionabstractIn this article, our recent efforts on directly modeling utterance-level aggregation for speaker and language recognition is summarized. First, an on-the-fly data loader for efficient network training is proposed. The data loader acts as a bridge between the full-length utterances and the network. It generates mini-batch samples on the fly, which allows batch-wise variable-length training and online data augmentation. Second, the traditional dictionary learning and Baum-Welch statistical accumulation mechanisms are applied to the network structure, and a learnable dictionary encoding (LDE) layer is introduced. The former accumulates discriminative statistics from the variable-length input sequence and outputs a single fixed-dimensional utterance-level representation. Experiments were conducted on four different datasets, namely NIST LRE 2007, AP17-OLR, SITW, and NIST SRE 2016. Experimental results show the effectiveness of the proposed batch-wise variable-length training with online data augmentation and the LDE layer, which significantly outperforms the baseline methods. Weicheng Cai, Jinkun Chen, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | A fast attribute reduction method for large formal decision contexts
Jinkun Chen, Ju-Sheng Mi, Bin Xie 0003, Yaojin Lin |
Int. J. Approx. Reason. | 1 |
| 2019 | Intuitionistic Fuzzy Rough Set-Based Granular Structures and Attribute Subset SelectionabstractAttribute subset selection is an important issue in data mining and information processing. However, most automatic methodologies consider only the relevance factor between samples while ignoring the diversity factor. This may not allow the utilization value of hidden information to be exploited. For this reason, we propose a hybrid model named intuitionistic fuzzy (IF) rough set to overcome this limitation. The model combines the technical advantages of rough set and IF set and can effectively consider the above-mentioned statistical factors. First, fuzzy information granules based on IF relations are defined and used to characterize the hierarchical structures of the lower and upper approximations of IF rough set within the framework of granular computing. Then, the computation of IF rough approximations and knowledge reduction in IF information systems are investigated. Third, based on the approximations of IF rough set, significance measures are developed to evaluate the approximation quality and classification ability of IF relations. Furthermore, a forward heuristic algorithm for finding one optimal reduct of IF information systems is developed using these measures. Finally, numerical experiments are conducted on public datasets to examine the effectiveness and efficiency of the proposed algorithm in terms of the number of selected attributes, computational time, and classification accuracy. Anhui Tan, Weizhi Wu 0001, Jiye Liang, Jinkun Chen, Jinjin Li 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2018 | Analysis of Length Normalization in End-to-End Speaker Verification SystemabstractThe classical i-vectors and the latest end-to-end deep speaker embeddings are the two representative categories of utterancelevel representations in automatic speaker verification systems.Traditionally, once i-vectors or deep speaker embeddings are extracted, we rely on an extra length normalization step to normalize the representations into unit-length hyperspace before back-end modeling.In this paper, we explore how the neural network learns length-normalized deep speaker embeddings in an end-to-end manner.To this end, we add a length normalization layer followed by a scale layer before the output layer of the common classification network.We conducted experiments on the verification task of the Voxceleb1 dataset.The results show that integrating this simple step in the end-to-end training pipeline significantly boosts the performance of speaker verification.In the testing stage of our L2-normalized end-to-end system, a simple inner-product can achieve the state-of-the-art. Weicheng Cai, Jinkun Chen, Ming Li 0026 |
INTERSPEECH | 2 |
| 2018 | A graph approach for knowledge reduction in formal contexts
Jinkun Chen, Ju-Sheng Mi, Yaojin Lin |
Knowl. Based Syst. | 1 |
| 2018 | Attribute reduction for multi-label learning with fuzzy rough set
Yaojin Lin, Chenxi Wang 0002, Jinkun Chen |
Knowl. Based Syst. | 4 |
| 2017 | Automatic emotional spoken language text corpus construction from written dialogs in fictionsabstractIn this paper, we propose a novel method to automatically construct emotional spoken language text corpus from written dialogs, and release a large scale Chinese emotional text dataset with short conversations extracted from thousands of fictions using the proposed method. The emotional spoken language transcript resources in Chinese are relatively limited. However, constructing a large scale supervised corpus manually is neither efficient nor low-cost. This motivates us to try alternative efficient and effective approaches. First, we build a small scale emotion dictionary manually instead of a large scale corpus. Each word in dictionary has an emotion tag. Then, we use the emotional words to search emotional dialogs heuristically in fictions and classify them automatically. Second, we share our work to boost the performance of emotion recognition on spoken languages using the proposed new database. The labeled dialogs can be used for supervised learning while the unlabeled ones provide better word embeddings for the semantic level emotion recognition. We use the dialogs corpus as an auxiliary dataset in speech emotion recognition. We carry out experiments on automatic speech recognition (ASR) generated texts from the speech signals in Chinese Natural Emotional Audio-Visual Database (CHEAVD). It is an eight emotion states recognition task. We obtain a baseline average macro precision (MAP) of 37.08% and accuracy of 31.13% in terms of text-based method. With the labeled dialogs to pre-train neural networks and over-sampling the minority classes, we achieve an optimized MAP of 47.50% and the accuracy of 43.91%, which outperforms the baseline by 10.42% and 12.78% respectively. Jinkun Chen, Ming Li 0026 |
ACII | 1 |
| 2017 | Attribute reduction of covering decision systems by hypergraph model
Jinkun Chen, Yaojin Lin, Guoping Lin, Jinjin Li 0001, Yan-Lan Zhang |
Knowl. Based Syst. | 1 |
| 2015 | Relations of reduction between covering generalized rough sets and concept lattices
Jinkun Chen, Jinjin Li 0001, Yaojin Lin, Guoping Lin, Zhou-Ming Ma |
Inf. Sci. | 1 |
| 2015 | The relationship between attribute reducts in rough sets and minimal vertex covers of graphs
Jinkun Chen, Yaojin Lin, Guoping Lin, Jinjin Li 0001, Zhou-Ming Ma |
Inf. Sci. | 1 |
| 2014 | A new nearest neighbor classifier via fusing neighborhood information
Yaojin Lin, Jinjin Li 0001, Menglei Lin, Jinkun Chen |
Neurocomputing | 4 |
| 2014 | Feature selection via neighborhood multi-granulation fusion
Yaojin Lin, Jinjin Li 0001, Peirong Lin, Guoping Lin, Jinkun Chen |
Knowl. Based Syst. | 5 |
| 2013 | Computing connected components of simple undirected graphs based on generalized rough sets
Jinkun Chen, Jinjin Li 0001, Yaojin Lin |
Knowl. Based Syst. | 1 |
| 2012 | An application of rough sets to graph theory
Jinkun Chen, Jinjin Li 0001 |
Inf. Sci. | 1 |