Chin-Teng Lin

dblp:34/1447 · DBLP profile ↗
← Back
19ranked-venue papers in the field
1as first author
11since 2021 · last 2026
0000-0001-8371-8197ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 8Database Systems & Data Management · 5 (1 first)Data Mining & Knowledge Discovery · 2Other / Interdisciplinary · 2Information Retrieval & Web Search · 1Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2026 When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse
abstract
Retrieval-augmented generation (RAG) is indispensable for enhancing large language models. However, RAGs are increasingly susceptible to poisoning attacks, in which adversarial documents are injected to manipulate generator outputs. Previous methods rely on output-side signals such as perplexity and consistency checks to detect such attacks. Nevertheless, our analysis reveals that deliberate attacks often induce false confidence, where poisoned outputs exhibit even lower perplexity than benign ones, rendering uncertainty-based detection ineffective. To address this challenge, we explore the internal dynamics of the generator and identify a distinctive signature termed Attention Collapse. Unlike the dispersed attention in benign generations, attacked generations exhibit a decrease in entropy as attention concentrates on poisoned documents. Building on these findings, we propose D-SCAN (Document-level Signal Collapse Analysis), a lightweight detection framework that monitors attention dynamics to identify attacked generations. Extensive experiments on multiple attack benchmarks demonstrate the effectiveness of our method. Moreover, D-SCAN can detect attacks even when they fail to alter the final answer. Code is available at https://github.com/yingtaoren/D-Scan.git.
Yingtao Ren, Yiwei Fu, Xiao Luo 0001, Chin-Teng Lin
SIGIR6
2026 Online Learning in Open Data Space
Zhi Cao 0001, Peijia Qin, Chin-Teng Lin, Xin Yao 0001
IEEE Trans. Knowl. Data Eng.3
2026 A CDGFN-Based Quantum Multisource Information Fusion With Its Application in Time Series Classification
abstract
Time series classification (TSC) is a critical area with broad applications. In the field of evidence theory, quantum evidence theory (QET) offers a promising framework for onedimensional TSC tasks, leveraging the capabilities of quantum basic probability amplitude (QBPA) to capture two-dimensional uncertainty. However, as the first step for the application of QET to TSC, how to construct QBPA still remains an open issue. In this paper, a novel approach to generate QBPA is devised. Specifically, we first apply the discrete Fourier transform (DFT) to the original data, extracting two-dimensional features embedded in the magnitude and phase from the frequency domain based on the front-few multi-frequency components, achieved by setting a threshold frequency index (TFI) to limit the frequencies considered. Next, we introduce the complex dual gaussian fuzzy number (CDGFN) as a carrier for QBPA, effectively representing two-dimensional uncertainty in the data. A CDGFN-based multisource information fusion (CDGFN-MSIF) algorithm for decision-making is proposed to combine information from different frequency components. Finally, the decisionmaking algorithm is validated on multiple time series datasets. Experimental results highlight the superior performance of the proposed approach over other state-of-the-art models, demonstrating its effectiveness and enhanced classification accuracy.
Junhao Yu, Fuyuan Xiao 0001, Zehong Cao, Chin-Teng Lin
IEEE Trans. Knowl. Data Eng.5
2024 RCAR-UNet: Retinal vessel segmentation network algorithm via novel rough attention mechanism
Weiping Ding 0001, Jiashuang Huang, Hengrong Ju, Chongsheng Zhang, Guang Yang 0006, Chin-Teng Lin
Inf. Sci.7
2024 Fractal Belief Rényi Divergence With its Applications in Pattern Classification
abstract
Multisource information fusion is a comprehensive and interdisciplinary subject. Dempster-Shafer (D-S) evidence theory copes with uncertain information effectively. Pattern classification is the core research content of pattern recognition, and multisource information fusion based on D-S evidence theory can be effectively applied to pattern classification problems. However, in D-S evidence theory, highly-conflicting evidence may cause counterintuitive fusion results. Belief divergence theory is one of the theories that are proposed to address problems of highly-conflicting evidence. Although belief divergence can deal with conflict between evidence, none of the existing belief divergence methods has considered how to effectively measure the discrepancy between two pieces of evidence with time evolutionary. In this study, a novel fractal belief Rényi (FBR) divergence is proposed to handle this problem. We assume that it is the first divergence that extends the concept of fractal to R/'enyi divergence. The advantage is measuring the discrepancy between two pieces of evidence with time evolution, which satisfies several properties and is flexible and practical in various circumstances. Furthermore, a novel algorithm for multisource information fusion based on FBR divergence, namely FBReD-based weighted multisource information fusion, is developed. Ultimately, the proposed multisource information fusion algorithm is applied to a series of experiments for pattern classification based on real datasets, where our proposed algorithm achieved superior performance.
Yingcheng Huang, Fuyuan Xiao 0001, Zehong Cao, Chin-Teng Lin
IEEE Trans. Knowl. Data Eng.4
2023 Online Ensemble of Ensemble OVA Framework for Class Evolution with Dominant Emerging Classes
abstract
In real-world data stream mining, the composition of classes undergoes unpredictable changes, giving rise to the challenge of class evolution, encompassing class emergence, disappearance, and reoccurrence. However, most existing approaches require the storage of past data to adapt their model. While some studies have focused on online learning approaches, they are built on an underlying assumption that the number of instances in any single class is consistently less than the sum of other classes. This assumption becomes invalid when a class emerges with a dominant amount, e.g., news about a pandemic outbreak, harming the performance of existing methods. In this paper, we thoroughly investigate this scenario and propose a novel online ensemble of ensemble one-versus-all framework (EEOF) to handle class evolution adaptively. The novel ensemble of ensemble architecture boosts diversity in each one-versus-all classifier. A novel adaptive model adaptation method is also designed to balance the error feedback between the emerging class and the other classes. A confidence-triggered fallback mode is integrated to prevent performance drop due to a wrong decision regarding class disappearance. Experimental studies are conducted on both synthetic and real-world data streams to show that our method achieves higher accuracy in diverse class evolution scenarios compared with the state-of-the-art method, particularly when classes emerge with dominant amounts.
Zhi Cao 0001, Chin-Teng Lin
ICDM3
2023 Special issue on Recent Advances in Fuzzy Deep Learning for Uncertain Medicine Data
Weiping Ding 0001, Jun Liu 0001, Chin-Teng Lin, Dariusz Mrozek
Inf. Sci.3
2023 Belief f-divergence for EEG complexity evaluation
Xingjian Song, Fuyuan Xiao 0001, Zehong Cao, Chin-Teng Lin
Inf. Sci.5
2023 Toward multi-target self-organizing pursuit in a partially observable Markov game
Lijun Sun 0002, Chao Lyu, Ye Shi 0001, Yuhui Shi 0001, Chin-Teng Lin
Inf. Sci.6
2023 A Complex Weighted Discounting Multisource Information Fusion With its Application in Pattern Classification
abstract
Complex evidence theory (CET) is an effective method for uncertainty reasoning in knowledge-based systems with good interpretability that has recently attracted much attention. However, approaches to improve the performance of uncertainty reasoning in CET-based expert systems remains an open issue. One key to performance improvement is the adequate management of conflict from multisource information. In this paper, a generalized correlation coefficient, namely, the complex evidential correlation coefficient (CECC), is proposed for the complex mass functions or complex basic belief assignments (CBBAs) in CET. On this basis, a complex conflict coefficient is proposed to measure the conflict between CBBAs; when CBBAs turn into classic BBAs, the complex correlation and conflict coefficients will degrade into traditional coefficients. The complex conflict coefficient satisfies nonnegativity, symmetry, boundedness, extreme consistency, and insensitivity to refinement properties, which are desirable for conflict measurement. Several numerical examples validate through comparisons the superiority of the complex conflict coefficient. In this context, a weighted discounting multisource information fusion algorithm, which is called the CECC-WDMSIF, is designed based on the CECC to improve the performance of CET-based expert systems. By applying the CECC-WDMSIF method to the pattern classification of diverse real-world datasets, it is demonstrated that the proposed CECC-WDMSIF outperforms well-known related approaches with higher classification accuracy and robustness.
Fuyuan Xiao 0001, Zehong Cao, Chin-Teng Lin
IEEE Trans. Knowl. Data Eng.3
2021 Text-line-up: Don't Worry About the Caret
Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein
ICDAR (3)3
2020 Current trends of granular data mining for biomedical data analysis
Weiping Ding 0001, Chin-Teng Lin, Alan Wee-Chung Liew, Isaac Triguero, Wenjian Luo
Inf. Sci.2
2019 Detecting Named Entities in Unstructured Bengali Manuscript Images
abstract
In this paper, we undertake a task to find named entities directly from unstructured handwritten document images without any intermediate text/character recognition. Here, we do not receive any assistance from natural language processing. Therefore, it becomes more challenging to detect the named entities. We work on Bengali script which brings some additional hurdles due to its own unique script characteristics. Here, we propose a new deep neural network-based architecture to extract the latent features from a text image. The embedding is then fed to a BLSTM (Bidirectional Long Short-Term Memory) layer. After that, the attention mechanism is adapted to an approach for named entity detection. We perform experimentation on two publicly-available offline handwriting repositories containing 420 Bengali handwritten pages in total. The experimental outcome of our system is quite impressive as it attains 95.43% balanced accuracy on overall named entity detection.
Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein
ICDAR3
2019 Semi-supervised feature learning for improving writer identification
Shiming Chen 0002, Yisong Wang 0004, Chin-Teng Lin, Weiping Ding 0001, Zehong Cao
Inf. Sci.3
2019 Active learning for regression using greedy sampling
Dongrui Wu, Chin-Teng Lin, Jian Huang 0001
Inf. Sci.2
2018 Privacy-Preserving Linear Regression for Brain-Computer Interface Applications
abstract
Many machine learning (ML) applications rely on large amounts of personal data for training and inference. Among the most intimate exploited data sources is electroencephalogram (EEG) data. The emergence of consumer -grade, low-cost brain -computer interfaces (BCIs) and corresponding software development kits' is bringing the use of BCI within reach of application developers. The access that BCI applications have to neural signals rightly raises privacy concerns. Application developers can easily gain knowledge beyond the professed scope from unprotected EEG signals, including passwords, ATM PINs, and other personal data. The challenge is how to engage in meaningful ML with EEG data while protecting the privacy of users.
Anisha Agarwal, Rafael Dowsley, Nicholas D. McKinney, Dongrui Wu, Chin-Teng Lin, Martine De Cock, Anderson C. A. Nascimento
IEEE BigData5
2018 Minority Oversampling in Kernel Adaptive Subspaces for Class Imbalanced Datasets
abstract
The class imbalance problem in machine learning occurs when certain classes are underrepresented relative to the others, leading to a learning bias toward the majority classes. To cope with the skewed class distribution, many learning methods featuring minority oversampling have been proposed, which are proved to be effective. To reduce information loss during feature space projection, this study proposes a novel oversampling algorithm, named minority oversampling in kernel adaptive subspaces (MOKAS), which exploits the invariant feature extraction capability of a kernel version of the adaptive subspace self-organizing maps. The synthetic instances are generated from well-trained subspaces and then their pre-images are reconstructed in the input space. Additionally, these instances characterize nonlinear structures present in the minority class data distribution and help the learning algorithms to counterbalance the skewed class distribution in a desirable manner. Experimental results on both real and synthetic data show that the proposed MOKAS is capable of modeling complex data distribution and outperforms a set of state-of-the-art oversampling algorithms.
Chin-Teng Lin, Tsung-Yu Hsieh, Yu-Ting Liu, Yang-Yin Lin, Chieh-Ning Fang, Yu-Kai Wang, Gary G. Yen, Nikhil R. Pal, Chun-Hsiang Chuang
IEEE Trans. Knowl. Data Eng.1
2014 Takagi-Sugeno-Kang type collaborative fuzzy rule based system
abstract
In this paper, a Takagi-Sugeno-Kang (TSK) type collaborative fuzzy rule based system is proposed with the help of knowledge learning ability of collaborative fuzzy clustering (CFC). The proposed method split a huge dataset into several small datasets and applying collaborative mechanism to interact each other and this process could be helpful to solve the big data issue. The proposed method applies the collective knowledge of CFC as input variables and the consequent part is a linear combination of the input variables. Through the intensive experimental tests on prediction problem, the performance of the proposed method is as higher as other methods. The proposed method only uses one half information of given dataset for training process and provide an accurate modeling platform while other methods use whole information of given dataset for training.
Kuang-Pen Chou, Mukesh Prasad, Yang-Yin Lin, Sudhanshu Joshi, Chin-Teng Lin, Jyh-Yeong Chang
CIDM5
2013 Fuzzy adaptive synchronization of time-reversed chaotic systems via a new adaptive control strategy
Shih-Yu Li, Cheng-Hsiung Yang, Shi-An Chen, Li-Wei Ko, Chin-Teng Lin
Inf. Sci.5