EDBT 2026 Demo / reviewers in the wild / expert
Chaowei Zhang 0001
dblp:188/1352-1
· DBLP profile ↗
26ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Systems, architecture and hardware · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Early Malicious Users Prediction Benchmark for Chinese Esports via Bullet Chats
Xiang Xing, Yi Zhu 0006, Chaowei Zhang 0001, Jipeng Qiang |
ICIC (22) | 4 |
| 2026 | Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance ReasoningabstractThe widespread proliferation of online content has intensified concerns about clickbait, deceptive or exaggerated headlines designed to attract attention. While Large Language Models (LLMs) offer a promising avenue for addressing this issue, their effectiveness is often hindered by Sycophancy, a tendency to produce reasoning that matches users' beliefs over truthful ones, which deviates from instruction-following principles. Rather than treating sycophancy as a flaw to be eliminated, this work proposes a novel approach that initially harnesses this behavior to generate contrastive reasoning from opposing perspectives. Specifically, we design a Self-renewal Opposing-stance Reasoning Generation (SORG) framework that prompts LLMs to produce high-quality ''agree'' and ''disagree'' reasoning pairs for a given news title without requiring ground-truth labels. To utilize the generated reasoning, we develop a local Opposing Reasoning-based Clickbait Detection (ORCD) model that integrates three BERT encoders to represent the title and its associated reasoning. The model leverages contrastive learning, guided by soft labels derived from LLM-generated credibility scores, to enhance detection robustness. Experimental evaluations on three benchmark datasets demonstrate that our method consistently outperforms LLM prompting, fine-tuned smaller language models, and state-of-the-art clickbait detection baselines. Our code is available in https://github.com/126541/ORCD. Chaowei Zhang 0001, Xiansheng Luo, Zewei Zhang, Yi Zhu 0006, Jipeng Qiang, Longwei Wang |
WWW | 1 |
| 2026 | Analyzing bullet chats for recommendation intent identification: Dataset and method
Yi Zhu 0006, Qinqin Han, Yun-Hao Yuan 0001, Chaowei Zhang 0001, Jipeng Qiang, Xindong Wu 0001 |
Artif. Intell. | 4 |
| 2026 | SubAttack: A word-level adversarial textual attack method via antonym substitutionabstractOver the past few years, various word-level textual attack approaches have been proposed to reveal the vulnerability in existing deep neural networks and even large language models (LLMs) for Natural Language Processing (NLP). The textual attack aims to fool existing models into making erroneous predictions by altering the text without affecting the user’s understanding. However, current methods either struggle to construct semantically preserved adversarial texts and altered the semantics of the original text, or fail to consider the semantic perturbation constraints and are prone to invalid adversarial examples. In this paper, we propose an efficient and effective framework SubAttack to address these issues. SubAttack is a word-level adversarial textual attack method via antonym substitution, which replaces semantic indicator keywords to generate high-quality adversarial samples with considering both semantically preservation and semantic perturbation. Specifically, the process first involves tokenizing the text and performing part-of-speech tagging Identifying the semantic indicator keywords. Then, the antonym ranking is designed to decide the substitutions of candidate words to fit the context. Finally, while retaining the original text, the ranked antonyms are integrated into the text and the instructions are added for both semantically preservation and semantic perturbation. Extensive experiments reveal that state-of-the-art (SOTA) LLMs (e.g. Llama and QWen) are still vulnerable to our SubAttack. Further experiments show that the adversarial examples crafted by SubAttack usually have higher quality, exhibit better fluency and barely affect human performance and can bring more robustness improvement to victim models by adversarial training. Chenqi Hua, Yi Zhu 0006, Chaowei Zhang 0001, Yun Li 0010, Yun-Hao Yuan 0001, Jipeng Qiang |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Dualmark: A novel dual watermarking approach for large language models
Zihao Qiang, Jifei Hao, Jipeng Qiang, Yi Zhu 0006, Chaowei Zhang 0001, Yan Liu 0038, Wei Li 0121 |
Inf. Process. Manag. | 5 |
| 2026 | Turning hallucinations into knowledge: Towards identifying clickbait using LLM-generated fallacies
Chaowei Zhang 0001, Zhicong Wang, Zewei Zhang, Yi Zhu 0006, Jipeng Qiang, Yuchao Huang |
Inf. Process. Manag. | 1 |
| 2025 | Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News DetectionabstractThe questionable responses caused by knowledge hallucination may lead to LLMs' unstable ability in decision-making. However, it has never been investigated whether the LLMs' hallucination is possibly usable for generating negative reasoning to assist fake news detection. In this paper, we propose a novel supervised self-reinforced reasoning rectification approach - SR^3 that not only yields common reasonable reasoning for news but also forces LLMs to generate the wrong understandings of news via LLMs reflection for semantic consistency learning. Upon that, we construct a negative reasoning-based news learning model called - NRFE, which leverages positive or negative news-reasoning pairs for learning the semantic consistency between them. To avoid the impact of label-implicated reasoning, we deploy a student model - NRFE-D that only takes news content as input to inspect the performance of our method by distilling the knowledge from NRFE. The experimental results verified on three popular fake news datasets demonstrate the superiority of our method compared with three kinds of baselines including prompting-based LLMs, fine-tuning-based PLMs, and other representative fake news detection methods. Chaowei Zhang 0001, Zongling Feng, Zewei Zhang, Jipeng Qiang, Guandong Xu, Yun Li 0010 |
AAAI | 1 |
| 2025 | ICE: Incremental Subspace Clustering of High-Dimensional Categorical DataabstractSubspace clustering is an effective way to analyze high-dimensional data. The main problems of the conventional subspace clustering techniques are as follows: first, conventional clustering methods can not describe categorical attribute space in more detail; second, most subspace-clustering techniques failure to process dynamic data effectively; finally, lack of effective noise recognition leads to the decline of the efficiency of incremental subspace-clustering analysis. We address the above problems by an incremental subspace-clustering algorithm — called ICE. With attribute subspace constructed by a rough set-based weight computing method, ICE obtains clustering results through initial and incremental clustering stage. Utilizing the original cluster results generated from initial clustering stage, we adopt merging and splitting operation to dynamic adjust cluster-structure in incremental clustering stage. Before achieving the final results, a polymerization-based noise recognition technique is employed to automatically identify noise from sparse clusters without human threshold intervention. We implement ICE on synthetic and real-world datasets. The experimental results reveal that incremental subspace-clustering method can achieves satisfactory performance on extensibility, accuracy and robustness. Ning Pang, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001 |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |
| 2025 | Similarity Metrics: Chebyshev Coulomb Force and Resultant Force for High-Dimensional DataabstractThe similarity metric has garnered widespread attention thanks to its potential applications in the fields of data mining, machine learning, and so on. Due to the interference of “distance concentration” caused by “Curse of dimensionality,” however, existing similarity metrics are inadequate in high-dimensional data analysis. In this study, we propose two innovative similarity metrics—Chebyshev Coulomb force and Chebyshev Coulomb resultant force—anchored on Chebyshev p-norms. In the initial phase, we eliminate dependency relationships among attributes by applying a metric matrix—and the theoretical analysis reveals that the Chebyshev p-norms is capable of mitigating the effect of “distance concentration” among high-dimensional data objects. Next, we devise two similarity metrics—Chebyshev Coulomb force and Chebyshev Coulomb resultant force—by adopting the metric matrix and Chebyshev p-norms. Chebyshev Coulomb force and Chebyshev Coulomb resultant force, being effective in characterizing the similarity among data objects, quantify the deviation of data objects from their respective dataset centers. Additionally, the two metrics alleviate the interference of “distance concentration.” Importantly, the discrepancy of data objects in attribute dimensions is captured by Chebyshev Coulomb force vector, rendering the similarity metric interpretable. By utilizing the UCI dataset, the experimental validation demonstrates the superiority of our similarity metrics, confirming their efficacy in mitigating the interference of “distance concentration.” Compared with the existing similarity metric approaches, the AUC index of outlier detection shows an average improvement of 8.18%—and the ARI, NMI, and F_score indices of clustering are revamped by averages 6.56%, 6.87%, and 6.01%, respectively. Jian Ying Liu, Chaowei Zhang 0001, Min Zhang 0049, Xiao Qin 0001, Jifu Zhang |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | Asynchronous Compaction Acceleration Scheme for Near-data Processing-enabled LSM-tree-based KV StoresabstractLSM-tree-based key-value stores (KV stores) convert random-write requests to sequence-write ones to achieve high I/O performance. Meanwhile, compaction operations in KV stores update SSTables in forms of reorganizing low-level data components to high-level ones, thereby guaranteeing an orderly data layout in each component. Repeated writes caused by compaction (a.k.a. write amplification) impacts I/O bandwidth and overall system performance. Near-data processing (NDP) is one of the effective approaches to addressing this write-amplification issue. Most NDP-based techniques adopt synchronous parallel schemes to perform a compaction task on both the host and its NDP-enabled device. In synchronous parallel compaction schemes, the execution time of compaction is determined by a subsystem that has lower compaction performance coupled by under-utilized computing resources in a NDP framework. To solve this problem, we propose an asynchronous parallel scheme named PStore to improve the compaction performance in KV stores. In PStore, we designed a multi-tasks queue and three priority-based scheduling methods. PStore elects proper compaction tasks to be offloaded in host- and device-side compaction modules. Our proposed cross-leveled compaction mechanism mitigates write amplification induced by asynchronous compaction. PStore featured with the asynchronous compaction mechanism fully utilizes computing resources in both host- and device-side subsystems. Compared with the two popular synchronous compaction modes based on KV stores (TStore and LevelDB), our PStore immensely improves the throughput by up to a factor of 14 and 10.52 with an average of a factor of 2.09 and 1.73, respectively. Hui Sun 0002, Bendong Lou, Deyan Kong, Chaowei Zhang 0001, Jianzhong Huang 0001, Yinliang Yue, Xiao Qin 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2023 | A multi-view ensemble clustering approach using joint affinity matrix
Xueying Niu, Chaowei Zhang 0001, Lihua Hu, Jifu Zhang |
Expert Syst. Appl. | 2 |
| 2023 | A computational approach for real-time detection of fake news
Chaowei Zhang 0001, Ashish Gupta 0004, Xiao Qin 0001, Yi Zhou 0009 |
Expert Syst. Appl. | 1 |
| 2023 | A multi-view subspace representation learning approach powered by subspace transformation relationship
Xueying Niu, Chaowei Zhang 0001, Lihua Hu, Jifu Zhang |
Knowl. Based Syst. | 2 |
| 2023 | Chinese Idiom ParaphrasingabstractAbstract Idioms are a kind of idiomatic expression in Chinese, most of which consist of four Chinese characters. Due to the properties of non-compositionality and metaphorical meaning, Chinese idioms are hard to be understood by children and non-native speakers. This study proposes a novel task, denoted as Chinese Idiom Paraphrasing (CIP). CIP aims to rephrase idiom-containing sentences to non-idiomatic ones under the premise of preserving the original sentence’s meaning. Since the sentences without idioms are more easily handled by Chinese NLP systems, CIP can be used to pre-process Chinese datasets, thereby facilitating and improving the performance of Chinese NLP tasks, e.g., machine translation systems, Chinese idiom cloze, and Chinese idiom embeddings. In this study, we can treat the CIP task as a special paraphrase generation task. To circumvent difficulties in acquiring annotations, we first establish a large-scale CIP dataset based on human and machine collaboration, which consists of 115,529 sentence pairs. In addition to three sequence-to-sequence methods as the baselines, we further propose a novel infill-based approach based on text infilling. The results show that the proposed method has better performance than the baselines based on the established CIP dataset. Jipeng Qiang, Yang Li 0186, Chaowei Zhang 0001, Yun Li 0010, Yi Zhu 0006, Yun-Hao Yuan 0001, Xindong Wu 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | A High-Dimensional Outlier Detection Approach Based on Local Coulomb ForceabstractTraditional outlier detections are inadequate for high-dimensional data analysis due to the interference of distance tending to be concentrated (“curse of dimensionality”). Inspired by the Coulomb’s law, we propose a new high-dimensional data similarity measure vector, which consists of outlier Coulomb force and outlier Coulomb resultant force. Outlier Coulomb force not only effectively gauges similarity measures among data objects, but also fully reflects differences among dimensions of data objects by vector projection in each dimension. More importantly, Coulomb resultant force can effectively measure deviations of data objects from a data center, making detection results interpretable. We introduce a new neighborhood outlier factor, which drives the development of a high-dimensional outlier detection algorithm. In our approach, attribute values with a high deviation degree is treated as interpretable information of outlier data. Finally, we implement and evaluate our algorithm using the UCI and synthetic datasets. Our experimental results show that the algorithm effectively alleviates the interference of “Curse of Dimensionality”. The findings confirm that high-dimensional outlier data originated by the algorithm are interpretable. Pengyun Zhu, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | RT-FEND: Spark-Based Real Time FakE News DetectionabstractFake news is a rampant societal and organizational problem with various social media outlets further aggravating its spread. There is a pressing demand to assist people to identify misinformation from massive amount of news data in a timely manner. Detecting Fake news in a timely manner is critical for mitigating its impact. In this research, we propose a novel approach for detecting fake news in real time, RT-FEND (Real Time- FakE News Detection), which relies on distributed computing paradigm. The proposed methodology utilizes event and topic extraction techniques along with a topic- merging mechanism to process real time news data and reduce the number of topics for managing the curse of dimensionality. We report the findings from several experiments to compare RT-FEND with other systems to benchmark in different system settings. RT-FEND approach is more performance-improved and time-efficient in detecting fake news when compared to other fake news detection baselines. Chaowei Zhang 0001, Ashish Gupta 0004, Hui Sun 0002, Yun Li 0010, Xiao Qin 0001 |
NAS | 1 |
| 2022 | Performance modeling for I/O-intensive applications on virtual machinesabstractAbstract Models for virtual machines running on cloud computing systems. Modeling system behaviors of clouds is a grand challenge because the resource utilization in VMs is heterogeneous due to variability in workload conditions. We address this challenging issue by uniquely (1) objectifying the usage prediction of virtualized resources and (2) predicting the performance trends of programs running on clouds. At the heart of the modeling system, we pay particular attention to CPU cores, disk size, main memory space, and input data volume, which serve as important factors for the developed prediction module. We devise two resource‐utilization prediction algorithms driven by two distinctive sets of I/O and CPU intensive benchmarks, where one algorithm deals with execution time and the other one revolves around input data size. We investigate the correlation between CPU/disk utilization and VM live migrations. Our system aims at not only providing performance optimization for virtualized resources but also ensuring service level agreement (SLA) and Quality of Service (QoS). The model fits the curve quite well, thereby advocating for the efficiency of the algorithm. The case studies conducted in this project draw the comparisons between the performance of striped and monolithic disks as well as bringing forth the problem of cache coherence that causes hindrance to the experiment. We also deal with the cache‐coherence problem to improve the accuracy of our prediction algorithms Tathagata Bhattacharya, Xiaopu Peng, Jianzhou Mao, Chaowei Zhang 0001, Taha Takreeti, Ye Wang 0024, Xiao Qin 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | MiCS-P: Parallel mutual-information computation of big categorical data on spark
Junli Li 0005, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001, Lihua Hu |
J. Parallel Distributed Comput. | 2 |
| 2022 | HIPA: A hybrid load balancing method in SSDs for improved parallelism performance
Hui Sun 0002, Chaowei Zhang 0001, Yinliang Yue, Xiao Qin 0001 |
J. Syst. Archit. | 3 |
| 2021 | Outlier detection from multiple data sources
Xujun Zhao, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001 |
Inf. Sci. | 3 |
| 2020 | Computing Mutual Information of Big Categorical Data and Its Application to Feature GroupingabstractThis paper develops a parallel computing system - MiCS - for mutual information of big categorical data on the Spark computing platform. The MiCS algorithm is conductive to processing a large amount and strong repeatability of mutual-information calculation among feature pairs by applying a column-wise transformation scheme. And to improve the efficiency of the MiCS and the utilization rate of Spark cluster resources, we adopt a virtual partitioning scheme to achieve balanced load while mitigating the data skewness problem in the Spark Shuffle process. Junli Li 0005, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001 |
ICDE | 2 |
| 2020 | A popularity-aware reconstruction technique in erasure-coded storage systems
Xiaopu Peng, Chaowei Zhang 0001, Taha Khalid Al Tekreeti, Jianzhou Mao, Xiao Qin 0001, Jianzhong Huang 0001 |
J. Parallel Distributed Comput. | 3 |
| 2019 | PUMA: Parallel subspace clustering of categorical data using multi-attribute weights
Ning Pang, Jifu Zhang, Chaowei Zhang 0001, Xiao Qin 0001, Jianghui Cai |
Expert Syst. Appl. | 3 |
| 2019 | Parallel Hierarchical Subspace Clustering of Categorical DataabstractParallel clustering is an important research area of big data analysis. The conventional Hierarchical Agglomerative Clustering (HAC) techniques are inadequate to handle big-scale categorical datasets due to two drawbacks. First, HAC consumes excessive CPU time and memory resources; and second, it is non-trivial to decompose clustering tasks into independent sub-tasks executed in parallel. We solve these two problems by a MapReduce-based hierarchical subspace-clustering algorithm - called PAPU - using LSH-based data partitioning. PAPU is conducive to partitioning a large-scale dataset into multiple independent sub-datasets, into which similar data objects are mapped. Advocating parallel computing, PAPU obtains sub-clusters corresponding to respective attribute subspaces from independent chunks in the local clustering phase. To improve the accuracy of approximated clustering results, PAPU measures various scale clusters by applying the hierarchical clustering scheme to iteratively merge sub-clusters during the global clustering phase. We implement PAPU on a 24-node Hadoop computing platform. The experimental results reveal that hierarchical subspace-clustering coupled with the data-partitioning strategy achieves high clustering efficiency on both synthetic and real-world large-scale datasets. The experiments also demonstrate that PAPU delivers superior performance in terms of extensibility and scalability (e.g., a nearly linear speedup). Ning Pang, Jifu Zhang, Chaowei Zhang 0001, Xiao Qin 0001 |
IEEE Trans. Computers | 3 |
| 2019 | GreenDB: Energy-Efficient Prefetching and Caching in Database ClustersabstractIn this study, we propose an energy-efficient database system called GreenDB running on clusters. GreenDB applies a workload-skewness strategy by managing hot nodes coupled with a set of cold nodes in a database cluster. GreenDB fetches popular data tables to hot nodes, aiming to keep cold nodes in the low-power mode in increased time periods. GreenDB is conducive to reducing the number of power-state transitions, thereby lowering energy-saving overhead. A prefetching model and an energy saving model are seamlessly integrated into GreenDB to facilitate the power management in database clusters. We quantitatively evaluate GreenDB's energy efficiency in terms of managing, fetching, and storing data. We compare GreenDB's prefetching strategy with the one implemented in Postgresql. Experimental results indicate that GreenDB conserves the energy consumption of the existing solution by up to 98.4 percent. The findings show that the energy efficiency of GreenDB can be optimized by tuning system parameters, including table size, hit rates, number of nodes, number of disks, and inter-arrival delays. Yi Zhou 0009, Shubbhi Taneja, Chaowei Zhang 0001, Xiao Qin 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Miner*: A Weighted Distance Sum based Outlier Mining System of Star Spectrum DataabstractExisting distance-based outlier mining methods do not consider the impact of each attribute's importance degree, thereby resulting in poor mining accuracies. To address this problem, we propose a new outlier mining algorithm – Miner* – that makes use of information entropy and Weighted Distance Sum to substantially improve mining accuracies. Miner* employs information entropy to determine weight values indicating the importance degrees of data attributes. An input dataset is reduced by Miner* through the neighbour-radius-based pruning technologies. Thus, Miner* obtains a candidate outlier set by removing any data objects that are unlikely to be outliers. Miner* calculates the weighted distance sum value Wkof each object in the candidate outlier set; Wkvalue ranks the top n to be regarded as outliers. Due to the sum of distance, which takes full advantage of the clustering characteristics of the dataset, edge distribution data objects and local outliers can be effectively mined out. To demonstrate the effectiveness of the Miner* algorithm, we implement Miner* in a prototype system to detect star spectrum data objects with abnormal characteristic lines. Our experimental results show that the algorithm in Miner* achieves high accuracy, high scalability, and low man-made influence by utilizing UCI and star spectrum dataset. Our results also confirm that Miner* is feasible and effective in mining spectrum data with abnormal characteristic lines from massive star spectrum dataset. Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001, Sulan Zhang |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |