VLDB 2026 Research / reviewers in the wild / expert
Chongsheng Zhang
dblp:82/4043
· DBLP profile ↗
49ranked-venue papers
23as first author
31since 2021 · last 2026
0000-0003-1632-7238ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 18 first-author · 18 since 2021Databases, data management, data science and information retrieval · 14 · 8 first-author · 7 since 2021Computer networks · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit DynamicsabstractYuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher, Christian Heumann, Chongsheng Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher, Christian Heumann, Chongsheng Zhang |
ACL (1) | 6 |
| 2026 | Parameter-Efficient and Adaptive Fine-Tuning for Long-Tailed Ancient Characters Recognition
Aouaidjia Kamel, Constantine Kotropoulos, Chongsheng Zhang |
ICDAR (3) | 4 |
| 2026 | F2C-Net: A privacy-preserving federated transformer-RL architecture for real-time control in multi-domain SD-IoT systems
Samra Zafar, Bakhtawar Zafar, Gaojuan Fan, Chongsheng Zhang |
Comput. Networks | 4 |
| 2026 | Improving action segmentation via explicit similarity measurement
Aouaidjia Kamel, Aofan Li, Chongsheng Zhang |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Edge-Optimized Lightweight and Transformer Backbones for Real-Time Road Damage Detection in IIoT SystemsabstractAccurate and efficient road damage detection is critical for maintaining urban infrastructure and ensuring public safety in Intelligent Internet of Things (IIoT) systems. There remains a significant challenge to achieving a balance between detection accuracy and real-time inference on resource-constrained edge devices despite advances in deep learning. This paper addresses this gap by enhancing the YOLOv9c object detection framework with two distinct backbone architectures: MobileNet V3-Small, which is a lightweight convolutional neural network optimized for edge deployment, and Swin Transformer, which is a hierarchical vision transformer that captures rich contextual features. We present a systematic, dual-backbone performance benchmark that quantifies the critical trade-off between computational efficiency and detection precision, which is essential for guiding IIoT deployment strategies. We conducted experiments on the Street View Road Damage Detection (SVRDD) dataset to evaluate detection accuracy, computational efficiency, and latency. The MobileNet backbone achieves the highest mean Average Precision ([email protected]) of 74.0% (a 1.5% gain over baseline) and recall of 68.8% (a 7.3% gain over baseline), demonstrating improved accuracy while maintaining a low inference time on a baseline GPU, indicating its suitability for deployment on IoT edge devices. Importantly, the MobileNet variant reduces the parameter count from 25.6 M to 2.54 M and the Giga Floating-point Operations Per Second (GFLOPs) from 102.3 to 0.49, making it more efficient for IIoT edge devices. Both backbones performed better than the YOLOv9c baseline model in terms of accuracy, thus providing scalable and practical solutions for real-time infrastructure monitoring. These findings contribute to the development of intelligent, efficient, and scalable object detection systems tailored for smart city and IIoT environments. Hafiz Muhammad Sanaullah Badar, Israr Hussain, Ali Kashif Bashir, Nazik Alturki, Gaojuan Fan, Chongsheng Zhang |
IEEE Internet Things J. | 6 |
| 2026 | ATTA-FL-Lite: Lightweight Byzantine-Robust Federated Learning for Resource-Constrained Medical IoT DevicesabstractFederated learning (FL) enables privacy-preserving analytics at the medical IoT (IoMT) edge but is vulnerable to model poisoning and distribution shift. We presentATTA-FL-Lite, a lightweight aggregation rule that admits a client update only when three tests are jointly satisfied: (i) scale conformity via a median–absolute–deviation (MAD)z-score, (ii) directional alignment via cosine similarity to a coordinate-wise median reference, and (iii) non-degradation via a validation-lossz-score computed on a small, centrally held clean set. If no update passes, a coordinate-wise median fallback is used. We provide sub-Gaussian tail bounds for the loss test and an expected one-step descent bound forL-smooth objectives under benign mean-alignment and an accepted-set composition assumption; the analysis does not require a positive cosine threshold. Experiments on MNIST, Fashion-MNIST, and PathMNIST, under IID and non-IID partitions with up to 40% adversaries across four attack families, show that ATTA-FL-Lite maintains accuracy representatively ≈ 0.78–0.98 across Tiny/Small/Medium CNNs, reliably filters magnitude/noise attacks, and remains competitive against sign-flip. Server runtime scales approximately linearly with the number of participating clients at fixed model and validation sizes. These results indicate that ATTA-FL-Lite offers practical robustness for FL in resource-constrained IoMT deployments without cryptographic overhead or trusted root data beyond a small validation set. Hafiz Muhammad Sanaullah Badar, Nadeem Iqbal 0003, Khalid Mahmood 0002, Khan Muhammad 0001, Gaojuan Fan, Chongsheng Zhang |
IEEE Internet Things J. | 6 |
| 2026 | Distance-Gradient-Based Convex Optimization for Efficient Near-Optimal Coverage in WSNsabstractCoverage optimization in Wireless Sensor Networks is a fundamental yet NP-hard problem that directly affects monitoring quality and efficiency. Existing solutions mainly rely on meta-heuristic algorithms that use fitness-based evaluations, which often incur high computational overhead, slow convergence, and limited scalability, particularly in real-time or high-precision monitoring scenarios. In this paper, we examine the relationship between effective coverage area and redundant distances in an analytical manner. We then propose reformulating WSN coverage optimization as a Distance-Gradient based convex optimization problem, which can be subsequently solved using the first-order Gradient Descent algorithm or the second-order quasi-Newton algorithm. Extensive comparative experiments against five representative meta-heuristic methods, the Virtual Force Algorithm (VFA) and a general convex optimization algorithm (CVX), demonstrate that our approach achieves near-optimal coverage while preserving network connectivity within milliseconds, highlighting its advantages over existing methods for WSNs coverage optimization. Gaojuan Fan, Feitao Li, Chongsheng Zhang, Hafiz Muhammad Sanaullah Badar, Christian Heumann |
IEEE Internet Things J. | 3 |
| 2026 | PoisonShield-FL-NIDS: A Robust Defense Against Poisoning Attacks in Federated Learning Intrusion DetectionabstractFederated learning (FL) has emerged as a privacy-preserving paradigm for collaborative intrusion detection in networked environments. However, it remains vulnerable to Poisoning Attacks (PA) wherein malicious clients can corrupt the global model through deceptive updates. To address this, we propose PoisonShield-FL-NIDS, a robust FL-based intrusion detection system that integrates client-side anomaly filtering with trust-aware aggregation to defend against poisoned contributions. Experimental evaluation under varying levels of adversarial influence demonstrates that PoisonShield-FL-NIDS achieves superior performance across key metrics, attaining 93% accuracy, 91% precision, 94% recall, and an AUC of 0.96, while maintaining a low robustness index RI < 0.05 even with 30% compromised clients. Compared to baseline FL models such as FL-CNN and FedACNN, our framework demonstrates faster convergence and higher resilience with only a marginal increase in communication overhead. Nadeem Iqbal 0003, Michael G. Madden, Gaojuan Fan, Chongsheng Zhang, Hafiz Muhammad Sanaullah Badar |
IEEE Internet Things J. | 5 |
| 2026 | ASRec: adaptive sequential recommendation with dynamic and periodic preferences capturing
Wenlong Hao, Ghufran Ahmad Khan, Gaojuan Fan, Chongsheng Zhang |
Knowl. Inf. Syst. | 4 |
| 2026 | Anomal-EFD: A self-supervised model for anomaly detection in dynamic IoT networks
Gaojuan Fan, Qingyi Huang, Hafiz Muhammad Sanaullah Badar, Chongsheng Zhang |
Peer Peer Netw. Appl. | 5 |
| 2025 | QuinNet: Quintuple u-shape networks for scale- and shape-variant lesion segmentation
Gaojuan Fan, Ruixue Xia, Funa Zhou, Chongsheng Zhang |
Appl. Intell. | 5 |
| 2025 | Spatio-temporal invariant descriptors for skeleton-based human action recognition
Aouaidjia Kamel, Chongsheng Zhang, Ioannis Pitas |
Inf. Sci. | 2 |
| 2025 | Open-set long-tailed recognition via orthogonal prototype learning and false rejection correction
Binquan Deng, Aouaidjia Kamel, Chongsheng Zhang |
Neural Networks | 3 |
| 2025 | UCR: A unified character-radical dual-supervision framework for accurate Chinese character recognition
Chongsheng Zhang |
Pattern Recognit. | 2 |
| 2025 | A Systematic Review on Long-Tailed LearningabstractLong-tailed data are a special type of multiclass imbalanced data with a very large amount of minority/tail classes that have a very significant combined influence. Long-tailed learning (LTL) aims to build high-performance models on datasets with long-tailed distributions that can identify all the classes with high accuracy, in particular the minority/tail classes. It is a cutting-edge research direction that has attracted a remarkable amount of research effort in the past few years. In this article, we present a comprehensive survey of the latest advances in long-tailed visual learning. We first propose a new taxonomy for LTL, which consists of eight different dimensions, including data balancing, neural architecture, feature enrichment, logits adjustment, loss function, bells and whistles, network optimization, and posthoc processing techniques. Based on our proposed taxonomy, we present a systematic review of LTL methods, discussing their commonalities and alignable differences. We also analyze the differences between imbalance learning and LTL. Finally, we discuss prospects and future directions in this field. Chongsheng Zhang, George Almpanidis, Gaojuan Fan, Binquan Deng, Ji Liu 0003, Aouaidjia Kamel, Paolo Soda, João Gama 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | On mask-based image set desensitization with recognition support
Ji Liu 0003, Chongsheng Zhang, Dejing Dou |
Appl. Intell. | 4 |
| 2024 | RCAR-UNet: Retinal vessel segmentation network algorithm via novel rough attention mechanism
Weiping Ding 0001, Jiashuang Huang, Hengrong Ju, Chongsheng Zhang, Guang Yang 0006, Chin-Teng Lin |
Inf. Sci. | 5 |
| 2024 | imFTP: Deep imbalance learning via fuzzy transition and prototypical learning
Yaxin Hou, Weiping Ding 0001, Chongsheng Zhang |
Inf. Sci. | 3 |
| 2024 | TextFuse: Fusing Deep Scene Text Detection Models for Enhanced Performance
Xianjin Shi, Guowen Peng, Xiajiong Shen, Chongsheng Zhang |
Multim. Tools Appl. | 4 |
| 2024 | Neighbor-Enhanced Representation Learning for Link Prediction in Dynamic Heterogeneous Attributed NetworksabstractDynamic link prediction aims to predict future connections among unconnected nodes in a network. It can be applied for friend recommendations, link completion, and other tasks. Network representation learning algorithms have demonstrated considerable effectiveness in various prediction tasks. However, most network representation learning algorithms are based on homogeneous networks and static networks for link prediction that do not consider rich semantic and dynamic information. Additionally, existing dynamic network representation learning methods neglect the neighborhood interaction structure of the node. In this work, we design a neighbor-enhanced dynamic heterogeneous attributed network embedding method (NeiDyHNE) for link prediction. In light of the impressive achievements of the heuristic methods, we learn the information of common neighbors and neighbors’ interaction in heterogeneous networks to preserve the neighbors proximity and common neighbors proximity. NeiDyHNE encodes the attributes and neighborhood structure of nodes as well as the evolutionary features of the dynamic network. More specifically, NeiDyHNE consists of the hierarchical structure attention module and the convolutional temporal attention module. The hierarchical structure attention module captures the rich features and semantic structure of nodes. The convolutional temporal attention module captures the evolutionary features of the network over time in dynamic heterogeneous networks. We evaluate our method and various baseline methods on the dynamic link prediction task. Experimental results demonstrate that our method is superior to baseline methods in terms of accuracy. Wei Wang 0012, Chongsheng Zhang, Weiping Ding 0001, Bin Wang 0062, Yaguan Qian, Zhen Han 0001, Chunhua Su |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | Quality-Aware Self-Training on Differentiable Synthesis of Rare Relational DataabstractData scarcity is a very common real-world problem that poses a major challenge to data-driven analytics. Although a lot of data-balancing approaches have been proposed to mitigate this problem, they may drop some useful information or fall into the overfitting problem. Generative Adversarial Network (GAN) based data synthesis methods can alleviate such a problem but lack of quality control over the generated samples. Moreover, the latent associations between the attribute set and the class labels in a relational data cannot be easily captured by a vanilla GAN. In light of this, we introduce an end-to-end self-training scheme (namely, Quality-Aware Self-Training) for rare relational data synthesis, which generates labeled synthetic data via pseudo labeling on GAN-based synthesis. We design a semantic pseudo labeling module to first control the quality of the generated features/samples, then calibrate their semantic labels via a classifier committee consisting of multiple pre-trained shallow classifiers. The high-confident generated samples with calibrated pseudo labels are then fed into a semantic classification network as augmented samples for self-training. We conduct extensive experiments on 20 benchmark datasets of different domains, including 14 industrial datasets. The results show that our method significantly outperforms state-of-the-art methods, including two recent GAN-based data synthesis schemes. Codes are available at https://github.com/yaxinhou/QAST. Chongsheng Zhang, Yaxin Hou, Ke Chen 0004, Shuang Cao, Gaojuan Fan, Ji Liu 0003 |
AAAI | 1 |
| 2023 | SACA-UNet:Medical Image Segmentation Network Based on Self-Attention and ASPPabstractIn recent years, deep learning based techniques have been successfully applied to medical image segmentation, which plays an important role in intelligent lesion analysis and disease diagnosis. At present, the mainstream segmentation models are primarily based on the U-Net model for extracting local features through multi-layer convolution, which lacks global information and the multi-scale semantic information interaction between the Encoder and Decoder process, leading to sub-optimal segmentation performance. To address such issues, in this work we propose a new medical image segmentation network, namely SACA-UNet, which improves the U-Net model via the self-attention and cross atrous spatial pyramid pooling (Cross-ASPP) mechanisms. In specific, SACA-UNet first utilizes the self- attention mechanism to capture the global feature, it next devises a Cross-ASPP module to extract and fuse features of varying reception fields to prompt multi-scale semantic interaction. We evaluate the segmentation performance of our proposed model on four benchmark datasets including the ISIC2018, BUSI, CVC- ClinicDB, and COVID-19 datasets, in terms of both the Dice coefficient and IoU metrics. Experimental results demonstrate that SACA-UNet remarkably outperforms the baseline methods. Gaojuan Fan, Chongsheng Zhang |
CBMS | 3 |
| 2023 | An empirical study on the joint impact of feature selection and data resampling on imbalance classification
Chongsheng Zhang, Paolo Soda, Jingjun Bi, Gaojuan Fan, George Almpanidis, Weiping Ding 0001 |
Appl. Intell. | 1 |
| 2023 | Correction to: An empirical study on the joint impact of feature selection and data resampling on imbalance classification
Chongsheng Zhang, Paolo Soda, Jingjun Bi, Gaojuan Fan, George Almpanidis, Weiping Ding 0001 |
Appl. Intell. | 1 |
| 2023 | A personalized federated learning-based fault diagnosis method for data suffering from network attacks
Funa Zhou, Chongsheng Zhang, Chenglin Wen, Tianzhen Wang |
Appl. Intell. | 3 |
| 2023 | A unified deep semi-supervised graph learning scheme based on nodes re-weighting and manifold regularizationabstractIn recent years, semi-supervised learning on graphs has gained importance in many fields and applications. The goal is to use both partially labeled data (labeled examples) and a large amount of unlabeled data to build more effective predictive models. Deep Graph Neural Networks (GNNs) are very useful in both unsupervised and semi-supervised learning problems. As a special class of GNNs, Graph Convolutional Networks (GCNs) aim to obtain data representation through graph-based node smoothing and layer-wise neural network transformations. However, GCNs have some weaknesses when applied to semi-supervised graph learning: (1) it ignores the manifold structure implicitly encoded by the graph; (2) it uses a fixed neighborhood graph and focuses only on the convolution of a graph, but pays little attention to graph construction; (3) it rarely considers the problem of topological imbalance. To overcome the above shortcomings, in this paper, we propose a novel semi-supervised learning method called Re-weight Nodes and Graph Learning Convolutional Network with Manifold Regularization (ReNode-GLCNMR). Our proposed method simultaneously integrates graph learning and graph convolution into a unified network architecture, which also enforces label smoothing through an unsupervised loss term. At the same time, it addresses the problem of imbalance in graph topology by adaptively reweighting the influence of labeled nodes based on their distances to the class boundaries. Experiments on 8 benchmark datasets show that ReNode-GLCNMR significantly outperforms the state-of-the-art semi-supervised GNN methods.1 Fadi Dornaika, Jingjun Bi, Chongsheng Zhang |
Neural Networks | 3 |
| 2023 | Object-centric Contour-aware Data Augmentation Using Superpixels of Varying GranularityabstractRegional dropout strategies have demonstrated to be very effective in improving both the performance and the generalization capability of deep learning models. However, when such strategies are performed in a totally random manner, the background noise and label mismatch problems arise. To tackle such problems, existing approaches typically focus on regions with the highest distinctiveness. Yet, there are two main drawbacks of existing approaches: (I) Many existing region-based augmentation methods can only use rectangular regions, resulting in the loss of object contour information; (II) Deterministic selection of the most discriminative regions leads to poor diversification in data augmentation. In fact, a trade-off is needed between diversification and concentration, which can decrease the undesirable noise. In this paper, we propose a novel object-centric contour-aware CutMix data augmentation strategy with arbitrary- shape and size superpixel supports, which is hereafter referred to as OcCaMix for short. It not only captures the most discriminative regions, but also effectively preserves the contour details of the objects. Moreover, it enables the search of natural object parts of different sizes. Extensive experiments on a large number of benchmark datasets show that OcCaMix significantly outperforms state-of-the-art CutMix based data augmentation methods in classification tasks. The source codes and trained models are available at https://github.com/DanielaPlusPlus/OcCaMix. Fadi Dornaika, Danyang Sun, Karim Hammoudi, Jinan Charafeddine, Adnane Cabani, Chongsheng Zhang |
Pattern Recognit. | 6 |
| 2022 | Parallel High Utility Itemset Mining
Gaojuan Fan, Huaiyuan Xiao, Chongsheng Zhang, George Almpanidis, Philippe Fournier-Viger, Hamido Fujita |
IEA/AIE | 3 |
| 2022 | Data-Driven Oracle Bone Rejoining: A Dataset and Practical Self-Supervised Learning SchemeabstractOracle Bone Inscriptions (OBI) is one of the oldest scripts in the world. The rejoining of Oracle Bone (OB) fragments is of vital importance to the research of ancient scripts and history. Although significant progress has been achieved in the past decades, the rejoining work still heavily relies on domain knowledge and manual work, thus remains a low efficient and time-consuming process Therefore, an automatic and practical algorithm/system for OB rejoining is of great value to the OBI community. To this end, we collect a real-world dataset for rejoining Oracle Bone fragments, namely OB-Rejoin, which consists of 998 OB rubbing images that suffer from low quality image problems, due to intrinsic underground eroding over time and extrinsic imaging conditions in the past. Moreover, a practical Self-Supervised Splicing Network, S3-Net, is proposed to rejoin the OB fragments based on shape similarity of their borderlines. Specifically, we first transform the manually annotated borderline strokes of OB images into times series style shape representations, which are fed as input to a Generative Adversarial Network for augmenting positive pairs of rejoinable OBs for each OB fragment that does not have rejoinable counterparts. A Siamese network is trained on such augmented data in a contrastive learning manner to retrieve the matching OB fragments of an unseen query from an OB fragment gallery. Experiments on the OB-Rejoin benchmark show that our data-driven approach outperforms two recent methods for time-series analysis. In order to demonstrate its practical potential, we deploy the proposed S3-Net method in real tests and ultimately discover dozens of new rejoinings missed by domain experts for decades. Chongsheng Zhang, Bin Wang 0063, Ke Chen 0004, Ruixing Zong, Bofeng Mo, Yi Men, George Almpanidis, Shanxiong Chen, Xiangliang Zhang 0001 |
KDD | 1 |
| 2022 | OBM-CNN: a new double-stream convolutional neural network for shield pattern segmentation in ancient oracle bones
Weize Gao, Shanxiong Chen, Chongsheng Zhang, Bofeng Mo, Xuxing Liu |
Appl. Intell. | 3 |
| 2021 | Street View Text Recognition With Deep Learning for Urban Scene Understanding in Intelligent Transportation SystemsabstractUnderstanding the surrounding scenes is one of the fundamental tasks in intelligent transportation systems (ITS), especially in unpredictable driving scenes or in developing regions/cities without digital maps. Street view is the most common scene during driving. Since streets are often full of shops with signboards, scene text recognition over the shop sign images in street views is of great significance and utility to urban scene understanding in ITS. To advance research in this field, (1) we build ShopSign, which is a large-scale scene text dataset of Chinese shop signs in street views. It contains 25,770 natural scene images, and 267,049 text instances. The images in ShopSign were captured in different scenes, from downtown to developing regions, and across 8 provinces and 20 cities in China, using more than 50 different mobile phones. It is very sparse and imbalanced in nature. (2) we carry out a comprehensive empirical study on the performance of state-of-the-art DL based scene text reading algorithms on ShopSign and three other Chinese scene text datasets, which has not been addressed in the literature before. Through comparative analysis, we demonstrate that language has a critical influence on scene text detection. Moreover, by comparing the accuracy of four scene text recognition algorithms, we show that there is a very large room for further improvements in street view text recognition to fit real-world ITS applications. Chongsheng Zhang, Weiping Ding 0001, Guowen Peng, Feifei Fu, Wei Wang 0012 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | AI-Powered Oracle Bone Inscriptions Recognition and Fragments RejoiningabstractOracle Bone Inscriptions (OBI) research is very meaningful for both history and literature. In this paper, we introduce our contributions in AI-Powered Oracle Bone (OB) fragments rejoining and OBI recognition. (1) We build a real-world dataset OB-Rejoin, and propose an effective OB rejoining algorithm which yields a top-10 accuracy of 98.39%. (2) We design a practical annotation software to facilitate OBI annotation, and build OracleBone-8000, a large-scale dataset with character-level annotations. We adopt deep learning based scene text detection algorithms for OBI localization, which yield an F-score of 89.7%. We propose a novel deep template matching algorithm for OBI recognition which achieves an overall accuracy of 80.9%. Since we have been cooperating closely with OBI domain experts, our effort above helps advance their research. The resources of this work are available at https://github.com/chongshengzhang/OracleBone. Chongsheng Zhang, Ruixing Zong, Shuang Cao, Yi Men, Bofeng Mo |
IJCAI | 1 |
| 2019 | Multi-Imbalance: An open-source software for multi-class imbalance learning
Chongsheng Zhang, Jingjun Bi, Shixin Xu, Enislay Ramentol, Gaojuan Fan, Hamido Fujita |
Knowl. Based Syst. | 1 |
| 2019 | On Incremental Learning for Gradient Boosting Decision Trees
Chongsheng Zhang, Xianjin Shi, George Almpanidis, Gaojuan Fan, Xiajiong Shen |
Neural Process. Lett. | 1 |
| 2018 | An empirical evaluation of high utility itemset mining algorithms
Chongsheng Zhang, George Almpanidis, Wanwan Wang, Changchang Liu |
Expert Syst. Appl. | 1 |
| 2018 | An empirical comparison on state-of-the-art multi-class imbalance learning algorithms and a new diversified ensemble learning scheme
Jingjun Bi, Chongsheng Zhang |
Knowl. Based Syst. | 2 |
| 2017 | Feature selection and resampling in class imbalance learning: Which comes first? An empirical study in the biological domainabstractClass imbalance exists in many applications of bioinformatics and biomedicine, while dimension reduction in the feature space is often needed when building prediction models on a dataset. When the above two issues need to be considered simultaneously for skewed/imbalanced datasets, practitioners and researchers in machine learning may raise the following question: should feature selection be conducted before or after the resampling methods for combating the skewness of a dataset? While feature selection and class imbalance learning have been widely studied in the literature, little study has jointly investigated them. This paper presents a first empirical study on the performance of the two opposing pipelines for binary imbalance learning, i.e., first feature selection then resampling, or first resampling then feature selection. We carry out the study on 35 publicly available datasets belonging to the biological field, using 9 feature selection methods, 6 resampling approaches for class imbalance learning, and 3 well-known classifiers. Our experiments reveal that, there is no constant winner between the two pipelines, practitioners should test both pipelines in order to derive the best classification model for imbalance learning, in particular, the resampling before feature selection pipeline should not be neglected; but we also show that, the feature selection before resampling pipeline outperforms the other in more cases than not. Chongsheng Zhang, Jingjun Bi, Paolo Soda |
BIBM | 1 |
| 2017 | An up-to-date comparison of state-of-the-art classification algorithms
Chongsheng Zhang, Changchang Liu, Xiangliang Zhang 0001, George Almpanidis |
Expert Syst. Appl. | 1 |
| 2016 | A parameter-free label propagation algorithm for person identification in stereo videos
Chongsheng Zhang, Jingjun Bi, Changchang Liu, Ke Chen 0004 |
Neurocomputing | 1 |
| 2014 | Real-Time Biomedical Instance SelectionabstractComputer-based medical systems play a very important role in medical applications because they can strongly support the physicians in the decision making process. The large amount of data nowadays available, although collected from high quality sources, usually contain irrelevant, redundant, or noisy information, suggesting that not all the training instances are useful for the classification task. To address this issue, we present here an instance selection method that, different from the existing approaches, selects in ``real-time" a subset of instances from the original training set on the basis of the information derived from each test instance to be classified. We apply our method to seven public benchmark datasets, achieving larger performances than a baseline classifier. Chongsheng Zhang, Roberto D'Ambrosio, Paolo Soda |
CBMS | 1 |
| 2014 | "Real-time" Instance Selection for Biomedical Data Classification
Chongsheng Zhang, Roberto D'Ambrosio, Paolo Soda |
DaWaK | 1 |
| 2014 | A system for efficient and simultaneous processing of moving K nearest neighbor and spatial keyword queriesabstractWe study the efficient, generic processing of moving K nearest neighbor (MKNN) and top-K spatial keyword (MKSK) queries. Such generic processing is attractive during high query loads. We propose GridVoronoi--an index that enables users to find the spatial nearest neighbor (NN) from uniformly distributed datasets in almost O(1) time. GridVoronoi is based upon Voronoi diagram which has proven to be highly efficient in exploring the local neighborhood of a given Voronoi cell. However, Voronoi diagram needs a method to promptly find out which Voronoi cell contains the query point. So we add a virtual (i.e., conceptual) grid to the Voronoi diagram. For any query point, GridVoronoi first uses the grid to compute which Voronoi cell contains the query, next utilizes Voronoi diagram to quickly find the NN and KNN (i.e., K nearest neighbors) of the query. Chongsheng Zhang |
SSDBM | 1 |
| 2014 | The anti-bouncing data stream model for web usage streams with intralinkings
Chongsheng Zhang, Florent Masseglia, Yves Lechevallier |
Inf. Sci. | 1 |
| 2012 | Discovering Highly Informative Feature Set over High DimensionsabstractFor many textual collections, the number of features is often overly large. These features can be very redundant, it is therefore desirable to have a small, succinct, yet highly informative collection of features that describes the key characteristics of a dataset. Information theory is one such tool for us to obtain this feature collection. With this paper, we mainly contribute to the improvement of efficiency for the process of selecting the most informative feature set over high-dimensional unlabeled data. We propose a heuristic theory for informative feature set selection from high dimensional data. Moreover, we design data structures that enable us to compute the entropies of the candidate feature sets efficiently. We also develop a simple pruning strategy that eliminates the hopeless candidates at each forward selection step. We test our method through experiments on real-world data sets, showing that our proposal is very efficient. Chongsheng Zhang, Florent Masseglia, Xiangliang Zhang 0001 |
ICTAI | 1 |
| 2012 | A Double-Ensemble Approach for Classifying Skewed Data Streams
Chongsheng Zhang, Paolo Soda |
PAKDD (1) | 1 |
| 2012 | Modeling and Clustering Users with Evolving Profiles in Usage StreamsabstractToday, there is an increasing need of data stream mining technology to discover important patterns on the fly. Existing data stream models and algorithms commonly assume that users' records or profiles in data streams will not be updated or revised once they arrive. Nevertheless, in various applications such as Web usage, the records/profiles of the users can evolve along time. This kind of streaming data evolves in two forms, the streaming of tuples or transactions as in the case of traditional data streams, and more importantly, the evolving of user records/profiles inside the streams. Such data streams bring difficulties on modeling and clustering for exploringusers' behaviors. In this paper, we propose three models to summarize this kind of data streams, which are the batch model, the Evolving Objects (EO) model and the Dynamic Data Stream (DDS) model. Through creating, updating and deleting user profiles, these models summarize the behaviors of each user as a profile object. Based upon these models, clustering algorithms are employed to discover interesting user groups from the profile objects. We have evaluated all the proposed models on a large real-world data set, showing that the DDS model summarizes the data streams with evolving tuples more efficiently and effectively, and provides better basis for clustering users than the other two models. Chongsheng Zhang, Florent Masseglia, Xiangliang Zhang 0001 |
TIME | 1 |
| 2010 | Discovering Highly Informative Feature Sets from Data Streams
Chongsheng Zhang, Florent Masseglia |
DEXA (1) | 1 |
| 2010 | ABS: The Anti Bouncing Model for Usage Data StreamsabstractUsage data mining is an important research area with applications in various fields. However, usage data is usually considered streaming, due to its high volumes and rates. Because of these characteristics, we only have access, at any point in time, to a small fraction of the stream. When the data is observed through such a limited window, it is challenging to give a reliable description of the recent usage data. We study the important consequences of these constraints, through the “bounce rate” problem and the clustering of usage data streams. Then, we propose the ABS (Anti-Bouncing Stream) model which combines the advantages of previous models but discards their drawbacks. First, under the same resource constraints as existing models in the literature, ABS can better model the recent data. Second, owing to its simple but effective management approach, the data in ABS is available at any time for analysis. We demonstrate its superiority through a theoretical study and experiments on two real-world data sets. Chongsheng Zhang, Florent Masseglia, Yves Lechevallier |
ICDM | 1 |
| 2008 | Mining Top-n Local Outliers in Constrained Spatial Networks
Chongsheng Zhang, Zhongbo Wu, Bo Qu, Hong Chen 0001 |
ADMA | 1 |