Zhiming Zhao

dblp:81/2506 · DBLP profile ↗
← Back
11ranked-venue papers in the field
1as first author
11since 2021 · last 2027
0000-0002-6717-9418ORCID · conflict

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 4Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 2 (1 first)Big Data, Cloud & Distributed Data Systems · 2Database Systems & Data Management · 1
YearPublicationVenuePosition
2027 HAT-KI: A Hierarchical Auxiliary Task framework with asymmetric Knowledge Injection for early rumor detection
Wenbo Xing, Yan Li 0104, Zhiming Zhao
Inf. Process. Manag.3
2025 How Good are LLMs at Retrieving Documents in a Specific Domain?
Nafis Tanveer Islam, Zhiming Zhao
FQAS2
2025 Enhancing Adversarial Transferability via Self-Ensemble Feature Alignment
abstract
Deep neural networks (DNNs) have demonstrated remarkable success in tasks such as image classification and object detection, but remain vulnerable to adversarial attacks. To enhance the adversarial transferability across different architectures (e.g., from CNNs to ViTs), existing attacks leverage various strategies such as input transformations, gradient rectification, custom optimization objectives, and model ensembles, but struggle with limited effectiveness under minimal knowledge (e.g., the number of surrogate models). In this work, we propose a novel self-ensemble feature alignment(SEFA) strategy that significantly boosts adversarial transferability with minimal resource overhead. Motivated by the observation that adversarial transferability correlates with feature similarity across models, we leverage Centered Kernel Alignment (CKA) to measure and investigate intermediate features in both inter- and intra-model representation spaces. By splitting a single model into multiple sub-networks and aligning their feature spaces, our method effectively enhances adversarial transferability without relying on additional surrogate models. Experiments on the ImageNet dataset demonstrate that our approach achieves an average ASR of 80.0% (ResNet-50 surrogate) and 75.9% (Inc-v3 surrogate) on ImageNet, which outperform the second-best method (i.e., BSR) by +8.3% and +4.7% respectively. Further, it can seamlessly integrate with existing attacks to further increase cross architecture transferability.
Zhiming Zhao, Qingming Li, Chunyi Zhou 0001, Shouling Ji
ICMR1
2023 Ocean Data Quality Assessment through Outlier Detection-enhanced Active Learning
abstract
Ocean and climate research benefits from global ocean observation initiatives such as Argo, GLOSS, and EMSO. The Argo network, dedicated to ocean profiling, generates a vast volume of observatory data. However, data quality issues from sensor malfunctions and transmission errors necessitate stringent quality assessment. Existing methods, including machine learning, fall short due to limited labeled data and imbalanced datasets. To address these challenges, we propose an Outlier Detection-Enhanced Active Learning (ODEAL) framework for ocean data quality assessment, employing Active Learning (AL) to reduce human experts’ workload in the quality assessment workflow and leveraging outlier detection algorithms for effective model initialization. We also conduct extensive experiments on five large-scale realistic Argo datasets to gain insights into our proposed method, including the effectiveness of AL query strategies and the initial set construction approach. The results suggest that our framework enhances quality assessment efficiency by up to 465.5% with the uncertainty-based query strategy compared to random sampling and minimizes overall annotation costs by up to 76.9% using the initial set built with outlier detectors.
Yiyang Qi, Ruyue Xin, Zhiming Zhao
IEEE Big Data4
2022 MOGPlay: A Decentralized Crowd Journalism Application for Democratic News Production
abstract
Media production and consumption behaviors are changing in response to new technologies and demands, giving birth to a new generation of social applications. Among them, crowd journalism represents a novel way of constructing democratic and trustworthy news relying on ordinary citizens arriving at breaking news locations and capturing relevant videos using their smartphones. The ARTICONF project [1] proposes a trustworthy, resilient, and globally sustainable toolset for developing decentralized applications (DApps). Leveraging the ARTICONF tools, we introduce a new DApp for crowd journalism called MOGPlay. MOGPlay collects and manages audio-visual content generated by citizens and provides a secure blockchain platform that rewards all stakeholders involved in professional news production. Besides live streaming, MOGPlay offers a marketplace for audio-visual content trading among citizens and free journalists with an internal token ecosystem. We discuss the functionality and implementation of the MOGPlay DApp and illustrate three pilot crowd journalism live scenarios that validate the prototype.
Inês Rito Lima, Cláudia Marinho, Vasco Filipe, Alexandre Ulisses, Nishant Saurabh, Antorweep Chakravorty, Zhiming Zhao, Atanas Hristov, Radu Prodan
ASONAM7
2022 An Adaptable Indexing Pipeline for Enriching Meta Information of Datasets from Heterogeneous Repositories
Siamak Farshidi, Zhiming Zhao
PAKDD (2)2
2022 Label entropy-based cooperative particle swarm optimization algorithm for dynamic overlapping community detection in complex networks
abstract
The real-world complex networks, such as biological, transportation, biomedical, web, and social networks, are usually dynamic and change over time. The communities which reflect the substructures hidden in the networks usually overlap each other, and detecting overlapping communities in the dynamic complex networks is a challenging task. Prior researchers have applied multiobjective optimization method to the detection of dynamic overlapping communities and achieved some excellent results. However, in terms of multiobjective processing, the prior studies all adopt the decomposition method based on weight parameters, and different weight parameters or different parameter values can easily affect the community detection results which further results in the uneven distribution of the detected results in the target space. To solve the above problems, a hybrid algorithm, that is, Collaborative Particle Swarm multiobjective Optimization-based Dynamic Overlapping Community Detection (CPSO-DOCD) algorithm is proposed in this paper. First, to improve the diversity of particles, the encoding/decoding of the particle and the cross inheritance and the variation of particle are redefined first based on label propagation. In each network snapshot, multiple particle swarms are initialized based on Community Overlap Propagation Algorithm (COPRA) to generate particles with uniform distribution. Multiple different objective functions are optimized using multiple particle swarms respectively to avoid the incorrect selection of weight parameters. In addition, a reference-point-based is adopted in the particle selecting stage to solve the uneven distribution of detected results in the target space. Second, a node label entropy-based particle swarm algorithm is proposed to improve the accuracy of community detection of current network snapshots. Finally, when one snapshot switches to another over time, a migration strategy based on COPRA local-search and clique generation is utilized to adjust the prior community detection results, which enables the former results can be adapted to the new network snapshots. The experiments are implemented based on four dynamic networks which are Cit-HepPh, Cit-HepTh, Emailed-EU-core-temporal, and CollegeMsg. The hypervolume value of the overlapping community detection result obtained by CPSO-DOCD is 0.5%–2% higher than MDOA, MCMOEA, SLPAD, and iLCD. Furthermore, CPSO-DOCD also performed better than MDOA, MCMOEA, SLPAD, and iLCD on C-metric values, and CPSO-DOCD can approach approximately to the Pareto frontier.
Wenchao Jiang, Shucan Pan, Chaohai Lu, Zhiming Zhao, Sui Lin, Meng Xiong, Zhongtang He
Int. J. Intell. Syst.4
2022 Metric learning-based whole health indicator model for industrial robots
abstract
Aiming at the problems of complex structure, high components coupling, and difficultly monitoring of the whole health status with the industrial robot, a metric learning-based whole health indicator model is proposed. First, according to the more obvious degradation characteristics of industrial robots during accelerated operation, the accelerated signal is segmented and then the time-domain features are extracted. Second, the long-term and short-term memory (LSTM) network combined with the multihead attention is used to construct the network model, and the metric learning method is adopted to learn the similarity measurement method of the industrial robot monitoring data. Finally, the similarity measure method got from metric learning is used to construct the whole health indicator, which describes the whole degradation trend of the industrial robot. The experiments are based on the real accelerated aging data set from industrial robots. The results show that the proposed model can effectively construct the whole health indicator for industrial robots. The average trend of the proposed model reaches 0.9769. The average monotonicity reaches 0.5666, which is 0.1748, 0.1577, and 0.1492 higher than the similarity measurement method based on Euclidean distance, Markov distance, and LSTM.
Ping Li 0045, Hanlin Zeng, Tiancai Liang, Wenchao Jiang, Zhiming Zhao
Int. J. Intell. Syst.6
2022 Real-time recognition and warning of mask wearing based on improved YOLOv5 R6.1
abstract
Since the new crown epidemic, mask-wearing has become a new normal in people's work and life. The inspection mechanism for mask-wearing at the entrance and exit of public places is seriously insufficient. The phenomenon of “pick-up on entry” has led to the severe formalization of mask-wearing inspection. Manual detection of mask-wearing in an open and dynamic crowded environment is unrealistic, which is not only time-consuming and labor-intensive but also cannot achieve early warning throughout the entire process. In response to this problem, this paper proposes a real-time recognition and early warning method for mask-wearing in an open, dynamic, complex environment based on improved YOLOv5 R6.1. First, replacing the first Conv structure of the backbone network in the YOLOv5 R6.1 model with an improved Stem structure to minimize the computational overhead while improving the performance. Then by normalizing the data, the random erasure data expansion technique is used to enhance the antiocclusion robustness of the algorithm. Finally, according to the mask-wearing specification in the training data set, optimizing and adjusting the anchor box parameters of the YOLOv5 R6.1 model to improve the model's ability to recognize small targets. The experiments are based on open data sets, and the results show that the mean precision (mAP), precision, and recall of this method reach 92.9%, 94.1%, and 88.5% on average, and the average frames per second (FPS) reaches 117. Moreover, the mAP and FPS are improved by an average of 6.5% and 474% compared with algorithms based on RetinaNet, Attention-Retina, Single Shot multibox Detector, Fast-RCNN, YOLOv4, and YOLOv5.
Shenghai Yuan 0001, Tiancai Liang, Wenchao Jiang, Sui Lin, Zhiming Zhao
Int. J. Intell. Syst.6
2022 Short-text feature expansion and classification based on nonnegative matrix factorization
abstract
In this paper, a non-negative matrix factorization feature expansion (NMFFE) approach was proposed to overcome the feature-sparsity issue when expanding features of short-text. First, we took the internal relationships of short texts and words into account when segmenting words from texts and constructing their relationship matrix. Second, we utilized the Dual regularization non-negative matrix tri-factorization (DNMTF) algorithm to obtain the words clustering indicator matrix, which was used to get the feature space by dimensionality reduction methods. Thirdly, words with close relationship were selected out from the feature space and added into the short-text to solve the sparsity issue. The experimental results showed that the accuracy of short text classification of our NMFFE algorithm increased 25.77%, 10.89%, and 1.79% on three data sets: Web snippets, Twitter sports, and AGnews, respectively compared with the Word2Vec algorithm and Char-CNN algorithm. It indicated that the NMFFE algorithm was better than the BOW algorithm and the Char-CNN algorithm in terms of classification accuracy and algorithm robustness.
Wenchao Jiang, Zhiming Zhao
Int. J. Intell. Syst.3
2021 Unsupervised Anomaly Detection in Data Quality Control
abstract
Data is one of the most valuable assets of an organization and has a tremendous impact on its long-term success and decision-making processes. Typically, organizational data error and outlier detection processes perform manually and reactively, making them time-consuming and prone to human errors. Additionally, rich data types, unlabeled data, and increased volume have made such data more complex. Accordingly, an automated anomaly detection approach is required to improve data management and quality control processes. This study introduces an unsupervised anomaly detection approach based on models comparison, consensus learning, and a combination of rules of thumb with iterative hyper-parameter tuning to increase data quality. Furthermore, a domain expert is considered a human in the loop to evaluate and check the data quality and to judge the output of the unsupervised model. An experiment has been conducted to assess the proposed approach in the context of a case study. The experiment results confirm that the proposed approach can improve the quality of organizational data and facilitate anomaly detection processes.
Lex Poon, Siamak Farshidi, Zhiming Zhao
IEEE BigData4