EDBT 2026 Demo / reviewers in the wild / expert
Zhuowei Wang 0001
dblp:77/7826
· DBLP profile ↗
42ranked-venue papers
11as first author
36since 2021 · last 2026
0000-0001-6479-5154ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chameleon: Benchmarking Detection and Backtracking on Commercial-Grade AI-Generated VideosabstractThe proliferation of AI-Generated Content (AIGC), especially deepfake videos, poses a severe threat to social trust by enabling fraud, privacy violations and disinformation. Existing AI-generated video detection (AGVD) benchmarks focus on open-source model generated videos, yet commercial closed-source models produce more realistic, temporally coherent videos that are underexplored in detection research. To fill this gap, we present Chameleon, a commercial-grade dataset with 1,700 AI-generated videos from 600 real-world sources across three key domains (News, Speech, Recommendation), featuring high resolution, rich annotations and 3D consistency metrics for dynamic scene spatial coherence, shifting detection from face-centric forgery to holistic scene forensics. This benchmark assesses models on two core tasks: accurate AI video detection in real-world conditions and forensic backtracking of original sources. Experimental results reveal critical limitations of existing methods in detecting and backtracking high-fidelity, spatiotemporally consistent videos from commercial closed-source models, highlighting current methods’ flawed forensic reasoning and establishing Chameleon as a vital challenge for AIGC security research. The code and data are available at https://github.com/lxixim/Chameleon. Xingming Liao, Meiyu Zeng, Canyu Chen, Nankai Lin, Zhuowei Wang 0001, Aimin Yang 0002 |
ICMR | 5 |
| 2026 | FGR: Frequency Aware and Geometric Structure-Guided Multi-modality Image Registration Framework
Qihao Ye, Zhuowei Wang 0001 |
MMM (2) | 2 |
| 2026 | Uni-DTR: A unified framework for balancing knowledge preservation and adaptation in exemplar-free class-incremental learning
Zhuowei Wang 0001 |
Neurocomputing | 2 |
| 2026 | Bidirectional-Graph Attention Networks Parallel Encoder for Data Imputation and Fault Diagnosis of Industrial RobotsabstractSafe operation is a key concern for industrial robots. However, due to hardware failures and unstable data transmission issues, the multivariate time-series data generated by these axes often contain missing or corrupted signals, which severely hinders downstream tasks such as fault diagnosis. Additionally, the substantial volume of industrial data demands considerable time for training time-series imputation models and subsequent classification models. To address these challenges, this study proposes a multitask approach that serves both the data imputation and fault diagnosis tasks for industrial robots. Specifically, the parameters trained in the imputation model can be transferred to the fault diagnosis model, enhancing its performance and efficiency. A multitask method named Bidirectional-Graph Attention Networks Parallel Encoder (Bi-GATPE) is proposed, which employs a bidirectional graph attention network to capture the spatial dependencies among the various variables of industrial robots. Subsequently, a parallel encoder with Diagonal-Filter Attention is designed to model temporal correlations. This dual approach improves the accuracy and training speed for both the imputation and fault diagnosis tasks. Experimental studies based on real industrial robot datasets demonstrate that by modifying the feature fusion layer of the imputation task and sharing the trained parameters with the fault diagnosis task, the proposed method significantly accelerates the convergence of the fault diagnosis model while also improving diagnostic accuracy. The experiments also indicate that our method shows merits in the imputation and fault diagnosis tasks. The source code of Bi-GATPE is available at:https://github.com/miten073/Bi-GATPE. Zhuowei Wang 0001, Chong Chen 0010, Tao Wang 0014, Zhiwen Yu 0002, Zhuyun Chen 0001 |
IEEE Internet Things J. | 1 |
| 2026 | Preserving overlapped information via parallel one-hop and multi-hop neighbor encoding for knowledge graph entity typing
Hongbin Zhang 0008, Zhenghao Huang, Ruihao Li 0006, Tao Wang 0014, Zhuowei Wang 0001, Lianglun Cheng |
Inf. Process. Manag. | 5 |
| 2026 | RWKV-SKF: A recurrent architecture with state-space and frequency-domain filtering for dissolved oxygen predicting and revealing influencing mechanisms
Peijian Zeng, Xingming Liao, Jianhui Xu, Shuisen Chen, Zhuowei Wang 0001, Xingda Chen 0003 |
Inf. Sci. | 5 |
| 2026 | ES-DETR: Real-time detection transformer with encover and soft-dropout
Yiqing He, Zefeng Zheng, Zhuowei Wang 0001, Lianglun Cheng |
Neural Networks | 3 |
| 2025 | Robust Unsupervised Outlier Detection in Mixed Data Using Hierarchical Reference SetsabstractUnsupervised outlier detection in mixed-attribute data poses significant challenges in healthcare and network security domains where data combine numerical, nominal, and ordinal features. Existing methods struggle with two critical limitations: they typically handle only single-type data and fail to distinguish scattered outliers from clustered outliers-locally dense microclusters that are globally abnormal but mask each other due to internal consistency. This paper proposes HAOD (Heterogeneous Attribute-based Outlier Detection), a hierarchical framework that constructs Natural Neighbor Sets (NNS) for adaptive local structure modeling and organizes them into Natural Neighbor Graph Reference Sets (NGS) for global connectivity representation. A dependency-aware mixed-distance metric unifies heterogeneous attributes by quantifying inter-attribute correlations. HAOD integrates Local Isolation Score (LIS) and Subset Isolation Score (SIS) to comprehensively detect both anomaly types without the masking effect. Experiments on NSL-KDD, UNSW-NB15, and Thyroid Disease datasets show HAOD outperforms eight baseline methods across AUC, precision, and average precision metrics. Ablation studies confirm both components are essential. The method operates parameter-free with$\mathbf{O}\left(\mathbf{n}^{2} \mathbf{d}\right)$complexity, offering a robust solution for heterogeneous monitoring systems. Xiaopeng Luo, Zhuowei Wang 0001 |
BIBM | 2 |
| 2025 | Improving Cognitive Capability of Large Language Model: A Multi-Step Symbolic Reasoning Approach
Jinkun Zhai, Chong Chen 0010, Zhuowei Wang 0001, Tao Wang 0014, Lianglun Cheng |
CogSci | 3 |
| 2025 | DynaCLIP: A Novel Framework for Video-Text Retrieval Via Dynamic Curriculum Learning and Adaptive Prompt Mixture-of-ExpertsabstractUnderstanding and modeling time remains a key challenge in today's video understanding systems. With language playing a central role in driving powerful generalization, foundational video-language models must inherently capture a sense of temporality. However, Video-text retrieval faces three challenges: Parameter efficiency bottlenecks (traditional prompt learning requires adjusting over 10 % of parameters, leading to a sharp increase in computational costs), limitations of static course strategies (manually defined difficulty thresholds cause pseudo-label biases in highly abstract actions), and insufficient adaptability to heterogeneous architectures (cross-modal hybrid expert frameworks rely on shared weights, restricting the flexible combination of heterogeneous models such as CLIP and ViT). To address these challenges, this paper proposes the Dynamic Curriculum Learning with Adaptive Prompt Mixture of Experts (DynaCLIP) framework. Through the collaborative optimization of the Adaptive Prompt Mixture of Experts (APMoE) module and dynamic curriculum learning, the parameter efficiency and dynamic adaptability are significantly improved. APMoE adopts orthogonal initialization expert pools and cross-modal routing networks. It freezes the pre-trained backbone networks (such as CLIP/BERT) on the premise of fine-tuning only less than 1 % of parameters (static expert pool + dynamic prompt generator), and generates instance-aware text encoder prompts by dynamically selecting Top-k experts and weighted fusion. Here, these instanceaware prompts are implemented as continuous soft prompts in the embedding space, rather than discrete natural language text. Meanwhile, a two-dimensional difficulty quantification system is constructed based on the verb abstraction level (VerbNet three-level classification) and BERT semantic similarity. Combined with the improved binary search strategy to balance high-confidence samples and diversity exploration, a cognitiveinspired course learning mechanism is formed. Experiments show that DynaCLIP achieves 7.4 % R@1 on the TEMPO dataset, 67.5 % accuracy in temporal inference. The code is available at: https://github.com/18162195164/DynaCLIP. Xingming Liao, Chengzhong Lin, Zhuowei Wang 0001, Peijian Zeng |
ICDM | 4 |
| 2025 | LSFNet: A Lightweight Spatial-Frequency Integrated Framework for Efficient Motion Deblurring
Chuxiu Guo, Zhuowei Wang 0001, Xingming Liao, Chengzhong Lin |
PRCV (4) | 2 |
| 2025 | Large language model assisted fine-grained knowledge graph construction for robotic fault diagnosis
Xingming Liao, Chong Chen 0010, Zhuowei Wang 0001, Ying Liu 0004, Tao Wang 0014, Lianglun Cheng |
Adv. Eng. Informatics | 3 |
| 2025 | Toward Latency-Efficient Multicast Coflow Scheduling for Reconfigurable Data Center NetworksabstractABSTRACT The emerging optical circuit technology, which can establish circuit connections among switches, has been proposed as a promising paradigm for data center networks. This paper investigates the problem of minimizing the completion time of multicast coflows in optical circuit switches (OCSs)‐based data center networks. The existing works either only focused on multicast coflow scheduling or focused solely on circuit scheduling in OCS‐based networks, which greatly limits their performance. Hence, in this paper, we study how to reduce the completion time of multicast coflows by considering circuit scheduling and coflow scheduling simultaneously. First, the problem of multicast coflow scheduling is formulated and proved to be NP‐hard. Then, a Delay‐efficient Multicast Coflow Scheduling (DMCS) algorithm is proposed by integrating multicast coflow scheduling with circuit scheduling. The proposed DMCS algorithm is proved to have an approximation ratio of at most , where represents the number of OCS. Through extensive simulations, it is shown that the proposed DMCS algorithm can achieve high performance compared to state‐of‐the‐art methods. Fulong Li, Fanlong Zhang, Yuhang Wu 0008, Zhuowei Wang 0001, Quan Chen 0003, Yongchao Tao |
Concurr. Comput. Pract. Exp. | 4 |
| 2025 | Wavelet guided real time detection transformer with sparse attention
Yiqing He, Zefeng Zheng, Zhuowei Wang 0001, Hanwei Wu, Yunyun Zhang, Lianglun Cheng |
Multim. Syst. | 3 |
| 2025 | A Reduced State-Space Generation Method for Concurrent Systems Based on CPN-PR ModelabstractColored Petri nets (CPNs) provide descriptions of the concurrent behaviors for software and hardware. Model checking based on CPNs is an effective method to simulate and verify the concurrent behavior in system design. However, the model-checking method traverses the full state space, which suffers from the state-space explosion problem. A reduced state-space generation method related to the property of concurrent systems is proposed. Specifically, we extend CPNs to define a property-related model (CPN-PR) and give a property-related analysis method whose results can be used to generate the CPN-PR model. A reduced state-space generation method is developed based on enabled binding element filtering rules. The stutter trace equivalence between the state spaces of CPN and CPN-PR has been proven by showing that the reduced state space may not change the model-checking result. A comparison experiment is conducted to demonstrate the effectiveness of our method. Wenjie Zhong, Tao Sun 0002, Jiantao Zhou 0002, Zhuowei Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Enhancement of Convolutional Neural Network for Protein-Protein Interaction Prediction Using Sequence Feature ExtensionabstractComplexes from protein-protein interaction (PPI) are one of the fundamental molecular parts to perform a variety of biological functions, being of great importance in studying protein functions and action mechanisms. In this paper, we summarize the common protein sequence coding methods and propose a novel multiple-channel encoding, especially for convolutional neural networks (CNN). The proposed encoding consists of basic sequence information and additional sequence characteristics, such as amino acid contents and local sequence fragments. This new composite encoding provides specific and combined features from the original sequence data to enhance the feature abstraction capability of the CNN model. Results of encoding testing indicated performance improvement of 8.46% than the original SSC encoding method, and 4.13%-10.88% compared with literature methods. In 5-fold cross-validation experiments of 718306 PPIs involved 16470 proteins, the overall performance of the proposed method can achieve an accuracy of 94.35% and 0. 8871 of MCC. The prediction was validated by carrying out molecular docking and enrichment analysis, suggesting potential possibilities of real PPIs. The proposed method may provide new insights into the identification techniques of PPIs and can help improve the PPI prediction method. . Yang Wang 0169, Yanjiao Zeng, Dongning Liu, Zhuowei Wang 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Hierarchical Multi-Frequency Transform for Sequential RecommendationabstractSequential Recommendation (SR) aims to understand user preferences by analyzing historical interactions with items. Recent approaches have shifted from the time domain to the frequency domain to potentially enhance preference modeling. While fast Fourier transform is a common choice for frequency transform, it may introduce issues like the Gibbs phenomenon, leading to potentially suboptimal model performance. To address this, we introduce discrete cosine transform into sequential recommendation and present a novel multi-frequency transformation sequential recommendation, named HMFTRec, within a hierarchical framework. Specifically, we develop a discrete cosine transform module base on channel attention. A hierarchical spectrum framework that combines Fourier and discrete cosine transforms is introduced to capture finer-grained frequency domain information and mitigate the Gibbs phenomenon to some extent. Furthermore, contrastive learning is employed to potentially enhance the quality of user embeddings learned from the frequency domain. Extensive experiments conducted on four widely recognized benchmark datasets demonstrate that our model significantly outperforms state-of-the-art approaches. Zhenyi Fan, Hongbin Zhang 0008, Guangyu Lin, Lianglun Cheng, Zhuowei Wang 0001, Chong Chen 0010 |
CSCWD | 5 |
| 2024 | Composited-Nested-Learning with Data Augmentation for Nested Named Entity RecognitionabstractNested Named Entity Recognition (NNER) focuses on addressing overlapped entity recognition. Compared to Flat Named Entity Recognition (FNER), annotated resources are scarce in the corpus for NNER. Data augmentation is an effective approach to address the insufficient annotated corpus. However, there is a significant lack of exploration in data augmentation methods for NNER. Due to the presence of nested entities in NNER, existing data augmentation methods cannot be directly applied to NNER tasks. Therefore, in this work, we focus on data augmentation for NNER and resort to more expressive structures, Composited-Nested-Label Classification (CNLC) in which constituents are combined by nested-word and nested-label, to model nested entities. The dataset is augmented using the Composited-Nested-Learning (CNL). In addition, we propose the Confidence Filtering Mechanism (CFM) for a more efficient selection of generated data. Experimental results demonstrate that this approach results in improvements in ACE2004 and ACE2005 and alleviates the impact of sample imbalance. Xingming Liao, Nankai Lin, Lianglun Cheng, Zhuowei Wang 0001, Chong Chen 0010 |
CSCWD | 5 |
| 2024 | Improving Distantly-Supervised Relation Extraction through Label PromptabstractDistantly supervised relation extraction (DSRE) aims to automatically identify relation facts from unstructured text. Most current DSRE works solve the noise problem based on the bag-level, but the denoising ability of these methods decreases when the bag consists of fewer sentences. In this study, we propose a Distantly supervised Relation extraction with Label Prompt (DRLP) framework. We use textual labels (such as label names) as label prompts to alleviate the problem of decreased denoising ability by utilizing the information of entities and relations in label names. During the training process, label prompts are directly connected to the sentences in the bag to provide a more comprehensive bag representation, and label prompts are randomly deleted based on the number of sentences in the bag. Moreover, we design a residual selective attention mechanism that minimizes the influence of spurious features and optimizes the utilization of label information. Our framework is evaluated on NYT-10d and NYT-10m, the results indicate that our method outperforms the state-of-the-art methods. Guangyu Lin, Hongbin Zhang 0008, Zhenyi Fan, Lianglun Cheng, Zhuowei Wang 0001, Chong Chen 0010 |
CSCWD | 5 |
| 2024 | Hybrid Transformer Architecture for Spectral Super-Resolution Reconstruction of Multispectral ImagesabstractSpectral super-resolution technology, which reconstructs 31-band hyper-spectral images from RGB natural scene images within the 400-700nm bands, has seen rapid growth. However, its fixed spectral resolution and spectral coverage limit its application in remote sensing imaging, particularly for aerial images with multi-band information. The lack of corresponding high-spectral image pairs has hindered research progress, leaving the potential spectral information of these remote sensing images untapped. In this study, we explore a hybrid transformer architecture for multispectral images that carry visible light and near-infrared informations to achieve spectral super-resolution. This network integrates both intra-row and intra-column attention mechanisms, along with a cross inter-row and inter-column attention mechanism, to precisely capture and process the spatial and spectral features in spectral images. In the case of two simulated datasets, the experimental results demonstrate favorable outcomes. In classification experiments using multimodal Pavia University datasets, the reconstructed hyper-spectral images exhibit superior performance with higher average accuracy (95.30%), overall accuracy (95.70%), and Kappa coefficient (93.50%). Genping Zhao, Yudan He, Zhuowei Wang 0001, Heng Wu 0002 |
IGARSS | 3 |
| 2024 | Communication-Aware Energy Consumption Model in Heterogeneous Computing SystemsabstractAbstract Large heterogeneous computing systems are composed of conventional central processing units and graphics processing units (GPUs) where communication plays a crucial role for system performance. This paper presents an energy consumption analytical model in terms of communication perception for the communication–computing pipeline characterization of discrete GPUs systems. We propose a dynamically adaptive energy-efficient task assignment approach, which harnesses particle swarm optimization. Static energy optimization is addressed by optimal task partition granularity. The experimental results demonstrate that the communication-based energy optimization algorithms can be more energy-saving than those without communication consideration. For some application benchmarks, the energy consumption can be saved by up to 31%. This implies the potential that the energy-saving optimization methods can be incorporated in system engineering processes. Zhuowei Wang 0001, Hao Wang 0003 |
Comput. J. | 1 |
| 2024 | Towards real-time non-preemptive multicast scheduling in reconfigurable data center networks
Fanlong Zhang, Jianglong Liu, Yuhang Wu 0008, Quan Chen 0003, Zhuowei Wang 0001 |
Peer Peer Netw. Appl. | 6 |
| 2024 | Deep Learning Acceleration Optimization of Stress Boundary Value Problem SolversabstractThe solution to boundary value problems is of great significance in industrial software applications. In this paper, we propose a novel deep learning method for simulating stress field distributions in simply supported beams, aiming to serve as a solver for stress boundary value problems. Our regression network, Stress-EA, utilizes the convolution encoder module and additive attention to accurately estimate the stress in the beam. By comparing the Stress-EA prediction results with the stress values calculated using ABAQUS, we achieve a mean absolute error (MAE) of less than 0.06. This indicates a high level of consistency between the stress values obtained from the two approaches. Moreover, the prediction time of Stress-EA is significantly shorter, taking only 0.0011s, compared to the calculation time of ABAQUS, which is 16.91s. This demonstrates the high accuracy and low computational latency of our model. Furthermore, our model exhibits smaller model parameters, requires less computation, and has a shorter prediction time compared to training results obtained using classic and advanced networks. To accelerate training, we utilize data parallel methods, achieving up to 1.89 speedup on a dual-GPU platform without compromising accuracy. This advancement enhances the computing efficiency for large-scale industrial software applications. Yongsheng Chen, Zhuowei Wang 0001, Lianglun Cheng |
IEEE Trans. Computers | 2 |
| 2024 | Dual-Branch Domain Adaptation Few-Shot Learning for Hyperspectral Image ClassificationabstractCross-domain few-shot learning (FSL) often employs adversarial domain adaptation techniques to address the issue of data distribution discrepancies between the source and target domains. However, forcing the alignment of two distinct domains may lead to distortions in class distribution alignment and result in a decrease in classification performance in hyperspectral image analysis. Moreover, existing cross-domain methods are often applied to satellite/airborne hyperspectral image as both the source and target domain. It is rarely explored whether the same cross-domain methods can be applied for cross applications where the source domain and target domain data could be both satellite/airborne hyperspectral image with lower spatial resolution and unmanned aerial vehicle (UAV) hyperspectral image with higher spatial resolution. To address these issues, this paper proposes a novel domain-adaptive FSL network with dual branches respectively aiming at domain fusion and domain separation. The domain fusion branch uses a conditional adversarial network to align the global distributions of the two domains, while the domain separation branch introduces gate mechanism for discriminative feature learning in each domain to achieve independent category distributions. During the experiment, the proposed method is evaluated by performing cross-transfer learning under the condition that low spatial resolution hyperspectral data and high spatial resolution hyperspectral data are used as source and target data alternately. The experimental results suggest that the proposed method not only mitigates the negative effects of forced alignment in domain fusion but also holds potential for cross-domain transfer learning between low and high spatial resolution hyperspectral images. Zhuowei Wang 0001, Shihui Zhao, Genping Zhao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Knowledge-integrated Multi-modal Movie Turning Point IdentificationabstractThe rapid development of artificial intelligence provides rich technologies and tools for the automated understanding of literary works. As a comprehensive carrier of storylines, movies are natural multimodal data sources that provide sufficient data foundations, and how to fully leverage the benefits of data remains a sustainable research hotspot. In addition, the efficient representation of multi-source data also poses new challenges for information fusion technology. Therefore, we propose a knowledge-enhanced turning points identification (KTPi) method for multimodal scene recognition. First, the BiLSTM method is used to encode scene text and integrate contextual information into scene representations to complete text sequence modeling. Then, the graph structure is used to model all scenes, which strengthens long-range semantic dependencies between scenes and enhances scene representations using graph convolution network. After, the self-supervised method is used to obtain the optimal number of neighboring nodes in sparse graph. Next, actor and verb knowledge involved in the scene text are added to the multimodal data to enhance the diversity of scene feature expressions. Finally, the teacher-student network strategy is used to train the KTPi model. Experimental results show that KTPi outperforms baseline methods in scene role recognition tasks, and ablation experiments show that incorporating knowledge into multimodal model can improve its performance. Depei Wang, Ruifeng Xu 0001, Lianglun Cheng, Zhuowei Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Alzheimer's disease classification using distilled multi-residual network
Xuehu Liang, Zhuowei Wang 0001 |
Appl. Intell. | 2 |
| 2023 | Multi-branch detection network based on trigger attention for pedestrian detection under occlusion
Zhuowei Wang 0001, Weida Lin, Lianglun Cheng, Yang Wang 0169 |
Appl. Intell. | 1 |
| 2023 | A memetic algorithm for high-strength covering array generationabstractAbstract Covering array generation (CAG) is the key research problem in combinatorial testing and is an NP‐complete problem. With the increasing complexity of software under test and the need for higher interaction covering strength t , the techniques for constructing high‐strength covering arrays are expected. This paper presents a hybrid heuristic memetic algorithm named QSSMA for high‐strength CAG problem. The sub‐optimal solution acceptance rate is introduced to generate multiple test cases after each iteration to improve the efficiency of constructing high‐covering strength test suites. The QSSMA method could successfully build high‐strength test suites for some instances where t up to 15 within one day cutoff time and report five new best test suite size records. Extensive experiments demonstrate that QSSMA is a competitive method compared to state‐of‐the‐art methods. Jiantao Zhou 0002, Fei-Yue Wang 0001, Kecheng Tang, Zhuowei Wang 0001 |
IET Softw. | 6 |
| 2023 | Augmenting Feature Representation with Gradient Penalty for Robust Text CategorizationabstractThe capabilities of deep models are constantly mined for extraction and representation of features among text classification tasks. However, these models are sensitive to changes in input data, resulting in poor robustness. Meanwhile, the model lacks information interaction and weak representation ability. In this work, for feature extraction, a joint model that consists of a convolutional neural network, a bidirectional gated recurrent unit, and an attention mechanism is proposed. This new model can improve versatility and fully discover category information in text. For feature representation, a projector under the supervised contrastive learning method is introduced. The method can improve the representation of an encoder and realize aggregation of the same category. Considering the robustness of the PCRA, the gradient penalty is added to a contrastive loss function. Experiments are performed on four datasets to assess the proposed model (PCRA and PCRA‐GP) using an accuracy metric. The experimental results show that our model is suitable for variable‐length and bilingual texts. Compared with the baseline model, it remains competitive, and it reaches SOTA on the 20 Newsgroups dataset. Moreover, the performance of the model is evaluated under different hyperparameters to clarify its working mechanism. Depei Wang, Lianglun Cheng, Zhuowei Wang 0001 |
Int. J. Intell. Syst. | 3 |
| 2023 | Constructing High Radix Quotient Digit Selection Tables for SRT Division and Square RootabstractHigh radix SRT division plays an important role in contemporary microprocessors as the quotient digit selection tables effectively reduce the computation complexity of the quotient digits. The quotient digit selection table is constructed according to the rounded lower and upper bounds of the overlapping regions in the traditional method. The table construction process lacks in mathematical rigor and consequently is susceptible to error. This paper proposes an algebraic method for computing the quotient digit selection tables. We characterize the quotient digit selection functions to construct quotient digit selection tables required for SRT division and SRT square root with any valid redundancy. The functions include the maximum and minimum legal quotient digit selections. We compute the truncations of the remainder$p$and divisor$d$when$d \in [1,2)$. We implement procedures to compute quotient digit selection tables by our functions. The computation of the quotient digit selection table for the case radix-4 with the quotient digit set$[-2,2]$is presented by using the minimum quotient digit selection function. Our functions can compute quotient digit selection tables in the design phase of SRT division and square root by given radix and redundancy. Zhuowei Wang 0001, Yan Wang 0037, Jiantao Zhou 0002 |
IEEE Trans. Computers | 3 |
| 2023 | Warp-Aware Adaptive Energy Efficiency Calibration for Multi-GPU SystemsabstractMassive GPU acceleration processors have been used in high-performance computing systems. The Dennard scaling has led to power and thermal constraints limiting the performance of such systems. The demand for both increased performance and energy efficiency is highly desired. This article presents a multilayer low-power optimization method for warps and tasks parallelisms. We present a dynamic frequency regulation scheme for performance parameters in terms of load balance and load imbalance. The method monitors the energy parameters in runtime and adjusts adaptively the voltage level to ensure performance efficiency with energy reduction. The experimental results show that the multilayer low-power optimization with dynamic frequency regulation can achieve 40% energy consumption reduction with only 1.6% performance degradation, thus reducing 59% maximum energy consumption. It can further save about 30% energy consumption in comparison with the single-layer energy optimization. Zhuowei Wang 0001, Lianglun Cheng, Hai Wan, Wuqing Zhao, Tao Wang 0014 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Distributed Deep Learning Optimization of Heat Equation Inverse Problem SolversabstractThe inversion problem of partial differential equation plays a crucial role in cyber–physical systems applications. This article presents a novel deep learning optimization approach to constructing a solver of heat equation inversion. To improve the computational efficiency in large-scale industrial applications, data and model parallelisms are incorporated on a platform of multiple GPUs. The advanced Ring-AllReduce architecture is harnessed to achieve an acceleration ratio of 3.46. Then, a new multi-GPUs distributed optimization method GradReduce is proposed based on Ring-AllReduce architecture. This method optimizes the original data communication mechanism based on mechanical time and frequency by introducing the gradient transmission scheme solved by linear programming. The experimental results show that the proposed method can achieve an acceleration ratio of 3.84 on a heterogeneous system platform with two CPUs and four GPUs. Zhuowei Wang 0001, Genping Zhao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | An elastic recommender process for cloud service recommendation scalabilityabstractAbstract Cloud computing services are ubiquitous in society and cloud recommender systems play a crucial role in intelligently selecting services for cloud users. Currently, recommendations are static with low scalability. Only one recommendation list is generated at a time and the recommender strategy in the recommendation cycle is not adjustable. This paper presents a new elastic recommender process (ERP) for cloud users. A Markov model is used to characterize the dynamic relationship between different user states. The ERP generates an elastic recommendation that can be used to dynamically adjust the recommender strategy to meet the user's needs based on their browsing records in the current service cycle without the recommender system's involvement. Experimental results show that the ERP improves the effectiveness of the recommender thus increasing the accuracy and diversity of its recommendations. Rui-dong Qi, Jiantao Zhou 0002, Zhuowei Wang 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Phenotypic Parameters Estimation of Plants Using Deep Learning-Based 3-D Reconstruction From Single RGB ImageabstractMonitoring crop growth is of great significance to obtain crop growth status information for development of smart agriculture. The traditional way to measure the phenotypic parameters of crops is labor-intensive and encounters inconvenient operations. In this study, we propose to obtain the phenotypic parameters of crops from 3-D reconstruction of plants from single RGB images using a data-driven plant phenotypic parameters estimation network (P3ES-Net) deep neural network, which enables to estimate the depth shift and camera focal length used for depth estimation and reconstruction of the 3-D model of plants. Based on the principles of the monocular ranging and pinhole imaging model, crop phenotypic parameters such as height, canopy size, and trunk diameter can then be calculated from the 3-D model. Experiments with four practical plants present that our method is able to achieve acceptable evaluation of the growth status of plants. Of more significance, it achieves particular superior depth estimation performance over a commercial depth camera, which is a very new on-sale depth camera using stereo vision and deep learning network. This potential performance throws light on the low-cost measurement of crop phenotypic parameters using RGB camera in monitoring crop growth. Genping Zhao, Weitao Cai, Zhuowei Wang 0001, Heng Wu 0002, Yeping Peng, Lianglun Cheng |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Attention-based BiLSTM fused CNN with gating mechanism model for Chinese long text classification
Lianglun Cheng, Zhuowei Wang 0001 |
Comput. Speech Lang. | 3 |
| 2021 | Activity-Driven Task Allocation in Energy-Constrained Heterogeneous GPUs SystemsabstractAs computing systems continue to increase in complexity, energy optimization plays a key role in the design and implementation of heterogeneous systems. Although the energy consumed by off-chip memory accounts for a large proportion of the total power consumed by the system as a whole, current research on energy optimization mainly focuses on optimizing the energy consumed by the processors. This article explores the coordinated optimization of the holistic performance of the processors and memory system for heterogeneous systems with energy constraints. A communication–computing pipeline model for parallel executions is characterized to optimize program performance by simultaneously scaling the voltage and frequency of the processors and memory using task allocation strategies. A synergistic load-balancing optimization approach is presented to resolve the load imbalance among graphics processing units. Our experimental results substantiate the effectiveness of the approach in terms of execution times and throughputs with the energy constraints. Zhuowei Wang 0001, Lianglun Cheng, Hao Wang 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | DHD-Net: A Novel Deep-Learning-based Dehazing NetworkabstractEliminating haze interference in images is still a challenging problem. In this paper, we consider more systematically the physical hazing mechanisms, combined with deep learning, propose a new end-to-end dehazing network called DHD-Net. For physical hazing mechanisms, we fuse the global atmosphere light, transmission maps, and the atmospheric scattering model for dehazing. For the estimation of global atmosphere light, We propose a deep learning-based haze density estimation algorithm (DL-HDE). We establish a new dataset, of which each data item consists of the hazy image, the transmission map, the haze-free image, and the dense-haze area mask. Our experimental results demonstrate that our proposed DHD-Net has better dehazing performance than state-of-the-art algorithms. Liangru Xie, Hao Wang 0003, Zhuowei Wang 0001, Lianglun Cheng |
IJCNN | 3 |
| 2019 | Energy optimization of parallel programs in a heterogeneous system by combining processor core-shutdown and dynamic voltage scaling
Zhuowei Wang 0001, Hao Wang 0003, Wuqing Zhao, Lianglun Cheng |
Future Gener. Comput. Syst. | 1 |
| 2018 | Three-level performance optimization for heterogeneous systems based on software prefetching under power constraints
Zhuowei Wang 0001, Wuqing Zhao, Hao Wang 0003, Lianglun Cheng |
Future Gener. Comput. Syst. | 1 |
| 2016 | An architecture-level graphics processing unit energy modelabstractSummary With the continued development of hardware and software, graphics processing unit (GPU) has been used in general purpose computational fields, while accelerating applications for CPUs. To achieve high computing performance, a GPU typically includes hundreds of computing units. The high density of computing resource on‐chip incurs high power consumption as well as engendering high performance. The power consumption problem has become one of the most important problems for the development of GPUs. Focusing on a CPU‐GPU heterogeneous parallel system, this research proposed an architecture‐level GPU energy model, with the aim of reducing system energy demand and improving system efficiency. Taking the influences of memory and temperature on GPU energy demand into account, a dynamic energy model based on division of computation and memory and a static energy model based on real‐time temperature perception were established. Validation, through the evaluation and comparative analysis of nine typical GPU programmes, demonstrated that these models could reduce chip energy demand under the performance constraint conditions imposed. Copyright © 2014 John Wiley & Sons, Ltd. Zhuowei Wang 0001, Lianglun Cheng, Wuqing Zhao, Naixue Xiong |
Concurr. Comput. Pract. Exp. | 1 |
| 2011 | GPU Acceleration for GRAPES Meteorological ModelabstractThere is heavy computation involved in Global / Regional Assimilation and Prediction System (GRAPES) which is a typical non-linear discrete system of the Earth's atmosphere developed for numerical weather prediction. Researchers have recently paid a lot of attention to the parallel acceleration of the GRAPES model by low-cost, low-power, high-performance Graphics Processor Units (GPUs). We implement in this paper the acceleration for the GRAPES model in GPUs, which is however not as efficient as supposed. On this basis, we further propose strategies to optimize the performance, including the ease of data transmission time, reducing the amount of device memory loaded and used to store various objects, and avoiding the branches of thread control flows. One can find from the experiments that, parallel acceleration with GPUs helps to improve the performance of GRAPES in terms of measures like accuracy and timeliness. Zhuowei Wang 0001, Xianbin Xu, Naixue Xiong, Laurence T. Yang, Wuqing Zhao |
HPCC | 1 |
| 2008 | A Weight-Based Dynamic Replica Replacement Strategy in Data GridsabstractData grid is a very important and useful technique to process the large number of data produced by scientific experiments and simulations. However, the high latency of the Internet turns to be the bottleneck in accessing the files in the grid. The replication strategy can shorten the time of getting the files by creating many replicas and storing the replicas in appropriate locations. Restricted by the storage capacity, it is important to design an effective algorithm for the replication replacement task. In this paper, we propose a novel replacement strategy which calculates the weight of replicas based on the access times in the future time window, the access bandwidth and the file size. Results from the simulation procedure show the better performance of our algorithm than former ones. Wuqing Zhao, Xianbin Xu, Naixue Xiong, Zhuowei Wang 0001 |
APSCC | 4 |