EDBT 2026 Demo / reviewers in the wild / expert
Guoqiang Zhong 0001
dblp:15/3875
· DBLP profile ↗
74ranked-venue papers
20as first author
30since 2021 · last 2026
0000-0002-2952-6642ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 18 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressively attentional architecture search
Xuanmian Liu, Xianping Qin, Fuchang Zhang, Rachid Hedjam, Guoqiang Zhong 0001 |
Neurocomputing | 6 |
| 2026 | A quantum neural network with built-in self-attention mechanism
Shangshang Shi, Ruimin Shang, Haiyong Zheng, Guoqiang Zhong 0001, Yongjian Gu |
Neurocomputing | 7 |
| 2026 | Progressively unfreezing perceptual GAN
Rachid Hedjam, Jinxuan Sun, Yang Chen 0036, Junyu Dong, Guoqiang Zhong 0001, Wei Xiang 0001 |
Multim. Syst. | 7 |
| 2026 | Beyond Semantics: Multiscale Interaction Network for Referring Camouflaged Object DetectionabstractReferring camouflaged object detection (Ref-COD) is an emerging and challenging task that aims to localize camouflaged objects in complex scenes based on a small set of referring images with salient objects. However, existing methods primarily focus on semantic alignment between the referring and camouflaged objects while overlooking scale discrepancies, leading to under-response when small references guide large objects and over-response when large references guide small ones. To overcome this limitation, we propose a novel Multi-scale Interaction Network (MINet), explicitly designed to handle feature interactions across different scales in Ref-COD. MINet begins with a Dual-Source Fusion Block (DSFB) for semantic fusion between the referring and camouflaged features. Then, the Intra-scale Interaction Block (IIB) enhances local saliency within each scale by modeling contextual importance. Next, the Cross-scale Interaction Block (CIB) performs offset-guided alignment to bridge spatial gaps in multiscale feature fusion. Finally, the Cross-scale Aggregation Decoder (CAD) integrates multiscale features, effectively decoding the aggregated information to produce accurate predictions. Extensive experiments on Ref-COD datasets demonstrate that our method achieves state-of-the-art performance, highlighting the importance of scale interaction in Ref-COD. Xiandong Wang, Tianqi Guo, Fengqin Yao, Shengke Wang, Junyu Dong, Guoqiang Zhong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | GBNet: Gated Boundary-Aware Network for Camouflaged Object DetectionabstractCamouflaged object detection involves identifying camouflaged objects visually blended into the surroundings, holding crucial significance in various visual applications. Existing methods primarily focus on leveraging boundary information to enhance camouflaged object detection. However, they often overlook the background interference near the object boundaries, which leads to coarse boundary predictions and results in suboptimal detection performance. In this paper, to address this problem, we propose GBNet, a gated boundary-aware network designed to enhance boundary precision and improve overall detection performance. Specifically, GBNet incorporates a boundary-enhanced module that selectively filters extraneous background information through a boundary gate block, ensuring the generation of high-quality boundary information. Additionally, a boundary-aware decoder is designed to enrich the representation ability of the decoder by injecting high-quality boundary features and aggregating contextual features. With meticulous design, GBNet excels in accurately segmenting camouflaged objects in challenging scenarios. Extensive experiments demonstrate that GBNet outperforms 19 state-of-the-art methods significantly across four widely-used benchmark datasets. The source code is publicly available at https://github.com/wooownn/GBNet. Xiandong Wang, Fengqin Yao, Guoqiang Zhong 0001, Shengke Wang, James T. Kwok |
IEEE Trans. Image Process. | 3 |
| 2026 | A Survey of Deep Learning for Time Series Forecasting: Taxonomy, Analysis and Future DirectionsabstractAs a critical branch of time series analysis, time series forecasting (TSF) focuses on predicting future trends based on historical data, and it plays a pivotal role in a wide range of applications, including meteorology, finance, and healthcare. Recently, deep learning has demonstrated significant potential in TSF. While several existing surveys have systematically summarized deep learning-based methods, we complement these foundational works by investigating emerging models, such as large language models (LLMs), and providing an in-depth comparative analysis of distinct models alongside the challenges currently facing the field. Specifically, we propose a hierarchical taxonomy based on model structural dependency, categorizing existing studies into model-specific and model-agnostic frameworks. The model-specific framework is further divided into discriminative and generative paradigms, accompanied by a detailed comparison of these distinct model types. Moreover, we systematically review prevalent time-series datasets across diverse domains, analyze their key statistics, and summarize evaluation metrics. Finally, we analyze the key challenges currently faced by TSF and explore potential future research directions. Through this systematic review and forward-looking analysis, we aim to provide novel perspectives and establish a clear classification framework of TSF methods. which compiles related papers and open-source code in TSF. This repository will be continuously updated to include the latest research advancements. Baolin Zhao, Mingchen Song, Guoqiang Zhong 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | SADN: Saliency-Based Attention Discrimination Network for Deepfake Detection
Dezhi Lu, Jiajia Dong, Rachid Hedjam, Guoqiang Zhong 0001 |
PRCV (3) | 5 |
| 2025 | Generative neural architecture search
Xiaotong Zhai, Guoqiang Zhong 0001, Tao Li 0031, Fuchang Zhang, Rachid Hedjam |
Neurocomputing | 3 |
| 2025 | Quantum Gated Recurrent Neural NetworksabstractThe exploration of quantum advantages with Quantum Neural Networks (QNNs) is an exciting endeavor. Recurrent neural networks, the widely used framework in deep learning, suffer from the gradient vanishing and exploding problem, which limits their ability to learn long-term dependencies. To address this challenge, in this work, we develop the sequential model of Quantum Gated Recurrent Neural Networks (QGRNNs). This model naturally integrates the gating mechanism into the framework of the variational ansatz circuit of QNNs, enabling efficient execution on near-term quantum devices. We present rigorous proof that QGRNNs can preserve the gradient norm of long-term interactions throughout the recurrent network, enabling efficient learning of long-term dependencies. Meanwhile, the architectural features of QGRNNs can effectively mitigate the barren plateau phenomenon. The effectiveness of QGRNNs in sequential learning is convincingly demonstrated through various typical tasks, including solving the adding problem, learning gene regulatory networks, and predicting stock prices. The hardware-efficient architecture and superior performance of our QGRNNs indicate their promising potential for finding quantum advantageous applications in the near term. Ruipeng Xing, Changheng Shao, Shangshang Shi, Guoqiang Zhong 0001, Yongjian Gu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Self-Prompt Mechanism for Few-Shot Image RecognitionabstractFew-shot learning poses a formidable challenge as it necessitates effective recognition of novel classes based on a limited set of examples. Recent studies have sought to address the challenge of rare samples by tuning visual features through the utilization of external text prompts. However, the performance of these methods is constrained due to the inherent modality gap between the prompt text and image features. Instead of naively utilizing the external semantic information generated from text to guide the training of the image encoder, we propose a novel self-prompt mechanism (SPM) to adaptively adjust the neural network according to unseen data. Specifically, SPM involves a systematic selection of intrinsic semantic features generated by the image encoder across spatial and channel dimensions, thereby engendering self-prompt information. Subsequently, upon backpropagation of this self-prompt information to the deeper layers of the neural network, it effectively steers the network toward the learning and adaptation of new samples. Meanwhile, we propose a novel parameter-efficient tuning method that exclusively fine-tunes the parameters relevant to self-prompt (prompts are no more than 2% of the total parameters), and the incorporation of additional learnable parameters as self-prompt ensures the retention of prior knowledge through frozen encoder weights. Therefore, our method is highly suited for few-shot recognition tasks that require both information retention and adaptive adjustment of network parameters with limited labeling data constraints. Extensive experiments demonstrate the effectiveness of the proposed SPM in both 5-way 1-shot and 5-way 5-shot settings for standard single-domain and cross-domain few-shot recognition datasets, respectively. Our code is available at https://github.com/codeshop715/SPM. Mingchen Song, Guoqiang Zhong 0001 |
AAAI | 3 |
| 2024 | TextGT: A Double-View Graph Transformer on Text for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is aimed at predicting the sentiment polarities of the aspects included in a sentence instead of the whole sentence itself, and is a fine-grained learning task compared to the conventional text classification. In recent years, on account of the ability to model the connectivity relationships between the words in one sentence, graph neural networks have been more and more popular to handle the natural language processing tasks, and meanwhile many works emerge for the ABSA task. However, most of the works utilizing graph convolution easily incur the over-smoothing problem, while graph Transformer for ABSA has not been explored yet. In addition, although some previous works are dedicated to using both GNN and Transformer to handle text, the methods of tightly combining graph view and sequence view of text is open to research. To address the above issues, we propose a double-view graph Transformer on text (TextGT) for ABSA. In TextGT, the procedure in graph view of text is handled by GNN layers, while Transformer layers deal with the sequence view, and these two processes are tightly coupled, alleviating the over-smoothing problem. Moreover, we propose an algorithm for implementing a kind of densely message passing graph convolution called TextGINConv, to employ edge features in graphs. Extensive experiments demonstrate the effectiveness of our TextGT over the state-of-the-art approaches, and validate the TextGINConv module. The source code is available at https://github.com/shuoyinn/TextGT. Guoqiang Zhong 0001 |
AAAI | 2 |
| 2024 | MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical ReasoningabstractThe tool-use Large Language Models (LLMs) that integrate with external Python interpreters have significantly enhanced mathematical reasoning capabilities for open-source LLMs, while tool-free methods chose another track: augmenting math reasoning data.However, a great method to integrate the above two research paths and combine their advantages remains to be explored.In this work, we firstly include new math questions via multi-perspective data augmenting methods and then synthesize code-nested solutions to them.The open LLMs (e.g., Llama-2) are finetuned on the augmented dataset to get the resulting models, MuMath-Code (µ-Math-Code).During the inference phase, our MuMath-Code generates code and interacts with the external python interpreter to get the execution results.Therefore, MuMath-Code leverages the advantages of both the external tool and data augmentation.To fully leverage the advantages of our augmented data, we propose a two-stage training strategy: In Stage-1, we finetune Llama-2 on pure CoT data to get an intermediate model, which then is trained on the code-nested data in Stage-2 to get the resulting MuMath-Code.Our MuMath-Code-7B achieves 83.8% on GSM8K and 52.4% on MATH, while MuMath-Code-70B model achieves new state-of-the-art performance among open methods-achieving 90.7% on GSM8K and 55.1% on MATH.Extensive experiments validate the combination of tool use and data augmentation, as well as our two-stage training strategy.We release the proposed dataset along with the associated code for public use: https://github.com/ youweihao-tal/MuMath-Code. Weihao You, Zhilong Ji, Guoqiang Zhong 0001, Jinfeng Bai |
EMNLP | 4 |
| 2024 | MF-Net: Multi-frequency intrusion detection network for Internet traffic data
Zhaoxu Ding, Guoqiang Zhong 0001, Xianping Qin, Qingyang Li 0007, Zhenlin Fan, Zhaoyang Deng, Wei Xiang 0001 |
Pattern Recognit. | 2 |
| 2024 | Matching Multi-Scale Feature Sets in Vision Transformer for Few-Shot ClassificationabstractRecently, Transformer-based few-shot classification methods are widely exploited. However, they only leverage feature information at a single scale, resulting in weak feature representations, which cannot fully capture the rich information contained in a limited number of images regarding diverse objects with different scales, even those belonging to the same category. To mitigate this issue, we propose a multi-scale feature sets matching scheme in vision Transformer for few-shot classification, and name it FSViT, which can sufficiently extract discriminative features from the few number of labeled support examples. Concretely, we establish a patch-based multi-scale feature representation based on the feature extractors of FSViT, where we introduce an attention-aware grid pooling operation to merge adjacent patches with various scales to obtain multi-scale feature sets. Moreover, we devise a multi-scale patch matching metric to aggregate the measurement of similarity over the multi-scale feature sets for few-shot classification. Extensive experiments demonstrate the effectiveness of the proposed FSViT in both 1-shot and 5-shot scenarios on standard single-domain and cross-domain few-shot classification, especially improving the state-of-the-art recognition accuracy by 1.27% and 1.33% on average on the Mini-ImageNet and CFAIR-FS datasets, respectively. The code of FSViT is available athttps://github.com/codeshop715/FSViT. Mingchen Song, Fengqin Yao, Guoqiang Zhong 0001, Zhong Ji, Xiaowei Zhang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | DewaterGAN: A Physics-Guided Unsupervised Image Water Removal for UAV in Coastal ZoneabstractAccurate recognition of marine species in drone-captured images is essential in maintaining the stability of coastal zone ecosystems. Unmanned aerial vehicle (UAV) remote sensing images usually lack paired supervised signals and suffer from color distortion and blurring due to the interaction of ambient light with cross-medium transmission between air and water. However, current algorithms mainly focus on supervised training methods and also ignore the interaction involved in the cross-medium transmission of light in water. In this article, for UAV in coastal zones, we propose an unsupervised image water removal model, named DewaterGAN, which is based solely on low-tide and high-tide images without paired supervised signals and also preserves color and texture in the water removal process. Specifically, our approach involves two key steps: an unsupervised training CycleGAN network accomplishes domain transitions from low-tide level to high-tide level, and a physics-based attention module guides image water removal and maintains authenticity. Additionally, we utilize evaluation metrics of image restoration peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) to quantitatively analyze the performance of the model. We also employed several non-reference metrics (UIQM, UCIQE, NIQE, BRISQUE, LIQE, ILNIQE, and CLIPIQA) to evaluate the visual quality of the image de-watering process. Extensive experiments conducted on both our water removal dataset and public datasets validate the efficacy of our model. The code is athttps://github.com/yfq-yy/Dewater.git. Fengqin Yao, Fuzhi Tang, Xiandong Wang, Shengke Wang, Guoqiang Zhong 0001, Jingfeng Zhang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | LGI-GT: Graph Transformers with Local and Global Operators InterleavingabstractSince Transformers can alleviate some critical and fundamental problems of graph neural networks (GNNs), such as over-smoothing, over-squashing and limited expressiveness, they have been successfully applied to graph representation learning and achieved impressive results. However, although there are many works dedicated to make graph Transformers (GTs) aware of the structure and edge information by specifically tailored attention forms or graph-related positional and structural encodings, few works address the problem of how to construct high-performing GTs with modules of GNNs and Transformers. In this paper, we propose a novel graph Transformer with local and global operators interleaving (LGI-GT), in which we further design a new method propagating embeddings of the [CLS] token for global information representation. Additionally, we propose an effective message passing module called edge enhanced local attention (EELA), which makes LGI-GT a full-attention GT. Extensive experiments demonstrate that LGI-GT performs consistently better than previous state-of-the-art GNNs and GTs, while ablation studies show the effectiveness of the proposed LGI scheme and EELA. The source code of LGI-GT is available at https://github.com/shuoyinn/LGI-GT. Guoqiang Zhong 0001 |
IJCAI | 2 |
| 2023 | Self-supervised generative learning for sequential data prediction
Guoqiang Zhong 0001, Zhaoyang Deng, Kang Zhang 0007, Kaizhu Huang |
Appl. Intell. | 2 |
| 2023 | Recurrent attention unit: A new gated recurrent unit for long-term memory of important parts in sequential data
Zhaoyang Niu, Guoqiang Zhong 0001, Guohua Yue, Li-Na Wang, Hui Yu 0001, Junyu Dong |
Neurocomputing | 2 |
| 2023 | Lightweight network learning with Zero-Shot Neural Architecture Search for UAV images
Fengqin Yao, Shengke Wang, Laihui Ding, Guoqiang Zhong 0001, Leon Bevan Bullock, Junyu Dong |
Knowl. Based Syst. | 4 |
| 2023 | Ocean Front Detection With Bi-Directional Progressive Fusion Attention NetworkabstractOcean fronts are a mesoscale phenomenon in the ocean. It is important for fisheries, environmental protection, and military activities. Therefore, more and more attention has been attracted to ocean front detection. However, the distribution of front and non-front pixels is highly unbalanced in remote sensing images, and it is not easy to establish an effective ocean front detection algorithm with high accuracy. To alleviate these problems, we model the problem of detecting ocean fronts as an edge detection task and design a new end-to-end bi-directional progressive fusion attention network (BPFANet). Specifically, BPFANet consists of an effective backbone and a bi-directional path. The whole backbone has four stage detection blocks (SD blocks), which capture the ocean front features at different scales. Each SD block contains a side branch structure, which includes a deep residual dilated convolution (DRDC) module to enrich multi-scale edge information and an attention module (AM) to enhance the feature representation in both the channel and spatial dimensions. In addition, the bi-directional path can progressively fuse the four SD blocks of ocean front information. To evaluate BPFANet, we perform experiments on the OFDS365 dataset and show its advantages over existing ocean front detection methods. Qingyang Li 0007, Cui Xie, Guoqiang Zhong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Quantum recurrent neural networks for sequential learning
Rongbing Han, Shangshang Shi, Ruimin Shang, Haiyong Zheng, Guoqiang Zhong 0001, Yongjian Gu |
Neural Networks | 8 |
| 2022 | Deep Discriminative Hashing Networks on Point-wise Labels for Image RetrievalabstractWith the continuous development of digital technology and the Internet, there are extremely large amounts of images generated every day. Hashing algorithms are gaining popularity for image retrieval tasks due to their high efficiency. However, most existing deep hashing methods adopt pairwise labels or triplet labels as their supervisory information, which is unfavorable when the data size is very large, because of the combinational property. We address this problem in this paper and propose a novel point-wise labels based hashing model named Deep Discriminative Hashing Networks (DDHNs). Moreover, we have designed a Hashing Discriminative Loss (HD-Loss) to improve the separability of the learned data representations. We validate that DDHNs outperform traditional hashing algorithms and existing deep hashing algorithms on multiple datasets. In particular, to testify the effect of HD-Loss, we have conducted ablation study and demonstrated that DDHNs perform better than DDHNs without HD-Loss on the used datasets. Zhenchao Chen, Guoqiang Zhong 0001, Jianzhang Qu, Haizhen Wang |
ICPR | 2 |
| 2022 | SteelyGAN: Semantic Unsupervised Symbolic Music Genre Transfer
Zhaoxu Ding, Xiang Liu 0012, Guoqiang Zhong 0001 |
PRCV (1) | 3 |
| 2022 | Weak Edge Identification Network for Ocean Front DetectionabstractOcean fronts have an important influence on global ocean–atmosphere interactions and marine fishery. Hence, it is of great significance to obtain the positions of the ocean fronts. However, current ocean front detection research confronts two challenges: scarcity of labeled data and limitations of ocean front detection algorithms. To address these two problems, we have collected and labeled an ocean front data set and proposed a new deep learning model for ocean front detection. For concreteness, due to the weak edge property of the ocean fronts, we formulate ocean front detection as a weak edge identification problem and propose the weak edge identification network (WEIN) for ocean front detection. WEIN consists of four convolutional blocks. Each block has a side output layer used to detect front edges at a specific image representation level. The side outputs are then fused to predict (detect) the locations of the ocean fronts. In this work, we adopt two metrics to measure the experimental results, i.e., the$F_{1}$-score and intersection over union (IoU). The experimental results with comparison to traditional and deep learning approaches demonstrate the superiority of WEIN for ocean front detection. Qingyang Li 0007, Guoqiang Zhong 0001, Cui Xie, Rachid Hedjam |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Face image-sketch synthesis via generative adversarial fusion
Jianyuan Sun, Hongchuan Yu, Jian J. Zhang 0001, Junyu Dong, Hui Yu 0001, Guoqiang Zhong 0001 |
Neural Networks | 6 |
| 2022 | Adaptive Gabor convolutional networks
Li-Na Wang, Guoqiang Zhong 0001, Wencong Jiao, Junyu Dong, Biao Shen, Dongdong Xia, Wei Xiang 0001 |
Pattern Recognit. | 3 |
| 2022 | Random Shapley Forests: Cooperative Game-Based Random Forests With ConsistencyabstractThe original random forests (RFs) algorithm has been widely used and has achieved excellent performance for the classification and regression tasks. However, the research on the theory of RFs lags far behind its applications. In this article, to narrow the gap between the applications and the theory of RFs, we propose a new RFs algorithm, called random Shapley forests (RSFs), based on the Shapley value. The Shapley value is one of the well-known solutions in the cooperative game, which can fairly assess the power of each player in a game. In the construction of RSFs, RSFs use the Shapley value to evaluate the importance of each feature at each tree node by computing the dependency among the possible feature coalitions. In particular, inspired by the existing consistency theory, we have proved the consistency of the proposed RFs algorithm. Moreover, to verify the effectiveness of the proposed algorithm, experiments on eight UCI benchmark datasets and four real-world datasets have been conducted. The results show that RSFs perform better than or at least comparable with the existing consistent RFs, the original RFs, and a classic classifier, support vector machines. Jianyuan Sun, Hui Yu 0001, Guoqiang Zhong 0001, Junyu Dong, Shu Zhang 0002, Hongchuan Yu |
IEEE Trans. Cybern. | 3 |
| 2021 | Differentiable Light-Weight Architecture SearchabstractIn this paper, we propose a differentiable light-weight architecture search method, called DLWAS. It can search for CNNs with fewer parameters and floating point operations (FLOP-s) while maintaining the SOTA performance. Concretely, we first build a new light-weight search space, which contains the latest and effective light-weight operations, limiting the parameters and FLOPs from the source of neural architecture search (NAS). Secondly, we propose an effective neural architecture optimization method, which results in a more sparse and robust topology for differentiable NAS. Experimental results show that DLWAS achieves an error rate of 2.76% on CIFAR10 comparable to the state-of-the-art methods, using only 2.4M params and 336M FLOPs. The parameters and FLOPs are reduced by 30% and 36% compared with the closest counterpart DARTS, respectively. On ImageNet, our model achieves 3.2% better top-1 accuracy than the SOTA MobileNet, while using fewer parameters and FLOPs. Yuxu Mao, Guoqiang Zhong 0001, Zhaoyang Deng |
ICME | 2 |
| 2021 | Deep Architecture Compression with Automatic Clustering of Similar Neurons
Xiang Liu 0012, Wenxue Liu, Li-Na Wang, Guoqiang Zhong 0001 |
PRCV (4) | 4 |
| 2021 | A review on the attention mechanism of deep learning
Zhaoyang Niu, Guoqiang Zhong 0001, Hui Yu 0001 |
Neurocomputing | 2 |
| 2020 | A Feature Fusion Network for Multi-modal Mesoscale Eddy Detection
Zhenlin Fan, Guoqiang Zhong 0001 |
ICONIP (1) | 2 |
| 2020 | BEDNet: Bi-directional Edge Detection Network for Ocean Front Detection
Qingyang Li 0007, Zhenlin Fan, Guoqiang Zhong 0001 |
ICONIP (4) | 3 |
| 2020 | MCRN: A New Content-Based Music Classification and Recommendation Network
Yuxu Mao, Guoqiang Zhong 0001, Haizhen Wang, Kaizhu Huang |
ICONIP (4) | 2 |
| 2020 | Feature Redirection Network for Few-Shot Classification
Guoqiang Zhong 0001, Yuxu Mao, Kaizhu Huang |
ICONIP (4) | 2 |
| 2020 | EDNet: A Mesoscale Eddy Detection Network with Multi-Modal DataabstractMesoscale eddies play an important role in the transportation and distribution of energy, material and heat in the global ocean. Therefore, mesoscale eddy detection has been researched for a long time. At present, several deep learning models have been proposed for mesoscale eddy detection. However, most of these methods only use single-modal data, while ignoring data of other modals closely related to mesoscale eddy detection. In this paper, we introduce a multi-modal mesoscale eddy dataset, consisting of the satellite data in three modals, i.e., sea surface height (SSH), sea surface temperature (SST) and velocity of flow. Furthermore, we propose an EDNet (Eddy Detection Network), which contains four modules, i.e., multi-modal data fusion module, deep fusion module, region proposal module and head module. We use multi-modal data fusion module to fuse multi-modal data, use deep fusion module to learn the feature representations of the fused multi-modal data and use the region proposal module to generate region proposals containing the mesoscale eddies. There are two branches in the head module, one for classifying and locating the mesoscale eddies, while the other for providing pixel-level instance segmentation of the mesoscale eddies. The experimental results show that EDNet based on multi-modal data fusion significantly improves the accuracy of mesoscale eddy detection over previous approaches. Zhenlin Fan, Guoqiang Zhong 0001, Hongxu Wei |
IJCNN | 2 |
| 2020 | MetaCGAN: A Novel GAN Model for Generating High Quality and Diversity Images with Few Training DataabstractGiven a large amount of data in the base classes and a small number of data in the new classes, meta-learning can learn prior experience from the base classes and transfer knowledge to the new classes by generating network parameters for data generation. In this paper, we propose a novel generative adversarial network called MetaCGAN for generating high quality and diversity images to achieve data augmentation for the new classes with few data. In particular, MetaCGAN consists of two modules, the conditional GAN (CGAN) and MetaNet modules. The CGAN module is our skeleton network that is applied to generate images, while the MetaNet module is our auxiliary network that is applied to provide deconvolutional weights for the generator of CGAN. Experimental results on the MNIST, Fashion MNIST and CelebA data sets demonstrate the superiority of MetaCGAN over baseline models. Both qualitative and quantitative results show that the MetaNet module can learn prior knowledge and transfer it from the base classes to the new classes, which is beneficial for generating high quality and diversity images to the new classes with few images. Guoqiang Zhong 0001 |
IJCNN | 2 |
| 2020 | DNA computing inspired deep networks design
Guoqiang Zhong 0001, Tao Li 0031, Wencong Jiao, Li-Na Wang, Junyu Dong, Cheng-Lin Liu 0001 |
Neurocomputing | 1 |
| 2020 | Generative adversarial networks with mixture of t-distributions noise for diverse image generation
Jinxuan Sun, Guoqiang Zhong 0001, Yang Chen 0036, Tao Li 0031, Kaizhu Huang |
Neural Networks | 2 |
| 2020 | Generative adversarial networks with decoder-encoder output noises
Guoqiang Zhong 0001, Youzhao Yang, Dahan Wang, Kaizhu Huang |
Neural Networks | 1 |
| 2019 | Learnable Gabor Convolutional Networks
Guoqiang Zhong 0001, Wencong Jiao, Biao Shen, Dongdong Xia |
ICONIP (4) | 1 |
| 2019 | Recovering Super-Resolution Generative Adversarial Network for Underwater Images
Yang Chen 0036, Jinxuan Sun, Wencong Jiao, Guoqiang Zhong 0001 |
ICONIP (4) | 4 |
| 2019 | AutoML for DenseNet Compression
Wencong Jiao, Tao Li 0031, Guoqiang Zhong 0001, Li-Na Wang |
ICONIP (3) | 3 |
| 2019 | Enhanced LSTM with Batch Normalization
Li-Na Wang, Guoqiang Zhong 0001, Shoujun Yan, Junyu Dong, Kaizhu Huang |
ICONIP (1) | 2 |
| 2018 | MusicCNNs: A New Benchmark on Content-Based Music Recommendation
Guoqiang Zhong 0001, Haizhen Wang, Wencong Jiao |
ICONIP (1) | 1 |
| 2018 | Convolutional Discriminant AnalysisabstractSoftmax regressor is arguably the most commonly used classifier in convolutional neural networks (CNNs). However, the cross-entropy based softmax loss only supervises the deep neural networks to learn effective representations of data, but does not explicitly enforce the separability between the classes. In this paper, we propose a novel convolutional neural network model, called convolutional discriminative analysis (CDA). Beyond the softmax loss, CDA employs a convolutional discriminant loss (CD-Loss), which minimizes the distance between the sample and its class center while maximizes the distance between the sample and its adversarial class center in the space of the learned deep representations. Extensive experiments on two benchmark data sets, Fashion-MNIST and CIFAR-10, demonstrate the superiority of CDA over traditional deep CNNs on the image classification tasks. Guoqiang Zhong 0001, Xu-Yao Zhang, Hongxu Wei |
ICPR | 1 |
| 2018 | Merging Neurons for Structure Compression of Deep NetworksabstractDeep neural networks are increasingly used in many fields, such as pattern recognition, computer vision, and natural language processing. However, how to apply deep neural networks in mobile settings has become an urgent issue, as mobile devices are getting more and more popularity. This is mainly due to the fact that mobile devices usually have very limited computation and storage resources, which prevents from running a large-scale deep network. This paper proposes a novel method for structure compression of deep neural networks. The main idea is to merge the neurons and connections of the original network using clustering methods. To the end, the new network after compression possesses much less parameters, which leads to reduced requirements for computation and storage resources. Experiments on benchmark data sets demonstrate that the proposed method can greatly improve the efficiency of deep neural networks, while retain their learning capability. Guoqiang Zhong 0001, Huiyu Zhou 0001 |
ICPR | 1 |
| 2018 | Deep Gabor Scattering Network for Image Classification
Benxiu Liu, Haizhen Wang, Guoqiang Zhong 0001, Junyu Dong |
PRCV (2) | 4 |
| 2018 | Natural image illuminant estimation via deep non-negative matrix factorisationabstractThe influence of environmental light sources affects the colour cast in natural images. In computer vision, biased colours have a significant influence on object recognition and classification. Illuminant estimation aims to eliminate these effects and obtain the image in canonical white light. In this study, the authors propose a deep non‐negative matrix factorisation (DeepNMF) method to estimate the illuminant of colour‐biased images. DeepNMF deeply factorises the input matrix into multiple layers, separating the image into patches and reshaping each channel of the patch as an [ R , G , B ] matrix. Based on the diagonal model, they assume that the final layer is the estimated illuminant of each patch. Mean pooling is then used to estimate the illuminant of the overall image. The angular error is used as a metric to test the authors’ method on three commonly used colour constancy datasets. The results show that the proposed method is comparable to state‐of‐the‐art methods, although it is simpler to implement. As the proposed method uses a single image as input, it does not require a learning process. Guoqiang Zhong 0001, Junyu Dong |
IET Image Process. | 2 |
| 2018 | An anchor-based spectral clustering methodabstractSpectral clustering is one of the most popular and important clustering methods in pattern recognition, machine learning, and data mining. However, its high computational complexity limits it in applications involving truly large-scale datasets. For a clustering problem with n samples, it needs to compute the eigenvectors of the graph Laplacian with O ( n 3 ) time complexity. To address this problem, we propose a novel method called anchor-based spectral clustering (ASC) by employing anchor points of data. Specifically, m ( m ≪ n ) anchor points are selected from the dataset, which can basically maintain the intrinsic (manifold) structure of the original data. Then a mapping matrix between the original data and the anchors is constructed. More importantly, it is proved that this data-anchor mapping matrix essentially preserves the clustering structure of the data. Based on this mapping matrix, it is easy to approximate the spectral embedding of the original data. The proposed method scales linearly relative to the size of the data but with low degradation of the clustering performance. The proposed method, ASC, is compared to the classical spectral clustering and two state-of-the-art accelerating methods, i.e., power iteration clustering and landmark-based spectral clustering, on 10 real-world applications under three evaluation metrics. Experimental results show that ASC is consistently faster than the classical spectral clustering with comparable clustering performance, and at least comparable with or better than the state-of-the-art methods on both effectiveness and efficiency. Qin Zhang 0008, Guoqiang Zhong 0001, Junyu Dong |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2018 | Banzhaf random forests: Cooperative game theory based random forests with consistency
Jianyuan Sun, Guoqiang Zhong 0001, Kaizhu Huang, Junyu Dong |
Neural Networks | 2 |
| 2018 | SLMOML: Online Metric Learning With Global ConvergenceabstractMetric and similarity learning are important approaches to classification and retrieval. To efficiently learn a distance metric or a similarity function, online learning algorithms have been widely applied. In general, however, existing online metric and similarity learning algorithms have limited performance in real-world classification and retrieval applications. In this paper, we introduce a convergent online metric learning model named scalable large margin online metric learning (SLMOML). SLMOML belongs to the passive-aggressive learning family. At each step, it adopts the LogDet divergence to maintain the closeness between two successively learned Mahalanobis matrices, and utilizes the hinge loss to enforce a large margin between relatively dissimilar samples. More importantly, the Mahalanobis matrix can be updated by closed-form solution at each step. Furthermore, if the initial matrix is positive semi-definite (PSD), the learned matrices in the following steps are always PSD. Based on the Karush–Kuhn–Tucker conditions and the built equivalence between the passive-aggressive learning family and the Bregman projections, we have proved the global convergence of SLMOML. Extensive experiments on classification and retrieval tasks demonstrate the effectiveness and efficiency of SLMOML. Guoqiang Zhong 0001, Sheng Li 0001, Yun Fu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Perception driven texture generationabstractThis paper investigates a novel task of generating texture images from perceptual descriptions. Previous work on texture generation focused on either synthesis from examples or generation from procedural models. Generating textures from perceptual attributes has not been extensively studied yet. Meanwhile, perceptual attributes, such as directionality, regularity and roughness are important factors for human observers to describe a texture. In this paper, we propose a joint deep network model that combines adversarial training and perceptual feature regression for texture generation, while only random noise and user-defined perceptual attributes are required as input. In this model, a preliminary trained convolutional neural network is essentially integrated with the adversarial framework, which can drive the generated textures to possess given perceptual attributes. An important aspect of the proposed model is that, if we change one of the input perceptual features, the corresponding appearance of the generated textures will also be changed. We designed several experiments to validate the effectiveness of the proposed method. The results show that the proposed method can produce high-quality texture images with desired perceptual properties. Yanhai Gan, Huifang Chi, Ying Gao 0005, Jun Liu 0055, Guoqiang Zhong 0001, Junyu Dong |
ICME | 5 |
| 2017 | Underwater image colour constancy based on DSNMFabstractDifferent wavelengths of light may undergo changes in underwater environment resulting in altered images. For example, the presence of floating particles causes underwater images to appear bluish and blurred. In this study, the authors propose a method called the deep sparse non‐negative matrix factorisation (DSNMF) to estimate the illumination of an underwater image. The image under observation is divided into patches and each channel of a single patch is reshaped as an [ R, G, B ] matrix. The DSNMF method deeply factorises each input matrix into multiple layers with a sparseness constraint. The last layer of the factorised matrix is used as the illumination of the patch. The sparseness constraint adjusts the appearance of the final image. After factorisation, the estimated illumination is applied to each patch of the original image to obtain the final image. Compared with state‐of‐the‐art underwater image enhancement methods using no reference image quality assessment, not only does the proposed method outperforms current techniques in terms of its visual effect and IQA, but is also simpler to implement. Guoqiang Zhong 0001, Junyu Dong |
IET Image Process. | 2 |
| 2017 | Random Multi-Graphs: A semi-supervised learning framework for classification of high dimensional data
Qin Zhang 0008, Jianyuan Sun, Guoqiang Zhong 0001, Junyu Dong |
Image Vis. Comput. | 3 |
| 2017 | Tandem hidden Markov models using deep belief networks for offline handwriting recognitionabstractUnconstrained offline handwriting recognition is a challenging task in the areas of document analysis and pattern recognition. In recent years, to sufficiently exploit the supervisory information hidden in document images, much effort has been made to integrate multi-layer perceptrons (MLPs) in either a hybrid or a tandem fashion into hidden Markov models (HMMs). However, due to the weak learnability of MLPs, the learnt features are not necessarily optimal for subsequent recognition tasks. In this paper, we propose a deep architecture-based tandem approach for unconstrained offline handwriting recognition. In the proposed model, deep belief networks are adopted to learn the compact representations of sequential data, while HMMs are applied for (sub-)word recognition. We evaluate the proposed model on two publicly available datasets, i.e., RIMES and IFN/ENIT, which are based on Latin and Arabic languages respectively, and one dataset collected by ourselves called Devanagari (an Indian script). Extensive experiments show the advantage of the proposed model, especially over the MLP-HMMs tandem approaches. Partha Pratim Roy 0001, Guoqiang Zhong 0001, Mohamed Cheriet |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2017 | Prediction of Sea Surface Temperature Using Long Short-Term MemoryabstractThis letter adopts long short-term memory (LSTM) to predict sea surface temperature (SST), and makes short-term prediction, including one day and three days, and long-term prediction, including weekly mean and monthly mean. The SST prediction problem is formulated as a time series regression problem. The proposed network architecture is composed of two kinds of layers: an LSTM layer and a full-connected dense layer. The LSTM layer is utilized to model the time series relationship. The full-connected layer is utilized to map the output of the LSTM layer to a final prediction. The optimal setting of this architecture is explored by experiments and the accuracy of coastal seas of China is reported to confirm the effectiveness of the proposed method. The prediction accuracy is also tested on the SST anomaly data. In addition, the model's online updated characteristics are presented. Qin Zhang 0008, Junyu Dong, Guoqiang Zhong 0001, Xin Sun 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2016 | Learning perceptual texture similarity and relative attributes from computational featuresabstractPrevious work has shown that perceptual texture similarity and relative attributes cannot be well described by computational features. In this paper, we propose to predict human's visual perception of texture images by learning a non-linear mapping from computational feature space to perceptual space. Hand-crafted features and deep features, which were successfully applied in texture classification tasks, were extracted and used to train Random Forest and rankSVM models against perceptual data from psychophysical experiments. Three texture datasets were used to test our proposed method and the experiments show that the predictions of such learnt models are in high correlation with human's results. Jianwen Lou, Lin Qi 0004, Junyu Dong, Hui Yu 0001, Guoqiang Zhong 0001 |
IJCNN | 5 |
| 2016 | Relational Fisher Analysis: A general framework for dimensionality reductionabstractIn this paper, we propose a novel and general framework for dimensionality reduction, called Relational Fisher Analysis (RFA). Unlike traditional dimensionality reduction methods, such as linear discriminant analysis (LDA) and marginal Fisher analysis (MFA), RFA seamlessly integrates relational information among data into the representation learning framework, which in general provides strong evidence for related data to belong to the same class. To address nonlinear dimensionality reduction problems, we extend RFA to its kernel version. Furthermore, the convergence of RFA is also proved in this paper. Extensive experiments on documents understanding and recognition, face recognition and other applications from the UCI machine learning repository demonstrate the effectiveness and efficiency of RFA. Guoqiang Zhong 0001, Yaxin Shi, Mohamed Cheriet |
IJCNN | 1 |
| 2016 | Deep hashing learning networksabstractHashing-based methods seek compact and efficient binary codes that preserve the similarity between data. For most existing hashing methods, an input (e.g. image) is first encoded as a vector of hand-crafted visual feature, followed by a hash projection and quantization step to obtain the compact binary vector. Most of hand-crafted features only encode low-level information of the input, the feature may not preserve semantic similarities of pairwise inputs. Meanwhile, the hash function learning process is independent with the feature representation, so that the feature may not be optimal for the hash projection. In this paper, we propose a supervised hashing learning method based on a well designed deep convolutional neural network, which tries to learn hashing code and compact representations of data simultaneously. Particularly, the proposed model learns binary codes by adding a compact sigmoid layer before the classifier layer. Experiments on several image data sets show that the proposed model outperforms other state-of-the-art hashing learning approaches. Guoqiang Zhong 0001, Pan Yang 0004, Sijiang Wang, Junyu Dong |
IJCNN | 1 |
| 2016 | Scalable large margin online metric learningabstractWe present a novel online metric learning model, called scalable large margin online metric learning (SLMOML). SLMOML belongs to the passive-aggressive family of learning models. In the formulation of SLMOML, we use the LogDet divergence to measure the closeness between two continuously learned matrices, which naturally ensures the positive semi-definiteness of the learned matrix at each iteration, provided the initial matrix is positive semi-definite. In addition, a hinge loss is used to maintain a large margin of distance between relatively dissimilar data. Using the Karush-Kuhn-Tucker (KKT) condition, the updating rule of SLMOML can be equivalently viewed as Bregman projections. Based on this fact, we have proved the global convergence of SLMOML. Extensive experiments on real world applications demonstrate the superiority of SLMOML over state-of-the-art metric learning and similarity learning approaches. Guoqiang Zhong 0001, Sheng Li 0001, Yun Fu 0001 |
IJCNN | 1 |
| 2016 | Improving short text classification by learning vector representations of both words and hidden topics
Heng Zhang 0028, Guoqiang Zhong 0001 |
Knowl. Based Syst. | 2 |
| 2015 | 3D Texture Recognition for RGB-D Images
Guoqiang Zhong 0001, Yaxin Shi, Junyu Dong |
CAIP (2) | 1 |
| 2015 | Stretching deep architectures for text recognitionabstractIn recent years, many deep architectures have been proposed for handwritten text recognition. However, most of the previous deep models need large scale training data and a long training time to obtain good results. In this paper, we propose a novel deep learning method based on “stretching” the projection matrices of stacked feature learning models. We call the proposed method “stretching deep architectures” (or SDA). In the implementation of SDA, stacked feature learning models are first learned layer by layer, and then the stretching technique is applied on the weight matrices between successive layers. As the feature learning models can be efficiently optimized and the stretching results can be easily computed, the training of SDA is very fast and no back propagation is needed. We have tested SDA on handwritten digits recognition, Arabic subword recognition and English letter recognition tasks. Extensive experiments demonstrate that SDA performs not only better than shallow feature learning models, but also state-of-the-art deep learning models. Yuchen Zheng 0001, Yajuan Cai, Guoqiang Zhong 0001, Youssouf Chherawala, Yaxin Shi, Junyu Dong |
ICDAR | 3 |
| 2015 | Is DeCAF Good Enough for Accurate Image Classification?
Yajuan Cai, Guoqiang Zhong 0001, Yuchen Zheng 0001, Kaizhu Huang, Junyu Dong |
ICONIP (2) | 2 |
| 2015 | Tensor representation learning based image patch analysis for text identification and recognition
Guoqiang Zhong 0001, Mohamed Cheriet |
Pattern Recognit. | 1 |
| 2014 | Low-Rank Tensor Learning with Discriminant Analysis for Action Classification and Image RecoveryabstractTensor completion is an important topic in the area of image processing and computer vision research, which is generally built on extraction of the intrinsic structure of the tensor data. Drawing on this fact, action classification, relying heavily on the extracted features of high-dimensional tensors, may indeed benefit from tensor completion techniques. In this paper, we propose a low-rank tensor completion method for action classification, as well as image recovery. Since there may exist distortion and corruption in the tensor representations of video sequences, we project the tensors into a subspace, which contains the invariant structure of the tensors. In order to integrate useful supervisory information for classification, we adopt a discriminant analysis criterion to learn the projection matrices. The resulting multi-variate optimization problem can be effectively solved using the augmented Lagrange multiplier (ALM) algorithm. Experiments demonstrate that our method results with better accuracy compared with some other state-of-the-art low-rank tensor representation learning approaches on the MSR Hand Gesture 3D database and the MSR Action 3D database. By denoising the Multi-PIE face database, our experimental setup testifies the proposed method can also be employed to recover images. Chengcheng Jia, Guoqiang Zhong 0001, Yun Raymond Fu |
AAAI | 2 |
| 2014 | Large Margin Low Rank Tensor AnalysisabstractWe present a supervised model for tensor dimensionality reduction, which is called large margin low rank tensor analysis (LMLRTA). In contrast to traditional vector representation-based dimensionality reduction methods, LMLRTA can take any order of tensors as input. And unlike previous tensor dimensionality reduction methods, which can learn only the low-dimensional embeddings with a priori specified dimensionality, LMLRTA can automatically and jointly learn the dimensionality and the low-dimensional representations from data. Moreover, LMLRTA delivers low rank projection matrices, while it encourages data of the same class to be close and of different classes to be separated by a large margin of distance in the low-dimensional tensor space. LMLRTA can be optimized using an iterative fixed-point continuation algorithm, which is guaranteed to converge to a local optimal solution of the optimization problem. We evaluate LMLRTA on an object recognition application, where the data are represented as 2D tensors, and a face recognition application, where the data are represented as 3D tensors. Experimental results show the superiority of LMLRTA over state-of-the-art approaches. Guoqiang Zhong 0001, Mohamed Cheriet |
Neural Comput. | 1 |
| 2013 | An Empirical Evaluation of Supervised Dimensionality Reduction for RecognitionabstractIn the literature, many dimensionality reduction methods have been proposed and applied to recognition tasks, including handwritten digits recognition, character recognition and string recognition. However, it is usually difficult for the researchers to decide which method is the optimal choice for the problem at hand. In this paper, we empirically compare some supervised dimensionality reduction methods on handwritten digits recognition, English letter recognition and ancient Arabic sub word recognition, to evaluate their performance on the recognition tasks. These compared methods include traditional linear dimensionality reduction approach (linear discriminant analysis, LDA), locality-based manifold learning approach (marginal Fisher analysis, MFA) and relational learning approach (probabilistic relational principal component analysis, PRPCA). Experimental results and statistical tests show that locality-based manifold learning approach (MFA) generally performs well in terms of recognition accuracy, but with high computational complexity, traditional linear dimensionality reduction approach (LDA) is efficient, but not necessarily to deliver the best result, relational learning approach (PRPCA) is promising, and more efforts should be dedicated to this area. Guoqiang Zhong 0001, Youssouf Chherawala, Mohamed Cheriet |
ICDAR | 1 |
| 2013 | Adaptive Error-Correcting Output Codes
Guoqiang Zhong 0001, Mohamed Cheriet |
IJCAI | 1 |
| 2013 | Error-correcting output codes based ensemble feature extraction
Guoqiang Zhong 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 1 |
| 2012 | Joint learning of error-correcting output codes and dichotomizers from data
Guoqiang Zhong 0001, Kaizhu Huang, Cheng-Lin Liu 0001 |
Neural Comput. Appl. | 1 |
| 2011 | Low Rank Metric Learning with Manifold RegularizationabstractIn this paper, we present a semi-supervised method to learn a low rank Mahalanobis distance function. Based on an approximation to the projection distance from a manifold, we propose a novel parametric manifold regularizer. In contrast to previous approaches that usually exploit side information only, our proposed method can further take advantages of the intrinsic manifold information from data. In addition, we focus on learning a metric of low rank directly, this is different from traditional approaches that often enforce the l1norm on the metric. The resulting configuration is convex with respect to the manifold structure and the distance function, respectively. We solve it with an alternating optimization algorithm, which proves effective to find a satisfactory solution. For efficient implementation, we even present a fast algorithm, in which the manifold structure and the distance function are learned independently without alternating minimization. Experimental results over 12 standard UCI data sets demonstrate the advantages of our method. Guoqiang Zhong 0001, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICDM | 1 |
| 2010 | Gaussian Process Latent Random FieldabstractIn this paper, we propose a novel supervised extension of GPLVM, called Gaussian process latent random field (GPLRF), by enforcing the latent variables to be a Gaussian Markov random field with respect to a graph constructed from the supervisory information. Guoqiang Zhong 0001, Wu-Jun Li, Dit-Yan Yeung, Xinwen Hou, Cheng-Lin Liu 0001 |
AAAI | 1 |
| 2010 | Learning ECOC and Dichotomizers Jointly from Data
Guoqiang Zhong 0001, Kaizhu Huang, Cheng-Lin Liu 0001 |
ICONIP (1) | 1 |