EDBT 2026 Demo / reviewers in the wild / expert
Wanqing Zhao
dblp:75/8384
· DBLP profile ↗
43ranked-venue papers
15as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Universal EEG Epilepsy Detection via Evidential Multi-View De-BiasingabstractEpilepsy is a widespread neurological disorder characterized by highly patient-specific EEG patterns. Existing EEG-based seizure detection methods either train individualized models for each patient or adapt models pre-trained on known patients to new ones. However, when encountering previously unseen patients, these methods typically require retraining or fine-tuning, which limits their practical utility in clinical settings. This limitation can be linked to biases caused by patient-specific variations, which obscure the underlying pathological patterns of seizures. To address this, we propose an evidential multi-view framework that reinforces the learning of core epileptic features by promoting consistency across multiple views and reducing reliance on high-uncertainty, patient-specific segments. Specifically, we introduce Bias-guided Fisher-Evidential Multi-View Learning (BF-EML) to guide the model toward discovering intrinsic seizure patterns. BF-EML employs a two-stage training architecture: In Stage 1, we use the Fisher Information Matrix to reorder EEG segments by uncertainty and deliberately train a biased feature generator on low-evidence segments. In Stage 2, we design a dual-branch network where the biased and unbiased branches are alternately trained, encouraging the unbiased branch to reduce its reliance on patient-specific biases. Finally, we introduce a shift-calibrated fusion strategy to enhance the consistency of pathogenic feature integration. Extensive experiments on public datasets and a clinical dataset demonstrate that our method achieves superior performance in both single- and multi-patient scenarios. Importantly, it generalizes well to unseen patients without the need for retraining. Ziqi Wen, Wanqing Zhao, Jie Zhao 0013, Wei Zhao 0019 |
AAAI | 3 |
| 2026 | Act-LLM: A whole-process chain for character-centric role-playing with large language models
Xiaoxu Han, Wanqing Zhao, Ziyu Guan, Jinye Peng 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Forecasting tourism stock index dynamics: a multiscale deep learning framework integrating emerging media data
Wanqing Zhao, Dao Lan |
Expert Syst. Appl. | 3 |
| 2026 | Multi-modal mutual-guidance conditional prompt learning for vision-language models
Shijun Yang, Xiang Zhang 0018, Wanqing Zhao, Qiyao Hu, Xianlin Peng |
Expert Syst. Appl. | 3 |
| 2025 | Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task LearningabstractGenerative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often undermine their effectiveness in downstream visual tasks. This paper introduces the Iterative Self-Training with Class-Aware Text-to-Image Synthesis (IST-CATS) framework, which addresses these challenges by integrating a class-aware text-to-image synthesis (CATS) component with an iterative self-training (IST) strategy. CATS innovatively introduces a class-aware chain approach to generate detailed descriptions. These descriptions act as prompts for a diffusion model, enabling the creation of a diverse of images accompanied by distinguishable objects against the background. The generated images can be easily pseudo-labeled by an unsupervised instance segmentation method, and then noisy pseudo labels can be effectively purified by a novel feature similarity-based filtering mechanism. The generated images underpin our IST, which progressively enhances vision models and refines pseudo labels through self-training and our proposed label filtering strategy (LabFilt). LabFilt meticulously improves the quality of pseudo labels by employing class-adaptive techniques at both the pixel and object levels, ensuring refined pseudo-label accuracy. IST-CATS demonstrates superior performance in object detection and semantic segmentation compared to traditional synthetic and semi/weakly-supervised methods, effectively addressing data collection and annotation challenges. Xiang Zhang 0018, Wanqing Zhao, Pengyang Li, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
AAAI | 2 |
| 2025 | Environment-Agnostic Pose: Generating Environment-Independent Object Representations for 6D Pose Estimation
Shaobo Zhang 0006, Wanqing Zhao, Wei Zhao 0019, Ziyu Guan, Jinye Peng 0001 |
ICCV | 3 |
| 2025 | Beyond Equal Views: Strength-Adaptive Evidential Multi-View Learning
Ziqi Wen, Jie Zhao 0013, Wanqing Zhao, Jinlong Yu, Haishun Chen, Ziyu Guan, Wei Zhao 0019 |
ACM Multimedia | 4 |
| 2025 | Semantic image segmentation via dynamic curriculum learning
Xiang Zhang 0018, Wanqing Zhao, Chenji Wang, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
Appl. Intell. | 2 |
| 2025 | An efficient parallel mesh generation method for finite element based analysis of large complex architecture
Wanqing Zhao, Chunnan Li, Tongkun Deng, Jun Wang 0078, Jinye Peng 0001 |
Comput. Aided Des. | 2 |
| 2025 | A unified data-driven approach under deep reinforcement learning with direct control responses for microgrid operationsabstractMicrogrid systems have now seen many integrations with energy storage systems (ESS) and renewable energy sources (RES) to supply cleaner and cheaper energy. A pressing challenge is how to optimally meet the requirements ranging from reducing operational costs and carbon footprints to relieving the grid constraints, alongside the consideration of uncertainties in the supply and demand. To embark on this challenge, this paper proposes a deep reinforcement learning (DRL) approach with direct control responses to optimize multi-objective microgrid operations. First, a new objective function is derived to build a direct response between the control action and the optimization objectives, aiming to improve the learning efficiency. Second, a unified control scheme is designed to study the combined use of past observations and predicted data for microgrid controls. Third, a realistic microgrid model is created to incorporate battery charging and discharging processes with dynamic efficiency and nonlinear battery degradation. Finally, the effectiveness of the proposed approach is validated through various simulations conducted on a US case study, with an additional Norwegian microgrid presented in the supplementary material. The results suggest that the annual reward in the US microgrid can be improved by 139.33% over the baseline (vanilla DQN with a conventional scheme) under perfect predictions, and by 125.45% under noisy predictions. • A direct-response reward function is derived to improve the learning efficiency. • A unified control scheme is designed to organize the state space. • A realistic microgrid model is created for power optimization. Fulong Yao, Wanqing Zhao, Matthew Forshaw, Wenju Zhou |
Knowl. Based Syst. | 2 |
| 2025 | Cross-Modal Guided Visual Representation Learning for Social Image RetrievalabstractSocial images are often associated with rich but noisy tags from community contributions. Although social tags can potentially provide valuable semantic training information for image retrieval, existing studies all fail to effectively filter noises by exploiting the cross-modal correlation between image content and tags. The current cross-modal vision-and-language representation learning methods, which selectively attend to the relevant parts of the image and text, show a promising direction. However, they are not suitable for social image retrieval since: (1) they deal with natural text sequences where the relationships between words can be easily captured by language models for cross-modal relevance estimation, while the tags are isolated and noisy; (2) they take (image, text) pair as input, and consequently cannot be employed directly for unimodal social image retrieval. This paper tackles the challenge of utilizing cross-modal interactions to learn precise representations for unimodal retrieval. The proposed framework, dubbed CGVR (Cross-modal Guided Visual Representation), extracts accurate semantic representations of images from noisy tags and transfers this ability to image-only hashing subnetwork by a carefully designed training scheme. To well capture correlated semantics and filter noises, it embeds a priori common-sense relationship among tags into attention computation for joint awareness of textual and visual context. Experiments show that CGVR achieves approximately 8.82 and 5.45 points improvement in MAP over the state-of-the-art on two widely used social image benchmarks. CGVR can serve as a new baseline for the image retrieval community. The code is provided at https://github.com/zhaowanqing/CGVR. Ziyu Guan, Wanqing Zhao, Hongmin Liu 0001, Yuta Nakashima, Noboru Babaguchi, Xiaofei He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Learning Cross-View Consistent 3D Keypoints for Object 6D Pose EstimationabstractAccurate 6D object pose estimation from RGB images is crucial for various computer vision applications, such as augmented reality, robotic manipulation and autonomous driving. Existing methods often rely on extensive labeled data, either manually annotated or synthetically generated, which can be laborious and impractical for real-world deployment. To address these challenges, we propose OK-POSE, a keypoint-based 6D object pose estimation method that leverages relative transformations between viewpoints for training. By utilizing pairs of images with object annotations and relative transformation information, OK-POSE automatically learns to detect 3D keypoints of objects, enabling geometrically and visually consistent pose estimation. The simplicity and accessibility of obtaining relative transformation information, which can be acquired from inexpensive binocular cameras or common smartphone devices, significantly reduce labeling costs and mitigate domain gap issues associated with synthetic data. Experimental results demonstrate that OK-POSE achieves competitive performance compared to methods relying on explicit 3D annotations or object 3D models. Moreover, we provide insights into the data collection process and introduce OK-POSE++, an enhanced version with optimized network architecture and loss functions, yielding further improvements in performance. Our approach offers a practical solution for 6D object pose estimation, suitable for real-world applications in scenarios where extensive 3D annotations or object models are unavailable. The code is released athttps://github.com/acmff22/OKPOSE. Shaobo Zhang 0006, Wanqing Zhao, Ziyu Guan, Wei Zhao 0019, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Enhancing Fake News Detection in Social Media via Label Propagation on Cross-modal Tweet GraphabstractFake news detection in social media has become increasingly important due to the rapid proliferation of personal media channels and the consequential dissemination of misleading information. Existing methods, which primarily rely on multimodal features and graph-based techniques, have shown promising performance in detecting fake news. However, they still face a limitation, i.e., sparsity in graph connections, which hinders capturing possible interactions among tweets. This challenge has motivated us to explore a novel method that densifies the graph's connectivity to capture denser interaction better. Our method constructs a cross-modal tweet graph using CLIP, which encodes images and text into a unified space, allowing us to extract potential connections based on similarities in text and images. We then design a Feature Contextualization Network with Label Propagation (FCN-LP) to model the interaction among tweets as well as positive or negative correlations between predicted labels of connected tweets. The propagated labels from the graph are weighted and aggregated for the final detection. To enhance the model's generalization ability to unseen events, we introduce a domain generalization loss that ensures consistent features between tweets on seen and unseen events. We use three publicly available fake news datasets, Twitter, PHEME, and Weibo, for evaluation. Our method consistently improves the performance over the state-of-the-art methods on all benchmark datasets and effectively demonstrates its aptitude for generalizing fake news detection in social media. Wanqing Zhao, Yuta Nakashima, Haiyuan Chen, Noboru Babaguchi |
ACM Multimedia | 1 |
| 2023 | Data to intelligence: The role of data-driven models in wastewater treatment
Majid Bahramian, Recep Kaan Dereli, Wanqing Zhao, Matteo Giberti, Eoin Casey |
Expert Syst. Appl. | 3 |
| 2022 | Class Guided Channel Weighting Network for Fine-Grained Semantic SegmentationabstractDeep learning has achieved promising performance on semantic segmentation, but few works focus on semantic segmentation at the fine-grained level. Fine-grained semantic segmentation requires recognizing and distinguishing hundreds of sub-categories. Due to the high similarity of different sub-categories and large variations in poses, scales, rotations, and color of the same sub-category in the fine-grained image set, the performance of traditional semantic segmentation methods will decline sharply. To alleviate these dilemmas, a new approach, named Class Guided Channel Weighting Network (CGCWNet), is developed in this paper to enable fine-grained semantic segmentation. For the large intra-class variations, we propose a Class Guided Weighting (CGW) module, which learns the image-level fine-grained category probabilities by exploiting second-order feature statistics, and use them as global information to guide semantic segmentation. For the high similarity between different sub-categories, we specially build a Channel Relationship Attention (CRA) module to amplify the distinction of features. Furthermore, a Detail Enhanced Guided Filter (DEGF) module is proposed to refine the boundaries of object masks by using an edge contour cue extracted from the enhanced original image. Experimental results on PASCAL VOC 2012 and six fine-grained image sets show that our proposed CGCWNet has achieved state-of-the-art results. Xiang Zhang 0018, Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
AAAI | 2 |
| 2022 | Deep objectness hashing using large weakly tagged photosabstractCNN-based hashing methods have greatly boosted the performance of image retrieval, under the strong supervision of large amounts of manually annotated labels. In recent years, a large number of social media images with user tags have been generated on the Internet. These images can be regarded as weakly labeled training data, which can provide rich samples for training hash network, and greatly reduce the cost of obtaining training data. However, there are noise and visual irrelevant tags in user tags, and different tags may describe different objects in the image. In the previous CNN-based hashing method, a training image usually corresponds to a manual label and generates a hash code. These methods are difficult to use the image described by user tags with noise. For solving the above problem, we propose a CNN-based objectness hash learning method using user tags as a guide for training. First of all, the user tags are roughly filtered to remove noise tags that are not related to the visual content of images. Secondly, we quantify user tags into a unified semantic space and extract the highest-frequency words of the semantic space from similar objectness areas as their labels. Then, these objectness areas with their labels are grouped into a series of triple units as training data. So that the generated hash code can inherit the semantic similarity of the objectness areas well, that is, the Hamming distance between hash codes generated by similar objectness areas is closer, the reverse is the farther. Experimental results on NUS-WIDE and Flickr datasets show that our method can effectively extract object-level semantic information from weak user tags, and improve the accuracy of image retrieval. Fei Xie 0007, Wanqing Zhao, Ziyu Guan, Hexu Wang, Qun Duan |
Neurocomputing | 2 |
| 2022 | Automatic learning for object detection
Xiang Zhang 0018, Hangzai Luo, Wanqing Zhao, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 4 |
| 2022 | TelecomNet: Tag-Based Weakly-Supervised Modally Cooperative Hashing Network for Image RetrievalabstractWe are concerned with using user-tagged images to learn proper hashing functions for image retrieval. The benefits are two-fold: (1) we could obtain abundant training data for deep hashing models; (2) tagging data possesses richer semantic information which could help better characterize similarity relationships between images. However, tagging data suffers from noises, vagueness and incompleteness. Different from previous unsupervised or supervised hashing learning, we propose a novel weakly-supervised deep hashing framework which consists of two stages: weakly-supervised pre-training and supervised fine-tuning. The second stage is as usual. In the first stage, we propose two formulations Tag-basEd weakLy-supErvised Modally COoperative hashing Network (TelecomNet) and Generalized TelecomNet (GTelecomNet). Rather than performing supervision on tags, TelecomNet first learns an observed semantic embedding vector for each image from attached tags and then uses it to guide hashing learning. GTelecomNet introduces a novel semantic network to exploit more precise semantic information. By carefully designing the optimization problem, they can well leverage tagging information and image content for hashing learning. The framework is general and does not depend on specific deep hashing methods. Empirical results on real world datasets show that they significantly increase the performance of state-of-the-art deep hashing methods. Wei Zhao 0019, Ziyu Guan, Xunlian Wu, Wanqing Zhao, Qiguang Miao, Xiaofei He 0001, Quan Wang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Guided Filter Network for Semantic Image SegmentationabstractThe existing publicly available datasets with pixel-level labels contain limited categories, and it is difficult to generalize to the real world containing thousands of categories. In this paper, we propose an approach to generate object masks with detailed pixel-level structures/boundaries automatically to enable semantic image segmentation of thousands of targets in the real world without manually labelling. A Guided Filter Network (GFN) is first developed to learn the segmentation knowledge from an existed dataset, and such GFN then transfers the learned segmentation knowledge to generate initial coarse object masks for the target images. These coarse object masks are treated as pseudo labels to self-optimize the GFN iteratively in the target images. Our experiments on six image sets have demonstrated that our proposed approach can generate object masks with detailed pixel-level structures/boundaries, whose quality is comparable to the manually-labelled ones. Our proposed approach also achieves better performance on semantic image segmentation than most existing weakly-supervised, semi-supervised, and domain adaptation approaches under the same experimental conditions. Xiang Zhang 0018, Wanqing Zhao, Wei Zhang 0016, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Keypoint-Graph-Driven Learning Framework for Object Pose EstimationabstractMany recent 6D pose estimation methods exploited object 3D models to generate synthetic images for training because labels come for free. However, due to the domain shift of data distributions between real images and synthetic images, the network trained only on synthetic images fails to capture robust features in real images for 6D pose estimation. We propose to solve this problem by making the network insensitive to different domains, rather than taking the more difficult route of forcing synthetic images to be similar to real images. Inspired by domain adaption methods, a Domain Adaptive Keypoints Detection Network (DAKDN) including a domain adaption layer is used to minimize the discrepancy of deep features between synthetic and real images. A unique challenge here is the lack of ground truth labels (i.e., keypoints) for real images. Fortunately, the geometry relations between keypoints are invariant under real/synthetic domains. Hence, we propose to use the domain-invariant geometry structure among keypoints as a "bridge" constraint to optimize DAKDN for 6D pose estimation across domains. Specifically, DAKDN employs a Graph Convolutional Network (GCN) block to learn the geometry structure from synthetic images and uses the GCN to guide the training for real images. The 6D poses of objects are calculated using Perspective-n-Point (PnP) algorithm based on the predicted keypoints. Experiments show that our method outperforms state-of-the-art approaches without manual poses labels and competes with approaches using manual poses labels. Shaobo Zhang 0006, Wanqing Zhao, Ziyu Guan, Xianlin Peng, Jinye Peng 0001 |
CVPR | 2 |
| 2021 | Enhanced three-dimensional U-Net with graph-based refining for segmentation of gastrointestinal stromal tumoursabstractAbstract The gastrointestinal stromal tumour (GIST) is a common mesenchymal tumour that lacks specificity of clinical manifestations. Therefore, preoperative accurate localization and accurate prediction of tumour risk are of important clinical value. At present, the diagnosis of GIST relies mainly on manual annotation of CT by professional doctors, which is inefficient and affected by subjective factors. A GIST segmentation algorithm is proposed based on a convolutional neural network to fuse multi‐scale features. The algorithm is applied to GIST segmentation with an improved 3‐D U‐Net method. Skip connections are introduced between encoders and decoders at different layers to account for the obvious differences in tumour size between different cases, which increases the path of information transmission in the network and solves the problem that U‐Net is too weak to simultaneously extract the features of different scales. In addition, due to the difficulty of tumour labelling and the correlation between small intestine segmentation and GIST segmentation, the model of small intestine segmentation is transferred to the model of GIST segmentation. Experiments show that the proposed method achieves better performance than that of the traditional U‐Net. Finally, the graph neural network is introduced to reduce the repetitive work of doctors in refining the segmentation results. Wanqing Zhao, Fei Xie 0007, Ziyu Guan, Wei Zhao 0019 |
IET Comput. Vis. | 3 |
| 2021 | Deep Multiple Instance Hashing for Fast Multi-Object Image SearchabstractMulti-keyword query is widely supported in text search engines. However, an analogue in image retrieval systems, multi-object query, is rarely studied. Meanwhile, traditional object-based image retrieval methods often involve multiple steps separately. In this work, we propose a weakly-supervised Deep Multiple Instance Hashing (DMIH) approach for multi-object image retrieval. Our DMIH approach, which leverages a popular CNN model to build the end-to-end relation between a raw image and the binary hash codes of its multiple objects, can support multi-object queries effectively and integrate object detection with hashing learning seamlessly. We treat object detection as a binary multiple instance learning (MIL) problem and such instances are automatically extracted from multi-scale convolutional feature maps. We also design a conditional random field (CRF) module to capture both the semantic and spatial relations among different class labels. For hashing training, we sample image pairs to learn their semantic relationships in terms of hash codes of the most probable proposals for owned labels as guided by object predictors. The two objectives benefit each other in a multi-task learning scheme. Finally, a two-level inverted index method is proposed to further speed up the retrieval of multi-object queries. Our DMIH approach outperforms state-of-the-arts on public benchmarks for object-based image retrieval and achieves promising results for multi-object queries. Wanqing Zhao, Ziyu Guan, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Learning Deep Network for Detecting 3D Object Keypoints and 6D PosesabstractThe state-of-art 6D object pose detection methods use convolutional neural networks to estimate objects' 6D poses from RGB images. However, they require huge numbers of images with explicit 3D annotations such as 6D poses, 3D bounding boxes and 3D keypoints, either obtained by manual labeling or inferred from synthetic images generated by 3D CAD models. Manual labeling for a large number of images is a laborious task, and we usually do not have the corresponding 3D CAD models of objects in real environment. In this paper, we develop a keypoint-based 6D object pose detection method (and its deep network) called Object Keypoint based POSe Estimation (OK-POSE). OK-POSE employs relative transformation between viewpoints for training. Specifically, we use pairs of images with object annotation and relative transformation information between their viewpoints to automatically discover objects' 3D keypoints which are geometrically and visually consistent. Then, the 6D object pose can be estimated using a keypoint-based geometric reasoning method with a reference viewpoint. The relative transformation information can be easily obtained from any cheap binocular cameras or most smartphone devices, thus greatly lowering the labeling cost. Experiments have demonstrated that OK-POSE achieves acceptable performance compared to methods relying on the object's 3D CAD model or a great deal of 3D labeling. These results show that our method can be used as a suitable alternative when there are no 3D CAD models or a large number of 3D annotations. Wanqing Zhao, Shaobo Zhang 0006, Ziyu Guan, Wei Zhao 0019, Jinye Peng 0001, Jianping Fan 0001 |
CVPR | 1 |
| 2020 | 6D object pose estimation via viewpoint relation reasoning
Wanqing Zhao, Shaobo Zhang 0006, Ziyu Guan, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 1 |
| 2020 | A fast hybrid computer vision technique for real-time embedded bus passenger flow calculation through camera
Ejaz Ul Haq, Huarong Xu, Wanqing Zhao, Jianping Fan 0001, Fazeel Abid |
Multim. Tools Appl. | 4 |
| 2019 | Answer Identification from Product Reviews for User Questions by Multi-Task Attentive NetworksabstractOnline Shopping has become a part of our daily routine, but it still cannot offer intuitive experience as store shopping. Nowadays, most e-commerce Websites offer a Question Answering (QA) system that allows users to consult other users who have purchased the product. However, users still need to wait patiently for others’ replies. In this paper, we investigate how to provide a quick response to the asker by plausible answer identification from product reviews. By analyzing the similarity and discrepancy between explicit answers and reviews that can be answers, a novel multi-task deep learning method with carefully designed attention mechanisms is developed. The method can well exploit large amounts of user generated QA data and a few manually labeled review data to address the problem. Experiments on data collected from Amazon demonstrate its effectiveness and superiority over competitive baselines. Long Chen 0007, Ziyu Guan, Wei Zhao 0019, Wanqing Zhao, Zhou Zhao 0001, Huan Sun 0001 |
AAAI | 4 |
| 2019 | Multilayer feature descriptors fusion CNN models for fine-grained visual recognitionabstractAbstract Fine‐grained image classification is a challenging topic in the field of computer vision. General models based on first‐order local features cannot achieve acceptable performance because the features are not so efficient in capturing fine‐grained difference. A bilinear convolutional neural network (CNN) model exhibits that a second‐order statistical feature is more efficient in capturing fine‐grained difference than a first‐order local feature. However, this framework only considers the extraction of a second‐order feature descriptor, using a single convolutional layer. The potential effective classification features of other convolutional layers are ignored, resulting in loss of recognition accuracy. In this paper, a multilayer feature descriptors fusion CNN model is proposed. It fully considers the second‐order feature descriptors and the first‐order local feature descriptor generated by different layers. Experimental verification was carried out on fine‐grained classification benchmark data sets, CUB‐200‐2011, Stanford Cars, and FGVC‐aircraft. Compared with the bilinear CNN model, the proposed method has improved accuracy by 0.8%, 1.1%, and 5.5%. Compared with the compact bilinear pooling model, there is an accuracy increase of 0.64%, 1.63%, and 1.45%, respectively. In addition, the proposed model effectively uses multiple 1×1 convolution kernels to reduce dimension. The experimental results show that the multilayer low‐dimensional second‐order feature descriptors fusion model has comparable recognition accuracy of the original model. Yong Hou, Hangzai Luo, Wanqing Zhao, Xiang Zhang 0018, Jun Wang 0078, Jinye Peng 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2019 | Plant recognition via leaf shape and margin features
Xiang Zhang 0018, Wanqing Zhao, Hangzai Luo, Long Chen 0007, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 2 |
| 2019 | Automated Model Construction for Combined Sewer Overflow Prediction Based on Efficient LASSO AlgorithmabstractThe prediction of combined sewer overflow (CSO) operation in urban environments presents a challenging task for water utilities. The operation of CSOs (most often in heavy rainfall conditions) prevents houses and businesses from flooding. However, sometimes, CSOs do not operate as they should, potentially bringing environmental pollution risks. Therefore, CSOs should be appropriately managed by water utilities, highlighting the need for adapted decision support systems. This paper proposes an automated CSO predictive model construction methodology using field monitoring data, as a substitute for the commonly established hydrological-hydraulic modeling approach for time-series prediction of CSO statuses. It is a systematic methodology factoring in all monitored field variables to construct time-series dependencies for CSO statuses. The model construction process is largely automated with little human intervention, and the pertinent variables together with their associated time lags for every CSO are holistically and automatically generated. A fast least absolute shrinkage and selection operator solution generating scheme is proposed to expedite the model construction process, where matrix inversions are effectively eliminated. The whole algorithm works in a stepwise manner, invoking either an incremental or decremental movement for including or excluding one model regressor into, or from, the predictive model at every step. The computational complexity is thereby analyzed with the pseudo code provided. Actual experimental results from both single-step ahead (i.e., 15 min) and multistep ahead predictions are finally produced and analyzed on a U.K. pilot area with various types of monitoring data made available, demonstrating the efficiency and effectiveness of the proposed approach. Wanqing Zhao, Thomas H. Beach, Yacine Rezgui |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2018 | Achieving Smart Water Network Management Through Semantically Driven Cognitive Systems
Thomas H. Beach, Shaun Howell, Julia Terlet, Wanqing Zhao, Yacine Rezgui |
PRO-VE | 4 |
| 2018 | Nurturing Virtual Collaborative Networks into Urban Resilience for Seismic Hazards Mitigation
Giulia Cere, Wanqing Zhao, Yacine Rezgui |
PRO-VE | 2 |
| 2018 | Tag-based Weakly-supervised Hashing for Image RetrievalabstractWe are concerned with using user-tagged images to learn proper hashing functions for image retrieval. The benefits are two-fold: (1) we could obtain abundant training data for deep hashing models; (2) tagging data possesses richer semantic information which could help better characterize similarity relationships between images. However, tagging data suffers from noises, vagueness and incompleteness. Different from previous unsupervised or supervised hashing learning, we propose a novel weakly-supervised deep hashing framework which consists of two stages: weakly-supervised pre-training and supervised fine-tuning. The second stage is as usual. In the first stage, rather than performing supervision on tags, the framework introduces a semantic embedding vector (sem-vector) for each image and performs learning of hashing and sem-vectors jointly. By carefully designing the optimization problem, it can well leverage tagging information and image content for hashing learning. The framework is general and does not depend on specific deep hashing methods. Empirical results on real world datasets show that when it is integrated with state-of-art deep hashing methods, the performance increases by 8-10%. Ziyu Guan, Fei Xie 0007, Wanqing Zhao, Long Chen 0007, Wei Zhao 0019, Jinye Peng 0001 |
IJCAI | 3 |
| 2018 | Locally linear spatial pyramid hash for large-scale image search
Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 1 |
| 2017 | Deep Multiple Instance Hashing for Object-based Image RetrievalabstractMulti-keyword query is widely supported in text search engines. However, an analogue in image retrieval systems, multi-object query, is rarely studied. Meanwhile, traditional object-based image retrieval methods often involve multiple steps separately and need expensive location labeling for detecting objects. In this work, we propose a weakly-supervised Deep Multiple Instance Hashing (DMIH) framework for object-based image retrieval. DMIH integrates object detection and hashing learning on the basis of a popular CNN model to build the end-to-end relation between a raw image and the binary hashing codes of multiple objects in it. Specifically, we cast the object detection of each object class as a binary multiple instance learning problem where instances are object proposals extracted from multi-scale convolutional feature maps. For hashing training, we sample image pairs to learn their semantic relationships in terms of hash codes of the most probable proposals for owned labels as guided by object predictors. The two objectives benefit each other in learning. DMIH outperforms state-of-the-arts on public benchmarks for object-based image retrieval and achieves promising results for multi-object queries. Wanqing Zhao, Ziyu Guan, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
IJCAI | 1 |
| 2017 | Spatial pyramid deep hashing for large-scale image retrieval
Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 1 |
| 2017 | MapReduce-based clustering for near-duplicate image identification
Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 1 |
| 2016 | A Heuristic Distributed Task Allocation Method for Multivehicle Multitask Problems and Its Application to Search and Rescue ScenarioabstractUsing distributed task allocation methods for cooperating multivehicle systems is becoming increasingly attractive. However, most effort is placed on various specific experimental work and little has been done to systematically analyze the problem of interest and the existing methods. In this paper, a general scenario description and a system configuration are first presented according to search and rescue scenario. The objective of the problem is then analyzed together with its mathematical formulation extracted from the scenario. Considering the requirement of distributed computing, this paper then proposes a novel heuristic distributed task allocation method for multivehicle multitask assignment problems. The proposed method is simple and effective. It directly aims at optimizing the mathematical objective defined for the problem. A new concept of significance is defined for every task and is measured by the contribution to the local cost generated by a vehicle, which underlies the key idea of the algorithm. The whole algorithm iterates between a task inclusion phase, and a consensus and task removal phase, running concurrently on all the vehicles where local communication exists between them. The former phase is used to include tasks into a vehicle's task list for optimizing the overall objective, while the latter is to reach consensus on the significance value of tasks for each vehicle and to remove the tasks that have been assigned to other vehicles. Numerical simulations demonstrate that the proposed method is able to provide a conflict-free solution and can achieve outstanding performance in comparison with the consensus-based bundle algorithm. Wanqing Zhao, Qinggang Meng, Paul W. H. Chung |
IEEE Trans. Cybern. | 1 |
| 2016 | Optimization of Potable Water Distribution and Wastewater Collection Networks: A Systematic Review and Future Research DirectionsabstractPotable water distribution networks (WDNs) and wastewater collection networks (WWCNs) are the two fundamental constituents of the complex urban water infrastructure. Such water networks require adapted design interventions as part of retrofitting, extension, and maintenance activities. Consequently, proper optimization methodologies are required to reduce the associated capital cost while also meeting the demands of acquiring clean water and releasing wastewater by consumers. In this paper, a systematic review of the optimization of both WDNs and WWCNs, from the preliminary stages of development through to the state-of-the-art, is jointly presented. First, both WDNs and WWCNs are conceptually and functionally described along with illustrative benchmarks. The optimization of water networks across both clean and waste domains is then systematically reviewed and organized, covering all levels of complexity from the formulation of cost functions and constraints, through to traditional and advanced optimization methodologies. The rationales behind employing these methodologies as well as their advantages and disadvantages are investigated. This paper then critically discusses current trends and identifies directions for future research by comparing the existing optimization paradigms within WDNs and WWCNs and proposing common research directions for optimizing water networks. Optimization of urban water networks is a multidisciplinary domain, within which this paper is anticipated to be of great benefit to researchers and practitioners. Wanqing Zhao, Thomas H. Beach, Yacine Rezgui |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2015 | An Efficient LS-SVM-Based Method for Fuzzy System ConstructionabstractThis paper proposes an efficient learning mechanism to build fuzzy rule-based systems through the construction of sparse least-squares support vector machines (LS-SVMs). In addition to the significantly reduced computational complexity in model training, the resultant LS-SVM-based fuzzy system is sparser while offers satisfactory generalization capability over unseen data. It is well known that the LS-SVMs have their computational advantage over conventional SVMs in the model training process; however, the model sparseness is lost, which is the main drawback of LS-SVMs. This is an open problem for the LS-SVMs. To tackle the nonsparseness issue, a new regression alternative to the Lagrangian solution for the LS-SVM is first presented. A novel efficient learning mechanism is then proposed in this paper to extract a sparse set of support vectors for generating fuzzy if-then rules. This novel mechanism works in a stepwise subset selection manner, including a forward expansion phase and a backward exclusion phase in each selection step. The implementation of the algorithm is computationally very efficient due to the introduction of a few key techniques to avoid the matrix inverse operations to accelerate the training process. The computational efficiency is also confirmed by detailed computational complexity analysis. As a result, the proposed approach is not only able to achieve the sparseness of the resultant LS-SVM-based fuzzy systems but significantly reduces the amount of computational effort in model training as well. Three experimental examples are presented to demonstrate the effectiveness and efficiency of the proposed learning mechanism and the sparseness of the obtained LS-SVM-based fuzzy systems, in comparison with other SVM-based learning techniques. Wanqing Zhao |
IEEE Trans. Fuzzy Syst. | 1 |
| 2013 | A Hybrid Learning Method for Constructing Compact Rule-Based Fuzzy ModelsabstractThe Takagi–Sugeno–Kang-type rule-based fuzzy model has found many applications in different fields; a major challenge is, however, to build a compact model with optimized model parameters which leads to satisfactory model performance. To produce a compact model, most existing approaches mainly focus on selecting an appropriate number of fuzzy rules. In contrast, this paper considers not only the selection of fuzzy rules but also the structure of each rule premise and consequent, leading to the development of a novel compact rule-based fuzzy model. Here, each fuzzy rule is associated with two sets of input attributes, in which the first is used for constructing the rule premise and the other is employed in the rule consequent. A new hybrid learning method combining the modified harmony search method with a fast recursive algorithm is hereby proposed to determine the structure and the parameters for the rule premises and consequents. This is a hard mixed-integer nonlinear optimization problem, and the proposed hybrid method solves the problem by employing an embedded framework, leading to a significantly reduced number of model parameters and a small number of fuzzy rules with each being as simple as possible. Results from three examples are presented to demonstrate the compactness (in terms of the number of model parameters and the number of rules) and the performance of the fuzzy models obtained by the proposed hybrid learning method, in comparison with other techniques from the literature. Wanqing Zhao, Qun Niu, Kang Li 0002, George W. Irwin |
IEEE Trans. Cybern. | 1 |
| 2013 | A New Gradient Descent Approach for Local Learning of Fuzzy Neural ModelsabstractThe majority of reported learning methods for Takagi-Sugeno-Kang (TSK) fuzzy neural models to date mainly focus on improvement of their accuracy. However, one of the key design requirements in building an interpretable fuzzy model is that each obtained rule consequent must match well with the system local behavior when all the rules are aggregated to produce the overall system output. This is one of the distinctive characteristics from black-box models such as neural networks. Therefore, how to find a desirable set of fuzzy partitions and, hence, identify the corresponding consequent models which can be directly explained in terms of system behavior, presents a critical step in fuzzy neural modeling. In this paper, a new learning approach considering both nonlinear parameters in the rule premises and linear parameters in the rule consequents is proposed. Unlike the conventional two-stage optimization procedure widely practiced in the field where the two sets of parameters are optimized separately, the consequent parameters are transformed into a dependent set on the premise parameters, thereby enabling the introduction of a new integrated gradient descent learning approach. Thus, a new Jacobian matrix is proposed and efficiently computed to achieve a more accurate approximation of the cost function by using the second-order Levenberg-Marquardt optimization method. Several other interpretability issues regarding the fuzzy neural model are also discussed and integrated into this new learning approach. Numerical examples are presented to illustrate the resultant structure of the fuzzy neural models and the effectiveness of the proposed new algorithm, and compared with the results from some well-known methods. Wanqing Zhao, Kang Li 0002, George W. Irwin |
IEEE Trans. Fuzzy Syst. | 1 |
| 2012 | Improved Structure Optimization for Fuzzy-Neural NetworksabstractFuzzy-neural-network-based inference systems are well-known universal approximators which can produce linguistically interpretable results. Unfortunately, their dimensionality can be extremely high due to an excessive number of inputs and rules, which raises the need for overall structure optimization. In the literature, various input selection methods are available, but they are applied separately from rule selection, often without considering the fuzzy structure. This paper proposes an integrated framework to optimize the number of inputs and the number of rules simultaneously. First, a method is developed to select the most significant rules, along with a refinement stage to remove unnecessary correlations. An improved information criterion is then proposed to find an appropriate number of inputs and rules to include in the model, leading to a balanced tradeoff between interpretability and accuracy. Simulation results confirm the efficacy of the proposed method. Barbara Pizzileo, Kang Li 0002, George W. Irwin, Wanqing Zhao |
IEEE Trans. Fuzzy Syst. | 4 |
| 2010 | An Integrated Method for the Construction of Compact Fuzzy Neural Models
Wanqing Zhao, Kang Li 0002, George W. Irwin, Minrui Fei |
ICIC (1) | 1 |