Weitao Wan

dblp:211/5806 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Identifying Iso-Cost Sweet Spot Configurations for Cloud OLAP Queries
Weitao Wan, Huanchen Zhang, Mingyu Gao 0001
APPT1
2024 Compressed Data Direct Computing for Databases
abstract
Directly performing operations on compressed data has been proven to be a big success facing Big Data problems in modern data management systems. These systems have demonstrated significant compression benefits and performance improvement for data analytics applications. However, current systems only focus on data queries, while a complete Big Data system must support both data query and data manipulation. To solve this problem, we develop CompressDB, which is a new storage engine that can support data processing for databases without decompression. CompressDB has the following advantages. First, CompressDB utilizes context-free grammar to compress data, and supports both data query and data manipulation. Second, for adaptability, we integrate CompressDB to file systems so that a wide range of databases can directly use CompressDB without any change. Third, we enable operation pushdown to storage so that we can perform data query and manipulation in storage systems without bringing large data to memory for high efficiency. We validate the efficacy of CompressDB supporting various kinds of database systems, including SQLite, MySQL, LevelDB, MongoDB, ClickHouse, and Neo4j. We evaluate our method using seven real-world datasets with various lengths, structures, and content in both single node and cluster environments. Experiments show that CompressDB achieves 40% throughput improvement and 44% latency reduction, along with 1.75 compression ratio on average.
Weitao Wan, Feng Zhang 0007, Chenyang Zhang 0005, Mingde Zhang, Jidong Zhai, Yunpeng Chai, Huanchen Zhang, Wei Lu 0015, Yuxing Chen 0003, Haixiang Li, Anqun Pan, Xiaoyong Du 0001
IEEE Trans. Knowl. Data Eng.1
2023 Weak-shot Object Detection through Mutual Knowledge Transfer
abstract
Weak-shot Object Detection methods exploit a fully-annotated source dataset to facilitate the detection performance on the target dataset which only contains image-level labels for novel categories. To bridge the gap between these two datasets, we aim to transfer the object knowledge between the source (S) and target (T) datasets in a bi-directional manner. We propose a novel Knowledge Transfer (KT) loss which simultaneously distills the knowledge of objectness and class entropy from a proposal generator trained on the S dataset to optimize a multiple instance learning module on the T dataset. By jointly optimizing the classification loss and the proposed KT loss, the multiple instance learning module effectively learns to classify object proposals into novel categories in the T dataset with the transferred knowledge from base categories in the S dataset. Noticing the predicted boxes on the T dataset can be regarded as an extension for the original annotations on the S dataset to refine the proposal generator in return, we further propose a novel Consistency Filtering (CF) method to reliably remove inaccurate pseudo labels by evaluating the stability of the multiple instance learning module upon noise injections. Via mutually transferring knowledge between the S and T datasets in an iterative manner, the detection performance on the target dataset is significantly improved. Extensive experiments on public benchmarks validate that the proposed method performs favourably against the state-of-the-art methods without increasing the model parameters or inference computational complexity.
Xuanyi Du, Weitao Wan, Chen Li 0031
CVPR2
2023 Shaping Deep Feature Space Towards Gaussian Mixture for Visual Classification
abstract
The softmax cross-entropy loss function has been widely used to train deep models for various tasks. In this work, we propose a Gaussian mixture (GM) loss function for deep neural networks for visual classification. Unlike the softmax cross-entropy loss, our method explicitly shapes the deep feature space towards a Gaussian Mixture distribution. With a classification margin and a likelihood regularization, the GM loss facilitates both high classification performance and accurate modeling of the feature distribution. The GM loss can be readily used to distinguish the adversarial examples based on the discrepancy between feature distributions of clean and adversarial examples. Furthermore, theoretical analysis shows that a symmetric feature space can be achieved by using the GM loss, which enables the models to perform robustly against adversarial attacks. The proposed model can be implemented easily and efficiently without introducing more trainable parameters. Extensive evaluations demonstrate that the method with the GM loss performs favorably on image classification, face recognition, and detection as well as recognition of adversarial examples generated by various attacks.
Weitao Wan, Jiansheng Chen 0001, Yuanyi Zhong, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Hierarchical Feature Embedding for Visual Tracking
Zhixiong Pi, Weitao Wan, Changxin Gao, Nong Sang, Chen Li 0025
ECCV (22)2
2022 CompressDB: Enabling Efficient Compressed Data Direct Processing for Various Databases
abstract
In modern data management systems, directly performing operations on compressed data has been proven to be a big success facing big data problems. These systems have demonstrated significant compression benefits and performance improvement for data analytics applications. However, current systems only focus on data queries, while a complete big data system must support both data query and data manipulation.
Feng Zhang 0007, Weitao Wan, Chenyang Zhang 0005, Jidong Zhai, Yunpeng Chai, Haixiang Li, Xiaoyong Du 0001
SIGMOD Conference2
2022 Co-attention dictionary network for weakly-supervised semantic segmentation
Weitao Wan, Jiansheng Chen 0001, Ming-Hsuan Yang 0001, Huimin Ma 0001
Neurocomputing1
2022 Frequency Diverse Array Introduced Into SAR GMTI to Mitigate Blind Velocity and Doppler Ambiguity
abstract
In this letter, a frequency diverse array (FDA) is introduced into synthetic aperture radar ground moving target indication (SAR GMTI) to mitigate both the blind velocity and Doppler ambiguity problems. Due to the$2\pi $periodicity of the echo phases, notches will occur periodically in clutter cancelers at nonzero velocities, leading to the blind velocity problem. In the proposed scheme, the dependence of the measured blind velocity on the transmission frequency is considered, and a new clutter canceler unaffected by the blind velocity problem is constructed via the integration of multiple cancelers with diverse frequencies. Moreover, to resolve the Doppler ambiguity, along-track interferometry (ATI) based double interferometry is proposed to extend the maximum unambiguous radial velocity (RV). Additionally, search-based clustering is adopted to enhance the precision of RV estimation. Finally, a moving target can be brought into focus and correctly relocated using the estimated RV. Numerical results verify the effectiveness of the proposed method.
Libing Huang, Xin Li 0127, Weitao Wan, Shunsheng Zhang, Wen-Qin Wang
IEEE Geosci. Remote. Sens. Lett.3
2022 Resolving Doppler Ambiguity of High-Speed Moving Targets via FDA-MIMO Radar
abstract
To address the problem of Doppler ambiguity in low pulse repetition frequency (PRF) radar, we propose the use of frequency diverse array multiple-input–multiple-output (FDA-MIMO) radar to detect high-speed moving targets. This method does not require the radar to transmit multiple-PRF pulses. The possible Doppler ambiguity is resolved by exploiting the multicarrier characteristics of FDA-MIMO radar. The effectiveness of the proposed method is verified by simulation results.
Weitao Wan, Shunsheng Zhang, Wen-Qin Wang
IEEE Geosci. Remote. Sens. Lett.1
2022 Exploring Query Processing on CPU-GPU Integrated Edge Device
abstract
Huge amounts of data have been generated on edge devices every day, which requires efficient data analytics and management. However, due to the limited computing capacity of these edge devices, query processing at the edge faces tremendous pressure. Fortunately, in recent years, hardware vendors have integrated heterogeneous coprocessors, such as GPUs, into the edge device, which can provide much more computing power. Furthermore, the CPU-GPU integrated edge device has shown significant benefits in a variety of situations. Therefore, the exploration of query processing on such CPU-GPU integrated edge devices becomes an urgent need. In this article, we develop a fine-grained query processing engine, called FineQuery, which can perform efficient query processing on CPU-GPU integrated edge devices. Particularly, FineQuery can take advantage of both architectural features of edge devices and query characteristics by performing fine-grained workload scheduling between the CPU and the GPU. Experiments show that on TPC-H workloads, FineQuery reduces 42.81% latency and improves 2.39× bandwidth utilization on average compared to the implementation of using only GPU or CPU. Furthermore, query processing at the edge can bring significant performance-per-cost benefits and energy efficiency. On average, FineQuery at the edge brings 21× performance-per-cost ratio and 4× energy efficiency compared with processing the data on a discrete GPU platform.
Jiesong Liu, Feng Zhang 0007, Hourun Li, Dalin Wang, Weitao Wan, Xiaokun Fang, Jidong Zhai, Xiaoyong Du 0001
IEEE Trans. Parallel Distributed Syst.5
2021 FineQuery: Fine-Grained Query Processing on CPU-GPU Integrated Architectures
abstract
Using heterogeneous coprocessors, such as GPUs, to accelerate complicated SQL queries has been proved to be effective in the database domain. Previous works show that taking advantage of the high parallelism and computing capacity of heterogeneous coprocessors can bring significant performance improvements. However, in the discrete memory architecture, the advantages of heterogeneous coprocessors will be weakened due to the low PCI-e bandwidth and high latency. Fortunately, hardware vendors have proposed a novel integration architecture design, which integrates CPU and GPU on the same chip. This integrated architecture allows the GPU and CPU to share the same unified memory, taking new opportunities for fine-grained collaboration between the GPU and CPU to optimize SQL queries. In this paper, we propose a query processing engine, called FineQuery, to optimize the execution of SQL queries on the integrated architecture. FineQuery can take advantage of both architectural features and query characteristics by performing fine-grained workload scheduling between the CPU and the GPU. Experimental results show that 1) on the integration architecture, FineQuery can reduce the latency by 25.30% and increase the bandwidth utilization by 39.46% on average. 2) FineQuery on the integrated architecture achieves $13.74\times$ the performance-per-cost ratio and $6.87\times$ energy efficiency over query processing on the discrete GPU platform.
Dalin Wang, Feng Zhang 0007, Weitao Wan, Hourun Li, Xiaoyong Du 0001
CLUSTER3
2021 Defending against Universal Adversarial Patches by Clipping Feature Norms
abstract
Physical-world adversarial attacks based on universal adversarial patches have been proved to be able to mislead deep convolutional neural networks (CNNs), exposing the vulnerability of real-world visual classification systems based on CNNs. In this paper, we empirically reveal and mathematically explain that the universal adversarial patches usually lead to deep feature vectors with very large norms in popular CNNs. Inspired by this, we propose a simple yet effective defending approach using a new feature norm clipping (FNC) layer which is a differentiable module that can be flexibly inserted in different CNNs to adaptively suppress the generation of large norm deep feature vectors. FNC introduces no trainable parameter and only very low computational overhead. However, experiments on multiple datasets validate that it can effectively improve the robustness of different CNNs towards white-box universal patch attacks while maintaining a satisfactory recognition accuracy for clean samples.
Youze Xue, Weitao Wan, Jiayu Bao, Huimin Ma 0001
ICCV5
2020 Adversarial Training with Bi-directional Likelihood Regularization for Visual Classification
Weitao Wan, Jiansheng Chen 0001, Ming-Hsuan Yang 0001
ECCV (24)1
2020 Image Captioning With End-to-End Attribute Detection and Subsequent Attributes Prediction
abstract
Semantic attention has been shown to be effective in improving the performance of image captioning. The core of semantic attention based methods is to drive the model to attend to semantically important words, or attributes. In previous works, the attribute detector and the captioning network are usually independent, leading to the insufficient usage of the semantic information. Also, all the detected attributes, no matter whether they are appropriate for the linguistic context at the current step, are attended to through the whole caption generation process. This may sometimes disrupt the captioning model to attend to incorrect visual concepts. To solve these problems, we introduce two end-to-end trainable modules to closely couple attribute detection with image captioning as well as prompt the effective uses of attributes by predicting appropriate attributes at each time step. The multimodal attribute detector (MAD) module improves the attribute detection accuracy by using not only the image features but also the word embedding of attributes already existing in most captioning models. MAD models the similarity between the semantics of attributes and the image object features to facilitate accurate detection. The subsequent attribute predictor (SAP) module dynamically predicts a concise attribute subset at each time step to mitigate the diversity of image attributes. Compared to previous attribute based methods, our approach enhances the explainability in how the attributes affect the generated words and achieves a state-of-the-art single model performance of 128.8 CIDEr-D on the MSCOCO dataset. Extensive experiments on the MSCOCO dataset show that our proposal actually improves the performances in both image captioning and attribute detection simultaneously. The codes are available at: https://github.com/ RubickH/Image-Captioning-with-MAD-and-SAP.
Jiansheng Chen 0001, Wanli Ouyang, Weitao Wan, Youze Xue
IEEE Trans. Image Process.4
2019 Information Entropy Based Feature Pooling for Convolutional Neural Networks
abstract
In convolutional neural networks (CNNs), we propose to estimate the importance of a feature vector at a spatial location in the feature maps by the network's uncertainty on its class prediction, which can be quantified using the information entropy. Based on this idea, we propose the entropy-based feature weighting method for semantics-aware feature pooling which can be readily integrated into various CNN architectures for both training and inference. We demonstrate that such a location-adaptive feature weighting mechanism helps the network to concentrate on semantically important image regions, leading to improvements in the large-scale classification and weakly-supervised semantic segmentation tasks. Furthermore, the generated feature weights can be utilized in visual tasks such as weakly-supervised object localization. We conduct extensive experiments on different datasets and CNN architectures, outperforming recently proposed pooling methods and attention mechanisms in ImageNet classification as well as achieving state-of-the-arts in weakly-supervised semantic segmentation on PASCAL VOC 2012 dataset.
Weitao Wan, Tianpeng Li, Jingqi Tian, Youze Xue
ICCV1
2019 MVSCRF: Learning Multi-View Stereo With Conditional Random Fields
abstract
We present a deep-learning architecture for multi-view stereo with conditional random fields (MVSCRF). Given an arbitrary number of input images, we first use a U-shape neural network to extract deep features incorporating both global and local information, and then build a 3D cost volume for the reference camera. Unlike previous learning based methods, we explicitly constraint the smoothness of depth maps by using conditional random fields (CRFs) after the stage of cost volume regularization. The CRFs module is implemented as recurrent neural networks so that the whole pipeline can be trained end-to-end. Our results show that the proposed pipeline outperforms previous state-of-the-arts on large-scale DTU dataset. We also achieve comparable results with state-of-the-art learning based methods on outdoor Tanks and Temples dataset without fine-tuning, which demonstrates our method's generalization ability.
Youze Xue, Weitao Wan, Tianpeng Li, Jiayu Bao
ICCV3
2019 Image Captioning with Attribute Refinement
abstract
Semantic attention has long been adopted to image captioning models to enhance the image captioning performances. The models pre-trained for attribute recognition are utilized to generate image attributes in image captioning. Generally, these models are not jointly trained with image captioning models. In this paper, we propose attribute refinement network, which incorporates attribute recognition with image captioning to boost the performance on both tasks. We model the correlation between attributes with the semantic information from image captioning to improve the recognition accuracy. In turn, better attribute recognition results effectively enhance image captioning performance. Our model achieves CIDEr-D/SPICE scores of 115.1 and 20.9 respectively on the MS COCO test set, comprehensively yields improvement over all compared methods.
Tianpeng Li, Weitao Wan
ICIP4
2019 Improving Human Parsing by Extracting Global Information Using the Non-Local Operation
abstract
Human parsing has recently attracted considerable interests due to its wide application potentials. However, developing an accurate human parsing system is still a challenge for researchers. In this paper, we demonstrate that global information are critical for accurate prediction by applying a non-local operation for effectively extracting global information. Meanwhile several training data refinement methodologies are proposed to further boost the performance. Benefiting from all the approaches, the proposed single human parsing model NLGINet achieves the state-of-the-art segmentation accuracy on two human parsing benchmark datasets LIP and Pascal-Person-Parts.
Tianpeng Li, Weitao Wan
ICIP2
2018 Rethinking Feature Distribution for Loss Functions in Image Classification
abstract
We propose a large-margin Gaussian Mixture (L-GM) loss for deep neural networks in classification tasks. Different from the softmax cross-entropy loss, our proposal is established on the assumption that the deep features of the training set follow a Gaussian Mixture distribution. By involving a classification margin and a likelihood regularization, the L-GM loss facilitates both a high classification performance and an accurate modeling of the training feature distribution. As such, the L-GM loss is superior to the softmax loss and its major variants in the sense that besides classification, it can be readily used to distinguish abnormal inputs, such as the adversarial examples, based on their features' likelihood to the training feature distribution. Extensive experiments on various recognition benchmarks like MNIST, CIFAR, ImageNet and LFW, as well as on adversarial examples demonstrate the effectiveness of our proposal.
Weitao Wan, Yuanyi Zhong, Tianpeng Li
CVPR1
2017 Combining Object-Based Attention and Attributes for Image Captioning
Weitao Wan, Tianpeng Li
ICIG (1)3
2017 Occlusion robust face recognition based on mask learning
abstract
Face occlusion has been a long standing challenging issue in face recognition. In the state-of-the-art deep Convolutional Neural Network (CNN) face recognition models, occluded facial parts are generally embedded into the learned features together with the non-occluded parts in an equivalent manner. As such, the discriminative power of the generated face representation may be weakened for occluded face images. To address this problem, we propose the MaskNet, a trainable module which can be included in existing CNN architectures. With end-to-end training supervised by only the personal identity labels, MaskNet learns a proper way of adaptively generating different feature map masks for different occluded face images. Intuitively, MaskNet automatically assigns higher weights to the hidden units activated by the non-occluded facial parts and lower weights to those that are activated by the occluded facial parts. Experiments on datasets consisting of real-life and synthetic occluded faces demonstrate that MaskNet can effectively improve the robustness of CNN models towards occlusions in face recognition.
Weitao Wan
ICIP1