VLDB 2026 Research / reviewers in the wild / expert
Wei Hu 0004
dblp:52/173-4
· DBLP profile ↗
33ranked-venue papers
7as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MIDI-Zero: A MIDI-driven Self-Supervised Learning Approach for Music RetrievalabstractContent-based Music Retrieval (CBMR) is a fundamental task in music information retrieval, encompassing sub-tasks including Audio Identification, Audio Matching, and Version Identification. Traditional methods typically analyze audio signals or spectrograms to extract features related to rhythm, melody, harmony, and timbre. However, with the rapid development of Music Transcription and digital music technologies, MIDI representation has emerged as a powerful alternative fo r music analysis. In this paper, we propose MIDI-Zero, a novel self-supervisedlearning framework for CBMR that operates entirely on MIDI representations. Unlike existing approaches, MIDI-Zero requires no external training data; all training data is automatically generated based on predefined task rules, eliminating the need for labeled datasets or external music collections. MIDI-Zero is designed to handle both symbolic music data and audio-based tasks by leveraging Music Transcription models. Its strong robustness ensures effectiveness even with low-quality transcriptions. Extensive experiments demonstrate that MIDI-Zero achieves competitive performance across various CBMR sub-tasks, particularly excelling in Audio Matching. Our approach simplifies the feature extraction process, bridges the gap between audio and symbolic music representations, and offers a versatile and scalable solution for music retrieval. Wei Hu 0004, Hongfeng Gao, Fan Zhang 0007 |
SIGIR | 2 |
| 2024 | LTGC: Long-Tail Recognition via Leveraging LLMs-Driven Generated ContentabstractLong-tail recognition is challenging because it requires the model to learn good representations from tail categories and address imbalances across all categories. In this paper, we propose a novel generative and fine-tuning framework, LTGC, to handle long-tail recognition via leveraging generated content. Firstly, inspired by the rich implicit knowledge in large-scale models (e.g., large language models, LLMs), LTGC leverages the power of these models to parse and reason over the original tail data to produce diverse tail-class content. We then propose several novel designs for LTGC to ensure the quality of the generated data and to efficiently fine-tune the model using both the generated and original data. The visualization demonstrates the effectiveness of the generation module in LTGC, which produces accurate and diverse tail data. Additionally, the experimental results demonstrate that our LTGC outperforms existing state-of-the-art methods on popular long-tailed benchmarks. Qihao Zhao, Yalun Dai, Hao Li 0075, Wei Hu 0004, Fan Zhang 0007, Jun Liu 0036 |
CVPR | 4 |
| 2024 | LTRL: Boosting Long-Tail Recognition via Reflective Learning
Qihao Zhao, Yalun Dai, Shen Lin 0006, Wei Hu 0004, Fan Zhang 0007, Jun Liu 0036 |
ECCV (67) | 4 |
| 2024 | AMG-Embedding: A Self-Supervised Embedding Approach for Audio IdentificationabstractAudio Identification aims to precisely retrieve exact matches from a vast music repository through a query audio snippet. The need for specificity and granularity has traditionally led to representing music audio using numerous short fixed-duration overlapped segment/shingle features in fingerprinting approaches. However, fingerprinting imposes constraints on scalability and efficiency, as hundreds or even thousands of embeddings are generated to represent a music audio. In this paper, we present an innovative self-supervised approach called Angular Margin Guided Embedding (AMG-Embedding). AMG-Embedding is built on a traditional fingerprinting encoder and aims to represent variable-duration non-overlapped segments as embeddings through a two-stage embedding and class-level learning process. AMG-Embedding significantly reduces the number of generated embeddings while achieving high-specific fragment-level audio identification simultaneously. Experimental results demonstrate that AMG-Embedding achieves retrieval accuracy comparable to the based fingerprinting approach while consuming less than 1/10th of its storage and retrieval time. The efficiency gains of our approach position it as a promising solution for scalable and efficient audio identification systems. Wei Hu 0004, Fan Zhang 0007 |
ACM Multimedia | 2 |
| 2024 | Virtual Agents in Immersive Virtual Reality Environments: Impact of Humanoid Avatars and Output Modalities on Shopping ExperienceabstractEmbodied virtual agents (EVAs) are used in various online shopping scenarios to enhance consumer experience. Humanoid avatars and output modalities of EVAs are two important factors affecting online shopping experience, which have been widely studied in shopping websites before. Moreover, immersive virtual reality (IVR) shopping is a future trend in online shopping, but research on EVAs in IVR environments is limited. The impact of them in an IVR shopping environment is not yet fully understood. Therefore, this study aims to investigate the effect of an EVA's humanoid avatar and output modalities on the online shopping experience in an IVR environment. First, 66 participants were invited to participate in a 2 (humanoid avatar: presence vs. absence) × 3 (output modalities: text-only, voice-only, and text and voice) within-group experiment to complete the specified purchase task. After each purchase task was completed, participants were required to fill in a subjective questionnaire to score the six different combinations of virtual agents (VAs). After all purchase tasks were completed, a follow-up interview was conducted to better obtain the participants’ preferences and detailed reasons for the six different combinations of VAs. The results showed that humanoid avatars significantly enhanced participants’ perceptions of warmth, communication, trust, and satisfaction. Additionally, output modalities significantly enhanced participants’ perception of warmth, communication, trust, comfort, and satisfaction. Finally, the humanoid avatar and output modalities had significant interactions in terms of comfort. The results of this research provide theoretical reference and guidance for the design of VAs in IVR shopping environment, and also provide inspiration for the design and research of VAs in other online shopping environments. Shangshang Zhu, Wei Hu 0004, Yenan Dong |
Int. J. Hum. Comput. Interact. | 2 |
| 2024 | OHD: An Online Category-Aware Framework for Learning With Noisy Labels Under Long-Tailed DistributionabstractRecently, many effective methods have emerged to address the robustness problem of Deep Neural Networks (DNNs) trained with noisy labels. However, existing work on learning with noisy labels (LNL) mainly focuses on balanced datasets, while real-world scenarios usually also exhibit a long-tailed distribution (LTD). In this paper, we propose an online category-aware approach to mitigate the impact of noisy labels and LTD on the robustness of DNNs. First, the category frequency of clean samples used to rebalance the feature space cannot be obtained directly in the presence of noisy samples. We design a novel category-aware Online Joint Distribution to dynamically estimate the category frequency of clean samples. Second, previous LNL methods were category-agnostic. These methods would easily be confused with noisy samples and tail categories’ samples under LTD. Based on this observation, we propose a Harmonizing Factor strategy to exploit more information from the category-aware online joint distribution. This strategy provides more accurate estimates of clean samples between noisy samples and samples with tail categories. Finally, we propose Dynamic Cost-sensitive Learning, which utilizes the loss and category frequency of the estimated clean samples to address both LNL and LTD. Compared to extensive state-of-the-art methods, our strategy consistently improves the generalization performance of DNNs on several synthetic datasets and two real-world datasets. Qihao Zhao, Fan Zhang 0007, Wei Hu 0004, Songhe Feng, Jun Liu 0036 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | SPGC: Shape-Prior-Based Generated Content Data Augmentation for Remote Sensing Object DetectionabstractWhile deep learning-based methods have made significant strides in remote sensing applications, the scarcity and inadequate quality of remote sensing images tend to curtail the improvement of follow-up research such as remote sensing object detection. However, the human visual system is able to quickly grasp the features of an unseen object given only a few examples, which is considered to be related to a strong shape bias. Inspired by how human toddlers learn shapes and the process of recognizing objects by shape, this paper proposes a novel method known as Shape-Prior based Generated Content (SPGC) data augmentation to overcome these challenges. Specifically, our method includes two main steps: shape data generation and stylization. Initially, the method begins with generating shape data regardless of training data availability. Next, we enhance the robustness of the generated shape data through stylization, forming a robust shape dataset. Stylization is further bifurcated into two scenarios: when training data is unseen, self-stylization is employed where the shape data simultaneously serves as content and style data, resulting in significant performance improvements. When training data is accessible, data-specific stylization is applied, with the shape data as content and training data as style, leading to more substantial enhancements than self-stylization. Experimental results on mainstream remote sensing object detection datasets including NWPU VHR-10, DIOR, and FAIR1M demonstrate that our method significantly improves performance and underscores its effectiveness. Yalun Dai, Fei Ma 0001, Wei Hu 0004, Fan Zhang 0007 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | MDCS: More Diverse Experts with Consistency Self-distillation for Long-tailed RecognitionabstractRecently, multi-expert methods have led to significant improvements in long-tail recognition (LTR). We summarize two aspects that need further enhancement to contribute to LTR boosting: (1) More diverse experts: (2) Lower model variance. However, the previous methods didn’t handle them well. To this end, we propose More Diverse experts with Consistency Self-distillation (MDCS) to bridge the gap left by earlier methods. Our MDCS approach consists of two core components: Diversity Loss (DL) and Consistency Self-distillation (CS). In detail, DL promotes diversity among experts by controlling their focus on different categories. To reduce the model variance, we employ KL divergence to distill the richer knowledge of weakly augmented instances for the experts’ self-distillation. In particular, we design Confident Instance Sampling (CIS) to select the correctly classified instances for CS to avoid biased/noisy knowledge. In the analysis and ablation study, we demonstrate that our method compared with previous work can effectively increase the diversity of experts, significantly reduce the variance of the model, and improve recognition accuracy. Moreover, the roles of our DL and CS are mutually reinforcing and coupled: the diversity of experts benefits from the CS, and the CS cannot achieve remarkable results without the DL. Experiments show our MDCS outperforms the state-of-the-art by 1% ~ 2% on five popular long-tailed benchmarks, including CIFAR10-LT, CIFAR100-LT, ImageNet-LT, Places-LT, and iNaturalist 2018. The code is available at https://github.com/fistyee/MDCS Qihao Zhao, Wei Hu 0004, Fan Zhang 0007, Jun Liu 0036 |
ICCV | 3 |
| 2023 | MixPro: Data Augmentation with MaskMix and Progressive Attention Labeling for Vision Transformer
Qihao Zhao, Yangyu Huang, Wei Hu 0004, Fan Zhang 0007, Jun Liu 0036 |
ICLR | 3 |
| 2023 | Pixel-Level Annotation of Specific Targets for Large-Scale Remote Sensing ImagesabstractCommon remote sensing images usually contain targets such as airports, stations, stadiums, mountains, lakes and so on. The localization of these targets is of great research importance. However, a large-scale annotation of these images is difficult to be obtained, because these labels can only be annotated manually, and the annotation is undoubtedly very time-consuming and laborious. In the paper, we propose a framework to generate localization annotations via semi-supervised training. Firstly, optical remote sensing images, with or without target objects, are automatically collected, and a binary classification model is trained on these images. Then the feature map, as a preliminary candidate region for targets, can be obtained from the model. Finally, a saliency detection method based on transformer is employed to further refine the candidate regions and get more precise localization annotations. The proposed framework is a semi-supervised approach to automatically generate large-scale precise location annotation datasets, which can further help many related researches. Guolong Liu, Wei Hu 0004, Fan Zhang 0007 |
IGARSS | 2 |
| 2023 | Cloud Detection Network Based on Scenario Synthesis and Transformer in Remote Sensing ImagesabstractCloud detection plays an essential part in optical remote sensing data processing. Cloud detection technology has made significant progress, with the rapid development of deep learning technology in image processing. Then, Since Transformer has excelled in the image semantic segmentation task, researchers have begun experimenting with introducing Transformer into cloud detection tasks to tackle the difficulties in acquiring information on the ground due to the cloud cover. Compared to other detection problems, cloud detection is difficult to obtain large amounts of annotated data for training models because of poor interpretability and fewer applications. In order to solve these problems, in this paper, we propose a cloud detection method based on scene synthesis and Transformer. To address the difficulty of thin cloud scenarios with small amounts of data, we segmented the cloud target by multi-channel merged region growing algorithm after Gabor filtering, and then randomly embeded the cloud scene into reasonable regions to synthesize the cloud detection dataset. In addition, we designed a Trans-CloudNet with combination loss to improve the class imbalance problem in the training data by increasing the inter-class weight coefficients while comparing the image similarity. Experiments demonstrate that the method we proposed has excellent performance on thin cloud detection. Yipeng Ru, Fan Zhang 0007, Wei Hu 0004 |
IGARSS | 3 |
| 2023 | Crop Classification of Multitemporal PolSAR Based on 3-D Attention Module With ViTabstractMulti-temporal polarimertic SAR is considered to be very effective in crop classification and cultivated land detection, which has received much attention from researchers. Currently, for most multi-temporal polarimetric SAR data classification methods, the simultaneous temporal-polarimetric-spatial feature extraction capability has not been exploited sufficiently. Also, the diversity of different time and different polarimetric features has not been taken into account sufficiently. In this paper, we propose a classification model that combines a dual-stream network as a temporal-polarimetric-spatial feature extraction module with Vision Transformer(ViT) called Temporal-Polarimetric-Spatial Transformer(TSPT) to address the above problems. Secondly, a 3 dimension(3D) convolutional attention module that enables the network to weight the temporal dimension, polarimetric feature dimension and spatial dimension is developed, according to their importance. Experimental results on both UAVSAR and RADARSAT-2 datasets show that the proposed method outperforms ResNet. Qiang Yin 0001, Wei Hu 0004, Carlos López-Martínez, Fan Zhang 0007 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | SGT: A Generalized Processing Model for 1-D Remote Sensing Signal ClassificationabstractThis paper proposes a generalized feature extraction framework for one-dimensional(1D) remote sensing data. This approach streamlines the processing for extracting features by eliminating the need for some preprocessing, such as data normalization, data filtering, and spectrogram generation, which explicitly encode domain-specific knowledge of the tasks. The main component of the new framework, called Shifted-Grad Transformer(SGT), includes the Shift module, Grad module, Smooth module, Raw embedding module, Transformer encoder module, and additional essential module. Extensive experiments on data sets such as Hyperspectral image data, Magnetic signal data, and other 1D data have demonstrated that the SGT performs significantly better than existing methods and provides a new solution to the 1D data processing problem. Our Training code and data are available at https://github.com/wfnian/SGT. Wei Hu 0004, Fangnian Wang, Qiang Yin 0001, Fan Zhang 0007 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | SeqFace: Learning discriminative features by using face sequencesabstractAbstract Deep convolutional neural networks (CNNs) have greatly improved the Face Recognition (FR) performance in recent years. Almost all CNNs in FR are trained on the carefully labeled datasets containing plenty of identities. However, such high‐quality datasets are very expensive to collect, which restricts many researchers to achieve state‐of‐the‐art performance. In this paper, a framework, called SeqFace, for learning discriminative face features is proposed. Besides a traditional identity training dataset, the designed SeqFace can train CNNs by using an additional dataset which includes a large number of face sequences collected from videos. Moreover, the label smoothing regularization (LSR) and a new proposed discriminative sequence agent (DSA) loss are employed to enhance the discrimination power of deep face features via making full use of the sequence data. Only with a single ResNet model, the method achieves very competitive performance on several face recognition benchmarks, including LFW, YTF, CFP, AgeDB, and MegaFace. The code and model are publicly available at the website https://github.com/huangyangyu/SeqFace . Wei Hu 0004, Yangyu Huang, Fan Zhang 0007, Ruirui Li 0001, Heng-Chao Li 0001 |
IET Image Process. | 1 |
| 2021 | P-DIFF+: Improving learning classifier with noisy labels by Noisy Negative Learning loss
Qihao Zhao, Wei Hu 0004, Yangyu Huang, Fan Zhang 0007 |
Neural Networks | 2 |
| 2020 | P-DIFF: Learning Classifier with Noisy Labels based on Probability Difference DistributionsabstractLearning deep neural network (DNN) classifier with noisy labels is a challenging task because the DNN can easily overfit on these noisy labels due to its high capability. In this paper, we present a very simple but effective training paradigm called P-DIFF, which can train DNN classifiers but obviously alleviate the adverse impact of noisy labels. Our proposed probability difference distribution implicitly reflects the probability of a training sample to be clean, then this probability is employed to re-weight the corresponding sample during the training process. P-DIFF can also achieve good performance even without prior-knowledge on the noise rate of training samples. Experiments on benchmark datasets also demonstrate that P-DIFF is superior to the state-of-the-art sample selection methods. Wei Hu 0004, Qihao Zhao, Yangyu Huang, Fan Zhang 0007 |
ICPR | 1 |
| 2020 | SAR Target Small Sample Recognition Based on CNN Cascaded Features and AdaBoost Rotation ForestabstractAutomatic target recognition (ATR) has made great progress with the development of deep learning. However, the target feature in synthetic aperture radar (SAR) image is not consistent with human vision, and the SAR training samples are always limited. These hard issues pose new challenges to the SAR ATR based on convolutional neural network (CNN). In this letter, we propose an improved CNN model to solve the limited sample issue via the feature augmentation and ensemble learning strategies. Normally, the high-level features that are more comprehensive and discriminative than the middle-level and low-level features are always employed for category discrimination. In order to make up the insufficient training features in the limited sample case, the cascaded features from optimally selected convolutional layers are concatenated to provide more comprehensive representation for the recognition. To take full advantage of these cascaded features, the ensemble learning-based classifier, namely, the AdaBoost rotation forest (RoF), is introduced to replace the original softmax layer to realize a more accurate limited sample recognition. Through the AdaBoost RoF method, not only are these features further enhanced by the rotation matrix but also a strong classifier is constructed by several weak classifiers with different adjusted weights. The experimental results on MSTAR data set show that the cascaded features and ensemble weak classifiers can fully exploit effective information in limited samples. Compared with the existing CNN method, the proposed method can improve the recognition accuracy by about 20% under the condition of ten training samples per class. Fan Zhang 0007, Yunchong Wang 0002, Yongsheng Zhou, Wei Hu 0004 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2019 | Noise-Tolerant Paradigm for Training Face Recognition CNNsabstractBenefit from large-scale training datasets, deep Convolutional Neural Networks(CNNs) have achieved impressive results in face recognition(FR). However, tremendous scale of datasets inevitably lead to noisy data, which obviously reduce the performance of the trained CNN models. Kicking out wrong labels from large-scale FR datasets is still very expensive, although some cleaning approaches are proposed. According to the analysis of the whole process of training CNN models supervised by angular margin based loss(AM-Loss) functions, we find that the distribution of training samples implicitly reflects their probability of being clean. Thus, we propose a novel training paradigm that employs the idea of weighting samples based on the above probability. Without any prior knowledge of noise, we can train high performance CNN models with largescale FR datasets. Experiments demonstrate the effectiveness of our training paradigm. The codes are available at https://github.com/huangyangyu/NoiseFace. Wei Hu 0004, Yangyu Huang, Fan Zhang 0007, Ruirui Li 0001 |
CVPR | 1 |
| 2019 | High Resolution SAR Image Synthesis with Hierarchical Generative Adversarial NetworksabstractGenerative adversarial network (GAN) is an artificial neural network based on unsupervised learning method. Due to its powerful model representation capabilities, GAN has been introduced to synthesize synthetic aperture radar (SAR) image data, for the real sample is difficult to acquire. Large-scale, high-resolution SAR images play an important role in promoting SAR applications, such as automatic target recognition and image interpretation. However, on account of the difficult training problem of GAN network, especially for SAR images with speckle noise, it is difficult to obtain high-resolution SAR images by simply transfer the net from optical image. Recent studies in other image fields have shown that hierarchical structure is an effective and useful way to decompose a generation task into several smaller subtasks. How to obtain more high-resolution SAR images from limited original samples through GAN is the target of our research. Therefore, in this paper, we introduce a hierarchical GAN network model to generate SAR images, through the multi-stage network, gradually improve the quality of the generated image, and finally obtain high-resolution images. The type and aspect of generated images are determined by the input of condition vectors in the last two stages. In addition, we introduce the triple loss, in which the background loss is used to imitating background clutter noise of SAR image, the condition loss is to make the generated images' type and aspect become controllable, and the global loss for getting higher image generation quality. The generated images show high similarity with the real samples. Henghua Huang, Fan Zhang 0007, Yongsheng Zhou, Qiang Yin 0001, Wei Hu 0004 |
IGARSS | 5 |
| 2019 | A Fast Inference Networks for SAR Target Few-Shot Learning Based on Improved Siamese NetworksabstractIn this paper, we improve the Siamese Networks for SAR target few-shot learning. SAR target recognition is an important branch of SAR application. It can efficiently extract target category information from complex SAR images and help humans quickly understand SAR images. However, many successful machine learning methods require large amounts of annotated data. So, few-shot learning is always a topical challenge for machine learning. We apply Siamese Networks to SAR target recognition with limited data and improved it. Our model consists of CNN encoder, similarity discriminator and classifier. Relevantly, it has two inputs and three outputs. CNN encoder is constrained by similarity discriminator and classifier. Furthermore, the larger difference from the Siamese Network is that the target category is outputted by the classifier, not by the similarity discriminator. Our method not only makes use of the advantage of metric learning to improve the accuracy of SAR target recognition with limited data, but also significantly reduces the prediction time consumption for the model based on metric learning. In the ten categories military vehicle classification task, there are only five samples for each category and a total of 2425 testing samples. Our method outperforms A-ConvNet and Siamese Networks by 15.8% and 8.41%. The prediction time consumption of Siamese Networks is 114.832s, while that of our method is 1.172s. Jiaxin Tang, Fan Zhang 0007, Yongsheng Zhou, Qiang Yin 0001, Wei Hu 0004 |
IGARSS | 5 |
| 2018 | Small Sample Learning Optimization for Resnet Based Sar Target RecognitionabstractDeep convolutional neural network (CNN) is an important branch of deep learning. Due to its strong ability of feature extraction, CNN models have been introduced to solve the problems of synthetic aperture radar automatic target recognition (SAR-ATR). However, labeled SAR images are difficult to acquire. Therefore, how to obtain a good recognition result from a small sample dataset is what we mainly focus on. In theory, a deeper network can bring a better training result. But it also brings more difficulties to the training process, especially with limited labeled training data. The residual learning which proposed in recent years can alleviate this problem effectively. In this paper, we use a deep residual network, and introduce the dropout layer into the building block to alleviate overfitting caused by limited SAR data. In order to improve the training effect, the new loss function center loss is adopted and combined with softmax loss as the supervision signal to train the deep CNN. The experimental results show that our method can achieve the classification accuracy of 99.67% with all training data, without data augmentation or pre-training. When data of the training dataset was reduced to 20%, we can still achieve a recognition result higher than 94%. Zhenzhen Fu, Fan Zhang 0007, Qiang Yin 0001, Ruirui Li 0001, Wei Hu 0004, Wei Li 0032 |
IGARSS | 5 |
| 2018 | Shipnet for Semantic Segmentation on VHR Maritime ImageryabstractFor VHR maritime images, sematic segmentation is a new research hotspot and plays an important role in coastline navigation, resource management and territory protection. Without enough labeled training data, it is a challenge to separate small objects on a large scale while segment the big area clearly. To deal with it, we propose a novel ShipNet and design a weighted loss function for simultaneous sea-land segmentation and ship detection. To prove the proposed method, we also built and opened a new dataset to the community which contains VHR multiscale maritime images. Compared with the FCN and ResNet, the proposed method got much better F1 scores 85.90% for ship class and 97.54% overall accuracy. Compared with multiscale FCN, the ShipNet could obtain details results like sharp edges. Even for images with bad quality, the ShipNet could also keep robust and get good results. Shihao Sun, Wei Hu 0004, Ruirui Li 0001 |
IGARSS | 4 |
| 2018 | Attention Enhanced ConvNet-RNN for Chinese Vehicle License Plate Recognition
Shiming Duan, Wei Hu 0004, Ruirui Li 0001, Wei Li 0032, Shihao Sun |
PRCV (2) | 2 |
| 2016 | A spaceborne SAR on-board processing simulator using mobile GPUabstractThis paper presents a new simulator for spaceborne SAR on-board imaging process on mobile GPUs. The system can generate raw data and perform imaging process in real time. Due to introducing low power GPUs, it has low power consumption, light weight and high computing capability. This simulator has the unique structure, which can guarantee its real-time processing. The experimented results indicate that it is possible to apply the simulator to simulating spaceborne SAR imaging process, and this portable simulator can be used in other areas, which could broaden new horizons in low power space. Hanyuan Tang, Fan Zhang 0007, Wei Hu 0004, Wei Li 0032 |
IGARSS | 4 |
| 2016 | Atomic-free optimization on GPU based SAR raw data simulationabstractSynthetic Aperture Radar (SAR) has been widely used in airborne remote sensing and satellite ocean observation fields to reduce the affect of weather condition and sun illumination. As technology developed, swath and resolution requirements are increased in terrain, which result in a huge increase in echo data and simulated time[1]. With the development of graphics processing unit (GPU), it can reduce simulated time effectively. In order to simulate the coherent integration, atomic operation is always used in GPU, which has a bad influence to simulated time. To optimize simulated time, in this article, we put forward three GPU optimistic strategies for atomic-free SAR raw data simulation. Xiaojie Yao, Fan Zhang 0007, Wei Hu 0004, Wei Li 0032 |
IGARSS | 4 |
| 2015 | Efficient SAR raw data parallel simulation based on multicore vector extensionabstractDue to independence of weather condition and sun illumination, Synthetic Aperture Radar (SAR) has been widely used in airborne remote sensing and satellite ocean observation fields. With the increase in swath and higher resolution requirements, SAR imaging algorithms require further study. However, as a research support, echo data has massive increase. So the high efficient echo simulation is required emergently. SAR echo simulation is optimized to accelerate, based on AVX/SSE of vector instruction set and OpenMP. Experiments demonstrate that in the case of using OpenMP and AVX / SSE combined, CPU simulation efficiency is improved by about 23 times, compared to the traditional one. Fan Zhang 0007, Lixiang Ma, Wei Hu 0004, Wei Li 0032 |
IGARSS | 5 |
| 2015 | Accelerating SAR imaging using vector extension on multi-core SIMD CPUabstractWith the development of synthetic aperture radar (SAR) technology in recent years, we have to face the huge amount of data. So, the fast image processing technology seems to be really important in the domain of SAR. Based on the situation above, a new method to accelerate SAR image processing is proposed in this paper. The proposed method employs SIMD (Single Instruction Multiple Data) instructions and OpenMP (Open Multiprocessing) technology on multi-core SIMD CPU to realize parallel optimization on image processing algorithms. The CS (Chirp Scaling) algorithm is also chosen to process SAR data in our experiment. The experimental results demonstrate that the proposed SIMD based implementation is able to increase performance more than 20 times when compared to the single-core baseline used. Fan Zhang 0007, Lixiang Ma, Wei Hu 0004, Wei Li 0032 |
IGARSS | 4 |
| 2015 | Collaborative-Representation-Based Nearest Neighbor Classifier for Hyperspectral ImageryabstractNovel collaborative representation (CR)-based nearest neighbor (NN) algorithms are proposed for hyperspectral image classification. The proposed methods are based on a CR computed by an ℓ2-norm minimization with a Tikhonov regularization matrix. More specific, a testing sample is represented as a linear combination of all the training samples, and the weights for representation are estimated by an ℓ2-norm minimization-derived closed-form solution. In the first strategy, the label of a testing sample is determined by majority voting of those with k largest representation weights. In the second strategy, local within-class CR is considered as an alternative, and the testing sample is assigned to the class producing the minimum representation residual. The experimental results show that the proposed algorithms achieve better performance than several previous algorithms, such as the original k-NN classifier and the local mean-based NN classifier. Wei Li 0032, Qian Du 0001, Fan Zhang 0007, Wei Hu 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2014 | Ray tracing via GPU rasterization
Wei Hu 0004, Yangyu Huang, Fan Zhang 0007, Guodong Yuan, Wei Li 0032 |
Vis. Comput. | 1 |
| 2013 | Effect of ionosphere refraction on spaceborne SAR imaging precisionabstractIonosphere has different stratification at different height, so it's index of refraction varies with height. When radar's echo signals transmit through the ionosphere, the real propagation path is not the ideal straight line, in fact, it will inflect to a certain extent. For the spaceborne synthetic aperture radar (SAR) system, the imaging precision has been decreased both in azimuth section and range section due to imprecise slant range history. In this paper, we simulate the real distance between satellite and the target according to Snell's law, and analyze the effect of ionospheric refraction on spaceborne SAR imaging performance at L-band C-band and X-band. Fan Zhang 0007, Wei Hu 0004 |
IGARSS | 4 |
| 2013 | Gpu rasterization based octree fast generation algorithm for terrain modelingabstractIn order to meet the need of continuous development of three dimensional (3D) Geographic Information System (GIS), we apply an algorithm that efficiently builds a compact sparse octree into GIS. The algorithm is based on the fast GPU hardware rasterization, which builds the octree structure from top level to bottom level at the same time. Therefore the graphics based parallel algorithm can accelerate the process of generate an octree structure for terrain geometric modeling. Otherwise, the algorithm also supports a fast dynamic update of the octree structure. The result shows that the algorithm can improve the efficiency of octree structure generation and large terrain data modeling. Fan Zhang 0007, Wei Hu 0004 |
IGARSS | 3 |
| 2012 | Edit Propagation via Edge-Aware Filtering
Wei Hu 0004, Zhao Dong 0001, Guodong Yuan |
J. Comput. Sci. Technol. | 1 |
| 2010 | Interactive volume caustics in single-scattering mediaabstractVolume caustics are intricate illumination patterns formed by light first interacting with a specular surface and subsequently being scattered inside a participating medium. Although this phenomenon can be simulated by existing techniques, image synthesis is usually non-trivial and time-consuming. Wei Hu 0004, Zhao Dong 0001, Ivo Ihrke, Thorsten Grosch, Guodong Yuan, Hans-Peter Seidel |
SI3D | 1 |