EDBT 2026 Demo / reviewers in the wild / expert
Wenhui Li 0002
dblp:95/2212-2 · also Wen-hui Li 0002
· DBLP profile ↗
42ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-6490-9852ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ExpPortrait: Novel Expression Generation for Fine-Grained Controllable Portrait AnimationabstractWe propose ExpPortrait, a diffusion-based method for fine-grained facial expression control, which achieves high-fidelity aggregation of input video information to generate high-quality, novel, and diverse facial expressions. To overcome the limitations of existing methods that, due to limited single-view facial information, often yield distorted textures and errors in unseen extreme facial expressions, we propose a dual-layer control mechanism that regulates the output expressions and poses through both image-level and parameter-level inputs. Our approach supports both single and multi-keyframe inputs to guide the generation of sequences with unseen extreme facial expressions. We integrate a motion module to address issues of insufficient inter-frame smoothness, flickering, and instability, and propose an appearance-consistent inference strategy to achieve more realistic facial textures and skin tones. Experiments demonstrate that ExpPortrait generates high-fidelity, unseen extreme facial expressions with superior temporal, 3D, and appearance consistency, outperforming state-of-the-art methods across multiple metrics. Hualiang Wei, Wenhui Li 0002 |
ICMR | 2 |
| 2026 | GDCR: Geometry-enhanced directional consistency representation for point cloud analysis
Zi-Ming Wang 0002, Boxiang Zhang, Yue Wang 0114, Taoli Du, Ying Wang 0024, Wenhui Li 0002 |
Expert Syst. Appl. | 7 |
| 2025 | MncCap: Mining Neural Composition for Zero-shot Image Captioning via Text-only TrainingabstractCurrent text-only image captioning methods leverage the shared feature space of CLIP to train zero-shot image captioning using text data only, leaving feature associations and contextual understanding not fully explored. Neurological studies have revealed that the anterior temporal lobes of the brain are responsible for binding attributes to specific individuals and the corresponding collective connections. Inspired by the above studies, we propose a novel Mining Neural Composition for zero-shot image captioning (MncCap) via text-only training to model the neural composition. During training, we combine the global and local fine-grained features provided by the text clues to achieve a stronger ability of contextual understanding. To express the relationship from discriminative information in the text, we propose a strategy of converting each candidate sentence into a text-tree. During inference, a pre-trained detector is used to obtain the ROI features in the image to improve the contextual integrity of the semantic features. Experimental results conducted on three image captioning benchmark datasets show that our framework achieves remarkable performance improvements. Tongtong Liu 0002, Qinxu Gao, Enhua Song, Wenhui Li 0002 |
ICASSP | 6 |
| 2025 | Global Static Pruning via Adaptive Sample Complexity AwarenessabstractDynamic pruning leverage the feature information of each input sample to dynamically adjust the network structure, generating multiple subnetworks suitable for different sample complexity. However, it inevitably introduces higher computational complexity and increased memory consumption. In addition, complex multi-stage pipelines are required to counteract the performance degradation caused by pruning. In this paper, a simple yet effective global static pruning method based on Adaptive Sample Complexity Awareness is proposed, called ASCA, which achieves model compression without pre-training and fine-tuning. Specifically, an adaptive sample complexity-aware static pruning method is proposed, which leverages task loss to guide the network in enhancing or suppressing the feature learning of samples with varying complexities. Then, a new mask binarization loss is proposed to automatically distinguish important and unimportant channels, avoiding the impact of hand-crafted thresholds on pruning performance. Extensive experiments demonstrate that ASCA outperforms state-of-the-art pruning methods on CIFAR-10 and ImageNet datasets. Yue Wang 0114, Taoli Du, Qinxu Gao, Ying Wang 0024, Wenhui Li 0002 |
ICASSP | 6 |
| 2025 | MMamba: Enhancing image deraining with Morton curve-driven locality learning
Zongzhi Ouyang, Wenhui Li 0002 |
Neurocomputing | 2 |
| 2025 | Semantic Hierarchy-Aware Hyperbolic Representations for Multi-Label Classification With Single Positive LabelsabstractSingle positive multi-label learning (SPML) aims to recognize multiple categories with limited supervision from one positive label in an image. With the emergence of pre-trained visual-language models such as CLIP, recent studies focused on capturing label-to-label dependencies. However, hierarchies with deeper layers of labels or more branches in label-to-label relationships cannot be well expressed in Euclidean space. To address the challenge, we introduce a semantic hierarchy-aware hyperbolic representations framework for single positive multi-label learning. Specifically, drawing inspiration from semantic hierarchical information, we introduce a label relation prior strategy to map single labels to other labels. The semantic chain of labels is extracted along the hierarchical path from the child node to the parent node. Furthermore, hyperbolic entailment constraints are adopted to enforce the semantic similarity between image-text pairs and the hierarchical consistency among labels in hyperbolic space. Experimental results conducted on four SPML benchmark datasets demonstrate that our SHHNet achieves state-of-the-art performance. Tongtong Liu 0002, Ying Wang 0024, Wenhui Li 0002 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Global Channel Pruning With Self-Supervised Mask LearningabstractNetwork pruning is widely used in model compression due to its simplicity and efficiency. Existing methods typically introduce sparse loss regularization to learn masks. However, this sparse regularization approach lacks a clear criterion for evaluating channel importance and relies on manually defined rules, leading to a decline in model performance. In this article, a Self-Supervised Mask Learning (SSML) method for global channel pruning is proposed, casting mask learning as a self-supervised binary classification task to automatically identify less important channels. Specifically, a dedicated pretext task is designed for the channelwise masks, which leverages the original network to generate pseudo-labels from the mask itself to guide mask learning. Then, a polarization mask loss function is proposed, transforming the discrete mask learning problem into a differentiable binary classification problem. The proposed loss function distinguishes the similarity between pseudo-labels and masks, clustering similar masks together in the feature space and separating dissimilar masks, ultimately allowing channels with masks of 0 to be safely removed without damaging the performance of the pruned model. In addition, SSML can train from scratch to yield a compact model. Extensive experiments on CIFAR-10, CIFAR-100 and ImageNet datasets demonstrate that SSML outperforms state-of-the-art methods. For instance, SSML prunes 52.7% of the FLOPs of ResNe34 on the ImageNet dataset with only 0.01% drop in Top-1 accuracy. Moreover, the generalization of SSML is verified on downstream tasks. Tongzhou Zhang 0001, Zi-Ming Wang 0002, Yue Wang 0114, Taoli Du, Wenhui Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | MVITP: Multi-View Image-Text Perception for Few-Shot Remote Sensing Image ClassificationabstractFew-shot learning has been extensively applied in current remote sensing image classification, enabling rapid identification of new classes by leveraging prior knowledge effectively. However, current methods mainly rely on image modality to address the issue of low intra-class similarity and high interclass similarity, while the utilization of multimodal methods in remote sensing tasks remains largely unexplored. Therefore, we propose a novel framework for few-shot remote sensing image classification, named multi-view image-text perception (MVITP). Specifically, it leverages maximum mutual information across multiple views to train an image encoder and generate image features. A text encoder is employed to generate text features. Next, we introduce a multimodal fusion encoder to capture the similarity between image features and text features. Finally, class predictions are further made by computing the similarity between the support set and the query set. We conduct experiments on three remote sensing datasets, demonstrating the outstanding performance of MVITP. Tongtong Liu 0002, Didi Jiao, Wenhui Li 0002 |
ICASSP | 4 |
| 2024 | Multi-fineness Boundaries and the Shifted Ensemble-aware Encoding for Point Cloud Semantic SegmentationabstractPoint cloud segmentation forms the foundation of 3D scene understanding. Boundaries, the intersections of regions, are prone to mis-segmentation. Current point cloud segmentation models exhibit unsatisfactory performance on boundaries. There is limited focus on explicitly addressing semantic segmentation of point cloud boundaries. We introduce a method called Multi-fineness Boundary Constraint (MBC) to tackle this challenge. By querying boundaries at various degrees of fineness and imposing feature constraints within these boundary areas, we enhance the discrimination between boundaries and non-boundaries, improving point cloud boundary segmentation. However, solely emphasizing boundaries may compromise the segmentation accuracy in broader non-boundary regions. To mitigate this, we introduce a new concept of point cloud space termed ensemble and a Shifted Ensemble-aware Perception (SEP) module. This module establishes information interactions between points with minimal computational cost, effectively capturing direct point-to-point long-range correlations within ensembles. It enhances segmentation performance for both boundaries and non-boundaries. Zi-Ming Wang 0002, Boxiang Zhang, Yue Wang 0114, Taoli Du, Wenhui Li 0002 |
ACM Multimedia | 6 |
| 2024 | MMI-ML: Maximize Mutual Information Between Different Views for Few-Shot Remote Sensing Image ClassificationabstractFew-shot learning is widely applied in the current stage for remote sensing image classification to use prior knowledge to identify new classes faster. However, since existing few-shot remote sensing image classification methods only process the feature vectors extracted from the complete image, ignoring the localized knowledge in the input incorporated into the target can better utilize the contextual information to capture more local details. To address these problems, we propose a metric learning framework based on maximizing mutual information between different views (MMI-ML). Specifically, we introduce a self-supervised model to train an embedding network with enhanced feature representation by maximizing mutual information of global features and local features at different scales. In addition, we design a new embedding network to make it more appropriate for the self-supervised model. Finally, we devise a new loss function in the training stage, which can effectively speed up the convergence of the model. We conduct comparative experiments on three public remote sensing datasets, and the experimental results show that the classification accuracy of the MMI-ML framework is improved by up to 3.22%. Yuanyuan Guan, Tongtong Liu 0002, Boxiang Zhang, Wenhui Li 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | ICSFF: Information Constraint on Self-Supervised Feature Fusion for Few-Shot Remote Sensing Image ClassificationabstractThe self-supervised few-shot remote sensing image classification task is to achieve efficient and accurate remote sensing image classification through autonomous learning and feature exploitation with limited data labels. However, at this stage, one common challenge in self-supervised learning is the significant disparity between the self-supervised learning task and the main classification task. This disparity can lead to a situation where the model overly emphasizes features or local information emphasized by the self-supervised task while neglecting the essential global semantic information relevant to the main classification task. To solve these problems, this paper proposes a few-shot remote sensing image classification framework based on information constraint on self-supervised feature fusion called ICSFF. Firstly, we train a supervised model to capture important semantic information. Then, we leverage this supervised information to constrain the learning process of the self-supervised model. We utilize an attention mechanism to integrate supervised information and self-supervised information through a graph structural feature fusion approach, resulting in enhanced feature representations. In addition, we design a new feature extractor called GCCANet. It helps the model to better utilize the key features by incorporating the global attention module, the group convolution, the residual operation, and the channel shuffle module techniques. We conduct comparative experiments on three public remote sensing datasets, and the experimental results show that ICSFF achieves outstanding performance in the remote sensing image classification task. Tongtong Liu 0002, Guanhua Chen 0003, Wenhui Li 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Mx2M: Masked Cross-Modality Modeling in Domain Adaptation for 3D Semantic SegmentationabstractExisting methods of cross-modal domain adaptation for 3D semantic segmentation predict results only via 2D-3D complementarity that is obtained by cross-modal feature matching. However, as lacking supervision in the target domain, the complementarity is not always reliable. The results are not ideal when the domain gap is large. To solve the problem of lacking supervision, we introduce masked modeling into this task and propose a method Mx2M, which utilizes masked cross-modality modeling to reduce the large domain gap. Our Mx2M contains two components. One is the core solution, cross-modal removal and prediction (xMRP), which makes the Mx2M adapt to various scenarios and provides cross-modal self-supervision. The other is a new way of cross-modal feature matching, the dynamic cross-modal filter (DxMF) that ensures the whole method dynamically uses more suitable 2D-3D complementarity. Evaluation of the Mx2M on three DA scenarios, including Day/Night, USA/Singapore, and A2D2/SemanticKITTI, brings large improvements over previous methods on many metrics. Boxiang Zhang, Zunran Wang, Yonggen Ling, Yuanyuan Guan, Shenghao Zhang 0001, Wenhui Li 0002 |
AAAI | 6 |
| 2023 | ShuffleTrans: Patch-wise weight shuffle for transparent object segmentation
Boxiang Zhang, Zunran Wang, Yonggen Ling, Yuanyuan Guan, Shenghao Zhang 0001, Wenhui Li 0002, Lei Wei 0002, Chunxu Zhang |
Neural Networks | 6 |
| 2022 | Memory-Based Jitter: Improving Visual Recognition on Long-Tailed Data with Diversity in MemoryabstractThis paper considers deep visual recognition on long-tailed data. To make our method general, we tackle two applied scenarios, i.e. , deep classification and deep metric learning. Under the long-tailed data distribution, the most classes (i.e., tail classes) only occupy relatively few samples and are prone to lack of within-class diversity. A radical solution is to augment the tail classes with higher diversity. To this end, we introduce a simple and reliable method named Memory-based Jitter (MBJ). We observe that during training, the deep model constantly changes its parameters after every iteration, yielding the phenomenon of weight jitters. Consequentially, given a same image as the input, two historical editions of the model generate two different features in the deeply-embedded space, resulting in feature jitters. Using a memory bank, we collect these (model or feature) jitters across multiple training iterations and get the so-called Memory-based Jitter. The accumulated jitters enhance the within-class diversity for the tail classes and consequentially improves long-tailed visual recognition. With slight modifications, MBJ is applicable for two fundamental visual recognition tasks, i.e., deep image classification and deep metric learning (on long-tailed data). Extensive experiments on five long-tailed classification benchmarks and two deep metric learning benchmarks demonstrate significant improvement. Moreover, the achieved performance are on par with the state of the art on both tasks. Jialun Liu, Wenhui Li 0002, Yifan Sun 0003 |
AAAI | 2 |
| 2022 | Learning Memory-Augmented Unidirectional Metrics for Cross-modality Person Re-identificationabstractThis paper tackles the cross-modality person re-identification (re-ID) problem by suppressing the modality discrepancy. In cross-modality re-ID, the query and gallery images are in different modalities. Given a training identity, the popular deep classification baseline shares the same proxy (i.e., a weight vector in the last classification layer) for two modalities. We find that it has considerable tolerance for the modality gap, because the shared proxy acts as an intermediate relay between two modalities. In response, we propose a Memory-Augmented Unidirectional Metric (MAUM) learning method consisting of two novel designs, i.e., unidirectional metrics, and memory-based augmentation. Specifically, MAUM first learns modality-specific proxies (MS-Proxies) independently under each modality. Afterward, MAUM uses the already-learned MS-Proxies as the static references for pulling close the features in the counterpart modality. These two unidirectional metrics (IR image to RGB proxy and RGB image to IR proxy) jointly alleviate the relay effect and benefit cross-modality association. The cross-modality association is further enhanced by storing the MS-Proxies into memory banks to increase the reference diversity. Importantly, we show that MAUM improves cross-modality re-ID under the modality-balanced setting and gains extra robustness against the modality-imbalance problem. Extensive experiments on SYSU-MMOI and RegDB datasets demonstrate the superiority of MAUM over the state-of-the-art. The code will be available. Jialun Liu, Yifan Sun 0003, Feng Zhu 0005, Hongbin Pei, Yi Yang 0001, Wenhui Li 0002 |
CVPR | 6 |
| 2022 | Feature Cloud: Improving Deep Visual Recognition With Probabilistic Feature AugmentationabstractThis paper considers deep visual recognition on long-tailed data. Under the long-tailed distribution, a small portion of the classes (head classes) occupy most training samples and the most classes (tail classes) only occupy relatively few samples. We observe that such long-tailed distribution significantly distorts the deeply-learned feature space, which consequentially compromises the deep visual recognition. Specifically, during training, each head class is prone to a relatively wide spatial distribution in the deep feature space, while each tail class is prone to a relatively small spatial distribution. In another word, the tail classes usually have much smaller spatial distribution than the head classes, distorting the overall feature space. In response, we propose to explicitly inflate the distribution of each tail class in the deep feature space, so that the tail classes will have comparable distribution range as the head classes. To this end, we replace each tail feature vector with a set of feature vectors on the fly. These feature vectors follow a probabilistic distribution learned from the head classes and yield a “feature cloud” surrounding the original tail feature. We show that the feature cloud effectively transfers the within-class diversity from the head classes onto the tail classes, maintaining an effect of probabilistic feature augmentation. An important advantage of the proposed feature cloud is that it is capable to bring general improvement to long-tailed visual recognition on two fundamental tasks,i.e., deep classification and deep representation learning, in spite of the significant differences between them. Extensive experiments on both deep metric learning benchmarks and deep image classification benchmarks validate the effectiveness of the proposed feature cloud. Jialun Liu, Yifan Sun 0003, Yijin Xu, Hongbin Pei, Wenhui Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | DOBNET: Dynamic Object Boundary-Refinement Network for Real-Time Instance SegmentationabstractMainstream real-time instance segmentation methods always predict masks in the ’detect-then-segment’ way and ignore the object boundaries, leading to resource wasting and indistinct masks. To overcome these drawbacks, we propose a Dynamic Object Boundary-refinement Network (DOBNet) to predict masks in the principle of SOLO [1]. In this method, we first adapt the OctConv [2] as the generator to produce two parallel dynamic convolutions for mask and boundary features, respectively. The Boundary Refinement Module then helps fuse the features from the two convolutions and thereby refine the final predictions with boundary information. Hence, our method attains a precise segmentation while maintaining real-time speed. More specifically, the architecture achieved 37.9 AP on the COCO test-dev2017 dataset with a speed of 31.8 FPS, as is shown in Table 1. The results are more accurate than the existing real-time method. Boxiang Zhang, Yuanyuan Guan, Hongru Liu, Wenhui Li 0002, Ying Wang 0024 |
ICME | 4 |
| 2021 | Global Attention Augmentation Ghost Module: More Features from Lightweight Global Attention ExtractionabstractRecently, in order to deploy neural networks on mobile devices, many studies have focused on reducing the number of parameters and computational complexity of neural networks. However, most existing methods do reduce the computational complexity of deep neural networks, but also greatly sacrifice their performance. To maintain the relative balance of computational complexity and performance of deep neural networks, this paper proposes a Global Attention Augmentation Ghost(GAAG) module, which decreases the number of parameters while bring performance improvements. Analyzed the network architecture of the Ghost module, we empirically show it is waste of computational resources that cheap operation in Ghost module produces more feature maps by linear transformations, only increasing the width of convolution neural network instead of extracting more useful feature information, and in a convolution layer composed of multiple feature blocks, the circulation of channel information is essential to better integrate the information of each feature blocks. Therefore, we propose a lightweight long-range dependency extraction block instead of cheap operation, which increases the ability to extract the non-local information of the model while keeping the computational cost almost invariable. Furthermore, we combine channel shuffle and channel attention to promote the fusion of local and non-local information. The proposed GAAG module is efficient yet effective and can be flexibly plugged into existing convolutional neural networks. Experiments conducted on the benchmark demonstrate that the GAAG module can perfectly replace the traditional convolutional layer in the baseline model. We extensively evaluate our GAAG module on image classification and object detection with backbones of ResNets. The experimental results show that our GAAG module can keep a good balance between lightweight and high performance compared with the similar model. Hongru Liu, Zhezhou Yu, Yuanyuan Guan, Boxiang Zhang, Wenhui Li 0002 |
ICTAI | 6 |
| 2021 | Multi-label classification by formulating label-specific features from simultaneous instance level and feature level
Yuanyuan Guan, Wenhui Li 0002, Boxiang Zhang, Manglai Ji |
Appl. Intell. | 2 |
| 2021 | Identification and Grading of Maize Drought on RGB Images of UAV Based on Improved U-NetabstractA prerequisite for solving many agricultural problems is to accurately estimate the area affected by crop disasters and its severity rating. In this letter, we propose a pipeline to segment the drought area and distinguish the severity rating of the maize on RGB images accessed by an unmanned aerial vehicle (UAV) through a semantic segmentation method based on deep learning. First, the ground truth is created through expert evaluation and visual interpretation with the aid of the Normalized Difference Vegetation Index (NDVI). The neural network structure that was used is based on U-Net. Some structural and parameter improvements on U-net were made using SE-ResNeXt-50 as the backbone with the atrous spatial pyramid pooling (ASPP) module. By using RGB images as the input of the neural network for training, the final trained network can work on RGB images captured by a consumer UAV. The experimental results showed that our pipeline achieved an F1-score of 0.9034 and a Jaccard index of 0.8287 on the test set. Chang Liu 0064, Huiying Li 0002, Anyang Su, Shengbo Chen, Wenhui Li 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2021 | Exploring contextual information for view-wised 3D model retrieval
Wenhui Li 0002, Yuting Su 0001, Zhenlan Zhao |
Multim. Tools Appl. | 1 |
| 2021 | Semi-supervised partial multi-label classification with low-rank and manifold constraints
Yuanyuan Guan, Boxiang Zhang, Wenhui Li 0002, Ying Wang 0024 |
Pattern Recognit. Lett. | 3 |
| 2020 | Deep Representation Learning on Long-Tailed Data: A Learnable Embedding Augmentation PerspectiveabstractThis paper considers learning deep features from long-tailed data. We observe that in the deep feature space, the head classes and the tail classes present different distribution patterns. The head classes have a relatively large spatial span, while the tail classes have a significantly small spatial span, due to the lack of intra-class diversity. This uneven distribution between head and tail classes distorts the overall feature space, which compromises the discriminative ability of the learned features. In response, we seek to expand the distribution of the tail classes during training, so as to alleviate the distortion of the feature space. To this end, we propose to augment each instance of the tail classes with certain disturbances in the deep feature space. With the augmentation, a specified feature vector becomes a set of probable features scattered around itself, which is analogical to an atomic nucleus surrounded by the electron cloud. Intuitively, we name it as ``feature cloud''. The intra-class distribution of the feature cloud is learned from the head classes, and thus provides higher intra-class variation to the tail classes. Consequentially, it alleviates the distortion of the learned feature space, and improves deep representation learning on long tailed data. Extensive experimental evaluations on person re-identification and face recognition tasks confirm the effectiveness of our method. Jialun Liu, Yifan Sun 0003, Chuchu Han, Zhaopeng Dou, Wenhui Li 0002 |
CVPR | 5 |
| 2020 | Parallel generated method of transcriptional regulatory networksabstractSummary Generated method of transcriptional regulatory networks remains an important research in biology. Many approaches have been proposed to construct transcriptional regulatory networks. However, with the increase of ChIP‐seq and RNA‐seq data, the speed of constructing transcriptional regulation networks is still a challenge. Moreover, parallel computing lacks application in constructing gene regulatory networks through the analysis of the relationships between transcription factors (TFs) and target genes (TGs). Therefore, in this paper, a parallel generated method of transcriptional regulatory networks was proposed. First, two datasets, Michigan Cancer Foundation – 7 (MCF‐7) and Cardiomyocytes (CM) were applied. Then, a parallel method was used to generate transcriptional regulatory network with their transcription factors (TFs) and target genes (TGs). Finally, experimental results showed that 61% regulatory relations in MCF‐7 were validated in the Gene Expression Omnibus (GEO), while 29% results needed further experimental verification. Besides, 56% regulatory relations in CM were consistent with GEO, while 33% results were not yet verified. Furthermore, speed of parallel algorithm was faster than traditional serial algorithm in generating transcriptional regulatory networks. Shuai Liu 0002, Na Ta 0005, Mengye Lu, Gaocheng Liu, Weiling Bai, Wenhui Li 0002 |
Concurr. Comput. Pract. Exp. | 6 |
| 2020 | MFENet: Multi-level feature enhancement network for real-time semantic segmentation
Boxiang Zhang, Wenhui Li 0002, Yuming Hui, Jiayun Liu, Yuanyuan Guan |
Neurocomputing | 2 |
| 2020 | A lane detection network based on IBN and attention
Wenhui Li 0002, Feng Qu, Jialun Liu, Fengdong Sun 0001, Ying Wang 0024 |
Multim. Tools Appl. | 1 |
| 2019 | Self-attention recurrent network for saliency detection
Fengdong Sun 0001, Wenhui Li 0002, Yuanyuan Guan |
Multim. Tools Appl. | 2 |
| 2019 | A robust lane detection method based on hyperbolic model
Wenhui Li 0002, Feng Qu, Ying Wang 0024 |
Soft Comput. | 1 |
| 2018 | Flame Detection Using Generic Color Model and Improved Block-Based PCA in Active Infrared CameraabstractIn this paper, we proposed an all-weather flame detection algorithm which could make full use of active infrared cameras presently installed in many public places for surveillance purposes. Firstly, according to the different spectral imaging results in day and night, we propose a video type classification algorithm (VTCA) via imaging clues. VTCA could help us select different flame visual features in color image and infrared image. Secondly, we use a generic YCbCr-color-space-based chrominance model to extract regions of interest (ROI) of flame. Thirdly, two flame dynamic features are used to verify the candidate ROIs, which are common flame flicker feature and an improved block-based PCA in consecutive frames. The experimental results show that the proposed flame detection model has been successfully applied to various situations, including day and night, indoor and outdoor on our test video datasets, and it gives a better performance compared with other state-of-the-art methods. Fengdong Sun 0001, Wenhui Li 0002, Peixun Liu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2018 | Extraction of digital terrain model based on regular mesh generation in mountainous areas
Wenhui Li 0002, Daifeng Han, Huiying Li 0002, Xuezhi Wang 0005, Jinlong Zhu |
Multim. Tools Appl. | 1 |
| 2012 | A hierarchical contour method for automatic 3D city reconstruction from LiDAR dataabstractRecent years LiDAR (Light Detection and Ranging) data is widely used for constructing 3D terrain models which provide realistic impressions of the urban environment. This paper presents a hierarchical contour method for building boundary extraction and 3D reconstruction from LiDAR Data. This method provides acceptable visualization for large-scale poor quality LiDAR data reconstruction. Huiying Li 0002, Zhi Wang 0009, Guan-liang Wu, Wenhui Li 0002, Cai Liu |
IGARSS | 5 |
| 2010 | Fusion of LiDAR data and orthoimage for automatic building reconstructionabstractRecent years LiDAR data is widely used for constructing 3D terrain models which provide realistic impressions of the urban environment. This paper presents an automatic method for extracting 3D building model by the fusion of LiDAR data, 2D building outlines and orthoimage. 2D building outlines is generated by classifying the LiDAR data to terrain and off-terrain points, then detecting building edges points through step-structure detector and generalization. 2D building boundaries are added on the DSM (Digital Surface Model) from LiDAR data to generate complex buildings by using CSG with the Boolean operations of union, intersection and differences. Huiying Li 0002, Shengbo Chen, Zhi Wang 0009, Wenhui Li 0002 |
IGARSS | 4 |
| 2008 | Semantic image classification using statistical local spatial relations model
Dongfeng Han, Wenhui Li 0002, Zongcheng Li |
Multim. Tools Appl. | 2 |
| 2007 | Hilbert-Huang Transform-based Local Regions DescriptorsabstractThis paper presents a new interest local regions descriptors method based on Hilbert-Huang Transform. The neighborhood of the interest local region is decomposed adaptively into oscillatory components called intrinsic mode functions (IMFs). Then the Hilbert transform is applied to each component and get the phase and amplitude information. The proposed descriptors samples the phase angles information and amalgamates them into 10 overlap squares with 8-bin orientation histograms. The experiments show that the proposed descriptors are better than SIFT and other standard descriptors. Essentially, the Hilbert-Huang Transform based descriptors can belong to the class of phase-based descriptors. So it can provides a better way to overcome the illumination changes. Additionally, the Hilbert-Huang transform is a new tool for analyzing signals and the proposed descriptors is a new attempt to the Hilbert-Huang transform. 1 Dongfeng Han, Wenhui Li 0002, Wu Guo |
BMVC | 2 |
| 2006 | Independent Components Analysis for Representation Interest Point Descriptors
Dongfeng Han, Wenhui Li 0002, Tianzhu Wang, Lingling Liu |
ICIC (1) | 2 |
| 2006 | Representation Interest Point Using Empirical Mode Decomposition and Independent Components Analysis
Dongfeng Han, Wenhui Li 0002, Xiaosuo Lu |
ISMIS | 2 |
| 2006 | Image segmentation by aggregation graph-cutsabstractIn this paper, we describe a fast semi-automatic segmentation algorithm using nodes aggregation and graph-cuts. The segmentation process is reliably computed automatically no additional users' efforts are required. It is convenient and efficient in practical applications. Experiments are given and outputs are encouraging. Dongfeng Han, Wenhui Li 0002, Tianzhu Wang, Wang Yi 0002, Yanjie She |
MMM | 2 |
| 2006 | Stochastic collision detection between deformable models using particle swarm optimization algorithmabstractWe present an efficient algorithm for detecting collisions and self-collisions between highly deformable mass models, which is a combination of newly developed stochastic method and particle swarm optimization (PSO) algorithm. In stochastic collision detection, user can balance performance and detection quality by sampling primitive pairs within the models. To accelerate detecting process in the primitive pair space, we introduce PSO algorithm to complete the optimization for the first time. And in the end of this paper, we give the precision and efficiency evaluation about the algorithm and find it might be a reasonable choice for deformable models in collision detection Tianzhu Wang, Wenhui Li 0002, Wang Yi 0002, Zihou Ge, Dongfeng Han |
MMM | 2 |
| 2006 | The Parametric Design Based on Organizational Evolutionary Algorithm
Chunhong Cao, Bin Zhang 0001, Limin Wang 0007, Wenhui Li 0002 |
PRICAI | 4 |
| 2005 | A Rule System for Heterogeneous Spatial Reasoning in Geographic Information System
Wenhui Li 0002 |
DEXA | 2 |
| 2005 | Heterogeneous Spatial Reasoning
Wenhui Li 0002 |
ECSQARU | 2 |
| 2005 | New Rules for Hybrid Spatial Reasoning
Wenhui Li 0002 |
IDEAL | 1 |