EDBT 2026 Demo / reviewers in the wild / expert
Guanghui Wang 0001
dblp:44/2323-1
· DBLP profile ↗
94ranked-venue papers
20as first author
31since 2021 · last 2026
0000-0003-3182-104XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 15 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond ZOH: Advanced Discretization Strategies for Vision Mamba
Fady Ibrahim, Guangjun Liu 0001, Guanghui Wang 0001 |
ICPR (11) | 3 |
| 2026 | Mix-domain contrastive learning with Mamba generator for unpaired H&E-to-IHC stain translation
Zhong Zhang 0012, Song Wang 0008, Guanghui Wang 0001 |
Knowl. Based Syst. | 5 |
| 2026 | PanoTPS-Net: Panoramic room layout estimation via thin plate spline transformation
Hatem Ibrahem, Ahmed Salem 0005, Qinmin Hu, Guanghui Wang 0001 |
Pattern Recognit. | 4 |
| 2026 | T-CACE: A Time-Conditioned Autoregressive Contrast Enhancement Multi-Task Framework for Contrast-Free Liver MRI Synthesis, Segmentation, and DiagnosisabstractMagnetic resonance imaging (MRI) is a leading modality for the diagnosis of liver cancer, significantly improving the classification of the lesion and patient outcomes. However, traditional MRI faces challenges including risks from contrast agent (CA) administration, time-consuming manual assessment, and limited annotated datasets. To address these limitations, we propose a Time-Conditioned Autoregressive Contrast Enhancement (T-CACE) framework for synthesizing multi-phase contrast-enhanced MRI (CEMRI) directly from non-contrast MRI (NCMRI). T-CACE introduces three core innovations: a conditional token encoding (CTE) mechanism that unifies anatomical priors and temporal phase information into latent representations; and a dynamic time-aware attention mask (DTAM) that adaptively modulates inter-phase information flow using a Gaussian-decayed attention mechanism, ensuring smooth and physiologically plausible transitions across phases. Furthermore, a constraint for temporal classification consistency (TCC) aligns the lesion classification output with the evolution of the physiological signal, further enhancing diagnostic reliability. Extensive experiments on two independent liver MRI datasets demonstrate that T-CACE outperforms state-of-the-art methods in image synthesis, segmentation, and lesion classification. This framework offers a clinically relevant and efficient alternative to traditional contrast-enhanced imaging, improving safety, diagnostic efficiency, and reliability for the assessment of liver lesion. Xiaojiao Xiao, Jianfeng Zhao 0004, Qinmin Hu, Guanghui Wang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | LFSRDiff: Light Field Image Super-Resolution via Diffusion ModelsabstractDiffusion models have become a rising star in image super-resolution (SR) tasks. However, it is not trivial to apply diffusion models for light field (LF) image SR, which requires maintaining the high-quality visual appearance of each sub-aperture image (SAI) and the angular consistency between the different SAIs. This paper proposes the first diffusion-based LF image SR model, namely LFSRDiff, by incorporating the LF disentanglement mechanism and residual modeling. Specifically, we introduce a disentangled U-Net (Distg U-Net) for diffusion models, enabling improved extraction and fusion of the spatial and angular information in LF images. Furthermore, we leverage residual modeling in diffusion to learn the residual between the upsampled low-resolution and the ground truth high-resolution, which significantly accelerates model training and yields superior results compared to direct learning. Extensive experiments conducted on the five datasets demonstrate the effectiveness of our approach, which can produce realistic SR results and achieve the highest perceptual metric in terms of LPIPS. Code is publicly available at https://github.com/chaowentao/LFSRDiff. Wentao Chao, Junli Zhao, Fuqing Duan, Guanghui Wang 0001 |
ICASSP | 4 |
| 2025 | ABKD: Pursuing a Proper Allocation of the Probability Mass in Knowledge Distillation via α-β-DivergenceabstractKnowledge Distillation (KD) transfers knowledge from a large teacher model to a smaller student model by minimizing the divergence between their output distributions, typically using forward Kullback-Leibler divergence (FKLD) or reverse KLD (RKLD). It has become an effective training paradigm due to the broader supervision information provided by the teacher distribution compared to one-hot labels. We identify that the core challenge in KD lies in balancing two mode-concentration effects: the Hardness-Concentration effect, which refers to focusing on modes with large errors, and the Confidence-Concentration effect, which refers to focusing on modes with high student confidence. Through an analysis of how probabilities are reassigned during gradient updates, we observe that these two effects are entangled in FKLD and RKLD, but in extreme forms. Specifically, both are too weak in FKLD, causing the student to fail to concentrate on the target class. In contrast, both are too strong in RKLD, causing the student to overly emphasize the target class while ignoring the broader distributional information from the teacher. To address this imbalance, we propose ABKD, a generic framework with $\alpha$-$\beta$-divergence. Our theoretical results show that ABKD offers a smooth interpolation between FKLD and RKLD, achieving a better trade-off between these effects. Extensive experiments on 17 language/vision datasets with 12 teacher-student settings confirm its efficacy. Guanghui Wang 0001, Zhiyong Yang 0001, Zitai Wang, Qianqian Xu 0001, Qingming Huang |
ICML | 1 |
| 2025 | Pyramid Hierarchical Masked Diffusion Model for Imaging SynthesisabstractMedical image synthesis plays a crucial role in clinical workflows, addressing the common issue of missing imaging modalities due to factors such as extended scan times, scan corruption, artifacts, patient motion, and intolerance to contrast agents. The paper presents a novel image synthesis network, the Pyramid Hierarchical Masked Diffusion Model (PHMDiff), which employs a multi-scale hierarchical approach for more detailed control over synthesizing high-quality images across different resolutions and layers. Specifically, this model utilizes randomly multi-scale high-proportion masks to speed up diffusion model training, and balances detail fidelity and overall structure. The integration of a Transformer-based Diffusion model process incorporates cross-granularity regularization, modeling the mutual information consistency across each granularity’s latent spaces, thereby enhancing pixel-level perceptual accuracy. Comprehensive experiments on two challenging datasets demonstrate that PHMDiff achieves superior performance in both the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM), highlighting its capability to produce high-quality synthesized images with excellent structural integrity. Ablation studies further confirm the contributions of each component. Furthermore, the PHMDiff model, a multi-scale image synthesis framework across and within medical imaging modalities, shows significant advantages over other methods. The source code will be released with the paper. The source code is available at https://github.com/xiaojiao929/PHMDiff Xiaojiao Xiao, Qinmin Hu, Guanghui Wang 0001 |
IJCNN | 3 |
| 2025 | EntroFormer: An entropy-based sparse vision transformer for real-time semantic segmentation
Song Wang 0008, Lin Wu 0001, Deyin Liu, Lei Gao 0001, Lin Qi 0001, Guanghui Wang 0001 |
Comput. Vis. Image Underst. | 7 |
| 2025 | Depth-Wise Convolutions in Vision Transformers for efficient training on small datasetsabstractThe Vision Transformer (ViT) leverages the Transformer’s encoder to capture global information by dividing images into patches and achieves superior performance across various computer vision tasks. However, the self-attention mechanism of ViT captures the global context from the outset, overlooking the inherent relationships between neighboring pixels in images or videos. Transformers mainly focus on global information while ignoring the fine-grained local details. Consequently, ViT lacks inductive bias during image or video dataset training. In contrast, convolutional neural networks (CNNs), with their reliance on local filters, possess an inherent inductive bias, making them more efficient and quicker to converge than ViT with less data. In this paper, we present a lightweight Depth-Wise Convolution module as a shortcut in ViT models, bypassing entire Transformer blocks to ensure the models capture both local and global information with minimal overhead. Additionally, we introduce two architecture variants, allowing the Depth-Wise Convolution modules to be applied to multiple Transformer blocks for parameter savings, and incorporating independent parallel Depth-Wise Convolution modules with different kernels to enhance the acquisition of local information. The proposed approach significantly boosts the performance of ViT models on image classification, object detection, and instance segmentation by a large margin, especially on small datasets, as evaluated on CIFAR-10, CIFAR-100, Tiny-ImageNet and ImageNet for image classification, and COCO for object detection and instance segmentation. The source code can be accessed at https://github.com/ZTX-100/Efficient_ViT_with_DW . Tianxiao Zhang, Wenju Xu, Bo Luo, Guanghui Wang 0001 |
Neurocomputing | 4 |
| 2025 | Robust 3D point clouds classification based on declarative defenders
Kaidong Li, Tianxiao Zhang, Cuncong Zhong, Guanghui Wang 0001 |
Neural Comput. Appl. | 5 |
| 2025 | MaskBlur: Spatial and Angular Data Augmentation for Light Field Image Super-ResolutionabstractData augmentation (DA) is an effective approach for enhancing model performance with limited data, such as light field (LF) image super-resolution (SR). LF images inherently possess rich spatial and angular information. Nonetheless, there is a scarcity of DA methodologies explicitly tailored for LF images, and existing works tend to concentrate solely on either the spatial or angular domain. This paper proposes a novel spatial and angular DA strategy named MaskBlur for LF image SR by concurrently addressing spatial and angular aspects. MaskBlur consists of spatial blur and angular dropout two components. Spatial blur is governed by a spatial mask, which controls where pixels are blurred, i.e., pasting pixels between the low-resolution and high-resolution domains. The angular mask is responsible for angular dropout, i.e., selecting which views to perform the spatial blur operation. By doing so, MaskBlur enables the model to treat pixels differently in the spatial and angular domains when super-resolving LF images rather than blindly treating all pixels equally. Extensive experiments demonstrate the efficacy of MaskBlur in significantly enhancing the performance of existing SR methods. We further extend MaskBlur to other LF image tasks such as denoising, deblurring, low-light enhancement, and real-world SR. Wentao Chao, Fuqing Duan, Yulan Guo, Guanghui Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | SuperLoRA: Parameter-Efficient Unified Adaptation of Large Foundation Models
Xiangyu Chen 0008, Jing Liu 0009, Ye Wang 0001, Pu Wang 0004, Matthew Brand, Guanghui Wang 0001, Toshiaki Koike-Akino |
BMVC | 6 |
| 2024 | Disentangled Representation Learning for Controllable Person Image GenerationabstractIn this paper, we propose a novel framework named DRL-CPG to learn disentangled latent representation for controllable person image generation, which can produce realistic person images with desired poses and human attributes (e.g. pose, head, upper clothes, and pants) provided by various source persons. Unlike the existing works leveraging the semantic masks to obtain the representation of each component, we propose to generate disentangled latent code via a novel attribute encoder with transformers trained in a manner of curriculum learning from a relatively easy step to a gradually hard one. A random component mask-agnostic strategy is introduced to randomly remove component masks from the person segmentation masks, which aims at increasing the difficulty of training and promoting the transformer encoder to recognize the underlying boundaries between each component. This enables the model to transfer both the shape and texture of the components. Furthermore, we propose a novel attribute decoder network to integrate multi-level attributes (e.g. the structure feature and the attribute representation) with well-designed Dual Adaptive Denormalization (DAD) residual blocks. Extensive experiments strongly demonstrate that the proposed approach is able to transfer both the texture and shape of different human parts and yield realistic results. To our knowledge, we are the first to learn disentangled latent representations with transformers for person image generation. Wenju Xu, Chengjiang Long, Yongwei Nie, Guanghui Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | PST-Net: Point Cloud Completion Network Based on Local Geometric Feature Reuse and Neighboring Recovery with Taylor ApproximationabstractTransformer has recently been introduced into point cloud completion and achieved inspirational performance on 3D point cloud generation. However, the low-level local geometric features are ignored in the existing feature extraction network, and this leads to the loss of geometry in the recovery results. Meanwhile, FoldingNet models that estimate neighborhood information from predicted center points cannot effectively recover geometric information. In this paper, we propose a point cloud completion network based on local geometric feature reuse and neighboring recovery with Taylor approximation (PST-Net). Specifically, a feature extraction network named Skip-DGCNN is constructed to integrate local and global geometric features to reduce the geometry loss during the feature extraction. In addition, we propose a computational model through Taylor approximation to recover the geometry information in the neighborhood of the prediction center. Moreover, we design the TSMB module corresponding to Taylor's approximation to maintain the end-to-end training mode. The proposed method is extensively evaluated and compared with previous methods on three datasets including PCN, ShapeNet-55 and ShapeNet-34. The proposed model outperforms the state-of-the-art (SOTA) and PoinTr on ShapeNet-55 and ShapeNet-34. The complexity analysis on the PCN dataset shows that the number of FLOPs of our approach is 60.79% lower than that of the SOTA. Visual comparisons demonstrate that the proposed method can effectively and accurately complete the geometry of missing parts. Yinchu Wang, Haijiang Zhu, Guanghui Wang 0001 |
IJCNN | 3 |
| 2023 | Edge-Aware Multi-task Network for Integrating Quantification Segmentation and Uncertainty Prediction of Liver Tumor on Multi-modality Non-contrast MRI
Xiaojiao Xiao, Qinmin Hu, Guanghui Wang 0001 |
MICCAI (4) | 3 |
| 2023 | Accumulated Trivial Attention Matters in Vision Transformers on Small DatasetsabstractVision Transformers has demonstrated competitive performance on computer vision tasks benefiting from their ability to capture long-range dependencies with multi-head self-attention modules and multi-layer perceptron. However, calculating global attention brings another disadvantage compared with convolutional neural networks, i.e. requiring much more data and computations to converge, which makes it difficult to generalize well on small datasets, which is common in practical applications. Previous works are either focusing on transferring knowledge from large datasets or adjusting the structure for small datasets. After carefully examining the self-attention modules, we discover that the number of trivial attention weights is far greater than the important ones and the accumulated trivial weights are dominating the attention in Vision Transformers due to their large quantity, which is not handled by the attention itself. This will cover useful non-trivial attention and harm the performance when trivial attention includes more noise, e.g. in shallow layers for some backbones. To solve this issue, we proposed to divide attention weights into trivial and non-trivial ones by thresholds, then Suppressing Accumulated Trivial Attention (SATA) weights by proposed Trivial WeIghts Suppression Transformation (TWIST) to reduce attention noise. Extensive experiments on CIFAR-100 and Tiny-ImageNet datasets show that our suppressing method boosts the accuracy of Vision Transformers by up to 2.3%. Code is available at https://github.com/xiangyu8/SATA. Xiangyu Chen 0008, Kaidong Li, Cuncong Zhong, Guanghui Wang 0001 |
WACV | 5 |
| 2022 | Robust Structured Declarative Classifiers for 3D Point Clouds: Defending Adversarial Attacks with Implicit GradientsabstractDeep neural networks for 3D point cloud classification, such as PointNet, have been demonstrated to be vulnerable to adversarial attacks. Current adversarial defenders often learn to denoise the (attacked) point clouds by reconstruction, and then feed them to the classifiers as input. In contrast to the literature, we propose a family of robust structured declarative classifiers for point cloud classification, where the internal constrained optimization mechanism can effectively defend adversarial attacks through implicit gradients. Such classifiers can be formulated using a bilevel optimization framework. We further propose an effective and efficient instantiation of our approach, namely, Lattice Point Classifier (LPC), based on structured sparse coding in the permutohedral lattice and 2D convolutional neural networks (CNNs) that is end-to-end trainable. We demonstrate state-of-the-art robust point cloud classification performance on ModelNet40 and ScanNet under seven different attackers. For instance, we achieve 89.51% and 83.16% test accuracy on each dataset under the recent JGBA attacker that outperforms DUP-Net and IF-Defense with PointNet by ~70%. The demo code is available at https://zhang-vislab.github.io. Kaidong Li, Cuncong Zhong, Guanghui Wang 0001 |
CVPR | 4 |
| 2022 | Aggregating Global Features into Local Vision TransformerabstractLocal Transformer-based classification models have recently achieved promising results with relatively low computational costs. However, the effect of aggregating spatial global information of local Transformer-based architecture is not clear. This work investigates the outcome of applying a global attention-based module named multi-resolution overlapped attention (MOA) in the local window-based transformer after each stage. The proposed MOA employs slightly larger and overlapped patches in the key to enable neighborhood pixel information transmission, which leads to significant performance gain. In addition, we thoroughly investigate the effect of the dimension of essential architecture components through extensive experiments and discover an optimum architecture design. Extensive experimental results CIFAR-10, CIFAR-100, and ImageNet-1K datasets demonstrate that the proposed approach outperforms previous vision Transformers with a comparatively fewer number of parameters. The source code and models are publicly available at: https://github.com/krushi1992/MOA-transformer Krushi Patel, Andrés M. Bur, Fengjun Li, Guanghui Wang 0001 |
ICPR | 4 |
| 2022 | Towards more effective PRM-based crowd counting via a multi-resolution fusion and attention network
Usman Sajid, Guanghui Wang 0001 |
Neurocomputing | 2 |
| 2022 | An unsupervised domain adaptation model based on dual-module adversarial training
Yiju Yang, Tianxiao Zhang, Taejoon Kim, Guanghui Wang 0001 |
Neurocomputing | 5 |
| 2022 | Semantic clustering based deduction learning for image recognition and classification
Wenchi Ma, Xuemin Tu, Bo Luo, Guanghui Wang 0001 |
Pattern Recognit. | 4 |
| 2022 | A discriminative channel diversification network for image classification
Krushi Patel, Guanghui Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Building Instance Mapping From ALS Point Clouds Aided by Polygonal MapsabstractBuilding region extraction from ALS point clouds has been widely studied, whereas instance-level building mapping has been overlooked and remains unsolved. In this study, we present a method to extract individual buildings from ALS point clouds with the help of widely accessible polygonal footprints. The key idea is to merge roof segments to a set of building candidates, from which correct instances are selected by finding optimal matches between polygonal footprints and building candidates. The method has three steps: roof segmentation, building candidate generation, and instance-polygon matching. The method is tested on two large-scale scenes of different building types and can generally achieve high instance-level building mapping accuracy (around 90%) when there are large positioning errors (6.0 m) among polygons. Future work will focus on classification errors in preprocessing, shape inconsistency between point clouds and polygons, and building footprint delineation and updating in postprocessing. Shaobo Xia, Sheng Xu 0003, Ruisheng Wang 0001, Jonathan Li 0001, Guanghui Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | A Domain Gap Aware Generative Adversarial Network for Multi-Domain Image TranslationabstractRecent image-to-image translation models have shown great success in mapping local textures between two domains. Existing approaches rely on a cycle-consistency constraint that supervises the generators to learn an inverse mapping. However, learning the inverse mapping introduces extra trainable parameters and it is unable to learn the inverse mapping for some domains. As a result, they are ineffective in the scenarios where (i) multiple visual image domains are involved; (ii) both structure and texture transformations are required; and (iii) semantic consistency is preserved. To solve these challenges, the paper proposes a unified model to translate images across multiple domains with significant domain gaps. Unlike previous models that constrain the generators with the ubiquitous cycle-consistency constraint to achieve the content similarity, the proposed model employs a perceptual self-regularization constraint. With a single unified generator, the model can maintain consistency over the global shapes as well as the local texture information across multiple domains. Extensive qualitative and quantitative evaluations demonstrate the effectiveness and superior performance over state-of-the-art models. It is more effective in representing shape deformation in challenging mappings with significant dataset variation across multiple domains. Wenju Xu, Guanghui Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | DRB-GAN: A Dynamic ResBlock Generative Adversarial Network for Artistic Style TransferabstractThe paper proposes a Dynamic ResBlock Generative Adversarial Network (DRB-GAN) for artistic style transfer. The style code is modeled as the shared parameters for Dynamic ResBlocks connecting both the style encoding network and the style transfer network. In the style encoding network, a style class-aware attention mechanism is used to attend the style feature representation for generating the style codes. In the style transfer network, multiple Dynamic ResBlocks are designed to integrate the style code and the extracted CNN semantic feature and then feed into the spatial window Layer-Instance Normalization (SW-LIN) decoder, which enables high-quality synthetic images with artistic style transfer. Moreover, the style collection conditional discriminator is designed to equip our DRB-GAN model with abilities for both arbitrary style transfer and collection style transfer during the training stage. No matter for arbitrary style transfer or collection style transfer, extensive experiments strongly demonstrate that our proposed DRB-GAN outperforms state-of-the-art methods and exhibits its superior performance in terms of visual quality and efficiency. Our source code is available at https://github.com/xuwenju123/DRB-GAN. Wenju Xu, Chengjiang Long, Ruisheng Wang 0001, Guanghui Wang 0001 |
ICCV | 4 |
| 2021 | Six-Channel Image Representation for Cross-Domain Object Detection
Tianxiao Zhang, Wenchi Ma, Guanghui Wang 0001 |
ICIG (1) | 3 |
| 2021 | Model Predictive Control of Nonlinear Latent Force Models: A Scenario-Based Approach
Thomas Woodruff, Iman Askari, Guanghui Wang 0001, Huazhen Fang |
ICRA | 3 |
| 2021 | Parallel Scale-wise Attention Network for Effective Scene Text RecognitionabstractThe paper proposes a new text recognition network for scene-text images. Many state-of-the-art methods employ the attention mechanism either in the text encoder or decoder for the text alignment. Although the encoder-based attention yields promising results, these schemes inherit noticeable limitations. They perform the feature extraction (FE) and visual attention (VA) sequentially, which bounds the attention mechanism to rely only on the FE final single-scale output. Moreover, the utilization of the attention process is limited by only applying it directly to the single scale feature-maps. To address these issues, we propose a new multi-scale and encoder-based attention network for text recognition that performs the multi-scale FE and VA in parallel. The multi-scale channels also undergo regular fusion with each other to develop the coordinated knowledge together. Quantitative evaluation and robustness analysis on the standard benchmarks demonstrate that the proposed network outperforms the state-of-the-art in most cases. Usman Sajid, Michael Chow, Taejoon Kim, Guanghui Wang 0001 |
IJCNN | 5 |
| 2021 | A Fine-Grained Visual Attention Approach for Fingerspelling Recognition in the WildabstractFingerspelling in sign language has been the means of communicating technical terms and proper nouns when they do not have dedicated sign language gestures. Automatic recognition of fingerspelling can help resolve communication barriers when interacting with deaf people. The main challenges prevalent in fingerspelling recognition are the ambiguity in the gestures and strong articulation of the hands. The automatic recognition model should address high inter-class visual similarity and high intra-class variation in the gestures. Most of the existing research in fingerspelling recognition has focused on the dataset collected in a controlled environment. The recent collection of a large-scale annotated fingerspelling dataset in the wild, from social media and online platforms, captures the challenges in a real-world scenario. In this work, we propose a fine-grained visual attention mechanism using the Transformer model for the sequence-to-sequence prediction task in the wild dataset. The fine-grained attention is achieved by utilizing the change in motion of the video frames (optical flow) in sequential context-based attention along with a Transformer encoder model. The unsegmented continuous video dataset is jointly trained by balancing the Connectionist Temporal Classification (CTC) loss and the maximum-entropy loss. The proposed approach can capture better fine-grained attention in a single iteration. Experiment evaluations show that it outperforms the state-of-the-art approaches. Kamala Gajurel, Cuncong Zhong, Guanghui Wang 0001 |
SMC | 3 |
| 2021 | SOSD-Net: Joint semantic object segmentation and depth estimation from monocular images
Lei He 0004, Jiwen Lu, Guanghui Wang 0001, Shiyu Song, Jie Zhou 0001 |
Neurocomputing | 3 |
| 2021 | Deep feature augmentation for occluded image classification
Feng Cen, Wuzhuang Li, Guanghui Wang 0001 |
Pattern Recognit. | 4 |
| 2020 | Multi-Resolution Fusion and Multi-scale Input Priors Based Crowd CountingabstractCrowd counting in still images is a challenging problem in practice due to huge crowd-density variations, large perspective changes, severe occlusion, and variable lighting conditions. The state-of-the-art patch rescaling module (PRM) based approaches prove to be very effective in improving the crowd counting performance. However, the PRM module requires an additional and compromising crowd-density classification process. To address these issues and challenges, the paper proposes a new multi-resolution fusion based end-to-end crowd counting network. It employs three deep-layers based columns/branches, each catering the respective crowd-density scale. These columns regularly fuse (share) the information with each other. The network is divided into three phases with each phase containing one or more columns. Three input priors are introduced to serve as an efficient and effective alternative to the PRM module, without requiring any additional classification operations. Along with the final crowd count regression head, the network also contains three auxiliary crowd estimation regression heads, which are strategically placed at each phase end to boost the overall performance. Comprehensive experiments on three benchmark datasets demonstrate that the proposed approach outperforms all the state-of-the-art models under the RMSE evaluation metric. The proposed approach also has better generalization capability with the best results during the cross-dataset experiments. Usman Sajid, Wenchi Ma, Guanghui Wang 0001 |
ICPR | 3 |
| 2020 | Why Layer-Wise Learning is Hard to Scale-up and a Possible Solution via Accelerated DownsamplingabstractLayer-wise learning, as an alternative to global backpropagation, is memory efficient and easy to interpret, analyze. Recent studies demonstrate that layer-wise learning can achieve state-of-the-art performance in image classification on various datasets. However, previous studies on layer-wise learning are limited to networks with simple hierarchical structures, and the performance decreases severely for deeper networks like ResNet. This paper, for the first time, reveals the fundamental reason that impedes the scale-up of layer-wise learning is the relatively poor separability of the feature space in shallow layers. This argument is empirically verified by controlling the intensity of the convolution operation in local layers. We discover that the poorly-separable features from shallow layers are mismatched with the strong supervision constraint throughout the entire network, making the layer-wise learning sensitive to network depth. The paper further proposes a downsampling acceleration approach to weaken the poor learning of shallow layers so as to transfer the learning emphasis to deep feature space where the separability matches better with the supervision restraint. Extensive experiments have been conducted to verify the finding and demonstrate the advantages of the proposed downsampling acceleration in improving the performance of layer-wise learning. Wenchi Ma, Kaidong Li, Guanghui Wang 0001 |
ICTAI | 4 |
| 2020 | Improving High Dynamic Range Image Based Light MeasurementabstractThis study proposes a fast high dynamic range imaging (HDRI) technique for light measurement to shorten the long capturing time of current camera-aided computational photography widely used in lighting practice. In comparison with the conventional meter measurement, HDRI-assisted lighting measurement is a remote, efficient, affordable yet time-consuming method. The fast HDRI technique increases the film speed (ISO) to speed up the process taking a sequence of low dynamic range images. Since increasing camera's film speed may introduce more image noise, the possible error rate of the proposed method is evaluated by applying Gaussian noise estimation and impulsive noise detection on the image with different film speeds. In addition, a new per-pixel calculation process is developed to retrieve the illuminance of a target scene with selected regions of interest, which can be used to assist human-centric lighting tasks. Extensive comparative experiments are also conducted to verify the accuracy and efficiency of the proposed method. Hankun Li, Hongyi Cai, Guanghui Wang 0001 |
SMC | 3 |
| 2020 | Classification of Noncoding RNA Elements Using Deep Convolutional Neural NetworksabstractThe paper proposes to employ deep convolutional neural networks (CNNs) to classify noncoding RNA (ncRNA) sequences. To this end, we first propose an efficient approach to convert the RNA sequences into images characterizing their base-pairing probability. As a result, classifying RNA sequences is converted to an image classification problem that can be efficiently solved by available CNN-based classification models. The paper also considers the folding potential of the ncRNAs in addition to their primary sequence. Based on the proposed approach, a benchmark image classification dataset is generated from the RFAM database of ncRNA sequences. In addition, three classical CNN models have been implemented and compared to demonstrate the superior performance and efficiency of the proposed approach. Extensive experimental results show the great potential of using deep learning approaches for RNA classification. Brian McClannahan, Krushi Patel, Usman Sajid, Cuncong Zhong, Guanghui Wang 0001 |
SMC | 5 |
| 2020 | Real-time Golf Ball Detection and Tracking Based on Convolutional Neural NetworksabstractThis paper focuses on the problem of real-time detection and tracking of a golf ball from video sequences. We propose an efficient and effective solution by integrating object detection and a discrete Kalman model. For ball detection, three classical convolutional neural network based detection models are implemented, including Faster R-CNN, YOLOv3, and YOLOv3 tiny. At the tracking stage, a discrete Kalman filter is employed to predict the location of the golf ball based on the previous observations. To increase the detection accuracy and speed, we propose to use image patches rather than the entire images for detection. In order to train the detection models and test the tracking algorithm, we collect and annotate a collection of golf ball dataset. Extensive experimental results are performed to demonstrate the effectiveness and superior performance of the proposed approach. Tianxiao Zhang, Yiju Yang, Zongbo Wang, Guanghui Wang 0001 |
SMC | 5 |
| 2020 | Plug-and-Play Rescaling Based Crowd Counting in Static ImagesabstractCrowd counting is a challenging problem especially in the presence of huge crowd diversity across images and complex cluttered crowd-like background regions, where most previous approaches do not generalize well and consequently produce either huge crowd underestimation or overestimation. To address these challenges, we propose a new image patch rescaling module (PRM) and three independent PRM employed crowd counting methods. The proposed frameworks use the PRM module to rescale the image regions (patches) that require special treatment, whereas the classification process helps in recognizing and discarding any cluttered crowd-like background regions which may result in overestimation. Experiments on three standard benchmarks and cross-dataset evaluation show that our approach outperforms the state-of-the-art models in the RMSE evaluation metric with an improvement up to 10.4%, and possesses superior generalization ability to new datasets. Usman Sajid, Guanghui Wang 0001 |
WACV | 2 |
| 2020 | Towards Learning Affine-Invariant Representations via Data-Efficient CNNsabstractIn this paper we propose integrating a priori knowledge into both design and training of convolutional neural networks (CNNs) to learn object representations that are invariant to affine transformations (i.e. translation, scale, rotation). Accordingly we propose a novel multi-scale maxout CNN and train it end-to-end with a novel rotation-invariant regularizer. This regularizer aims to enforce the weights in each 2D spatial filter to approximate circular patterns. In this way, we manage to handle affine transformations in training using convolution, multi-scale maxout, and circular filters. Empirically we demonstrate that such knowledge can significantly improve the data-efficiency as well as generalization and robustness of learned models. For instance, on the Traffic Sign data set and trained with only 10 images per class, our method can achieve 84.15% that outperforms the state-of-the-art by 29.80% in terms of test accuracy. Wenju Xu, Guanghui Wang 0001, Alan Sullivan |
WACV | 2 |
| 2020 | Self-Orthogonality Module: A Network Architecture Plug-in for Learning Orthogonal FiltersabstractIn this paper, we investigate the empirical impact of orthogonality regularization (OR) in deep learning, either solo or collaboratively. Recent works on OR showed some promising results on the accuracy. In our ablation study, however, we do not observe such significant improvement from existing OR techniques compared with the conventional training based on weight decay, dropout, and batch normalization. To identify the real gain from OR, inspired by the locality sensitive hashing (LSH) in angle estimation, we propose to introduce an implicit self-regularization into OR to push the mean and variance of filter angles in a network towards 90° and 0° simultaneously to achieve (near) orthogonality among the filters, without using any other explicit regularization. Our regularization can be implemented as an architectural plug-in and integrated with an arbitrary network. We reveal that OR helps stabilize the training process and leads to faster convergence and better generalization. Wenchi Ma, Yuanwei Wu, Guanghui Wang 0001 |
WACV | 4 |
| 2020 | Adaptively Denoising Proposal Collection for Weakly Supervised Object Localization
Wenju Xu, Yuanwei Wu, Wenchi Ma, Guanghui Wang 0001 |
Neural Process. Lett. | 4 |
| 2020 | MDFN: Multi-scale deep feature learning network for object detection
Wenchi Ma, Yuanwei Wu, Feng Cen, Guanghui Wang 0001 |
Pattern Recognit. | 4 |
| 2020 | ZoomCount: A Zooming Mechanism for Crowd Counting in Static ImagesabstractThis paper proposes a novel approach for crowd counting in low to high density scenarios in static images. Current approaches cannot handle huge crowd diversity well and thus perform poorly in extreme cases, where the crowd density in different regions of an image is either too low or too high, leading to crowd underestimation or overestimation. The proposed solution is based on the observation that detecting and handling such extreme cases in a specialized way leads to better crowd estimation. Additionally, existing methods find it hard to differentiate between the actual crowd and the cluttered background regions, resulting in further count overestimation. To address these issues, we propose a simple yet effective modular approach, where an input image is first subdivided into fixed-size patches and then fed to a four-way classification module labeling each image patch as low, medium, high-dense or no-crowd. This module also provides a count for each label, which is then analyzed via a specifically devised novel decision module to decide whether the image belongs to any of the two extreme cases (very low or very high density) or a normal case. Images, specified as high- or low-density extreme or a normal case, pass through dedicated zooming or normal patch-making blocks respectively before routing to the regressor in the form of fixed-size patches for crowd estimate. Extensive experimental evaluations demonstrate that the proposed approach outperforms the state-of-the-art methods on four benchmarks under most of the evaluation criteria. Usman Sajid, Hasan Sajid, Guanghui Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Joint Correlation Filtering for Visual TrackingabstractCorrelation filtering-based visual tracking has achieved impressive success in terms of both tracking accuracy and computational efficiency. In this paper, a novel correlation filtering approach is proposed by means of joint learning to bridge the gap between the circulant filtering and the classical filtering methods. The circulant structure of tracking and the information from successive frames are simultaneously exploited in the proposed work. A new formulation for the correlation filter learning is proposed to enhance the discrimination of the learned filter by integrating both the kernel and the image feature domains. The proposed approach is computationally efficient since a closed-form solution is derived for the new formulation. Extensive experiments are conducted on two popular tracking benchmarks, and the experimental results demonstrate that the proposed tracker outperforms most of the state-of-the-art trackers. Yao Sui, Guanghui Wang 0001, Li Zhang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Boosting Occluded Image Classification via Subspace Decomposition-Based Estimation of Deep FeaturesabstractClassification of partially occluded images is a highly challenging computer vision problem even for the cutting-edge deep learning technologies. To achieve a robust image classification for occluded images, this article proposes a novel scheme using the subspace decomposition-based estimation (SDBE). The proposed SDBE-based classification scheme first employs a base convolutional neural network to extract the deep feature vector (DFV) and then utilizes the SDBE to compute the DFV of the original occlusion-free image for classification. The SDBE is performed by projecting the DFV of the occluded image onto the linear span of a class dictionary (CD) along the linear span of an occlusion error dictionary (OED). The CD and OED are constructed, respectively, by concatenating the DFVs of a training set and the occlusion error vectors of an extra set of image pairs. Two implementations of the SDBE are studied in this article: 1) the l1-norm and 2) the squared l2-norm regularized least-squares estimates. By employing the ResNet-152, pretrained on the ImageNet Large-Scale Visual Recognition Challenge 2012 (ILSVRC2012) training set, as the base network, the proposed SBDE-based classification scheme is extensively evaluated on the Caltech-101 and ILSVRC2012 datasets. Extensive experimental results demonstrate that the proposed SDBE-based scheme dramatically boosts the classification accuracy for occluded images, and achieves around 22.25% increase in classification accuracy under 20% occlusion on the ILSVRC2012 dataset. Feng Cen, Guanghui Wang 0001 |
IEEE Trans. Cybern. | 2 |
| 2019 | Exploiting the Anisotropy of Correlation Filter Learning for Visual Tracking
Yao Sui, Guanghui Wang 0001, Yafei Tang, Li Zhang 0023 |
Int. J. Comput. Vis. | 3 |
| 2019 | Stacked Wasserstein Autoencoder
Wenju Xu, Shawn Shahriar Keshmiri, Guanghui Wang 0001 |
Neurocomputing | 3 |
| 2019 | 3D reconstruction for ultrasonic C-scan images of tissue-mimicking phantom based on an improved K-nearest neighbor filtering
Haijiang Zhu, Longbiao He, Guanghui Wang 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Sparse subspace clustering via Low-Rank structure propagation
Yao Sui, Guanghui Wang 0001, Li Zhang 0023 |
Pattern Recognit. | 2 |
| 2019 | Toward learning a unified many-to-many mapping for diverse image translation
Wenju Xu, Shawn Shahriar Keshmiri, Guanghui Wang 0001 |
Pattern Recognit. | 3 |
| 2019 | Adversarially Approximated Autoencoder for Image Generation and ManipulationabstractRegularized autoencoders learn the latent codes, a structure with the regularization under the distribution, which enables them the capability to infer the latent codes given observations and generate new samples given the codes. However, they are sometimes ambiguous as they tend to produce reconstructions that are not necessarily a faithful reproduction of the inputs. The main reason is to enforce the learned latent code distribution to match a prior distribution while the true distribution remains unknown. To improve the reconstruction quality and learn the latent space a manifold structure, this paper presents a novel approach using the adversarially approximated autoencoder (AAAE) to investigate the latent codes with adversarial approximation. Instead of regularizing the latent codes by penalizing on the distance between the distributions of the model and the target, AAAE learns the autoencoder flexibly and approximates the latent space with a simpler generator. The ratio is estimated using a generative adversarial network to enforce the similarity of the distributions. In addition, the image space is regularized with an additional adversarial regularizer. The proposed approach unifies two deep generative models for both latent space inference and diverse generation. The learning scheme is realized without regularization on the latent codes, which also encourages faithful reconstruction. Extensive validation experiments on four real-world datasets demonstrate the superior performance of AAAE. In comparison to the state-of-the-art approaches, AAAE generates samples with better quality and shares the properties of a regularized autoencoder with a nice latent manifold structure. Wenju Xu, Shawn Shahriar Keshmiri, Guanghui Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | BPGrad: Towards Global Optimality in Deep Learning via Branch and PruningabstractUnderstanding the global optimality in deep learning (DL) has been attracting more and more attention recently. Conventional DL solvers, however, have not been developed intentionally to seek for such global optimality. In this paper we propose a novel approximation algorithm, BPGrad, towards optimizing deep models globally via branch and pruning. Our BPGrad algorithm is based on the assumption of Lipschitz continuity in DL, and as a result it can adaptively determine the step size for current gradient given the history of previous updates, wherein theoretically no smaller steps can achieve the global optimality. We prove that, by repeating such branch-and-pruning procedure, we can locate the global optimality within finite iterations. Empirically an efficient solver based on BPGrad for DL is proposed as well, and it outperforms conventional DL solvers such as Adagrad, Adadelta, RMSProp, and Adam in the tasks of object recognition, detection, and segmentation. Yuanwei Wu, Guanghui Wang 0001 |
CVPR | 3 |
| 2018 | Spindle-Net: CNNs for Monocular Depth Inference with Dilation Kernel MethodabstractLearning depth from a single image is an important issue in computer vision. To solve this problem, encoder-decoder architect is usually employed as a powerful architecture to learn the dense corresponding function. In this work, we propose a symmetrical Spindle network of the encoder-decoder to learn the fine-grained depth. Unlike traditional convolution neural network, we first boost up the feature maps from low-dimension space to a high-dimension space, then extract the features for monocular depth learning. In order to overcome limitation of the computer memory, a single image super-resolution technique is proposed to replace the boosting process by fusing local cues in edge direction. Given the super-resolution images, the monocular depth learning needs more global information than most architectures for pixel-wise predictions. To address this issue, dilation kernel method is proposed to enlarge the receptive field in each layer. For the task of the super-resolution, the proposed method achieves better performance than the state-of-the-art methods. Extensive experiments on the monocular depth inference demonstrate that the Spindle network could achieve comparable performance on the NYU and Make3D datasets, compared with the state-of-the-art algorithms. The proposed method reveals a new perspective to learn the depth from a single image, which shows a promising generality to other pixel-wise prediction problems. Lei He 0004, Guanghui Wang 0001 |
ICPR | 3 |
| 2018 | MDCN: Multi-Scale, Deep Inception Convolutional Neural Networks for Efficient Object DetectionabstractObject detection in challenging situations such as scale variation, occlusion, and truncation depends not only on feature details but also on contextual information. Most previous networks emphasize too much on detailed feature extraction through deeper and wider networks, which may enhance the accuracy of object detection to certain extent. However, the feature details are easily being changed or washed out after passing through complicated filtering structures. To better handle these challenges, the paper proposes a novel framework, multi-scale, deep inception convolutional neural network (MDCN), which focuses on wider and broader object regions by activating feature maps produced in the deep part of the network. Instead of incepting inner layers in the shallow part of the network, multi-scale inceptions are introduced in the deep layers. The proposed framework integrates the contextual information into the learning process through a single-shot network structure. It is computational efficient and avoids the hard training problem of previous macro feature extraction network designed for shallow layers. Extensive experiments demonstrate the effectiveness and superior performance of MDCN over the state-of-the-art models. Wenchi Ma, Yuanwei Wu, Zongbo Wang, Guanghui Wang 0001 |
ICPR | 4 |
| 2018 | An Efficient Approach for Polyps Detection in Endoscopic Videos Based on Faster R-CNNabstractPolyp has long been considered as one of the major etiologies to colorectal cancer which is a fatal disease around the world, thus early detection and recognition of polyps plays an crucial role in clinical routines. Accurate diagnoses of polyps through endoscopes operated by physicians becomes a chanllenging task not only due to the varying expertise of physicians, but also the inherent nature of endoscopic inspections. To facilitate this process, computer-aid techniques that emphasize on fully-conventional image processing and novel machine learning enhanced approaches have been dedicatedly designed for polyp detection in endoscopic videos or images. Among all proposed algorithms, deep learning based methods take the lead in terms of multiple metrics in evolutions for algorithmic performance. In this work, a highly effective model, namely the faster region-based convolutional neural network (Faster R-CNN) is implemented for polyp detection. In comparison with the reported results of the state-of-the-art approaches on polyps detection, extensive experiments demonstrate that the Faster R-CNN achieves very competing results, and it is an efficient approach for clinical practice. Xi Mo, Ke Tao, Guanghui Wang 0001 |
ICPR | 4 |
| 2018 | Visual Tracking via Subspace Learning: A Discriminative Approach
Yao Sui, Yafei Tang, Li Zhang 0023, Guanghui Wang 0001 |
Int. J. Comput. Vis. | 4 |
| 2018 | Correlation Filter Learning Toward Peak Strength for Visual TrackingabstractThis paper presents a novel visual tracking approach to correlation filter learning toward peak strength of correlation response. Previous methods leverage all features of the target and the immediate background to learn a correlation filter. Some features, however, may be distractive to tracking, like those from occlusion and local deformation, resulting in unstable tracking performance. This paper aims at solving this issue and proposes a novel algorithm to learn the correlation filter. The proposed approach, by imposing an elastic net constraint on the filter, can adaptively eliminate those distractive features in the correlation filtering. A new peak strength metric is proposed to measure the discriminative capability of the learned correlation filter. It is demonstrated that the proposed approach effectively strengthens the peak of the correlation response, leading to more discriminative performance than previous methods. Extensive experiments on a challenging visual tracking benchmark demonstrate that the proposed tracker outperforms most state-of-the-art methods. Yao Sui, Guanghui Wang 0001, Li Zhang 0023 |
IEEE Trans. Cybern. | 2 |
| 2018 | Learning Depth From Single Images With Deep Neural Network Embedding Focal LengthabstractLearning depth from a single image, as an important issue in scene understanding, has attracted a lot of attention in the past decade. The accuracy of the depth estimation has been improved from conditional Markov random fields, non-parametric methods, to deep convolutional neural networks most recently. However, there exist inherent ambiguities in recovering 3D from a single 2D image. In this paper, we first prove the ambiguity between the focal length and monocular depth learning, and verify the result using experiments, showing that the focal length has a great influence on accurate depth recovery. In order to learn monocular depth by embedding the focal length, we propose a method to generate synthetic varying-focal-length dataset from fixed-focal-length datasets, and a simple and effective method is implemented to fill the holes in the newly generated images. For the sake of accurate depth recovery, we propose a novel deep neural network to infer depth through effectively fusing the middle-level information on the fixed-focal-length dataset, which outperforms the state-of-the-art methods built on pretrained VGG. Furthermore, the newly generated varying-focallength dataset is taken as input to the proposed network in both learning and inference phases. Extensive experiments on the fixed- and varying-focal-length datasets demonstrate that the learned monocular depth with embedded focal length is significantly improved compared to that without embedding the focal length information. Lei He 0004, Guanghui Wang 0001, Zhanyi Hu |
IEEE Trans. Image Process. | 2 |
| 2018 | Exploiting Spatial-Temporal Locality of Tracking via Structured Dictionary LearningabstractIn this paper, a novel spatial-temporal locality is proposed and unified via a discriminative dictionary learning framework for visual tracking. By exploring the strong local correlations between temporally obtained target and their spatially distributed nearby background neighbors, a spatial-temporal locality is obtained. The locality is formulated as a subspace model and exploited under a unified structure of discriminative dictionary learning with a subspace structure. Using the learned dictionary, the target and its background can be described and distinguished effectively through their sparse codes. As a result, the target is localized by integrating both the descriptive and the discriminative qualities. Extensive experiments on various challenging video sequences demonstrate the superior performance of proposed algorithm over the other state-of-the-art approaches. Yao Sui, Guanghui Wang 0001, Li Zhang 0023, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Salient object detection based on global multi-scale superpixel contrastabstractSalient object detection, as a necessary step of many computer vision applications, has attracted extensive attention in recent years. A novel salient object detection method is proposed based on multi‐superpixel‐scale contrast. Saliency value of each superpixel is measured with a global score, which is computed using the region's colour contrast and the spatial distances to all other regions in the image. High‐level information is also incorporated to improve the performance, and the saliency maps are fused across multiple levels to yield a reliable final result using the modified multi‐layer cellular automata. The proposed algorithm is evaluated and compared with five state‐of‐the‐art approaches on three publicly standard datasets. Both quantitative and qualitative experimental results demonstrate the effectiveness and efficiency of the proposed method. Jin-Fu Yang, Guanghui Wang 0001, Ming-Ai Li |
IET Comput. Vis. | 3 |
| 2016 | Tracking Completion
Yao Sui, Guanghui Wang 0001, Yafei Tang, Li Zhang 0023 |
ECCV (8) | 2 |
| 2016 | Real-Time Visual Tracking: Promoting the Robustness of Correlation Filter Learning
Yao Sui, Guanghui Wang 0001, Yafei Tang, Li Zhang 0023 |
ECCV (8) | 3 |
| 2016 | Fast and Robust Object Tracking with Adaptive DetectionabstractObject detection and tracking is an important research topic in computer vision with numerous practical applications. Although great progress has been made both in object detection and tracking, it is still a big challenge in automatic real-time applications. In this paper, a fast and robust approach is proposed by integrating an adaptive object detection technique within a kernelized correlation filter (KCF) framework. The KCF tracker is automatically initialized via salient object detection and localization. An adaptive object detection strategy is proposed to refine the location and boundary of the object when the tracking confidence value is below a certain threshold. In addition, a reliable post-processing technique is designed to accurately localize the object from a saliency map. Extensive quantitative and qualitative experiments on the challenging datasets have been performed to verify the proposed approach, which also demonstrates that our approach greatly outperforms the state-of-the-art methods in terms of tracking speed and accuracy. Sushil Pratap Bharati, Soumyaroop Nandi, Yuanwei Wu, Yao Sui, Guanghui Wang 0001 |
ICTAI | 5 |
| 2016 | Textual Ontology and Visual Features Based Search for a Paleontology Digital LibraryabstractThe Treatise on Invertebrate Paleontology is the most reliable information source of invertebrate paleontology research. Based on this Treatise, an Invertebrate Paleontology Knowledgebase (IPKB) has been built as a digital library to provide these data through a web interface. However, the search functions provided by the old IPKB system are only based on textual information, while some more important information, such as textual ontology and fossil images, are not considered at all. In order to overcome this limitation, and provide more reliable and flexible search options, we develop a new hybrid search function for the current IPKB system. In particular, we propose an approach to extract textual ontology information for each genus, as well as build a fossil image dataset where each image is tagged with its genus name. Based on the data from both sources, a hybrid search system is developed by integrating both textual and visual features, and thus, more search options are available to users and the searching results are significantly improved. Ranjith Sompalli, Guanghui Wang 0001, Bo Luo |
ICTAI | 3 |
| 2016 | Triangulation and metric of lines based on geometric error
Fuchao Wu, Ming Zhang 0031, Guanghui Wang 0001, Zhanyi Hu |
Comput. Vis. Image Underst. | 3 |
| 2016 | A novel feature extraction method for scene recognition based on Centered Convolutional Restricted Boltzmann Machines
Jingyu Gao, Jin-Fu Yang, Guanghui Wang 0001, Ming-Ai Li |
Neurocomputing | 3 |
| 2015 | Scene and place recognition using a hierarchical latent topic model
Jin-Fu Yang, Guanghui Wang 0001, Ming-Ai Li |
Neurocomputing | 3 |
| 2015 | A comparative experimental study of image feature detectors and descriptors
Dibyendu Mukherjee, Q. M. Jonathan Wu, Guanghui Wang 0001 |
Mach. Vis. Appl. | 3 |
| 2013 | Robust rank-4 affine factorization for structure from motionabstractThe paper focuses on 3D structure and motion factorization from uncalibrated image sequences. A rank-4 affine factorization algorithm and a robust structure and motion factorization scheme are proposed to handle outlying and missing data. The novelty and main contribution of the paper are as follows: (i) The rank-4 factorization algorithm is a new addition to previous affine factorization family using rank-3 constraint; (ii) the outliers and image uncertainty are estimated directly from the image reprojection residuals; and (iii) the robust factorization scheme is proved empirically to be more efficient and accurate than other robust algorithms. Extensive experiments on synthetic data and real images validate the proposed approach. Guanghui Wang 0001, John S. Zelek, Q. M. Jonathan Wu, Ruzena Bajcsy |
WACV | 1 |
| 2012 | Structure and Motion Recovery Based on Spatial-and-Temporal-Weighted FactorizationabstractThis paper focuses on the problem of structure and motion recovery from uncalibrated image sequences. It has been empirically proven that image measurement uncertainties can be modeled spatially and temporally by virtue of reprojection residuals. Consequently, a spatial-and-temporal-weighted factorization (STWF) algorithm is proposed to handle significant noise contained in the tracking data. This paper presents three novelties and contributions. First, the image reprojection residual of a feature point is demonstrated to be generally proportional to the error magnitude associated with the image point. Second, the error distributions are estimated from a different perspective, that of the reprojection residuals. The image errors are modeled both spatially and temporally to cope with different kinds of uncertainties. Previous studies have considered only the spatial information. Third, based on the estimated error distributions, an STWF algorithm is proposed to improve the overall accuracy and robustness of traditional approaches. Unlike existing approaches, the proposed technique does not require prior information of image measurement and is easy to implement. Extensive experiments on synthetic data and real images validate the proposed method. Guanghui Wang 0001, John S. Zelek, Q. M. Jonathan Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Tracking and Pairing Vehicle Headlight in Night ScenesabstractTraffic surveillance is an important topic in computer vision and intelligent transportation systems and has intensively been studied in the past decades. However, most of the state-of-the-art methods concentrate on daytime traffic monitoring. In this paper, we propose a nighttime traffic surveillance system, which consists of headlight detection, headlight tracking and pairing, and camera calibration and vehicle speed estimation. First, a vehicle headlight is detected using a reflection intensity map and a reflection suppressed map based on the analysis of the light attenuation model. Second, the headlight is tracked and paired by utilizing a simple yet effective bidirectional reasoning algorithm. Finally, the trajectories of the vehicle's headlight are employed to calibrate the surveillance camera and estimate the vehicle's speed. Experimental results on typical sequences show that the proposed method can robustly detect, track, and pair the vehicle headlight in night scenes. Extensive quantitative evaluations and related comparisons demonstrate that the proposed method outperforms state-of-the-art methods. Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001, Xinge You |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2010 | Stereo matching algorithm based on curvelet decomposition and modified support weightsabstractWe present a novel multiresolution analysis based stereo matching method using curvelets and modified adaptive support weight. Multiresolution analysis has long been applied to stereo correspondence. However, previous methods suffer from false matches arising from textureless region or repetitive textures and fattening effect due to area based matching. In the proposed approach, we have reduced false matches by using curvelet coefficients in different scales and orientations. Curvelet coefficients can uniquely represent different image points and increase matching accuracy. The fattening effect is reduced using support weights modified for curvelets. The proposed method is verified and compared with state-of-the art methods by extensive tests, and good results are obtained. Dibyendu Mukherjee, Guanghui Wang 0001, Q. M. Jonathan Wu |
ICASSP | 2 |
| 2010 | Quasi-perspective Projection Model: Theory and Application to Structure and Motion Factorization from Uncalibrated Image Sequences
Guanghui Wang 0001, Q. M. Jonathan Wu |
Int. J. Comput. Vis. | 1 |
| 2010 | Image matching using enclosed region detector
Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001, Xinge You |
J. Vis. Commun. Image Represent. | 3 |
| 2010 | The quasi-perspective model: Geometric properties and 3D reconstruction
Guanghui Wang 0001, Q. M. Jonathan Wu |
Pattern Recognit. | 1 |
| 2010 | An Adaptive Computational Model for Salient Object DetectionabstractSalient object detection is a basic technique for many computer vision applications. In this paper, we propose an adaptive computational model to detect the salient object in color images. Firstly, three human observation behaviors and scalable subtractive clustering techniques are used to construct attention Gaussian mixture model (AGMM) and background Gaussian mixture model (BGMM). Secondly, the Bayesian framework is employed to classify each pixel into salient object or background object. Thirdly, expectation-maximization (EM) algorithm is utilized to update the parameters of AGMM, BGMM, and Bayesian framework based on the detection results. Finally, the classification and update procedures are repeated until the detection results evolve to a steady state. Experiments on a variety of images demonstrate the robustness of the proposed method. Extensive quantitative evaluations and comparisons demonstrate that the proposed method significantly outperforms state-of-the-art methods. Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001, Hai Bing Yin |
IEEE Trans. Multim. | 3 |
| 2009 | Two-View Geometry and Reconstruction under Quasi-perspective Projection
Guanghui Wang 0001, Q. M. Jonathan Wu |
ACCV (2) | 1 |
| 2009 | Vehicle Headlights Detection Using Markov Random Fields
Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001 |
ACCV (1) | 3 |
| 2009 | What can we learn about the scene structure from three orthogonal vanishing points in images
Guanghui Wang 0001, Hung-Tat Tsui, Q. M. Jonathan Wu |
Pattern Recognit. Lett. | 1 |
| 2009 | Perspective 3-D Euclidean Reconstruction With Varying Camera ParametersabstractThe paper addresses the problem of 3-D Euclidean structure and motion recovery from video sequences based on perspective factorization. It is well known that projective depth recovery and camera calibration are two essential and difficult steps in metric reconstruction. We focus on the difficulties and propose two new algorithms to improve the performance of perspective factorization. First, we propose to initialize the projective depths via a projective structure reconstructed from two views with large camera movement, and optimize the depths iteratively by minimizing reprojection residues. The algorithm is more accurate than previous methods and converges quickly. Second, we propose a self-calibration method based on the Kruppa constraint to deal with more general camera model. The Euclidean structure can be recovered from factorization of the normalized tracking matrix. Extensive experiments on synthetic data and real sequences are performed to validate the proposed method and good improvements are observed. Guanghui Wang 0001, Q. M. Jonathan Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Quasi-perspective projection with applications to 3D factorization from uncalibrated image sequencesabstractThe paper addresses the problem of factorization-based 3D reconstruction from uncalibrated image sequences. We propose a quasi-perspective projection model and apply the model to structure and motion recovery of rigid and nonrigid objects based on factorization of tracking matrix. The novelty and contribution of the paper lies in three aspects. First, under the assumption that the camera is far away from the object with small rotations, we propose and prove that the imaging process can be modeled by quasi-perspective projection. The model is more accurate than affine since the projective depths are implicitly embedded. Second, we apply the model to the factorization algorithm and establish the framework of rigid and nonrigid factorization under quasi-perspective assumption. Third, we propose a new and robust method to recover the transformation matrix that upgrades the factorization to the Euclidean space. The proposed method is validated and evaluated on synthetic and real image sequences and good improvements over existing solutions are observed. Guanghui Wang 0001, Q. M. Jonathan Wu |
CVPR | 1 |
| 2008 | Structure and motion factorization under quasi-perspective projection with missing data in tracking matrixabstractThe paper is focused on the problem of structure and motion factorization from uncalibrated image sequences. Based on our early study on quasi-perspective projection, we give an analysis on the imaging errors of different projection models and propose to adopt power factorization algorithm to deal with missing data problem. The main contribution lies in two aspects. First, we carry out an error analysis of the quasi-perspective projection and prove that it is more accurate than affine model under small camera movements. Second, we propose to utilize power factorization to factorize the tracking matrix. Compared with SVD-based method, the algorithm can work with incomplete tracking data, and it is computationally cheaper than other methods. The proposed method is evaluated on synthetic and real image sequences and better results are observed. Guanghui Wang 0001, Q. M. Jonathan Wu, Wei Huang 0029 |
ICPR | 1 |
| 2008 | Adaptive semantic Bayesian framework for image attentionabstractImage attention is the basic technique for many computer vision applications. In this paper, we propose an adaptive Bayesian framework to detect the image attention in color image. Firstly, three simple semantics and subtractive clustering are used to construct attention Gaussians mixture model (AGMM) and background Gaussians mixture model (BGMM). Secondly, the Bayesian framework is utilized to classify each pixel into attention objects and background objects. Thirdly, EM algorithm is used to update the parameters of AGMM, BGMM, and Bayesian framework according to the detection results. Finally, the above classification and update procedures are repeated until the detection results become steady. Experimental results on typical images exhibit the robustness of the proposed method. Wei Zhang 0025, Q. M. Jonathan Wu, Guanghui Wang 0001 |
ICPR | 3 |
| 2008 | Rotation constrained power factorization for structure from motion of nonrigid objects
Guanghui Wang 0001, Hung-Tat Tsui, Q. M. Jonathan Wu |
Pattern Recognit. Lett. | 1 |
| 2008 | Single view based pose estimation from circle or parallel lines
Guanghui Wang 0001, Q. M. Jonathan Wu, Zhengqiao Ji |
Pattern Recognit. Lett. | 1 |
| 2008 | Kruppa equation based camera calibration from homography induced by remote plane
Guanghui Wang 0001, Q. M. Jonathan Wu, Wei Zhang 0025 |
Pattern Recognit. Lett. | 1 |
| 2008 | Stratification Approach for 3-D Euclidean Reconstruction of Nonrigid Objects From Uncalibrated Image SequencesabstractThis paper addresses the problem of 3-D reconstruction of nonrigid objects from uncalibrated image sequences. Under the assumption of affine camera and that the nonrigid object is composed of a rigid part and a deformation part, we propose a stratification approach to recover the structure of nonrigid objects by first reconstructing the structure in affine space and then upgrading it to the Euclidean space. The novelty and main features of the method lies in several aspects. First, we propose a deformation weight constraint to the problem and prove the invariability between the recovered structure and shape bases under this constraint. The constraint was not observed by previous studies. Second, we propose a constrained power factorization algorithm to recover the deformation structure in affine space. The algorithm overcomes some limitations of a previous singular-value-decomposition-based method. It can even work with missing data in the tracking matrix. Third, we propose to separate the rigid features from the deformation ones in 3-D affine space, which makes the detection more accurate and robust. The stratification matrix is estimated from the rigid features, which may relax the influence of large tracking errors in the deformation part. Extensive experiments on synthetic data and real sequences validate the proposed method and show improvements over existing solutions. Guanghui Wang 0001, Q. M. Jonathan Wu |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2007 | Pose Estimation from Circle or Parallel Lines in a Single Image
Guanghui Wang 0001, Q. M. Jonathan Wu, Zhengqiao Ji |
ACCV (2) | 1 |
| 2007 | Structure and motion of nonrigid object under perspective projection
Guanghui Wang 0001, Hung-Tat Tsui, Zhanyi Hu |
Pattern Recognit. Lett. | 1 |
| 2006 | Euclidean reconstruction of a circular truncated cone only from its uncalibrated contours
Yihong Wu 0002, Guanghui Wang 0001, Fuchao Wu, Zhanyi Hu |
Image Vis. Comput. | 2 |
| 2005 | Single view metrology from scene constraints
Guanghui Wang 0001, Zhanyi Hu, Fuchao Wu, Hung-Tat Tsui |
Image Vis. Comput. | 1 |
| 2005 | Camera calibration and 3D reconstruction from a single view based on scene constraints
Guanghui Wang 0001, Hung-Tat Tsui, Zhanyi Hu, Fuchao Wu |
Image Vis. Comput. | 1 |
| 2005 | Reconstruction of structured scenes from two uncalibrated images
Guanghui Wang 0001, Hung-Tat Tsui, Zhanyi Hu |
Pattern Recognit. Lett. | 1 |
| 2004 | Single View Based Measurement on Space Planes
Guanghui Wang 0001, Zhanyi Hu, Fuchao Wu |
J. Comput. Sci. Technol. | 1 |
| 2003 | The impossibility of affine reconstruction from perspective image pairs obtained by a translating camera with varying parameters
Zhanyi Hu, Fuchao Wu, Guanghui Wang 0001 |
Pattern Recognit. Lett. | 3 |