Zongqing Lu 0001

dblp:99/965-1 · DBLP profile ↗
← Back
42ranked-venue papers
3as first author
24since 2021 · last 2025
0000-0002-1191-9069ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 UV Gaussians: Joint learning of mesh deformation and Gaussian textures for human avatar modeling
Yujiao Jiang, Qingmin Liao, Xiaoyu Li 0002, Qi Zhang 0029, Chaopeng Zhang, Zongqing Lu 0001, Ying Shan
Knowl. Based Syst.7
2024 SMPLX-Lite: A Realistic and Drivable Avatar Benchmark with Rich Geometry and Texture Annotations
abstract
Recovering photorealistic and drivable full-body avatars is crucial for numerous applications, including virtual reality, 3D games, and tele-presence. Most methods, whether reconstruction or generation, require large numbers of human motion sequences and corresponding textured meshes. To easily learn a drivable avatar, a reasonable parametric body model with unified topology is paramount. However, existing human body datasets either have images or textured models and lack parametric models which fit clothes well. We propose a new parametric model SMPLX-Lite-D, which can fit detailed geometry of the scanned mesh while maintaining stable geometry in the face, hand and foot regions. We present SMPLX-Lite dataset, the most comprehensive clothing avatar dataset with multi-view RGB sequences, keypoints annotations, textured scanned meshes, and textured SMPLX-Lite-D models. With the SMPLX-Lite dataset, we train a conditional variational autoencoder model that takes human pose and facial keypoints as input, and generates a photorealistic drivable human avatar.
Yujiao Jiang, Qingmin Liao, Xiangru Lin, Zongqing Lu 0001, Yuxi Zhao, Hanqing Wei, Jingrui Ye, Yu Zhang 0166, Zhijing Shao
ICME5
2024 FusionDreamer: Consistent Images Generation from Sparse-view Images
abstract
We introduce FusionDreamer, a diffusion-based approach that can generate high-quality multiview-consistent images from any number of input images. Recent diffusion-based methods like Zero123 and SyncDreamer showcase the capability to generate novel views from a single input view through the utilization of pre-trained large-scale 2D diffusion models, but they are constrained by the restriction to a solitary view. This limitation results in the underutilization of valuable information in potential multiview scenarios, leading to challenges in accurately predicting intricate details and complex shapes. To address this limitation, we propose a Geometer-aware Pixel-wise Attention Mechanism, extending these methods to utilize multiview information when generating novel views. Experiments show that a simple extension of the single view generation model is not enough, and FusionDreamer can consider multiple views and obtain novel views with higher consistency and better quality.
Risheng Huang, Hao-Zhi Huang 0001, Zongqing Lu 0001
ICME4
2024 MetaMask: Improving Few-Shot Semantic Segmentation via Multi-Mask Calibriation
abstract
Few-shot Semantic Segmentation (FSS) aims to develop models that can segment previously unseen classes with only a few annotations. Recent approaches employ a "multi-mask" framework, which initially generates various mask proposals from query images and then matches related mask proposals to get the final output guided by support images. Despite its promise, this framework is limited by the quality of mask proposals for unseen classes and a naive mask matching process. To address such limitations, in this paper, we propose a meta-learning-based method called MetaMask. First, MetaMask builds a Support-Guided Latent Object Segmenter (SG-LOS) module, which incorporates unseen class information into mask proposal generation for query images, where episodic training is used to enhance mask generation for latent unseen classes. Second, MetaMask improves the mask-matching mechanism through our proposed Contrastive Mask Matching (CMM) module with a cross-image multi-level contrastive learning strategy, bolstering feature embedding spaces. Our method shows competitive results on two main benchmarks: 69.9% mIoU on Pascal-5ione-shot setting and 49.6% mIoU COCO-20ione-shot setting, marginally outperforming our baseline by 6.6% and 5.4%, setting a new state-of-the-art on the both Pascal-5iand COCO-20idatasets.
Li Dinghang, Zongqing Lu 0001, Weiliang Zheng, Qingmin Liao, Fan Lyu
IJCNN2
2024 Multi-Dimensional Attention on Cost Volume for Stereo Matching
abstract
Stereo matching is a fundamental research topic in computer vision tasks, and the careful processing of cost volume plays a vital role in stereo matching solutions. Previous convolutional networks have deep-layer structures but could only aggregate local regions, leading to suboptimal matching performance in areas with edges or weak textures, etc. Considering the global perception capability of the attention mechanism, we for the first time propose global attention modules directly operating on the cost volume for cost aggregation. Our proposed attention module is named Multi-Dimensional Attention (MDA) and it includes two submodules: the Cross-Disparity Attention (CDA) and the Intra-Disparity Attention (IDA). CDA accomplishes cost aggregation under different disparities, and IDA is further categorized into Channel-Wise Attention (CWA) and Disparity-Wise Attention (DWA), focusing on the similarity of structure and disparity variations within a fixed disparity. For evaluation, we conduct experiments on four publicly available datasets including KITTI 2012, KITTI 2015, Scene Flow and Middlebury, and results show that our proposed method achieves state-of-the-art (SoTA) performance in stereo matching tasks.
Zhou Jiale, Wenqin Huang, Qingmin Liao, Zongqing Lu 0001
IJCNN4
2024 Cross-Patch Relation Enhanced for Weakly Supervised Semantic Segmentation
abstract
Weakly Supervised Semantic Segmentation (WSSS) using only image-level labels relies on Class Activation Map (CAM) to produce pixel-level pseudo segmentation labels, but it struggles with limited object region activation, resulting in low-quality annotations. To address this issue, a local-to-global framework is employed to enable the model to capture details from patches randomly cropped from input images. However, the pseudo-masks generated by this approach still have an issue with object incompleteness. We notice that it is caused by the neglect of semantic relations among patches, which capture abundant contextual information. Under this observation, we present a Cross-Patch Relation Enhanced Network to improve the quality of the CAMs, leading to the generation of better pseudo segmentation labels. Specifically, a cross-patch relation attention (including the class-prototype extraction and the class-feature aggregation) is proposed to alleviate the intra-class inconsistency due to variations of contextual information across local patches. The class-prototype extraction module gathers contextual relation from all local class-region embeddings. Besides, class-feature aggregation improves class-level representations of multiple patches through feature aggregation. Extensive experimental results on two public datasets have demonstrated the effectiveness of the proposed method. Our method achieves competitive scores with state-of-the-art methods for weakly supervised semantic segmentation on both PASCAL VOC 2012 and MS-COCO 2014 benchmarks.
Huiqing Su, Wenqin Huang, Qingmin Liao, Zongqing Lu 0001
IJCNN4
2024 AdaPKC: PeakConv with Adaptive Peak Receptive Field for Radar Semantic Segmentation
abstract
Deep learning-based radar detection technology is receiving increasing attention in areas such as autonomous driving, UAV surveillance, and marine monitoring. Among recent efforts, PeakConv (PKC) provides a solution that can retain the peak response characteristics of radar signals and play the characteristics of deep convolution, thereby improving the effect of radar semantic segmentation (RSS). However, due to the use of a pre-set fixed peak receptive field sampling rule, PKC still has limitations in dealing with problems such as inconsistency of target frequency domain response broadening, non-homogeneous and time-varying characteristic of noise/clutter distribution. Therefore, this paper proposes an idea of adaptive peak receptive field, and upgrades PKC to AdaPKC based on this idea. Beyond that, a novel fine-tuning technology to further boost the performance of AdaPKC-based RSS networks is presented. Through experimental verification using various real-measured radar data (including publicly available low-cost millimeter-wave radar dataset for autonomous driving and self-collected Ku-band surveillance radar dataset), we found that the performance of AdaPKC-based models surpasses other SoTA methods in RSS tasks. The code is available at https://github.com/lihua199710/AdaPKC.
Youcheng Zhang, ZijunHu, Pengcheng Pi, Zongqing Lu 0001, Qingmin Liao
NeurIPS6
2024 Dual Correlation Network for Efficient Video Semantic Segmentation
abstract
Video data bring a big challenge to semantic segmentation due to the large volume of data and strong inter-frame redundancy. In this paper, we propose a dual local and global correlation network tailored for efficient video semantic segmentation. It consists of three modules: 1) a local attention based module, which measures correlation and achieves feature aggregation in a local region between key frame and non-key frame; 2) a consistent constraint module, which considers long-range correlation among pixels from a global view for promoting intra-frame semantic consistency of non-key frame; and 3) a key frame decision module, which selects key frames adaptively based on the ability of feature transferring. Extensive experiments on the Cityscapes and Camvid video datasets demonstrate that our proposed method could reduce inference time significantly while maintaining high accuracy. The implementation is available at https://github.com/An01168/DCNVSS.
Shumin An, Qingmin Liao, Zongqing Lu 0001, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.3
2023 Trans-Cycle: Unpaired Image-to-Image Translation Network by Transformer
Mengze Pan, Zongqing Lu 0001, Qingmin Liao
ICANN (6)3
2023 LLA-Flow: A Lightweight Local Aggregation on Cost Volume for Optical Flow Estimation
abstract
Lack of texture often causes ambiguity in matching, and handling this issue is an important challenge in optical flow estimation. Some methods insert stacked transformer modules that allow the network to use global information of cost volume for estimation. But the global information aggregation often incurs serious memory and time costs during training and inference, which hinders model deployment. We draw inspiration from the traditional local region constraint and design the local similarity aggregation (LSA) and the shifted local similarity aggregation (SLSA). The aggregation for cost volume is implemented with lightweight modules that act on the feature maps. Experiments on the final pass of Sintel show the lower cost required for our approach while maintaining competitive performance.
Zongqing Lu 0001, Qingmin Liao
ICIP2
2023 Patchmatch Stereo++: Patchmatch Binocular Stereo with Continuous Disparity Optimization
abstract
Current deep-learning-based stereo matching algorithms achieve remarkably low error rates but they suffer from the edge ambiguity effect. The primary reason is that they treat disparity estimation as a labeling problem, constructing a cost volume based on uniform discrete pixel-wise labels. It is insufficient to model the continuous disparity probability distribution (DPD), which harms the accuracy of complex regions. Moreover, current cost aggregation strategies cannot process unstructured disparity candidates very well, which is one of the bottlenecks limiting continuous modeling. We propose Patchmatch Stereo++, inspired by the traditional Patchmatch Stereo to achieve better continuous disparity optimization in deep-learning-based methods. Firstly, to model accurate continuous DPD, we introduce an adaptive dense sub-pixel sampling strategy to binocular stereo and approximate a continuous unstructured DPD for every pixel. Secondly, we design a convolution-based optimizer that can accept unstructured disparity candidates to parse the above continuous DPD in an adaptive manner and perform updates accordingly. Extensive experiments demonstrate our method has the best performance among existing stereo matching networks at the edges, both quantitatively and qualitatively. At the time of submission, compared with published works pre-trained on SceneFlow, we rank 1st in the foreground of KITTI and 2nd on SceneFlow, ETH3D under various metrics.The source code will be released.
Wenjia Ren, Qingmin Liao, Zhijing Shao, Xiangru Lin, Xin Yue, Yu Zhang 0166, Zongqing Lu 0001
ACM Multimedia7
2023 APANet: Adaptive Prototypes Alignment Network for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation aims to segment novel-class objects in a given query image with only a few labeled support images. Most advanced solutions exploit a metric learning framework that performs segmentation through matching each query feature to a learned class-specific prototype. However, this framework suffers from biased classification due to incomplete feature comparisons. To address this issue, we present an adaptive prototype representation by introducing class-specific and class-agnostic prototypes and thus construct complete sample pairs for learning semantic alignment with query features. The complementary features learning manner effectively enriches feature comparison and helps yield an unbiased segmentation model in the few-shot setting. It is implemented with a two-branch end-to-end network (i.e., a class-specific branch and a class-agnostic branch), which generates prototypes and then combines query features to perform comparisons. In addition, the proposed class-agnostic branch is simple yet effective. In practice, it can adaptively generate multiple class-agnostic prototypes for query images and learn feature alignment in a self-contrastive manner. Extensive experiments on PASCAL-5$^{i}$and COCO-20$^{i}$demonstrate the superiority of our method. At no expense of inference efficiency, our model achieves state-of-the-art results in both 1-shot and 5-shot settings for semantic segmentation.
Bin-Bin Gao, Zongqing Lu 0001, Jing-Hao Xue, Chengjie Wang 0001, Qingmin Liao
IEEE Trans. Multim.3
2022 Efficient Semantic Segmentation via Self-Attention and Self-Distillation
abstract
Lightweight models are pivotal in efficient semantic segmentation, but they often suffer from insufficient context information due to limited convolution and small receptive field. To address this problem, we propose a tailored approach to efficient semantic segmentation by leveraging two complementary distillation schemes for supplementing context information to small networks: 1) a self-attention distillation scheme, which transfers long-range context knowledge adaptively from large teacher networks to small student networks; and 2) a layer-wise context distillation scheme, which transfers structured context from deep layers to shallow layers within student networks for promoting semantic consistency of the shallow layers. Extensive experiments on the ADE20K, Cityscapes, and Camvid datasets well demonstrate the effectiveness of our proposal.
Shumin An, Qingmin Liao, Zongqing Lu 0001, Jing-Hao Xue
IEEE Trans. Intell. Transp. Syst.3
2022 Deep Learning in Lane Marking Detection: A Survey
abstract
Lane marking detection is a fundamental but crucial step in intelligent driving systems. It can not only provide relevant road condition information to prevent lane departure but also assist vehicle positioning and forehead car detection. However, lane marking detection faces many challenges, including extreme lighting, missing lane markings, and obstacle obstructions. Recently, deep learning-based algorithms draw much attention in intelligent driving society because of their excellent performance. In this paper, we review deep learning methods for lane marking detection, focusing on their network structures and optimization objectives, the two key determinants of their success. Besides, we summarize existing lane-related datasets, evaluation criteria, and common data processing techniques. We also compare the detection performance and running time of various methods, and conclude with some current challenges and future trends for deep learning-based lane marking detection algorithm.
Youcheng Zhang, Zongqing Lu 0001, Xuechen Zhang 0003, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Intell. Transp. Syst.2
2021 Lane Line Detection based on Parallel Spatial Separation Convolution
Xile Shen, Zongqing Lu 0001, Youcheng Zhang, Jing-Hao Xue
BMVC2
2021 Parallax Contextual Representations For Stereo Matching
abstract
In this work, we study the context aggregation in stereo matching from a new parallax perspective. Unlike previous works, we propose to characterize and augment a pixel with its parallax contextual representation (PCR), which has not been explored before. We also propose a new concept called disparity prototype to describe the overall representation of a disparity plane. Our proposed PCR module consists of three steps: 1) divide disparity planes for a rough estimation of disparity; 2) estimate the disparity prototypes for each disparity plane; 3) derive PCR-augmented representations with disparity prototypes. Extensive experiments on various datasets using different networks validate the effectiveness of our proposal.
Qingmin Liao, Zongqing Lu 0001, Jing-Hao Xue
ICIP3
2021 A Region-Based Descriptor Network for Uniformly Sampled Keypoints
abstract
Matching keypoint pairs of different images is a basic task of computer vision. Most methods require customized extremum point schemes to obtain the coordinates of feature points with high confidence, which often need complex algorithmic design or a network with higher training difficulty and also ignore the possibility that flat regions can be used as candidate regions of matching points. In this paper, we design a region-based descriptor by combining the context features of a deep network. The new descriptor can give a robust representation of a point even in flat regions. By the new descriptor, we can obtain more high confidence matching points without extremum operation. The experimental results show that our proposed method achieves a performance comparable to state-of-the-art.
Zongqing Lu 0001, Qingmin Liao
ICIP2
2021 CLUSAC: Clustering Sample Consensus for Fundamental Matrix Estimation
abstract
In the process of model fitting for fundamental matrix estimation, RANSAC and its variants disregard and fail to reduce the interference of outliers. These methods select correspondences and calculate the model scores from the original dataset. In this work, we propose an inlier filtering method that can filter inliers from the original dataset. Using the filtered inliers can substantially reduce the interference of outliers. Based on the filtered inliers, we propose a new algorithm called CLUSAC, which calculates model quality scores on all filtered inliers. Our approach is evaluated through estimating the fundamental matrix in the dataset kusvod2, and it shows superior performance to other compared RANSAC variants in terms of precision.
Xuanyu Xiao, Zongqing Lu 0001, Jing-Hao Xue
ICIP2
2021 Disparity Estimation with Scene Depth Cues
abstract
The cost volume plays a pivotal role in stereo matching, usually working as an optimization object. However, we find it also can provide effective scene prior to guide the disparity learning, as it reflects well the depth relationship between scenario objects. Inspired by this new perspective, we propose the CSA module, which consists of a new correlation and selection (CS) layer and a new aggregation layer. The CS layer can regulate the matching costs and re-encode the feature information into the correlation volume. The aggregation layer can preserve better the depth cues of the refined cost volume, through a convolution network and a unimodalization operation. The proposed module can be trained in a supervised manner, making the extraction of scene depth cues more accurate. Extensive experiments on the Sceneflow and KITTI datasets have demonstrated that with our module embedded, SOTA networks can achieve substantially better performance.
Zongqing Lu 0001, Qingmin Liao, Jing-Hao Xue
ICME2
2021 Better Stereo Matching From Simple Yet Effective Wrangling of Deep Features
abstract
Cost volume plays a pivotal role in stereo matching. Most recent works focused on deep feature extraction and cost refinement for a more accurate cost volume. Unlike them, we probe from a different perspective: feature wrangling. We find that simple wrangling of deep features can effectively improve the construction of cost volume and thus the performance of stereo matching. Specifically, we develop two simple yet effective wrangling techniques of deep features, spatially a differentiable feature transformation and channel-wise a memory-economical feature expansion, for better cost construction. Exploiting the local ordering information provided by a differentiable rank transform, we achieve an enhancement of the search for correspondence; with the help of disparity division, our feature expansion allows for more features into the cost volume with no extra memory required. Equipped with these two feature wrangling techniques, our simple network can perform outstandingly on the widely used KITTI and Sceneflow datasets.
Zongqing Lu 0001, Qingmin Liao, Jing-Hao Xue
ICME2
2021 How Video Super-Resolution and Frame Interpolation Mutually Benefit
abstract
Video super-resolution (VSR) and video frame interpolation (VFI) are inter-dependent for enhancing videos of low resolution and low frame rate. However, most studies treat VSR and temporal VFI as independent tasks. In this work, we design a spatial-temporal super-resolution network based on exploring the interaction between VSR and VFI. The main idea is to improve the middle frame of VFI by the super-resolution (SR) frames and feature maps from VSR. In the meantime, VFI also provides extra information for VSR and thus, through interacting, the SR of consecutive frames of the original video can also be improved by the feedback from the generated middle frame. Drawing on this, our approach leverages a simple interaction of VSR and VFI and achieves state-of-the-art performance on various datasets. Due to such a simple strategy, our approach is universally applicable to any existing VSR or VFI networks for effectively improving their video enhancement performance.
Chengcheng Zhou, Zongqing Lu 0001, Linge Li, Qiangyu Yan, Jing-Hao Xue
ACM Multimedia2
2021 ReFlowNet: Revisiting Coarse-to-fine Learning of Optical Flow
Leyang Xu, Zongqing Lu 0001
PRCV (1)2
2021 Guest Editorial: Special issue on deep learning with small samples
Jing-Hao Xue, Jufeng Yang, Yan Yan 0001, Yujiu Yang 0001, Zongqing Lu 0001, Zhanyu Ma
Neurocomputing6
2021 Ripple-GAN: Lane Line Detection With Ripple Lane Line Detection Network and Wasserstein GAN
abstract
With artificial intelligence technology being advanced by leaps and bounds, intelligent driving has attracted a huge amount of attention recently in research and development. In intelligent driving, lane line detection is a fundamental but challenging task particularly under complex road conditions. In this paper, we propose a simple yet appealing network called Ripple Lane Line Detection Network (RiLLD-Net), to exploit quick connections and gradient maps for effective learning of lane line features. RiLLD-Net can handle most common scenes of lane line detection. Then, in order to address challenging scenarios such as occluded or complex lane lines, we propose a more powerful network called Ripple-GAN, by integrating RiLLD-Net, confrontation training of Wasserstein generative adversarial networks, and multi-target semantic segmentation. Experiments show that, especially for complex or obscured lane lines, Ripple-GAN can produce a superior detection performance to other state-of-the-art methods.
Youcheng Zhang, Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Intell. Transp. Syst.2
2020 Noise-Sampling Cross Entropy Loss: Improving Disparity Regression Via Cost Volume Aware Regularizer
abstract
Recent end-to-end deep neural networks for disparity regression have achieved the state-of-the-art performance. However, many well-acknowledged specific properties of disparity estimation are omitted in these deep learning algorithms. Especially, matching cost volume, one of the most important procedure, is treated as a normal intermediate feature for the following softargmin regression, lacking explicit constraints compared with those traditional algorithms. In this paper, inspired by previous canonical definition of cost volume, we propose the noise-sampling cross entropy loss function to regularize the cost volume produced by deep neural networks to be unimodal and coherent. Extensive experiments validate that the proposed noise-sampling cross entropy loss can not only help neural networks learn more informative cost volume, but also lead to better stereo matching performance compared with several representative algorithms.
Zongqing Lu 0001, Xuechen Zhang 0003, Qingmin Liao
ICIP2
2019 Exploring Discriminative Features in Mueller Matrix Images for Electrospinning Classification
abstract
Polarization images, which are captured in lights with different polarization angles, can extract more detail information about samples. Generally, there are two ways to make use of polarization images: direct processing of original polarization images and processing of Mueller matrix (MM) images. Since MM has clear physical meaning and each element in it represents a specific characteristic about samples, studying the relationship between the elements in MM and samples is a meaningful topic. In this paper, an importance sorting algorithm is proposed to explore discriminative elements in MM. Firstly, a linear weighted feature fusion method is proposed and three distances are defined to form the target function. Then, a convex quadratic programming model is built, with an algorithm to search the optimal solution. Finally, discriminative elements are choosed for classification according to the optimal weight vector. Experiments conducted on an electrospinning dataset show that the proposed method not only provides a consistent importance order of elements in MM, but also helps to find discriminative feature combinations for classification, which is useful for explaining of the polarization characteristics of samples. The source code is available at: https://github.com/madd2014/ImportanceSort.
Zongqing Lu 0001, Youcheng Zhang, Qingmin Liao
ICIP2
2019 Estimating Human Shape Under Clothing from Single Frontal View Point Cloud of a Dressed Human
abstract
Estimating human shape under clothing is a challenging task. We propose the first method to estimate accurate shape parameters from single-frame frontal view point cloud. To account for casual clothing, we personalize the original SMPL model to describe clothing as deviation from naked human parametric model, define a novel method to search for corresponding vertex pairs, and design a novel objective function that enforces point cloud vertices to remain outside of the naked body shape and tightly cling the personalized shape. Consolidating these three parts, our method integrates the advantages of free deformation method and model-based method. Our method is more effective than previous works in dealing with casual clothing situation. We evaluate the accuracy of estimated shape on noisy point cloud data captured by a commodity depth sensor.
Zongqing Lu 0001, Qingmin Liao
ICIP2
2019 A New Object Scene Flow Algorithm Based on Support Points Selection and Robust Moving Object Proposal
abstract
Recent algorithms of object scene flow estimation suffer from low computational efficiency or unstable moving object proposals. To tackle these two problems simultaneously, in this paper we propose a new, efficient and robust algorithm for object scene flow estimation, through making two technical contributions. Firstly to improve the efficiency, we propose to select only a few pixels termed support points for matching cost calculation rather than using all pixels. The support points are defined as those pixels with high confidence in feature matching. Secondly to attain stable moving object proposals, we propose a motion magnitude-adaptive thresholding scheme for ego-motion outlier detection, after patch matching on CNN-extracted high quality features. These two contributions, though simple, ensure a remarkable improvement in both efficiency and accuracy from the original object scene flow method, as well as making the proposed algorithm a strong practicable alternative to much more sophisticated state-of-the-art competitors.
Zhengyang Sun, Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
ICME2
2019 A New Approach to Automatic Clothing Matting from Mannequins
abstract
It is crucial to extract retail clothes from images of mannequins when building a database of clothing images for virtual try-on systems. However, clothes often have complex texture and translucent material, such as holes and laces. It is thus difficult to extract clothes as foreground by existing generic natural image matting methods. Hence in this paper, we present a novel approach to automatic clothing matting from mannequins, with auxiliary information from a rough background image of the mannequin only. Experiments show that we can achieve remarkable improvement on the alpha matte near challenging regions of complex texture and translucent material of clothes. Moreover, our approach can automatically generate trimaps to facilitate the development and evaluation of other image matting algorithms.
Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
ICME2
2019 A New Rotation-Invariant Deep Network for 3D Object Recognition
abstract
When inputs are rotated, most 3D convolutional neural networks (CNNs) will have their performance much dropped, especially for those models with voxelized input of 3D objects. The newly proposed Spherical CNNS, with the concept of the rotation-equivariant spherical correlation, aims to achieve rotation invariance. Inspired by this, we propose a new rotation-invariant deep network to recognize rotated 3D objects. Specifically, we adopt the spherical representation and the spherical correlation S^2 layer of Spherical CNNs, for their capacity of representing 3D objects and rotation equivariance. In the meantime, we improve the computational efficiency and expressiveness of Spherical CNNs, by replacing its time-consuming and depth-limited SO(3) layer with a PointNet-style network architecture. Hence our proposed network can maintain the equivariance as the network grows deeper while substantially reducing its runtime, leading to a much better efficiency and expressiveness of rotation-invariant representation. Experimental results show that our network performs better than or comparable to the state-of-the-art methods in the ModelNet40 classification challenge.
Yachi Zhang, Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
ICME2
2018 Binarized features with discriminant manifold filters for robust single-sample face recognition
Wanping Zhang, Zongqing Lu 0001, Weifeng Li 0001, Qingmin Liao
Signal Process. Image Commun.4
2017 Wavelet-based single image super-resolution with an overall enhancement procedure
abstract
In this paper, we address the problem of generating a super-resolution image based on a dictionary of low- and high-resolution exemplars from a single input image in wavelet domain with a overall enhancement procedure. Most methods extract different kinds of features in low-resolution image and high-resolution images to establish the mapping relation. But in this paper, we implement wavelet-transform to extract the same kind of feature to make the mapping more reasonable. Meanwhile we implement local Lipschitz regularity constraint and structure-keeping constraint to preserve the local singularity and edge in our method. Compared with current state-of-art methods on standard images, our method obtains both visual and PSNR improvement.
Zongqing Lu 0001, Quan Zou 0001, Fei Zhou 0001, Qingmin Liao
ICASSP1
2017 Locality Sensitive Hashing based deepmatching for optical flow estimation
abstract
DeepMatching (DM) is one of the state-of-art matching algorithms to compute quasi-dense correspondences between images. Recent optical flow methods use DeepMatching to find initial image correspondences and achieves outstanding performance. However, the key building block of DeepMatching, the correlation map computation, is time-consuming. In this paper, we propose a new algorithm, LSHDM, which addresses the problem by employing Locality Sensitive Hashing (LSH) to DeepMatching. The computational complexity is greatly reduced for the correlation map computation step. Experiments show that image matching can be accelerated by our approach in ten times or more compared to DeepMatching, while retaining comparable accuracy for optical flow estimation.
Zongqing Lu 0001, Qingmin Liao, Danyi Li
ICASSP2
2017 Face recognition via weighted sparse representation using metric learning
abstract
Face recognition methods utilizing Sparse Representation based Classification (SRC) and Collaborative Representation based Classification (CRC) have recently attracted a great deal of attention due to inherent simplicity and efficiency. In this paper, we introduce the Large Margin Nearest Neighbor (LMNN), which learns a Mahalanobis distance metric that is applied, to SRC and CRC as the locality constraint. Next, a locality LMNN Weighted Sparse Representation based Classification (LMNN-WSRC) and a locality LMNN Weighted Collaborative Representation based Classification (LMNN-WCRC) are proposed. Our methods utilize both linearity and data locality. For a query face image, our target is to exploit the appropriate distance metric as the locality constraint that could focus more on those truly related images in the code book. Experimental results on the Extended Yale B database and the AR database show that our methods are more effective than SRC, Weighted SRC (WSRC) and CRC.
Zongqing Lu 0001, Bokun Xu, Qingmin Liao
ICME1
2017 Weighted contourlet binary patterns and image-based fisher linear discriminant for face recognition
Weifeng Li 0001, Yinyan Jiang, Zongqing Lu 0001, Qingmin Liao
Neurocomputing5
2016 Visual enhancement using sparsity-based image decomposition for low backlight displays
abstract
We propose a power-constrained image enhancement system to maintain human visual perception when the LCD or LED display is under low backlight. Adopting the low backlight mode can save the electricity and lengthen the battery using time. First, we deduce the relationship between the image and the backlight for maintaining the same visual perceptual quality. Then, we propose a sparsity-based image decomposition to separate the intensity image into base layer and detail layer. Afterwards, we refer to the image-backlight relationship to compensate the base layer, while we also adopt texture-aw are boosting to enhance the detail layer. Experimental simulated results show that our system outperforms than the compared systems.
Chih-Tsung Shen, Zongqing Lu 0001, Yi-Ping Hung, Soo-Chang Pei
ISCAS2
2015 Texture classification using uniform rotation invariant gradient
abstract
In this paper, we present a novel descriptor called uniform rotation invariant gradient(URIG) aiming at texture classification under variant rotation and illumination condition. Instead of using URIG directly, a 2D descriptor can be formulated combining URIG with average of local pixels. Given a texture image, such 2D descriptors are extracted from every pixel followed by clustering. The centers of clustering can be viewed as a texton dictionary over which a histogram is computed as the representation of given texture image. Experiments are carried out on Outex and CUReT databases comparing to state-of-the-art approaches. Our proposed method achieved promising performance against illumination and rotation changes with least cost for representing histogram dimension.
Wenteng Zhao, Zongqing Lu 0001, Qingmin Liao
ICIP2
2014 Vanishing point estimation for challenging road images
abstract
In this paper, we present an efficient vanishing point detection method for challenging road images. This detection process is based on the geometrical features of the roads. The slope distribution of the line segments is analyzed to reduce the spurious lines. A distance-based weighting scheme is also utilized to eliminate the voting noise in the voting stage. The proposed algorithm has been tested on a natural data set from Defense Advanced Research Projects Agency (DARPA). Experimental results with both quantitative and qualitative analyses are provided, which demonstrate the superiority of the proposed method over some state-of-the-art methods.
Qingyun She, Zongqing Lu 0001, Qingmin Liao
ICIP2
2014 Local texture based optical flow for complex brightness variations
abstract
In real-world scenarios, complex brightness variations are commonly seen, due to shadows, global illumination changes and nonlinear camera responses, etc. Classical optical flow methods based on brightness or gradient constancy assumption tends to fail under these circumstances. This work proposes an image texture descriptor called LSOT, based on the local spatial structure of a pixel and the relative ordinal information. Then a texture constancy assumption is embedded into a variational optical flow estimation framework as a data term, in order to cope with complex brightness variations. In addition, a non-local regularization term is used to improve the accuracy of the obtained flow fields. The energy functional is optimized using a primal-dual algorithm in a coarse-to-fine warping fashion. Experimental results on synthetic and real image sequences demonstrate the superior performance of the proposed method.
Zongqing Lu 0001, Qingmin Liao
ICIP2
2014 Generalized Weber-face for illumination-robust face recognition
Yong Wu 0003, Yinyan Jiang, Yicong Zhou, Weifeng Li 0001, Zongqing Lu 0001, Qingmin Liao
Neurocomputing5
2009 An Effective Method for Foreground Segmentation of Video
abstract
In this paper, we propose a novel foreground segmentation approach for applications using static cameras. The foreground segmentation is modeled as an energy function optimum process, where energy function is based on Markov Random Field (MRF) and efficiently optimized by Gibbs sampling. The essence of our method is that we fuse four foreground/background models based on color and texture. This allows composing a robust likelihood term that not only reflects the appearance of foreground/background, but also models the shadow removal process, together with a spatial contrast term and a better temporal persistence term, which achieves a more accurate segmentation. This method has been run on both indoor and outdoor sequences, and the results have proved its effectiveness.
Jianfeng Shen, Zongqing Lu 0001, Qingmin Liao
ICIG2
2009 A variational approach to automatic segmentation of RNFL on OCT data sets of the retina
abstract
Optical coherence tomography (OCT) as a new imaging technology is gaining popularity in the diagnosis of ocular diseases. It enable clinicians to perform accurate, objective, and reproducible measurements of the retinal nerve fibre layer (RNFL) whose thickness is closely related to many ocular diseases. Automatic segmenting RNFL is a challenging image processing problem, which is a critical job for final thickness estimation. We modeled the OCT data sets as probability density fields and introduced a level set model to outline the RNFL region within the retina. We also introduced the symmetrized Kullback-Leibler distance to describe the difference of two density functions. The new approach can deal with the typical problems of OCT image analysis: speckle noise and faint structure in an efficient way.
Zongqing Lu 0001, Qingmin Liao
ICIP1