VLDB 2026 Research / reviewers in the wild / expert
Lu Yang 0006
dblp:58/2893-6
· DBLP profile ↗
33ranked-venue papers
13as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 17 · 8 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Position Encoding Mechanism in Diffusion U-Net for Training-free High-resolution Image GenerationabstractDenoising higher-resolution latents using a pre-trained U-Net often results in repetitive and disordered image patterns. In this work, we are motivated to reveal the intrinsic cause of such pattern disruption in high-resolution image generation. Through theoretical analysis and empirical studies, we reveal that the pre-trained U-Net fails to provide sufficient positional information for tokens at high-resolution. Specifically, 1) zero-padding serves as a critical mechanism for position encoding but lacks robustness across varying resolutions; and 2) tokens located farther from the feature map boundaries have increasing difficulty acquiring positional awareness, leading to pattern disruptions. Inspired by these findings, we propose a novel training-free approach for high-resolution generation, introducing a Progressive Boundary Complement (PBC) method. It creates dynamic virtual image boundaries inside the feature map to supplement position information at high resolution, enabling high-quality and rich-content high-resolution image synthesis. Extensive experiments show that our method significantly improves high-resolution image synthesis in terms of visual quality and content richness, achieving state-of-the-art performance. Pu Cao, Yiyang Ma, Lu Yang 0006, Yonghao Dang, Jianqin Yin |
AAAI | 4 |
| 2026 | Controllable Generation With Text-to-Image Diffusion Models: A SurveyabstractIn the rapidly advancing realm of visual generation, diffusion models have revolutionized the landscape, marking a significant shift in capabilities with their impressive text-guided generative functions. However, relying solely on text for conditioning these models does not fully cater to the varied and complex requirements of different applications and scenarios. Acknowledging this shortfall, a variety of studies aim to control pre-trained text-to-image (T2I) models to support novel conditions. In this survey, we undertake a thorough review of the literature on controllable generation with T2I diffusion models, covering both the theoretical foundations and practical advancements in this domain. Our review begins with a brief introduction to the basics of denoising diffusion probabilistic models (DDPMs) and widely used T2I diffusion models. Additionally, we provide a detailed overview of research in this area, categorizing it from the condition perspective into three directions: generation with specific conditions, generation with multiple conditions, and universal controllable generation. For each category, we analyze the underlying control mechanisms and review representative methods based on their core techniques. Pu Cao, Qing Song 0006, Lu Yang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Large-Scale Omnidirectional Person Positioning
Lu Yang 0006, Liulei Li, Jianan Wei, Pu Cao, Wenguan Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | ParsingFormer for pixel-wise hierarchical human representation learning
Pu Cao, Zhixiang Lv, Junyi Ji, Shan Li 0001, Qing Song 0006, Lu Yang 0006 |
Pattern Recognit. | 6 |
| 2026 | Quality transformer for human parsing
Lu Yang 0006, Pu Cao, Shan Li 0001, Qing Song 0006 |
Pattern Recognit. | 2 |
| 2025 | Image is All You Need to Empower Large-scale Diffusion Models for In-Domain GenerationabstractIn-domain generation aims to perform a variety of tasks within a specific domain, such as unconditional generation, text-to-image, image editing, 3D generation, and more. Early research typically required training specialized generators for each unique task and domain, often relying on fully-labeled data. Motivated by the powerful generative capabilities and broad applications of diffusion models, we are driven to explore leveraging label-free data to empower these models for in-domain generation. Fine-tuning a pre-trained generative model on domain data is an intuitive but challenging way and often requires complex manual hyper-parameter adjustments since the limited diversity of the training data can easily disrupt the model’s original generative capabilities. To address this challenge, we propose a guidance-decoupled prior preservation mechanism to achieve high generative quality and controllability by image-only data, inspired by preserving the pre-trained model from a denoising guidance perspective. We decouple domain-related guidance from the conditional guidance used in classifier-free guidance mechanisms to preserve open-world control guidance and unconditional guidance from the pre-trained model. We further propose an efficient domain knowledge learning technique to train an additional text-free UNet copy to predict domain guidance. Besides, we theoretically illustrate a multi-guidance in-domain generation pipeline for a variety of generative tasks, leveraging multiple guidances from distinct diffusion models and conditions. Extensive experiments demonstrate the superiority of our method in domain-specific synthesis and its compatibility with various diffusion-based control methods and applications. Pu Cao, Lu Yang 0006, Tianrui Huang, Qing Song 0006 |
CVPR | 3 |
| 2025 | E4C: Enhance Editability for Text-Based Image Editing by Harnessing Efficient CLIP GuidanceabstractDiffusion-based image editing involves both preserving the source image content and generating new content or applying modifications. Although current editing approaches have made improvements under text guidance, they have two key drawbacks: overemphasis on retaining original image info, neglecting editability and text alignment, and inability to handle both structure-consistent and non-rigid editing tasks. In this paper, we propose a zero-shot image editing method, named Enhance Editability for text-based image Editing via Efficient CLIP guidance (E4C), which presents an innovative adaptive feature sharing mechanism to enable multi-task editing. Additionally, a novel random gateway mechanism is designed to efficiently introduce CLIP guidance into the multi-step sampling of diffusion, achieving high congruence between editing results and target text. Comprehensive quantitative and qualitative experiments demonstrate that our method effectively resolves the text alignment issues prevalent in existing methods while maintaining the fidelity to the source image, and performs well across a wide range of editing tasks. Tianrui Huang, Pu Cao, Lu Yang 0006, Chun Liu 0004, Mengjie Hu 0002, Zhiwei Liu 0004, Qing Song 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Constructing Balanced Training Samples: A New Perspective on Long-Tailed ClassificationabstractThe most significant characteristic of long-tailed classification is that severe sample imbalance causes the model to be biased towards the head category. While the long-tailed distribution of multimedia dataset remains a constant, we can enhance the acquisition of balanced training samples and corresponding features during the learning process. This paper innovatively designs a sample provider to construct balanced training samples to enhance the acquisition of comprehensive features, and proposes a Siamese-based parameter-sharing framework to handle data with long-tailed distributions. Specifically, one branch of the Siamese network is introduced to classify samples with conventional random cropping sampling, another branch integrates the advantages of constructed balanced samples and hybrid optimization to capture the balanced features to identify more precise category boundaries. This combination not only facilitates the learning of long-tailed distribution but also strengthens the model's extraction of balanced features through the incorporation of contrastive learning. Most significantly, extensive experiments on CIFAR10-LT, CIFAR100-LT, ImageNet-LT and iNaturalist 2018 datasets demonstrate our model not only achieves superior performance but also retains the benefits of end-to-end training. Specifically, our method achieves 60.7% accuracy on ImageNet-LT with an end-to-end ResNeXt-50 backbone. Wenyi Zhao, Wei Li 0243, Lu Yang 0006, Zhenhao Liang, Enwen Hu, Weidong Zhang 0007 |
IEEE Trans. Multim. | 4 |
| 2024 | What Decreases Editing Capability? Domain-Specific Hybrid Refinement for Improved GAN InversionabstractRecently, inversion methods have been exploring the incorporation of additional high-rate information from pretrained generators (such as weights or intermediate features) to improve the refinement of inversion and editing results from embedded latent codes. While such techniques have shown reasonable improvements in reconstruction, they often lead to a decrease in editing capability, especially when dealing with complex images that contain occlusions, detailed backgrounds, and artifacts. To address this problem, we propose a novel refinement mechanism called Domain-Specific Hybrid Refinement (DHR), which draws on the advantages and disadvantages of two mainstream refinement techniques. We find that the weight modulation can gain favorable editing results but is vulnerable to these complex image areas and feature modulation is efficient at reconstructing. Hence, we divide the image into two domains and process them with these two methods separately. We first propose a Domain-Specific Segmentation module to automatically segment images into in-domain and out-of-domain parts according to their invertibility and editability without additional data annotation, where our hybrid refinement process aims to maintain the editing capability for in-domain areas and improve fidelity for both of them. We achieve this through Hybrid Modulation Refinement, which respectively refines these two domains by weight modulation and feature modulation. Our proposed method is compatible with all latent code embedding methods. Extension experiments demonstrate that our approach achieves state-of-the-art in real image inversion and editing. Code is available at https: //github.com/caopulan/Domain-Specific_ Hybrid_Refinement_Inversion. Pu Cao, Lu Yang 0006, Dongxv Liu, Xiaoya Yang, Tianrui Huang, Qing Song 0006 |
WACV | 2 |
| 2024 | Deep Learning Technique for Human Parsing: A Survey and Outlook
Lu Yang 0006, Wenhe Jia, Shan Li 0001, Qing Song 0006 |
Int. J. Comput. Vis. | 1 |
| 2024 | MtArtGPT: A Multi-Task Art Generation System With Pre-Trained TransformerabstractInstruction tuning large language models are making rapid advances in the field of artificial intelligence where GPT-4 models have exhibited impressive multi-modal perception capabilities. Such models have been used as the core assistant for many tasks including art generation. However, high-quality art generation relies heavily on human prompt engineering which is in general uncontrollable. To address these issues, we propose a multi-task AI generated content (AIGC) system for art generation. Specifically, a dense representation manager is designed to process multi-modal user queries and generate dense and applicable prompts to GPT. To enhance artistic sophistication of the whole system, we fine-tune the GPT model by a meticulously collected prompt-art dataset. Furthermore, we introduce artistic benchmarks for evaluating the system based on professional knowledge. Experiments demonstrate the advantages of our proposed MtArtGPT system. Ruolin Zhu, Zixing Zhu, Lu Yang 0006, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Learning What and Where to Learn: A New Perspective on Self-Supervised LearningabstractSelf-supervised learning (SSL) has demonstrated its power in generalized model acquisition by leveraging the discriminative semantic and explicit positional information of unlabeled datasets. Unfortunately, mainstream contrastive learning-based methods excessive focus on semantic information and ignore the position is also the carrier of image content, resulting in inadequate data utilization and extensive computational consumption. To address these issues, we present an efficient SSL framework, learning What and Where to learn (W2SSL), to aggregate semantic and position features. Concretely, we devise a spatially-coupled sampling manner to process images through pre-defined rules, which integrates the advantage of semantic (What) and positional (Where) features into framework to enrich the diversity of feature representation capabilities and improve data utilization. Besides, a spectrum of latent vectors is obtained by mapping the positional features, which implicitly explores the relationship between these vectors. Whereafter, the corresponding discriminative and contrastive optimization objectives are seamlessly embedded in the framework via a cascade paradigm to explore semantic and positional features. The proposed W2SSL is verified on different types of datasets, which demonstrates that it still outperforms state-of-the-art SSL methods even with half the computational consumption. Code will be available at https://github.com/WilyZhao8/W2SSL. Wenyi Zhao, Lu Yang 0006, Weidong Zhang 0007, Yongqin Tian, Wenhe Jia, Wei Li 0243, Mu Yang, Xipeng Pan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Frequency-Based Matcher for Long-Tailed Semantic SegmentationabstractThe successful application of semantic segmentation technology in the real world has been among the most exciting achievements in the computer vision community over the past decade. Although the long-tailed phenomenon has been investigated in many fields,e.g., classification and object detection, it has not received enough attention in semantic segmentation and has become a nonnegligible obstacle to applying semantic segmentation technology in autonomous driving and virtual reality. Therefore, in this work, we focus on a relatively underexplored task setting,long-tailed semantic segmentation(LTSS). We first establish three representative datasets from different aspects, i.e., scene, object, and human. We further propose a dual-metric evaluation system and construct the LTSS benchmark to demonstrate the performance of semantic segmentation methods and long-tailed solutions. We also propose a transformer-based algorithm to improve LTSS,frequency-based matcher, which solves the oversuppression problem by one-to-many matching and automatically determines the number of matching queries for each class. Given the comprehensiveness of this work and the importance of the issues revealed, this work aims to promote the empirical study of semantic segmentation tasks. Our datasets, codes, and models will be publicly available. Shan Li 0001, Lu Yang 0006, Pu Cao, Liulei Li, Huadong Ma |
IEEE Trans. Multim. | 2 |
| 2023 | Large-Scale Person Detection and Localization using Overhead Fisheye CamerasabstractLocation determination finds wide applications in daily life. Instead of existing efforts devoted to localizing tourist photos captured by perspective cameras, in this article, we focus on devising person positioning solutions using overhead fisheye cameras. Such solutions are advantageous in large field of view (FOV), low cost, anti-occlusion, and unaggressive work mode (without the necessity of cameras carried by persons). However, related studies are quite scarce, due to the paucity of data. To stimulate research in this exciting area, we present LOAF, the first large-scale overhead fisheye dataset for person detection and localization. LOAF is built with many essential features, e.g., i) the data cover abundant diversities in scenes, human pose, density, and location; ii) it contains currently the largest number of annotated pedestrian, i.e., 457K bounding boxes with groundtruth location information; iii) the body-boxes are labeled as radius-aligned so as to fully address the positioning challenge. To approach localization, we build a fisheye person detection network, which exploits the fisheye distortions by a rotation-equivariant training strategy and predict radius-aligned human boxes end-to-end. Then, the actual locations of the detected persons are calculated by a numerical solution on the fisheye model and camera altitude data. Extensive experiments on LOAF validate the superiority of our fisheye detector w.r.t. previous methods, and show that our whole fisheye positioning solution is able to locate all persons in FOV with an accuracy of0.5 m, within 0.1 s. Lu Yang 0006, Liulei Li, Xueshi Xin, Qing Song 0006, Wenguan Wang |
ICCV | 1 |
| 2023 | TIVE: A toolbox for identifying video instance segmentation errors
Wenhe Jia, Lu Yang 0006, Zilong Jia, Wenyi Zhao, Qing Song 0006 |
Neurocomputing | 2 |
| 2023 | Rethinking the activation function in lightweight network
Lu Yang 0006, Qing Song 0006, Zimeng Fan 0003, Chun Liu 0004, Mengjie Hu 0002 |
Multim. Tools Appl. | 1 |
| 2023 | Embedding Global Contrastive and Local Location in Self-Supervised LearningabstractSelf-supervised representation learning (SSL) typically suffers from inadequate data utilization and feature-specificity due to the suboptimal sampling strategy and the monotonous optimization method. Existing contrastive-based methods alleviate these issues through exceedingly long training time and large batch size, resulting in non-negligible computational consumption and memory usage. In this paper, we present an efficient self-supervised framework, called GLNet. The key insights of this work are the novel sampling and ensemble learning strategies embedded in the self-supervised framework. We first propose a location-based sampling strategy to integrate the complementary advantages of semantic and spatial characteristics. Whereafter, a Siamese network with momentum update is introduced to generate representative vectors, which are used to optimize the feature extractor. Finally, we particularly embed global contrastive and local location tasks in the framework, which aims to leverage the complementarity between the high-level semantic features and low-level texture features. Such complementarity is significant for mitigating the feature-specificity and improving the generalizability, thus effectively improving the performance of downstream tasks. Extensive experiments on representative benchmark datasets demonstrate that GLNet performs favorably against the state-of-the-art SSL methods. Specifically, GLNet improves MoCo-v3 by 2.4% accuracy on ImageNet dataset, while improves 2% accuracy and consumes only 75% training time on the ImageNet-100 dataset. In addition, GLNet is appealing in its compatibility with popular SSL frameworks. Code is available at GLNet. Wenyi Zhao, Chongyi Li, Weidong Zhang 0007, Lu Yang 0006, Peixian Zhuang, Lingqiao Li, Kefeng Fan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Quality-Aware Network for Human ParsingabstractHow to estimate the quality of the network output is an important issue, and currently there is no effective solution in the field of human parsing. To solve this problem, this work proposes a statistical method based on the output probability map to calculate the pixel classification quality, which is called pixel score. In addition, the Quality-Aware Module (QAM) is proposed to fuse the different quality information, the purpose of which is to estimate the quality of human parsing results. We combine QAM with a concise and effective network design to propose Quality-Aware Network (QANet) for human parsing. Benefiting from the superiority of QAM and QANet, we achieve the best performance on three multiple and one single human parsing benchmarks, including CIHP, MHP-v2, Pascal-Person-Part, ATR and LIP. Without increasing the training and inference time, QAM improves the AP$^\text{r}$criterion by more than 10 points in the multiple human parsing task. QAM can be extended to other tasks with good quality estimation,e.ginstance segmentation. Specifically, QAM improves Mask R-CNN by$\scriptstyle \sim$1% mAP on COCO and LVISv1.0 datasets. Based on the proposed QAM and QANet, our overall system wins 1st place in CVPR2021 L2ID High-resolution Human Parsing (HRHP) Challenge, and 2nd in CVPR2021 PIC Short-video Face Parsing (SFP) Challenge. Code and models are available athttps://github.com/soeaver/QANet. Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011, Zhiwei Liu 0004, Songcen Xu, Zhihao Li 0002 |
IEEE Trans. Multim. | 1 |
| 2022 | Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence LearningabstractOur target is to learn visual correspondence from unlabeled videos. We develop Liir, a locality-aware inter-and intra-video reconstruction method that fills in three missing pieces, i.e., instance discrimination, location awareness, and spatial compactness, of self-supervised correspondence learning puzzle. First, instead of most existing efforts focusing on intra-video self-supervision only, we exploit cross-video affinities as extra negative samples within a unified, inter-and intra-video reconstruction scheme. This enables instance discriminative representation learning by contrasting desired intra-video pixel association against negative inter-video correspondence. Second, we merge position information into correspondence matching, and design a position shifting strategy to remove the side-effect of position encoding during inter-video affinity computation, making our Liir location-sensitive. Third, to make full use of the spatial continuity nature of video data, we impose a compactness-based constraint on correspondence matching, yielding more sparse and reliable solutions. The learned representation surpasses self-supervised state-of-the-arts on label propagation tasks including objects, semantic parts, and keypoints. Liulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang 0006, Jianwu Li, Yi Yang 0001 |
CVPR | 4 |
| 2022 | A Survey on Long-Tailed Visual Recognition
Lu Yang 0006, Qing Song 0006 |
Int. J. Comput. Vis. | 1 |
| 2022 | Double parallel branches FCOS for human detection in a crowd
Qing Song 0006, Lu Yang 0006, Xueshi Xin, Chun Liu 0004, Mengjie Hu 0002 |
Multim. Tools Appl. | 3 |
| 2021 | Multi-initialization Optimization Network for Accurate 3D Human Pose and Shape Estimationabstract3D human pose and shape recovery from a monocular RGB image is a challenging task. Existing learning based methods highly depend on weak supervision signals, e.g. 2D and 3D joint location, due to the lack of in-the-wild paired 3D supervision. However, considering the 2D-to-3D ambiguities existed in these weak supervision labels, the network is easy to get stuck in local optima when trained with such labels. In this paper, we reduce the ambituity by optimizing multiple initializations. Specifically, we propose a three-stage framework named Multi-Initialization Optimization Network (MION). In the first stage, we strategically select different coarse 3D reconstruction candidates which are compatible with the 2D keypoints of input sample. Each coarse reconstruction can be regarded as an initialization leads to one optimization branch. In the second stage, we design a mesh refinement transformer (MRT) to respectively refine each coarse reconstruction result via a self-attention mechanism. Finally, a Consistency Estimation Network (CEN) is proposed to find the best result from mutiple candidates by evaluating if the visual evidence in RGB image matches a given 3D reconstruction. Experiments demonstrate that our Multi-Initialization Optimization Network outperforms existing 3D mesh based methods on multiple public benchmarks. Zhiwei Liu 0004, Xiangyu Zhu 0001, Lu Yang 0006, Ming Tang 0001, Zhen Lei 0001, Guibo Zhu, Xuetao Feng, Yan Wang 0068, Jinqiao Wang |
ACM Multimedia | 3 |
| 2021 | CPM R-CNN: Calibrating Point-guided Misalignment in Object DetectionabstractIn object detection, offset-guided and point-guided regression dominate anchor-based and anchor-free method separately. Recently, point-guided approach is introduced to anchor-based method. However, we observe points predicted by this way are misaligned with matched region of proposals and score of localization, causing a notable gap in performance. In this paper, we propose CPM R-CNN which contains three efficient modules to optimize anchor- based point-guided method. According to sufficient evaluations on the COCO dataset, CPM R-CNN is demonstrated efficient to improve the localization accuracy by calibrating mentioned misalignment. Compared with Faster R-CNN and Grid R-CNN based on ResNet-101 with FPN, our approach can substantially improve detection mAP by 3.3% and 1.5% respectively without whistles and bells. Moreover, our best model achieves improvement by a large margin to 49.9% on COCO test-dev. Code is available at https://github.com/zhubinQAQ/CPM-R-CNN. Qing Song 0006, Lu Yang 0006, Zhihui Wang 0011, Chun Liu 0004, Mengjie Hu 0002 |
WACV | 3 |
| 2021 | Attacks on state-of-the-art face recognition using attentional adversarial attack generative networkabstractAbstract With the broad use of face recognition, its weakness gradually emerges that it is able to be attacked. Therefore, it is very important to study how face recognition networks are subject to attacks. Generating adversarial examples is an effective attack method, which misleads the face recognition system through obfuscation attack (rejecting a genuine subject) or impersonation attack (matching to an impostor). In this paper, we introduce a novel GAN, Attentional Adversarial Attack Generative Network (A3GN), to generate adversarial examples that mislead the network to identify someone as the target person not misclassify inconspicuously. For capturing the geometric and context information of the target person, this work adds a conditional variational autoencoder and attention modules to learn the instance-level correspondences between faces. Unlike traditional two-player GAN, this work introduces a face recognition network as the third player to participate in the competition between generator and discriminator which allows the attacker to impersonate the target person better. The generated faces which are hard to arouse the notice of onlookers can evade recognition by state-of-the-art networks and most of them are recognized as the target person. Lu Yang 0006, Qing Song 0006, Yingqi Wu |
Multim. Tools Appl. | 1 |
| 2021 | Hier R-CNN: Instance-Level Human Parts Detection and A New BenchmarkabstractDetecting human parts at instance-level is an essential prerequisite for the analysis of human keypoints, actions, and attributes. Nonetheless, there is a lack of a large-scale, rich-annotated dataset for human parts detection. We fill in the gap by proposing COCO Human Parts. The proposed dataset is based on the COCO 2017, which is the first instance-level human parts dataset, and contains images of complex scenes and high diversity. For reflecting the diversity of human body in natural scenes, we annotate human parts with (a) location in terms of a bounding-box, (b) various type including face, head, hand, and foot, (c) subordinate relationship between person and human parts, (d) fine-grained classification into right-hand/left-hand and left-foot/right-foot. A lot of higher-level applications and studies can be founded upon COCO Human Parts, such as gesture recognition, face/hand keypoint detection, visual actions, human-object interactions, and virtual reality. There are a total of 268,030 person instances from the 66,808 images, and 2.83 parts per person instance. We provide a statistical analysis of the accuracy of our annotations. In addition, we propose a strong baseline for detecting human parts at instance-level over this dataset in an end-to-end manner, call Hier(archy) R-CNN. It is a simple but effective extension of Mask R-CNN, which can detect human parts of each person instance and predict the subordinate relationship between them. Codes and dataset are publicly available (https://github.com/soeaver/Hier-R-CNN). Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011, Mengjie Hu 0002, Chun Liu 0004 |
IEEE Trans. Image Process. | 1 |
| 2020 | Renovating Parsing R-CNN for Accurate Multiple Human Parsing
Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011, Mengjie Hu 0002, Chun Liu 0004, Xueshi Xin, Wenhe Jia, Songcen Xu |
ECCV (12) | 1 |
| 2020 | Accelerate neural style transfer with super-resolution
Zuoxin Li, Fuqiang Zhou, Lu Yang 0006, Juan Li 0006 |
Multim. Tools Appl. | 3 |
| 2019 | Parsing R-CNN for Instance-Level Human AnalysisabstractInstance-level human analysis is common in real-life scenarios and has multiple manifestations, such as human part segmentation, dense pose estimation, human-object interactions, etc. Models need to distinguish different human instances in the image panel and learn rich features to represent the details of each instance. In this paper, we present an end-to-end pipeline for solving the instance-level human analysis, named Parsing R-CNN. It processes a set of human instances simultaneously through comprehensive considering the characteristics of region-based approach and the appearance of a human, thus allowing representing the details of instances. Parsing R-CNN is very flexible and efficient, which is applicable to many issues in human instance analysis. Our approach outperforms all state-of-the-art methods on CIHP (Crowd Instance-level Human Parsing), MHP v2.0 (Multi-Human Parsing) and DensePose-COCO datasets. Based on the proposed Parsing R-CNN, we reach the 1st place in the COCO 2018 Challenge DensePose Estimation task. Code and models are publicly available. Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011 |
CVPR | 1 |
| 2019 | Attention Inspiring Receptive-Fields Network for Learning Invariant RepresentationsabstractIn this paper, we describe a simple and highly efficient module for image classification, which we term the "Attention Inspiring Receptive-fields" (Air) module. We effectively convert the spatial attention mechanism into a plug-in module. In addition, we reveal the relationship between the spatial attention mechanism and the receptive fields, indicating that the proper use of the spatial attention mechanism can effectively increase the receptive fields of the module, which is able to enhance translation invariance and scale invariance of the network. By integrating the Air module into advanced convolutional neural networks (such as ResNet and ResNeXt), we can construct AirNet architectures for learning invariant representations and gain significant improvements on challenging data sets. We present extensive experiments on CIFAR and ImageNet data sets to verify the effectiveness and feature invariance of the Air module and explore more concise and efficient designs of the proposed module. On ImageNet classification, our AirNet-50 and AirNet-101 (ResNet-50/101 with Air module) achieve 1.69% and 1.50% top-1 accuracy improvement with a small amount of extra computation and parameters compared with the original ResNet. We make models and code public available https://github.com/soeaver/AirNet-PyTorch. We further demonstrate that AirNet has a good ability for transfer learning and measure the performance on Microsoft Common Objects in Context object detection, instance segmentation, and pose estimation. Lu Yang 0006, Qing Song 0006, Yingqi Wu, Mengjie Hu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Detector-in-Detector: Multi-level Analysis for Human-Parts
Lu Yang 0006, Qing Song 0006, Fuqiang Zhou |
ACCV (2) | 2 |
| 2018 | Fast Single Shot Instance Segmentation
Zuoxin Li, Fuqiang Zhou, Lu Yang 0006 |
ACCV (4) | 3 |
| 2018 | Cross Connected Network for Efficient Image Recognition
Lu Yang 0006, Qing Song 0006, Zuoxin Li, Yingqi Wu, Mengjie Hu 0002 |
ACCV (1) | 1 |
| 2017 | Maximum linear matching: Intelligent and automatic wavelength calibration methodabstractSummary Wavelength calibration is a necessary means to ensure the normal operation of the spectrometer. In general, calibration light sources that can emit fixed wavelength (such as mercury‐argon lamp) are utilized to generate spectral line on the linear array charge‐coupled device detector. Thus, it is a prerequisite for wavelength calibration to match the pixel position of the well‐divided spectral line with light of different wavelength emitted by calibration light source correctly. In this paper, we aim to present a method to make calibration procedure intelligent and automatic by machine. We establish a pixel‐wavelength model based on bipartite graph and propose the maximum linear matching (MLM) algorithm to find the correct set of pixel‐wavelength automatically. Meanwhile, we calculate precision and recall to measure the effectiveness of MLM and analyze the practical application of MLM by comparing it with conventional artificial methods. Experiments show that the autocalibration method based on MLM can calibrate many types of grating spectrometer more accurately and more reliably. With MLM, we report 98.33% precision and 94.59% recall on 8 groups of experiments. Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011 |
Concurr. Comput. Pract. Exp. | 1 |