Qing Song 0006

dblp:150/5904-6 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 18 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Controllable Generation With Text-to-Image Diffusion Models: A Survey
abstract
In the rapidly advancing realm of visual generation, diffusion models have revolutionized the landscape, marking a significant shift in capabilities with their impressive text-guided generative functions. However, relying solely on text for conditioning these models does not fully cater to the varied and complex requirements of different applications and scenarios. Acknowledging this shortfall, a variety of studies aim to control pre-trained text-to-image (T2I) models to support novel conditions. In this survey, we undertake a thorough review of the literature on controllable generation with T2I diffusion models, covering both the theoretical foundations and practical advancements in this domain. Our review begins with a brief introduction to the basics of denoising diffusion probabilistic models (DDPMs) and widely used T2I diffusion models. Additionally, we provide a detailed overview of research in this area, categorizing it from the condition perspective into three directions: generation with specific conditions, generation with multiple conditions, and universal controllable generation. For each category, we analyze the underlying control mechanisms and review representative methods based on their core techniques.
Pu Cao, Qing Song 0006, Lu Yang 0006
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 ParsingFormer for pixel-wise hierarchical human representation learning
Pu Cao, Zhixiang Lv, Junyi Ji, Shan Li 0001, Qing Song 0006, Lu Yang 0006
Pattern Recognit.5
2026 Quality transformer for human parsing
Lu Yang 0006, Pu Cao, Shan Li 0001, Qing Song 0006
Pattern Recognit.6
2026 Trajectory-Aware Attack: Explainable Adversarial Attack Against Multiple Object Trackers
abstract
Multi-Object Tracking (MOT) aims to build moving trajectories of objects within video sequences and serves as a critical component in autonomous driving systems. Recently, several studies have revealed the vulnerability of existing MOT methods by investigating adversarial attacks against MOT, raising significant safety concerns for real-world applications. These methods attack trackers by deliberately inserting false alarms, which mislead trajectories to drift from their correct paths. However, current MOT attack methods fail to propose efficient strategies for generating false alarms, as they either rely on computationally intensive optimization to determine the placement of false alarms, or crudely insert a large number of heuristically designed false alarms. In this paper, we propose an explainable and effective false alarm generation module, named Target Generating Module (TGM), that adaptively determines the location and size of false alarms by leveraging historical trajectory information. Based on this module, we design an attack method targeting mainstream MOT approaches, namedTrajectory-Aware Attack (TA Attack). TA Attack achieves effective disruption of MOT systems by combining detection erasure and false alarm generation, requiring only a few frames to successfully compromise trajectories. To exhibit the flexibility and effectiveness of our method, we conduct experiments using four multi-object trackers (ByteTrack, SORT, CenterTrack and FairMOT) which are enabled by two representative detectors (YOLOX and CenterNet). The results demonstrate our method achieves state of the art performance with 74.87% attack success rate on BDD100K, 81.7% attack success rate on MOT17 and 83.87% attack success rate on MOT20 while 4 frames being attacked averagely, revealing the vulnerability of association mechanism in MOT methods.
Mengjie Hu 0002, Yufei Ding 0003, Chun Liu 0004, Qing Song 0006
IEEE Trans. Multim.7
2025 Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
abstract
In-domain generation aims to perform a variety of tasks within a specific domain, such as unconditional generation, text-to-image, image editing, 3D generation, and more. Early research typically required training specialized generators for each unique task and domain, often relying on fully-labeled data. Motivated by the powerful generative capabilities and broad applications of diffusion models, we are driven to explore leveraging label-free data to empower these models for in-domain generation. Fine-tuning a pre-trained generative model on domain data is an intuitive but challenging way and often requires complex manual hyper-parameter adjustments since the limited diversity of the training data can easily disrupt the model’s original generative capabilities. To address this challenge, we propose a guidance-decoupled prior preservation mechanism to achieve high generative quality and controllability by image-only data, inspired by preserving the pre-trained model from a denoising guidance perspective. We decouple domain-related guidance from the conditional guidance used in classifier-free guidance mechanisms to preserve open-world control guidance and unconditional guidance from the pre-trained model. We further propose an efficient domain knowledge learning technique to train an additional text-free UNet copy to predict domain guidance. Besides, we theoretically illustrate a multi-guidance in-domain generation pipeline for a variety of generative tasks, leveraging multiple guidances from distinct diffusion models and conditions. Extensive experiments demonstrate the superiority of our method in domain-specific synthesis and its compatibility with various diffusion-based control methods and applications.
Pu Cao, Lu Yang 0006, Tianrui Huang, Qing Song 0006
CVPR5
2025 Semantic Guided Matting Net
abstract
Abstract Human matting refers to extracting human parts from natural images with high quality, including human detail information such as hair, glasses, hats, etc. This technology plays an essential role in image synthesis and visual effects in the film industry. When the green screen is not available, the existing human matting methods need the help of additional inputs (such as trimap, background image, etc.), or the model with high computational cost and complex network structure, which brings great difficulties to the application of human matting in practice. To alleviate such problems, we use a segmentation network as the foundation and use multiple branches to achieve human segmentation, contour detail extraction, and information fusion. We also propose a foreground probability map module, which uses the feature maps in the segmentation network to pre-estimate the foreground probabilities of each pixel and obtain Semantic Guided Matting Net. Under the condition that only a single image is needed as the input, the human matting task can be realized by making full use of the semantic information in the image. We validate our method on the P3M-10k dataset. Compared with the benchmark, our method has made significant improvements in various evaluation indicators.
Qing Song 0006, Wenfeng Sun, Donghan Yang, Mengjie Hu 0002, Chun Liu 0004
Comput. J.1
2025 Beyond-Skeleton: Zero-shot Skeleton Action Recognition enhanced by supplementary RGB visual information
Yingchun Niu, Chun Liu 0004, Mengjie Hu 0002, Qing Song 0006
Expert Syst. Appl.6
2025 E4C: Enhance Editability for Text-Based Image Editing by Harnessing Efficient CLIP Guidance
abstract
Diffusion-based image editing involves both preserving the source image content and generating new content or applying modifications. Although current editing approaches have made improvements under text guidance, they have two key drawbacks: overemphasis on retaining original image info, neglecting editability and text alignment, and inability to handle both structure-consistent and non-rigid editing tasks. In this paper, we propose a zero-shot image editing method, named Enhance Editability for text-based image Editing via Efficient CLIP guidance (E4C), which presents an innovative adaptive feature sharing mechanism to enable multi-task editing. Additionally, a novel random gateway mechanism is designed to efficiently introduce CLIP guidance into the multi-step sampling of diffusion, achieving high congruence between editing results and target text. Comprehensive quantitative and qualitative experiments demonstrate that our method effectively resolves the text alignment issues prevalent in existing methods while maintaining the fidelity to the source image, and performs well across a wide range of editing tasks.
Tianrui Huang, Pu Cao, Lu Yang 0006, Chun Liu 0004, Mengjie Hu 0002, Zhiwei Liu 0004, Qing Song 0006
IEEE Trans. Circuits Syst. Video Technol.7
2025 Multi-Object Tracking With Separation in Deep Space
abstract
In deep space environment, some objects may split into several small fragments during movement, and these deep space objects often appear as points in satellite images. In this article, we conduct research on multi-object tracking (MOT) for these objects. First, we propose a simulation dataset, ScatterDataset, which simulates the movement and separation of objects in deep space background. By assigning two IDs to a trajectory, we describe the trajectory’s relationship before and after separation. Second, we present an end-to-end motion association model, ScatterNet, which encodes the position information of trajectories and detections into motion features. These features are processed through temporal aggregation by a Transformer encoder and spatial aggregation by a graph network; then, we get the association results by calculating the similarity between these features. Finally, we introduce a tracker, ScatterTracker, which is suitable for tracking in scenarios with object separation. Experiments with state-of-the-art tracking methods on ScatterDataset demonstrate that our approach has achieved significant performance improvements in deep space scenarios. The code is available at:https://github.com/wht-bupt/ScatterTrack.
Mengjie Hu 0002, Binyu Li, Shixiang Cao, Tao Zhan 0002, Xiaotong Zhu, Chun Liu 0004, Qing Song 0006
IEEE Trans. Geosci. Remote. Sens.10
2024 What Decreases Editing Capability? Domain-Specific Hybrid Refinement for Improved GAN Inversion
abstract
Recently, inversion methods have been exploring the incorporation of additional high-rate information from pretrained generators (such as weights or intermediate features) to improve the refinement of inversion and editing results from embedded latent codes. While such techniques have shown reasonable improvements in reconstruction, they often lead to a decrease in editing capability, especially when dealing with complex images that contain occlusions, detailed backgrounds, and artifacts. To address this problem, we propose a novel refinement mechanism called Domain-Specific Hybrid Refinement (DHR), which draws on the advantages and disadvantages of two mainstream refinement techniques. We find that the weight modulation can gain favorable editing results but is vulnerable to these complex image areas and feature modulation is efficient at reconstructing. Hence, we divide the image into two domains and process them with these two methods separately. We first propose a Domain-Specific Segmentation module to automatically segment images into in-domain and out-of-domain parts according to their invertibility and editability without additional data annotation, where our hybrid refinement process aims to maintain the editing capability for in-domain areas and improve fidelity for both of them. We achieve this through Hybrid Modulation Refinement, which respectively refines these two domains by weight modulation and feature modulation. Our proposed method is compatible with all latent code embedding methods. Extension experiments demonstrate that our approach achieves state-of-the-art in real image inversion and editing. Code is available at https: //github.com/caopulan/Domain-Specific_ Hybrid_Refinement_Inversion.
Pu Cao, Lu Yang 0006, Dongxv Liu, Xiaoya Yang, Tianrui Huang, Qing Song 0006
WACV6
2024 Deep Learning Technique for Human Parsing: A Survey and Outlook
Lu Yang 0006, Wenhe Jia, Shan Li 0001, Qing Song 0006
Int. J. Comput. Vis.4
2024 UV R-CNN: Stable and efficient dense human pose estimation
Wenhe Jia, Xuhan Zhu, Mengjie Hu 0002, Chun Liu 0004, Qing Song 0006
Multim. Tools Appl.6
2024 CoT-MISR:Marrying convolution and transformer for multi-image super-resolution
Qing Song 0006, Mingming Xiu, Yang Nie, Mengjie Hu 0002, Chun Liu 0004
Multim. Tools Appl.1
2024 Faster learning of temporal action proposal via sparse multilevel boundary generator
Qing Song 0006, Mengjie Hu 0002, Chun Liu 0004
Multim. Tools Appl.1
2023 Large-Scale Person Detection and Localization using Overhead Fisheye Cameras
abstract
Location determination finds wide applications in daily life. Instead of existing efforts devoted to localizing tourist photos captured by perspective cameras, in this article, we focus on devising person positioning solutions using overhead fisheye cameras. Such solutions are advantageous in large field of view (FOV), low cost, anti-occlusion, and unaggressive work mode (without the necessity of cameras carried by persons). However, related studies are quite scarce, due to the paucity of data. To stimulate research in this exciting area, we present LOAF, the first large-scale overhead fisheye dataset for person detection and localization. LOAF is built with many essential features, e.g., i) the data cover abundant diversities in scenes, human pose, density, and location; ii) it contains currently the largest number of annotated pedestrian, i.e., 457K bounding boxes with groundtruth location information; iii) the body-boxes are labeled as radius-aligned so as to fully address the positioning challenge. To approach localization, we build a fisheye person detection network, which exploits the fisheye distortions by a rotation-equivariant training strategy and predict radius-aligned human boxes end-to-end. Then, the actual locations of the detected persons are calculated by a numerical solution on the fisheye model and camera altitude data. Extensive experiments on LOAF validate the superiority of our fisheye detector w.r.t. previous methods, and show that our whole fisheye positioning solution is able to locate all persons in FOV with an accuracy of0.5 m, within 0.1 s.
Lu Yang 0006, Liulei Li, Xueshi Xin, Qing Song 0006, Wenguan Wang
ICCV5
2023 TIVE: A toolbox for identifying video instance segmentation errors
Wenhe Jia, Lu Yang 0006, Zilong Jia, Wenyi Zhao, Qing Song 0006
Neurocomputing6
2023 Fast and robust for texture-less feature registration via adaptive heterogeneous kernels
Yuandong Ma, Qing Song 0006, Hezheng Lin, Chun Liu 0004, Mengjie Hu 0002, Xiaotong Zhu
Knowl. Based Syst.2
2023 Rethinking the activation function in lightweight network
Lu Yang 0006, Qing Song 0006, Zimeng Fan 0003, Chun Liu 0004, Mengjie Hu 0002
Multim. Tools Appl.2
2023 A continuation method for image registration based on dynamic adaptive kernel
Yuandong Ma, Hezheng Lin, Chun Liu 0004, Mengjie Hu 0002, Qing Song 0006
Neural Networks6
2023 A Lightweight Neural Learning Algorithm for Real-Time Facial Feature Tracking System via Split-Attention and Heterogeneous Convolution
Yuandong Ma, Qing Song 0006, Mengjie Hu 0002, Xiaotong Zhu
Neural Process. Lett.2
2023 Correction: A Lightweight Neural Learning Algorithm for Real-Time Facial Feature Tracking System via Split-Attention and Heterogeneous Convolution
Yuandong Ma, Qing Song 0006, Mengjie Hu 0002, Xiaotong Zhu
Neural Process. Lett.2
2023 STDFormer: Spatial-Temporal Motion Transformer for Multiple Object Tracking
abstract
Mainstream multi-object tracking methods exploit appearance information and/or motion information to achieve interframe association. However, dealing with similar appearance and occlusion is a challenge for appearance information, while motion information is limited by linear assumptions and is prone to failure in nonlinear motion patterns. In this work, we disregard appearance clues and propose a pure motion tracker to address the above issues. It dexterously utilizes Transformer to estimate complex motion and achieves high-performance tracking with low computing resources. Furthermore, contrastive learning is introduced to optimize feature representation for robust association. Specifically, we first exploit the long-range modeling capability of Transformer to mine intention information in temporal motion and decision information in spatial interaction and introduce prior detection to constrain the range of motion estimation. Then, we introduce contrastive learning as an auxiliary task to extract reliable motion features to compute affinity and introduce bidirectional matching to improve the affinity computation distribution. In addition, given that both tasks are dedicated to narrowing the embedding distance between the motion features of the tracked object and the detection features, we design a joint-motion-and-association framework to unify the above two tasks in one framework for optimization. The experimental results achieved with three benchmark datasets, MOT17, MOT20 and DanceTrack, verify the effectiveness of our proposed method. Compared with state-of-the-art methods, the proposed STDFormer sets a new state-of-the-art on DanceTrack and achieves competitive performance on MOT17 and MOT20. This demonstrates the advantage of our method in handling associations under similar appearance, occlusion or nonlinear motion. At the same time, the significant advantages of the proposed method over Transformer-based and contrastive learning-based methods suggest a new direction for the application of Transformer and contrastive learning in MOT. In addition, to verify the generalization of STDFormer in unmanned aerial vehicle (UAV) videos, we also evaluate STDFormer on VisDrone2019. The results show that STDFormer achieves state-of-the-art performance on VisDrone2019, which proves that it can handle small-scale object associations in UAV videos well. The code is available at https://github.com/Xiaotong-Zhu/STDFormer.
Mengjie Hu 0002, Xiaotong Zhu, Shixiang Cao, Chun Liu 0004, Qing Song 0006
IEEE Trans. Circuits Syst. Video Technol.6
2023 Deep-Chain Echo State Network With Explainable Temporal Dependence for Complex Building Energy Prediction
abstract
Building energy prediction plays critical roles in the study of green building and smart city. The most challenging issue is to predict energy demand profiles over multiple time steps, which may have inconsistent timescales. Due to complex temporal dependence, existing prediction approaches cannot satisfy certain requirements in multistep (MS) or multitimescale (MTS) applications. In this article, a deep-chain echo state network (DCESN) is proposed to enhance the mapping capability for the MS demand prediction. The DCESN composed of many submodules of echo state network (ESN) belongs to the single-input multi-output (SIMO) model, and neuron states generated by sequential steps are utilized to prevent accumulative error in the recursive chain. Due to the novel learning mechanism of DCESN, numerical coefficients of temporal dependence are presented to explain the short-term and long-term impacts on future energy consumption. Experimental results in four cases indicate that the proposed DCESN has promising performance on MS and MTS prediction, and temporal dependence can be explained in a visible way. Comparative results of DCESN, sliding-window ESN, and long-short term memory (LSTM) demonstrate that the proposed learning mechanism could prevent error accumulation effectively.
Ruiqi Jiang, Shaoxiong Zeng, Qing Song 0006, Zhou Wu 0001
IEEE Trans. Ind. Informatics3
2023 Quality-Aware Network for Human Parsing
abstract
How to estimate the quality of the network output is an important issue, and currently there is no effective solution in the field of human parsing. To solve this problem, this work proposes a statistical method based on the output probability map to calculate the pixel classification quality, which is called pixel score. In addition, the Quality-Aware Module (QAM) is proposed to fuse the different quality information, the purpose of which is to estimate the quality of human parsing results. We combine QAM with a concise and effective network design to propose Quality-Aware Network (QANet) for human parsing. Benefiting from the superiority of QAM and QANet, we achieve the best performance on three multiple and one single human parsing benchmarks, including CIHP, MHP-v2, Pascal-Person-Part, ATR and LIP. Without increasing the training and inference time, QAM improves the AP$^\text{r}$criterion by more than 10 points in the multiple human parsing task. QAM can be extended to other tasks with good quality estimation,e.ginstance segmentation. Specifically, QAM improves Mask R-CNN by$\scriptstyle \sim$1% mAP on COCO and LVISv1.0 datasets. Based on the proposed QAM and QANet, our overall system wins 1st place in CVPR2021 L2ID High-resolution Human Parsing (HRHP) Challenge, and 2nd in CVPR2021 PIC Short-video Face Parsing (SFP) Challenge. Code and models are available athttps://github.com/soeaver/QANet.
Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011, Zhiwei Liu 0004, Songcen Xu, Zhihao Li 0002
IEEE Trans. Multim.2
2022 A Survey on Long-Tailed Visual Recognition
Lu Yang 0006, Qing Song 0006
Int. J. Comput. Vis.3
2022 Double parallel branches FCOS for human detection in a crowd
Qing Song 0006, Lu Yang 0006, Xueshi Xin, Chun Liu 0004, Mengjie Hu 0002
Multim. Tools Appl.1
2021 CPM R-CNN: Calibrating Point-guided Misalignment in Object Detection
abstract
In object detection, offset-guided and point-guided regression dominate anchor-based and anchor-free method separately. Recently, point-guided approach is introduced to anchor-based method. However, we observe points predicted by this way are misaligned with matched region of proposals and score of localization, causing a notable gap in performance. In this paper, we propose CPM R-CNN which contains three efficient modules to optimize anchor- based point-guided method. According to sufficient evaluations on the COCO dataset, CPM R-CNN is demonstrated efficient to improve the localization accuracy by calibrating mentioned misalignment. Compared with Faster R-CNN and Grid R-CNN based on ResNet-101 with FPN, our approach can substantially improve detection mAP by 3.3% and 1.5% respectively without whistles and bells. Moreover, our best model achieves improvement by a large margin to 49.9% on COCO test-dev. Code is available at https://github.com/zhubinQAQ/CPM-R-CNN.
Qing Song 0006, Lu Yang 0006, Zhihui Wang 0011, Chun Liu 0004, Mengjie Hu 0002
WACV2
2021 Attacks on state-of-the-art face recognition using attentional adversarial attack generative network
abstract
Abstract With the broad use of face recognition, its weakness gradually emerges that it is able to be attacked. Therefore, it is very important to study how face recognition networks are subject to attacks. Generating adversarial examples is an effective attack method, which misleads the face recognition system through obfuscation attack (rejecting a genuine subject) or impersonation attack (matching to an impostor). In this paper, we introduce a novel GAN, Attentional Adversarial Attack Generative Network (A3GN), to generate adversarial examples that mislead the network to identify someone as the target person not misclassify inconspicuously. For capturing the geometric and context information of the target person, this work adds a conditional variational autoencoder and attention modules to learn the instance-level correspondences between faces. Unlike traditional two-player GAN, this work introduces a face recognition network as the third player to participate in the competition between generator and discriminator which allows the attacker to impersonate the target person better. The generated faces which are hard to arouse the notice of onlookers can evade recognition by state-of-the-art networks and most of them are recognized as the target person.
Lu Yang 0006, Qing Song 0006, Yingqi Wu
Multim. Tools Appl.2
2021 Hier R-CNN: Instance-Level Human Parts Detection and A New Benchmark
abstract
Detecting human parts at instance-level is an essential prerequisite for the analysis of human keypoints, actions, and attributes. Nonetheless, there is a lack of a large-scale, rich-annotated dataset for human parts detection. We fill in the gap by proposing COCO Human Parts. The proposed dataset is based on the COCO 2017, which is the first instance-level human parts dataset, and contains images of complex scenes and high diversity. For reflecting the diversity of human body in natural scenes, we annotate human parts with (a) location in terms of a bounding-box, (b) various type including face, head, hand, and foot, (c) subordinate relationship between person and human parts, (d) fine-grained classification into right-hand/left-hand and left-foot/right-foot. A lot of higher-level applications and studies can be founded upon COCO Human Parts, such as gesture recognition, face/hand keypoint detection, visual actions, human-object interactions, and virtual reality. There are a total of 268,030 person instances from the 66,808 images, and 2.83 parts per person instance. We provide a statistical analysis of the accuracy of our annotations. In addition, we propose a strong baseline for detecting human parts at instance-level over this dataset in an end-to-end manner, call Hier(archy) R-CNN. It is a simple but effective extension of Mask R-CNN, which can detect human parts of each person instance and predict the subordinate relationship between them. Codes and dataset are publicly available (https://github.com/soeaver/Hier-R-CNN).
Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011, Mengjie Hu 0002, Chun Liu 0004
IEEE Trans. Image Process.2
2020 Renovating Parsing R-CNN for Accurate Multiple Human Parsing
Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011, Mengjie Hu 0002, Chun Liu 0004, Xueshi Xin, Wenhe Jia, Songcen Xu
ECCV (12)2
2019 Parsing R-CNN for Instance-Level Human Analysis
abstract
Instance-level human analysis is common in real-life scenarios and has multiple manifestations, such as human part segmentation, dense pose estimation, human-object interactions, etc. Models need to distinguish different human instances in the image panel and learn rich features to represent the details of each instance. In this paper, we present an end-to-end pipeline for solving the instance-level human analysis, named Parsing R-CNN. It processes a set of human instances simultaneously through comprehensive considering the characteristics of region-based approach and the appearance of a human, thus allowing representing the details of instances. Parsing R-CNN is very flexible and efficient, which is applicable to many issues in human instance analysis. Our approach outperforms all state-of-the-art methods on CIHP (Crowd Instance-level Human Parsing), MHP v2.0 (Multi-Human Parsing) and DensePose-COCO datasets. Based on the proposed Parsing R-CNN, we reach the 1st place in the COCO 2018 Challenge DensePose Estimation task. Code and models are publicly available.
Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011
CVPR2
2019 Attention Inspiring Receptive-Fields Network for Learning Invariant Representations
abstract
In this paper, we describe a simple and highly efficient module for image classification, which we term the "Attention Inspiring Receptive-fields" (Air) module. We effectively convert the spatial attention mechanism into a plug-in module. In addition, we reveal the relationship between the spatial attention mechanism and the receptive fields, indicating that the proper use of the spatial attention mechanism can effectively increase the receptive fields of the module, which is able to enhance translation invariance and scale invariance of the network. By integrating the Air module into advanced convolutional neural networks (such as ResNet and ResNeXt), we can construct AirNet architectures for learning invariant representations and gain significant improvements on challenging data sets. We present extensive experiments on CIFAR and ImageNet data sets to verify the effectiveness and feature invariance of the Air module and explore more concise and efficient designs of the proposed module. On ImageNet classification, our AirNet-50 and AirNet-101 (ResNet-50/101 with Air module) achieve 1.69% and 1.50% top-1 accuracy improvement with a small amount of extra computation and parameters compared with the original ResNet. We make models and code public available https://github.com/soeaver/AirNet-PyTorch. We further demonstrate that AirNet has a good ability for transfer learning and measure the performance on Microsoft Common Objects in Context object detection, instance segmentation, and pose estimation.
Lu Yang 0006, Qing Song 0006, Yingqi Wu, Mengjie Hu 0002
IEEE Trans. Neural Networks Learn. Syst.2
2018 Detector-in-Detector: Multi-level Analysis for Human-Parts
Lu Yang 0006, Qing Song 0006, Fuqiang Zhou
ACCV (2)3
2018 Cross Connected Network for Efficient Image Recognition
Lu Yang 0006, Qing Song 0006, Zuoxin Li, Yingqi Wu, Mengjie Hu 0002
ACCV (1)2
2017 Maximum linear matching: Intelligent and automatic wavelength calibration method
abstract
Summary Wavelength calibration is a necessary means to ensure the normal operation of the spectrometer. In general, calibration light sources that can emit fixed wavelength (such as mercury‐argon lamp) are utilized to generate spectral line on the linear array charge‐coupled device detector. Thus, it is a prerequisite for wavelength calibration to match the pixel position of the well‐divided spectral line with light of different wavelength emitted by calibration light source correctly. In this paper, we aim to present a method to make calibration procedure intelligent and automatic by machine. We establish a pixel‐wavelength model based on bipartite graph and propose the maximum linear matching (MLM) algorithm to find the correct set of pixel‐wavelength automatically. Meanwhile, we calculate precision and recall to measure the effectiveness of MLM and analyze the practical application of MLM by comparing it with conventional artificial methods. Experiments show that the autocalibration method based on MLM can calibrate many types of grating spectrometer more accurately and more reliably. With MLM, we report 98.33% precision and 94.59% recall on 8 groups of experiments.
Lu Yang 0006, Qing Song 0006, Zhihui Wang 0011
Concurr. Comput. Pract. Exp.2