Qiang Nie

dblp:115/5257 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 12 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 EmbryoDiff: A Conditional Diffusion Framework with Multi-Focal Feature Fusion for Fine-Grained Embryo Developmental Stage Recognition
abstract
Identification of fine-grained embryo developmental stages during In Vitro Fertilization (IVF) is crucial for assessing embryo viability. Although recent deep learning methods have achieved promising accuracy, existing discriminative models fail to utilize the distributional prior of embryonic development to improve accuracy. Moreover, their reliance on single-focal information leads to incomplete embryonic representations, making them susceptible to feature ambiguity under cell occlusions. To address these limitations, we propose EmbryoDiff, a two-stage diffusion-based framework that formulates the task as a conditional sequence denoising process. Specifically, we first train and freeze a frame-level encoder to extract robust multi-focal features. In the second stage, we introduce a Multi-Focal Feature Fusion Strategy that aggregates information across focal planes to construct a 3D-aware morphological representation, effectively alleviating ambiguities arising from cell occlusions. Building on this fused representation, we derive complementary semantic and boundary cues and design a Hybrid Semantic-Boundary Condition Block to inject them into the diffusion-based denoising process, enabling accurate embryonic stage classification. Extensive experiments on two benchmark datasets show that our method achieves state-of-the-art results. Notably, with only a single denoising step, our model obtains the best average test performance, reaching 82.8% and 81.3% accuracy on the two datasets, respectively.
Zhengjie Zhang, Junyu Shi, Lijiang Liu, Qiang Nie
AAAI6
2026 Research on Energy Efficient Collaborative Control Method of Autonomous Consistency for Urban Rail Trains
Ruxun Xu, Qiang Nie, Decang Li, Jianjun Meng
IEEE Trans. Intell. Transp. Syst.3
2025 One Polyp Identifies All: One-Shot Polyp Segmentation with SAM via Cascaded Priors and Iterative Prompt Evolution
abstract
Polyp segmentation is vital for early colorectal cancer detection, yet traditional fully supervised methods struggle with morphological variability and domain shifts, requiring frequent retraining. Additionally, reliance on large-scale annotations is a major bottleneck due to the time-consuming and error-prone nature of polyp boundary labeling. Recently, vision foundation models like Segment Anything Model (SAM) have demonstrated strong generalizability and fine-grained boundary detection with sparse prompts, effectively addressing key polyp segmentation challenges. However, SAM's prompt-dependent nature limits automation in medical applications, since manually inputting prompts for each image is labor-intensive and time-consuming. We propose OP-SAM, a One-shot Polyp segmentation framework based on SAM that automatically generates prompts from a single annotated image, ensuring accurate and generalizable segmentation without additional annotation burdens. Our method introduces Correlation-based Prior Generation (CPG) for semantic label transfer and Scale-cascaded Prior Fusion (SPF) to adapt to polyp size variations as well as filter out noisy transfers. Instead of dumping all prompts at once, we devise Euclidean Prompt Evolution (EPE) for iterative prompt refinement, progressively enhancing segmentation quality. Extensive evaluations across five datasets validate OP-SAM's effectiveness. Notably, on Kvasir, it achieves 76.93% IoU, surpassing the state-of-the-art by 11.44%.
Xiaohan Xing, Jianbang Liu 0002, Fan Bai 0008, Qiang Nie, Max Q.-H. Meng
ICCV6
2025 GenM3: Generative Pretrained Multi-Path Motion Model for Text Conditional Human Motion Generation
Junyu Shi, Lijiang Liu, Jinni Zhou, Qiang Nie
ICCV6
2025 RMG: Real-Time Expressive Motion Generation with Self-collision Avoidance for 6-DOF Companion Robotic Arms
abstract
The six-degree-of-freedom (6-DOF) robotic arm has gained widespread application in human-coexisting environments. While previous research has predominantly focused on functional motion generation, the critical aspect of expressive motion in human-robot interaction remains largely unexplored. This paper presents a novel real-time motion generation planner that enhances interactivity by creating expressive robotic motions between arbitrary start and end states within predefined time constraints. Our approach involves three key contributions: first, we develop a mapping algorithm to construct an expressive motion dataset derived from human dance movements; second, we train motion generation models in both Cartesian and joint spaces using this dataset; third, we introduce an optimization algorithm that guarantees smooth, collision-free motion while maintaining the intended expressive style. Experimental results demonstrate the effectiveness of our method, which can generate expressive and generalized motions in under 0.5 seconds while satisfying all specified constraints.
Jiansheng Li, Haotian Song, Haoang Li, Jinni Zhou, Qiang Nie
IROS5
2025 Time-Lapse Video-Based Embryo Grading via Complementary Spatial-Temporal Pattern Mining
Junyu Shi, Yanmei Xiao, Manxi Jiang, Qiang Nie
MICCAI (13)8
2025 Towards Balanced Representation Learning with Semantic Anchor Regularization
Chengjie Wang 0001, Qiang Nie, Yong Liu 0032, Xi Jiang 0009, Yanqi Ge, Yunsheng Wu, Feng Zheng 0001, Lizhuang Ma
Int. J. Comput. Vis.2
2024 Unsupervised Continual Anomaly Detection with Contrastively-Learned Prompt
abstract
Unsupervised Anomaly Detection (UAD) with incremental training is crucial in industrial manufacturing, as unpredictable defects make obtaining sufficient labeled data infeasible. However, continual learning methods primarily rely on supervised annotations, while the application in UAD is limited due to the absence of supervision. Current UAD methods train separate models for different classes sequentially, leading to catastrophic forgetting and a heavy computational burden. To address this issue, we introduce a novel Unsupervised Continual Anomaly Detection framework called UCAD, which equips the UAD with continual learning capability through contrastively-learned prompts. In the proposed UCAD, we design a Continual Prompting Module (CPM) by utilizing a concise key-prompt-knowledge memory bank to guide task-invariant 'anomaly' model predictions using task-specific 'normal' knowledge. Moreover, Structure-based Contrastive Learning (SCL) is designed with the Segment Anything Model (SAM) to improve prompt learning and anomaly segmentation results. Specifically, by treating SAM's masks as structure, we draw features within the same mask closer and push others apart for general feature representations. We conduct comprehensive experiments and set the benchmark on unsupervised continual anomaly detection and segmentation, demonstrating that our method is significantly better than anomaly detection methods, even with rehearsal training. The code will be available at https://github.com/shirowalker/UCAD.
Jiaqi Liu 0004, Qiang Nie, Bin-Bin Gao, Yong Liu 0032, Jinbao Wang 0001, Chengjie Wang 0001, Feng Zheng 0001
AAAI3
2024 Beyond Prototypes: Semantic Anchor Regularization for Better Representation Learning
abstract
One of the ultimate goals of representation learning is to achieve compactness within a class and well-separability between classes. Many outstanding metric-based and prototype-based methods following the Expectation-Maximization paradigm, have been proposed for this objective. However, they inevitably introduce biases into the learning process, particularly with long-tail distributed training data. In this paper, we reveal that the class prototype is not necessarily to be derived from training features and propose a novel perspective to use pre-defined class anchors serving as feature centroid to unidirectionally guide feature learning. However, the pre-defined anchors may have a large semantic distance from the pixel features, which prevents them from being directly applied. To address this issue and generate feature centroid independent from feature learning, a simple yet effective Semantic Anchor Regularization (SAR) is proposed. SAR ensures the inter-class separability of semantic anchors in the semantic space by employing a classifier-aware auxiliary cross-entropy loss during training via disentanglement learning. By pulling the learned features to these semantic anchors, several advantages can be attained: 1) the intra-class compactness and naturally inter-class separability, 2) induced bias or errors from feature learning can be avoided, and 3) robustness to the long-tailed problem. The proposed SAR can be used in a plug-and-play manner in the existing models. Extensive experiments demonstrate that the SAR performs better than previous sophisticated prototype-based methods. The implementation is available at https://github.com/geyanqi/SAR.
Yanqi Ge, Qiang Nie, Yong Liu 0020, Chengjie Wang 0001, Feng Zheng 0001, Wen Li 0001, Lixin Duan
AAAI2
2024 Distribution-Aware Calibration for Object Detection with Noisy Bounding Boxes
Jinpeng Li 0004, Jiancheng Huang, Qiang Nie, Yong Liu 0032, Bin-Bin Gao, Qiong Wang 0001, Pheng-Ann Heng, Guangyong Chen
BMVC5
2024 LORS: Low-Rank Residual Structure for Parameter-Efficient Network Stacking
abstract
Deep learning models, particularly those based on transformers, often employ numerous stacked structures, which possess identical architectures and perform similar functions. While effective, this stacking paradigm leads to a substantial increase in the number of parameters, posing challenges for practical applications. In today's landscape of increasingly large models, stacking depth can even reach dozens, further exacerbating this issue. To mitigate this problem, we introduce LORS (LOw-rank Residual Structure). LORS allows stacked modules to share the majority of parameters, requiring a much smaller number of unique ones per module to match or even surpass the performance of using entirely distinct ones, thereby significantly reducing parameter usage. We validate our method by applying it to the stacked decoders of a query-based object detector, and conduct extensive experiments on the widely used MS COCO dataset. Experimental results demonstrate the effectiveness of our method, as even with a 70% reduction in the parameters of the decoder, our method still enables the model to achieve comparable or even better performance than its original.
Qiang Nie, Weifu Fu, Yuhuan Lin, Guangpin Tao, Yong Liu 0032, Chengjie Wang 0001
CVPR2
2024 Tuning-Free Image Customization with Image and Text Guidance
Pengzhi Li, Qiang Nie, Xi Jiang 0009, Yuhuan Lin, Yong Liu 0032, Jinlong Peng, Chengjie Wang 0001, Feng Zheng 0001
ECCV (76)2
2023 HopFIR: Hop-wise GraphFormer with Intragroup Joint Refinement for 3D Human Pose Estimation
abstract
2D-to-3D human pose lifting is fundamental for 3D human pose estimation (HPE), for which graph convolutional networks (GCNs) have proven inherently suitable for modeling the human skeletal topology. However, the current GCN-based 3D HPE methods update the node features by aggregating their neighbors’ information without considering the interaction of joints in different joint synergies. Although some studies have proposed importing limb information to learn the movement patterns, the latent synergies among joints, such as maintaining balance are seldom investigated. We propose the Hop-wise GraphFormer with Intragroup Joint Refinement (HopFIR) architecture to tackle the 3D HPE problem. HopFIR mainly consists of a novel hop-wise GraphFormer (HGF) module and an intragroup joint refinement (IJR) module. The HGF module groups the joints by k-hop neighbors and applies a hop-wise transformer-like attention mechanism to these groups to discover latent joint synergies. The IJR module leverages the prior limb information for peripheral joint refinement. Extensive experimental results show that HopFIR outperforms the SOTA methods by a large margin, with a mean per-joint position error (MPJPE) on the Human3.6M dataset of 32.67 mm. We also demonstrate that the state-of-the-art GCN-based methods can benefit from the proposed hop-wise attention mechanism with a significant improvement in performance: SemGCN [42] and MGCN [49] are improved by 8.9% and 4.5%, respectively.
Kai Zhai, Qiang Nie, Bo Ouyang, Shanlin Yang
ICCV2
2023 NeRF-Loc: Visual Localization with Conditional Neural Radiance Field
abstract
We propose a novel visual re-localization method based on direct matching between the implicit 3D descriptors and the 2D image with transformer. A conditional neural radiance field(NeRF) is chosen as the 3D scene representation in our pipeline, which supports continuous 3D descriptors generation and neural rendering. By unifying the feature matching and the scene coordinate regression to the same framework, our model learns both generalizable knowledge and scene prior respectively during two training stages. Furthermore, to improve the localization robustness when domain gap exists between training and testing phases, we propose an appearance adaptation layer to explicitly align styles between the 3D model and the query image. Experiments show that our method achieves higher localization accuracy than other learning-based approaches on multiple benchmarks. Code is available at https://github.com/JenningsL/nerf-loc.
Qiang Nie, Yong Liu 0032, Chengjie Wang 0001
ICRA2
2023 Lifting 2D Human Pose to 3D with Domain Adapted 3D Body Concept
Qiang Nie, Ziwei Liu 0002, Yun-Hui Liu 0001
Int. J. Comput. Vis.1
2022 Deterministic Point Cloud Registration via Novel Transformation Decomposition
abstract
Given a set of putative 3D-3D point correspondences, we aim to remove outliers and estimate rigid transformation with 6 degrees of freedom (DOF). Simultaneously estimating these 6 DOF is time-consuming due to high-dimensional parameter space. To solve this problem, it is common to decompose 6 DOF, i.e. independently compute 3-DOF rotation and 3-DOF translation. However, high non-linearity of 3-DOF rotation still limits the algorithm efficiency, especially when the number of correspondences is large. In contrast, we propose to decompose 6 DOF into$(2+1)$and$(1+2)\ DOF$. Specifically,$(2+1)DOF$represent 2-DOF rotation axis and 1-DOF displacement along this rotation axis.$(1+2)\ DOF$indicate 1-DOF rotation angle and 2-DOF displacement orthogonal to the above rotation axis. To compute these DOF, we design a novel two-stage strategy based on inlier set maximization. By leveraging branch and bound, we first search for$(2+1)\ DOF$, and then the remaining$(1+2)\ DOF$. Thanks to the proposed transformation decomposition and two-stage search strategy, our method is deterministic and leads to low computational complexity. We extensively compare our method with state-of-the-art approaches. Our method is more accurate and robust than the approaches that provide similar efficiency to ours. Our method is more efficient than the approaches whose accuracy and robustness are comparable to ours.
Wen Chen 0021, Haoang Li, Qiang Nie, Yun-Hui Liu 0001
CVPR3
2022 Weakly Supervised 3D Scene Segmentation with Region-Level Boundary Awareness and Instance Discrimination
Kangcheng Liu, Yuzhi Zhao, Qiang Nie, Zhi Gao 0005, Ben M. Chen
ECCV (28)3
2022 SoftPatch: Unsupervised Anomaly Detection with Noisy Data
abstract
Although mainstream unsupervised anomaly detection (AD) algorithms perform well in academic datasets, their performance is limited in practical application due to the ideal experimental setting of clean training data. Training with noisy data is an inevitable problem in real-world anomaly detection but is seldom discussed. This paper considers label-level noise in image sensory anomaly detection for the first time. To solve this problem, we proposed a memory-based unsupervised AD method, SoftPatch, which efficiently denoises the data at the patch level. Noise discriminators are utilized to generate outlier scores for patch-level noise elimination before coreset construction. The scores are then stored in the memory bank to soften the anomaly detection boundary. Compared with existing methods, SoftPatch maintains a strong modeling ability of normal data and alleviates the overconfidence problem in coreset. Comprehensive experiments in various noise scenes demonstrate that SoftPatch outperforms the state-of-the-art AD methods on the MVTecAD and BTAD benchmarks and is comparable to those methods under the setting without noise.
Xi Jiang 0009, Jinbao Wang 0001, Qiang Nie, Yong Liu 0032, Chengjie Wang 0001, Feng Zheng 0001
NeurIPS4
2021 Development of a Vision-Based Robotic Manipulation System for Transferring of Oocytes
abstract
Embryos/oocytes vitrification is an essential cryopreservation technique in IVF (in vitro fertilization) clinics. The reliable and effective transferring of embryos/oocytes is crucial to the subsequent steps in the whole procedure of vitrification. After each transferring, the straw needs to be replaced with a new one. Due to the uncertainties in the fabrication and installation, the exact knowledge of the kinematic model of the straw is usually unknown, and the relationship between the microscope and the straw is also unknown without calibration beforehand. In such situation, automatically transferring the oocytes from micropipette to the narrow tip of straw (0.7mm) is very challenging. In this paper, a new vision-guided robotic system is developed to automate the transferring of the oocyte without calibration. To this end, the unknown depth information is estimated then compensated by constructing a deep vision network through microscope image, and an approximate Jacobian control algorithm is also proposed to servo control the end tip of the uncalibrated straw to contact the micropipette with the vision feedback. After that, the oocyte is automatically transferred from the micropipette to the straw to finalize the task. The stability of the closed-loop control system is rigorously proved with Lyapunov methods, and the effectiveness of the developed robot is validated in experiments.
Shu Miao, Qiang Nie, Xin Jiang 0001, Xulin Sun, Jianjun Dai, Yun-Hui Liu 0001, Xiang Li 0009
IROS3
2021 View Transfer on Human Skeleton Pose: Automatically Disentangle the View-Variant and View-Invariant Information for Pose Representation Learning
Qiang Nie, Yun-Hui Liu 0001
Int. J. Comput. Vis.1
2020 Unsupervised 3D Human Pose Representation with Viewpoint and Pose Disentanglement
Qiang Nie, Ziwei Liu 0002, Yun-Hui Liu 0001
ECCV (19)1
2019 View-Invariant Human Action Recognition Based on a 3D Bio-Constrained Skeleton Model
abstract
Skeleton-based human action recognition has been a hot topic in recent years. Most existing studies are based on the skeleton data obtained from Kinect, which is noisy and unstable, in particular, in the case of occlusions. To cope with the noisy skeleton data and variation of viewpoints, this paper presents a view-invariant method for human action recognition by recovering the corrupted skeletons based on a 3D bio-constrained skeleton model and visualizing those body-level motion features obtained during the recovery process with images. The bio-constrained skeleton model is defined with two types of constraints: 1) constant bone lengths and 2) motion limits of joints. Based on the bio-constrained model, an effective method is proposed for skeleton recovery. Two types of new motion features, the Euclidean distance matrix between joints (JEDM), which contains the global structure information of the body, and the local dynamic variation of the joint Euler angles (JEAs) are used in describing human action. These two types of features are encoded into different motion images, which are fed into a two-stream convolutional neural network for learning different action patterns. The experiments on three benchmark datasets achieve better accuracy than the state-of-the-art approaches, which demonstrates the effectiveness of the proposed method.
Qiang Nie, Jiangliu Wang, Yun-Hui Liu 0001
IEEE Trans. Image Process.1