EDBT 2026 Demo / reviewers in the wild / expert
Zhongwei Qiu
dblp:246/5883
· DBLP profile ↗
20ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0001-9700-9658ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional disentangled information bottleneck for multi-view molecular graph representation learning
Jiaxin Dai, Dongmei Fu, Zhongwei Qiu, Lingwei Ma |
Expert Syst. Appl. | 3 |
| 2026 | Taylor-Sensus Network: Embracing Noise to Enlighten Uncertainty for Scientific DataabstractUncertainty estimation is vital for machine learning with scientific data. Although existing methods effectively address inherent model uncertainty, they frequently overlook explicit modeling of complex noise in data devoid of temporal or spatial dependencies. This gap is especially challenging in structured scientific data, where such dependencies are commonly lacking. To address these challenges in scientific research, we propose the Taylor-Sensus Network (TSNet). TSNet innovatively uses a Taylor series expansion to model complex, heteroscedastic noise and proposes a deep Taylor block for aware noise distribution. TSNet includes a noise-aware contrastive learning module and a data density perception module for aleatoric and epistemic uncertainty. Additionally, an uncertainty combination operator is used to integrate these uncertainties, and the network is trained using a novel heteroscedastic mean square error loss. TSNet demonstrates superior performance over mainstream and state-of-the-art methods in experiments, highlighting its potential in scientific research and noise resistance. Guangxuan Song, Dongmei Fu, Zhongwei Qiu, Jintao Meng 0003 |
IEEE Trans. Big Data | 3 |
| 2025 | Bridging Local Inductive Bias and Long-Range Dependencies With Pixel-Mamba for End-To-End Whole Slide Image Analysis
Zhongwei Qiu, Hanqing Chao, Tiancheng Lin 0004, Wanxing Chang, Zijiang Yang 0009, Wenpei Jiao, Yunshuo Zhang, Yelin Yang, Yun Bian, Ke Yan 0006, Dakai Jin, Le Lu 0001 |
ICCV | 1 |
| 2025 | Knowledge graph information bottleneck enhanced molecular representation learning
Jiaxin Dai, Dongmei Fu, Zhongwei Qiu, Lingwei Ma |
Neural Networks | 3 |
| 2025 | Gaussian-Based Swap Operator for Context-Aware Extraction of Building Boundary VectorsabstractAccurate extraction of building vector boundaries holds paramount importance within the domains of urban planning and Geographic Information Systems (GIS), providing indispensable support for urban construction endeavors and resource management initiatives. CNNs, while proficient in local feature extraction, often falter in capturing holistic, global image characteristics. Transformers excel in contextual feature comprehension but demand substantial computational resources and parameterization, impeding practical deployment. To address these challenges, this paper introduces an innovative computational operator known as G-Swap, which integrates Gaussian-distance-based feature correlation considerations, thereby significantly augmenting contextual comprehension within the computational framework. Additionally, a universal architecture for boundary vector extraction is proposed in this paper, comprising three primary components: 1) an Enhanced Backbone, integrating the G-Swap operator to enhance the backbone while bolstering model expressiveness; 2) a Decoder module, tasked with discriminating corner and edge features; and 3) a Two-branch Detection Head. Empirical experiments conducted on the Vectorizing World Building Dataset (VWB) underscore the model’s superior performance. Our G-Swap achieved F1 scores of 91.2% for vertices and 80.1% for edges, surpassing the previous state-of-the-art by 2.1% and 2.0% respectively. Moule Lin, Weipeng Jing 0001, Weitao Zou, Zhongwei Qiu, Chao Li 0066 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | MM-NeRF: Multimodal-Guided 3D Multi-Style Transfer of Neural Radiance Fieldabstract3D style transfer aims to generate stylized views of 3D scenes with specified styles, which requires high-quality generating and keeping multi-view consistency. Existing methods still suffer the challenges of high-quality stylization with texture details and stylization with multimodal guidance. In this paper, we reveal that the common training method of stylization with NeRF, which generates stylized multi-view supervision by 2D style transfer models, causes the same object in supervision to show various states (color tone, details, etc.) in different views, leading NeRF to tend to smooth the texture details, further resulting in low-quality rendering for 3D multi-style transfer. To tackle these problems, we propose a novel Multimodal-guided 3D Multi-style transfer of NeRF, termed MM-NeRF. First, MM-NeRF projects multimodal guidance into a unified space to keep the multimodal styles consistency and extracts multimodal features to guide the 3D stylization. Second, a novel multi-head learning scheme is proposed to relieve the difficulty of learning multi-style transfer, and a multi-view style consistent loss is proposed to track the inconsistency of multi-view supervision data. Finally, a novel incremental learning mechanism is proposed to generalize MM-NeRF to any new style with small costs. Extensive experiments on several real-world datasets show that MM-NeRF achieves high-quality 3D multi-style stylization with multimodal guidance, and keeps multi-view consistency and style consistency between multimodal guidance. Zijiang Yang 0009, Zhongwei Qiu, Chang Xu 0002, Dongmei Fu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | IHCSurv: Effective Immunohistochemistry Priors for Cancer Survival Analysis in Gigapixel Multi-stain Whole Slide Images
Yejia Zhang, Hanqing Chao, Zhongwei Qiu, Nishchal Sapkota, Pengfei Gu, Danny Ziyi Chen, Le Lu 0001, Ke Yan 0006, Dakai Jin, Yun Bian |
MICCAI (4) | 3 |
| 2024 | Pedestrian-Centric 3D Pre-collision Pose and Shape Estimation from Dashcam PerspectiveabstractPedestrian pre-collision pose is one of the key factors to determine the degree of pedestrian-vehicle injury in collision. Human pose estimation algorithm is an effective method to estimate pedestrian emergency pose from accident video. However, the pose estimation model trained by the existing daily human pose datasets has poor robustness under specific poses such as pedestrian pre-collision pose, and it is difficult to obtain human pose datasets in the wild scenes, especially lacking scarce data such as pedestrian pre-collision pose in traffic scenes. In this paper, we collect pedestrian-vehicle collision pose from the dashcam perspective of dashcam and construct the first Pedestrian-Vehicle Collision Pose dataset (PVCP) in a semi-automatic way, including 40k+ accident frames and 20K+ pedestrian pre-collision pose annotation (2D, 3D, Mesh). Further, we construct a Pedestrian Pre-collision Pose Estimation Network (PPSENet) to estimate the collision pose and shape sequence of pedestrians from pedestrian-vehicle accident videos. The PPSENet first estimates the 2D pose from the image (Image to Pose, ITP) and then lifts the 2D pose to 3D mesh (Pose to Mesh, PTM). Due to the small size of the dataset, we introduce a pre-training model that learns the human pose prior on a large number of pose datasets, and use iterative regression to estimate the pre-collision pose and shape of pedestrians. Further, we classify the pre-collision pose sequence and introduce pose class loss, which achieves the best accuracy compared with the existing relevant \textit{state-of-the-art} methods. Code and data are available for research at https://github.com/wmj142326/PVCP. Meijun Wang, Zhongwei Qiu, Pengxiaorui |
NeurIPS | 3 |
| 2024 | Multi-scale contrastive adaptor learning for segmenting anything in underperformed scenes
Zhongwei Qiu, Dongmei Fu |
Neurocomputing | 2 |
| 2023 | DMIS: Dynamic Mesh-Based Importance Sampling for Training Physics-Informed Neural NetworksabstractModeling dynamics in the form of partial differential equations (PDEs) is an effectual way to understand real-world physics processes. For complex physics systems, analytical solutions are not available and numerical solutions are widely-used. However, traditional numerical algorithms are computationally expensive and challenging in handling multiphysics systems. Recently, using neural networks to solve PDEs has made significant progress, called physics-informed neural networks (PINNs). PINNs encode physical laws into neural networks and learn the continuous solutions of PDEs. For the training of PINNs, existing methods suffer from the problems of inefficiency and unstable convergence, since the PDE residuals require calculating automatic differentiation. In this paper, we propose Dynamic Mesh-based Importance Sampling (DMIS) to tackle these problems. DMIS is a novel sampling scheme based on importance sampling, which constructs a dynamic triangular mesh to estimate sample weights efficiently. DMIS has broad applicability and can be easily integrated into existing methods. The evaluation of DMIS on three widely-used benchmarks shows that DMIS improves the convergence speed and accuracy in the meantime. Especially in solving the highly nonlinear Schrödinger Equation, compared with state-of-the-art methods, DMIS shows up to 46% smaller root mean square error and five times faster convergence speed. Code is available at https://github.com/MatrixBrain/DMIS. Zijiang Yang 0009, Zhongwei Qiu, Dongmei Fu |
AAAI | 2 |
| 2023 | PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation with Progressive Video TransformersabstractExisting methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then performs single-person PSE with temporal model. However, the global spatio-temporal context among spatial instances can not be captured. In this paper, we propose a new end-to-end multi-person 3D Pose and Shape estimation framework with progressive Video Transformer, termed PSVT. In PSVT, a spatio-temporal encoder (STE) captures the global feature dependencies among spatial objects. Then, spatio-temporal pose decoder (STPD) and shape decoder (STSD) capture the global dependencies between pose queries and feature tokens, shape queries and feature tokens, respectively. To handle the variances of objects as time proceeds, a novel scheme of progressive decoding is used to update pose and shape queries at each frame. Besides, we propose a novel pose-guided attention (PGA) for shape decoder to better predict shape parameters. The two components strengthen the decoder of PSVT to improve performance. Extensive experiments on the four datasets show that PSVT achieves stage-of-the-art results. Zhongwei Qiu, Qiansheng Yang, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Chang Xu 0002, Dongmei Fu, Jingdong Wang 0001 |
CVPR | 1 |
| 2023 | Stable Diffusion is UnstableabstractRecently, text-to-image models have been thriving. Despite their powerful generative capacity, our research has uncovered a lack of robustness in this generation process. Specifically, the introduction of small perturbations to the text prompts can result in the blending of primary subjects with other categories or their complete disappearance in the generated images. In this paper, we propose **Auto-attack on Text-to-image Models (ATM)**, a gradient-based approach, to effectively and efficiently generate such perturbations. By learning a Gumbel Softmax distribution, we can make the discrete process of word replacement or extension continuous, thus ensuring the differentiability of the perturbation generation. Once the distribution is learned, ATM can sample multiple attack samples simultaneously. These attack samples can prevent the generative model from generating the desired subjects without tampering with the category keywords in the prompt. ATM has achieved a 91.1\% success rate in short-text attacks and an 81.2\% success rate in long-text attacks. Further empirical analysis revealed three attack patterns based on: 1) variability in generation speed, 2) similarity of coarse-grained characteristics, and 3) polysemy of words. The code is available at https://github.com/duchengbin8/Stable_Diffusion_is_Unstable Chengbin Du, Yanxi Li 0001, Zhongwei Qiu, Chang Xu 0002 |
NeurIPS | 3 |
| 2023 | HAP: Structure-Aware Masked Image Modeling for Human-Centric PerceptionabstractModel pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight, we further incorporate an intuitive human structure prior - human parts - into pre-training. Specifically, we employ this prior to guide the mask sampling process. Image patches, corresponding to human part regions, have high priority to be masked out. This encourages the model to concentrate more on body structure information during pre-training, yielding substantial benefits across a range of human-centric perception tasks. To further capture human characteristics, we propose a structure-invariant alignment loss that enforces different masked views, guided by the human part prior, to be closely aligned for the same image. We term the entire method as HAP. HAP simply uses a plain ViT as the encoder yet establishes new state-of-the-art performance on 11 human-centric benchmarks, and on-par result on one dataset. For example, HAP achieves 78.1% mAP on MSMT17 for person re-identification, 86.54% mA on PA-100K for pedestrian attribute recognition, 78.2% AP on MS COCO for 2D pose estimation, and 56.0 PA-MPJPE on 3DPW for 3D pose and shape estimation. Junkun Yuan, Xinyu Zhang 0015, Hao Zhou 0039, Jian Wang 0066, Zhongwei Qiu, Zhiyin Shao, Shaofeng Zhang, Sifan Long 0001, Kun Kuang 0001, Junyu Han, Errui Ding, Lanfen Lin, Fei Wu 0001, Jingdong Wang 0001 |
NeurIPS | 5 |
| 2023 | Learning Degradation-Robust Spatiotemporal Frequency-Transformer for Video Super-ResolutionabstractVideo Super-Resolution (VSR) aims to restore high-resolution (HR) videos from low-resolution (LR) videos. Existing VSR techniques usually recover HR frames by extracting pertinent textures from nearby frames with known degradation processes. Despite significant progress, grand challenges remain to effectively extract and transmit high-quality textures from high-degraded low-quality sequences, such as blur, additive noises, and compression artifacts. This work proposes a novel degradation-robust Frequency-Transformer (FTVSR++) for handling low-quality videos that carry out self-attention in a combined space-time-frequency domain. First, video frames are split into patches and each patch is transformed into spectral maps in which each channel represents a frequency band. It permits a fine-grained self-attention on each frequency band so that real visual texture can be distinguished from artifacts. Second, a novel dual frequency attention (DFA) mechanism is proposed to capture the global and local frequency relations, which can handle different complicated degradation processes in real-world scenarios. Third, we explore different self-attention schemes for video processing in the frequency domain and discover that a "divided attention" which conducts joint space-frequency attention before applying temporal-frequency attention, leads to the best video enhancement quality. Extensive experiments on three widely-used VSR datasets show that FTVSR++ outperforms state-of-the-art methods on different low-quality videos with clear visual margins. Zhongwei Qiu, Huan Yang 0005, Jianlong Fu, Daochang Liu, Chang Xu 0002, Dongmei Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Weakly-supervised pre-training for 3D human pose estimation via perspective knowledgeabstractModern deep learning-based 3D pose estimation approaches require plenty of 3D pose annotations. However, existing 3D datasets lack diversity, which limits the performance of current methods and their generalization ability. Although existing methods utilize 2D pose annotations to help 3D pose estimation, they mainly focus on extracting 2D structural constraints from 2D poses, ignoring the 3D information hidden in the images. In this paper, we propose a novel method to extract weak 3D information directly from 2D images without 3D pose supervision. Firstly, we utilize 2D pose annotations and perspective prior knowledge to generate the relative depth of human joints . Then, we collect a 2D pose dataset (MCPC) and generate relative depth labels. Based on MCPC, we propose a weakly-supervised pre-training (WSP) strategy to distinguish the depth relationship between two points in an image. WSP enables the learning of the relative depth of two keypoints on lots of in-the-wild images, which is more capable of predicting depth and generalization ability for 3D human pose estimation. After fine-tuning the pose model on 3D pose datasets, WSP achieves state-of-the-art results on two widely-used benchmarks. Zhongwei Qiu, Kai Qiu 0001, Jianlong Fu, Dongmei Fu |
Pattern Recognit. | 1 |
| 2022 | Learning Spatiotemporal Frequency-Transformer for Compressed Video Super-Resolution
Zhongwei Qiu, Huan Yang 0005, Jianlong Fu, Dongmei Fu |
ECCV (18) | 1 |
| 2022 | Dynamic Graph Reasoning for Multi-person 3D Pose EstimationabstractMulti-person 3D pose estimation is a challenging task because of occlusion and depth ambiguity, especially in the cases of crowd scenes. To solve these problems, most existing methods explore modeling body context cues by enhancing feature representation with graph neural networks or adding structural constraints. However, these methods are not robust for their single-root formulation that decoding 3D poses from a root node with a pre-defined graph. In this paper, we propose GR-M3D, which models the Multi-person 3D pose estimation with dynamic Graph Reasoning. The decoding graph in GR-M3D is predicted instead of pre-defined. In particular, It firstly generates several data maps and enhances them with a scale and depth aware refinement module (SDAR). Then multiple root keypoints and dense decoding paths for each person are estimated from these data maps. Based on them, dynamic decoding graphs are built by assigning path weights to the decoding paths, while the path weights are inferred from those enhanced data maps. And this process is named dynamic graph reasoning (DGR). Finally, the 3D poses are decoded according to dynamic decoding graphs for each detected person. GR-M3D can adjust the structure of the decoding graph implicitly by adopting soft path weights according to input data, which makes the decoding graphs be adaptive to different input persons to the best extent and more capable of handling occlusion and depth ambiguity than previous methods. We empirically show that the proposed bottom-up approach even outperforms top-down methods and achieves state-of-the-art results on three 3D pose datasets. Zhongwei Qiu, Qiansheng Yang, Jian Wang 0066, Dongmei Fu |
ACM Multimedia | 1 |
| 2022 | IVT: An End-to-End Instance-guided Video Transformer for 3D Pose EstimationabstractVideo 3D human pose estimation aims to localize the 3D coordinates of human joints from videos. Recent transformer-based approaches focus on capturing the spatiotemporal information from sequential 2D poses, which cannot model the contextual depth feature effectively since the visual depth features are lost in the step of 2D pose estimation. In this paper, we simplify the paradigm into an end-to-end framework, Instance-guided Video Transformer (IVT), which enables learning spatiotemporal contextual depth information from visual features effectively and predicts 3D poses directly from video frames. In particular, we firstly formulate video frames as a series of instance-guided tokens and each token is in charge of predicting the 3D pose of a human instance. These tokens contain body structure information since they are extracted by the guidance of joint offsets from the human center to the corresponding body joints. Then, these tokens are sent into IVT for learning spatiotemporal contextual depth. In addition, we propose a cross-scale instance-guided attention mechanism to handle the variational scales among multiple persons. Finally, the 3D poses of each person are decoded from instance-guided tokens by coordinate regression. Experiments on three widely-used 3D pose estimation benchmarks show that the proposed IVT achieves state-of-the-art performances. Zhongwei Qiu, Qiansheng Yang, Jian Wang 0066, Dongmei Fu |
ACM Multimedia | 1 |
| 2020 | DGCN: Dynamic Graph Convolutional Network for Efficient Multi-Person Pose EstimationabstractMulti-person pose estimation aims to detect human keypoints from images with multiple persons. Bottom-up methods for multi-person pose estimation have attracted extensive attention, owing to the good balance between efficiency and accuracy. Recent bottom-up methods usually follow the principle of keypoints localization and grouping, where relations between keypoints are the keys to group keypoints. These relations spontaneously construct a graph of keypoints, where the edges represent the relations between two nodes (i.e., keypoints). Existing bottom-up methods mainly define relations by empirically picking out edges from this graph, while omitting edges that may contain useful semantic relations. In this paper, we propose a novel Dynamic Graph Convolutional Module (DGCM) to model rich relations in the keypoints graph. Specifically, we take into account all relations (all edges of the graph) and construct dynamic graphs to tolerate large variations of human pose. The DGCM is quite lightweight, which allows it to be stacked like a pyramid architecture and learn structural relations from multi-level features. Our network with single DGCM based on ResNet-50 achieves relative gains of 3.2% and 4.8% over state-of-the-art bottom-up methods on COCO keypoints and MPII dataset, respectively. Zhongwei Qiu, Kai Qiu 0001, Jianlong Fu, Dongmei Fu |
AAAI | 1 |
| 2019 | Learning Recurrent Structure-Guided Attention Network for Multi-person Pose EstimationabstractMulti-person pose estimation aims to localize tens of human joints (e.g., elbow, wrist, etc.) from multiple human bodies in an image. Existing approaches mainly adopt a two stage pipeline, which usually consists of a human detector (i.e., generating a bounding box for each person) and a single person pose estimator (i.e., generating human joints from each bounding box). However, these approaches neglect the challenges of large pose variations and heavy occlusions in each bounding box, which often results in imprecise human joint localization. In this paper, we propose a structure-guided attention network (SGAN) for multi-person pose estimation. Specifically, a structured pose representation is encoded by learning a joint confidence map and a joint association map, which can be further refined by a structure-guided attention network (SGAN) in a recurrent way. Note that SGAN enables a deep neural network to take initial pose estimation as references, and to discover multi-scale pose features as completion, and thus the learning of pose structures can be reinforced. Extensive experiments show the best single-model results against the state-of-the-art approaches, with a relative 3.5% mAP gain in the challenging COCO Keypoint dataset. Zhongwei Qiu, Kai Qiu 0001, Jianlong Fu, Dongmei Fu |
ICME | 1 |