Zitian Wang

dblp:46/7299 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection
abstract
The rise of autonomous vehicles has significantly increased the demand for robust 3D object detection systems. While cameras and LiDAR sensors each offer unique advantages-cameras provide rich texture information and LiDAR offers precise 3D spatial data-relying on a single modality often leads to performance limitations. This paper introduces MV2DFusion, a multi-modal detection framework that integrates the strengths of both worlds through an advanced query-based fusion mechanism. By introducing an image query generator to align with image-specific attributes and a point cloud query generator, MV2DFusion effectively combines modality-specific object semantics without biasing toward one single modality. Then the sparse fusion process can be accomplished based on the valuable object semantics, ensuring efficient and accurate object detection across various scenarios. Our framework's flexibility allows it to integrate with any image and point cloud-based detectors, showcasing its adaptability and potential for future advancements. Extensive evaluations on the nuScenes and Argoverse2 datasets demonstrate that MV2DFusion achieves state-of-the-art performance, particularly excelling in long-range detection scenarios.
Zitian Wang, Zehao Huang, Yulu Gao, Naiyan Wang, Si Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 PIESCCA: A Chinese Dataset for Pharmacovigilance Information Extraction Based on Standardised Case Causality Assessment of National Medical Products Administration
abstract
Adverse drug reactions (ADRs) can reduce therapeutic efficacy and pose serious risks. Pharmacovigilance monitors and prevents ADRs, but current evaluations rely on manual text assessments, causing inefficiency and potential errors. To address this, we introduce PIESCCA, a Chinese dataset for pharmacovigilance information extraction aligned with the Standardised Case Causality Assessment (SCCA) guidelines from the NMPA. PIESCCA frames causality assessment as four NLP tasks: named entity recognition (NER), relation extraction (RE), event extraction (EE), and question answering (QA). It is curated from over 50,000 real-world reports collected in the past five years, with$2,100+$representative cases selected and annotated under strict quality standards. We benchmark PIESCCA using four LLMs and seven PLM-based models, establishing a foundation for automated pharmacovigilance research. Our code and data are available at https://github.com/deeomnjobs/PIESCCA.
Kang Xu 0001, Peihan Cai, Liqun Lu, Zitian Wang, Lingli Jiao, Danhua Ma, Zhenjiang Dong
BIBM4
2025 Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
abstract
Preference alignment has emerged as an effective strategy to enhance the performance of Multimodal Large Language Models (MLLMs) following supervised fine-tuning. While existing preference alignment methods predominantly target hallucination factors, they overlook the factors essential for multi-modal comprehension capabilities, often narrowing their improvements on hallucination mitigation. To bridge this gap, we propose Instruction-oriented Preference Alignment (IPA), a scalable framework designed to automatically construct alignment preferences grounded in instruction fulfillment efficacy. Our method involves an automated preference construction coupled with a dedicated verification process that identifies instruction-oriented factors, avoiding significant variability in response representations. Additionally, IPA incorporates a progressive preference collection pipeline, further recalling challenging samples through model self-evolution and reference-guided refinement. Experiments conducted on Qwen2VL-7B demonstrate IPA's effectiveness across multiple benchmarks, including hallucination evaluation, visual question answering, and text understanding tasks, highlighting its capability to enhance general comprehension.
Zitian Wang, Yue Liao, Kang Rong, Fengyun Rao, Si Liu 0001
ICCV1
2025 MH-V2X: Robust Multi-hypothesis V2X Perception via Probabilistic Fusion
Zitian Wang
PRCV (17)1
2025 MSAttnFlow: Normalizing flow for unsupervised anomaly detection with multi-scale attention
Zhengnan Hu, Zhou-Ping Yin, Erli Meng, Leyan Zhu, Zitian Wang
Pattern Recognit.8
2024 Eliminating Cross-modal Conflicts in BEV Space for LiDAR-Camera 3D Object Detection
abstract
Recent 3D object detectors typically utilize multi-sensor data and unify multi-modal features in the shared bird’s-eye view (BEV) representation space. However, our empirical findings indicate that previous methods have limitations in generating fusion BEV features free from cross-modal conflicts. These conflicts encompass extrinsic conflicts caused by BEV feature construction and inherent conflicts stemming from heterogeneous sensor signals. Therefore, we propose a novel Eliminating Conflicts Fusion (ECFusion) method to explicitly eliminate the extrinsic/inherent conflicts in BEV space and produce improved multi-modal BEV features. Specifically, we devise a Semantic-guided Flow-based Alignment (SFA) module to resolve extrinsic conflicts via unifying spatial distribution in BEV space before fusion. Moreover, we design a Dissolved Query Recovering (DQR) mechanism to remedy inherent conflicts by preserving objectness clues that are lost in the fusion BEV feature. In general, our method maximizes the effective information utilization of each modality and leverages inter-modal complementarity. Our method achieves state-of-the-art performance in the highly competitive nuScenes 3D object detection dataset. The code is released at https://github.com/fjhzhixi/ECFusion.
Jiahui Fu 0003, Chen Gao 0005, Zitian Wang, Lirong Yang, Beipeng Mu, Si Liu 0001
ICRA3
2024 Multi-Person Pose Regression With Distribution-Aware Single-Stage Models
abstract
Understanding human posture is a challenging topic, which encompasses several tasks, e.g., pose estimation, body mesh recovery and pose tracking. In this article, we propose a novel Distribution-Aware Single-stage (DAS) model for the pose-related tasks. The proposed DAS model estimates human position and localizes joints simultaneously, which requires only a single pass. Meanwhile, we utilize normalizing flow to enable DAS to learn the true distribution of joint locations, rather than making simple Gaussian or Laplacian assumptions. This provides a pivotal prior and greatly boosts the accuracy of regression-based methods, thus making DAS achieve comparable performance to the volumetric-based methods. We also introduce a recursively update strategy to progressively approach the regression target, reducing the difficulty of regression and improving the regression performance. We further adapt DAS to multi-person mesh recovery and pose tracking tasks and achieve considerable performance on both tasks. Comprehensive experiments on CMU Panoptic and MuPoTS-3D demonstrate the superior efficiency of DAS, specifically 1.5 times speedup over previous best method, and its state-of-the-art accuracy for multi-person pose estimation. Extensive experiments on 3DPW and PoseTrack2018 indicate the effectiveness and efficiency of DAS for human body mesh recovery and pose tracking, respectively, which prove the generality of our proposed DAS model.
Leyan Zhu, Zitian Wang, Si Liu 0001, Xuecheng Nie, Luoqi Liu, Bo Li 0006
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Object as Query: Lifting any 2D Object Detector to 3D Detection
abstract
3D object detection from multi-view images has drawn much attention over the past few years. Existing methods mainly establish 3D representations from multi-view images and adopt a dense detection head for object detection, or employ object queries distributed in 3D space to localize objects. In this paper, we design Multi-View 2D Objects guided 3D Object Detector (MV2D), which can lift any 2D object detector to multi-view 3D object detection. Since 2D detections can provide valuable priors for object existence, MV2D exploits 2D detectors to generate object queries conditioned on the rich image semantics. These dynamically generated queries help MV2D to recall objects in the field of view and show a strong capability of localizing 3D objects. For the generated queries, we design a sparse cross attention module to force them to focus on the features of specific objects, which suppresses interference from noises. The evaluation results on the nuScenes dataset demonstrate the dynamic object queries and sparse feature aggregation can promote 3D detection capability. MV2D also exhibits a state-of-the-art performance among existing methods. We hope MV2D can serve as a new baseline for future research. Code is available at https://github.com/tusen-ai/MV2D.
Zitian Wang, Zehao Huang, Jiahui Fu 0003, Naiyan Wang, Si Liu 0001
ICCV1
2023 VcT: Visual Change Transformer for Remote Sensing Image Change Detection
abstract
Given two remote sensing images, the goal of visual change detection task is to detect significantly changed areas between them. Existing visual change detectors usually adopt CNNs or Transformers for feature representation learning and focus on learning effective representation for the changed regions between images. Although good performance can be obtained by enhancing the features of the change regions, however, these works are still limited mainly due to the ignorance of mining the unchanged background context information. It is known that one main challenge for change detection is how to obtain the consistent representations for two images involving different variations, such as spatial variation, sunlight intensity, etc. In this work, we demonstrate that carefully mining the common background information provides an important cue to learn the consistent representations for the two images which thus obviously facilitates the visual change detection problem. Based on this observation, we propose a novel Visual change Transformer (VcT) model for visual change detection problem. To be specific, a shared backbone network is first used to extract the feature maps for the given image pair. Then, each pixel of feature map is regarded as a graph node and the graph neural network is proposed to model the structured information for coarse change map prediction. Top-K reliable tokens can be mined from the map and refined by using the clustering algorithm. Then, these reliable tokens are enhanced by first utilizing self/cross-attention schemes and then interacting with original features via an anchor-primary attention learning module. Finally, the prediction head is proposed to get a more accurate change map. Extensive experiments on multiple benchmark datasets validated the effectiveness of our proposed VcT model. The source code and pre-trained models are available at https://github.com/Event-AHU/VcT_Remote_Sensing_Change_Detection.
Bo Jiang 0002, Zitian Wang, Xixi Wang 0005, Lan Chen 0003, Xiao Wang 0014, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Distribution-Aware Single-Stage Models for Multi-Person 3D Pose Estimation
abstract
In this paper, we present a novel Distribution-Aware Single-stage (DAS) model for tackling the challenging multi-person 3D pose estimation problem. Different from existing top-down and bottom-up methods, the proposed DAS model simultaneously localizes person positions and their corresponding body joints in the 3D camera space in a one-pass manner. This leads to a simplified pipeline with enhanced efficiency. In addition, DAS learns the true distribution of body joints for the regression of their positions, rather than making a simple Laplacian or Gaussian assumption as previous works. This provides valuable priors for model prediction and thus boosts the regression-based scheme to achieve competitive performance with volumetric-base ones. Moreover, DAS exploits a recur-sive update strategy for progressively approaching to regression target, alleviating the optimization difficulty and further lifting the regression performance. DAS is implemented with a fully Convolutional Neural Network and end-to-end learnable. Comprehensive experiments on benchmarks CMU Panoptic and MuPoTS-3D demonstrate the superior efficiency of the proposed DAS model, specifically 1.5x speedup over previous best model, and its stat-of-the-art accuracy for multi-person 3D pose estimation.
Zitian Wang, Xuecheng Nie, Xiaochao Qu, Yunpeng Chen, Si Liu 0001
CVPR1
2022 Human-Centric Relation Segmentation: Dataset and Solution
abstract
Vision and language understanding techniques have achieved remarkable progress, but currently it is still difficult to well handle problems involving very fine-grained details. For example, when the robot is told to "bring me the book in the girl's left hand", most existing methods would fail if the girl holds one book respectively in her left and right hand. In this work, we introduce a new task named human-centric relation segmentation (HRS), as a fine-grained case of HOI-det. HRS aims to predict the relations between the human and surrounding entities and identify the relation-correlated human parts, which are represented as pixel-level masks. For the above exemplar case, our HRS task produces results in the form of relation triplets 〈girl [left hand], hold, book 〉 and exacts segmentation masks of the book, with which the robot can easily accomplish the grabbing task. Correspondingly, we collect a new Person In Context (PIC) dataset for this new task, which contains 17,122 high-resolution images and densely annotated entity segmentation and relations, including 141 object categories, 23 relation categories and 25 semantic human parts. We also propose a Simultaneous Matching and Segmentation (SMS) framework as a solution to the HRS task. It contains three parallel branches for entity segmentation, subject object matching and human parsing respectively. Specifically, the entity segmentation branch obtains entity masks by dynamically-generated conditional convolutions; the subject object matching branch detects the existence of any relations, links the corresponding subjects and objects by displacement estimation and classifies the interacted human parts; and the human parsing branch generates the pixelwise human part labels. Outputs of the three branches are fused to produce the final HRS results. Extensive experiments on PIC and V-COCO datasets show that the proposed SMS method outperforms baselines with the 36 FPS inference speed. Notably, SMS outperforms the best performing baseline m-KERN with only 17.6 percent time cost. The dataset and code will be released at http://picdataset.com/challenge/index/.
Si Liu 0001, Zitian Wang, Yulu Gao, Lejian Ren, Yue Liao, Guanghui Ren, Bo Li 0006, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Lvio-Fusion: A Self-adaptive Multi-sensor Fusion SLAM Framework Using Actor-critic Method
abstract
State estimation with sensors is essential for mobile robots. Due to different performance of sensors in different environments, how to fuse measurements of various sensors is a problem. In this paper, we propose a tightly coupled multi-sensor fusion framework, Lvio-Fusion, which fuses stereo camera, Lidar, IMU, and GPS based on the graph optimization. Especially for urban traffic scenes, we introduce a segmented global pose graph optimization with GPS and loop-closure, which can eliminate accumulated drifts. Additionally, we creatively use a actor-critic method in reinforcement learning to adaptively adjust sensors’ weight. After training, actor-critic agent can provide the system better and dynamic sensors’ weight. We evaluate the performance of our system on public datasets and compare it with other state-of-the-art methods, which shows that the proposed method achieves high estimation accuracy and robustness to various environments. And our implementations are open source and highly scalable.
Yupeng Jia, Haiyong Luo, Fang Zhao 0003, Guanlin Jiang, Jiaquan Yan, Zhuqing Jiang, Zitian Wang
IROS8
2021 Attack Traffic Detection Based on LetNet-5 and GRU Hierarchical Deep Neural Network
Zitian Wang, Zesong Wang 0001, FangZhou Yi
WASA (3)1
2018 Joint and Competitive Caching Designs in Large-Scale Multi-Tier Wireless Multicasting Networks
abstract
Caching and multicasting are two promising methods to support massive content delivery in multi-tier wireless networks. In this paper, we consider a random caching and multicasting scheme with caching distributions in the two tiers as design parameters, to achieve efficient content dissemination in a two-tier large-scale cache-enabled wireless multicasting network. First, we derive tractable expressions for the successful transmission probabilities in the general region as well as the high signal-to-noise ratio (SNR) and high user density region, respectively, utilizing tools from stochastic geometry. Then, for the case of a single operator for the two tiers, we formulate the optimal joint caching design problem to maximize the successful transmission probability in the asymptotic region, which is nonconvex in general. By using the block successive approximate optimization technique, we develop an iterative algorithm, which is shown to converge to a stationary point. Next, for the case of two different operators, one for each tier, we formulate the competitive caching design game where each tier maximizes its successful transmission probability in the asymptotic region. We show that the game has a unique Nash equilibrium (NE) and adopt an iterative algorithm, which is shown to converge to the NE under a mild condition. Finally, by numerical simulations, we show that the proposed designs achieve significant gains over existing schemes.
Ying Cui 0001, Zitian Wang, Yang Yang 0033, Feng Yang 0006, Lianghui Ding, Liang Qian
IEEE Trans. Commun.2
2017 Joint and Competitive Caching Designs in Large-Scale Multi-Tier Wireless Multicasting Networks
abstract
Caching and multicasting are two promising methods to support massive content delivery in multi-tier wireless networks. In this paper, we consider a random caching and multicasting scheme with caching distributions in the two tiers as design parameters, to achieve efficient content dissemination in a two- tier large-scale cache-enabled wireless multicasting network. First, we derive tractable expressions for the successful transmission probabilities in the general region as well as the high SNR and high user density region, respectively, utilizing tools from stochastic geometry. Then, for the case of a single operator for the two tiers, we formulate the optimal joint caching design problem to maximize the successful transmission probability in the asymptotic region, which is nonconvex in general. By using the block successive approximate optimization technique, we develop an iterative algorithm, which is shown to coverage to a stationary point. Next, for the case of two different operators, one for each tier, we formulate the competitive caching design game where each tier maximizes its successful transmission probability in the asymptotic region. We show that the game has a unique Nash equilibrium (NE) and develop an iterative algorithm, which is shown to converge to the NE under a mild condition. Finally, by numerical simulations, we show that the proposed designs achieve significant gains over existing schemes.
Zitian Wang, Zhehan Cao, Ying Cui 0001, Yang Yang 0033
GLOBECOM1
2015 Stock market trend prediction using dynamical Bayesian factor graph
Zitian Wang, Shaohua Tan
Expert Syst. Appl.2
2014 Efficient and effective Bayesian network local structure learning
Yunhai Tong, Zitian Wang, Shaohua Tan
Frontiers Comput. Sci.3
2012 Continuous variable based Bayesian network structure learning from financial factors
abstract
In this paper, for the discovery the interrelationship of financial factors, we present a two-step accelerated method in learning the structure of Bayesian networks without making parametric assumptions for continuous domains. Our approach divides the high dimensional space into an uniform grid, over which the density can be estimated in an efficient way by using compact support kernels. Local scores are then estimated by the iterative Monte Carlo approximation method with rigorous relative error control. Empirical studies on 15 US financial factors show the efficiency and effectiveness of our method.
Zitian Wang, Bingwu Liu, Shaohua Tan
CIFEr2
2012 Linear non-Gaussian causal discovery from a composite set of major US macroeconomic factors
Zitian Wang, Shaohua Tan
Expert Syst. Appl.2
2009 Automatic linear causal relationship identification for financial factor modeling
Zitian Wang, Shaohua Tan
Expert Syst. Appl.1
2009 Identifying idiosyncratic stock return indicators from large financial factor set via least angle regression
Zitian Wang, Shaohua Tan
Expert Syst. Appl.1