Jiguang Zhang

dblp:195/8013 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
24since 2021 · last 2026
0000-0002-8212-1361ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 18 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
abstract
Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance. However, these pipelines rely on implicit modeling that uses frame-level or fragmented video features, failing to capture the temporal coherence across event sequences and comprehensive semantics within visual contexts. To address this, we propose an explicit temporal-semantic modeling framework called Context-Aware Cross-Modal Interaction (CACMI), which leverages both latent temporal characteristics within videos and linguistic semantics from text corpus. Specifically, our model consists of two core components: Cross-modal Frame Aggregation aggregates relevant frames to extract temporally coherent, event-aligned textual features through cross-modal retrieval; and Context-aware Feature Enhancement utilizes query-guided attention to integrate visual dynamics with pseudo-event semantics. Extensive experiments on the ActivityNet Captions and YouCook2 datasets demonstrate that CACMI achieves the state-of-the-art performance on dense video captioning task.
Mingda Jia, Weiliang Meng, Zenghuang Fu, Ju Xin, Rongtao Xu, Jiguang Zhang, Xiaopeng Zhang 0001
AAAI9
2026 VPGS : Virtual plant reconstruction and rendering based on 3D Gaussian Splatting
Weilong Ding 0001, Yun Sang, Lifeng Xu, Jiguang Zhang
Comput. Graph.6
2026 Explicit to implicit presentation for 3D unbounded open scenes reconstruction: the survey
Jiguang Zhang, Weiliang Meng, Zhaohui Zhang 0002, Xiaopeng Zhang 0001
Expert Syst. Appl.2
2026 Adaptive in Adapter: Boosting Open-Vocabulary Semantic Segmentation With Adaptive Dropout Adapter
abstract
Open-vocabulary semantic segmentation is a challenging multimedia task that requires segmentation and recognition of unseen word classes during the testing phase. Recent works bridge the gap between closed and open-vocabulary recognition by introducing large-scale visual language models such as CLIP with cross-modal alignment capabilities. To preserve multimodal alignment capabilities, it is common to freeze the parameters of the CLIP and then add additional learnable components such as adapters to expand to downstream tasks. However, for the open-vocabulary semantic segmentation task, the plain adapter suffers from overfitting the closed-vocabulary classes and impairs performance on the open-vocabulary unseen classes. In addition, since CLIP is trained to perform image-level alignment can cause the network to over-focus on partially discriminative regions, resulting in incomplete segmentation masks. To alleviate the above problems, we introduce adaptive dropout adapters to release theAdaptiveInAdapter (i.e.AIA) from the following two aspects:i)A Generalization Feature Selection Adapter (GFSA) is proposed to improve the generalization of network over unseen classes.ii)A Discriminative Region Mask Adapter (DRMA) is proposed for retrofitting CLIP backbone, has provided region free biased features for segmentation mask generation. Meanwhile, our proposed AIA achieves the current state-of-the-art performance on several open-vocabulary semantic segmentation benchmarks. Code is available athttps://github.com/clearxu/AIA.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Jiguang Zhang, Xiaoqiang Teng, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.6
2026 Robust detection in complex construction sites: HiPA-DETR with weather-aware and cross-domain generalization
Zenghuang Fu, Muyang Zhang, Changwei Wang 0001, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Vis. Comput.8
2025 Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
abstract
Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual aliasing because of easy overlap. However, existing methods lack precisely modeling and full exploitation of these discriminative regions. In addition, the lack of pixel-level correspondence supervision in the VPR dataset hinders further improvement of the local feature matching capability in the re-ranking stage. In this paper, we propose the Focus on Local (FoL) approach to stimulate the performance of image retrieval and re-ranking in VPR simultaneously by mining and exploiting reliable discriminative local regions in images and introducing pseudo-correlation supervision. First, we design two losses, Extraction-Aggregation Spatial Alignment Loss (SAL) and Foreground-Background Contrast Enhancement Loss (CEL), to explicitly model reliable discriminative local regions and use them to guide the generation of global representations and efficient re-ranking. Second, we introduce a weakly-supervised local feature training strategy based on pseudo-correspondences obtained from aggregating global features to alleviate the lack of local correspondences ground truth for the VPR task. Third, we suggest an efficient re-ranking pipeline that is efficiently and precisely based on discriminative region guidance. Finally, experimental results show that our FoL achieves the state-of-the-art on multiple VPR benchmarks in both image retrieval and re-ranking stages and also significantly outperforms existing two-stage VPR methods in terms of computational efficiency.
Changwei Wang 0001, Shunpeng Chen, Rongtao Xu, Jiguang Zhang, Haoran Yang 0003, Yu Zhang 0133, Kexue Fu 0001, Shide Du, Zhiwei Xu 0005, Longxiang Gao, Li Guo 0004, Shibiao Xu
AAAI6
2025 Novel View Synthesis Under Large-Deviation Viewpoint for Autonomous Driving
abstract
Novel view synthesis is a critical task in autonomous driving. Although 3D Gaussian Splatting (3D-GS) has shown success in generating novel views, it faces challenges in maintaining high-quality rendering when viewpoints deviate significantly from the training set. This difficulty primarily stems from complex lighting conditions and geometric inconsistencies in texture-less regions. To address these issues, we propose an attention-based illumination model that leverages light fields from neighboring views, enhancing the realism of synthesized images. Additionally, we propose a geometry optimization method using planar homography to improve geometric consistency in texture-less regions. Our experiments demonstrate substantial improvements in synthesis quality for large-deviation viewpoints, validating the effectiveness of our approach.
Jiguang Zhang, Shibiao Xu, Chengwei Pan
AAAI2
2025 AccidentX: A Large-Scale Multimodal BEV Dataset for Traffic Accident Analysis and Prevention
abstract
With the rapid development and widespread application of autonomous driving technology, the accurate analysis and prevention of traffic accidents have become critical challenges. However, current traffic accident datasets are often constrained by limited scale and diversity, impeding progress in this field. To address these limitations, we introduce AccidentX, a large-scale multimodal dataset specifically curated for comprehensive traffic accident analysis and prevention. Our AccidentX comprises over 10,000 bird’s-eye view (BEV) videos generated using the CARLA simulator, with detailed annotations covering a wide range of traffic scenarios. In comparison to existing datasets such as nuScenes, our AccidentX offers seven times more video frames and leverages Vision-Language Models (VLMs) and GPT-4o for enhanced scene understanding and decision-making. We also establish a benchmark for state-of-the-art Multimodal Large Language Models (MLLMs) on AccidentX, fostering further research and innovation within the community. AccidentX will be made available as a fully open source resource for the advancement of the autonomous driving safety algorithm community.
Muyang Zhang, Mingda Jia, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
IROS7
2025 C2Fi-NeRF: Coarse to fine inversion NeRF for 6D pose estimation
Jiguang Zhang, Zhaohui Zhang 0002, Xuxiang Feng, Shibiao Xu, Rongtao Xu, Changwei Wang 0001, Kexue Fu 0001, Jiaxi Sun, Weilong Ding 0001
Expert Syst. Appl.1
2025 ROMOT: Referring-expression-comprehension open-set multi-object tracking
Wei Li 0237, Bowen Li 0014, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Vis. Comput.5
2024 MIM-HD: Making Smaller Masked Autoencoder Better with Efficient Distillation
abstract
Self-supervised learning and knowledge distillation intersect to achieve exceptional performance on downstream tasks across diverse network capacities. This paper introduces MIM-HD, which implements enhancements for masked image modeling (MIM) distillation, in two key aspects. First, a vision transformer head-level relation adaptive distillation approach is proposed, allowing the student to dynamically draw multi-source knowledge from the teacher based on its evolving state, compatible with scenarios where teacher-student transformer block head count differs. Second, to address the overemphasis on the encoder and neglect of the decoder role in maintaining representation consistency in previous MIM distillations, a dual-view decoding strategy for latent visual representations is introduced, reusing the teacher’s decoder to alleviate MIM burdens on smaller networks. MIM-HD effectiveness is demonstrated through evaluations on ADE20K (mIoU) and ImageNet-1K (Acc), achieving +1.4% and +0.5% improved performance, respectively, compared to state-of-the-art methods, with substantial advantages on smaller pre-training datasets. Moreover, MIM-HD achieves superior efficiency, reducing pre-training epochs from 300 to 100.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Li Guo 0004, Jiguang Zhang, Xiaoqiang Teng, Wenbo Xu 0003
ECAI7
2024 HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
abstract
Infrared small object detection is an important computer vision task involving the recognition and localization of tiny objects in infrared images, which usually contain only a few pixels. However, it encounters difficulties due to the diminutive size of the objects and the generally complex backgrounds in infrared images. In this paper, we propose a deep learning method, HCF-Net, that significantly improves infrared small object detection performance through multiple practical modules. Specifically, it includes the parallelized patch-aware attention (PPA) module, dimension-aware selective integration (DASI) module, and multi-dilated channel refiner (MDCR) module. The PPA module uses a multi-branch feature extraction strategy to capture feature information at different scales and levels. The DASI module enables adaptive channel selection and fusion. The MDCR module captures spatial features of different receptive field ranges through multiple depth-separable convolutional layers. Extensive experimental results on the SIRST infrared single-frame image dataset show that the proposed HCF-Net performs well, surpassing other traditional and deep learning models. Code is available at https://github.com/zhengshuchen/HCFNet.
Shibiao Xu, ShuChen Zheng, Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Xiaoqiang Teng, Ao Li 0002, Li Guo 0004
ICME6
2024 Fake-GPT: Detecting Fake Image via Large Language Model
Yuming Fan, Dongming Yang, Jiguang Zhang, Bang Yang, Yuexian Zou
PRCV (8)3
2024 FEKNN: A Wi-Fi Indoor Localization Method Based on Feature Enhancement and KNN
Bowen Li 0014, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
WASA (1)5
2024 SocialVis: Dynamic social visualization in dense scenes via real-time multi-object tracking and proximity graph construction
abstract
Abstract To monitor and assess social dynamics and risks at large gatherings, we propose “SocialVis,” a comprehensive monitoring system based on multi‐object tracking and graph analysis techniques. Our SocialVis includes a camera detection system that operates in two modes: a real‐time mode, which enables participants to track and identify close contacts instantly, and an offline mode that allows for more comprehensive post‐event analysis. The dual functionality not only aids in preventing mass gatherings or overcrowding by enabling the issuance of alerts and recommendations to organizers, but also allows for the generation of proximity‐based graphs that map participant interactions, thereby enhancing the understanding of social dynamics and identifying potential high‐risk areas. It also provides tools for analyzing pedestrian flow statistics and visualizing paths, offering valuable insights into crowd density and interaction patterns. To enhance system performance, we designed the SocialDetect algorithm in conjunction with the BYTE tracking algorithm. This combination is specifically engineered to improve detection accuracy and minimize ID switches among tracked objects, leveraging the strengths of both algorithms. Experiments on both public and real‐world datasets validate that our SocialVis outperforms existing methods, showing improvement in detection accuracy and reduction in ID switches in dense pedestrian scenarios.
Bowen Li 0014, Wei Li 0237, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds5
2024 HIDE: Hierarchical iterative decoding enhancement for multi-view 3D human parameter regression
abstract
Abstract Parametric human modeling are limited to either single‐view frameworks or simple multi‐view frameworks, failing to fully leverage the advantages of easily trainable single‐view networks and the occlusion‐resistant capabilities of multi‐view images. The prevalent presence of object occlusion and self‐occlusion in real‐world scenarios leads to issues of robustness and accuracy in predicting human body parameters. Additionally, many methods overlook the spatial connectivity of human joints in the global estimation of model pose parameters, resulting in cumulative errors in continuous joint parameters.To address these challenges, we propose a flexible and efficient iterative decoding strategy. By extending from single‐view images to multi‐view video inputs, we achieve local‐to‐global optimization. We utilize attention mechanisms to capture the rotational dependencies between any node in the human body and all its ancestor nodes, thereby enhancing pose decoding capability. We employ a parameter‐level iterative fusion of multi‐view image data to achieve flexible integration of global pose information, rapidly obtaining appropriate projection features from different viewpoints, ultimately resulting in precise parameter estimation. Through experiments, we validate the effectiveness of the HIDE method on the Human3.6M and 3DPW datasets, demonstrating significantly improved visualization results compared to previous methods.
Weitao Lin, Jiguang Zhang, Weiliang Meng, Xianglong Liu 0007, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds2
2024 Key-point-guided adaptive convolution and instance normalization for continuous transitive face reenactment of any person
abstract
Abstract Face reenactment technology is widely applied in various applications. However, the reconstruction effects of existing methods are often not quite realistic enough. Thus, this paper proposes a progressive face reenactment method. First, to make full use of the key information, we propose adaptive convolution and instance normalization to encode the key information into all learnable parameters in the network, including the weights of the convolution kernels and the means and variances in the normalization layer. Second, we present continuous transitive facial expression generation according to all the weights of the network generated by the key points, resulting in the continuous change of the image generated by the network. Third, in contrast to classical convolution, we apply the combination of depth‐ and point‐wise convolutions, which can greatly reduce the number of weights and improve the efficiency of training. Finally, we extend the proposed face reenactment method to the face editing application. Comprehensive experiments demonstrate the effectiveness of the proposed method, which can generate a clearer and more realistic face from any person and is more generic and applicable than other methods.
Shibiao Xu, Miao Hua, Jiguang Zhang, Zhaohui Zhang 0002, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds3
2024 AG-SDM: Aquascape generation based on stable diffusion model with low-rank adaptation
abstract
Abstract As an amalgamation of landscape design and ichthyology, aquascape endeavors to create visually captivating aquatic environments imbued with artistic allure. Traditional methodologies in aquascape, governed by rigid principles such as composition and color coordination, may inadvertently curtail the aesthetic potential of the landscapes. In this paper, we propose Aquascape Generation based on Stable Diffusion Models (AG‐SDM), prioritizing aesthetic principles and color coordination to offer guiding principles for real artists in Aquascape creation. We meticulously curated and annotated three aquascape datasets with varying aspect ratios to accommodate diverse landscape design requirements regarding dimensions and proportions. Leveraging the Fréchet Inception Distance (FID) metric, we trained AGFID for quality assessment. Extensive experiments validate that our AG‐SDM excels in generating hyper‐realistic underwater landscape images, closely resembling real flora, and achieves state‐of‐the‐art performance in aquascape image generation.
Muyang Zhang, Yuewei Xian, Wei Li 0237, Jiaming Gu, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds7
2024 SkinFormer: Learning Statistical Texture Representation With Transformer for Skin Lesion Segmentation
abstract
Accurate skin lesion segmentation from dermoscopic images is of great importance for skin cancer diagnosis. However, automatic segmentation of melanoma remains a challenging task because it is difficult to incorporate useful texture representations into the learning process. Texture representations are not only related to the local structural information learned by CNN, but also include the global statistical texture information of the input image. In this paper, we propose a transFormer network (SkinFormer) that efficiently extracts and fuses statistical texture representation for Skin lesion segmentation. Specifically, to quantify the statistical texture of input features, a Kurtosis-guided Statistical Counting Operator is designed. We propose Statistical Texture Fusion Transformer and Statistical Texture Enhance Transformer with the help of Kurtosis-guided Statistical Counting Operator by utilizing the transformer's global attention mechanism. The former fuses structural texture information and statistical texture information, and the latter enhances the statistical texture of multi-scale features. Extensive experiments on three publicly available skin lesion datasets validate that our SkinFormer outperforms other SOAT methods, and our method achieves 93.2% Dice score on ISIC 2018. It can be easy to extend SkinFormer to segment 3D images in the future.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE J. Biomed. Health Informatics3
2023 FeaCo: Reaching Robust Feature-Level Consensus in Noisy Pose Conditions
abstract
Collaborative perception offers a promising solution to overcome challenges such as occlusion and long-range data processing. However, limited sensor accuracy leads to noisy poses that misalign observations among vehicles. To address this problem, we propose the FeaCo, which achieves robust Feature-level Consensus among collaborating agents in noisy pose conditions without additional training. We design an efficient Pose-error Rectification Module (PRM) to align derived feature maps from different vehicles, reducing the adverse effect of noisy pose and bandwidth requirements. We also provide an effective multi-scale Cross-level Attention Module (CAM) to enhance information aggregation and interaction between various scales. Our FeaCo outperforms all other localization rectification methods, as validated on both the collaborative perception simulation dataset OPV2V and real-world dataset V2V4Real, reducing heading error and enhancing localization accuracy across various error levels. Our code is available at: https://github.com/jmgu0212/FeaCo.git.
Jiaming Gu, Muyang Zhang, Weiliang Meng, Shibiao Xu, Jiguang Zhang, Xiaopeng Zhang 0001
ACM Multimedia6
2023 Toward Accurate and Efficient Road Extraction by Leveraging the Characteristics of Road Shapes
abstract
Automatically extracting roads from very high resolution (VHR) remote sensing images is of great importance in a wide range of remote sensing applications. However, complex shapes of roads (i.e., long, geometrically deformed, and thin) always affected the extraction accuracy, which is one of the challenges of road extraction. Based on the insight into road shape characteristics, we propose a novel road shape aware network (RSANet) to achieve efficient and accurate road extraction. First, we introduce the Efficient Strip Transformer Module (ESTM) to efficiently capture the global context to model the long-distance dependence required by the long roads. Second, we design a Geometric Deformation Estimation Module (GDEM) to adaptively extract the context from the shape deformation caused by shooting roads from different perspectives. Third, we provide a simple but effective Road Edge Focal Loss (REF loss) to make the network focus on optimizing the pixels around the road to alleviate the unbalanced distribution of foreground and background pixels caused by the roads being too thin. Finally, we conduct extensive evaluations on public datasets to verify the effectiveness of RSANet and each of the proposed components. Experiments validate that our RSANet outperforms state-of-the-art methods for road extraction in remote sensing images.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Jiguang Zhang, Xiaopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 RSSFormer: Foreground Saliency Enhancement for Remote Sensing Land-Cover Segmentation
abstract
High spatial resolution (HSR) remote sensing images contain complex foreground-background relationships, which makes the remote sensing land cover segmentation a special semantic segmentation task. The main challenges come from the large-scale variation, complex background samples and imbalanced foreground-background distribution. These issues make recent context modeling methods sub-optimal due to the lack of foreground saliency modeling. To handle these problems, we propose a Remote Sensing Segmentation framework (RSSFormer), including Adaptive TransFormer Fusion Module, Detail-aware Attention Layer and Foreground Saliency Guided Loss. Specifically, from the perspective of relation-based foreground saliency modeling, our Adaptive Transformer Fusion Module can adaptively suppress background noise and enhance object saliency when fusing multi-scale features. Then our Detail-aware Attention Layer extracts the detail and foreground-related information via the interplay of spatial attention and channel attention, which further enhances the foreground saliency. From the perspective of optimization-based foreground saliency modeling, our Foreground Saliency Guided Loss can guide the network to focus on hard samples with low foreground saliency responses to achieve balanced optimization. Experimental results on LoveDA datasets, Vaihingen datasets, Potsdam datasets and iSAID datasets validate that our method outperforms existing general semantic segmentation methods and remote sensing segmentation methods, and achieves a good compromise between computational overhead and accuracy. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/RSSFormer-TIP2023.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Image Process.3
2022 Correction to: A hybrid convolutional architecture for accurate image manipulation localization at the pixel-level
Jiguang Zhang, Shibiao Xu
Multim. Tools Appl.2
2021 A hybrid convolutional architecture for accurate image manipulation localization at the pixel-level
Jiguang Zhang, Shibiao Xu
Multim. Tools Appl.2
2018 Accurate blind deblurring using salientpatch-based prior for large-size images
Chengcheng Ma, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Runping Xi, G. Hemanth Kumar, Xiaopeng Zhang 0001
Multim. Tools Appl.2
2018 Real-time pedestrian detection via hierarchical convolutional feature
Dongming Yang, Jiguang Zhang, Shibiao Xu, Shuiying Ge, G. Hemanth Kumar, Xiaopeng Zhang 0001
Multim. Tools Appl.2