EDBT 2026 Demo / reviewers in the wild / expert
Jing Zhang 0038
dblp:05/3499-38
· DBLP profile ↗
23ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0002-6998-0268ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Knowledge-Guided Prompt Learning for Tropical Cyclone Intensity EstimationabstractTropical cyclones (TCs) are one of the most destructive climatic phenomena, making accurate estimation of their intensity crucial for assessing disaster risks. However, due to the complex and indistinguishable structure of their satellite images, existing deep learning methods struggle to differentiate between images of varying intensity, which leads to unsatisfactory accuracy in intensity estimation. In this article, we propose a novel knowledge-guided prompt learning (KGPL) method, KGPL, for TC intensity estimation. KGPL utilizes a pretrained VLM to encode both historical intensity and current satellite image for estimating TC intensity. To minimize domain disparities and facilitate the transfer of a general large model to the TC domain, we introduce a cross-modal interactive prompt learning strategy. Specifically, we embed shared prompts in the text and vision encoders, aiming to learn domain knowledge and promote collaboration and interaction between these two modalities. Furthermore, we design a subregion contrastive learning strategy, which sets constraints on the intensity differences of convective activity in different subregions and guides the model to focus on learning strong convective areas such as the eye and eyewall. Extensive experiments show that our KGPL achieves a significant 41.1% reduction in root mean square error (RMSE) with only one-ninth of the trainable parameters compared to the state-of-the-art method, which validates the effectiveness of our method for TC intensity estimation. The code and data are available athttps://github.com/LiYue-TC/KGPL. Wenhui Li 0001, Yue Li 0042, Dan Song 0006, Jing Zhang 0038, Zhiqiang Wei 0002, Anan Liu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Multimodal High-Order Relationship Inference Network for Fashion Compatibility Modeling in Internet of Multimedia ThingsabstractRecent progress in artificial intelligence (AI) have broadened various intelligent application scenarios on the Internet of Multimedia Things (IoMT). Due to the urgent demands for intelligence in the online fashion industry, using AI techniques to explore user’s clothing collocations from fashion data generated by the IoMT system is a challenging task. Under the background, fashion compatibility modeling (FCM), which aims to estimate the matching degree of a given outfit, has attracted great attention in the multimedia analysis field. However, most of the studies often fail to fully leverage multimodal content or ignore the sparse associations between fashion items. In this article, we propose a novel multimodal high-order relationship inference network (MHRIN) for FCM task. In MHRIN, we focus on enriching multimodal representations of fashion items by means of incorporating the category correlations and injecting high-order item–item connectivity. Concretely, considering that fashion collocations depend on the semantic relevance patterns between categories, we design a category correlations learning module to adaptively learn category representations. On this basis, multiple modality representations are aggregated by a hierarchical multimodal fusion module to generate visual-semantic embeddings. To address the item-item matching interactions issue, we further refine the final representations by a high-order message propagation module to absorb rich connection information. Experiments on the publicly available data set demonstrate the superiority of our MHRIN over state-of-the-art methods. Peiguang Jing, Jing Zhang 0038, Yun Li 0006, Yuting Su 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Research on type-aware fashion compatibility prediction based on a hybrid attention mechanism
Yun Li 0006, GuoXiang Li, Jing Zhang 0038, Peiguang Jing |
Multim. Tools Appl. | 3 |
| 2024 | Context-aware focal alignment network for micro-video multi-label classification
Weiheng Yao, Peiguang Jing, Jing Zhang 0038, Kim Fung Tsang, Shuqiang Wang |
Pattern Anal. Appl. | 4 |
| 2024 | Multi-Task Spatial-Temporal Transformer for Multi-Variable Meteorological ForecastingabstractThis study delves into multi-variable meteorological spatial-temporal prediction, focusing on the simultaneous forecasting of key meteorological parameters such as temperature, wind speed, and atmospheric pressure. The core challenge of this task lies in identifying commonalities across different variables while capturing their unique features and the interactions among them. To address this, we propose a novel multi-task learning framework tailored for multi-variable meteorological forecasting. Our framework integrates a convolutional variable-specific visual representation module and a variable-interactive spatial-temporal inference module. The former extracts distinct variable information independently for each variable, while the latter employs a tri-level attention mechanism across space, time, and variables to uncover both commonalities and interactions among the variables. An adaptive multi-loss optimization strategy and a local information aggregation module are introduced to balance task optimization complexities and enhance representation stability. Comprehensive experiments across various meteorological prediction tasks confirm the effectiveness of our methods, showcasing superior performance over existing approaches. Tianbao Li 0001, Anan Liu, Dan Song 0006, Wenhui Li 0001, Jing Zhang 0038, Zhiqiang Wei 0002, Yuting Su 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | MobileSky: Real-Time Sky Replacement for Mobile ARabstractWe present MobileSky, the first automatic method for real-time high-quality sky replacement for mobile AR applications. The primary challenge of this task is how to extract sky regions in camera feed both quickly and accurately. While the problem of sky replacement is not new, previous methods mainly concern extraction quality rather than efficiency, limiting their application to our task. We aim to provide higher quality, both spatially and temporally consistent sky mask maps for all camera frames in real time. To this end, we develop a novel framework that combines a new deep semantic network called FSNet with novel post-processing refinement steps. By leveraging IMU data, we also propose new sky-aware constraints such as temporal consistency, position consistency, and color consistency to help refine the weakly classified part of the segmentation output. Experiments show that our method achieves an average of around 30 FPS on off-the-shelf smartphones and outperforms the state-of-the-art sky replacement methods in terms of execution speed and quality. In the meantime, our mask maps appear to be visually more stable across frames. Our fast sky replacement method enables several applications, such as AR advertising, art making, generating fantasy celestial objects, visually learning about weather phenomena, and advanced video-based visual effects. To facilitate future research, we also create a new video dataset containing annotated sky regions with IMU data. Xinjie Wang 0003, Qingxuan Lv, Jing Zhang 0038, Zhiqiang Wei 0002, Junyu Dong, Hongbo Fu 0001, Zhipeng Zhu, Xiaogang Jin 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | FusionDeformer: text-guided mesh deformation using diffusion models
Hao Xu 0049, Xiangjun Tang, Jing Zhang 0038, Zhebin Zhang, Chen Li 0062, Xiaogang Jin 0001 |
Vis. Comput. | 4 |
| 2024 | Publisher Correction: FusionDeformer: text-guided mesh deformation using diffusion models
Hao Xu 0049, Xiangjun Tang, Jing Zhang 0038, Zhebin Zhang, Chen Li 0062, Xiaogang Jin 0001 |
Vis. Comput. | 4 |
| 2023 | LPFF: A Portrait Dataset for Face Generators Across Large PosesabstractExisting face generators exhibit exceptional performance on faces in small to medium poses (with respect to frontal faces) but struggle to produce realistic results for large poses. The distorted rendering results on large poses in 3D-aware generators further show that the generated 3D face shapes are far from the distribution of 3D faces in reality. We find that the above issues are caused by the training dataset’s pose imbalance. To this end, we present LPFF, a large-pose Flickr face dataset comprised of 19,590 high-quality real large-pose portrait images. We utilize our dataset to train a 2D face generator that can process large-pose face images, as well as a 3D-aware generator that can generate realistic human face geometry. To better validate our pose-conditional 3D-aware generators, we develop a new FID measure to evaluate the 3D-level performance. Through this novel FID measure and other experiments, we show that LPFF can help 2D face generators extend their latent space and better manipulate the large-pose data, and help 3D-aware face generators achieve better view consistency and more realistic 3D reconstruction results. Jing Zhang 0038, Hongbo Fu 0001, Xiaogang Jin 0001 |
ICCV | 2 |
| 2022 | Tripartite Graph Regularized Latent Low-Rank Representation for Fashion Compatibility PredictionabstractIn recent years, an increasing online shopping demand has greatly promoted the innovation and development of the fashion industry. Visual fashion analysis has become a prospective research topic in computer vision and multimedia fields. Among these studies, fashion compatibility analysis is required in many real applications, such as fashion recommendation, matching, and retrieval. However, learning fashion compatibility is nontrivial, not only due to the uncertain and sparse dependencies among fashion items but also the latent and mutual associations among multiple factors such as color, texture, style, and functionality. To better predict fashion compatibility, in this paper, we proposed a tripartite graph regularized latent low-rank representation method, named TGRLLR, for fashion compatibility prediction. In TGRLLR, to learn more low-dimensional and effective representations, we considered the latent low-rank representation by decomposing the original feature matrix in both the column and row directions to tackle the problem of insufficient observations. On this basis, we simultaneously exploited different regularization strategies to encode the structured correlations among features, the high-order relationships among items, and the geometrical structures of outfits for more informative representations. Extensive experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed method compared with state-of-the-art methods. Peiguang Jing, Jing Zhang 0038, Liqiang Nie, Jing Liu 0002, Yuting Su 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Deep low-rank matrix factorization with latent correlation estimation for micro-video multi-label classification
Yuting Su 0001, Junyu Xu, Daozheng Hong, Fugui Fan, Jing Zhang 0038, Peiguang Jing |
Inf. Sci. | 5 |
| 2021 | MV-LFN: Multi-view based local information fusion network for 3D shape recognitionabstract3D shape recognition is a challenging task due to the difficulty of representing the complex structure of 3D shapes. Recently, the view-based approaches that utilize the multiple views rendered from the shape for visual information extraction and feature aggregation to generate a global shape descriptor , achieved promising performance. However, the view-based approaches commonly ignore the exploration and utilization of local information in the multiple views, which influences the effectiveness of generated features. In this paper, we design a novel Multi-view based Local Information Fusion Network (MV-LFN) for the 3D shape recognition task. The local correlation attention mechanism (LCAM) is introduced to exploit the local correlations in the feature maps for generating a more effective view descriptor. Then, we hierarchically aggregate the multi-view feature maps to generate a shape super matrix (SSM). The local information is effectively extracted and maintained during the multi-view aggregation process, and the discrimination of shape descriptors is significantly improved. We conduct comparative experiments on the ModelNet and ShapeNetCore55 databases. The experimental performances effectively validate the superiority of MV-LFN. Jing Zhang 0038, Dangdang Zhou, Yue Zhao 0042, Weizhi Nie, Yuting Su 0001 |
Vis. Informatics | 1 |
| 2020 | ABSNet: Aesthetics-Based Saliency Network Using Multi-Task Convolutional NetworkabstractAs a smart visual attention mechanism to analyze visual scenes, visual saliency has been shown to closely correlate with semantic information such as faces. Although many semantic-information-guided saliency models have been proposed, to the best of our knowledge, no semantic information in affective domain has been employed for saliency detection. Aesthetic, the affective perceptual quality that integrates factors like scene composition and contrast, can certainly benefit visual attention that highly depends on these visual factors. In this letter, we propose an end-to-end multi-task framework called aesthetics-based saliency network (ABSNet). We use three commonly-used shared backbones and design two distinct branches for each task. Mean square error (MSE) loss and Earth Mover's Distance (EMD) loss are jointly adopted to alternately train the shared network and individual branch for different tasks, facilitating the proposed model to extract more effective features for visual perception. Moreover, our model is resolution-friendly to predict saliency for images of arbitrary size. It has been shown that the proposed multi-task method is superior over single-task version and outperforms state-of-the-art saliency methods. Jing Liu 0002, Jincheng Lv, Jing Zhang 0038, Yuting Su 0001 |
IEEE Signal Process. Lett. | 4 |
| 2019 | Tensor-driven low-rank discriminant analysis for image set classification
Jing Zhang 0038, Zhengnan Li, Peiguang Jing, Ye Liu 0002, Yuting Su 0001 |
Multim. Tools Appl. | 1 |
| 2019 | Visual attribute detction for pedestrian detection
Jing Zhang 0038, Fuwu Li, Weizhi Nie, Wenhui Li 0001, Yuting Su 0001 |
Multim. Tools Appl. | 1 |
| 2019 | A structure-transfer-driven temporal subspace clustering for video summarization
Jing Zhang 0038, Peiguang Jing, Jing Liu 0002, Yuting Su 0001 |
Multim. Tools Appl. | 1 |
| 2018 | Graph regularized low-rank tensor representation for feature selection
Yuting Su 0001, Peiguang Jing, Jing Zhang 0038, Jing Liu 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Low-Rank Regularized Heterogeneous Tensor Decomposition for Subspace ClusteringabstractThis letter proposes a low-rank regularized heterogeneous tensor decomposition (LRRHTD) algorithm for subspace clustering, in which various constrains in different modes are incorporated to enhance the robustness of the proposed model. Specifically, due to the presence of noise and redundancy in the original tensor, LRRHTD seeks a set of orthogonal factor matrices for all but the last mode to map the high-dimensional tensor into a low-dimensional latent subspace. Furthermore, by imposing a low-rank constraint on the last mode, which is relaxed by using a nuclear norm, the lowest rank representation that reveals the global structure of samples is obtained for the purpose of clustering. We develop an effective algorithm based on the augmented Lagrange multiplier to optimize our model. Experiments on two public datasets demonstrate that our method reaches convergence within a small number of iterations and achieves promising results in comparison with the state of the arts. Jing Zhang 0038, Peiguang Jing, Jing Liu 0002, Yuting Su 0001 |
IEEE Signal Process. Lett. | 1 |
| 2017 | Automatic report generation based on multi-modal information
Jing Zhang 0038, Weizhi Nie, Yuting Su 0001 |
Multim. Tools Appl. | 1 |
| 2016 | A Tensor-Driven Temporal Correlation Model for Video Sequence ClassificationabstractThe task of video sequence classification plays a critical role in the development of computer vision. Considering this fact, this letter proposes a novel tensor decomposition method called tensor-driven temporal correlation in which general tensors are used as input for video sequence classification. Because distortion and redundancy may exist in the tensor representations of video sequences, we project the original tensor into subspaces spanned by spatial basis matrices in the proposed formulation. Moreover, to better preserve the temporal smoothness between consecutive slices of the tensor, the basis matrices are jointly learned by introducing an autoregressive model. An experiment on the commonly used Cambridge hand-gesture database demonstrates that our proposed method reaches convergence within a small number of iterations during the training stage and achieves promising results compared with state-of-the-art methods. Jing Zhang 0038, Chuan-Zhong Xu, Peiguang Jing, Chengqian Zhang, Yuting Su 0001 |
IEEE Signal Process. Lett. | 1 |
| 2010 | A Novel Source MPEG-2 Video Identification AlgorithmabstractWith the availability of powerful multimedia editing software, all types of personalized image and video resources are available in networks. Multimedia forensics technology has become a new topic in the field of information security. In this paper, a new source video system identification algorithm is proposed based on the features in the video stream; it takes full advantage of the different characteristics in the rate control module and the motion prediction module, which are two open parts in the MPEG-2 video compression standard, and combines a support vector machine classifier to build an intelligent computing system for video source identification. The experiments show this proposed algorithm can effectively identify video streams that come from a number of video coding systems. Yuting Su 0001, Junyu Xu, Jing Zhang 0038, Qingzhong Liu |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2010 | Exposing Digital Video Logo-Removal Forgery by Inconsistency of BlurabstractA novel approach for detecting video logo-removal forgery is proposed by measuring inconsistency of blur. Our approach is based on the assumption that if a digital video undergoes logo-removal forgery; the blurriness of the forged region is expected to be different as compared to the nontampered parts of the video. Blurriness is first estimated by analyzing the spatial and temporal statistical property of logo areas, and suspicious areas are roughly located; then features are extracted and a fine classification is implemented by applying support vector machine (SVM) to extract features. If the suspicious areas and the reference areas are classified into different classes, the video is judged as a forged video. Experimental results show that our method is robust to video lossy compression for logo-removal forgery detection with the advantages of high classification accuracy and low computation cost. Yuting Su 0001, Jing Zhang 0038, Qingzhong Liu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2007 | Watermarking for Authentication of LZ-77 Compressed Documents
Yanfang Du, Jing Zhang 0038, Yuting Su 0001 |
IWDW | 2 |