VLDB 2026 Research / reviewers in the wild / expert
Jingran Zhang
dblp:247/9299
· DBLP profile ↗
14ranked-venue papers
3as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Design of Scalable Hybrid Boost Converter With Optimal Number of Stages for High Conversion Ratio and High Power EfficiencyabstractThis paper presents a scalable hybrid boost converter (SHBC) with adjusted number of stages (N) to acquire high conversion ratio (CR) and high power efficiency. Firstly, unlike most converters whereCRis related only to duty cycle (D), theCRof SHBC is related to bothDand N. This allows the converter to provide a highCRwhile keepingDat a reasonable value for easy loop control. Secondly, N flying capacitors in series with the inductor can divide output voltage ($V_{\mathrm {OUT}}$) to use low voltage devices with small figure of metric (FOM) for high power efficiency. Moreover, N flying capacitors in parallel with the inductor can shunt the inductor current to reduce the inductor DCR loss, which also helps improve power efficiency. Thirdly, when many N values meet the requirements of a specific application scenario, an optimal number of stages (NOPT) is proposed to minimize the total loss and optimize the power efficiency further. The prototype design of SHBC with$\mathrm{N}_{\mathrm {OPT}} =2$has been discussed and fabricated by$0.18~\mu $m BCD process. The measurement results demonstrate its operation is normal in theCRrange from 4 to 8. Besides, a 93.34% peak efficiency is also achieved with$I_{\mathrm {OUT}} =0.125$A andCR= 4. Kai Yu 0008, Yuhong Deng, Sizhen Li, Jingran Zhang, Mo Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Coreset Learning-Based Sparse Black-Box Adversarial Attack for Video RecognitionabstractIn recent years, researchers have explored the use of sparse black-box video adversarial attacks, which involve selecting keyframes to reduce computational complexity and improve efficiency in generating perturbations. However, the current sparse strategy is not optimized for attack and detection steps, resulting in inaccurate frame selection. Some researchers have used reinforcement learning to train an agent to select keyframes, but this method requires additional training. To address these challenges, we propose a plug-and-play black-box sparse attack algorithm called CLVA based on the coreset concept of active learning. Our algorithm treats a video as a mini-dataset and employs the K-Center-Greedy algorithm to compute the distances between frames. We then select the frame that meets the distance condition as the key frame. We conducted extensive experiments using two attack algorithms on five mainstream recognition models and three video recognition datasets. Our results demonstrate that CLVA significantly accelerates the black-box video attack algorithm while achieving state-of-the-art performance in sparsity, time, and success rate compared to recent sparse attack algorithms. The implementation code of our CLVA method is available athttps://github.com/machineNo6/CLVA. Jiefu Chen, Xing Xu 0001, Jingran Zhang, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | SDN: Semantic Decoupling Network for Temporal Language GroundingabstractTemporal language grounding (TLG) is one of the most challenging cross-modal video understanding tasks, which aims at retrieving the most relevant video segment from an untrimmed video according to a natural language sentence. The existing methods can be separated into two dominant types: 1) proposal-based and 2) proposal-free methods, where the former conduct contextual interactions and the latter localizes timestamps flexibly. However, the constant-scale candidates in proposal-based methods limit the localization precision and bring extra computational costs. In contrast, the proposal-free methods perform well on high-precision metrics-based on the fine-grained features but suffer from a lack of coarse-grained interactions, which cause degeneration when the video becomes complex. In this article, we propose a novel framework termed semantic decoupling network (SDN) that combines the advantages of proposal-based and proposal-free methods and overcomes their defects. It contains three key components: 1) semantic decoupling module (SDM); 2) context modeling block (CMB); and 3) semantic cross-level aggregation module (SCAM). By capturing the video-text contexts in multilevel semantics, the SDM and CMB effectively utilize the benefits of proposal-based methods. Meanwhile, the SCAM maintains the merit of proposal-free methods in that it localizes timestamps precisely. The experiments on three challenge datasets, i.e., Charades-STA, TACoS, and ActivityNet-Caption, show that our proposed SDN method significantly outperforms recent state-of-the-art methods, especially the proposal-free methods. Extensive analyses, as well as the implementation code of the proposed SDN method, are provided at https://github.com/CFM-MSG/Code_SDN. Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Fumin Shen, Zuo Cao, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Remote sensing imaging analysis and ubiquitous cloud-based mobile edge computing based intelligent forecast of forest tourism demand
Jingran Zhang, Wukui Wang |
Distributed Parallel Databases | 2 |
| 2022 | Semi-supervised Video Paragraph Grounding with Contrastive EncoderabstractVideo events grounding aims at retrieving the most relevant moments from an untrimmed video in terms of a given natural language query. Most previous works focus on Video Sentence Grounding (VSG), which localizes the moment with a sentence query. Recently, researchers extended this task to Video Paragraph Grounding (VPG) by retrieving multiple events with a paragraph. However, we find the existing VPG methods may not perform well on context modeling and highly rely on video-paragraph annotations. To tackle this problem, we propose a novel VPG method termed Semi-supervised Video-Paragraph TRansformer (SVPTR), which can more effectively exploit contextual information in paragraphs and significantly reduce the dependency on annotated data. Our SVPTR method consists of two key components: (1) a base model VPTR that learns the video-paragraph alignment with contrastive encoders and tackles the lack of sentence-level contextual interactions and (2) a semi-supervised learning framework with multimodal feature perturbations that reduces the requirements of annotated training data. We evaluate our model on three widely-used video grounding datasets, i.e., ActivityNet-Caption, Charades-CD-OOD, and TACoS. The experimental results show that our SVPTR method establishes the new state-of-the-art performance on all datasets. Even under the conditions of fewer annotations, it can also achieve competitive results compared with recent VPG methods. Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Fumin Shen, Zuo Cao, Heng Tao Shen |
CVPR | 3 |
| 2022 | SA-FGDEM: A Self-adaptive E-Learning Performance Prediction ModelabstractWith the rapid development of computer computing power and the severe challenges brought by the COVID-19, e-learning, as the optimal solution for most students and other learner groups, plays an extremely important role in maintaining the normal operation of educational institutions. As the user community continues to expand, it has become increasingly important to guarantee the quality of teaching and learning. One way to ensure the quality of online education is to construct e-learning behavior data to build learning performance predictors. Still, most studies have ignored the intrinsic correlation between e-learning behaviors. Therefore, this study proposes an adaptive feature fusion-based e-learning performance prediction model (SA-FGDEM) relying on the theoretical model of learning behav-ior classification. The experimental results show that the feature space mined by fine-grained differential evolution algorithm and the adaptive feature fusion combined with differential evolution algorithm can support e-learning performance prediction more effectively and is better than the benchmark method. Mingtao Ye, Liping Wang 0016, Jingran Zhang, Guodao Zhang |
DSAA | 3 |
| 2022 | GTLR: Graph-Based Transformer with Language Reconstruction for Video Paragraph GroundingabstractVideo Paragraph Grounding aims at retrieving multiple relevant moments from an untrimmed video with a given natural language paragraph query. However, the complex paragraph query brings more challenges to the multimodal fusion and context modeling, which limited the performance of existing VPG methods. To this end, we propose a novel framework for VPG in this paper, termed Graph-based Transformer with Language Reconstruction (GTLR). It consists of three components: (1) Multimodal Graph Encoder conducting the graph reasoning for video-text fusion. (2) Event-wise Decoder predicting the timestamps based on multiple sentence-level features. (3) Language Reconstructor rebuilding the paragraph queries and making our model explainable. We adopt two benchmarks, i.e., ActivityNet-Caption and Charades-STA, to evaluate our model and conduct comprehensive experiments to analyze the effectiveness of each component. The experimental results show that our GTLR method outperforms recent state-of-the-art methods. Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Fumin Shen, Zuo Cao |
ICME | 3 |
| 2022 | DHHN: Dual Hierarchical Hybrid Network for Weakly-Supervised Audio-Visual Video ParsingabstractThe Weakly-Supervised Audio-Visual Video Parsing (AVVP) task aims to parse a video into temporal segments and predict their event categories in terms of modalities, labeling them as either audible, visible, or both. Since the temporal boundaries and modalities annotations are not provided, only video-level event labels are available, this task is more challenging than conventional video understanding tasks.Most previous works attempt to analyze videos by jointly modeling the audio and video data and then learning information from the segment-level features with fixed lengths. However, such a design exist two defects: 1) The various semantic information hidden in temporal lengths is neglected, which may lead the models to learn incorrect information; 2) Due to the joint context modeling, the unique features of different modalities are not fully explored. In this paper, we propose a novel AVVP framework termedDual Hierarchical Hybrid Network (DHHN) to tackle the above two problems. Our DHHN method consists of three components: 1) A hierarchical context modeling network for extracting different semantics in multiple temporal lengths; 2) A modality-wise guiding network for learning unique information from different modalities; 3) A dual-stream framework generating audio and visual predictions separately. It maintains the best adaptions on different modalities, further boosting the video parsing performance. Extensive quantitative and qualitative experiments demonstrate that our proposed method establishes the new state-of-the-art performance on the AVVP task. Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Jingkuan Song, Fumin Shen, Huimin Lu 0001, Heng Tao Shen |
ACM Multimedia | 4 |
| 2022 | MAVT-FG: Multimodal Audio-Visual Transformer for Weakly-supervised Fine-Grained RecognitionabstractWeakly-supervised fine-grained recognition aims to detect potential differences between subcategories at a more detailed scale without using any manual annotations. While most recent works focus on classical image-based fine-grained recognition that recognizes subcategories at image-level, video-based fine-grained recognition is much more challenging and specifically needed. In this paper, we propose a Multimodal Audio-Visual Transformer for Weakly-supervised Fine-Grained Recognition (MAVT-FG) model which incorporates audio-visual modalities. Specifically, MAVT-FG consists of Audio-Visual Dual-Encoder for feature extraction, Cross-Decoder for Audio-Visual Fusion (DAVF) to exploit inherent cues and correspondences between two modalities, and Search-and-Select Fine-grained Branch (SSFG) to capture the most discriminative regions. Furthermore, we construct a new benchmark: Fine-grained Birds of Audio-Visual (FGB-AV) for audio-visual weakly-supervised fine-grained recognition at video-level. Experimental results show that our method achieves superior performance and outperforms other state-of-the-art methods. Xiaotong Song, Jingran Zhang |
ACM Multimedia | 4 |
| 2022 | Modeling Two-Stream Correspondence for Visual Sound SeparationabstractVisual sound separation (VSS) aims to obtain each sound component from the mixed audio signals with the guidance of visual information. Existing works mainly capture the global-level audio-visual correspondence and exploit various visual features to enhance the appearance and motion features of visual modality. However, they commonly neglect the intrinsic properties of the audio modality, resulting in less effective audio feature extraction and unbalanced audio-visual correspondence. To tackle this problem, we propose a novel end-to-end framework termed Modeling Two-Stream Correspondence (MTSC) for VSS by explicitly extracting the timbre and content features in audio modality. The proposed MTSC method employs a two-stream architecture to enhance audio-visual correspondence for both the appearance-timbre and motion-content features. Moreover, with the advanced two-stream pipeline, more lightweight appearance and motion features for visual modality are exploited. Extensive experiments conducted on two benchmark musical instrument datasets demonstrate that with the above properties, our MTSC method remarkably outperforms seven state-of-the-art VSS approaches. The implementation code and extensive experimental results of the proposed MTSC method are provided athttps://github.com/CFM-MSG/MTSC-VSS. Xing Xu 0001, Jingran Zhang, Fumin Shen, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Enhancing Audio-Visual Association with Self-Supervised Curriculum LearningabstractThe recent success of audio-visual representations learning can be largely attributed to their pervasive concurrency property, which can be used as a self-supervision signal and extract correlation information. While most recent works focus on capturing the shared associations between the audio and visual modalities, they rarely consider multiple audio and video pairs at once and pay little attention to exploiting the valuable information within each modality. To tackle this problem, we propose a novel audio-visual representation learning method dubbed self-supervised curriculum learning (SSCL) under the teacher-student learning manner. Specifically, taking advantage of contrastive learning, a two-stage scheme is exploited, which transfers the cross-modal information between teacher and student model as a phased process. The proposed SSCL approach regards the pervasive property of audiovisual concurrency as latent supervision and mutually distills the structure knowledge of visual to audio data. Notably, the SSCL method can learn discriminative audio and visual representations for various downstream applications. Extensive experiments conducted on both action video recognition and audio sound recognition tasks show the remarkably improved performance of the SSCL method compared with the state-of-the-art self-supervised audio-visual representation learning methods. Jingran Zhang, Xing Xu 0001, Fumin Shen, Huimin Lu 0001, Xin Liu 0011, Heng Tao Shen |
AAAI | 1 |
| 2021 | Video Representation Learning with Graph Contrastive AugmentationabstractContrastive-based self-supervised learning for image representations has significantly closed the gap with supervised learning. A natural extension of image-based contrastive learning methods to the video domain is to fully exploit the temporal structure presented in videos. We propose a novel contrastive self-supervised video representation learning framework, termed Graph Contrastive Augmentation (GCA), by constructing a video temporal graph and devising a graph augmentation that is designed to enhance the correlation across frames of videos and developing a new view for exploring temporal structure in videos. Specifically, we construct the temporal graph in the video by leveraging the relational knowledge behind the correlated sequence video features. Afterwards, we apply the proposed graph augmentation to generate another graph view by cooperating random corruption of the original graph to enhance the diversity of the intrinsic structure of the temporal graph. To this end, we provide two different kinds of contrastive learning methods to train our framework using temporal relationships concealed in videos as self-supervised signals. We perform empirical experiments on downstream tasks, action recognition and video retrieval, using the learned video representation, and the results demonstrate that with the graph view of temporal structure, our proposed GCA remarkably improves performance against or on par with the recent methods. Jingran Zhang, Xing Xu 0001, Fumin Shen, Yazhou Yao, Jie Shao 0001, Xiaofeng Zhu 0001 |
ACM Multimedia | 1 |
| 2021 | Adversarial Attack Against Urban Scene Segmentation for Autonomous VehiclesabstractUnderstanding the surrounding environment is crucial for autonomous vehicles to make correct driving decisions. In particular, urban scene segmentation is a significant integral module commonly equipped in the perception system of autonomous vehicles to understand the real scene like a human. Any missegmentation of the driving scenario can potentially result in uncontrollable consequences such as serious accidents or the exception of the perception system. In this article, we investigate the vulnerability of the popular scene segmentation models designed with the backbones of deep neural networks (DNNs), which have been shown to be sensitive to adversarial attacks. Specifically, we propose an iterative projected gradient-based attack method that can effectively fool several DNN-based segmentation models with a remarkably higher attacking successful rate, and much smaller adversarial perturbations. Moreover, we also develop an adversarial training algorithm with min-max optimization style to enrich the robustness of the scene segmentation models. Extensive experiments on the Cityscape benchmark dataset consisting of large-scale urban scene images for autonomous vehicles demonstrate the effectiveness of our proposed attack method, as well as the benefit of the adversarial training scheme for the scene segmentation models. Xing Xu 0001, Jingran Zhang, Yujie Li 0001, Yichuan Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Temporal Reasoning Graph for Activity RecognitionabstractDespite great success has been achieved in activity analysis, it still has many challenges. Most existing works in activity recognition pay more attention to designing efficient architecture or video sampling strategy. However, due to the property of fine-grained action and long term structure in video, activity recognition is expected to reason temporal relation between video sequences. In this paper, we propose an efficient temporal reasoning graph (TRG) to simultaneously capture the appearance features and temporal relation between video sequences at multiple time scales. Specifically, we construct learnable temporal relation graphs to explore temporal relation on the multi-scale range. Additionally, to facilitate multi-scale temporal relation extraction, we design a multi-head temporal adjacent matrix to represent multi-kinds of temporal relations. Eventually, a multi-head temporal relation aggregator is proposed to extract the semantic meaning of those features convolving through the graphs. Extensive experiments are performed on widely-used large-scale datasets, such as Something-Something, Charades and Jester, and the results show that our model can achieve stateof- the-art performance. Further analysis shows that temporal relation reasoning with our TRG can extract discriminative features for activity recognition. Jingran Zhang, Fumin Shen, Xing Xu 0001, Heng Tao Shen |
IEEE Trans. Image Process. | 1 |