VLDB 2026 Research / reviewers in the wild / expert
Shaochen Jiang
dblp:271/4174
· DBLP profile ↗
15ranked-venue papers
0as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ESBR: Event-Conditioned Structured Boundary Reasoning for Overlapping Event ExtractionabstractOverlapping event extraction aims to identify triggers, arguments, and semantic roles when multiple events share textual components within the same sentence. In agricultural text, this task is particularly challenging because shared arguments may play different roles across events, while compound terms and long-span expressions often exhibit ambiguous boundaries. Existing methods still suffer from insufficient event conditioning and weak exploitation of boundary cues, which easily lead to role confusion and noisy span candidates during decoding. To address these issues, we propose ESBR, a structured model that integrates event-conditioned encoding, structured span inference, hierarchical boundary reasoning, and boundary-guided decoding. We further enhance training and inference stability with uncertainty-aware weighting and adaptive thresholding. Experiments on the public benchmark FewFC and our constructed agricultural benchmark FewAgri show that ESBR consistently outperforms competitive baselines, with especially notable gains on argument identification and role classification. Bo Kong 0002, Zihua Song, Shaochen Jiang, Liruizhi Jia, Shengquan Liu |
ICIC | 4 |
| 2026 | STEP: Stable Gradient Projection for Continual LearningabstractContinual learning (CL) aims to enable networks to learn continuously from sequentially arriving task streams while avoiding catastrophic forgetting (CF) of previously learned tasks. In recent years, Orthogonal gradient projection (OGP)-based CL methods have garnered significant attention from the research community due to their remarkable performance. However, existing OGP approaches overlook two critical issues: (1) representation matrices are typically constructed via random sampling, which introduces misclassified and class-imbalanced samples into the projection basis, contaminating important gradient directions and degrading stability; and (2) task-specific output scale variations induce domain drift, resulting in projection bias that weakens orthogonal constraints across tasks. To address these limitations, we propose Stable Gradient Projection for Continual Learning (STEP), a plug-and-play enhancement framework for OGP-based CL that integrates Correctness-aware Balanced Sampling (CBS) to construct purified and class-balanced projection subspaces using only correctly classified samples, and Sigmoid Attention Constraint (SAC) to enforce consistent output scaling via a sigmoid-based gating mechanism, thereby mitigating scale-induced projection bias. Extensive experiments on Split CIFAR-100, CIFAR-100 Superclass, and 5-Datasets demonstrate that STEP consistently improves state-of-the-art OGP methods, achieving up to +1.4% average accuracy (ACC) gains on Split CIFAR-100, improving backward transfer (BWT) from − 0.37 to − 0.09 for GPM and from − 1.06 to − 0.73 for SGP, and attaining 93.28% ACC with positive BWT (0.17) on 5-Datasets. These results validate STEP as a simple yet effective strategy for enhancing stability–plasticity balance in OGP-based CL. Longlong Zhai, Jiao Tian, Yanjun Qin, Shaochen Jiang, Chong Peng 0001, Panpan Zheng |
ICMR | 6 |
| 2026 | LENS-Net: Low-energy spiking neural network for remote sensing saliency
Longlong Zhai, Marcin Pietron, Roberto Corizzo, Zhaoru Guo, Yongke Li, Chong Peng 0001, Shaochen Jiang, Panpan Zheng |
Neurocomputing | 8 |
| 2025 | GCDN: A Novel YOLOv11-Based Approach for Cotton Pest and Disease Detection
Xinkang Li, Shaochen Jiang |
ICIC (2) | 3 |
| 2025 | ME-CWNER: Multi-metadata Embedding Based Chinese Named Entity Recognition for Wheat Diseases and Pests
Shouhao Yu, Shengquan Liu, Shaochen Jiang, Ruizhi Jiali, Bo Kong 0002 |
ICIC (24) | 3 |
| 2025 | Multi-Scale Feature Recognition and Lightweight Cotton Pest Detection Network Based on YOLOv10abstractIn traditional cotton pest detection, manual identification is a slow, labor-intensive process limited by vision, thereby affecting accuracy. This study introduces an enhanced algorithm based on YOLOv10n, namely MCDN-YOLOv10n. Its numerous key improvements are noteworthy: Firstly, we introduce the Global and Local Cross-scale Attention (GGCA) mechanism to optimize and upgrade the C2f structure in the backbone network (C2f-GA), enabling more efficient feature extraction. Secondly, we replace SPPF with FocalNets to ensure precise target localization. Lastly, we adopt the Real-Time Deformable Detection Transformer (RT-DETR) detection head, significantly enhancing the detection capability for small and occluded targets. On the publicly available CottonInsect image dataset, our MCDN model performs exceptionally well, outperforming other mainstream algorithms. Specifically, the mean Average Precision (mAP) has improved by 2.1%, reaching an impressive 95.7%, with an accuracy increase of 3.0%. Simultaneously, we have successfully reduced the number of parameters to 4.2MB, and the inference time has also been significantly improved, dropping sharply from the original value to just 0.23 milliseconds, representing a speed increase of approximately 40% compared to the original model. Xinkang Li, Shaochen Jiang |
IJCNN | 3 |
| 2025 | Dual Fusion with Auxiliary Loss Hashing for Cross-Modal Retrieval
Shenao Shao, Shaochen Jiang, Beibei Gao |
PRCV (5) | 3 |
| 2025 | UniAVLM: Unified Large Audio-Visual Language Models for Comprehensive Video Understanding
Lecheng Yan, Chenyang Lyu, Wenxi Li, Younes Samih, Shaochen Jiang |
PRICAI (5) | 5 |
| 2025 | Mamba: A Contrastive Learning and Data Augmentation-Based Model for Improving Deepfake GeneralizationabstractIn Deepfake face detection, the variety of forgery techniques and limited datasets pose challenges for learning discriminative features for authenticity judgment. Addressing this issue is crucial for improving generalization ability, especially as models are likely to encounter unknown forgery techniques in real-world scenarios. To tackle this, we propose CDA-Net, a novel contrastive data augmentation network that enhances both detection accuracy and generalization. CDA-Net integrates three key modules: Contrast-Driven Feature Aggregation (CDFA), Dual-Perspective Normalization (DPN), and Multi-Scale Mamba (MS-Mamba).CDFA improves feature contrasts through sliding window scanning, enabling detection of subtle forgery boundaries without explicit supervision. DPN uses CrossNorm for data augmentation, avoiding overfitting, while SelfNorm adaptively reduces style shifts between training and testing sets, focusing on authenticity-critical features. MS-Mamba applies a multi-directional scanning mechanism for long-range relationship comparisons, boosting contrastive learning ability.Extensive evaluations across five benchmark datasets demonstrate that CDA-Net achieves superior average performance compared to existing state-of-the-art methods, highlighting its strong generalization ability to detect unknown forgery techniques and adapt to real-world scenarios. Shaochen Jiang, Sijia He, Hongmeng Lu |
SMC | 2 |
| 2025 | WTFN: Wavelet Convolution and Transformer Fusion Network With Spatial-Spectral Features for Hyperspectral Image ClassificationabstractIn hyperspectral image (HSI) classification tasks, effectively extracting the spatial-spectral features of the image is crucial. However, existing convolutional neural network (CNN)-based methods are limited by the fixed convolution kernel size, which means CNN can only extract local spatial features, neglecting the global features of the HSI. On the other hand, Transformer-based methods perform excellently in extracting global spectral features but are weaker at capturing texture and edge features. Therefore, to fully exploit the complementary advantages of CNN and Transformer methods, this paper proposes a dual-branch wavelet convolution and Transformer fusion network (WTFN). The two branches of WTFN are designed to capture local texture features and global contextual relationships, respectively. The attention cross fusion (ACF) module deeply enhances the information interaction of the spatial-spectral features of HSI. Extensive experimental results on three public datasets show that the WTFN method outperforms several state-of-the-art methods in classification performance. Hongmeng Lu, Shaochen Jiang |
SMC | 2 |
| 2025 | On the Taxonomy, Tasks, and Open-Challenges for Multimodal Large Language ModelsabstractIn recent years, the field of Artificial Intelligence has witnessed the emergence of Multimodal Large Language Models (MLLMs) that have significantly advanced the state-of-the-art in understanding and generating content across various data modalities. These models, capable of processing and integrating information from text, images, audio, and video, have opened new avenues for research and applications. Distinguished by their ability to understand and generation information with diverse modalities, such as text, image, audio and many others, MLLMs mark a significant step towards the final aim of Artificial General Intelligence (AGI). This comprehensive survey provides an in-depth examination of MLLMs, highlighting their evolutionary trajectory, current state-of-the-art developments, and prospective future directions. Specifically, we show taxonomy of MLLMs by their modalities to be processed and model architecture for aligning multiple modalities. Besides, we also present discussion regarding the different types of tasks related to MLLMs. The paper further delves into the pressing challenges confronted in this domain, such as data scarcity, computational complexity, ethical dilemmas, and privacy considerations. We analyze these issues in the context of both development and deployment of MLLMs. The survey comprehensively demonstrate and summarise the recent advances of the transformative influence of MLLMs while acknowledging their potential limitations, thereby outlining a prospective roadmap for future research endeavors in this rapidly developing field. Lecheng Yan, Jiahui Geng, Minghao Wu, Zhanyu Wang, Wenxi Li, Tianbo Ji, Shaochen Jiang, Chenyang Lyu |
SMC | 9 |
| 2024 | Real-Time Smoke Detection Network Based on Multi-Scale Feature Recognition and Lightweight Architecture DesignabstractForest fires have significantly impacted global ecosystems and human societies, necessitating the development of efficient and accurate early smoke detection technologies for fires. However, current smoke detection technologies face multiple challenges in real-time applications, including large parameter size, high computational complexity, and low detection accuracy in complex scenes. Therefore, based on YOLOv8, we propose a lightweight, high-precision, real-time smoke detection network, MLSD(MultiScale Lightweight Smoke Detection Network). First, in order to reduce the computational complexity and the number of parameters of the model, we propose a lightweight detection head called EISDH(Efficient Information Sharing Detection Head). Second, to reduce the extraction of redundant features, we innovatively propose the C2f-PConv module. Third, to enhance the extraction capabilities of multi-scale and subtle smoke features in complex visual scenes, the downsampling module ADown was innovatively integrated into the model. MLSD demonstrates superior performance on three testing benchmarks. Notably, on the custom dataset FFES(Forest Fire Early Smoke Dataset), compared to the baseline, MLSD improves its mAP50 by 1.7% and its accuracy by 2.9% while reducing the number of parameters by 6.1M and decreasing GFLOPs by 12.4G. Ganggang Li, Shaochen Jiang |
SMC | 3 |
| 2024 | PSA-Swin Transformer: Image Classification on Small-Scale DatasetsabstractThis paper introduces the PSA-Swin Transformer, a novel framework for image classification on small-scale datasets, highlighting the challenges of training effective models in resource-limited environments. Recognizing the limitations of current deep learning methods that rely heavily on large-scale datasets and extensive pre-training, we propose a method for handling small datasets. Our model can effectively han-dle smaller data volumes without the need for pre-training weights. The key to our approach is in the introduction of an efficient positional embedding (EPE) module, which improves parameter utilization and network expressiveness through a grouped convolutional architecture and shuffling operations for dynamic information exchange. In addition, we integrated the Polarized Self-Attention (PSA) module in Windows Multi-Head Self-Attention (W-MSA) and named the new module PSA-W-MSA; PSA addresses the complexity of learning element-specific attention by combining polarization filtering with augmentation techniques. Through a series of experiments on the Mini-Imagenet dataset, the PSA-Swin Transformer demonstrates notable performance, especially in environments where high-quality annotated data is scarce or costly to acquire. Our research results are expected to make progress in areas that require efficient and accurate image classification using limited resources. Chao Shao, Shaochen Jiang |
SMC | 2 |
| 2024 | MSViT: Training Multiscale Vision Transformers for Image RetrievalabstractThe recently developed vision transformer (ViT) has achieved promising results on image retrieval compared to convolutional neural networks. However, most of these vision transformer-based image retrieval methods use the original ViT model to extract global features, ignoring the importance of local features for image retrieval. In this work, we propose a vision transformer-based multiscale feature fusion image retrieval method (MSViT) to achieve the fusion of global features with local features. The challenge of this research work is how to learn the feature representation ability of transformer model, so as to improve the performance of image retrieval model. First, a transformer-based two-branch network structure is proposed to obtain different scale features by processing image patches with different granularities. Second, we present a multiscale feature fusion strategy, which can efficiently and effectively fuse the feature information of different sizes on two branches. Finally, to more fully utilize the label information to supervise the network training process, we optimize the construction rules for the triplet data. The comparison of experimental results with ten CNN-based and six transformer-based image retrieval methods on four publicly available image datasets shows that our method outperforms the state-of-the-art methods. And ablation experiments show that the designed multiscale feature fusion strategy and improved triplet loss function have an implicit improvement on the performance of MSViT. Xue Li 0008, Shaochen Jiang, Hongchun Lu, Ziyang Li 0010 |
IEEE Trans. Multim. | 3 |
| 2023 | MAFH: Multilabel aware framework for bit-scalable cross-modal hashing
Xue Li 0008, Hongchun Lu, Shaochen Jiang, Ziyang Li 0010, Peiyun Yao |
Knowl. Based Syst. | 4 |