VLDB 2026 Research / reviewers in the wild / expert
Lingtong Min
dblp:156/8392
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-3970-7823ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task-Adapter++: Task-specific adaptation with order-aware alignment for few-shot action recognition
Congqi Cao, Peiheng Han, Yueran Zhang, Yating Yu, Qinyi Lv, Lingtong Min, Yanning Zhang 0001 |
Pattern Recognit. | 6 |
| 2025 | Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIPabstractZero-shot action recognition (ZSAR) requires collaborative multi-modal spatiotemporal understanding. However, finetuning CLIP directly for ZSAR yields suboptimal performance, given its inherent constraints in capturing essential temporal dynamics from both vision and text perspectives, especially when encountering novel actions with fine-grained spatiotemporal discrepancies. In this work, we propose Spatiotemporal Dynamic Duo (STDD), a novel CLIP-based framework to comprehend multi-modal spatiotemporal dynamics synergistically. For the vision side, we propose an efficient Space-time Cross Attention, which captures spatiotemporal dynamics flexibly with simple yet effective operations applied before and after spatial attention, without adding additional parameters or increasing computational complexity. For the semantic side, we conduct spatiotemporal text augmentation by comprehensively constructing an Action Semantic Knowledge Graph (ASKG) to derive nuanced text prompts. The ASKG elaborates on static and dynamic concepts and their interrelations, based on the idea of decomposing actions into spatial appearances and temporal motions. During the training phase, the frame-level video representations are meticulously aligned with prompt-level nuanced text representations, which are concurrently regulated by the video representations from the frozen CLIP to enhance generalizability. Extensive experiments validate the effectiveness of our approach, which consistently surpasses state-of-the-art approaches on popular video benchmarks (i.e., Kinetics-600, UCF101, and HMDB51) under challenging ZSAR settings. Yating Yu, Congqi Cao, Yueran Zhang, Qinyi Lv, Lingtong Min |
AAAI | 5 |
| 2025 | Autoregressive Denoising Score Matching Is a Good Video Anomaly DetectorabstractVideo anomaly detection (VAD) is an important computer vision problem. Thanks to the mode coverage capabilities of generative models, the likelihood-based paradigm is catching growing interest, as it can model normal distribution and detect out-of-distribution anomalies. However, these likelihood-based methods are blind to the anomalies located in local modes near the learned distribution. To handle these ``unseen" anomalies, we dive into three gaps uniquely existing in VAD regarding scene, motion and appearance. Specifically, we first build a noise-conditioned score transformer for denoising score matching. Then, we introduce a scene-dependent and motion-aware score function by embedding the scene condition of input sequences into our model and assigning motion weights based on the difference between key frames of input sequences. Next, to solve the problem of blindness in principle, we integrate unaffected visual information via a novel autoregressive denoising score matching mechanism for inference. Through autoregressively injecting intensifying Gaussian noise into the denoised data and estimating the corresponding score function, we compare the denoised data with the original data to get a difference and aggregate it with the score function for an enhanced appearance perception and accumulate the abnormal context. With all three gaps considered, we can compute a more comprehensive anomaly indicator. Experiments on three popular VAD benchmarks demonstrate the state-of-the-art performance of our method. Hanwen Zhang 0017, Congqi Cao, Qinyi Lv, Lingtong Min, Yanning Zhang 0001 |
ICCV | 4 |
| 2025 | CM-YOLO: Context Modulated Representation Learning for Ship DetectionabstractShip detection is essential for both military and civilian applications. Existing ship detection methods focus on prominent offshore ships, paying less attention to complex nearshore ships, which are easily confused with the intricate background. Utilizing contextual information, such as location and shape, can enhance ship detection and classification in complex environments. In this article, we propose a context modulated representation learning-based detection method termed as CM-YOLO. It adopts the classical detector design framework, which includes the backbone, neck, and head. The input image is sequentially processed through these components to obtain the detection results. Our method specifically optimizes ship detection in complex scenarios. To achieve this, we propose a dual path context enhancement neck (DCEN) to extract contextual information for ship detection. The neck builds on the path augmentation feature pyramid network with the proposed dual path context enhancement (DCE) module, which is designed to enhance feature representations by incorporating high-level semantic information. It captures long-range dependencies across both channel and spatial dimensions while suppressing irrelevant features. Additionally, to enhance the scale-aware capability of the head for detecting multiscale ships in complex environments, we introduce the multicontext boosted (MCB) detection head. The MCB can flexibly adjust the receptive field and extracts relevant context for ships of various scales using multiple large-kernel convolutions. We conduct experiments on three commonly used ship datasets: Seaships7000, ShipRSImageNet, DIOR-ship, and HRSC2016. Experiment results demonstrate that CM-YOLO achieves excellent performance compared with other leading ship detection methods. Lingtong Min, Feiyang Dou, Dian Shao, Binglu Wang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Adaptive Fusion Learning for Compositional Zero-Shot RecognitionabstractCompositional Zero-Shot Learning (CZSL) aims to learn visual concepts (i.e., attributes and objects) from seen compositions and combine them to predict unseen compositions. Existing visual encoders in CZSL typically use traditional visual encoders (i.e., CNN and Transformer) or image encoders from Visual-Language Models (VLMs) to encode image features. However, traditional visual encoders need more multi-modal textual information, and image encoders of VLMs exhibit dependence on pre-training data, making them less effective when used independently for predicting unseen compositions. To overcome this limitation, we propose a novel approach based on the joint modeling of traditional visual encoders and VLMs visual encoders to enhance the prediction ability for uncommon and unseen compositions. Specifically, we design an adaptive fusion module that automatically adjusts the weighted parameters of similarity scores between traditional and VLMs methods during training, and these weighted parameters are inherited during the inference process. Given the significance of disentangling attributes and objects, we design a Multi-Attribute Object Module that, during the training phase, incorporates multiple pairs of attributes and objects as prior knowledge, leveraging this rich prior knowledge to facilitate the disentanglement of attributes and objects. Building upon this, we select the text encoder from VLMs to construct the Adaptive Fusion Network. We conduct extensive experiments on the Clothing16 K, UT-Zappos50 K, and C-GQA datasets, achieving excellent performance on the Clothing16 K and UT-Zappos50 K datasets. Lingtong Min, Ziman Fan, Shunzhou Wang, Feiyang Dou, Xin Li 0042, Binglu Wang |
IEEE Trans. Multim. | 1 |
| 2024 | Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action Recognition
Congqi Cao, Yueran Zhang, Yating Yu, Qinyi Lv, Lingtong Min, Yanning Zhang 0001 |
ACM Multimedia | 5 |
| 2024 | Rough-Fuzzy Graph Learning Domain Adaptation for Fake News DetectionabstractThe widespread dissemination of fake news across the internet has profound detrimental consequences for society, governments, and citizens. To address this pressing issue, numerous machine learning-based models have been developed for detecting fake news. However, the challenge of acquiring sufficient labeled news data in a new domain, coupled with the presence of inconsistent data distribution, necessitates the integration of unsupervised domain adaptation (DA) methods to enhance the reliability of cross-domain fake news detection. In this article, a rough-fuzzy graph learning DA for fake news detection is proposed. First, a rough-fuzzy graph learning method is proposed to effectively handle the representation of cross-domain sample uncertainty structural information, thereby learning a more discriminative subspace. Second, a rough-fuzzy region division strategy is designed to perform different analysis on target domain samples, thus achieving a more accurate description of the relationships between cross-domain samples. Furthermore, considering that domain private features may negatively affect the knowledge transfer process, a sparse structure preserving strategy is proposed to better capture shared general features across domains. Experimental evaluations conducted on three news datasets demonstrate the efficacy of the proposed method in cross-domain fake news detection. Jiao Shi, Yu Lei 0002, Lingtong Min |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | VS-TransGRU: A Novel Transformer-GRU-Based Framework Enhanced by Visual-Semantic Fusion for Egocentric Action AnticipationabstractEgocentric action anticipation is a challenging task that aims to make advanced predictions of future actions from current and historical observations in the first-person view. Most existing methods focus on improving the model architecture and loss function based on the visual input and recurrent neural network to boost the anticipation performance. However, these methods, which merely consider visual information and rely on a single network architecture, gradually reach a performance plateau. In order to fully understand what has been observed and capture the dependencies between current observations and future actions well enough, we propose a novel visual-semantic fusion enhanced and Transformer-GRU-based action anticipation framework in this paper. Firstly, high-level semantic information is introduced to improve the performance of action anticipation for the first time. We propose to use the semantic features generated based on the class labels or directly from the visual observations to augment the original visual features. Secondly, to take advantage of both the parallel and autoregressive models, we design a Transformer-based encoder for long-term sequential modeling and a GRU-based decoder for flexible iteration decoding. This hybrid architecture allows for better performance with fewer parameters and computations. Thirdly, an effective visual-semantic fusion module is proposed to make up for the semantic gap and fully utilize the complementarity of different modalities. Extensive experiments on two large-scale first-person view datasets and two third-person datasets validate the effectiveness of our proposed method, which achieves new state-of-the-art performance, outperforming previous approaches by a large margin. The code will be released after acceptance athttps://github.com/sunze992/VS-TransGRU. Congqi Cao, Ze Sun, Qinyi Lv, Lingtong Min, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Aerial-Ground Integrated Vehicular Networks: A UAV-Vehicle Collaboration PerspectiveabstractUnmanned aerial vehicle mounted base stations (UAV-BSs) are expected to become an integral component of future intelligent transportation systems, which can provide seamless coverage for vehicles on highways with poor cellular infrastructures. Motivated by the above, this paper proposes an aerial-ground integrated vehicular networking architecture, based on which a UAV-vehicle collaboration perspective is proposed. Specifically, an emerging vehicle-to-UAV (V2U) and vehicle-to-vehicle (V2V) collaboration framework is first presented to facilitate diverse vehicular applications. Next, we investigate the coverage radius maximization problem by optimizing the UAV-BS altitude. Meanwhile, by taking the channel state information (CSI) feedback delay into account, we formulate a V2U communication sum rate maximization problem by optimizing the power control and spectrum allocation, which is constrained by the capacity and reliability requirements. Then, we derive the closed-form expression of optimal UAV-BS altitude. Afterwards, we decouple the formulated sum rate maximization problem, and devise an efficient algorithm with polynomial complexity, where the optimal power control and spectrum sharing are solved. Finally, simulation results demonstrate that the maximum coverage radius and optimal UAV-BS altitude can be achieved by our proposed scheme in different urban environments. In addition, our designed scheme can effectively improve the V2U communication sum rate in comparison with the current works. Yixin He 0001, Dawei Wang 0001, Fanghui Huang, Ruonan Zhang 0001, Lingtong Min |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Uplink Secrecy Performance of RIS-Based RF/FSO Three-Dimension Heterogeneous NetworksabstractIn this paper, a novel reconfigurable intelligent surface (RIS)-assisted HAP-UAV secure multi-user mixed radio frequency (RF)/free space optical (FSO) system is proposed. Specifically, the Gamma-Gamma distribution is utilized to characterize the atmospheric turbulence effect for the FSO link from UAV to HAP, while the Rayleigh and Nakagami-$m$distribution fading are applied to simulate the legitimate and wiretap RF links, respectively. We present the closed-form expressions for the probability density functions, the cumulative distribution functions, and the secrecy outage probability (SOP) of the end-to-end signal-to-noise ratio (SNR) in terms of Meijer’s G-function. To gain more insight into secrecy performance, we further obtain the closed-form expressions for the asymptotic SOP, the asymptotic probability of positive secrecy capacity (PPSC), the diversity gain, and the coding gain at high SNR regions. We can observe that the secrecy performance depends on the weaker channel between the RF and FSO, and is closely related to the number of RIS elements, the number of terrestrial users, the atmospheric turbulence factor, pointing error parameters, and the fading parameter of Nakagami-$m$distributed wiretap link. Finally, numerical results validate the derived results and demonstrate that the proposed design achieves superior secrecy performance over the benchmarks. Dawei Wang 0001, Zhongxiang Wei, Keping Yu, Lingtong Min, Shahid Mumtaz |
IEEE Trans. Wirel. Commun. | 5 |
| 2023 | Secrecy Performance Analysis of RIS-Aided Hybrid RF/FSO NetworksabstractThe proposed study introduces a reconfigurable intelligent surface (RIS)-aided hybrid radio frequency (RF)/free space optical (FSO) system with an unmanned aerial vehicle (UAV) relay to enable an ultra-dense sixth-generation (6G) network. The channels for RF and FSO are represented by Rayleigh and Gamma-Gamma probability distributions, correspondingly. Additionally, the network includes a terrestrial eavesdropper that follows the Nakagami-m distribution, attempting to breach confidential information. To counter this threat, RIS technology is used to enhance the hybrid system's secrecy. The study conducts a closed-form analysis of the secrecy outage probability (SOP) and obtains its asymptotic expression for determining the diversity order and coding gain. Theoretical findings have been confirmed through thorough numerical simulations implemented with the Monte-Carlo approach. The findings demonstrate the RIS technology's effectiveness in enhancing the network's secrecy performance. Dawei Wang 0001, Lingtong Min, Yixin He 0001, Li Zhen, Keping Yu |
GLOBECOM | 3 |
| 2023 | SFRNet: Fine-Grained Oriented Object Recognition via Separate Feature RefinementabstractFine-grained oriented object recognition (FGO2R) is a practical need for intellectually interpreting remote sensing images. It aims at realizing fine-grained classification and precise localization with oriented bounding boxes, simultaneously. Our considerations for the task are general but decisive: (i) the extraction of subtle differences carries a big weight in differentiating fine-grained classes, and (ii) oriented localization prefers rotation-sensitive features. In this article, we propose a network with separate feature refinement (SFRNet), in which two transformer-based branches are designed to perform function-specific feature refinement for fine-grained classification and oriented localization, separately. To highlight the discriminative information advantageous to fine-grained classification, we propose a spatial and channel transformer (SC-Former) to capture both the long-range spatial interactions and the key correlations hidden in the feature channels. Besides, we design a Multi-RoI loss (MRL) following the protocol of deep metric learning to enhance the separability of fine-grained classes further. For oriented localization, we integrate the oriented response convolution with the transformer structure (namely, OR-Former) to assist in encoding rotation information during regression. Extensive experimental results validate the effectiveness and robustness of our SFRNet. Without bells and whistles, our SFRNet achieves state-of-the-art performance on the large-scale FAIR1M datasets (FAIR1M-1.0 and FAIR1M-2.0). Code will be available at https://github.com/Ranchosky/SFRNet. Gong Cheng 0003, Qingyang Li 0001, Guangxing Wang 0001, Xingxing Xie, Lingtong Min, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Cross-Spatial Pixel Integration and Cross-Stage Feature Fusion-Based Transformer Network for Remote Sensing Image Super-ResolutionabstractRemote sensing image super-resolution (RSISR) plays a vital role in enhancing spatial detials and improving the quality of satellite imagery. Recently, Transformer-based models have shown competitive performance in RSISR. To mitigate the quadratic computational complexity resulting from global self-attention, various methods constrain attention to a local window, enhancing its efficiency. Consequently, the receptive fields in a single attention layer are inadequate, leading to insufficient context modeling. Furthermore, while most transform-based approaches reuse shallow features through skip connections, relying solely on these connections treats shallow and deep features equally, impeding the model’s ability to characterize them. To address these issues, we propose a novel transformer architecture called Cross-Spatial Pixel Integration and Cross-Stage Feature Fusion Based Transformer Network (SPIFFNet) for RSISR. Our proposed model effectively enhances context cognition and understanding of the entire image, facilitating efficient integration of features cross-stages. The model incorporates Cross-Spatial Pixel Integration Attention (CSPIA) to introduce contextual information into a local window, while Cross-Stage Feature Fusion Attention (CSFFA) adaptively fuses features from the previous stage to improve feature expression in line with the requirements of the current stage. We conducted comprehensive experiments on multiple benchmark datasets, demonstrating the superior performance of our proposed SPIFFNet in terms of both quantitative metrics and visual quality when compared to state-of-the-art methods. Our code is available at https://github.com/Dr-Lyt/SPIFFNet. Lingtong Min, Binglu Wang, Le Zheng, Yongqiang Zhao 0001, Le Yang 0008, Teng Long 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |