VLDB 2026 Research / reviewers in the wild / expert
Shuyuan Lin
dblp:85/4569
· DBLP profile ↗
24ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0003-3247-4625ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SC-Net: Robust Correspondence Learning via Spatial and Cross-Channel ContextabstractRecent research has focused on using convolutional neural networks (CNNs) as the backbones in two-view correspondence learning, demonstrating significant superiority over methods based on multilayer perceptrons. However, CNN backbones that are not tailored to specific tasks may fail to effectively aggregate global context and oversmooth dense motion fields in scenes with large disparity. To address these problems, we propose a novel network named SC-Net, which effectively integrates bilateral context from both spatial and channel perspectives. Specifically, we design an adaptive focused regularization module (AFR) to enhance the model's position-awareness and robustness against spurious motion samples, thereby facilitating the generation of a more accurate motion field. We then propose a bilateral field adjustment module (BFA) to refine the motion field by simultaneously modeling long-range relationships and facilitating interaction across spatial and channel dimensions. Finally, we recover the motion vectors from the refined field using a position-aware recovery module (PAR) that ensures consistency and precision. Extensive experiments demonstrate that SC-Net outperforms state-of-the-art methods in relative pose estimation and outlier removal tasks on YFCC100M and SUN3D datasets. Shuyuan Lin, Hailiang Liao, Qiang Qi, Taotao Lai, Jian Weng 0001 |
AAAI | 1 |
| 2026 | MCI-Net: A Robust Multi-Domain Context Integration Network for Point Cloud RegistrationabstractRobust and discriminative feature learning is critical for high-quality point cloud registration. However, existing deep learning–based methods typically rely on Euclidean neighborhood-based strategies for feature extraction, which struggle to effectively capture the implicit semantics and structural consistency in point clouds. To address these issues, we propose a multi-domain context integration network (MCI-Net) that improves feature representation and registration performance by aggregating contextual cues from diverse domains. Specifically, we propose a graph neighborhood aggregation module, which constructs a global graph to capture the overall structural relationships within point clouds. We then propose a progressive context interaction module to enhance feature discriminability by performing intra-domain feature decoupling and inter-domain context interaction. Finally, we design a dynamic inlier selection method that optimizes inlier weights using residual information from multiple iterations of pose estimation, thereby improving the accuracy and robustness of registration. Extensive experiments on indoor RGB-D and outdoor LiDAR datasets show that the proposed MCI-Net significantly outperforms existing state-of-the-art methods, achieving the highest registration recall of 96.4% on 3DMatch. Shuyuan Lin, Wenwu Peng, Qiang Qi, Miaohui Wang, Jian Weng 0001 |
AAAI | 1 |
| 2026 | MSTDiff: Multiscale-Aware Transformer Diffusion Network for Video Object DetectionabstractVideo object detection is a fundamental yet challenging task in computer vision. Recently, DETR-based methods have gained prominence in this domain owing to their powerful global modeling capabilities. However, these methods are still confronted with two key limitations: frame-agnostic initialization of object queries and scale-agnostic attention mechanisms, which hinder their capability to capture the appearance variations of dynamic objects and model the temporal consistency across frames. To alleviate these limitations, we propose a multiscale-aware transformer diffusion network (MSTDiff), a novel framework designed for the video object detection task, including two technical improvements over existing methods. First, we design a diffusion-driven adaptive query module, which models the object query distribution through a diffusion process conditioned on input frames, enabling an adaptive and content-aware initialization of object queries. Second, we develop a multiscale-aware transformer encoder module, which combines multi-head convolutional units with attention mechanisms to enhance multi-scale feature representations while preserving global dependence modeling. We conduct extensive experiments on the public ImageNet VID dataset, and the results demonstrate that our MSTDiff achieves 87.7% mAP with ResNet-101, outperforming most previous state-of-the-art video object detection methods. Qiang Qi, Wenqi Shang, Shuyuan Lin |
AAAI | 5 |
| 2026 | Perceive More with Less: LiDAR Point Cloud Compression at Just Recognizable Distortion for 3D Scene UnderstandingabstractExisting LiDAR point cloud (LPC) data coding methods primarily focus on balancing compression efficiency and reconstruction quality according to the human vision system (HVS). However, these methods rarely consider the requirements of downstream scene understanding tasks from the perspective of the machine vision system (MVS). To address this challenge, we explore the maximum degree of LPC compression that has negligible impact on perception accuracy, called LPC-based just recognizable compression distortion (lpcJRCD). Specifically, we introduce a novel point-wise quantization approach for constructing a MVS-based LiDAR dataset and present a new lpcJRCD-guided intelligent compression framework tailored for MVS applications. To enhance MVS-based LPC compression efficiency, we develop a dual-feature interaction (DFI) module that fuses point and voxel features. Additionally, we propose a mask-based loss function to ensure accurate point-wise quality level prediction. Experimental results demonstrate the effectiveness of our proposed model in reducing the average bit rate by up to 94.98% while preserving perception accuracy in autonomous vehicles. Miaohui Wang, Runnan Huang, Taojun Liu, Shuyuan Lin, Ye Liu 0005, Yun Song |
AAAI | 4 |
| 2026 | LLHA-Net: A hierarchical attention network for two-view correspondence learning
Shuyuan Lin, Xiao Chen 0021, Guobao Xiao, Feiran Huang |
Pattern Recognit. | 1 |
| 2025 | DTSNet: A Denoising Teacher-Student Network with Reverse Distillation for Anomaly DetectionabstractKnowledge distillation has emerged as a promising method for unsupervised anomaly detection. However, the overgeneralization of the student network often reduces the distinction between teacher and student representations for anomalous samples, leading to detection failures. To address this problem, we propose a Denoising Teacher-Student Network (DTSNet), which integrates two teacher networks: a normal teacher and an anomalous teacher. These networks guide the student network in feature-space denoising, enabling it to restore anomalous features and amplify the representational disparity for anomalies. Furthermore, we propose an Attention-guided Perturbation Reconstruction (APR) module, which facilitates the student network to focus on critical pixel regions, enhancing its feature representation capability. Experimental results demonstrate that the proposed DTSNet outperforms several state-of-the-art methods on the MVTec AD, VisA, and BTAD datasets. Source code is available at http://www.linshuyuan.com. Taixiang Lin, Shuyuan Lin |
ICME | 2 |
| 2025 | MGCA-Net: Multi-Graph Contextual Attention Network for Two-View Correspondence LearningabstractTwo-view correspondence learning is a key task in computer vision, which aims to establish reliable matching relationships for applications such as camera pose estimation and 3D reconstruction. However, existing methods have limitations in local geometric modeling and cross-stage information optimization, which make it difficult to accurately capture the geometric constraints of matched pairs and thus reduce the robustness of the model. To address these challenges, we propose a Multi-Graph Contextual Attention Network (MGCA-Net), which consists of a Contextual Geometric Attention (CGA) module and a Cross-Stage Multi-Graph Consensus (CSMGC) module. Specifically, CGA dynamically integrates spatial position and feature information via an adaptive attention mechanism and enhances the capability to capture both local and global geometric relationships. Meanwhile, CSMGC establishes geometric consensus via a cross-stage sparse graph network, ensuring the consistency of geometric information across different stages. Experimental results on two representative YFCC100M and SUN3D datasets show that MGCA-Net significantly outperforms existing SOTA methods in the outlier rejection and camera pose estimation tasks. Source code is available at http://www.linshuyuan.com. Shuyuan Lin, Mengtin Lo, Haosheng Chen 0001, Qiangqiang Wu |
IJCAI | 1 |
| 2025 | Two-View Correspondence Pruning via Channel-Spatial Interaction and Bidirectional Consensus InteractionabstractAccurately identifying correct correspondences in two images is a crucial task in computer vision. Current methods predominantly use PointCN blocks as feature extraction backbones and learn local-global consensus through a progressive learning strategy. However, such methods have two main drawbacks: First, PointCN blocks, composed of multilayer perceptrons and normalization layers, process spatial positions independently, leading to limited interaction between channel-wise and spatial-wise dimensions. Second, the progressive learning strategy primarily focuses on unidirectional transfer from local to global consensus, yet neglects the bidirectional interaction between local and global consensus. To address these issues, we propose the Channel-Spatial interaction and Bidirectional Consensus interaction-Based Network (CSBCNet), which contains three innovative blocks: Channel-Spatial Interaction (CSI), Local Consensus Mining (LCM), and Global Consensus-Aware Attention (GCAA). Specifically, CSI enhances interaction between channel-wise and spatial-wise dimensions through a dual-path attention mechanism, addressing the limited interaction caused by the independent processing of spatial positions in PointCN blocks. LCM extracts reliable local consensus by modeling geometric structures and spatial continuity within correspondences. GCAA captures global consensus by aggregating correspondences that are highly likely to be correct ones, and achieves bidirectional interaction between local and global consensus through cross attention. Experiments demonstrate our CSBCNet's superior performance in camera pose estimation and correspondence pruning. Notably, when the CSI block is applied to the existing OANet and MS2DGNet networks, it achieves significant performance improvements of 10.27% and 7.5%, respectively, on the mAP5° metric on the camera pose estimation task. Xiangui Huang, Taotao Lai, Yizhang Liu, Shuyuan Lin |
ACM Multimedia | 4 |
| 2025 | Adaptive Graph Attention-Guided Parallel Sampling and Embedded Selection for Multi-Model FittingabstractMulti-model fitting is a fundamental challenge in computer vision, where real-world data often contains severe gross outliers and pseudo-outliers. Existing methods rely on inefficient sequential hypothesize-and-verify frameworks that require a predefined number of models and inlier thresholds-parameters that are difficult to determine in practical scenes. To overcome these limitations, we propose a novel Adaptive Graph Attention-guided parallel multi-model fitting method (AGASAC) that jointly learns local and global features, performs parallel hypothesis sampling, and executes confidence-embedded model selection. Specifically, we design a dual-confidence graph attention module that models data relationships using an adaptive graph attention network. This module computes minimal-set confidence and quality confidence to guide the multi-model fitting process, eliminating manual parameter tuning. Additionally, we propose a parallel discriminative sampling module that leverages minimal-set confidence to concurrently sample hypotheses. By enforcing a quantized consensus constraint, this module maximizes inter-model variance while minimizing intra-model discrepancy. It enables computationally efficient hypothesis generation and pseudo-outlier suppression. To obtain high-quality models, we present a quality-embedded selection module that integrates quality confidence into the joint optimization of model selection and data clustering. Extensive experiments show that the proposed method achieves a lower transfer error of 0.39 pixels and a 36.92% runtime reduction, surpassing state-of-the-art methods. The code is available at https://github.com/YWY-Vivian/AGASAC. Wenyu Yin, Shuyuan Lin, David Suter, Hanzi Wang |
ACM Multimedia | 2 |
| 2025 | Progressive Enhancement Dehazing for object detection in extreme weather
Zhiying Li 0003, Junhao Wu 0003, Shuyuan Lin, Zheng Wang 0013, Xiao-Bo Jin, Guanggang Geng, Feiran Huang, Jian Weng 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Diverse Consensuses Paired with Motion Estimation-Based Multi-Model FittingabstractMulti-model fitting aims to robustly estimate the parameters of various model instances in data contaminated by noise and outliers. Most previous works employ only a single type of consensus or implicit fusion model to represent the correlation between data points and model hypotheses. This approach often results in unrealistic and incorrect model fitting in the presence of noise and uncertainty. In this paper, we propose a novel method of diverse Consensuses paired with Motion estimation-based multi-Model Fitting (CMMF), which leverages three types of diverse consensuses along with inter-model collaboration to enhance the effectiveness of multi-model fusion. We design a Tangent Consensus Residual Reconstruction (TCRR) module to capture motion structure information of two points at the pixel level. Additionally, we introduce a Cross Consensus Affinity (CCA) framework to strengthen the correlation between data points and model hypotheses. To address the challenge of multi-body motion estimation, we propose a Nested Consensus Clustering (NCC) strategy, which formulates multi-model fitting as a motion estimation problem. It explicitly establishes motion collaboration between models and ensures that multiple models are well-fitted. Extensive quantitative and qualitative experiments are conducted on four public datasets (i.e., AdelaideRMF-F, Hopkins155, KITTI, MTPV62), and the results demonstrate that our proposed method outperforms several state-of-the-art methods. Wenyu Yin, Shuyuan Lin, Yang Lu 0009, Hanzi Wang |
ACM Multimedia | 2 |
| 2024 | GNN-fused CapsNet with multi-head prediction for diabetic retinopathy grading
Yongjia Lei, Shuyuan Lin, Zhiying Li 0003, Taotao Lai |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Robust Heterogeneous Model Fitting for Multi-source Image Correspondences
Shuyuan Lin, Feiran Huang, Taotao Lai, Jian-Huang Lai, Hanzi Wang, Jian Weng 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Multi-Motion Segmentation via Co-Attention-Induced Heterogeneous Model FittingabstractMotion segmentation is an essential task in artificial intelligence and computer vision. However, scene motion in real-world intelligent systems usually integrates multiple types of models, so specifying only one type of basic model may lead to the failure of scene-motion segmentation tasks. In this paper, we propose a novel and efficient heterogeneous model-fitting-based motion segmentation method (HMFMS) to accurately segment moving objects. HMFMS includes a new co-attention-induced heterogeneous model construction algorithm (HMC), an adaptive heterogeneous model refinement algorithm (HMR), and a heterogeneous model segmentation algorithm (HMS). First, we propose HMC to generate high-quality accumulated correlation matrices, by evaluating the quality of heterogeneous model hypotheses, based on the density estimation technique. Next, we propose HMR to construct sparse affinity matrices from the accumulated correlation matrices by applying information theory, effectively suppressing the values of correlations between different objects. Finally, we fuse the sparse affinity matrices and perform motion segmentation by using HMS, to obtain more accurate segmentation results. Experimental results show that HMFMS obtains superior performance on four challenging datasets (i.e., Hopkins155, Hopkins12, MTPV62 and KT3DMoSeg), compared with several subspace-based and model-fitting-based motion segmentation methods. More remarkably, HMFMS outperforms the state-of-the-art MCMS method by 57.1% and 1.8 times in terms of accuracy and computational efficiency on the representative KT3DMoSeg, respectively. Shuyuan Lin, Anjia Yang, Taotao Lai, Jian Weng 0001, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Multi-Stage Network With Geometric Semantic Attention for Two-View Correspondence LearningabstractThe removal of outliers is crucial for establishing correspondence between two images. However, when the proportion of outliers reaches nearly 90%, the task becomes highly challenging. Existing methods face limitations in effectively utilizing geometric transformation consistency (GTC) information and incorporating geometric semantic neighboring information. To address these challenges, we propose a Multi-Stage Geometric Semantic Attention (MSGSA) network. The MSGSA network consists of three key modules: the multi-branch (MB) module, the GTC module, and the geometric semantic attention (GSA) module. The MB module, structured with a multi-branch design, facilitates diverse and robust spatial transformations. The GTC module captures transformation consistency information from the preceding stage. The GSA module categorizes input based on the prior stage's output, enabling efficient extraction of geometric semantic information through a graph-based representation and inter-category information interaction using Transformer. Extensive experiments on the YFCC100M and SUN3D datasets demonstrate that MSGSA outperforms current state-of-the-art methods in outlier removal and camera pose estimation, particularly in scenarios with a high prevalence of outliers. Source code is available at https://github.com/shuyuanlin. Shuyuan Lin, Xiao Chen 0021, Guobao Xiao, Hanzi Wang, Feiran Huang, Jian Weng 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | DQFORMER: Dynamic Query Transformer for Lane DetectionabstractLane detection is one of the most important tasks in self-driving. The critical purpose of lane detection is the prediction of lane shapes. Meanwhile, it is challenging and difficult to determine lane instance positions before predicting lane shapes in an image. In this paper, we propose a top-down method called Dynamic Query Transformer (DQFormer), which uses a Dynamic Lane Queries (DLQs) module to predict lane shapes. Specifically, to accurately predict lane shapes, we propose a new framework for generating dynamic weights based on DLQs, which can focus on the context of lane shapes dynamically. Unlike existing transformer-based methods, the proposed DQFormer does not require setting a fixed number of lane queries, so it is suitable for various scenes. In addition, we further propose a Line Voting Module (LVM) which collects votes from other lanes to enhance lane features, to determine lane instance positions. Extensive experiments demonstrate that DQFormer outperforms several state-of-the-art methods on two popular lane detection benchmarks (i.e., CULane and TuSimple). Shuyuan Lin, Runqing Jiang, Yang Lu 0009, Hanzi Wang |
ICASSP | 2 |
| 2023 | Robust model estimation by using preference analysis and information theory principles
Taotao Lai, Weice Wang, Yizhang Liu, Shuyuan Lin |
Appl. Intell. | 5 |
| 2023 | Efficient sampling using feature matching and variable minimal structure size
Taotao Lai, Alireza Sadri, Shuyuan Lin, Riqing Chen, Hanzi Wang |
Pattern Recognit. | 3 |
| 2022 | SCINet: Semantic Cue Infusion Network for Lane DetectionabstractNowadays, lane detection plays an important role in autonomous driving. However, the task of lane detection still faces many challenges, such as external no-visual-clue and internal sparse supervisory signals. In this work, we propose a novel Semantic Cue Infusion Network (SCINet) that uses semantic cues to aid lane detection. Specifically, in order to overcome the external no-visual-clue condition, we introduce a strategy that utilizes semantic cues as additional supervisory signals, which facilitate SCINet to collect region-aware features in the shared layer. We also design a hypernetwork with the additional signals used as a critical part to generate dynamic weights for downstream output heads. Furthermore, we design two Slice Attention Modules (SAMs) based on the interdependencies between slices to improve the robustness of SCINet in distinguishing features between lanes and background. Experiments on two popular lane detection benchmarks (i.e., TuSimple and CULane) show that SCINet significantly outperforms several state-of-the-art methods. Shuyuan Lin, Yang Lu 0009, Hanzi Wang |
ICIP | 2 |
| 2022 | Triplet Relationship Guided Sampling Consensus for Robust Model EstimationabstractRANSAC (RANdom SAmple Consensus) is a widely used robust estimator for estimating a geometric model from feature matches in an image pair. Unfortunately, it becomes less effective when initial input feature matches (i.e., input data) are corrupted by a large number of outliers. In this paper, we propose a new robust estimator (called TRESAC) for model estimation, where data subsets are sampled with the guidance of the triplet relationships, which involve high relevance and local geometric consistency. Each triplet consists of three data, whose relationships satisfy spatial consistency constraints. Therefore, the triplet relationships can be used to effectively initialize and refine the sampling process. With the advantage of the triplet relationships, TRESAC significantly alleviates the influence of outliers and also improves the computational efficiency of model estimation. Experimental results on four challenging datasets show that TRESAC can achieve superior performance on both estimation accuracy and computational efficiency against several other state-of-the-art methods. Hanlin Guo, Yang Lu 0009, Guobao Xiao, Shuyuan Lin, Hanzi Wang |
IEEE Signal Process. Lett. | 4 |
| 2022 | Co-Clustering on Bipartite Graphs for Robust Model FittingabstractRecently, graph-based methods have been widely applied to model fitting. However, in these methods, association information is invariably lost when data points and model hypotheses are mapped to the graph domain. In this paper, we propose a novel model fitting method based on co-clustering on bipartite graphs (CBG) to estimate multiple model instances in data contaminated with outliers and noise. Model fitting is reformulated as a bipartite graph partition behavior. Specifically, we use a bipartite graph reduction technique to eliminate some insignificant vertices (outliers and invalid model hypotheses), thereby improving the reliability of the constructed bipartite graph and reducing the computational complexity. We then use a co-clustering algorithm to learn a structured optimal bipartite graph with exact connected components for partitioning that can directly estimate the model instances (i.e., post-processing steps are not required). The proposed method fully utilizes the duality of data points and model hypotheses on bipartite graphs, leading to superior fitting performance. Exhaustive experiments show that the proposed CBG method performs favorably when compared with several state-of-the-art fitting methods. Shuyuan Lin, Hailing Luo, Yan Yan 0001, Guobao Xiao, Hanzi Wang |
IEEE Trans. Image Process. | 1 |
| 2020 | Extraction and Classification of TCM Medical Records Based on BERT and Bi-LSTM With Attention MechanismabstractTraditional Chinese Medicine (TCM) medical records contain huge amounts of valuable medical information. However, in terms of text mining and utilization of TCM medical records, it is always difficult to extract and classify this information effectively. It is critical to identify a method of extracting and classifying the text from TCM medical records automatically. The method used in this paper attempts to apply a short medical record classification model based on BERT and Bi-LSTM with Attention mechanism. BERT prepossessing was used to obtain the short text vector as the input of the model. Result shows that the BERT-Bi-LSTM-Attention model achieves a highest average F1 value of 89.52% in the extraction and classification of TCM medical records, and therefore represents a significant improvement in modeling. Ye Hui, Shuyuan Lin, YiQian Qu |
BIBM | 3 |
| 2019 | Hypergraph Optimization for Multi-Structural Geometric Model FittingabstractRecently, some hypergraph-based methods have been proposed to deal with the problem of model fitting in computer vision, mainly due to the superior capability of hypergraph to represent the complex relationship between data points. However, a hypergraph becomes extremely complicated when the input data include a large number of data points (usually contaminated with noises and outliers), which will significantly increase the computational burden. In order to overcome the above problem, we propose a novel hypergraph optimization based model fitting (HOMF) method to construct a simple but effective hypergraph. Specifically, HOMF includes two main parts: an adaptive inlier estimation algorithm for vertex optimization and an iterative hyperedge optimization algorithm for hyperedge optimization. The proposed method is highly efficient, and it can obtain accurate model fitting results within a few iterations. Moreover, HOMF can then directly apply spectral clustering, to achieve good fitting performance. Extensive experimental results show that HOMF outperforms several state-of-the-art model fitting methods on both synthetic data and real images, especially in sampling efficiency and in handling data with severe outliers. Shuyuan Lin, Guobao Xiao, Yan Yan 0001, David Suter, Hanzi Wang |
AAAI | 1 |
| 2009 | Spanning Tree Based Attribute Clustering
Yifeng Zeng, Jorge Cordero Hernandez, Shuyuan Lin |
PAKDD | 3 |