EDBT 2026 Demo / reviewers in the wild / expert
Shuting Dong
dblp:304/7544
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-View Clustering with Granularity-Aware Pseudo SupervisionabstractModern multi-view clustering (MVC) is dominated by two paradigms: multi-view fusion and pseudo-label-guided learning. Pseudo-labeling methods can suffer from confirmation bias; their reliance on a fixed-granularity supervision from an initial clustering can cause learned embeddings to drift from the data's true structure and lose discriminative power. Conversely, fusion methods excel at integrating information but often struggle to robustly differentiate between high-quality and noisy views, which can obscure final cluster boundaries and degrade performance. To address these complementary challenges, we propose GAPS (Granularity-Aware Pseudo Supervision), a novel MVC framework. GAPS introduces a granularity-aware supervision mechanism that generates a full hierarchy of pseudo-labels, enabling the selection of a supervision level that best aligns with the data's intrinsic multi-scale structure. Furthermore, to ensure a high-quality supervisory signal, it incorporates a reliability-aware view selection strategy using a novel Separation-Compactness Index (SCI) to identify and leverage the most informative view for pseudo-label generation. This dual approach ensures the supervisory signal is both structurally adaptive and derived from the most reliable source, leading to highly effective final representations. Extensive experiments on synthetic and real-world datasets demonstrate the effectiveness and superiority of GAPS over other competitors. Jie Yang 0052, Cheng-You Lu, Zhongli Wang 0001, Hsiang-Ting Chen, Guangkui Xu, Shuting Dong, Xinyan Liang, Bingbing Jiang 0001 |
AAAI | 7 |
| 2026 | High-Frequency Prioritized Sparse Attention Network for Image RestorationabstractImage restoration aims to restore high-quality images from degraded inputs caused by factors such as motion blur, defocus blur, and rain, where the primary difference between degraded and high-quality images lies in their high-frequency components. Despite the critical role of high frequencies in restoration, few methods explicitly prioritize computational resources for high frequencies over low frequencies. To address this issue, we propose a High-Frequency Prioritized Sparse Attention Network (HFP-SAN), a novel architecture for image restoration tasks. We explicitly prioritize high-frequency components by designing a symmetric encoder-decoder framework integrated with High-Frequency Selective Sparse Attention (HFSSA) modules while handling low-frequency components with a smaller residual network, thereby proportionally allocating computational resources based on their relative importance. HFSSA incorporates a Frequency-Selective Matching (FSM) algorithm to focus attention on strongly correlated high-frequency regions, mitigating computation on areas with weak correlations and irrelevant areas. Additionally, we introduce a dynamically adjustable high-frequency mask that guides the network to focus on the severely degraded regions, further refining restoration quality. The above designs ensure the final reconstructed image is a high-quality product. Experiments demonstrate that our HFP-SAN achieves state-of-the-art performance across multiple image restoration tasks, both quantitatively and qualitatively. Shuting Dong, Zhe Wu 0006, Hongyang Wei, Mingzhi Chen 0003, Guanghao Li 0003, Haolong Qian, Hanyang Peng, Chun Yuan 0003 |
IEEE Trans. Multim. | 1 |
| 2025 | ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle MatricesabstractThe Transformer architecture has revolutionized various fields since it was proposed, where positional encoding plays an essential role in effectively capturing sequential order and context. Therefore, Rotary Positional Encoding (RoPE) was proposed to alleviate these issues, which integrates positional information by rotating the embeddings in the attention mechanism. However, RoPE utilizes manually defined rotation matrices, a design choice that favors computational efficiency but limits the model’s flexibility and adaptability. In this work, we propose ComRoPE, which generalizes RoPE by defining it in terms of trainable commuting angle matrices. Specifically, we demonstrate that pairwise commutativity of these matrices is essential for RoPE to achieve scalability and positional robustness. We formally define the RoPE Equation, which is an essential condition that ensures consistent performance with position offsets. Based on the theoretical analysis, we present two types of trainable commuting angle matrices as sufficient solutions to the RoPE equation, which significantly improve performance, surpassing the current state-of-the-art method by 1.6% at training resolution and 2.9% at higher resolution on the ImageNet-1K dataset. Furthermore, our framework shows versatility in generalizing to existing RoPE formulations and offering new insights for future positional encoding research. To ensure reproducibility, the source code and instructions are available at https://github.com/Longin-Yu/ComRoPE. Tangyu Jiang, Shuning Jia, Shannan Yan, Shunning Liu, Haolong Qian, Guanghao Li 0003, Shuting Dong, Chun Yuan 0003 |
CVPR | 8 |
| 2025 | Vpr-Cloak: a First Look at Privacy Cloak Against Visual Place Recognition
Shuting Dong, Mingzhi Chen 0003, Guanghao Li 0003, Zhe Wu 0006, Ming Tang 0006, Chun Yuan 0003 |
ICCV | 1 |
| 2025 | SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-ThoughtabstractChain-of-Thought (CoT) prompting improves the reasoning performance of large language models (LLMs) by encouraging step-by-step thinking. However, CoT-based methods depend on intermediate reasoning steps, which limits scalability and generalization. Recent work explores recursive reasoning, where LLMs reuse internal layers across iterations to refine latent representations without explicit CoT supervision. While promising, these approaches often require costly pretraining and lack a principled framework for how reasoning should evolve across iterations.
We address this gap by introducing **Flow Chain-of-Thought (Flow CoT)**, a reasoning paradigm that models recursive inference as a progressive trajectory of latent cognitive states. Flow CoT frames each iteration as a distinct cognitive stage—deepening reasoning across iterations without relying on manual supervision. To realize this, we propose **SCOUT** (*Stepwise Cognitive Optimization Using Teachers*), a lightweight fine-tuning framework that enables Flow CoT-style reasoning without the need for pretraining. SCOUT uses progressive distillation to align each iteration with a teacher of appropriate capacity, and a cross-attention-based retrospective module that integrates outputs from previous iterations while preserving the model’s original computation flow.
Experiments across eight reasoning benchmarks show that SCOUT consistently improves both accuracy and explanation quality, achieving up to 1.8\% gains under fine-tuning. Qualitative analyses further reveal that SCOUT enables progressively deeper reasoning across iterations—refining both belief formation and explanation granularity. These results not only validate the effectiveness of SCOUT, but also demonstrate the practical viability of Flow CoT as a scalable framework for enhancing reasoning in LLMs. Guanghao Li 0003, Mingfeng Chen, Shuting Dong, Ming Tang 0006, Chun Yuan 0003 |
NeurIPS | 6 |
| 2024 | Deep Homography Estimation for Visual Place RecognitionabstractVisual place recognition (VPR) is a fundamental task for many applications such as robot localization and augmented reality. Recently, the hierarchical VPR methods have received considerable attention due to the trade-off between accuracy and efficiency. They usually first use global features to retrieve the candidate images, then verify the spatial consistency of matched local features for re-ranking. However, the latter typically relies on the RANSAC algorithm for fitting homography, which is time-consuming and non-differentiable. This makes existing methods compromise to train the network only in global feature extraction. Here, we propose a transformer-based deep homography estimation (DHE) network that takes the dense feature map extracted by a backbone network as input and fits homography for fast and learnable geometric verification. Moreover, we design a re-projection error of inliers loss to train the DHE network without additional homography labels, which can also be jointly trained with the backbone network to help it extract the features that are more suitable for local matching. Extensive experiments on benchmark datasets show that our method can outperform several state-of-the-art methods. And it is more than one order of magnitude faster than the mainstream hierarchical VPR methods using RANSAC. The code is released at https://github.com/Lu-Feng/DHE-VPR. Shuting Dong, Bingxi Liu 0001, Xiangyuan Lan, Dongmei Jiang, Chun Yuan 0003 |
AAAI | 2 |
| 2024 | Towards Seamless Adaptation of Pre-trained Models for Visual Place RecognitionabstractRecent studies show that vision models pre-trained in generic visual learning tasks with large-scale data can provide useful feature representations for a wide range of visual perception problems. However, few attempts have been made to exploit pre-trained foundation models in visual place recognition (VPR). Due to the inherent difference in training objectives and data between the tasks of model pre-training and VPR, how to bridge the gap and fully unleash the capability of pre-trained models for VPR is still a key issue to address. To this end, we propose a novel method to realize seamless adaptation of pre-trained models for VPR. Specifically, to obtain both global and local features that focus on salient landmarks for discriminating places, we design a hybrid adaptation method to achieve both global and local adaptation efficiently, in which only lightweight adapters are tuned without adjusting the pre-trained model. Besides, to guide effective adaptation, we propose a mutual nearest neighbor local feature loss, which ensures proper dense local features are produced for local matching and avoids time-consuming spatial verification in re-ranking. Experimental results show that our method outperforms the state-of-the-art methods with less training data and training time, and uses about only 3% retrieval runtime of the two-stage VPR methods with RANSAC-based spatial verification. It ranks 1st on the MSLS challenge leaderboard (at the time of submission). The code is released at https://github.com/Lu-Feng/SelaVPR. Xiangyuan Lan, Shuting Dong, Yaowei Wang 0001, Chun Yuan 0003 |
ICLR | 4 |
| 2024 | SuperVLAD: Compact and Robust Image Descriptors for Visual Place RecognitionabstractVisual place recognition (VPR) is an essential task for multiple applications such as augmented reality and robot localization. Over the past decade, mainstream methods in the VPR area have been to use feature representation based on global aggregation, as exemplified by NetVLAD. These features are suitable for large-scale VPR and robust against viewpoint changes. However, the VLAD-based aggregation methods usually learn a large number of (e.g., 64) clusters and their corresponding cluster centers, which directly leads to a high dimension of the yielded global features. More importantly, when there is a domain gap between the data in training and inference, the cluster centers determined on the training set are usually improper for inference, resulting in a performance drop. To this end, we first attempt to improve NetVLAD by removing the cluster center and setting only a small number of (e.g., only 4) clusters. The proposed method not only simplifies NetVLAD but also enhances the generalizability across different domains. We name this method SuperVLAD. In addition, by introducing ghost clusters that will not be retained in the final output, we further propose a very low-dimensional 1-Cluster VLAD descriptor, which has the same dimension as the output of GeM pooling but performs notably better. Experimental results suggest that, when paired with a transformer-based backbone, our SuperVLAD shows better domain generalization performance than NetVLAD with significantly fewer parameters. The proposed method also surpasses state-of-the-art methods with lower feature dimensions on several benchmark datasets. The code is available at https://github.com/lu-feng/SuperVLAD. Xinyao Zhang 0001, Canming Ye, Shuting Dong, Xiangyuan Lan, Chun Yuan 0003 |
NeurIPS | 4 |
| 2023 | Frequency Reciprocal Action and Fusion for Single Image Super-ResolutionabstractFrequency-based methods have recently received much attention due to their impressive restoration of detail and structure in single image super-resolution (SISR). However, most of these methods mainly use frequency information as auxiliary means but ignore exploring the correlations and pixel distribution differences among various frequencies. To address the limitations, we propose a novel Frequency Reciprocal Action and Fusion Network (FRAF) that explores various frequency correlations and differences. Specifically, we design a Frequency Reciprocal Action (FRA) module, which safely enhances valid spatial information and decreases un-necessary repetition by reciprocal action among various spatial frequencies, to generate refined high- and low-frequency features. These refined frequency features are then progressively to guide the details and structure recovery, respectively. Furthermore, we develop a Detail and Structure Fusion (DSF) module to adaptively select, enhance and fuse the features to output the final HR image. This way ensures the final image is a high-quality product with rich details and a clear structure. Experimental results demonstrate that our method achieves superior performance over state-of-the-art (SOTA) approaches on both quantitative and qualitative evaluations. Shuting Dong, Chun Yuan 0003 |
ICASSP | 1 |
| 2023 | AANet: Aggregation and Alignment Network with Semi-hard Positive Sample Mining for Hierarchical Place RecognitionabstractVisual place recognition (VPR) is one of the research hotspots in robotics, which uses visual information to locate robots. Recently, the hierarchical two-stage VPR methods have become popular in this field due to the trade-off between accuracy and efficiency. These methods retrieve the top-k candidate images using the global features in the first stage, then re-rank the candidates by matching the local features in the second stage. However, they usually require additional al-gorithms (e.g. RANSAC) for geometric consistency verification in re-ranking, which is time-consuming. Here we propose a Dynamically Aligning Local Features (DALF) algorithm to align the local features under spatial constraints. It is significantly more efficient than the methods that need geometric consistency verification. We present a unified network capable of extracting global features for retrieving candidates via an aggregation module and aligning local features for re-ranking via the DALF alignment module. We call this network AANet. Meanwhile, many works use the simplest positive samples in triplet for weakly supervised training, which limits the ability of the network to recognize harder positive pairs. To address this issue, we propose a Semi-hard Positive Sample Mining (ShPSM) strategy to select appropriate hard positive images for training more robust VPR networks. Extensive experiments on four benchmark VPR datasets show that the proposed AANet can outperform several state-of-the-art methods with less time consumption. The code is released at https://github.com/Lu-Feng/AANet. Shuting Dong, Baifan Chen, Chun Yuan 0003 |
ICRA | 3 |
| 2023 | DFVSR: Directional Frequency Video Super-Resolution via Asymmetric and Enhancement Alignment NetworkabstractRecently, techniques utilizing frequency-based methods have gained significant attention, as they exhibit exceptional restoration capabilities for detail and structure in video super-resolution tasks. However, most of these frequency-based methods mainly have three major limitations: 1) insufficient exploration of object motion information, 2) inadequate enhancement for high-fidelity regions, and 3) loss of spatial information during convolution. In this paper, we propose a novel network, Directional Frequency Video Super-Resolution (DFVSR), to address these limitations. Specifically, we reconsider object motion from a new perspective and propose Directional Frequency Representation (DFR), which not only borrows the property of frequency representation of detail and structure information but also contains the direction information of the object motion that is extremely significant in videos. Based on this representation, we propose a Directional Frequency-Enhanced Alignment (DFEA) to use double enhancements of task-related information for ensuring the retention of high-fidelity frequency regions to generate the high-quality alignment feature. Furthermore, we design a novel Asymmetrical U-shaped network architecture to progressively fuse these alignment features and output the final output. This architecture enables the intercommunication of the same level of resolution in the encoder and decoder to achieve the supplement of spatial information. Powered by the above designs, our method achieves superior performance over state-of-the-art models on both quantitative and qualitative evaluations. Shuting Dong, Zhe Wu 0006, Chun Yuan 0003 |
IJCAI | 1 |
| 2023 | Enhanced Image Deblurring: An Efficient Frequency Exploitation and Preservation NetworkabstractMost of these frequency-based deblurring methods mainly have two major limitations: (1) insufficient exploitation of frequency information, (2) inadequate preservation of frequency information. In this paper, we propose a novel Efficient Frequency Exploitation and Preservation Network (EFEP) to address these limitations. Firstly, we propose a novel Frequency-Balanced Exploitation Encoder (FBE-Encoder) to sufficiently exploit frequency information. We insert a novel Frequency-Balanced Navigator (FBN) module in the encoder, which establishes a dynamic balance that adaptively explores and integrates the correlations between frequency features and other features presented in the network. And it also can highlight the most important regions in frequency features. Secondly, considering the limitation that frequency information is inevitably lost in deep network architectures, we present an Enhanced Selective Frequency Decoder (ESF-Decoder) that not only effectively reduces spatial information redundancy, but also fully explores the different importance of various frequency information to ensure the supplement of valid spatial information and weaken the invalid information. Thirdly, each encoder/decoder block of the EFEP consists of multiple Contrastive Residual Blocks (CRBs), which are designed to explicitly compute and incorporate feature distinctions. Powered by the above designs, our EFEP outperforms state-of-the-art models on both quantitative and qualitative evaluations. Shuting Dong, Zhe Wu 0006, Chun Yuan 0003 |
ACM Multimedia | 1 |
| 2021 | A dynamic predictor selection algorithm for predicting stock market movement
Shuting Dong, Jianxin Wang 0001, Hongze Luo, Fang-Xiang Wu |
Expert Syst. Appl. | 1 |