VLDB 2026 Research / reviewers in the wild / expert
Kunyang Sun
dblp:256/0999
· DBLP profile ↗
12ranked-venue papers
1as first author
11since 2021 · last 2026
0009-0008-5294-7948ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal ModelingabstractVideo shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VMM) and a Dark-aware Semantic Block (DSB), extracting text-guided features to explicitly differentiate shadows from dark objects. Furthermore, we introduce adaptive mask reweighting to downweight penumbra regions during training and apply edge masks at the final decoder stage for better supervision. For temporal modeling of variable shadow shapes, we propose a Tokenized Temporal Block (TTB) that decouples spatiotemporal learning. TTB summarizes cross-frame shadow semantics into learnable temporal tokens, enabling efficient sequence encoding with minimal computation overhead. Comprehensive Experiments on multiple benchmark datasets demonstrate state-of-the-art accuracy and real-time inference efficiency. Kunyang Sun, Rui Yao 0006, Hancheng Zhu, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao, Yong Zhou 0003 |
AAAI | 2 |
| 2026 | CLIPDet3D: Vision-Language Collaborative Distillation for 3D Object DetectionabstractMulti-view 3D object detection plays a vital role in autonomous driving systems due to its ability to perceive complex scenes accurately. However, real-world driving data often exhibits a long-tailed distribution, causing significant drops in detection accuracy for rare categories in existing methods. To mitigate this issue, we propose CLIPDet3D, a novel vision-language collaborative framework for multi-view 3D object detection. First, to tackle the difficulty of capturing the semantic information of rare categories, a Vision-Language Collaborative Learning strategy is proposed to incorporate class-level semantic priors from CLIP. Second, a Depth Feature Contrastive Distillation module is designed to overcome the large depth estimation error for rare categories by aligning depth features between a teacher and a student network. Furthermore, to alleviate the difficulty in focusing on regions of rare categories, a Dual-Stream Prompt Attention mechanism is devised to inject learnable prompts and compute attention along both horizontal and vertical BEV directions. Evaluations on the nuScenes dataset demonstrate that CLIPDet3D achieves state-of-the-art accuracy while maintaining efficient inference. Jiaqi Zhao 0001, Huanfeng Hu, Yong Zhou 0003, Wen-Liang Du 0002, Kunyang Sun, Rui Yao 0006, Qigong Sun |
AAAI | 5 |
| 2026 | A unified multi-stream diffusion framework for robust video camouflaged object detection
Yuyao Ke, Rui Yao 0006, Kunyang Sun, Hancheng Zhu, Jiaqi Zhao 0001, Bing Liu 0016 |
Neural Networks | 3 |
| 2026 | Dynamic Prompt Memory Network for video shadow detection
Rui Yao 0006, Hancheng Zhu, Kunyang Sun, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik |
Pattern Recognit. | 4 |
| 2026 | SpaceFormer: Spatial Position Contextual Semantics Embedding for Multi-View 3D Object Detectionabstract3D object detection aims to accurately localize and recognize objects in 3D space. It serves as a fundamental task for reliable perception in intelligent transportation systems, enabling the monitoring of diverse traffic participants such as vehicles, pedestrians, cyclists, and public transport. Recently, transformer-based methods have gained significant attention in multi-view 3D object detection due to their strong global reasoning capabilities. However, their limited capacity to model spatial positional information hinders accurate object localization, especially in complex and large-scale scenes. To address this limitation, SpaceFormer is proposed as a novel transformer-based multi-view 3D object detector. Specifically, a Contextual Visual Prompts Learning strategy is proposed to enhance the perception of small and sparse traffic participants by incorporating contextual priors. To further suppress background interference, a Semantics-guided Depth Estimation method is proposed to refine depth representations using high-level semantic information. Furthermore, a Spatial Position Embedding mechanism is proposed to improve the spatial localization capability of the transformer by integrating geometric position and polar spatial embedding. Extensive experiments on the nuScenes benchmark demonstrate that SpaceFormer achieves state-of-the-art performance with 55.5% mAP and 62.9% NDS. These improvements indicate not only methodological advances but also practical benefits for intelligent transportation systems, enhancing safety, reliability, and efficiency in real-world deployments. Jiaqi Zhao 0001, Huanfeng Hu, Wen-Liang Du 0002, Yong Zhou 0003, Kunyang Sun, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T TrackingabstractTo reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label noise triggered by similar object noise can further affect the tracking performance. In this paper, we propose GDSTrack, a novel approach that introduces dynamic graph fusion and temporal diffusion to address the above challenges in self-supervised RGB-T tracking. GDSTrack dynamically fuses the modalities of neighboring frames, treats them as distractor noise, and leverages the denoising capability of a generative model. Specifically, by constructing an adjacency matrix via an Adjacency Matrix Generator (AMG), the proposed Modality-guided Dynamic Graph Fusion (MDGF) module uses a dynamic adjacency matrix to guide graph attention, focusing on and fusing the object’s coherent regions. Temporal Graph-Informed Diffusion (TGID) models MDGF features from neighboring frames as interference, and thus improving robustness against similar-object noise. Extensive experiments conducted on four public RGB-T tracking datasets demonstrate that GDSTrack outperforms the existing state-of-the-art methods. The source code is available at https://github.com/LiShenglana/GDSTrack. Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Kunyang Sun, Bing Liu 0016, Zhiwen Shao, Jiaqi Zhao 0001 |
IJCAI | 5 |
| 2025 | Progressively Generated Text-Assisted Image Aesthetic Quality AssessmentabstractImage Aesthetic Quality Assessment (IAQA) aims to simulate user perceptions to judge the aesthetic quality of images. Due to the high subjectivity of users and the complexity of image aesthetics, modeling IAQA solely at the image level is a compromise. Consequently, existing methods mainly focus on multimodal-based models and achieve effective performance. These methods explore aesthetic comments on images to characterize users and serve as auxiliary text information for multimodal modeling. Unfortunately, this may suffer from two limitations. One limitation is that aesthetic comments are often unavailable for an unknown image in the test phase, and another limitation is that the semantic information of these comments may be uncertain and fuzzy. Therefore, this paper proposes a progressively generated text-assisted image aesthetic quality assessment method, aiming to address the lack of aesthetic comments and the fuzziness of aesthetic judgments in these comments. Specifically, we first adopt a Multimodal Large Language Model (MLLM) to generate aesthetic comments on images by simulating user perceptions and utilize the generated comments to characterize their aesthetic perception to assist in the pre-training of our multimodal-based IAQA model. Then, we design an attribute prediction module to determine the attribute levels of aesthetic judgments and utilize text template construction to further generate explicit descriptions of image aesthetics. Finally, we leverage the generated attribute descriptions to further assist in training our IAQA model. By progressively generating textual auxiliary descriptions of aesthetics for images, the proposed model can gradually determine the aesthetic quality of the images. Massive experimental results indicate that the proposed method outperforms existing mainstream methods on multiple IAQA datasets. Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Kunyang Sun, Leida Li |
IEEE Trans. Fuzzy Syst. | 5 |
| 2023 | MATE: Masked Autoencoders are Online 3D Test-Time LearnersabstractOur MATE is the first Test-Time-Training (TTT) method designed for 3D data, which makes deep networks trained for point cloud classification robust to distribution shifts occurring in test data. Like existing TTT methods from the 2D image domain, MATE also leverages test data for adaptation. Its test-time objective is that of a Masked Autoencoder: a large portion of each test point cloud is removed before it is fed to the network, tasked with reconstructing the full point cloud. Once the network is updated, it is used to classify the point cloud. We test MATE on several 3D object classification datasets and show that it significantly improves robustness of deep networks to several types of corruptions commonly occurring in 3D point clouds. We show that MATE is very efficient in terms of the fraction of points it needs for the adaptation. It can effectively adapt given as few as 5% of tokens of each test sample, making it extremely lightweight. Our experiments show that MATE also achieves competitive performance by adapting sparsely on the test data, which further reduces its computational overhead, making it ideal for real-time applications. Muhammad Jehanzeb Mirza, Inkyu Shin, Wei Lin 0019, Andreas Schriebl, Kunyang Sun, Jaesung Choe, Mateusz Kozinski, Horst Possegger, In-So Kweon, Kuk-Jin Yoon, Horst Bischof |
ICCV | 5 |
| 2022 | CycDA: Unsupervised Cycle Domain Adaptation to Learn from Image to Video
Wei Lin 0019, Anna Kukleva, Kunyang Sun, Horst Possegger, Hilde Kuehne, Horst Bischof |
ECCV (3) | 3 |
| 2022 | LIMO: Latent Inceptionism for Targeted Molecule GenerationabstractGeneration of drug-like molecules with high binding affinity to target proteins remains a difficult and resource-intensive task in drug discovery. Existing approaches primarily employ reinforcement learning, Markov sampling, or deep generative models guided by Gaussian processes, which can be prohibitively slow when generating molecules with high binding affinity calculated by computationally-expensive physics-based methods. We present Latent Inceptionism on Molecules (LIMO), which significantly accelerates molecule generation with an inceptionism-like technique. LIMO employs a variational autoencoder-generated latent space and property prediction by two neural networks in sequence to enable faster gradient-based reverse-optimization of molecular properties. Comprehensive experiments show that LIMO performs competitively on benchmark tasks and markedly outperforms state-of-the-art techniques on the novel task of generating drug-like compounds with high binding affinity, reaching nanomolar range against two protein targets. We corroborate these docking-based results with more accurate molecular dynamics-based calculations of absolute binding free energy and show that one of our generated drug-like compounds has a predicted $K_D$ (a measure of binding affinity) of $6 \cdot 10^{-14}$ M against the human estrogen receptor, well beyond the affinities of typical early-stage drug candidates and most FDA-approved drugs to their respective targets. Code is available at https://github.com/Rose-STL-Lab/LIMO. Peter Eckmann, Kunyang Sun, Bo Zhao 0028, Mudong Feng, Michael K. Gilson, Rose Yu |
ICML | 2 |
| 2021 | ECS-Net: Improving Weakly Supervised Semantic Segmentation by Using Connections Between Class Activation MapsabstractImage-level weakly supervised semantic segmentation is a challenging task. As classification networks tend to capture notable object features and are insensitive to over-activation, class activation map (CAM) is too sparse and rough to guide segmentation network training. Inspired by the fact that erasing distinguishing features force networks to collect new ones from non-discriminative object regions, we using relationships between CAMs to propose a novel weakly supervised method. In this work, we apply these features, learned from erased images, as segmentation super-vision, driving network to study robust representation. In specifically, object regions obtained by CAM techniques are erased on images firstly. To provide other regions with seg-mentation supervision, Erased CAM Supervision Net (ECS-Net) generates pixel-level labels by predicting segmentation results of those processed images. We also design the rule of suppressing noise to select reliable labels. Our experiments on PASCAL VOC 2012 dataset show that without data annotations except for ground truth image-level labels, our ECS-Net achieves 67.6% mIoU on test set and 66.6% mIoU on val set, outperforming previous state-of-the-art methods. Kunyang Sun, Haoqing Shi, Zhengming Zhang 0001, Yongming Huang 0001 |
ICCV | 1 |
| 2020 | BlendMask: Top-Down Meets Bottom-Up for Instance SegmentationabstractInstance segmentation is one of the fundamental vision tasks. Recently, fully convolutional instance segmentation methods have drawn much attention as they are often simpler and more efficient than two-stage approaches like Mask R-CNN. To date, almost all such approaches fall behind the two-stage Mask R-CNN method in mask precision when models have similar computation complexity, leaving great room for improvement. In this work, we achieve improved mask prediction by effectively combining instance-level information with semantic information with lower-level fine-granularity. Our main contribution is a blender module which draws inspiration from both top-down and bottom-up instance segmentation approaches. The proposed BlendMask can effectively predict dense per-pixel position-sensitive instance features with very few channels, and learn attention maps for each instance with merely one convolution layer, thus being fast in inference. BlendMask can be easily incorporate with the state-of-the-art one-stage detection frameworks and outperforms Mask R-CNN under the same training schedule while being faster. A light-weight version of BlendMask achieves 36.0 mAP at 27 FPS evaluated on a single 1080Ti. Because of its simplicity and efficacy, we hope that our BlendMask could serve as a simple yet strong baseline for a wide range of instance-wise prediction tasks. Hao Chen 0041, Kunyang Sun, Zhi Tian, Chunhua Shen, Yongming Huang 0001, Youliang Yan |
CVPR | 2 |