EDBT 2026 Demo / reviewers in the wild / expert
Yinghui Sun
dblp:86/10304
· DBLP profile ↗
20ranked-venue papers
3as first author
20since 2021 · last 2027
0000-0003-1456-2859ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | MDB-DPL:Multi-scale deep block-Diagonal dictionary pair learning with feature adaptation for SAR target recognition
Yinghui Sun, Peng Fu 0003, Xizhan Gao, Quan-Sen Sun |
Signal Process. | 2 |
| 2026 | BiNAR: A Bi-Modal Framework for Non-Aligned RGB-IR 3D Reconstruction via Gaussian Splatting
Zhongwen Wang, Han Ling, Yinghui Sun, Quan-Sen Sun |
WACV | 4 |
| 2026 | Joint discriminative feature boosting and adaptive sample reweighting for multi-omics clustering
Yinghui Sun, Min-Ling Zhang, Quan-Sen Sun |
Pattern Recognit. | 1 |
| 2025 | TPCH: Tensor-interacted Projection and Cooperative Hashing for Multi-view ClusteringabstractIn recent years, anchor and hash-based multi-view clustering methods have gained attention for their efficiency and simplicity in handling large-scale data. However, existing methods often overlook the interactions among multi-view data and higher-order cooperative relationships during projection, negatively impacting the quality of hash representation in low-dimensional spaces, clustering performance, and sensitivity to noise. To address this issue, we propose a novel approach named Tensor-Interacted Projection and Cooperative Hashing for Multi-View Clustering(TPCH). TPCH stacks multiple projection matrices into a tensor, taking into account the synergies and communications during the projection process. By capturing higher-order multi-view information through dual projection and Hamming space, TPCH employs an enhanced tensor nuclear norm to learn more compact and distinguishable hash representations, promoting communication within and between views. Experimental results demonstrate that this refined method significantly outperforms state-of-the-art methods in clustering on five large-scale multi-view datasets. Moreover, in terms of CPU time, TPCH achieves substantial acceleration compared to the most advanced current methods. Zhongwen Wang, Xingfeng Li 0004, Yinghui Sun, Quan-Sen Sun, Yuan Sun 0016, Han Ling, Jian Dai 0002, Zhenwen Ren |
AAAI | 3 |
| 2025 | OCSplats: Observation Completeness Quantification and Label Noise Separation in 3DGSabstract3D Gaussian Splatting (3DGS) has become one of the most promising 3D reconstruction technologies. However, label noise in real-world scenarios-such as moving objects, non-Lambertian surfaces, and shadows-often leads to reconstruction errors. Existing 3DGS-Bsed anti-noise reconstruction methods either fail to separate noise effectively or require scene-specific fine-tuning of hyperparameters, making them difficult to apply in practice. This paper re-examines the problem of anti-noise reconstruction from the perspective of epistemic uncertainty, proposing a novel framework, OCSplats. By combining key technologies such as hybrid noise assessment and observation-based cognitive correction, the accuracy of noise classification in areas with cognitive differences has been significantly improved. Moreover, to address the issue of varying noise proportions in different scenarios, we have designed a label noise classification pipeline based on dynamic anchor points. This pipeline enables OCSplats to be applied simultaneously to scenarios with vastly different noise proportions without adjusting parameters. Extensive experiments demonstrate that OCSplats always achieve leading reconstruction performance and precise label noise classification in scenes of different complexity levels. Han Ling, Yinghui Sun, Quan-Sen Sun |
ICCV | 3 |
| 2025 | Incomplete Multi-View Clustering With Paired and Balanced Dynamic Anchor LearningabstractCompared to static anchor selection, existing dynamic anchor learning could automatically learn more flexible anchors to improve the performance of large-scale multi-view clustering. Despite improving the flexibility of anchors, these methods do not pay sufficient attention to the alignment and fairness of learned anchors. Specifically, within each cluster, the positions and quantities of cross-view anchors may not align, or even anchor absence in some clusters, leading to severe anchor misalignment and imbalance issues. These issues result in inaccurate graph fusion and a reduction in clustering performance. Besides, in practical applications, missing information caused by sensor malfunctions or data losses could further exacerbate anchor misalignment and imbalance. To overcome such challenges, a novel Incomplete Multi-view Clustering withPaired and Balanced Dynamic Anchor Learning (PBDAL)is proposed to ensure the alignment and fairness of anchors. Unlike existing unsupervised anchor learning, we first design a paired and balanced dynamic anchor learning scheme to supervise dynamic anchors to be aligned and fair in each cluster. Meanwhile, we develop an enhanced bipartite graph tensor learning to refine paired and balanced anchors. Our superiority, effectiveness, and efficiency are all validated by performing extensive experiments on multiple public datasets. Xingfeng Li 0004, Yuangang Pan, Yuan Sun 0016, Quan-Sen Sun, Yinghui Sun, Ivor W. Tsang, Zhenwen Ren |
IEEE Trans. Multim. | 5 |
| 2024 | ADFactory: An Effective Framework for Generalizing Optical Flow With NeRFabstractA significant challenge facing current optical flow meth-ods is the difficulty in generalizing them well to the real world. This is mainly due to the lack of large-scale real-world datasets, and existing self-supervised methods are limited by indirect loss and occlusions, resulting in fuzzy outcomes. To address this challenge, we introduce a novel optical flow training framework: automatic data factory (ADF). ADF only requires RGB images as input to effectively train the optical flow network on the target data do-main. Specifically, we use advanced NeRF technology to reconstruct scenes from photo groups collected by a monoc-ular camera, and then calculate optical flow labels between camera pose pairs based on the rendering results. To elimi-nate erroneous labels caused by defects in the scene reconstructed by NeRF, we screened the generated labels from multiple aspects, such as optical flow matching accuracy, radiation field confidence, and depth consistency. The fil-tered labels can be directly used for network supervision. Experimentally, the generalization ability of ADF on KITTI surpasses existing self-supervised optical flow and monoc-ular scene flow algorithms. In addition, ADF achieves impressive results in real-world zero-point generalization evaluations and surpasses most supervised methods11Code: https://github.com/HanLingsgjk/UnifiedGeneralization. Han Ling, Quan-Sen Sun, Yinghui Sun, Xingfeng Li 0004 |
CVPR | 3 |
| 2024 | Fast Unpaired Multi-view Clustering
Xingfeng Li 0004, Yuangang Pan, Yinghui Sun, Quan-Sen Sun, Ivor W. Tsang, Zhenwen Ren |
IJCAI | 3 |
| 2024 | Interactive Segmentation by Considering First-Click Intentional AmbiguityabstractInteractive segmentation task (IS) aims at taking into account the influence of user preferences on the basis of general semantic segmentation in order to obtain the specific target-of-interest. Given the fact that most of the related algorithms generate a single mask only, the robustness of which might be constrained due to the diversity of user intention in the early interaction stage, namely the vague selection of object part/whole object/adherent object, especially when there's only one click. To handle this, we propose a novel framework called Diversified Interactive Segmentation Network (DISNet) in which we revisit the peculiarity of first-click: given an input image, DISNet outputs multiple candidate masks under the guidance of first-click only, then a Dual-attentional Mask Correction (DAMC) module is utilized to measure the complex mutual effect within first-click, all-clicks and image features. Moreover, we design a new sampling strategy to generate GT masks with rich semantic relations. Performance analysis plus adequate ablation studies has demonstrated the efficacy of our methods, which further exemplifies the decisive role of first-click in the realm of IS. Kangpeng Hu, Quan-Sen Sun, Yinghui Sun, Tao Wang 0020 |
ACM Multimedia | 3 |
| 2024 | Improved Weighted Tensor Schatten p-Norm for Fast Multi-view Graph ClusteringabstractRecently, tensor Schatten p-norm has achieved impressive performance for fast multi-view clustering [57]. This primarily ascribes the superiority of tensor Schatten p-norm in exploring high-order structure information among views. Whereas, 1) tensor Schatten p-norm treats different singular values equally, such that the larger singular values corresponding to certain significant feature information (i.e., prior information) have not been utilized fully; 2) tensor Schatten p-norm also ignore ranking the core entries of core tensor, which may contain noise information; 3) existing methods select fixed anchors or averagely update anchors to construct the neighbor bipartite graphs, greatly limiting the flexibility and expression of anchors. To break these limitations, we propose a novel Improved Weighted Tensor Schatten p-Norm for Fast Multi-view Graph Clustering (IWTSN-FMGC). Specifically, to eliminate the interference of the first two limitations, we propose an improved weighted tensor Schatten p-norm to dynamically rank core tensor and automatically shrink singular values. To this end, improved weighted tensor Schatten p-norm has the potential to more effectively leverage low-rank structures and prior information, thereby enhancing robustness compared to current tensor Schatten p-norm methods. Further, the designed adaptive neighbor bipartite graph learning can more flexibly and expressively encode the local manifold structure information than existing anchor selection and averaged anchor updating. Extensive experiments validate our effectiveness and superiority across multiple benchmark datasets. Yinghui Sun, Xingfeng Li 0004, Quan-Sen Sun, Min-Ling Zhang, Zhenwen Ren |
ACM Multimedia | 1 |
| 2024 | Enforced Block Diagonal Graph Learning for Multikernel ClusteringabstractThe existing multikernel graph clustering (MKGC) methods have emerged with notable success on nonlinear clustering tasks since the graph learning can effectively capture graph structure similarity between sample pair. Ideally, a high-quality graph should enjoy the good block diagonal property, i.e. the intercluster similarities correspond to zeros, while intracluster similarities represent nonzeros. Meanwhile, the number of diagonal blocks for a graph equals to the number of clusters on the dataset. However, most of the existing MKGC methods design a corpulent three-parts graph leaning process that poses challenges for hyperparameter tuning, time cost, and clustering performance. To overcome these challenging issues, we propose an enforced block diagonal graph learning for multikernel clustering (EBDGL-MKC) method, where we pursue a high-quality block diagonal graph via well-designed one-part graph leaning scheme rather than three parts. Inspired by symmetric matrix factorization (SMF), we first design a one-part block diagonal graph learning scheme to learn multiple block diagonal graphs, by exploring an explicit theoretical connection between the clustering partition of kernel$k$-means and the excellent block diagonal graph. Then, these block diagonal graphs are stacked into a low-rank tensor for exploiting the high-order structure information hidden in the nonlinear data. After that, an effective alternate algorithm with convergence proof is performed on extensive experiments to demonstrate the superiority of EBDGL compared with the state-of-the-art multikernel clustering (MKC) methods. Xingfeng Li 0004, Yinghui Sun, Quan-Sen Sun, Zhenwen Ren |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Query-Guided Prototype Optimization for Few-Shot ClassificationabstractWith limited labeled samples, few-shot classification poses a challenge to standard deep models and has attracted a surge of concern. Metric learning based approaches stand out for the minimalist and efficient design, aiming to classify query samples using supervision of support sets in the metric space. Prototypical Network has pioneered the use of mean feature embeddings to represent each support class and leaned on the computed prototypes for query classification. However, inherent bias exists in the mean prototypes generating from scarce support samples versus the actual class prototypes, which induces subsequent inference deviation. In this paper, we propose to diminish the bias by leveraging the semantic information of query samples to guide prototype optimization. Specifically, we exploit the semantic correlation between the local of initial mean prototypes and the global of query samples to generate query-guided masks, thus tailoring optimized prototypes that vary by query samples. This exploration of correlation is first utilized to alleviate the prototype bias problem and shows great brevity compared to existing methods. Extensive experiments are conducted on three few-shot image classification benchmark datasets, and demonstrate the effectiveness of our proposed method. Yinghui Sun, Xiaobo Shen 0001, Quan-Sen Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Unit Correlation With Interactive Feature for Robust and Effective TrackingabstractFor robust and effective tracking, most efforts strive to design a powerful representation target model, while we are inspired by the idea of“knowing oneself and knowing others”to major in both the target and non-target features. In this work, we propose a unit correlation with interactive feature tracker (UCIF), which utilizes feature interaction and independent correlation operation to improve robustness and effectiveness. Specifically, we first propose a feature integration network, in which the feature enhancement module concentrates on enhancing the tracker's representation ability for both target and non-target. The feature interaction module is in charge of completing the interactive learning between target and non-target features. Then, considering the potential risk of blurring spatial information in regular correlation operation, a unit correlation network is presented, where the convolution sampling strategy can integrate the target features as well as reduce the computation costs. The unit kernel for correlation operation can protect the target spatial information. The channel ranking module suppresses background interference via weight assignment. Extensive experiments are conducted on both the short-term and long-term challenging benchmarks, including OTB2015, NFS, UAV123, TrackingNet, GOT-10 k, TLP, LaSOT and VOT-LT2019. Our tracker achieves remarkable performance in robustness and effectiveness. Ze Zhou 0002, Yinghui Sun, Quan-Sen Sun, Zhenwen Ren |
IEEE Trans. Multim. | 2 |
| 2023 | Learning Optical Expansion from Scale MatchingabstractThis paper address the problem of optical expansion (OE). OE describes the object scale change between two frames, widely used in monocular 3D vision tasks. Previous methods estimate optical expansion mainly from optical flow results, but this two-stage architecture makes their results limited by the accuracy of optical flow and less robust. To solve these problems, we propose the concept of 3D optical flow by integrating optical expansion into the 2D optical flow, which is implemented by a plug-and-play module, namely TPCV. TPCV implements matching features at the correct location and scale, thus allowing the simultaneous optimization of optical flow and optical expansion tasks. Experimentally, we apply TPCV to the RAFT optical flow baseline. Experimental results show that the baseline optical flow performance is substantially improved. Moreover, we apply the optical flow and optical expansion results to various dynamic 3D vision tasks, including motion-in-depth, time-to-collision, and scene flow, often achieving significant improvement over the prior SOTA. Code is available at https://github.com/HanLingsgjk/TPCV. Han Ling, Yinghui Sun, Quan-Sen Sun, Zhenwen Ren |
CVPR | 2 |
| 2023 | Distribution Consistency based Fast Anchor Imputation for Incomplete Multi-view ClusteringabstractIn practical scenarios, partial missing of multi-view data is very common, such as register information missing from social network analysis, which results in incomplete multi-view clustering (IMVC). How to fill missing data fast and efficiently plays a vital role in improving IMVC, carrying a significant challenge. Existing IMVC methods always use all observed data to fill in missing data, resulting in high complexity and poor imputation quality due to a lack of guidance from consistent distribution. To break the existing limitations, we propose a novel Distribution Consistency based Fast Anchor Imputation for Incomplete Multi-view Clustering (DCFAI-IMVC) method. Specifically, to eliminate the interference of redundant and fraudulent features in the original space, incomplete data are first projected into a consensus latent space, where we dynamically learn a small number of anchors to achieve fast and good imputation. Then, we employ global distribution information of the observed embedding representations to further ensure the consistent distribution between the learned anchors and the observed embedding representations. Ultimately, a tensor low-rank constraint is imposed on bipartite graphs to investigate the high-order correlations hidden in data. DCFAI-IMVC enjoys linear complexity in terms of sample number, which gives it great potential to handle large-scale IMVC tasks. By performing extensive experiments, our effectiveness, superiority, and efficiency are all validated on multiple public datasets with recent advances. Xingfeng Li 0004, Yinghui Sun, Quan-Sen Sun, Jia Dai, Zhenwen Ren |
ACM Multimedia | 2 |
| 2023 | Two-directional two-dimensional fractional-order embedding canonical correlation analysis for multi-view dimensionality reduction and set-based video recognition
Yinghui Sun, Xizhan Gao, Sijie Niu, Dong Wei 0007, Zhen Cui 0001 |
Expert Syst. Appl. | 1 |
| 2023 | Attacking the tracker with a universal and attractive patch as fake targetabstractAdversarial attacks in visual object tracking aim to drop tracking performance through injecting imperceptible perturbations to the input of the tracker. Current methods usually superimpose perturbation maps on the input images, and advocate blinding the tracker via occluding the real targets to achieve attack effect. From the perspective of attraction, we alternatively propose a novel idea of attacking the tracker, which advocates using perturbation patches to act as fake targets to attract the tracker's attention. For this purpose, we establish a multi-conditional objective function to generate our ideal patch in an offline iterative manner. For invisibility, we integrate the constraint of patch value into the function for unified optimization. For universality, in addition to adopting large-scale and high-diversity training samples , we also incorporate the video-agnostic condition into this function. To make the patch attractive like a fake target, we elaborately design the non-overlapping area to determine the patch position, and generate matched fake labels to mislead the tracker to track the patch. In online attacking, it only needs to paste the optimized patch onto the video frames, the tracker will be successfully attracted by our patch, achieving attack effect. Extensive experimental results on 8 popular tracking datasets demonstrate that our method can obtain exceptional attack performance in both non-targeted and targeted attack. Additionally, the experiments on transferability illustrate our optimized patches can be directly applied to other trackers with different architectures. Ze Zhou 0002, Yinghui Sun, Quan-Sen Sun, Zhenwen Ren |
Inf. Sci. | 2 |
| 2023 | Consensus Cluster Center Guided Latent Multi-Kernel ClusteringabstractExisting multi-kernel clustering (MKC) methods usually focus on constructing a fixed dimension consensus-partition from base kernels to demonstrate their superior in integrating complementary information. Despite their success, they still suffer from the following limitations: (1) The size of consensus-partition is always fixed as the upper bound ($k=n$) or lower bound ($k=c$), where$n$,$c$, and$k$are the number of samples, clusters, and partition dimension, respectively, resulting in suboptimal partition;(2)The learned consensus-partition cannot make full use of the global distribution information hidden in data. To address these issues, we propose a latent consensus-partition learning framework for MKC, namelyConsensus Cluster Center Guided Latent Multi-kernel Clustering(C3LMC), including two methods,i.e., C3LMCKand C3LMCH. For C3LMCK, we flexibly search for a more proper dimension of consensus-partition in a latent embedding space rather than the fixed partition dimension. Meanwhile, the generation of latent consensus-partition is guided by a consensus cluster center of base kernels, such that global distribution information hidden in base kernels can be captured fully. However, C3LMCKsuffers from$\boldsymbol {\mathcal {O}}(n^{2})$computational complexity and memory complexity. Thus, we further propose C3LMCHto handle large-scale data by reducing both kinds of complexities to$\boldsymbol {\mathcal {O}}(n)$. Two solvers with convergence proof are developed to validate our effectiveness, superiority, and efficiency on multiple public datasets with the recent advances. Xingfeng Li 0004, Yinghui Sun, Quan-Sen Sun, Zhenwen Ren |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Only Once Attack: Fooling the Tracker With Adversarial TemplateabstractAdversarial attacks in visual object tracking aims to fool trackers via injecting invisible perturbations for the video frames. Most adversarial methods advocate generating perturbations for each video frame, but frequent attacks may increase the computational load and the risk of exposure. Unfortunately, less works are about only attacking the initial frame and their attack effects are insufficient. To tackle this, we focus on the initialization phase of tracking and propose an only once attack framework. It can effectively fool the tracker via only generating invisible perturbations for the initial template, rather than each frame. Specifically, considering the tracking mechanism of the Siamese-based trackers, we design the minimum score-based and the minimum IoU-based loss functions. Both of them are used for training the UNet-based perturbation generator instead of the tracker, achieving the non-targeted attack. Additionally, we propose the location and direction offsets as the base attacks of sophisticated targeted attack. Combined with the two basic attacks, the tracker can be easily hijacked to move towards the fake target predefined by users. Extensive experimental results demonstrate that our only once attack framework costs the least number of attacks yet achieves better attack effect, with the maximum drop of 68.7%. The transferability experiments illustrate that our attack framework with good generalization ability can be directly applicable to the CNN-based, Siamese-based, deep discriminative-based and Transformer-based trackers, without retraining. Ze Zhou 0002, Yinghui Sun, Quan-Sen Sun, Zhenwen Ren |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Dynamic Incomplete Multi-view Imputing and ClusteringabstractIncomplete multi-view clustering (IMVC) is deemed a significant research topic in multimedia to handle data loss situations. Current late fusion incomplete multi-view clustering methods have attracted intensive attention owing to their superiority in using consensus partition for effective and efficient imputation and clustering. However, 1) their imputation quality and clustering performance depend heavily on the static prior partition, such as predefined zeros filling, destroying the diversity of different views; 2) the size of base partitions is too small, which would lose advantageous details of base kernels to decrease clustering performance. To address these issues, we propose a novel IMVC method, named Dynamic Incomplete Multi-view Imputing and Clustering (DIMIC). Concretely, the observed views dynamically generate a consensus proxy with the guidance of a shared cluster matrix for more effective imputation and clustering, rather than a fixed predefined partition matrix. Furthermore, the proper size of base partitions is employed to protect sufficient kernel details for further enhancing the quality of the consensus proxy. By designing a solver with a linear computational and memory complexity on extensive experiments, our effectiveness, superiority, and efficiency are validated on multiple public datasets with recent advances. Xingfeng Li 0004, Quan-Sen Sun, Zhenwen Ren, Yinghui Sun |
ACM Multimedia | 4 |