EDBT 2026 Demo / reviewers in the wild / expert
Siran Dai
dblp:360/0801
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0005-9214-1883ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Representation and self-supervised learning · 30% Video understanding and tracking · 21% Generative modeling · 16% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 62% Data mining · 38% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
contrastive learning |
1.8 | 2 | 2026 | Semantic Concentration for Self-Supervised Dense Representations Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026 Regularized Contrastive Partial Multi-view Outlier Detection · ACM Multimedia 2024 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
dense representation learning |
1.0 | 1 | 2026 | Semantic Concentration for Self-Supervised Dense Representations Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Generative modeling › diffusion model
diffusion-based data augmentation |
1.0 | 1 | 2026 | HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models · AAAI 2026 |
Machine learning › Generative modeling
diffusion model |
1.0 | 1 | 2026 | HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models · AAAI 2026 |
Computer vision › Image recognition and object detection › image classification
fine-grained image classification |
1.0 | 1 | 2026 | HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models · AAAI 2026 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked video modeling |
0.9 | 1 | 2025 | When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning · CVPR 2025 |
Computer vision › Video understanding and tracking › video representation learning
self-supervised video representation learning |
0.9 | 1 | 2025 | When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning · CVPR 2025 |
Computer vision › Video understanding and tracking
temporal alignment |
0.9 | 1 | 2025 | When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning · CVPR 2025 |
Computer vision › Video understanding and tracking
video representation learning |
0.9 | 1 | 2025 | When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning · CVPR 2025 |
Data mining
anomaly detection |
0.8 | 1 | 2024 | Regularized Contrastive Partial Multi-view Outlier Detection · ACM Multimedia 2024 |
Data mining › anomaly detection › outlier detection
multi-view outlier detection |
0.8 | 1 | 2024 | Regularized Contrastive Partial Multi-view Outlier Detection · ACM Multimedia 2024 |
Information retrieval
retrieval models |
0.8 | 1 | 2024 | Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval · ACM Multimedia 2024 |
Information retrieval
similarity measure |
0.8 | 1 | 2024 | Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval · ACM Multimedia 2024 |
Information retrieval › multimedia analysis and retrieval
video retrieval |
0.8 | 1 | 2024 | Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval · ACM Multimedia 2024 |
Machine learning › Learning theory › ranking
AUC optimization |
0.7 | 1 | 2023 | DRAUC: An Instance-wise Distributionally Robust AUC Optimization Framework · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › robustness
distributionally robust optimization |
0.7 | 1 | 2023 | DRAUC: An Instance-wise Distributionally Robust AUC Optimization Framework · NeurIPS 2023 |
Machine learning › Learning theory
generalization bounds |
0.7 | 1 | 2023 | DRAUC: An Instance-wise Distributionally Robust AUC Optimization Framework · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | DRAUC: An Instance-wise Distributionally Robust AUC Optimization Framework · NeurIPS 2023 |
Visual content generation and editing
image generation |
0.3 | 1 | 2026 | HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models · AAAI 2026 |
Information retrieval › evaluation › effectiveness metrics
average precision |
0.2 | 1 | 2024 | Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval · ACM Multimedia 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.2 | 1 | 2023 | DRAUC: An Instance-wise Distributionally Robust AUC Optimization Framework · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › dataset bias
label bias |
0.2 | 1 | 2023 | DRAUC: An Instance-wise Distributionally Robust AUC Optimization Framework · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
classifier-free guidance · 2.0classifier guidance · 2.0neighbor alignment · 1.5cross-view relation transfer · 1.5contrastive loss · 1.5object-aware filter · 1.0noise-tolerant ranking loss · 1.0cross-attention · 1.0self-distillation · 0.9masked autoencoder · 0.9spreading regularization · 0.8pairwise loss · 0.8hierarchical learning · 0.8QuadLinear-AP loss · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion ModelsabstractGenerative diffusion models show promise for data augmentation. However, applying them to fine-grained tasks presents a significant challenge: ensuring synthetic images accurately capture the subtle, category-defining features critical for high fidelity. Standard approaches, such as text-based Classifier-Free Guidance (CFG), often lack the required specificity, potentially generating misleading examples that degrade fine-grained classifier performance. To address this, we propose Hierarchically Guided Fine-grained Augmentation (HiGFA). HiGFA leverages the temporal dynamics of the diffusion sampling process. It employs strong text and transformed contour guidance with fixed strengths in the early-to-mid sampling stages to establish overall scene, style, and structure. In the final sampling stages, HiGFA activates a specialized fine-grained classifier guidance and dynamically modulates the strength of all guidance signals based on prediction confidence. This hierarchical, confidence-driven orchestration enables HiGFA to generate diverse yet faithful synthetic images by intelligently balancing global structure formation with precise detail refinement. Experiments on several FGVC datasets demonstrate the effectiveness of HiGFA. Zhiguang Lu, Qianqian Xu 0001, Peisong Wen, Siran Dai, Qingming Huang |
AAAI | 4 |
| 2026 | Semantic Concentration for Self-Supervised Dense Representations LearningabstractRecent advances in image-level self-supervised learning (SSL) have made significant progress, yet learning dense representations for patches remains challenging. Mainstream methods encounter an over-dispersion phenomenon that patches from the same instance/category scatter, harming downstream performance on dense tasks. This work reveals that image-level SSL avoids over-dispersion by involving implicit semantic concentration. Specifically, the non-strict spatial alignment ensures intra-instance consistency, while shared patterns, i.e., similar parts of within-class instances in the input space, ensure inter-image consistency. Unfortunately, these approaches are infeasible for dense SSL due to their spatial sensitivity and complicated scene-centric data. These observations motivate us to explore explicit semantic concentration for dense SSL. First, to break the strict spatial alignment, we propose to distill the patch correspondences. Facing noisy and imbalanced pseudo labels, we propose a noise-tolerant ranking loss. The core idea is extending the Average Precision (AP) loss to continuous targets, such that its decision-agnostic and adaptive focusing properties prevent the student model from being misled. Second, to discriminate the shared patterns from complicated scenes, we propose the object-aware filter to map the output space to an object-based space. Specifically, patches are represented by learnable prototypes of objects via cross-attention. Last but not least, empirical studies across various tasks soundly support the effectiveness of our method. Peisong Wen, Qianqian Xu 0001, Siran Dai, Runmin Cong, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation LearningabstractThe past decade has witnessed notable achievements in self-supervised learning for video tasks. Recent efforts typically adopt the Masked Video Modeling (MVM) paradigm, leading to significant progress on multiple video tasks. However, two critical challenges remain: 1) Without human annotations, the random temporal sampling introduces uncertainty, increasing the difficulty of model training. 2) Previous MVM methods primarily recover the masked patches in the pixel space, leading to insufficient information compression for downstream tasks. To address these challenges jointly, we propose a self-supervised framework that leverages Temporal Correspondence for video Representation learning (T-CoRe). For challenge 1), we propose a sandwich sampling strategy that selects two auxiliary frames to reduce reconstruction uncertainty in a two-side-squeezing manner. Addressing challenge 2), we introduce an auxiliary branch into a self-distillation architecture to restore representations in the latent space, generating high-level semantic representations enriched with temporal information. Experiments of T-CoRe consistently present superior performance across several downstream tasks, demonstrating its effectiveness for video representation learning. The code is available at https://github.com/yafeng19/T-CORE. Yang Liu 0350, Qianqian Xu 0001, Peisong Wen, Siran Dai, Qingming Huang |
CVPR | 4 |
| 2025 | Exploring Structural Degradation in Dense Representations for Self-supervised LearningabstractIn this work, we observe a counterintuitive phenomenon in self-supervised learning (SSL): longer training may impair the performance of dense prediction tasks (e.g., semantic segmentation). We refer to this phenomenon as Self-supervised Dense Degradation (SDD) and demonstrate its consistent presence across sixteen state-of-the-art SSL methods with various losses, architectures, and datasets. When the model performs suboptimally on dense tasks at the end of training, measuring the performance during training becomes essential. However, evaluating dense performance effectively without annotations remains an open challenge.
To tackle this issue, we introduce a Dense representation Structure Estimator (DSE), composed of a class-relevance measure and an effective dimensionality measure. The proposed DSE is both theoretically grounded and empirically validated to be closely correlated with the downstream performance. Based on this metric, we introduce a straightforward yet effective model selection strategy and a DSE-based regularization method. Experiments on sixteen SSL methods across four benchmarks confirm that model selection improves mIoU by $3.0\\%$ on average with negligible computational cost. Additionally, DSE regularization consistently mitigates the effects of dense degradation. Code is available at \url{https://github.com/EldercatSAM/SSL-Degradation}. Siran Dai, Qianqian Xu 0001, Peisong Wen, Yang Liu 0350, Qingming Huang |
NeurIPS | 1 |
| 2024 | Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video RetrievalabstractThe rapid growth of online video resources has significantly promoted the development of video retrieval methods. As a standard evaluation metric for video retrieval, Average Precision (AP) assesses the overall rankings of relevant videos at the top list, making the predicted scores a reliable reference for the users. However, recent video retrieval methods utilize pair-wise losses that treat all sample pairs equally, leading to an evident gap between the training objective and evaluation metric. To effectively bridge this gap, in this work, we aim to address two primary challenges: a) The current similarity measure and AP-based loss are suboptimal for video retrieval; b) The noticeable noise from frame-to-frame matching introduces ambiguity in estimating the AP loss. In response to these challenges, we propose the Hierarchical learning framework for Average-Precision-oriented Video Retrieval (HAP-VR). For the former challenge, we develop the TopK-Chamfer Similarity and QuadLinear-AP loss to measure and optimize video-level similarities in terms of AP. For the latter challenge, we suggest constraining the frame-level similarities to achieve an accurate AP loss estimation. Experimental results present that HAP-VR outperforms existing methods on several benchmark datasets, providing a feasible solution for video retrieval tasks and thus offering potential benefits for the multi-media application. Yang Liu 0350, Qianqian Xu 0001, Peisong Wen, Siran Dai, Qingming Huang |
ACM Multimedia | 4 |
| 2024 | Regularized Contrastive Partial Multi-view Outlier DetectionabstractIn recent years, multi-view outlier detection (MVOD) methods have advanced significantly, aiming to identify outliers within multi-view datasets. A key point is to better detect class outliers and class-attribute outliers, which only exist in multi-view data. However, existing methods either is not able to reduce the impact of outliers when learning view-consistent information, or struggle in cases with varying neighborhood structures. Moreover, most of them do not apply to partial multi-view data in real-world scenarios. To overcome these drawbacks, we propose a novel method named Regularized Contrastive Partial Multi-view Outlier Detection (RCPMOD). In this framework, we utilize contrastive learning to learn view-consistent information and distinguish outliers by the degree of consistency. Specifically, we propose (1) An outlier-aware contrastive loss with a potential outlier memory bank to eliminate their bias motivated by a theoretical analysis. (2) A neighbor alignment contrastive loss to capture the view-shared local structural correlation. (3) A spreading regularization loss to prevent the model from overfitting over outliers. With the Cross-view Relation Transfer technique, we could easily impute the missing view samples based on the features of neighbors. Experimental results on four benchmark datasets demonstrate that our proposed approach could outperform state-of-the-art competitors under different settings. Qianqian Xu 0001, Yangbangyan Jiang, Siran Dai, Qingming Huang |
ACM Multimedia | 4 |
| 2023 | DRAUC: An Instance-wise Distributionally Robust AUC Optimization FrameworkabstractThe Area Under the ROC Curve (AUC) is a widely employed metric in long-tailed classification scenarios. Nevertheless, most existing methods primarily assume that training and testing examples are drawn i.i.d. from the same distribution, which is often unachievable in practice. Distributionally Robust Optimization (DRO) enhances model performance by optimizing it for the local worst-case scenario, but directly integrating AUC optimization with DRO results in an intractable optimization problem. To tackle this challenge, methodically we propose an instance-wise surrogate loss of Distributionally Robust AUC (DRAUC) and build our optimization framework on top of it. Moreover, we highlight that conventional DRAUC may induce label bias, hence introducing distribution-aware DRAUC as a more suitable metric for robust AUC learning. Theoretically, we affirm that the generalization gap between the training loss and testing error diminishes if the training set is sufficiently large. Empirically, experiments on corrupted benchmark datasets demonstrate the effectiveness of our proposed method. Code is available at: https://github.com/EldercatSAM/DRAUC. Siran Dai, Qianqian Xu 0001, Zhiyong Yang 0001, Xiaochun Cao, Qingming Huang |
NeurIPS | 1 |