EDBT 2026 Demo / reviewers in the wild / expert
Yeyu Chai
dblp:386/3618
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
3D vision · 50% Representation and self-supervised learning · 27% Vision and language · 23% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 100% | |
| Computer networks
1 paper |
Physical-layer communications · 100% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Multimedia analysis and retrieval
cross-modal retrieval |
2.0 | 2 | 2026 | Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network for Multi-Grained Cross-Modal Retrieval · IEEE Trans. Image Process. 2026 Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D Retrieval · AAAI 2026 |
Computer vision › 3D vision › geometric deep learning
3d representation learning |
1.0 | 1 | 2026 | Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D Retrieval · AAAI 2026 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning |
1.0 | 1 | 2026 | Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D Retrieval · AAAI 2026 |
Physical-layer communications › coding theory
joint source-channel coding |
1.0 | 1 | 2026 | Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network for Multi-Grained Cross-Modal Retrieval · IEEE Trans. Image Process. 2026 |
Computer vision › 3D vision › point cloud analysis › point cloud learning
point cloud representation learning |
0.9 | 1 | 2025 | Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
semantic decoupling · 2.0probabilistic mapping network · 2.0multi-grained alignment · 2.0hyperbolic geometry · 2.0hierarchical ordering loss · 2.0contrastive learning · 2.0riemannian attention · 0.9multi-scale attention · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D RetrievalabstractWith the daily influx of 3D data on the internet, text-3D retrieval has gained increasing attention. However, current methods face two major challenges: Hierarchy Representation Collapse (HRC) and Redundancy-Induced Saliency Dilution (RISD). HRC compresses abstract-to-specific and whole-to-part hierarchies in Euclidean embeddings, while RISD averages noisy fragments, obscuring critical semantic cues and diminishing the model’s ability to distinguish hard negatives. To address these challenges, we introduce the Hyperbolic Hierarchical Alignment Reasoning Network (H2ARN) for text-3D retrieval. H2ARN embeds both text and 3D data in a Lorentz-model hyperbolic space, where exponential volume growth inherently preserves hierarchical distances. A hierarchical ordering loss constructs a shrinking entailment cone around each text vector, ensuring that the matched 3D instance falls within the cone, while an instance-level contrastive loss jointly enforces separation from non-matching samples. To tackle RISD, we propose a contribution-aware hyperbolic aggregation module that leverages Lorentzian distance to assess the relevance of each local feature and applies contribution-weighted aggregation guided by hyperbolic geometry, enhancing discriminative regions while suppressing redundancy without additional supervision. We also release the expanded T3DR-HIT v2 benchmark, which contains 8,935 text-to-3D pairs, 2.6 times the original size, covering both fine-grained cultural artefacts and complex indoor scenes. Wenrui Li 0001, Yidan Lu, Yeyu Chai, Rui Zhao 0010, Hengyu Man, Xiaopeng Fan 0001 |
AAAI | 3 |
| 2026 | Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network for Multi-Grained Cross-Modal RetrievalabstractCross-modal retrieval is essential for exploring semantic correlations between multimodal data. However, existing approaches face challenges in resolving semantic ambiguity and transferring knowledge with sparse sample generalization. To address these challenges, we propose a new Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network (SKPMN). Specifically, the Semantic Decoupling and Distinction (SDD) module decomposes complex word-region relationships into relevance-driven representations. The Deep Probability Mapping (DPM) module introduces a paradigm shift by mapping multimodal features into probabilistic distributions, capturing the semantic similarities and the potential uncertainties that define sparse or ambiguous relationships. By combining the Attention Probabilistic Mapping (APM) module, the model can effectively transfer knowledge across similar samples while emphasizing critical distinctions, significantly enhancing generalization to sparse and ambiguous samples. Finally, the multi-grained alignment strategy establishes a novel integration of fine-grained patch-to-word alignment and coarse-grained global alignment. Experimental results show that SKPMN achieves superior retrieval accuracy across major benchmark datasets. Furthermore, we implement a channel resource allocation technique that allocates more transmission resources to semantically significant information. In resource-constrained environments, our approach leverages Joint Source-Channel Coding (JSCC) to enhance the efficiency of visual feature transmission. Wenrui Li 0001, Yeyu Chai, Liang-Jian Deng, Ruiqin Xiong, Xiaopeng Fan 0001, Yonghong Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Riemann-based Multi-scale Attention Reasoning Network for Text-3D RetrievalabstractDue to the challenges in acquiring paired Text-3D data and the inherent irregularity of 3D data structures, combined representation learning of 3D point clouds and text remains unexplored. In this paper, we propose a novel Riemann-based Multi-scale Attention Reasoning Network (RMARN) for text-3D retrieval. Specifically, the extracted text and point cloud features are refined by their respective Adaptive Feature Refiner (AFR). Furthermore, we introduce the innovative Riemann Local Similarity (RLS) module and the Global Pooling Similarity (GPS) module. However, as 3D point cloud data and text data often possess complex geometric structures in high-dimensional space, the proposed RLS employs a novel Riemann Attention Mechanism to reflect the intrinsic geometric relationships of the data. Without explicitly defining the manifold, RMARN learns the manifold parameters to better represent the distances between text-point cloud samples. To address the challenges of lacking paired text-3D data, we have created the large-scale Text-3D Retrieval dataset T3DR-HIT, which comprises over 3,380 pairs of text and point cloud data. T3DR-HIT contains coarse-grained indoor 3D scenes and fine-grained Chinese artifact scenes, consisting of 1,380 and over 2,000 text-3D pairs, respectively. Experiments on our custom datasets demonstrate the superior performance of the proposed method. Wenrui Li 0001, Wei Han 0002, Yandu Chen, Yeyu Chai, Yidan Lu, Xiaopeng Fan 0001 |
AAAI | 4 |