VLDB 2026 Research / reviewers in the wild / expert
Diego Uribe Mora
dblp:439/6270
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0001-7287-188XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems
cold-start recommendation |
1.0 | 1 | 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026 |
Recommender systems › video recommendation
short-video recommendation |
1.0 | 1 | 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026 |
Machine learning › Efficient and distributed learning
attention efficiency |
0.3 | 1 | 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026 |
Methods — techniques the papers use, named apart from their topics
transformer · 2.0temporal folding · 2.0semantic ID · 2.0global query integration · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence ModelingabstractCapturing user interests across extensive watch histories is critical for short-form video recommendation, yet scaling sequence length is limited by two bottlenecks: the semantic sparsity of atomic Video IDs and the quadratic computational complexity of Transformers. Traditional orthogonal Video IDs fail to capture content relationships and demand large embedding tables, while the quadratic complexity of self-attention restricts the maximum sequence length under strict industrial latency and resource constraints. In this work, we present a production-deployed framework for modeling ultra-long user behavior sequences at a billion-user scale. We first address the representation bottleneck by adopting content-native Semantic IDs. By utilizing depth-truncated, coarse-grained Semantic IDs, we shrink the embedding table size from corpus cardinality. This compact representation naturally generalizes to cold-start content through shared semantic prefixes. Second, to overcome the sequence scaling barrier, we introduce a Global-Aware Compression Transformer that leverages non-parametric temporal folding and unified global query integration to effectively condense the sequence, alleviating both the memory and computational bottlenecks of standard self-attention. Offline profiling on our computing infrastructure demonstrates an order-of-magnitude reduction in peak memory footprint and a drastic decrease in computational overhead. This efficiency gain enables supporting longer sequence lengths at an affordable cost in production, yielding substantial online gains in satisfied user engagement and satisfied content consumption in large-scale online A/B tests. Ruixiao Sun, Diego Uribe Mora, Zhimeng Jiang, Yuanzhen Lin, Yuening Li, Danfeng Guo, Zhizhong Chen, Liang Liu 0017 |
SIGIR | 2 |