VLDB 2026 Research / reviewers in the wild / expert
Ruixiao Sun
dblp:232/1881
· DBLP profile ↗
5ranked-venue papers
4as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 63% Vision and language · 37% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems
cold-start recommendation |
1.0 | 1 | 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026 |
Recommender systems › video recommendation
short-video recommendation |
1.0 | 1 | 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026 |
Machine learning › Efficient and distributed learning
active learning |
0.9 | 1 | 2025 | Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion · KDD (2) 2025 |
Computer vision › Vision and language
multimodal fusion |
0.9 | 1 | 2025 | Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion · KDD (2) 2025 |
Machine learning › Efficient and distributed learning
attention efficiency |
0.3 | 1 | 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026 |
Audio and music processing › acoustic signal processing › audio signal reconstruction › audio restoration
audio enhancement |
0.3 | 1 | 2025 | Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion · KDD (2) 2025 |
Methods — techniques the papers use, named apart from their topics
transformer · 2.0temporal folding · 2.0semantic ID · 2.0global query integration · 2.0mid-fusion · 1.7kNN-based latent space broadening · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence ModelingabstractCapturing user interests across extensive watch histories is critical for short-form video recommendation, yet scaling sequence length is limited by two bottlenecks: the semantic sparsity of atomic Video IDs and the quadratic computational complexity of Transformers. Traditional orthogonal Video IDs fail to capture content relationships and demand large embedding tables, while the quadratic complexity of self-attention restricts the maximum sequence length under strict industrial latency and resource constraints. In this work, we present a production-deployed framework for modeling ultra-long user behavior sequences at a billion-user scale. We first address the representation bottleneck by adopting content-native Semantic IDs. By utilizing depth-truncated, coarse-grained Semantic IDs, we shrink the embedding table size from corpus cardinality. This compact representation naturally generalizes to cold-start content through shared semantic prefixes. Second, to overcome the sequence scaling barrier, we introduce a Global-Aware Compression Transformer that leverages non-parametric temporal folding and unified global query integration to effectively condense the sequence, alleviating both the memory and computational bottlenecks of standard self-attention. Offline profiling on our computing infrastructure demonstrates an order-of-magnitude reduction in peak memory footprint and a drastic decrease in computational overhead. This efficiency gain enables supporting longer sequence lengths at an affordable cost in production, yielding substantial online gains in satisfied user engagement and satisfied content consumption in large-scale online A/B tests. Ruixiao Sun, Diego Uribe Mora, Zhimeng Jiang, Yuanzhen Lin, Yuening Li, Danfeng Guo, Zhizhong Chen, Liang Liu 0017 |
SIGIR | 1 |
| 2025 | Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data ExpansionabstractTransformer-based multimodal models are widely used in industrialscale recommendation, search, and advertising systems for content understanding and relevance ranking.Enhancing labeled training data quality and cross-modal fusion significantly improves model performance, influencing key metrics such as quality view rates and ad revenue.High-quality annotations are crucial for advancing content modeling, yet traditional statistical-based active learning (AL) methods face limitations: they struggle to detect overconfident misclassifications and are less effective in distinguishing semantically similar items in deep neural networks.Additionally, audio information plays an increasing role, especially in short-video platforms, yet most pretrained multimodal architectures primarily focus on text and images.While training from scratch across all three modalities is possible, it sacrifices the benefits of leveraging existing pretrained visual-language (VL) and audio models.To address these challenges, we propose kNN-based Latent Space Broadening (LSB) to enhance AL efficiency, achieving an up to 9% recall improvement at 80% precision on proprietary datasets.Additionally, we introduce Vision-Language Modeling with Audio Enhancement (VLMAE), a mid-fusion approach integrating audio into VL models, yielding up * Author corresponded for this research. Yu Sun 0088, Ruixiao Sun, Chunhui Liu 0002, Fangming Zhou, Ze Jin, Xiang Shen 0001, Zhuolin Hao, Hongyu Xiong |
KDD (2) | 3 |
| 2023 | Physical-informed deep learning framework for CO2-injected EOR compositional simulation
Ruixiao Sun, Huanquan Pan, Hongyu Xiong, Hamdi A. Tchelepi |
Eng. Appl. Artif. Intell. | 1 |
| 2021 | TRANSIT-GYM: A Simulation and Evaluation Engine for Analysis of Bus Transit SystemsabstractPublic-transit systems face a number of operational challenges: (a) changing ridership patterns requiring optimization of fixed line services, (b) optimizing vehicle-to-trip assignments to reduce maintenance and operation codes, and (c) ensuring equitable and fair coverage to areas with low ridership. Optimizing these objectives presents a hard computational problem due to the size and complexity of the decision space. State-of-the-art methods formulate these problems as variants of the vehicle routing problem and use data-driven heuristics for optimizing the procedures. However, the evaluation and training of these algorithms require large datasets that provide realistic coverage of various operational uncertainties. This paper presents a dynamic simulation platform, called TRANSIT-GYM, that can bridge this gap by providing the ability to simulate scenarios, focusing on variation of demand models, variations of route networks, and variations of vehicle-to-trip assignments. The central contribution of this work is a domain-specific language and associated experimentation tool-chain and infrastructure to enable subject-matter experts to intuitively specify, simulate, and analyze large-scale transit scenarios and their parametric variations. Of particular significance is an integrated microscopic energy consumption model that also helps to analyze the energy cost of various transit decisions made by the transportation agency of a city. Ruixiao Sun, Rongze Gui, Himanshu Neema, Yuche Chen, Juliette Ugirumurera, Joseph Severino, Philip Pugliese, Aron Laszka, Abhishek Dubey |
SMARTCOMP | 1 |
| 2019 | Novel single-valued neutrosophic decision-making approaches based on prospect theory and their applications in physician selection
Ruixiao Sun, Junhua Hu, Xiaohong Chen 0001 |
Soft Comput. | 1 |