Ruixiao Sun

dblp:232/1881 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 63% Vision and language · 37%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems
cold-start recommendation
1.012026
Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026
Recommender systems › video recommendation
short-video recommendation
1.012026
Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026
Machine learning › Efficient and distributed learning
active learning
0.912025
Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion · KDD (2) 2025
Computer vision › Vision and language
multimodal fusion
0.912025
Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion · KDD (2) 2025
Machine learning › Efficient and distributed learning
attention efficiency
0.312026
Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026
Machine learning › Efficient and distributed learning
model compression
0.312026
Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling · SIGIR 2026
Audio and music processing › acoustic signal processing › audio signal reconstruction › audio restoration
audio enhancement
0.312025
Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion · KDD (2) 2025

Methods — techniques the papers use, named apart from their topics

transformer · 2.0temporal folding · 2.0semantic ID · 2.0global query integration · 2.0mid-fusion · 1.7kNN-based latent space broadening · 1.7
YearPublicationVenuePosition
2026 Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling
abstract
Capturing user interests across extensive watch histories is critical for short-form video recommendation, yet scaling sequence length is limited by two bottlenecks: the semantic sparsity of atomic Video IDs and the quadratic computational complexity of Transformers. Traditional orthogonal Video IDs fail to capture content relationships and demand large embedding tables, while the quadratic complexity of self-attention restricts the maximum sequence length under strict industrial latency and resource constraints. In this work, we present a production-deployed framework for modeling ultra-long user behavior sequences at a billion-user scale. We first address the representation bottleneck by adopting content-native Semantic IDs. By utilizing depth-truncated, coarse-grained Semantic IDs, we shrink the embedding table size from corpus cardinality. This compact representation naturally generalizes to cold-start content through shared semantic prefixes. Second, to overcome the sequence scaling barrier, we introduce a Global-Aware Compression Transformer that leverages non-parametric temporal folding and unified global query integration to effectively condense the sequence, alleviating both the memory and computational bottlenecks of standard self-attention. Offline profiling on our computing infrastructure demonstrates an order-of-magnitude reduction in peak memory footprint and a drastic decrease in computational overhead. This efficiency gain enables supporting longer sequence lengths at an affordable cost in production, yielding substantial online gains in satisfied user engagement and satisfied content consumption in large-scale online A/B tests.
Ruixiao Sun, Diego Uribe Mora, Zhimeng Jiang, Yuanzhen Lin, Yuening Li, Danfeng Guo, Zhizhong Chen, Liang Liu 0017
SIGIR1
2025 Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion
abstract
Transformer-based multimodal models are widely used in industrialscale recommendation, search, and advertising systems for content understanding and relevance ranking.Enhancing labeled training data quality and cross-modal fusion significantly improves model performance, influencing key metrics such as quality view rates and ad revenue.High-quality annotations are crucial for advancing content modeling, yet traditional statistical-based active learning (AL) methods face limitations: they struggle to detect overconfident misclassifications and are less effective in distinguishing semantically similar items in deep neural networks.Additionally, audio information plays an increasing role, especially in short-video platforms, yet most pretrained multimodal architectures primarily focus on text and images.While training from scratch across all three modalities is possible, it sacrifices the benefits of leveraging existing pretrained visual-language (VL) and audio models.To address these challenges, we propose kNN-based Latent Space Broadening (LSB) to enhance AL efficiency, achieving an up to 9% recall improvement at 80% precision on proprietary datasets.Additionally, we introduce Vision-Language Modeling with Audio Enhancement (VLMAE), a mid-fusion approach integrating audio into VL models, yielding up * Author corresponded for this research.
Yu Sun 0088, Ruixiao Sun, Chunhui Liu 0002, Fangming Zhou, Ze Jin, Xiang Shen 0001, Zhuolin Hao, Hongyu Xiong
KDD (2)3
2023 Physical-informed deep learning framework for CO2-injected EOR compositional simulation
Ruixiao Sun, Huanquan Pan, Hongyu Xiong, Hamdi A. Tchelepi
Eng. Appl. Artif. Intell.1
2021 TRANSIT-GYM: A Simulation and Evaluation Engine for Analysis of Bus Transit Systems
abstract
Public-transit systems face a number of operational challenges: (a) changing ridership patterns requiring optimization of fixed line services, (b) optimizing vehicle-to-trip assignments to reduce maintenance and operation codes, and (c) ensuring equitable and fair coverage to areas with low ridership. Optimizing these objectives presents a hard computational problem due to the size and complexity of the decision space. State-of-the-art methods formulate these problems as variants of the vehicle routing problem and use data-driven heuristics for optimizing the procedures. However, the evaluation and training of these algorithms require large datasets that provide realistic coverage of various operational uncertainties. This paper presents a dynamic simulation platform, called TRANSIT-GYM, that can bridge this gap by providing the ability to simulate scenarios, focusing on variation of demand models, variations of route networks, and variations of vehicle-to-trip assignments. The central contribution of this work is a domain-specific language and associated experimentation tool-chain and infrastructure to enable subject-matter experts to intuitively specify, simulate, and analyze large-scale transit scenarios and their parametric variations. Of particular significance is an integrated microscopic energy consumption model that also helps to analyze the energy cost of various transit decisions made by the transportation agency of a city.
Ruixiao Sun, Rongze Gui, Himanshu Neema, Yuche Chen, Juliette Ugirumurera, Joseph Severino, Philip Pugliese, Aron Laszka, Abhishek Dubey
SMARTCOMP1
2019 Novel single-valued neutrosophic decision-making approaches based on prospect theory and their applications in physician selection
Ruixiao Sun, Junhua Hu, Xiaohong Chen 0001
Soft Comput.1