Xuanyuan Luo

dblp:236/4468 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0008-8412-3222ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Learning theory · 53% Optimization for machine learning · 40% Deep learning architectures and training · 7%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 77% Hardware accelerators and domain-specific architectures · 23%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
generalization bounds
1.022022
Generalization Bounds for Gradient Methods via Discrete and Continuous Prior · NeurIPS 2022
On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex Learning · ICLR 2020
Recommender systems
click-through rate prediction
1.012026
HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction · SIGIR 2026
Recommender systems › click-through rate prediction
feature interaction
1.012026
HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction · SIGIR 2026
Recommender systems › sequential recommendation
user behavior sequence modeling
1.012026
HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction · SIGIR 2026
Machine learning › Learning theory › generalization
generalization analysis
0.612022
Generalization Bounds for Gradient Methods via Discrete and Continuous Prior · NeurIPS 2022
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.612022
Generalization Bounds for Gradient Methods via Discrete and Continuous Prior · NeurIPS 2022
Machine learning › Learning theory › generalization bounds
PAC-Bayes bounds
0.612022
Generalization Bounds for Gradient Methods via Discrete and Continuous Prior · NeurIPS 2022
Machine learning › Optimization for machine learning › stochastic optimization
noisy gradient descent
0.412020
On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex Learning · ICLR 2020
Machine learning › Optimization for machine learning
non-convex optimization
0.412020
On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex Learning · ICLR 2020
Machine learning › Deep learning architectures and training
transformer
0.312026
HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction · SIGIR 2026
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.312026
Make It Long, Keep It Fast: End-to-End 10k-Sequence Modeling at Billion Scale on Douyin · WWW 2026
Machine learning › Optimization for machine learning
stochastic gradient descent
0.212022
Generalization Bounds for Gradient Methods via Discrete and Continuous Prior · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

transformer · 2.0alternating optimization · 2.0gradient descent · 1.0sequence parallelism · 1.0distributed training · 1.0langevin dynamics · 0.6PAC-Bayesian framework · 0.6stability analysis · 0.4
YearPublicationVenuePosition
2026 HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction
abstract
Industrial large-scale recommendation models (LRMs) face the challenge of jointly modeling long-range user behavior sequences and heterogeneous non-sequential features under strict efficiency constraints. However, most existing architectures employ a decoupled pipeline: long sequences are first compressed with a query-token based sequence compressor like LONGER, followed by fusion with dense features through token-mixing modules like RankMixer, which thereby limits both the representation capacity and the interaction flexibility. This paper presents HyFormer, a unified hybrid transformer architecture that tightly integrates long-sequence modeling and feature interaction into a single backbone. From the perspective of sequence modeling, we revisit and redesign query tokens in LRMs, and frame the LRM modeling task as an alternating optimization process that integrates two core components: Query Decoding which expands non-sequential features into Global Tokens and performs long sequence decoding over layer-wise key-value representations of long behavioral sequences; and Query Boosting which enhances cross-query and cross-sequence heterogeneous interactions via efficient token mixing. The two complementary mechanisms are performed iteratively to refine semantic representations across layers. Extensive experiments on billion-scale industrial datasets demonstrate that HyFormer consistently outperforms strong LONGER and RankMixer baselines under comparable parameter and FLOPs budgets, while exhibiting superior scaling behavior with increasing parameters and FLOPs. Large-scale online A/B tests in high-traffic production systems further validate its effectiveness, showing significant gains over deployed state-of-the-art models. These results highlight the practicality and scalability of HyFormer as a unified modeling framework for industrial LRMs.
Yunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin, Xuanyuan Luo, Zhe Wang 0060, Shikang Wu, Yuchao Zheng 0002, Jingjian Lin
SIGIR5
2026 Make It Long, Keep It Fast: End-to-End 10k-Sequence Modeling at Billion Scale on Douyin
Jia-Qi Yang 0001, Zhishan Zhao, Beichuan Zhang 0002, Xuanyuan Luo, Jinan Ni, Yuhang Qi, Zhifang Fan, Hangyu Wang, Qiwei Chen, Feng Zhang 0047
WWW6
2022 Generalization Bounds for Gradient Methods via Discrete and Continuous Prior
abstract
Proving algorithm-dependent generalization error bounds for gradient-type optimization methods has attracted significant attention recently in learning theory. However, most existing trajectory-based analyses require either restrictive assumptions on the learning rate (e.g., fast decreasing learning rate), or continuous injected noise (such as the Gaussian noise in Langevin dynamics). In this paper, we introduce a new discrete data-dependent prior to the PAC-Bayesian framework, and prove a high probability generalization bound of order $O(\frac{1}{n}\cdot \sum_{t=1}^T(\gamma_t/\varepsilon_t)^2\left\|{\mathrm{g}_t}\right\|^2)$ for Floored GD (i.e. a version of gradient descent with precision level $\varepsilon_t$), where $n$ is the number of training samples, $\gamma_t$ is the learning rate at step $t$, $\mathrm{g}_t$ is roughly the difference of the gradient computed using all samples and that using only prior samples. $\left\|{\mathrm{g}_t}\right\|$ is upper bounded by and and typical much smaller than the gradient norm $\left\|{\nabla f(W_t)}\right\|$. We remark that our bound holds for nonconvex and nonsmooth scenarios. Moreover, our theoretical results provide numerically favorable upper bounds of testing errors (e.g., $0.037$ on MNIST). Using similar technique, we can also obtain new generalization bounds for a certain variant of SGD. Furthermore, we study the generalization bounds for gradient Langevin Dynamics (GLD). Using the same framework with a carefully constructed continuous prior, we show a new high probability generalization bound of order $O(\frac{1}{n} + \frac{L^2}{n^2}\sum_{t=1}^T(\gamma_t/\sigma_t)^2)$ for GLD. The new $1/n^2$ rate is due to the concentration of the difference between the gradient of training samples and that of the prior.
Xuanyuan Luo, Bei Luo
NeurIPS1
2020 On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex Learning
Jian Li 0015, Xuanyuan Luo, Mingda Qiao
ICLR2