EDBT 2026 Demo / reviewers in the wild / expert
Xintian Han
dblp:233/4167
· DBLP profile ↗
6ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-1432-5095ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Probabilistic and Bayesian machine learning · 32% Optimization for machine learning · 19% Trustworthy machine learning · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 33% Information retrieval · 33% Data stream processing · 33% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
indexing |
0.9 | 1 | 2025 | Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization Retriever · KDD (2) 2025 |
Data stream processing › streaming data retrieval
stream indexing |
0.9 | 1 | 2025 | Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization Retriever · KDD (2) 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
survival analysis |
0.5 | 1 | 2021 | Inverse-Weighted Survival Games · NeurIPS 2021 |
Machine learning › Trustworthy machine learning
calibration |
0.4 | 1 | 2020 | X-CAL: Explicit Calibration for Survival Analysis · NeurIPS 2020 |
Bioinformatics and computational biology
survival analysis |
0.4 | 1 | 2020 | X-CAL: Explicit Calibration for Survival Analysis · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › relational model
random graph model |
0.3 | 1 | 2018 | A Note on Quickly Sampling a Sparse Matrix with Low Rank Expectation · J. Mach. Learn. Res. 2018 |
Machine learning › Graph learning
stochastic block model |
0.3 | 1 | 2018 | A Note on Quickly Sampling a Sparse Matrix with Low Rank Expectation · J. Mach. Learn. Res. 2018 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2025 | Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization Retriever · KDD (2) 2025 |
Machine learning › Representation and self-supervised learning
vector quantization |
0.3 | 1 | 2025 | Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization Retriever · KDD (2) 2025 |
Algorithms and data structures › randomized algorithms
sampling |
0.1 | 1 | 2018 | A Note on Quickly Sampling a Sparse Matrix with Low Rank Expectation · J. Mach. Learn. Res. 2018 |
Methods — techniques the papers use, named apart from their topics
vector quantization · 1.7approximate nearest neighbor search · 1.7maximum likelihood estimation · 0.9differentiable calibration objective · 0.9poisson sampling · 0.7inverse probability weighting · 0.5game-theoretic training · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RankMixer: Scaling Up Ranking Models in Industrial RecommendersabstractRecent progress on large language models (LLMs) has spurred interest in scaling up recommendation systems, yet two practical obstacles remain. First, training and serving cost on industrial Recommenders must respect strict latency bounds and high QPS demands. Second, most human-designed feature-crossing modules in ranking models were inherited from the CPU era and fail to exploit modern GPUs, resulting in low Model Flops Utilization (MFU) and poor scalability. We introduce RankMixer, a hardware-aware model design tailored towards a unified and scalable feature-interaction architecture. RankMixer retains the transformer's high parallelism while replacing quadratic self-attention with multi-head token mixing module for higher efficiency. Besides, RankMixer maintains both the modeling for distinct feature subspaces and cross-feature-space interactions with Per-token FFNs. We further extend it to one billion parameters with a Sparse-MoE variant for higher ROI. A dynamic routing strategy is adapted to address the inadequacy and imbalance of experts training. Experiments show RankMixer's superior scaling abilities on a trillion-scale production dataset. By replacing previously diverse handcrafted low-MFU modules with RankMixer, we boost the model MFU from 4.5% to 45%, and scale our online ranking model parameters by two orders of magnitude while maintaining roughly the same inference latency. We verify RankMixer's universality with online A/B tests across two core application scenarios (Recommendation and Advertisement). Finally, we launch 1B Dense-Parameters RankMixer for full traffic serving without increasing the serving cost, which improves user active days by 0.3% and total in-app usage duration by 1.08%. Zhifang Fan, Xiaoxie Zhu, Hangyu Wang, Xintian Han, Xinmin Wang, Wenlin Zhao, Huizhi Yang, Zhe Chen 0015, Yuchao Zheng 0002, Qiwei Chen, Feng Zhang 0047, Peng Xu 0017, Zuotao Liu |
CIKM | 6 |
| 2025 | Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization RetrieverabstractRetrievers, which form one of the most important recommendation stages, are responsible for efficiently selecting possible positive samples to the later stages under strict latency limitations. Because of this, large-scale systems always rely on approximate calculations and indexes to roughly shrink candidate scale, with a simple ranking model. Most of the existing methods mainly focus on incorporating complicated ranking models. However, index structure is not improved, which also bottlenecks the whole effectiveness. In this paper, we propose a novel index structure: streaming Vector Quantization model, as a new generation of retrieval paradigm. Streaming VQ attaches items with indexes in real time, granting it immediacy. Moreover, through meticulous verification of possible variants, it achieves additional benefits like index balancing and reparability, enabling it to support complicated ranking models as existing approaches. Streaming VQ has been deployed and replaced all major retrievers in Douyin and Douyin Lite, resulting in remarkable user engagement gain. Xingyan Bin, Jianfei Cui, Wujie Yan, Zhichen Zhao, Xintian Han, Chongyang Yan, Feng Zhang 0047, Zuotao Liu |
KDD (2) | 5 |
| 2021 | Inverse-Weighted Survival GamesabstractDeep models trained through maximum likelihood have achieved state-of-the-art results for survival analysis. Despite this training scheme, practitioners evaluate models under other criteria, such as binary classification losses at a chosen set of time horizons, e.g. Brier score (BS) and Bernoulli log likelihood (BLL). Models trained with maximum likelihood may have poor BS or BLL since maximum likelihood does not directly optimize these criteria. Directly optimizing criteria like BS requires inverse-weighting by the censoring distribution. However, estimating the censoring model under these metrics requires inverse-weighting by the failure distribution. The objective for each model requires the other, but neither are known. To resolve this dilemma, we introduce Inverse-Weighted Survival Games. In these games, objectives for each model are built from re-weighted estimates featuring the other model, where the latter is held fixed during training. When the loss is proper, we show that the games always have the true failure and censoring distributions as a stationary point. This means models in the game do not leave the correct distributions once reached. We construct one case where this stationary point is unique. We show that these games optimize BS on simulations and then apply these principles on real world cancer and critically-ill patient data. Xintian Han, Mark Goldstein, Aahlad Manas Puli, Thomas Wies, Adler J. Perotte, Rajesh Ranganath |
NeurIPS | 1 |
| 2020 | Deep Survival Analysis: The Impact of Feature Missingness
Shreyas Bhave, Xintian Han, Rajesh Ranganath, Adler J. Perotte |
AMIA | 2 |
| 2020 | X-CAL: Explicit Calibration for Survival AnalysisabstractSurvival analysis models the distribution of time until an event of interest, such as discharge from the hospital or admission to the ICU. When a model’s predicted number of events within any time interval is similar to the observed number, it is called well-calibrated. A survival model’s calibration can be measured using, for instance, distributional calibration (D-CALIBRATION) [Haider et al., 2020] which computes the squared difference between the observed and predicted number of events within different time intervals. Classically, calibration is addressed in post-training analysis. We develop explicit calibration (X-CAL), which turns D-CALIBRATION into a differentiable objective that can be used in survival modeling alongside maximum likelihood estimation and other objectives. X-CAL allows us to directly optimize calibration and strike a desired trade-off between predictive power and calibration. In our experiments, we fit a variety of shallow and deep models on simulated data, a survival dataset based on MNIST, on length-of-stay prediction using MIMIC-III data, and on brain cancer data from The Cancer Genome Atlas. We show that the models we study can be miscalibrated. We give experimental evidence on these datasets that X-CAL improves D-CALIBRATION without a large decrease in concordance or likelihood. Mark Goldstein, Xintian Han, Aahlad Manas Puli, Adler J. Perotte, Rajesh Ranganath |
NeurIPS | 2 |
| 2018 | A Note on Quickly Sampling a Sparse Matrix with Low Rank ExpectationabstractGiven matrices $X,Y \in R^{n \times K}$ and $S \in R^{K \times K}$ with positive elements, this paper proposes an algorithm fastRG to sample a sparse matrix $A$ with low rank expectation $E(A) = XSY^T$ and independent Poisson elements. This allows for quickly sampling from a broad class of stochastic blockmodel graphs (degree-corrected, mixed membership, overlapping) all of which are specific parameterizations of the generalized random product graph model defined in Section 2.2. The basic idea of fastRG is to first sample the number of edges $m$ and then sample each edge. The key insight is that because of the the low rank expectation, it is easy to sample individual edges. The naive “element-wise” algorithm requires $O(n^2)$ operations to generate the $n\times n$ adjacency matrix $A$. In sparse graphs, where $m = O(n)$, ignoring log terms, fastRG runs in time $O(n)$. An implementation in R is available on github. A computational experiment in Section 2.4 simulates graphs up to $n=10,000,000$ nodes with $m = 100,000,000$ edges. For example, on a graph with $n=500,000$ and $m = 5,000,000$, fastRG runs in less than one second on a 3.5 GHz Intel i5. Karl Rohe, Xintian Han, Norbert Binkiewicz |
J. Mach. Learn. Res. | 3 |