Xu Min

dblp:08/2810 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0002-5952-0794ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2024 MIBR: Bridging Domains through Diverse Interests for Cross-Domain Sequential Recommendation
abstract
Cross-Domain Sequential Recommendation (CDSR) aims to enhance personalized user experiences by leveraging user behaviors across multiple domains. Existing methods primarily focus on fusing information from various domains and modeling global user preferences, but often struggle with negative transfer, where knowledge from one domain impairs recommendation performance in another. For example, a user may enjoy watching sports games in the video domain but have no interest in participating in sports activities. Consequently, this interest does not extend to purchasing related sports gear. In such cases, a recommendation system suggesting sports gear based on the user’s viewing preferences may not elicit a positive response. To tackle this issue, we propose a novel method called Multi-Interest Bridge Recommender (MIBR). In light of the cross-domain scenario, where user preferences are not entirely consistent across domains, we design a Multi-Interest Extraction (MIE) module to capture the diversity of user interests based on a soft clustering approach. In the meantime, we design a cross-domain bridging (CDB) module, with the goal of mitigating the issue of negative transfer. CDB leverages the extracted interests as a bridge for inter-domain information transfer, enabling each domain to adaptively extract relevant information from diverse interests while ignoring unrelated ones. Extensive experiments on three popular datasets reveal MIBR’s significant superiority over baselines, e.g., with up to a 59.27% uplift in terms of HR@10 over C2DSR on the Movie-Book dataset.
Chengzhe Zhang, Xu Min, Weichang Wu, Jun Zhou 0011, Ye Yuan 0001, Guoren Wang
IEEE Big Data2
2024 Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations
abstract
Bayesian flow networks (BFNs) iteratively refine the parameters, instead of the samples in diffusion models (DMs), of distributions at various noise levels through Bayesian inference. Owing to its differentiable nature, BFNs are promising in modeling both continuous and discrete data, while simultaneously maintaining fast sampling capabilities. This paper aims to understand and enhance BFNs by connecting them with DMs through stochastic differential equations (SDEs). We identify the linear SDEs corresponding to the noise-addition processes in BFNs, demonstrate that BFN’s regression losses are aligned with denoise score matching, and validate the sampler in BFN as a first-order solver for the respective reverse-time SDE. Based on these findings and existing recipes of fast sampling in DMs, we propose specialized solvers for BFNs that markedly surpass the original BFN sampler in terms of sample quality with a limited number of function evaluations (e.g., 10) on both image and text datasets. Notably, our best sampler achieves an increase in speed of $5\sim20$ times for free.
Shen Nie, Xu Min, Jun Zhou 0011, Chongxuan Li
ICML4
2024 Lower Bounds of Uniform Stability in Gradient-Based Bilevel Algorithms for Hyperparameter Optimization
abstract
Gradient-based bilevel programming leverages unrolling differentiation (UD) or implicit function theorem (IFT) to solve hyperparameter optimization (HO) problems, and is proven effective and scalable in practice. To understand their generalization behavior, existing works establish upper bounds on the uniform stability of these algorithms, while their tightness is still unclear. To this end, this paper attempts to establish stability lower bounds for UD-based and IFT-based algorithms. A central technical challenge arises from the dependency of each outer-level update on the concurrent stage of inner optimization in bilevel programming. To address this problem, we introduce lower-bounded expansion properties to characterize the instability in update rules which can serve as general tools for lower-bound analysis. These properties guarantee the hyperparameter divergence at the outer level and the Lipschitz constant of inner output at the inner level in the context of HO. Guided by these insights, we construct a quadratic example that yields tight lower bounds for the UD-based algorithm and meaningful bounds for a representative IFT-based algorithm. Our tight result indicates that uniform stability has reached its limit in stability analysis for the UD-based algorithm.
Rongzhen Wang, Chenyu Zheng, Guoqiang Wu, Xu Min, Jun Zhou 0011, Chongxuan Li
NeurIPS4
2024 Exploring Multi-Scenario Multi-Modal CTR Prediction with a Large Scale Dataset
abstract
Click-through rate (CTR) prediction plays a crucial role in recommendation systems, with significant impact on user experience and platform revenue generation. Despite the various public CTR datasets available due to increasing interest from both academia and industry, these datasets have limitations. They cover a limited range of scenarios and predominantly focus on ID-based features, neglecting the vital role of multi-modal features for effective multi-scenario CTR prediction. Moreover, their scale is modest compared to real-world industrial datasets, hindering robust and comprehensive evaluation of complex models. To address these challenges, we introduce a large-scale Multi-Scenario Multi-Modal CTR dataset named AntM2 C, built from real industrial data from Alipay. This dataset offers an impressive breadth and depth of information, covering CTR data from four diverse business scenarios, including advertisements, consumer coupons, mini-programs, and videos. Unlike existing datasets, AntM2 C provides not only ID-based features but also five textual features and one image feature for both users and items, supporting more delicate multi-modal CTR prediction. AntM2 C is also substantially larger than existing datasets, comprising 100 million CTR data. This scale allows for robust and comprehensive evaluation and comparison of CTR prediction models. We employ AntM2 C to construct several typical CTR tasks, including multi-scenario modeling, item and user cold-start modeling, and multi-modal modeling. Initial experiments and comparisons with baseline methods have shown that AntM2 C presents both new challenges and opportunities for CTR models, with the potential to significantly advance CTR research. The AntM2 C dataset is available at https://www.atecup.cn/OfficalDataSet.
Zhaoxin Huan, Ke Ding 0001, Ang Li 0043, Xu Min, Yong He 0009, Liang Zhang 0045, Jun Zhou 0011, Linjian Mo, Jinjie Gu, Zhongyi Liu 0001, Leon Wenliang Zhong, Chenliang Li 0005, Fajie Yuan
SIGIR5
2023 SeqGen: A Sequence Generator via User Side Information for Behavior Sparsity in Recommendation
abstract
In real-world industrial advertising systems, user behavior sparsity is a key issue that affects online recommendation performance. We observe that users with rich behaviors can obtain better recommendation results than those with sparse behaviors in a conversion-rate (CVR) prediction model. Inspired by this phenomenon, we propose a new method SeqGen, in an effort to exploit user side information to bridge the gap between rich and sparse behaviors. SeqGen is a learnable and pluggable module, which can be easily integrated into any CVR model and no longer requires two-stage training as in previous works. In particular, SeqGen learns a mapping relationship between the user side information and behavior sequences, only on the basis of the users with long behavior sequences. After that, SeqGen can generate rich sequence features for users with sparse behaviors based on their side information, so as to alleviate the issue of user behavior sparsity. The generated sequence features will then be fed into the classifier tower of an arbitrary CVR model together with the original sequence features. To the best of our knowledge, our approach constitutes the first attempt to exploit user side information for addressing the user behavior sparsity issue. We validate the effectiveness of SeqGen on the publicly available dataset MovieLens-1M, and our method receives an improvement of up to 0.5% in terms of the AUC score. More importantly, we successfully deploy SeqGen in the commercial advertising system Xlight of Alipay, which improves the grouped AUC of the CVR model by 0.6% and brings a boost of 0.49% in terms of the conversion rate on A/B testing.
Xu Min, Yong He 0009, Jun Zhou 0011
CIKM1
2023 DAE: Distribution-Aware Embedding for Numerical Features in Click-Through Rate Prediction
abstract
Numerical features are an important type of input for CTR prediction models. Recently, several discretization and numerical transformation methods have been proposed to deal with numerical features. However, existing approaches do not fully consider compatibility with different distributions. Here, we propose a novel numerical feature embedding framework, called Distribution-Aware Embedding (DAE), which is applicable to various numerical feature distributions. First, DAE efficiently approximates the cumulative distribution function by estimating the expectation of the order statistics. Then, the distribution information is applied to the embedding layer by nonlinear interpolation. Finally, to capture both local and global information, we aggregate the embeddings at multiple scales to obtain the final representation. Empirical results validate the effectiveness of DAE compared to the baselines, while demonstrating the adaptability to different CTR models and distributions.
Xu Min, Zeyu Ke, Yong He 0009, Liang Zhang 0045, Xin Dong 0012, Linjian Mo
CIKM3
2023 SAMD: An Industrial Framework for Heterogeneous Multi-Scenario Recommendation
abstract
Industrial recommender systems usually need to serve multiple scenarios at the same time. In practice, there are various heterogeneous scenarios, since users frequently engage in scenarios with varying intentions and the items within each scenario typically belong to diverse categories. Existing works of multi-scenario recommendation mainly focus on modeling homogeneous scenarios which have similar data distributions. They equally transfer knowledge to each scenario without considering the diversity of heterogeneous scenarios. In this paper, we argue that the heterogeneity in multi-scenario recommendations is a key problem that needs to be solved. To this end, we propose an industrial framework named Scenario-Aware Model-Agnostic Meta Distillation (SAMD) for the multi-scenario recommendation. SAMD aims to provide scenario-aware and model-agnostic knowledge sharing across heterogeneous scenarios by modeling scenarios' relationship and conducting heterogeneous knowledge distillation. Specifically, SAMD first measures the comprehensive representation of each scenario and then proposes a novel meta distillation paradigm to conduct scenario-aware knowledge sharing. The meta network first establishes the potential scenarios' relationships and generates the strategies of knowledge sharing for each scenario. Then the heterogeneous knowledge distillation utilizes scenario-aware strategies to share knowledge across heterogeneous scenarios through intermediate features distillation without the restriction of the model architecture. In this way, SAMD shares knowledge across heterogeneous scenarios in a scenario-aware and model-agnostic manner, which addresses the problem of heterogeneity. Compared with other state-of-the-art methods, extensive offline experiments, and online A/B testing demonstrate the superior performance of the proposed SAMD framework, especially in heterogeneous scenarios.
Zhaoxin Huan, Ang Li 0043, Xu Min, Jieyu Yang, Yong He 0009, Jun Zhou 0011
KDD4
2023 Uncertainty-based Heterogeneous Privileged Knowledge Distillation for Recommendation System
abstract
In industrial recommendation systems, both data sizes and computational resources vary across different scenarios. For scenarios with limited data, data sparsity can lead to a decrease in model performance. Heterogeneous knowledge distillation-based transfer learning can be used to transfer knowledge from models in data-rich domains. However, in recommendation systems, the target domain possesses specific privileged features that significantly contribute to the model. While existing knowledge distillation methods have not taken these features into consideration, leading to suboptimal transfer weights. To overcome this limitation, we propose a novel algorithm called Uncertainty-based Heterogeneous Privileged Knowledge Distillation (UHPKD). Our method aims to quantify the knowledge of both the source and target domains, which represents the uncertainty of the models. This approach allows us to derive transfer weights based on the knowledge gain, which captures the difference in knowledge between the source and target domains. Experiments conducted on both public and industrial datasets demonstrate the superiority of our UHPKD algorithm compared to other state-of-the-art methods.
Ang Li 0043, Jian Hu 0002, Ke Ding 0001, Jun Zhou 0011, Yong He 0009, Xu Min
SIGIR7
2022 Cost and Care Insight: An Interactive and Scalable Hierarchical Learning System for Identifying Cost Saving Opportunities
David Koepke, Bibo Hao, Jing Mei, Xu Min, Rachna Gupta, Rajashree Joshi, Fiona McNaughton, Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001
ICIC (1)5
2022 TAMOR: Tier-Aware Multi-objective Recommendation for Ant Fortune Financial Marketing
Xu Min, Jun Zhou 0011, Changxun Fan, Junlin Yu
ECML/PKDD (6)1
2020 Embracing Disease Progression with a Learning System for Real World Evidence Discovery
Zefang Tang, Lun Hu, Xu Min, Jing Mei, Kenney Ng, Shaochun Li, Pengwei Hu 0001, Zhu-Hong You
ICIC (2)3
2017 Patient outcome prediction via convolutional neural networks based on multi-granularity medical concept embedding
abstract
The large availability of biomedical data brings opportunities and challenges to health care. Representation of medical concepts has been well studied in many applications, such as medical informatics, cohort selection, risk prediction, and health care quality measurement. In this paper, we propose an efficient multichannel convolutional neural network (CNN) model based on multi-granularity embeddings of medical concepts named MG-CNN, to examine the effect of individual patient characteristics including demographic factors and medical comorbidities on total hospital costs and length of stay (LOS) by using the Hospital Quality Monitoring System (HQMS) data. The proposed embedding method leverages prior medical hierarchical ontology and improves the quality of embedding for rare medical concepts. The embedded vectors are further visualized by the t-Distributed Stochastic Neighbor Embedding (t-SNE) technique to demonstrate the effectiveness of grouping related medical concepts. Experimental results demonstrate that our MG-CNN model outperforms traditional regression methods based on the one-hot representation of medical concepts, especially in the outcome prediction tasks for patients with low-frequency medical events. In summary, MG-CNN model is capable of mining potential knowledge from the clinical data and will be broadly applicable in medical research and inform clinical decisions.
Yujuan Feng, Xu Min, Ning Chen 0002, Xiaolei Xie, Ting Chen 0006
BIBM2
2017 Chromatin accessibility prediction via convolutional long short-term memory networks with k-mer embedding
abstract
MOTIVATION: Experimental techniques for measuring chromatin accessibility are expensive and time consuming, appealing for the development of computational approaches to predict open chromatin regions from DNA sequences. Along this direction, existing methods fall into two classes: one based on handcrafted k -mer features and the other based on convolutional neural networks. Although both categories have shown good performance in specific applications thus far, there still lacks a comprehensive framework to integrate useful k -mer co-occurrence information with recent advances in deep learning. RESULTS: We fill this gap by addressing the problem of chromatin accessibility prediction with a convolutional Long Short-Term Memory (LSTM) network with k -mer embedding. We first split DNA sequences into k -mers and pre-train k -mer embedding vectors based on the co-occurrence matrix of k -mers by using an unsupervised representation learning approach. We then construct a supervised deep learning architecture comprised of an embedding layer, three convolutional layers and a Bidirectional LSTM (BLSTM) layer for feature learning and classification. We demonstrate that our method gains high-quality fixed-length features from variable-length sequences and consistently outperforms baseline methods. We show that k -mer embedding can effectively enhance model performance by exploring different embedding strategies. We also prove the efficacy of both the convolution and the BLSTM layers by comparing two variations of the network architecture. We confirm the robustness of our model to hyper-parameters by performing sensitivity analysis. We hope our method can eventually reinforce our understanding of employing deep learning in genomic studies and shed light on research regarding mechanisms of chromatin accessibility. AVAILABILITY AND IMPLEMENTATION: The source code can be downloaded from https://github.com/minxueric/ismb2017_lstm . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online.
Xu Min, Wanwen Zeng, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001
Bioinform.1
2017 Predicting enhancers with deep convolutional neural networks
abstract
BACKGROUND: With the rapid development of deep sequencing techniques in the recent years, enhancers have been systematically identified in such projects as FANTOM and ENCODE, forming genome-wide landscapes in a series of human cell lines. Nevertheless, experimental approaches are still costly and time consuming for large scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable. RESULTS: To facilitate the identification of enhancers, we propose a computational framework, named DeepEnhancer, to distinguish enhancers from background genomic sequences. Our method purely relies on DNA sequences to predict enhancers in an end-to-end manner by using a deep convolutional neural network (CNN). We train our deep learning model on permissive enhancers and then adopt a transfer learning strategy to fine-tune the model on enhancers specific to a cell line. Results demonstrate the effectiveness and efficiency of our method in the classification of enhancers against random sequences, exhibiting advantages of deep learning over traditional sequence-based classifiers. We then construct a variety of neural networks with different architectures and show the usefulness of such techniques as max-pooling and batch normalization in our method. To gain the interpretability of our approach, we further visualize convolutional kernels as sequence logos and successfully identify similar motifs in the JASPAR database. CONCLUSIONS: DeepEnhancer enables the identification of novel enhancers using only DNA sequences via a highly accurate deep learning model. The proposed computational framework can also be applied to similar problems, thereby prompting the use of machine learning methods in life sciences.
Xu Min, Wanwen Zeng, Shengquan Chen, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001
BMC Bioinform.1
2016 DeepEnhancer: Predicting enhancers by convolutional neural networks
abstract
Enhancers are crucial to the understanding of mechanisms underlying gene transcriptional regulation. Although having been successfully applied in such projects as ENCODE and Roadmap to generate landscape of enhancers in human cell lines, high-throughput biological experimental techniques are still costly and time consuming for even larger scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable. In this paper, we propose a computational framework, named DeepEnhancer, to classify enhancers from background genomic sequences. We construct convolutional neural networks of various architectures and compare the classification performance with traditional sequence-based classifiers. We first train the deep learning model on the FANTOM5 permissive enhancer dataset, and then fine-tune the model on ENCODE cell type-specific enhancer datasets by adopting the transfer learning strategy. Experimental results demonstrate that DeepEnhancer has superior efficiency and effectiveness in classification tasks, and the use of max-pooling and batch normalization is beneficial to higher accuracy. To make our approach more understandable, we propose a strategy to visualize the convolutional kernels as sequence logos and compare them against the JASPAR database using TOMTOM. In summary, DeepEnhancer allows researchers to train highly accurate deep models and will be broadly applicable in computational biology.
Xu Min, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001
BIBM1
2014 Real-time object tracking via optimal feature subspace
abstract
In this paper, we present a real-time tracking approach based on the Optimal Feature Subspace (OFS). OFS is an optimal subspace of a random feature space, which can best represent the target and making it most distinguished in the whole scene. Initially, we randomly crop patches inside the bounding box to generate an efficient feature template set. Then a greedy algorithm fusing the cues of both target and background is proposed to seek the OFS at every frame. In the forthcoming frame, considering the correlation of different dimensions, we compute the Mahalanobis distance of candidate patches to the appearance model in the obtained subspace to locate the target. The experimental results on several challenging video clips demonstrate that our approach outperforms the state-of-the-art methods, in terms of both speed and robustness.
Xu Min, Yu Zhou 0016, Xiang Bai
ICIP1