EDBT 2026 Demo / reviewers in the wild / expert
Xinliang Zhu
dblp:160/9968
· DBLP profile ↗
18ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-4544-2078ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GENIUS: A Generative Framework for Universal Multimodal SearchabstractGenerative retrieval is an emerging approach in information retrieval that generates identifiers (IDs) of target data based on a query, providing an efficient alternative to traditional embedding-based retrieval methods. However, existing models are task-specific and fall short of embedding-based retrieval in performance. This paper proposes GENIUS, a universal generative retrieval framework supporting diverse tasks across multiple modalities and domains. At its core, GENIUS introduces modality-decoupled semantic quantization, transforming multimodal data into discrete IDs encoding both modality and semantics. Moreover, to enhance generalization, we propose a query augmentation that interpolates between a query and its target, allowing GENIUS to adapt to varied query forms. Evaluated on the M-BEIR benchmark, it surpasses prior generative methods by a clear margin. Unlike embedding-based retrieval, GENIUS consistently maintains high retrieval speed across database size, with competitive performance across multiple benchmarks. With additional re-ranking, GENIUS often achieves results close to those of embedding-based methods while preserving efficiency. Sungyeon Kim, Xinliang Zhu, Muhammet Bastan, Douglas Gray 0001, Suha Kwak |
CVPR | 2 |
| 2024 | Planes, Trains and Automobiles: Leverage Multimodal In-Mission Signals for Shopping JourneysabstractModern search systems offer multiple ways for expressing information needs, including image, voice, and text. Consequently, an increasing number of users seamlessly transition between these modalities to convey their intents. This emerging trend presents new opportunities for utilizing queries in different modalities to help users complete their search journeys efficiently. In this proposal, we introduce an approach to segmenting a multimodal query stream into missions, demonstrate how these in-mission queries can enhance search ranking, and outline key areas for future research. Viet Ha-Thuc, Shasha Li 0001, Arnau Ramisa, Xinliang Zhu |
CIKM | 4 |
| 2024 | Generative AI and Retrieval-Augmented Generation (RAG) Systems for EnterpriseabstractThis workshop introduces generative AI applications for enterprise, with a focus on retrieval-augmented generation (RAG) systems. Generative AI is a field of artificial intelligence that can create new content and solve complex problems. RAG systems are a novel generative AI technique that combines information retrieval with text generation to generate rich and diverse responses. RAG systems can leverage enterprise data, which is often specific, structured, and dynamic, to provide customized solutions for various domains. However, enterprise data also poses challenges such as scalability, security, and data quality. This workshop convenes researchers and practitioners to explore RAG and other generative AI systems in real-world enterprise scenarios, fostering knowledge exchange, collaboration, and identification of future directions. Relevant to the CIKM community, the workshop intersects with core areas of data science and machine learning, offering potential benefits across various domains. Anbang Xu, Min Du 0003, Pritam Gundecha, Xinliang Zhu, May Wang, Ping Li 0001 |
CIKM | 6 |
| 2024 | Bringing Multimodality to Amazon Visual Search SystemabstractImage to image matching has been well studied in the computer vision community. Previous studies mainly focus on training a deep metric learning model matching visual patterns between the query image and gallery images. In this study, we show that pure image-to- image matching suffers from false positives caused by matching to local visual patterns. To alleviate this issue, we propose to leverage recent advances in vision-language pretraining research. Specifically, we introduce additional image-text alignment losses into deep metric learning, which serve as constraints to the image-to-image matching loss. With additional alignments between the text (e.g., product title) and image pairs, the model can learn concepts from both modalities explicitly, which avoids matching low-level visual features. We progressively develop two variants, a 3-tower and a 4-tower model, where the latter takes one more short text query input. Through extensive experiments, we show that this change leads to a substantial improvement to the image to image matching problem. We further leveraged this model for multimodal search, which takes both image and reformulation text queries to improve search quality. Both offline and online experiments show strong improvements on the main metrics. Specifically, we see 4.95% relative improvement on image matching click through rate with the 3-tower model and 1.13% further improvement from the 4-tower model. Xinliang Zhu, Sheng-Wei Huang, Han Ding 0004, Kelvin Chen, Tal Neiman, Ouye Xie, Son Tran, Benjamin Z. Yao, Douglas Gray 0001, Anuj Bindal, Arnab Dhua |
KDD | 1 |
| 2024 | Multimodal Representation and Retrieval [MRR 2024]abstractMultimodal data is available in many applications like e-commerce production listings, social media posts and short videos. However, existing algorithms dealing with those types of data still focus on uni-modal representation learning by vision-language alignment and cross-modal retrieval. In this workshop, we target to bring a new retrieval problem where both queries and documents are multimodal. With the popularity of vision language modeling, large language models (LLMs), retrieval augmented generation (RAG), and multimodal LLM, we see a lot of new opportunities for multimodal representation and retrieval tasks. This event will be a comprehensive half-day workshop focusing on the subject of multimodal representation and retrieval. The agenda includes keynote speeches, oral presentations, and an interactive panel discussion. Xinliang Zhu, Arnab Dhua, Douglas Gray 0001, Ismet Zeki Yalniz, Mohamed Elhoseiny 0001, Bryan A. Plummer |
SIGIR | 1 |
| 2022 | Hierarchical Transformer for Survival Prediction Using Multimodality Whole Slide Images and GenomicsabstractLearning good representation of giga-pixel level whole slide pathology images (WSI) for downstream tasks is critical. Previous studies employ multiple instance learning (MIL) to represent WSIs as bags of sampled patches because, for most occasions, only slide-level labels are available, and only a tiny region of the WSI is disease-positive area. However, WSI representation learning still remains an open problem due to: (1) patch sampling on a higher resolution may be incapable of depicting microenvironment information such as the relative position between the tumor cells and surrounding tissues, while patches at lower resolution lose the fine-grained detail; (2) extracting patches from giant WSI results in large bag size, which tremendously increases the computational cost. To solve the problems, this paper proposes a hierarchical-based multimodal transformer framework that learns a hierarchical mapping between pathology images and corresponding genes. Precisely, we randomly extract instant-level patch features from WSIs with different magnification. Then a co-attention mapping between imaging and genomics is learned to uncover the pairwise interaction and reduce the space complexity of imaging features. Such early fusion makes it computationally feasible to use MIL Transformer for the survival prediction task. Our architecture requires fewer GPU resources compared with benchmark methods while maintaining better WSI representation ability. We evaluate our approach on five cancer types from the Cancer Genome Atlas database and achieved an average c-index of 0.673, outperforming the state-of-the-art multimodality methods. Chunyuan Li, Xinliang Zhu, Jiawen Yao, Junzhou Huang |
ICPR | 2 |
| 2022 | Hierarchical Proxy-based Loss for Deep Metric LearningabstractProxy-based metric learning losses are superior to pair-based losses due to their fast convergence and low training complexity. However, existing proxy-based losses focus on learning class-discriminative features while overlooking the commonalities shared across classes which are potentially useful in describing and matching samples. Moreover, they ignore the implicit hierarchy of categories in real-world datasets, where similar subordinate classes can be grouped together. In this paper, we present a framework that leverages this implicit hierarchy by imposing a hierarchical structure on the proxies and can be used with any existing proxy-based loss. This allows our model to capture both class-discriminative features and class-shared characteristics without breaking the implicit data hierarchy. We evaluate our method on five established image retrieval datasets such as In-Shop and SOP. Results demonstrate that our hierarchical proxy-based loss framework improves the performance of existing proxy-based losses, especially on large datasets which exhibit strong hierarchical structure. Zhibo Yang 0002, Muhammet Bastan, Xinliang Zhu, Douglas Gray 0001, Dimitris Samaras |
WACV | 3 |
| 2020 | WeightAln: Weighted Homologous Alignment for Protein Structure Property PredictionabstractAccurately predicting protein structure properties is essential in analyzing the structure and function of a protein, such as secondary structure, solvent accessibility, and dihedral angles. Multiple Sequence Alignment (MSA), which is a sequence alignment of multiple homologous protein sequences for the target protein, is widely used in the protein structure property prediction. The most popular strategy to exploit MSA is converting it into a position-specific scoring matrice (PSSM), then inputs the PSSM to the relevant prediction networks. PSSM is obtained by simply counting the frequency of amino acids presented at each position in the corresponding MSA, which means, each sequence in the MSA has the same weight to the target protein. However, simply setting the weights of homologous protein sequences of a protein as same cannot sufficiently model the complex relationships between them. Moreover, some sequences within the MSA are redundant, which raises a tantalizing question: can we generate a different weight for each sequence in the MSA and use the weighted PSSM to improve the performance of protein structure property prediction? To help answer this question, we present WeightAln framework, which to our knowledge, is the first attempt to generate learnable MSA weights for protein prediction tasks. We prove the effectiveness of our method by conducting extensive experiments on three protein structure property prediction tasks. Yuzhi Guo, Jiaxiang Wu 0001, Hehuan Ma, Xinliang Zhu, Junzhou Huang |
BIBM | 5 |
| 2020 | Label-Driven Reconstruction for Domain Adaptation in Semantic Segmentation
Weizhi An, Sheng Wang 0001, Xinliang Zhu, Chaochao Yan, Junzhou Huang |
ECCV (27) | 4 |
| 2020 | Whole slide images based cancer survival prediction using attention guided deep multiple instance learning networks
Jiawen Yao, Xinliang Zhu, Jitendra Jonnagaddala, Nicholas J. Hawkins, Junzhou Huang |
Medical Image Anal. | 2 |
| 2019 | Deep Multi-instance Learning for Survival Prediction from Whole Slide Images
Jiawen Yao, Xinliang Zhu, Junzhou Huang |
MICCAI (1) | 2 |
| 2018 | Graph CNN for Survival Analysis on Whole Slide Pathological Images
Ruoyu Li 0002, Jiawen Yao, Xinliang Zhu, Yeqing Li, Junzhou Huang |
MICCAI (2) | 3 |
| 2017 | WSISA: Making Survival Prediction from Whole Slide Histopathological ImagesabstractImage-based precision medicine techniques can be used to better treat cancer patients. However, the gigapixel resolution of Whole Slide Histopathological Images (WSIs) makes traditional survival models computationally impossible. These models usually adopt manually labeled discriminative patches from region of interests (ROIs) and are unable to directly learn discriminative patches from WSIs. We argue that only a small set of patches cannot fully represent the patients survival status due to the heterogeneity of tumor. Another challenge is that survival prediction usually comes with insufficient training patient samples. In this paper, we propose an effective Whole Slide Histopathological Images Survival Analysis framework (WSISA) to overcome above challenges. To exploit survival-discriminative patterns from WSIs, we first extract hundreds of patches from each WSI by adaptive sampling and then group these images into different clusters. Then we propose to train an aggregation model to make patient-level predictions based on cluster-level Deep Convolutional Survival (DeepConvSurv) prediction results. Different from existing state-of-the-arts image-based survival models which extract features using some patches from small regions of WSIs, the proposed framework can efficiently exploit and utilize all discriminative patterns in WSIs to predict patients survival status. To the best of our knowledge, this has not been shown before. We apply our method to the survival predictions of glioma and non-small-cell lung cancer using three datasets. Results demonstrate the proposed framework can significantly improve the prediction performance compared with the existing state-of-the-arts survival methods. Xinliang Zhu, Jiawen Yao, Feiyun Zhu, Junzhou Huang |
CVPR | 1 |
| 2017 | Deep Correlational Learning for Survival Prediction from Multi-modality Data
Jiawen Yao, Xinliang Zhu, Feiyun Zhu, Junzhou Huang |
MICCAI (2) | 2 |
| 2016 | Deep convolutional neural network for survival analysis with pathological imagesabstractTraditional Cox proportional hazard model for survival analysis are based on structured features like patients' sex, smoke years, BMI, etc. With the development of medical imaging technology, more and more unstructured medical images are available for diagnosis, treatment and survival analysis. Traditional survival models utilize these unstructured images by extracting human-designed features from them. However, we argue that those hand-crafted features have limited abilities in representing highly abstract information. In this paper, we for the first time develop a deep convolutional neural network for survival analysis (DeepConvSurv) with pathological images. The deep layers in our model could represent more abstract information compared with hand-crafted features from the images. Hence, it will improve the survival prediction performance. From our extensive experiments on the National Lung Screening Trial (NLST) lung cancer data, we show that the proposed DeepConvSurv model improves significantly compared with four state-of-the-art methods. Xinliang Zhu, Jiawen Yao, Junzhou Huang |
BIBM | 1 |
| 2016 | Imaging-genetic data mapping for clinical outcome prediction via supervised conditional Gaussian graphical modelabstractImaging-genetic data mapping is important for clinical outcome prediction like survival analysis. In this paper, we propose a supervised conditional Gaussian graphical model (SuperCGGM) to uncover survival associated mapping between pathological images and genetic data. The proposed method integrates heterogeneous modal data into the survival model by weighted projection within the data. To obtain a sparse solution, we employ l-1 regularization to the partial log likelihood loss function and propose a cyclic coordinate ascent algorithm to solve it. It also gives a way to bridge the gap between the supervised model with conditional Gaussian graphical model (CGGM). Compared to nine state-of-the-art methods like SuperPCA, CGGM, etc., our method is superior due to its ability of integrating diverse information from heterogeneous modal data in a supervised way. The extensive experiments also show the strong power of SuperCGGM in mapping survival associated image and gene expression signatures. Xinliang Zhu, Jiawen Yao, Guanghua Xiao, Jaime Rodriguez-Canales, Edwin R. Parra, Carmen Behrens, Ignacio I. Wistuba, Junzhou Huang |
BIBM | 1 |
| 2016 | Imaging Biomarker Discovery for Lung Cancer Survival Prediction
Jiawen Yao, Sheng Wang 0001, Xinliang Zhu, Junzhou Huang |
MICCAI (2) | 3 |
| 2015 | 10, 000+ Times Accelerated Robust Subset SelectionabstractSubset selection from massive data with noised information is increasingly popular for various applications. This problem is still highly challenging as current methods are generally slow in speed and sensitive to outliers. To address the above two issues, we propose an accelerated robust subset selection (ARSS) method. Extensive experiments on ten benchmark datasets verify that our method not only outperforms state of the art methods, but also runs 10,000+ times faster than the most related method. Feiyun Zhu, Bin Fan 0001, Xinliang Zhu, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
AAAI | 3 |