EDBT 2026 Demo / reviewers in the wild / expert
Yi Xu 0011
dblp:14/5580-11
· DBLP profile ↗
21ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0002-0604-8481ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Computer networks · 4 · 4 first-authorDatabases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMsabstractNicholas E. Corrado, Julian Katz-Samuels, Adithya M Devraj, Hyokun Yun, Chao Zhang, Yi Xu, Yi Pan, Bing Yin, Trishul Chilimbi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nicholas Corrado, Julian Katz-Samuels, Adithya M. Devraj, Hyokun Yun, Yi Xu 0011, Trishul Chilimbi |
ACL (1) | 6 |
| 2024 | Better Representations via Adversarial Training in Pre-Training: A Theoretical PerspectiveabstractPre-training is known to generate universal representations for downstream tasks in large-scale deep learning such as large language models. Existing literature, e.g., Kim et al. (2020), empirically observe that the downstream tasks can inherit the adversarial robustness of the pre-trained model. We provide theoretical justifications for this robustness inheritance phenomenon. Our theoretical results reveal that feature purification plays an important role in connecting the adversarial robustness of the pre-trained model and the downstream tasks in two-layer neural networks. Specifically, we show that (i) with adversarial training, each hidden node tends to pick only one (or a few) feature; (ii) without adversarial training, the hidden nodes can be vulnerable to attacks. This observation is valid for both supervised pre-training and contrastive learning. With purified nodes, it turns out that clean training is enough to achieve adversarial robustness in downstream tasks. Yue Xing 0002, Xiaofeng Lin 0005, Qifan Song, Yi Xu 0011, Belinda Zeng, Guang Cheng 0003 |
AISTATS | 4 |
| 2024 | Diffusion Models for Multi-Task Generative ModelingabstractDiffusion-based generative modeling has been achieving state-of-the-art results on various generation tasks. Most diffusion models, however, are limited to a single-generation modeling. Can we generalize diffusion models with the ability of multi-modal generative training for more generalizable modeling? In this paper, we propose a principled way to define a diffusion model by constructing a unified multi-modal diffusion model in a common {\em diffusion space}. We define the forward diffusion process to be driven by an information aggregation from multiple types of task-data, {\it e.g.}, images for a generation task and labels for a classification task. In the reverse process, we enforce information sharing by parameterizing a shared backbone denoising network with additional modality-specific decoder heads. Such a structure can simultaneously learn to generate different types of multi-modal data with a multi-task loss, which is derived from a new multi-modal variational lower bound that generalizes the standard diffusion model. We propose several multi-modal generation settings to verify our framework, including image transition, masked-image training, joint image-label and joint image-representation generative modeling. Extensive experimental results on ImageNet indicate the effectiveness of our framework for various multi-modal generative modeling, which we believe is an important research direction worthy of more future explorations. Changyou Chen, Han Ding 0004, Bunyamin Sisman, Yi Xu 0011, Ouye Xie, Benjamin Z. Yao, Son Dinh Tran, Belinda Zeng |
ICLR | 4 |
| 2024 | Robust Multi-Task Learning with Excess RisksabstractMulti-task learning (MTL) considers learning a joint model for multiple tasks by optimizing a convex combination of all task losses. To solve the optimization problem, existing methods use an adaptive weight updating scheme, where task weights are dynamically adjusted based on their respective losses to prioritize difficult tasks. However, these algorithms face a great challenge whenever label noise is present, in which case excessive weights tend to be assigned to noisy tasks that have relatively large Bayes optimal errors, thereby overshadowing other tasks and causing performance to drop across the board. To overcome this limitation, we propose Multi-Task Learning with Excess Risks (ExcessMTL), an excess risk-based task balancing method that updates the task weights by their distances to convergence instead. Intuitively, ExcessMTL assigns higher weights to worse-trained tasks that are further from convergence. To estimate the excess risks, we develop an efficient and accurate method with Taylor approximation. Theoretically, we show that our proposed algorithm achieves convergence guarantees and Pareto stationarity. Empirically, we evaluate our algorithm on various MTL benchmarks and demonstrate its superior performance over existing methods in the presence of label noise. Our code is available at https://github.com/yifei-he/ExcessMTL. Shiji Zhou, Hyokun Yun, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi, Han Zhao 0002 |
ICML | 5 |
| 2024 | Shopping MMLU: A Massive Multi-Task Online Shopping Benchmark for Large Language ModelsabstractOnline shopping is a complex multi-task, few-shot learning problem with a wide and evolving range of entities, relations, and tasks. However, existing models and benchmarks are commonly tailored to specific tasks, falling short of capturing the full complexity of online shopping. Large Language Models (LLMs), with their multi-task and few-shot learning abilities, have the potential to profoundly transform online shopping by alleviating task-specific engineering efforts and by providing users with interactive conversations. Despite the potential, LLMs face unique challenges in online shopping, such as domain-specific concepts, implicit knowledge, and heterogeneous user behaviors. Motivated by the potential and challenges, we propose Shopping MMLU, a diverse multi-task online shopping benchmark derived from real-world Amazon data. Shopping MMLU consists of 57 tasks covering 4 major shopping skills: concept understanding, knowledge reasoning, user behavior alignment, and multi-linguality, and can thus comprehensively evaluate the abilities of LLMs as general shop assistants. With Shoppping MMLU, we benchmark over 20 existing LLMs and uncover valuable insights about practices and prospects of building versatile LLM-based shop assistants. Shopping MMLU can be publicly accessed at https://github.com/KL4805/ShoppingMMLU. In addition, with Shopping MMLU, we are hosting a competition in KDD Cup 2024 with over 500 participating teams. The winning solutions and the associated workshop can be accessed at our website https://amazon-kddcup24.github.io/. Yilun Jin, Zheng Li 0018, Tianyu Cao 0001, Yifan Gao 0001, Pratik Jayarao, Xin Liu 0039, Ritesh Sarkhel, Xianfeng Tang, Wenju Xu, Jingfeng Yang 0001, Qingyu Yin, Priyanka Nigam, Yi Xu 0011, Kai Chen 0005, Qiang Yang 0001, Meng Jiang 0001 |
NeurIPS | 18 |
| 2023 | SST: Semantic and Structural Transformers for Hierarchy-aware Language Models in E-commerceabstractHierarchies are common structures used to organize data, such as e-commerce hierarchies associated with product data. With these product hierarchies, we aim to learn hierarchy-aware product text embeddings to improve fine-tuning performance on a variety of downstream e-commerce tasks. Existing methods leverage hierarchies by either aligning the text embeddings to separate hierarchical embeddings or by aligning the hierarchical information implicitly within a unified text Transformer. Although these models optimize to predict hierarchy information, performing further fine-tuning on new tasks is non-trivial. To bridge this gap, we propose a pre-training architecture to implicitly encode the hierarchy within the product text and then directly leverage a sub-set of the pre-training model during fine-tuning. Pre-training is done through Semantic and Structural Transformers (SST) where the Semantic-Transformer first encodes the product text into a contextual embedding, which is then used by the Structural-Transformer to infer the product’s path in the hierarchy. Fine-tuning is done using only the initial Semantic-Transformer, now that hierarchy-aware text embeddings are learned. With this design, we eliminate the need of linking each fine-tuning dataset with corresponding hierarchies. This leads to fine-tuning performance improvements on critical e-commerce downstream tasks over the existing state-of-the-art hierarchy models, even when hierarchy data $is$ available during fine-tuning. Moreover, this improvement is consistent even after augmenting our baseline models to support fine-tuning. We conclude by discussing how such implicit structural encodings can be leveraged beyond the e-commerce domain. Karan Samel, Houyu Zhang, Jun Ma 0029, Haoming Jiang, Qing Ping, Sheng Wang 0012, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi |
IEEE Big Data | 7 |
| 2023 | Unsupervised Multi-Modal Representation Learning for High Quality Retrieval of Similar Products at E-commerce ScaleabstractIdentifying similar products in e-commerce is useful in discovering relationships between products, making recommendations, and increasing diversity in search results. Product representation learning is the first step to define a generalized product similarity metric for search. The second step is to extend similarity search to a large scale (e.g., e-commerce catalog scale) without sacrificing quality. In this work, we present a solution that interweaves both steps, i.e., learn representations suited to high quality retrieval using contrastive learning (CL) and retrieve similar items from a large search space using approximate nearest neighbor search (ANNS) to trade-off quality for speed. We propose a CL training strategy for learning uni-modal encoders suited to multi-modal similarity search for e-commerce. We study ANNS retrieval by generating Pareto Frontiers (PFs) without requiring labels. Our CL training strategy doubles retrieval@1 metric across categories (e.g., from 36% to 88% in category C). We also demonstrate that ANNS engine optimization using PFs help select configurations appropriately (e.g., we achieve 6.8× search speed with just 2% drop from the maximum retrieval accuracy in medium size datasets). Kushal Kumar, Tarik Arici, Tal Neiman, Shioulin Sam, Yi Xu 0011, Hakan Ferhatosmanoglu, Ismail B. Tutar |
CIKM | 6 |
| 2023 | ForeSeer: Product Aspect Forecasting Using Temporal Graph EmbeddingabstractDeveloping text mining approaches to mine aspects from customer reviews has been well-studied due to its importance in understanding customer needs and product attributes. In contrast, it remains unclear how to predict the future emerging aspects of a new product that currently has little review information. This task, which we named product aspect forecasting, is critical for recommending new products, but also challenging because of the missing reviews. Here, we propose ForeSeer, a novel textual mining and product embedding approach progressively trained on temporal product graphs for this novel product aspect forecasting task. ForeSeer transfers reviews from similar products on a large product graph and exploits these reviews to predict aspects that might emerge in future reviews. A key novelty of our method is to jointly provide review, product, and aspect embeddings that are both time-sensitive and less affected by extremely imbalanced aspect frequencies. We evaluated ForeSeer on a real-world product review system containing 11,536,382 reviews and 11,000 products over 3 years. We observe that ForeSeer substantially outperformed existing approaches with at least 49.1% AUPRC improvement under the real setting where aspect associations are not given. ForeSeer further improves future link prediction on the product graph and the review aspect association prediction. Collectively, Foreseer offers a novel framework for review forecasting by effectively integrating review text, product network, and temporal information, opening up new avenues for online shopping recommendation and e-commerce applications. Zixuan Liu 0001, Gaurush Hiranandani, Kun Qian 0018, Edward W. Huang, Yi Xu 0011, Belinda Zeng, Karthik Subbian, Sheng Wang 0012 |
CIKM | 5 |
| 2023 | Understanding and Constructing Latent Modality Structures in Multi-Modal Representation LearningabstractContrastive loss has been increasingly used in learning representations from multiple modalities. In the limit, the nature of the contrastive loss encourages modalities to exactly match each other in the latent space. Yet it remains an open question how the modality alignment affects the downstream task performance. In this paper, based on an information-theoretic argument, we first prove that exact modality alignment is sub-optimal in general for down-stream prediction tasks. Hence we advocate that the key of better performance lies in meaningful latent modality structures instead of perfect modality alignment. To this end, we propose three general approaches to construct latent modality structures. Specifically, we design 1) a deep feature separation loss for intra-modality regularization; 2) a Brownian-bridge loss for inter-modality regularization; and 3) a geometric consistency loss for both intra- and intermodality regularization. Extensive experiments are conducted on two popular multi-modal representation learning frameworks: the CLIP-based two-tower model and the ALBEF-based fusion model. We test our model on a variety of tasks including zero/few-shot image classification, image-text retrieval, visual question answering, visual reasoning, and visual entailment. Our method achieves consistent improvements over existing methods, demonstrating the effectiveness and generalizability of our proposed approach on latent modality structure regularization. Changyou Chen, Han Zhao 0002, Liqun Chen 0001, Qing Ping, Son Dinh Tran, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi |
CVPR | 7 |
| 2023 | Graph-Aware Language Model Pre-Training on a Large Graph Corpus Can Help Multiple Graph ApplicationsabstractModel pre-training on large text corpora has been demonstrated effective for various downstream applications in the NLP domain. In the graph mining domain, a similar analogy can be drawn for pre-training graph models on large graphs in the hope of benefiting downstream graph applications, which has also been explored by several recent studies. However, no existing study has ever investigated the pre-training of text plus graph models on large heterogeneous graphs with abundant textual information (a.k.a. large graph corpora) and then fine-tuning the model on different related downstream applications with different graph schemas. To address this problem, we propose a framework of graph-aware language model pre-training (GaLM) on a large graph corpus, which incorporates large language models and graph neural networks, and a variety of fine-tuning methods on downstream applications. We conduct extensive experiments on Amazon's real internal datasets and large public datasets. Comprehensive empirical results and in-depth analysis demonstrate the effectiveness of our proposed methods along with lessons learned. Da Zheng 0004, Jun Ma 0029, Houyu Zhang, Vassilis N. Ioannidis, Xiang Song 0003, Qing Ping, Sheng Wang 0012, Carl Yang 0001, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi |
KDD | 10 |
| 2022 | Multi-modal Alignment using Representation CodebookabstractAligning signals from different modalities is an important step in vision-language representation learning as it affects the performance of later stages such as cross-modality fusion. Since image and text typically reside in different regions of the feature space, directly aligning them at instance level is challenging especially when features are still evolving during training. In this paper, we propose to align at a higher and more stable level using cluster representation. Specifically, we treat image and text as two “views” of the same entity, and encode them into a joint vision-language coding space spanned by a dictionary of cluster centers (codebook). We contrast positive and negative samples via their cluster assignments while simultaneously optimizing the cluster centers. To further smooth out the learning process, we adopt a teacher-student distillation paradigm, where the momentum teacher of one view guides the student learning of the other. We evaluated our approach on common vision language benchmarks and obtain new SoTA on zero-shot cross modality retrieval while being competitive on various other transfer tasks. Jiali Duan, Liqun Chen 0001, Son Tran, Yi Xu 0011, Belinda Zeng, Trishul Chilimbi |
CVPR | 5 |
| 2022 | Vision-Language Pre-Training with Triple Contrastive LearningabstractVision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information (MI) between an image and its matched text. However, simply performing cross-modal alignment (CMA) ignores data potential within each modality, which may result in degraded representations. For instance, although CMA-based models are able to map image-text pairs close together in the embedding space, they fail to ensure that similar inputs from the same modality stay close by. This problem can get even worse when the pre-training data is noisy. In this paper, we propose triple contrastive learning (TCL) for vision-language pre-training by leveraging both cross-modal and intra-modal self-supervision. Besides CMA, TCL introduces an intra-modal contrastive objective to provide complementary benefits in representation learning. To take advantage of localized and structural information from image and text input, TCL further maximizes the average MI between local regions of image/text and their global summary. To the best of our knowledge, ours is the first work that takes into account local structure information for multi-modality representation learning. Experimental evaluations show that our approach is competitive and achieves the new state of the art on various common downstream vision-language tasks such as image-text retrieval and visual question answering. Jiali Duan, Son Tran, Yi Xu 0011, Sampath Chanda, Liqun Chen 0001, Belinda Zeng, Trishul Chilimbi, Junzhou Huang |
CVPR | 4 |
| 2022 | Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias PerspectiveabstractContrastive learning (CL) has been the de facto technique for self-supervised representation learning (SSL), with impressive empirical success such as multi-modal representation learning. However, traditional CL loss only considers negative samples from a minibatch, which could cause biased gradients due to the non-decomposibility of the loss. For the first time, we consider optimizing a more generalized contrastive loss, where each data sample is associated with an infinite number of negative samples. We show that directly using minibatch stochastic optimization could lead to gradient bias. To remedy this, we propose an efficient Bayesian data augmentation technique to augment the contrastive loss into a decomposable one, where standard stochastic optimization can be directly applied without gradient bias. Specifically, our augmented loss defines a joint distribution over the model parameters and the augmented parameters, which can be conveniently optimized by a proposed stochastic expectation-maximization algorithm. Our framework is more general and is related to several popular SSL algorithms. We verify our framework on both small scale models and several large foundation models, including SSL of ImageNet and SSL for vision-language representation learning. Experiment results indicate the existence of gradient bias in all cases, and demonstrate the effectiveness of the proposed method on improving previous state of the arts. Remarkably, our method can outperform the strong MoCo-v3 under the same hyper-parameter setting with only around half of the minibatch size; and also obtains strong results in the recent public benchmark ELEVATER for few-shot image classification. Changyou Chen, Yi Xu 0011, Liqun Chen 0001, Jiali Duan, Yiran Chen 0001, Son Tran, Belinda Zeng, Trishul Chilimbi |
NeurIPS | 3 |
| 2020 | CAM: Uninteresting Speech Detector
Weiyi Lu, Yi Xu 0011, Belinda Zeng |
INTERSPEECH | 2 |
| 2014 | User grouping and scheduling for large scale MIMO systems with two-stage precodingabstractIn this paper, we consider the design of user grouping and scheduling for large-scale multiple-input multiple-output (MIMO) frequency-division-duplexing (FDD) systems. Based on a recently proposed two-stage precoding framework, we first propose an improved K-means user grouping scheme which allocates the users to different pre-beamforming groups using the second-order channel statistics, and then a user grouping scheme that considers both load balancing and precoding design. After user groups are so determined, we present a dynamic user scheduling scheme where second-stage precoding is designed based on instantaneous channel conditions. We demonstrate the efficacy of the proposed schemes through simulations. Yi Xu 0011, Guosen Yue, Narayan Prasad, Sampath Rangarajan, Shiwen Mao |
ICC | 1 |
| 2014 | Relay-Assisted Multiuser Video Streaming in Cognitive Radio NetworksabstractDue to the drastic increase in wireless video traffic, the capacity of the existing and future wireless networks will be greatly stressed, while interference will become the dominant capacity-limiting factor. In this paper, we investigate relay-assisted downlink multiuser video streaming in a cognitive radio (CR) cellular network. We incorporate zero-forcing precoding to allow transmitters collaboratively send encoded (mixed) signals to all CR users, such that undesired signals will be canceled and the desired signal can be decoded at each CR user. We present a stochastic programming formulation of the problem, as well as a problem reformulation that greatly reduces computational complexity. In the cases of a single licensed channel and multiple licensed channels with channel bonding, we develop an optimal distributed algorithm with proven convergence and convergence speed. In the case of multiple channels without channel bonding, we develop a greedy algorithm with a proven performance bound. The algorithms are evaluated with simulations and are shown to achieve considerable gains over two heuristic schemes. Yi Xu 0011, Donglin Hu, Shiwen Mao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Distributed Interference Alignment in Cognitive Radio NetworksabstractIn this paper, we investigate the problem of incorporating two advanced physical layer technologies, i.e., multiple-input and multiple- output (MIMO) and distributed interference alignment, in cognitive radio (CR) networks. We present a cooperative spectrum leasing scheme for primary and secondary users to trade off between data transmission and revenue collection/payment. A Stackelberg game is formulated, where the primary user is the leader and the secondary users are followers. With backward induction, we derive the unique Stackelberg Equilibrium, where no player can gain by unilaterally changing strategy, as well as the optimal strategies. We find spectrum leasing is always beneficial to enhance the utilities of primary and secondary users. The proposed scheme outperforms a no-spectrum-leasing scheme and a cooperative scheme presented in the literature with considerable gains, which demonstrate the benefits of spectrum leasing and distributed interference alignment and validate the efficacy of the proposed scheme. Yi Xu 0011, Shiwen Mao |
ICCCN | 1 |
| 2012 | On interference alignment in multi-user OFDM systemsabstractMulti-user Orthogonal Frequency Division Multiplexing (OFDM) have been widely adopted to combat the detrimental effects of wireless channels and enhance system throughput. Recently, interference alignment is proposed to exploit interference to enable concurrent transmissions of multiple signals. In this paper, we investigate how to incorporate interference alignment in multi-user OFDM systems. We first reveal the unique characteristics and challenges brought about by using interference alignment in diagonal channels. We then derive a performance bound for the multi-user OFDM/interference alignment system under practical constraints (i.e., a finite number of subcarriers), and show how to achieve this bound with a decomposition approach. The superior performance of the proposed scheme is validated with simulations. Yi Xu 0011, Shiwen Mao |
GLOBECOM | 1 |
| 2012 | On adopting Interleave Division Multiple Access in two-tier femtocell networks: The uplink caseabstractA femtocell base station (FBS) is designed to cater for the demand of ever-increasing wireless data traffic, typically in the indoor environment. Among the many technical problems, interference management is particularly a challenging one for fully harvesting the high potential of femtocell networks. In this paper, we address the interference management problem with an iterative multi-user detection approach, and propose to adopt Interleave Division Multiple Access (IDMA) for the uplink of two-tier femtocell networks by exploiting the processing capability of FBS. We consider three IDMA-based schemes, namely, FBS Decode, FBS Forward, and FBS Select, and evaluate their performance with simulations. Numerical results show that the proposed schemes achieve considerable throughput gain over traditional techniques and are highly suited for the uplink of two-tier femtocell networks. Yi Xu 0011, Shiwen Mao, Xin Su 0001 |
ICC | 1 |
| 2010 | Cooperative Spectrum Sensing in Cognitive Radio under Noise UncertaintyabstractCooperative energy spectrum sensing has been proved effective to detect the spectrum holes in Cognitive Radio (CR). However, its performance may suffer from the noise uncertainty, which is portrayed by the SNR wall in some literatures. In this paper we analyze the spectrum sensing performance under noise uncertainty and propose a new approach to obtain the SNR wall. In addition, a suboptimal liner cooperative sensing algorithm with wavelet denoising is proposed to reduce the impact of noise uncertainty. Analysis and numerical results show that cooperative sensing and wavelet denoising can significantly improve the sensing performance under noise uncertainty condition. Yi Xu 0011, Xin Su 0001, Jing Wang 0001 |
VTC Spring | 2 |
| 2010 | Cooperative Spectrum Sensing with Wavelet Denoising in Cognitive RadioabstractCooperative energy spectrum sensing has been proved effective to detect the spectrum holes in Cognitive Radio (CR). However, few studies make mention of wavelet transform based signal processing before energy detection. In this paper, a novel cooperative energy spectrum sensing algorithm with 1-D or 2-D wavelet denoising is proposed. The 1-D wavelet denoising is performed at each sensing node, while the 2-D wavelet denoising is implemented at the sensing station. Simulation results show that the spectrum sensing performance is improved significantly by the proposed cooperative spectrum sensing algorithm with wavelet denoising as opposed to the conventional method. Yi Xu 0011, Xin Su 0001, Jing Wang 0001 |
VTC Spring | 2 |