EDBT 2026 Demo / reviewers in the wild / expert
Xiangning Chen
dblp:56/7393
· DBLP profile ↗
26ranked-venue papers
8as first author
16since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorComputer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Red Teaming Language Model Detectors with Language ModelsabstractAbstract The prevalence and strong capability of large language models (LLMs) present significant safety and ethical risks if exploited by malicious users. To prevent the potentially deceptive usage of LLMs, recent work has proposed algorithms to detect LLM-generated text and protect LLMs. In this paper, we investigate the robustness and reliability of these LLM detectors under adversarial attacks. We study two types of attack strategies: 1) replacing certain words in an LLM’s output with their synonyms given the context; 2) automatically searching for an instructional prompt to alter the writing style of the generation. In both strategies, we leverage an auxiliary LLM to generate the word replacements or the instructional prompt. Different from previous works, we consider a challenging setting where the auxiliary LLM can also be protected by a detector. Experiments reveal that our attacks effectively compromise the performance of all detectors in the study with plausible generations, underscoring the urgent need to improve the robustness of LLM-generated text detection systems. Code is available at https://github.com/shizhouxing/LLM-Detector-Robustness. Zhouxing Shi, Fan Yin, Xiangning Chen, Kai-Wei Chang 0001, Cho-Jui Hsieh |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | Symbol tuning improves in-context learning in language modelsabstractJerry Wei, Le Hou, Andrew Lampinen, Xiangning Chen, Da Huang, Yi Tay, Xinyun Chen, Yifeng Lu, Denny Zhou, Tengyu Ma, Quoc Le. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jerry W. Wei, Le Hou, Andrew K. Lampinen, Xiangning Chen, Yi Tay, Yifeng Lu, Denny Zhou, Tengyu Ma 0001, Quoc V. Le |
EMNLP | 4 |
| 2023 | Symbolic Discovery of Optimization AlgorithmsabstractWe present a method to formulate algorithm discovery as program search, and apply it to discover optimization algorithms for deep neural network training. We leverage efficient search techniques to explore an infinite and sparse program space. To bridge the large generalization gap between proxy and target tasks, we also introduce program selection and simplification strategies.
Our method discovers a simple and effective optimization algorithm, $\textbf{Lion}$ ($\textit{Evo$\textbf{L}$ved S$\textbf{i}$gn M$\textbf{o}$me$\textbf{n}$tum}$). It is more memory-efficient than Adam as it only keeps track of the momentum. Different from adaptive optimizers, its update has the same magnitude for each parameter calculated through the sign operation.
We compare Lion with widely used optimizers, such as Adam and Adafactor, for training a variety of models on different tasks. On image classification, Lion boosts the accuracy of ViT by up to 2\% on ImageNet and saves up to 5x the pre-training compute on JFT. On vision-language contrastive learning, we achieve 88.3\% $\textit{zero-shot}$ and 91.1\% $\textit{fine-tuning}$ accuracy on ImageNet, surpassing the previous best results by 2\% and 0.1\%, respectively. On diffusion models, Lion outperforms Adam by achieving a better FID score and reducing the training compute by up to 2.3x. For autoregressive, masked language modeling, and fine-tuning, Lion exhibits a similar or better performance compared to Adam. Our analysis of Lion reveals that its performance gain grows with the training batch size. It also requires a smaller learning rate than Adam due to the larger norm of the update produced by the sign function. Additionally, we examine the limitations of Lion and identify scenarios where its improvements are small or not statistically significant. Xiangning Chen, Esteban Real, Hieu Pham 0001, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, Quoc V. Le |
NeurIPS | 1 |
| 2023 | Why Does Sharpness-Aware Minimization Generalize Better Than SGD?abstractThe challenge of overfitting, in which the model memorizes the training data and fails to generalize to test data, has become increasingly significant in the training of large neural networks. To tackle this challenge, Sharpness-Aware Minimization (SAM) has emerged as a promising training method, which can improve the generalization of neural networks even in the presence of label noise. However, a deep understanding of how SAM works, especially in the setting of nonlinear neural networks and classification tasks, remains largely missing. This paper fills this gap by demonstrating why SAM generalizes better than Stochastic Gradient Descent (SGD) for a certain data model and two-layer convolutional ReLU networks. The loss landscape of our studied problem is nonsmooth, thus current explanations for the success of SAM based on the Hessian information are insufficient. Our result explains the benefits of SAM, particularly its ability to prevent noise learning in the early stages, thereby facilitating more effective learning of features. Experiments on both synthetic and real data corroborate our theory. Zixiang Chen, Yiwen Kou, Xiangning Chen, Cho-Jui Hsieh, Quanquan Gu |
NeurIPS | 4 |
| 2022 | Towards Efficient and Scalable Sharpness-Aware MinimizationabstractRecently, Sharpness-Aware Minimization (SAM), which connects the geometry of the loss landscape and generalization, has demonstrated a significant performance boost on training large-scale models such as vision transformers. However, the update rule of SAM requires two sequential (non-parallelizable) gradient computations at each step, which can double the computational overhead. In this paper, we propose a novel algorithm LookSAM - that only periodically calculates the inner gradient ascent, to significantly reduce the additional training cost of SAM. The empirical results illustrate that LookSAM achieves similar accuracy gains to SAM while being tremendously faster - it enjoys comparable computational complexity with first-order optimizers such as SGD or Adam. To further evaluate the performance and scalability of LookSAM, we incorporate a layer-wise modification and perform experiments in the large-batch training scenario, which is more prone to converge to sharp local minima. Equipped with the proposed algorithms, we are the first to successfully scale up the batch size when training Vision Transformers (ViTs). With a 64k batch size, we are able to train ViTs from scratch in minutes while maintaining competitive performance. The code is available here: https://github.com/yong-6/LookSAM Yong Liu 0020, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh, Yang You 0001 |
CVPR | 3 |
| 2022 | When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
Xiangning Chen, Cho-Jui Hsieh, Boqing Gong |
ICLR | 1 |
| 2022 | Concurrent Adversarial Learning for Large-Batch Training
Yong Liu 0020, Xiangning Chen, Minhao Cheng, Cho-Jui Hsieh, Yang You 0001 |
ICLR | 2 |
| 2022 | Learning to Schedule Learning rate with Graph Neural Networks
Yuanhao Xiong, Li-Cheng Lan, Xiangning Chen, Cho-Jui Hsieh |
ICLR | 3 |
| 2022 | Random Sharpness-Aware MinimizationabstractCurrently, Sharpness-Aware Minimization (SAM) is proposed to seek the parameters that lie in a flat region to improve the generalization when training neural networks. In particular, a minimax optimization objective is defined to find the maximum loss value centered on the weight, out of the purpose of simultaneously minimizing loss value and loss sharpness. For the sake of simplicity, SAM applies one-step gradient ascent to approximate the solution of the inner maximization. However, one-step gradient ascent may not be sufficient and multi-step gradient ascents will cause additional training costs. Based on this observation, we propose a novel random smoothing based SAM (R-SAM) algorithm. To be specific, R-SAM essentially smooths the loss landscape, based on which we are able to apply the one-step gradient ascent on the smoothed weights to improve the approximation of the inner maximization. Further, we evaluate our proposed R-SAM on CIFAR and ImageNet datasets. The experimental results illustrate that R-SAM can consistently improve the performance on ResNet and Vision Transformer (ViT) training. Yong Liu 0020, Siqi Mai, Minhao Cheng, Xiangning Chen, Cho-Jui Hsieh, Yang You 0001 |
NeurIPS | 4 |
| 2022 | 2.5D visual relationship detection
Yu-Chuan Su, Soravit Changpinyo, Xiangning Chen, Sathish Thoppay, Cho-Jui Hsieh, Lior Shapira, Radu Soricut, Hartwig Adam, Matthew Brown 0001, Ming-Hsuan Yang 0001, Boqing Gong |
Comput. Vis. Image Underst. | 3 |
| 2022 | Cross-domain Recommendation with Bridge-Item EmbeddingsabstractWeb systems that provide the same functionality usually share a certain amount of items. This makes it possible to combine data from different websites to improve recommendation quality, known as the cross-domain recommendation task. Despite many research efforts on this task, the main drawback is that they largely assume the data of different systems can be fully shared . Such an assumption is unrealistic different systems are typically operated by different companies, and it may violate business privacy policy to directly share user behavior data since it is highly sensitive. In this work, we consider a more practical scenario to perform cross-domain recommendation. To avoid the leak of user privacy during the data sharing process, we consider sharing only the information of the item side, rather than user behavior data. Specifically, we transfer the item embeddings across domains, making it easier for two companies to reach a consensus (e.g., legal policy) on data sharing since the data to be shared is user-irrelevant and has no explicit semantics. To distill useful signals from transferred item embeddings, we rely on the strong representation power of neural networks and develop a new method named as NATR (short for N eural A ttentive T ransfer R ecommendation ). We perform extensive experiments on two real-world datasets, demonstrating that NATR achieves similar or even better performance than traditional cross-domain recommendation methods that directly share user-relevant data. Further insights are provided on the efficacy of NATR in using the transferred item embeddings to alleviate the data sparsity issue. Chen Gao 0001, Yong Li 0008, Fuli Feng, Xiangning Chen, Xiangnan He 0001, Depeng Jin |
ACM Trans. Knowl. Discov. Data | 4 |
| 2021 | Robust and Accurate Object Detection via Adversarial LearningabstractData augmentation has become a de facto component for training high-performance deep image classifiers, but its potential is under-explored for object detection. Noting that most state-of-the-art object detectors benefit from fine-tuning a pre-trained classifier, we first study how the classifiers’ gains from various data augmentations transfer to object detection. The results are discouraging; the gains diminish after fine-tuning in terms of either accuracy or robustness. This work instead augments the fine-tuning stage for object detectors by exploring adversarial examples, which can be viewed as a model-dependent data augmentation. Our method dynamically selects the stronger adversarial images sourced from a detector’s classification and localization branches and evolves with the detector to ensure the augmentation policy stays current and relevant. This model-dependent augmentation generalizes to different object detectors better than AutoAugment, a model-agnostic augmentation policy searched based on one particular detector. Our approach boosts the performance of state-of-the-art EfficientDets by +1.1 mAP on the COCO object detection benchmark. It also improves the detectors’ robustness against natural distortions by +3.8 mAP and against domain shift by +1.3 mAP. Xiangning Chen, Cihang Xie, Mingxing Tan, Li Zhang 0003, Cho-Jui Hsieh, Boqing Gong |
CVPR | 1 |
| 2021 | RANK-NOSH: Efficient Predictor-Based Architecture Search via Non-Uniform Successive HalvingabstractPredictor-based algorithms have achieved remarkable performance in the Neural Architecture Search (NAS) tasks. However, these methods suffer from high computation costs, as training the performance predictor usually requires training and evaluating hundreds of architectures from scratch. Previous works along this line mainly focus on reducing the number of architectures required to fit the predictor. In this work, we tackle this challenge from a different perspective - improve search efficiency by cutting down the computation budget of architecture training. We propose NOn-uniform Successive Halving (NOSH), a hierarchical scheduling algorithm that terminates the training of underperforming architectures early to avoid wasting budget. To effectively leverage the non-uniform supervision signals produced by NOSH, we formulate predictor-based architecture search as learning to rank with pairwise comparisons. The resulting method - RANK-NOSH, reduces the search budget by ~ 5× while achieving competitive or even better performance than previous state-of-the-art predictor-based methods on various spaces and datasets. Xiangning Chen, Minhao Cheng, Xiaocheng Tang, Cho-Jui Hsieh |
ICCV | 2 |
| 2021 | DrNAS: Dirichlet Neural Architecture Search
Xiangning Chen, Minhao Cheng, Xiaocheng Tang, Cho-Jui Hsieh |
ICLR | 1 |
| 2021 | Rethinking Architecture Selection in Differentiable NAS
Minhao Cheng, Xiangning Chen, Xiaocheng Tang, Cho-Jui Hsieh |
ICLR | 3 |
| 2021 | Learning to Recommend With Multiple Cascading BehaviorsabstractMost existing recommender systems leverage user behavior data of one type only, such as the purchase behavior in E-commerce that is directly related to the business Key Performance Indicator (KPI) of conversion rate. Besides the key behavioral data, we argue that other forms of user behaviors also provide valuable signal, such as views, clicks, adding a product to shopping carts and so on. They should be taken into account properly to provide quality recommendation for users. In this work, we contribute a new solution named short for Neural Multi-Task Recommendation (NMTR) for learning recommender systems from user multi-behavior data. We develop a neural network model to capture the complicated and multi-type interactions between users and items. In particular, our model accounts for the cascading relationship among different types of behaviors (e.g., a user must click on a product before purchasing it). To fully exploit the signal in the data of multiple types of behaviors, we perform a joint optimization based on the multi-task learning framework, where the optimization on a behavior is treated as a task. Extensive experiments on two real-world datasets demonstrate that NMTR significantly outperforms state-of-the-art recommender systems that are designed to learn from both single-behavior data and multi-behavior data. Further analysis shows that modeling multiple behaviors is particularly useful for providing recommendation for sparse users that have very few interactions. Chen Gao 0001, Xiangnan He 0001, Dahua Gan, Xiangning Chen, Fuli Feng, Yong Li 0008, Tat-Seng Chua, Lina Yao 0001, Yang Song 0001, Depeng Jin |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Stabilizing Differentiable Architecture Search via Perturbation-based RegularizationabstractDifferentiable architecture search (DARTS) is a prevailing NAS solution to identify architectures. Based on the continuous relaxation of the architecture space, DARTS learns a differentiable architecture weight and largely reduces the search cost. However, its stability has been challenged for yielding deteriorating architectures as the search proceeds. We find that the precipitous validation loss landscape, which leads to a dramatic performance drop when distilling the final architecture, is an essential factor that causes instability. Based on this observation, we propose a perturbation-based regularization - SmoothDARTS (SDARTS), to smooth the loss landscape and improve the generalizability of DARTS-based methods. In particular, our new formulations stabilize DARTS-based methods by either random smoothing or adversarial attack. The search trajectory on NAS-Bench-1Shot1 demonstrates the effectiveness of our approach and due to the improved stability, we achieve performance gain across various search spaces on 4 datasets. Furthermore, we mathematically show that SDARTS implicitly regularizes the Hessian norm of the validation loss, which accounts for a smoother loss landscape and improved performance. Xiangning Chen, Cho-Jui Hsieh |
ICML | 1 |
| 2020 | Efficient Neural Interaction Function Search for Collaborative FilteringabstractIn collaborative filtering (CF), interaction function (IFC) play the important role of capturing interactions among items and users. The most popular IFC is the inner product, which has been successfully used in low-rank matrix factorization. However, interactions in real-world applications can be highly complex. Thus, other operations (such as plus and concatenation), which may potentially offer better performance, have been proposed. Nevertheless, it is still hard for existing IFCs to have consistently good performance across different application scenarios. Motivated by the recent success of automated machine learning (AutoML), we propose in this paper the search for simple neural interaction functions (SIF) in CF. By examining and generalizing existing CF approaches, an expressive SIF search space is designed and represented as a structured multi-layer perceptron. We propose an one-shot search algorithm that simultaneously updates both the architecture and learning parameters. Experimental results demonstrate that the proposed method can be much more efficient than popular AutoML approaches, can obtain much better prediction performance than state-of-the-art CF approaches, and can discover distinct IFCs for different data sets and tasks.1 Quanming Yao, Xiangning Chen, James T. Kwok, Yong Li 0008, Cho-Jui Hsieh |
WWW | 2 |
| 2019 | Neural Multi-task Recommendation from Multi-behavior DataabstractMost existing recommender systems leverage user behavior data of one type, such as the purchase behavior data in E-commerce. We argue that other types of user behavior data also provide valuable signal, such as views, clicks, and so on. In this work, we contribute a new solution named NMTR (short for Neural Multi-Task Recommendation) for learning recommender systems from user multi-behavior data. In particular, our model accounts for the cascading relationship among different types of behaviors (e.g., a user must click on a product before purchasing it). We perform a joint optimization based on the multi-task learning framework, where the optimization on a behavior is treated as a task. Extensive experiments on the real-world dataset demonstrate that NMTR significantly outperforms state-of-the-art recommender systems that are designed to learn from both single-behavior data and multi-behavior data. Chen Gao 0001, Xiangnan He 0001, Dahua Gan, Xiangning Chen, Fuli Feng, Yong Li 0008, Tat-Seng Chua, Depeng Jin |
ICDE | 4 |
| 2019 | Neural Feature Search: A Neural Architecture for Automated Feature EngineeringabstractFeature engineering is a crucial step for developing effective machine learning models. Traditionally, feature engineering is performed manually, which requires much domain knowledge and is time-consuming. In recent years, many automated feature engineering methods have been proposed. These methods improve the accuracy of a machine learning model by automatically transforming the original features into a set of new features. However, existing methods either lack ability to perform high-order transformations or suffer from the feature space explosion problem. In this paper, we present Neural Feature Search (NFS), a novel neural architecture for automated feature engineering. We utilize a recurrent neural network based controller to transform each raw feature through a series of transformation functions. The controller is trained through reinforcement learning to maximize the expected performance of the machine learning algorithm. Extensive experiments on public datasets illustrate that our neural architecture is effective and outperforms the existing state-of-the-art automated feature engineering methods. Our architecture can efficiently capture potentially valuable high-order transformations and mitigate the feature explosion problem. Xiangning Chen, Bo Qiao 0001, Wei Wu 0011, Murali Chintalapati, Dongmei Zhang 0001, Qingwei Lin, Chuan Luo 0002, Hongyu Zhang 0002, Yong Xu 0010, Yingnong Dang, Kaixin Sui, Xu Zhang 0024 |
ICDM | 1 |
| 2019 | DeepAPF: Deep Attentive Probabilistic Factorization for Multi-site Video RecommendationabstractExisting web video systems recommend videos according to users' viewing history from its own website. However, since many users watch videos in multiple websites, this approach fails to capture these users' interests across sites. In this paper, we investigate the user viewing behavior in multiple sites based on a large scale real dataset. We find that user interests are comprised of cross-site consistent part and site-specific part with different degrees of the importance. Existing linear matrix factorization recommendation model has limitation in modeling such complicated interactions. Thus, we propose a model of Deep Attentive Probabilistic Factorization (DeepAPF) to exploit deep learning method to approximate such complex user-video interaction. DeepAPF captures both cross-site common interests and site-specific interests with non-uniform importance weights learned by the attentional network. Extensive experiments show that our proposed model outperforms by 17.62%, 7.9% and 8.1% with the comparison of three state-of-the-art baselines. Our study provides insight to integrate user viewing records from multiple sites via the trusted third party, which gains mutual benefits in video recommendation. Huan Yan 0003, Xiangning Chen, Chen Gao 0001, Yong Li 0008, Depeng Jin |
IJCAI | 2 |
| 2019 | Cross-domain Recommendation Without Sharing User-relevant DataabstractWeb systems that provide the same functionality usually share a certain amount of items. This makes it possible to combine data from different websites to improve recommendation quality, known as the cross-domain recommendation task. Despite many research efforts on this task, the main drawback is that they largely assume the data of different systems can be fully shared. Such an assumption is unrealistic - different systems are typically operated by different companies, and it may violate business privacy policy to directly share user behavior data since it is highly sensitive. Chen Gao 0001, Xiangning Chen, Fuli Feng, Xiangnan He 0001, Yong Li 0008, Depeng Jin |
WWW | 2 |
| 2017 | Research on RBF neural network in simulation of MBR membrane pollution simulationabstractMembrane pollution is the main obstacle to the popularization and application of MBR. In order to solve the problem that the influence factors of membrane fouling are more complicated, three kinds of membrane fouling factors with the contribution rate of more than 95% are selected by principal component analysis(PCA) method: The mixed solution suspended solids (MLSS), operating pressure (ΔP) and temperature (T) . The three influencing factors of MBR membrane were simulated and the membrane flux was used as output parameter. The predictive model of membrane fouling based on RBF neural network was established to realize the predictive control of membrane fouling. The whole experimental process has certain theoretical value and practical significance, and it should play an active role in guiding the actual project of MBR. Xiangning Chen, Chunqing Li |
ICIS | 1 |
| 2009 | A multi-dimensional evidence-based candidate gene prioritization approach for complex diseases-schizophrenia as a caseabstractMOTIVATION: During the past decade, we have seen an exponential growth of vast amounts of genetic data generated for complex disease studies. Currently, across a variety of complex biological problems, there is a strong trend towards the integration of data from multiple sources. So far, candidate gene prioritization approaches have been designed for specific purposes, by utilizing only some of the available sources of genetic studies, or by using a simple weight scheme. Specifically to psychiatric disorders, there has been no prioritization approach that fully utilizes all major sources of experimental data. RESULTS: Here we present a multi-dimensional evidence-based candidate gene prioritization approach for complex diseases and demonstrate it in schizophrenia. In this approach, we first collect and curate genetic studies for schizophrenia from four major categories: association studies, linkage analyses, gene expression and literature search. Genes in these data sets are initially scored by category-specific scoring methods. Then, an optimal weight matrix is searched by a two-step procedure (core genes and unbiased P-values in independent genome-wide association studies). Finally, genes are prioritized by their combined scores using the optimal weight matrix. Our evaluation suggests this approach generates prioritized candidate genes that are promising for further analysis or replication. The approach can be applied to other complex diseases. AVAILABILITY: The collected data, prioritized candidate genes, and gene prioritization tools are freely available at http://bioinfo.mc.vanderbilt.edu/SZGR/. Jingchun Sun, Peilin Jia, Ayman H. Fanous, Bradley Todd Webb, Edwin J. C. G. van den Oord, Xiangning Chen, József Bukszár, Kenneth S. Kendler, Zhongming Zhao |
Bioinform. | 6 |
| 2001 | The cell structure of the next generation ATMabstractThe purpose of ATM is to provide a high-speed, low-delay multiplexing and switching network to support any type of user traffic. It has not achieved this goal after 10 years of implementation. The reason may be various, but the fatal one is its low efficiency when supporting data traffic. Upgrading is the only way that ATM could keep its existence. We propose a novel architecture to improve the current ATM performance. By defining a new cell structure, the new ATM reinforces the end-to-end error control ability, and greatly increases the data transmission efficiency. The introduction of the micro cell further enhances its ability to support a low-speed real-time service. The new architecture can also support bi-directional point-to-multi-point transmission and can be used in the wireless environment directly. To verify its applicability, various simulations are carried out. We introduce our simulation tests on CBR transmission delay, Web response delay, sustained data transfer rate, and link utilization. The analysis and simulation show that, such a modification can overcome the shortcoming of the existing ATM protocol and upgrade its ability to support more services effectively. Xiangning Chen, Shixin Cheng |
ICC | 1 |
| 2001 | TCP Reno and Vegas performance in wireless ad hoc networksabstractWe use a simulation method to analyze the performance of TCP Reno, NewReno, SACK and Vegas in wireless ad hoc networks. We propose an alternative adapted SACK option format (ASACK) and a new technique, round trip time (RTT) notification (RN), to improve the performance of TCP Reno and Vegas in wireless ad hoc networks respectively. An associativity-based routing (ABR) protocol is also implemented to study how the route stability affects the performance of TCP. Shixin Cheng, Xiangning Chen |
ICC | 3 |