VLDB 2026 Research / reviewers in the wild / expert
Xuan Cao
dblp:185/4253
· DBLP profile ↗
19ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Request-Only Optimization for Recommendation SystemsabstractRecommendation systems represent one of the largest machine learning applications on the planet -- industry-scale recommendation models are trained with petabytes of data and serve billions of users every day. To utilize the rich user signals in the long user history, these models have been scaled up to unprecedented complexity, up to trillions of floating-point operations (TFLOPs) per example. This scale, coupled with the huge amount of training data, necessitates new storage and training algorithms to efficiently improve the quality of these complex recommendation systems. Lucy Liao, Huihui Cheng, Yanzun Huang, Keke Zhai, Pengchao Wang, Timothy Shi, Xuan Cao, Renqin Cai, Zhaojie Gong, Omkar Vichare, Rui Jian, Leon Gao, Shiyan Deng, Wenlei Xie, Jiaqi Zhai |
SIGIR | 12 |
| 2025 | CompNET: Boosting image recognition and writer identification via complementary neural network post-processingabstractIn current classification tasks, an important method to improve accuracy is to pre-train the model using a large-scale domain-specific dataset. However, many tasks such as writer identification (writerID) lack suitable large-scale datasets in practical scenarios. To address this issue, this paper proposes a method that can improve prediction accuracy without relying on significant pre-training but leveraging the diversity of probability distributions predicted by multiple networks, and enhancing the top-1 accuracy through complementary post-processing. Specifically, top-k distributions are sampled from the multiple probability mass functions separately. When the distribution differences of top-k are maximized, the intersection other than the correct category can be narrowed down. Finally, the correct target with suboptimal probability can be rectified by the only intersection. Furthermore, our method has exhibited an intriguing trait during experimentation. Its prediction accuracy enhances concurrently with the incorporation of novel SOTA methods, ultimately surpassing the performance of these new methods. Bocheng Zhao, Xuan Cao, Wenxing Zhang, Xujie Liu, Qiguang Miao, Yunan Li 0001 |
Pattern Recognit. | 2 |
| 2025 | Cross-corpus speech emotion recognition using semi-supervised domain adaptation network
Mao-shen Jia, Xuan Cao, Jiawei Ru, Xinfeng Zhang 0002 |
Speech Commun. | 3 |
| 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsabstractLarge-scale recommendation systems are characterized by their reliance on high cardinality, heterogeneous features and the need to handle tens of billions of user actions on a daily basis. Despite being trained on huge volume of data with thousands of features, most Deep Learning Recommendation Models (DLRMs) in industry fail to scale with compute. Inspired by success achieved by Transformers in language and vision domains, we revisit fundamental design choices in recommendation systems. We reformulate recommendation problems as sequential transduction tasks within a generative modeling framework (``Generative Recommenders''), and propose a new architecture, HSTU, designed for high cardinality, non-stationary streaming recommendation data. HSTU outperforms baselines over synthetic and public datasets by up to 65.8% in NDCG, and is 5.3x to 15.2x faster than FlashAttention2-based Transformers on 8192 length sequences. HSTU-based Generative Recommenders, with 1.5 trillion parameters, improve metrics in online A/B tests by 12.4% and have been deployed on multiple surfaces of a large internet platform with billions of users. More importantly, the model quality of Generative Recommenders empirically scales as a power-law of training compute across three orders of magnitude, up to GPT-3/LLaMa-2 scale, which reduces carbon footprint needed for future model developments, and further paves the way for the first foundation models in recommendations. Jiaqi Zhai, Lucy Liao, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He 0008, Yinghai Lu |
ICML | 6 |
| 2024 | Generating Explanations for Autonomous Robots Using Assumption-Alignment TrackingabstractAs the techniques of autonomous robots advance, there is an increasing demand for robots to provide explanations for their behavior. There are two commonly used explanation types. The first type emphasizes that a robot's policy is the best (or only) option that satisfies a specific property produced by its decision-making algorithms. The second explanation type is used when a robot fails and describes the cause of an error state that led to the failure. This paper proposes a new explanation type derived from a robot's proficiency self-assessment. The proposed explanation type not only supplements the first explanation type under typical operating conditions but also includes the second explanation type when the robot fails. The proposed explanation type is based on assumption-alignment tracking (AAT), a novel method for robot proficiency self-assessment. AAT provides three pieces of information for explanation generation: (1) assessment of assumptions veracity on which the robot's generators rely; (2) proficiency assessment measured by the probability that the robot will successfully accomplish its task; (3) counterfactual proficiency assessment computed by hypothetically varying assumptions. The information provided by AAT fits the situation awareness-based framework for explainable artificial intelligence. Examples of generated explanations are demonstrated using a simulated robot setting up a table with different blocks. Xuan Cao, Jacob W. Crandall, Michael A. Goodrich |
SMC | 1 |
| 2023 | Proficiency Self-Assessment without Breaking the Robot: Anomaly Detection using Assumption-Alignment Tracking from Safe ExperimentsabstractProficiency self-assessment (PSA), the ability to assess how well one can carry out a task, is a desirable capability of autonomous robot systems. Prior work has proposed assumption-alignment tracking (AAT) for performing PSA, and has shown that it can accurately predict robot performance in real-time given a dataset obtained from both normal and abnormal training runs. Obtaining data in abnormal conditions (i.e., conditions in which the robot is not prepared to operate) is difficult and is often not possible. As a result, many realistic datasets contain very few data points for abnormal conditions, making it difficult to apply AAT. This paper hypothesizes that a one-class classifier can be built to detect anomalies using only data collected under normal conditions. Two metrics, difference and separation, are proposed and used to demonstrate that AAT feature vectors from different running conditions tend to form distinct clusters that are identifiable by mainstream one-class classification algorithms. Thus, one-class classifiers trained on AAT feature vectors from normal data can detect anomalous conditions. Furthermore, preliminary results suggest that a few abnormal data points, if available, can be used to classify the abnormality type and, in turn, the degree to which the anomalies will likely impact robot performance. Empirical results from both a simulated navigation robot and a Sawyer robot manipulating blocks show the efficacy of the approach. Xuan Cao, Jacob W. Crandall, Ethan Pedersen, Alvika Gautam, Michael A. Goodrich |
ICRA | 1 |
| 2023 | ggpicrust2: an R package for PICRUSt2 predicted functional profile analysis and visualizationabstractSUMMARY: Microbiome research is now moving beyond the compositional analysis of microbial taxa in a sample. Increasing evidence from large human microbiome studies suggests that functional consequences of changes in the intestinal microbiome may provide more power for studying their impact on inflammation and immune responses. Although 16S rRNA analysis is one of the most popular and a cost-effective method to profile the microbial compositions, marker-gene sequencing cannot provide direct information about the functional genes that are present in the genomes of community members. Bioinformatic tools have been developed to predict microbiome function with 16S rRNA gene data. Among them, PICRUSt2 (Phylogenetic Investigation of Communities by Reconstruction of Unobserved States) has become one of the most popular functional profile prediction tools, which generates community-wide pathway abundances. However, no state-of-art inference tools are available to test the differences in pathway abundances between comparison groups. We have developed ggpicrust2, an R package, for analyzing functional profiles derived from 16S rRNA sequencing. This powerful tool enables researchers to conduct extensive differential abundance analyses and generate visually appealing visualizations that effectively highlight functional signals. With ggpicrust2, users can obtain publishable results and gain deeper insights into the functional composition of their microbial communities. AVAILABILITY AND IMPLEMENTATION: The package is open-source under the MIT and file license and is available at CRAN and https://github.com/cafferychen777/ggpicrust2. Its shiny web is available at https://a95dps-caffery-chen.shinyapps.io/ggpicrust2_shiny/. Jiahao Mai, Xuan Cao, Aaron Burberry, Fabio Cominelli |
Bioinform. | 3 |
| 2023 | Robot Proficiency Self-Assessment Using Assumption-Alignment TrackingabstractA robot is proficient if its performance for its task(s) satisfies a specific standard. While the design of autonomous robots often emphasizes such proficiency, another important attribute of autonomous robot systems is their ability to evaluate their own proficiency. A robot should be able to conduct proficiency self-assessment (PSA), i.e. assess how well it can perform a task before, during, and after it has attempted the task. We propose the assumption-alignment tracking (AAT) method, which provides time-indexed assessments of the veracity of robot generators' assumptions, for designing autonomous robots that can effectively evaluate their own performance. AAT can be considered as a general framework for using robot sensory data to extract useful features, which are then used to build data-driven PSA models. We develop various AAT-based data-driven approaches to PSA from different perspectives. First, we use AAT for estimating robot performance. AAT features encode how the robot's current running condition varies from the normal condition, which correlates with the deviation level between the robot's current performance and normal performance. We use the k-nearest neighbor algorithm to model that correlation. Second, AAT features are used for anomaly detection. We treat anomaly detection as a one-class classification problem where only data from the robot operating in normal conditions are used in training, decreasing the burden on acquiring data in various abnormal conditions. The cluster boundary of data points from normal conditions, which serves as the decision boundary between normal and abnormal conditions, can be identified by mainstream one-class classification algorithms. Third, we improve PSA models that predict robot success/failure by introducing meta-PSA models that assess the correctness of PSA models. The probability that a PSA model's prediction is correct is conditioned on four features: 1) the mean distance from a test sample to its nearest neighbors in the training set; 2) the predicted probability of success made by the PSA model; 3) the ratio between the robot's current performance and its performance standard; and 4) the percentage of the task the robot has already completed. Meta-PSA models trained on the four features using a Random Forest algorithm improve PSA models with respect to both discriminability and calibration. Finally, we explore how AAT can be used to generate a new type of explanation of robot behavior/policy from the perspective of a robot's proficiency. AAT provides three pieces of information for explanation generation: (1) veracity assessment of the assumptions on which the robot's generators rely; (2) proficiency assessment measured by the probability that the robot will successfully accomplish its task; and (3) counterfactual proficiency assessment computed with the veracity of some assumptions varied hypothetically. The information provided by AAT fits the situation awareness-based framework for explainable artificial intelligence. The efficacy of AAT is comprehensively evaluated using robot systems with a variety of robot types, generators, hardware, and tasks, including a simulated robot navigating in a maze-based (discrete time) Markov chain environment, a simulated robot navigating in a continuous environment, and both a simulated and a real-world robot arranging blocks of different shapes and colors in a specific order on a table. Xuan Cao, Alvika Gautam, Tim Whiting, Skyler Smith, Michael A. Goodrich, Jacob W. Crandall |
IEEE Trans. Robotics | 1 |
| 2022 | A Method for Designing Autonomous Robots that Know Their LimitsabstractWhile the design of autonomous robots often emphasizes developing proficient robots, another important attribute of autonomous robot systems is their ability to evaluate their own proficiency and limitations. A robot should be able to assess how well it can perform a task before, during, and after it attempts the task. Thus, we consider the following question: How can we design autonomous robots that know their own limits? Toward this end, this paper presents an approach, called assumption-alignment tracking (AAT), for designing autonomous robots that can effectively evaluate their own limits. In AAT, the robot combines (a) measures of how well its decision-making algorithms align with its environment and hardware systems with (b) its past experiences to assess its ability to succeed at a given task. The effectiveness of AAT in assessing a robot's limits are illustrated in a robot navigation task. Alvika Gautam, Tim Whiting, Xuan Cao, Michael A. Goodrich, Jacob W. Crandall |
ICRA | 3 |
| 2022 | Adapted Metrics for Measuring Competency and Resilience for Autonomous Robot Systems in Discrete Time Markov ChainsabstractAutonomous robot systems are often designed to achieve specific goals. This paper restricts attention to a specific type of goal, namely reaching a desired state within a certain time bound. For such goals, a robot system’s competency and resilience can be defined as the probability of reaching the desired state as a function of the time bound under a nominal unperturbed condition and under known perturbation conditions, respectively. Two metrics taken from prior work for measuring competency and resilience, power and efficiency, are modified so that they do not require subjective parameters. This paper formalizes the adapted metrics for discrete time Markov chains. The adapted metrics are applied to a best-of-N case study that is solved by a graph-based approach and modeled as a discrete time Markov chain. The case study demonstrates that the modified metrics allow power-efficiency trade-offs to be more easily visualized than the cluttered visualizations produced by the original metrics. Xuan Cao, Puneet Jain, Michael A. Goodrich |
SMC | 1 |
| 2021 | A Hierarchical Retrieval Method Based on Hash Table for Audio Fingerprinting
Mao-shen Jia, Xuan Cao |
ICIC (1) | 3 |
| 2020 | Adversarial Refinement Network for Human Motion Prediction
Xianjin Chao, Yanrui Bin, Wenqing Chu, Xuan Cao, Yanhao Ge, Chengjie Wang 0001, Feiyue Huang, Howard Leung |
ACCV (2) | 4 |
| 2020 | Adversarial Semantic Data Augmentation for Human Pose Estimation
Yanrui Bin, Xuan Cao, Xinya Chen, Yanhao Ge, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Changxin Gao, Nong Sang |
ECCV (19) | 2 |
| 2019 | Communication-Aware Container Placement and Reassignment in Large-Scale Internet Data CentersabstractContainerization has been used in many applications for isolation purposes due to its lightweight, scalable, and highly portable properties. However, to apply containerization in large-scale Internet data centers faces a big challenge. Services in data centers are always instantiated as a group of containers, which often generate heavy communication workloads and therefore resulting in inefficient communications and downgraded service performance. Although assigning the containers of the same service to the same server can reduce the communication overhead, this may cause heavily imbalanced resource utilization since containers of the same service are usually intensive to the same resource. To reduce communication cost as well as balance the resource utilization in large-scale data centers, we further explore the container distribution issues in a real industrial environment and find that such conflict lies in two phases-container placement and container reassignment. The objective of this paper is to address the container distribution problem in these two phases. For the container placement problem, we propose an efficient communication aware worst fit decreasing algorithm to place a set of new containers into data centers. For the container reassignment problem, we propose a two-stage algorithm called Sweep&Search to optimize a given initial distribution of containers by migrating containers among servers. We implement the proposed algorithms in Baidu's data centers and conduct extensive evaluations. Compared with the state-of-the-art strategies, the evaluation results show that our algorithms perform better up to 70% and increase the overall service throughput up to 90% simultaneously. Yuchao Zhang 0004, Yusen Li, Ke Xu 0002, Dan Wang 0002, Wendong Wang 0003, Xuan Cao, Qingqing Liang |
IEEE J. Sel. Areas Commun. | 8 |
| 2019 | Balance of mechanical forces drives endothelial gap formation and may facilitate cancer and immune-cell extravasationabstractThe formation of gaps in the endothelium is a crucial process underlying both cancer and immune cell extravasation, contributing to the functioning of the immune system during infection, the unfavorable development of chronic inflammation and tumor metastasis. Here, we present a stochastic-mechanical multiscale model of an endothelial cell monolayer and show that the dynamic nature of the endothelium leads to spontaneous gap formation, even without intervention from the transmigrating cells. These gaps preferentially appear at the vertices between three endothelial cells, as opposed to the border between two cells. We quantify the frequency and lifetime of these gaps, and validate our predictions experimentally. Interestingly, we find experimentally that cancer cells also preferentially extravasate at vertices, even when they first arrest on borders. This suggests that extravasating cells, rather than initially signaling to the endothelium, might exploit the autonomously forming gaps in the endothelium to initiate transmigration. Jorge Escribano, Michelle B. Chen, Emad Moeendarbary, Xuan Cao, Vivek Shenoy, José Manuel García-Aznar, Roger D. Kamm, Fabian Spill |
PLoS Comput. Biol. | 4 |
| 2018 | Sparse Photometric 3D Face Reconstruction Guided by Morphable ModelsabstractWe present a novel 3D face reconstruction technique that leverages sparse photometric stereo (PS) and latest advances on face registration / modeling from a single image. We observe that 3D morphable faces approach [21] provides a reasonable geometry proxy for light position calibration. Specifically, we develop a robust optimization technique that can calibrate per-pixel lighting direction and illumination at a very high precision without assuming uniform surface albedos. Next, we apply semantic segmentation on input images and the geometry proxy to refine hairy vs. bare skin regions using tailored filter. Experiments on synthetic and real data show that by using a very small set of images, our technique is able to reconstruct fine geometric details such as wrinkles, eyebrows, whelks, pores, etc, comparable to and sometimes surpassing movie quality productions. Xuan Cao, Anpei Chen, Xin Chen 0040, Jingyi Yu 0001 |
CVPR | 1 |
| 2018 | Index Shard Replication Strategies for Improving Resource Utilization in Large Scale Search EnginesabstractWith the rapid growth of the Web scale, large scale search engines have to set up a huge number of machines to place the index files of the Web contents. The index files are normally divided into smaller index shards which are often replicated so that queries can be processed in parallel. We observe that the index shard replication strategy could have a significant impact on the resource utilization of machines. In this paper, we investigate the index shard replication problem with the goal of improving the resource utilization of machines in search engine datacenters. We formulate the problem as a variant of the Multi-Dimensional Vector Bin Packing Problem and propose several strategies to approximate the optimal solution. The proposed strategies are evaluated by extensive experiments using data from real commercial search engines. The results demonstrate the effectiveness of the proposed strategies. Our work also yields many insights about the impact of different input properties on the performance and the key factors that each strategy is sensitive to. We believe that this paper will provide valuable guidance to the choice of the index shard replication strategy in practice. Yusen Li, Xueyan Tang, Wentong Cai 0001, Jiancong Tong, Xiaoguang Liu 0001, Gang Wang 0001, Chuansong Gao, Xuan Cao, Guanhui Geng |
ICPP | 8 |
| 2017 | A Communication-Aware Container Re-Distribution Approach for High Performance VNFsabstractContainers have been used in many applications for isolation purposes due to the lightweight, scalable and highly portable properties. However, to apply containers in virtual network functions (VNFs) faces a big challenge because high-performance VNFs often generate frequent communication workloads among containers while the container communications are generally not efficient. Compared with hardware modification solutions, properly distributing containers among hosts is an efficient and low-cost way to reduce communication overhead. However, we observe that this approach yields a trade-off between the communication overhead and the overall throughput of the cluster. In this paper, we focus on the communication-aware container redistribution problem to optimize the communication overhead and the overall throughput jointly for VNF clusters. We propose a solution called FreeContainer which utilizes a novel two-stage algorithm to re-distribute containers among hosts. We implement FreeContainer in Baidu clusters with 6000 servers and 35 services deployed. Extensive experiments on real networks are conducted to evaluate the performance of the proposed approach. The results show that FreeContainer can increase the overall throughput up to 90% with significant reduction on communication overhead. Yuchao Zhang 0004, Yusen Li, Ke Xu 0002, Dan Wang 0002, Xuan Cao, Qingqing Liang |
ICDCS | 6 |
| 2016 | A game-theoretic approach to sub-vertex registration
Zheng Geng, Xuan Cao, Renjing Pei, Xiangbing Meng |
Pattern Recognit. Lett. | 3 |