EDBT 2026 Demo / reviewers in the wild / expert
Yuxin Wen
dblp:205/4237
· DBLP profile ↗
35ranked-venue papers
14as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 11 first-author · 25 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Log-Likelihood Chain Framework for Defending Against LDP Data Poisoning AttacksabstractLocal differential privacy (LDP) provides strict privacy guarantee in a distributed environment. Recent studies demonstrated that LDP protocols are vulnerable to data poisoning attacks where an attacker can manipulate the perturbed result on the local side and send bogus data to skew the final estimate on the server. Unfortunately, existing attack detections do not create an effective attack indicator and rely on particular characteristics of LDP protocols. As a result, they typically exhibit limited detection performance. In this paper, we use log-likelihood as the attack indicator and propose a chain-style detection to enhance the detection effectiveness, in which the attack impact could propagate along the chain and exhibit clear anomaly signal even under stealthy attack scenarios. The experimental results show that our detection consistently outperforms the existing methods. Using four datasets containing categorical and numerical data separately, our detection achieves an F1 score exceeding 96% in most cases. It even remains above 0.9 under stealthy attack settings, outperforming the state-of-the-art detection by up to 0.25. Yuxin Wen, Haonan Yan, Yahong Chen, Zhe Sun 0005, Hui Li 0006, Xiaodong Lin 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Stable-SCore: A Stable Registration-based Framework for 3D Shape CorrespondenceabstractEstablishing character shape correspondence is a critical and fundamental task in computer vision and graphics, with diverse applications including re-topology, attribute transfer, and shape interpolation. Current dominant functional map methods, while effective in controlled scenarios, struggle in real situations with more complex challenges such as non-isometric shape discrepancies. In response, we revisit registration-for-correspondence methods and tap their potential for more stable shape correspondence estimation. To overcome their common issues including unstable deformations and the necessity for careful pre-alignment or high-quality initial 3D correspondences, we introduce Stable-SCore: A Stable Registration-based Framework for 3D Shape Correspondence. We first re-purpose a foundation model for 2D character correspondence that ensures reliable and stable 2D mappings. Crucially, we propose a novel Semantic Flow Guided Registration approach that leverages 2D correspondence to guide mesh deformations. Our framework significantly surpasses existing methods in challenging scenarios, and brings possibilities for a wide array of real applications, as demonstrated in our results. Haolin Liu 0004, Xiaohang Zhan, Zizheng Yan, Zhongjin Luo, Yuxin Wen, Xiaoguang Han 0001 |
CVPR | 5 |
| 2025 | A Technical Report on "Erasing the Invisible": The 2024 NeurIPS Competition on Stress Testing Image WatermarksabstractAI-generated images have become pervasive, raising critical concerns around content authenticity, intellectual property, and the spread of misinformation. Invisible watermarks offer a promising solution for identifying AI-generated images, preserving content provenance without degrading visual quality. However, their real-world robustness remains uncertain due to the lack of standardized evaluation protocols and large-scale stress testing. To bridge this gap, we organized “Erasing the Invisible,” a NeurIPS 2024 competition and newly established benchmark designed to systematically stress testing the resilience of watermarking techniques. The competition introduced two attack tracks—Black-box and Beige-box—that simulate practical scenarios with varying levels of attacker knowledge on watermarks, providing a comprehensive assessment of watermark robustness. The competition attracted significant global participation, with 2,722 submissions from 298 teams. Through a rigorous evaluation pipeline featuring real-time feedback and human-verified final rankings, participants developed and demonstrated new attack strategies that revealed critical vulnerabilities in state-of-the-art watermarking methods. On average, the top-5 teams in both tracks could remove watermarks from $\geq$ 89% of the images while preserving high visual quality, setting strong baselines for future research on watermark attacks and defenses. To support continued progress in this field, we summarize the insights and lessons learned from this competition in this paper, and release the benchmark dataset, evaluation toolkit, and competition results. “Erasing the Invisible” establishes a valuable open resource for advancing more robust watermarking techniques and strengthening content provenance in the era of generative AI. Mucong Ding, Bang An 0001, Tahseen Rabbani, Chenghao Deng, Anirudh Satheesh, Souradip Chakraborty, Mehrdad Saberi, Yuxin Wen, Kyle Sang, Aakriti Agrawal, Xuandong Zhao, Mary-Anne Hartley, Lei Li 0005, Yu-Xiang Wang 0003, Vishal M. Patel, Soheil Feizi, Tom Goldstein, Furong Huang |
NeurIPS | 8 |
| 2025 | Quantifying Cross-Modality Memorization in Vision-Language ModelsabstractUnderstanding what and how neural networks memorize during training is crucial, both from the perspective of unintentional memorization of potentially sensitive information and from the standpoint of effective knowledge acquisition for real-world, knowledge-intensive tasks. While previous studies primarily investigate memorization within a single modality, such as text memorization in large language models or image memorization in diffusion models, unified multimodal models are becoming increasingly prevalent in practical applications. In this work, we focus on the unique characteristics of cross-modality memorization and conduct a systematic study centered on vision-language models. To facilitate controlled experiments, we first introduce a synthetic persona dataset comprising diverse synthetic person images and textual descriptions. We quantify factual knowledge memorization and cross-modal transferability by training models on a single modality and evaluating their performance in the other. Our results reveal that facts learned in one modality transfer to the other, but a significant gap exists between recalling information in the source and target modalities. Furthermore, we observe that this gap exists across various scenarios, including more capable models, machine unlearning, and the multi-hop case. At the end, we propose a baseline method to mitigate this challenge. We hope our study can inspire future research on developing more robust multimodal learning techniques to enhance cross-modal transferability. Yuxin Wen, Yangsibo Huang, Tom Goldstein, Ravi Kumar 0001, Badih Ghazi, Chiyuan Zhang |
NeurIPS | 1 |
| 2025 | MTGlu: A Multi-scale Time-Frequency Approach for Personalized Glucose Prediction in Type 1 Diabetes
Lantian Fu, Xuxue Sun, Zhenning Bao, Yuxin Wen |
PAKDD (2) | 5 |
| 2025 | Multimodal mixing convolutional neural network and transformer for Alzheimer's disease recognition
Junde Chen, Sherry Yun Wang, Adnan Zeb, Md Suzauddola, Yuxin Wen |
Expert Syst. Appl. | 5 |
| 2025 | Privacy-enhanced data distillation with probability distribution matching
Ke Pan 0001, Yuxin Wen, Yiming Wang 0010, Maoguo Gong, Hui Li 0006, Shanfeng Wang |
Neurocomputing | 2 |
| 2024 | NEFTune: Noisy Embeddings Improve Instruction FinetuningabstractWe show that language model finetuning can be improved, sometimes dramatically, with a simple augmentation.
NEFTune adds noise to the embedding vectors during training.
Standard finetuning of LLaMA-2-7B using Alpaca achieves $29.79$\% on AlpacaEval, which rises to $64.69$\% using noisy embeddings. NEFTune also improves over strong baselines on modern instruction datasets.
Models trained with Evol-Instruct see a $10$\% improvement, with ShareGPT an $8$\% improvement, and with OpenPlatypus an $8$\% improvement.
Even powerful models further refined with RLHF such as LLaMA-2-Chat benefit from additional training with NEFTune. Particularly, we see these improvements on the conversational abilities of the instruction model and not on traditional tasks like those on the OpenLLM Leaderboard, where performance is the same. Neel Jain, Ping-Yeh Chiang, Yuxin Wen, John Kirchenbauer, Hong-Min Chu, Gowthami Somepalli, Brian R. Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Aniruddha Saha, Micah Goldblum, Jonas Geiping, Tom Goldstein |
ICLR | 3 |
| 2024 | On the Reliability of Watermarks for Large Language ModelsabstractAs LLMs become commonplace, machine-generated text has the potential to flood the internet with spam, social media bots, and valueless content. _Watermarking_ is a simple and effective strategy for mitigating such harms by enabling the detection and documentation of LLM-generated text. Yet a crucial question remains: How reliable is watermarking in realistic settings in the wild? There, watermarked text may be modified to suit a user's needs, or entirely rewritten to avoid detection. We study the robustness of watermarked text after it is re-written by humans, paraphrased by a non-watermarked LLM, or mixed into a longer hand-written document. We find that watermarks remain detectable even after human and machine paraphrasing. While these attacks dilute the strength of the watermark, paraphrases are statistically likely to leak n-grams or even longer fragments of the original text, resulting in high-confidence detections when enough tokens are observed. For example, after strong human paraphrasing the watermark is detectable after observing 800 tokens on average, when setting a $1\mathrm{e}{-5}$ false positive rate. We also consider a range of new detection schemes that are sensitive to short spans of watermarked text embedded inside a large document, and we compare the robustness of watermarking to other kinds of detectors. John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, Tom Goldstein |
ICLR | 3 |
| 2024 | Detecting, Explaining, and Mitigating Memorization in Diffusion ModelsabstractRecent breakthroughs in diffusion models have exhibited exceptional image-generation capabilities. However, studies show that some outputs are merely replications of training data. Such replications present potential legal challenges for model owners, especially when the generated content contains proprietary information. In this work, we introduce a straightforward yet effective method for detecting memorized prompts by inspecting the magnitude of text-conditional predictions. Our proposed method seamlessly integrates without disrupting sampling algorithms, and delivers high accuracy even at the first generation step, with a single generation per prompt. Building on our detection strategy, we unveil an explainable approach that shows the contribution of individual words or tokens to memorization. This offers an interactive medium for users to adjust their prompts. Moreover, we propose two strategies i.e., to mitigate memorization by leveraging the magnitude of text-conditional predictions, either through minimization during inference or filtering during training. These proposed strategies effectively counteract memorization while maintaining high-generation quality. Code is available at https://github.com/YuxinWenRick/diffusion_memorization. Yuxin Wen, Chen Chen 0043, Lingjuan Lyu |
ICLR | 1 |
| 2024 | WAVES: Benchmarking the Robustness of Image WatermarksabstractIn the burgeoning age of generative AI, watermarks act as identifiers of provenance and artificial content. We present WAVES (Watermark Analysis via Enhanced Stress-testing), a benchmark for assessing image watermark robustness, overcoming the limitations of current evaluation methods. WAVES integrates detection and identification tasks and establishes a standardized evaluation protocol comprised of a diverse range of stress tests. The attacks in WAVES range from traditional image distortions to advanced, novel variations of diffusive, and adversarial attacks. Our evaluation examines two pivotal dimensions: the degree of image quality degradation and the efficacy of watermark detection after attacks. Our novel, comprehensive evaluation reveals previously undetected vulnerabilities of several modern watermarking algorithms. We envision WAVES as a toolkit for the future development of robust watermarks. Bang An 0001, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal, Yuancheng Xu, Chenghao Deng, Sicheng Zhu, Abdirisak Mohamed, Yuxin Wen, Tom Goldstein, Furong Huang |
ICML | 9 |
| 2024 | Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMsabstractLarge language models can memorize and repeat their training data, causing privacy and copyright risks. To mitigate memorization, we introduce a subtle modification to the next-token training objective that we call the goldfish loss. During training, a randomly sampled subsets of tokens are excluded from the loss computation. These dropped tokens are not memorized by the model, which prevents verbatim reproduction of a complete chain of tokens from the training set. We run extensive experiments training billion-scale LLaMA-2 models, both pre-trained and trained from scratch, and demonstrate significant reductions in extractable memorization with little to no impact on downstream benchmarks.
_Code and checkpoints: https://github.com/ahans30/goldfish-loss_ Abhimanyu Hans, John Kirchenbauer, Yuxin Wen, Neel Jain, Hamid Kazemi, Prajwal Singhania, Gowthami Somepalli, Jonas Geiping, Abhinav Bhatele, Tom Goldstein |
NeurIPS | 3 |
| 2024 | Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained ModelsabstractIt is commonplace to produce application-specific models by fine-tuning large pre-trained models using a small bespoke dataset. The widespread availability of foundation model checkpoints on the web poses considerable risks, including the vulnerability to backdoor attacks. In this paper, we unveil a new vulnerability: the privacy backdoor attack. This black-box privacy attack aims to amplify the privacy leakage that arises when fine-tuning a model: when a victim fine-tunes a backdoored model, their training data will be leaked at a significantly higher rate than if they had fine-tuned a typical model. We conduct extensive experiments on various datasets and models, including both vision-language models (CLIP) and large language models, demonstrating the broad applicability and effectiveness of such an attack. Additionally, we carry out multiple ablation studies with different fine-tuning methods and inference strategies to thoroughly analyze this new threat. Our findings highlight a critical privacy concern within the machine learning community and call for a re-evaluation of safety protocols in the use of open-source pre-trained models. Yuxin Wen, Leo Marchyok, Sanghyun Hong 0001, Jonas Geiping, Tom Goldstein, Nicholas Carlini |
NeurIPS | 1 |
| 2024 | Democratizing AI: Open-source Scalable LLM Training on GPU-based SupercomputersabstractTraining and fine-tuning large language models (LLMs) with hundreds of billions to trillions of parameters requires tens of thousands of GPUs, and a highly scalable software stack. In this work, we present a novel four-dimensional hybrid parallel algorithm implemented in a highly scalable, portable, open-source framework called AxoNn. We describe several performance optimizations in AxoNN to improve matrix multiply kernel performance, overlap non-blocking collectives with computation, and performance modeling to choose performance optimal configurations. These have resulted in unprecedented scaling and peak flop/s (bf16) for training of GPT-style transformer models on Perlmutter (620.1 Petaflop/s), Frontier (1.381 Exaflop/s) and Alps (1.423 Exaflop/s). While the abilities of LLMs improve with the number of trainable parameters, so do privacy and copyright risks caused by memorization of training data, which can cause disclosure of sensitive or private information at inference time. We highlight this side effect of scale through experiments that explore “catastrophic memorization,” where models are sufficiently large to memorize training data in a single pass, and present an approach to prevent it. As part of this study, we demonstrate fine-tuning of a 405-billion parameter LLM using AxoNN on Frontier. Prajwal Singhania, Aditya K. Ranjan, John Kirchenbauer, Jonas Geiping, Yuxin Wen, Neel Jain, Abhimanyu Hans, Manli Shu, Aditya Tomar, Tom Goldstein, Abhinav Bhatele |
SC | 6 |
| 2024 | Surface Reconstruction From Point Clouds: A Survey and a BenchmarkabstractReconstruction of a continuous surface of two-dimensional manifold from its raw, discrete point cloud observation is a long-standing problem in computer vision and graphics research. The problem is technically ill-posed, and becomes more difficult considering that various sensing imperfections would appear in the point clouds obtained by practical depth scanning. In literature, a rich set of methods has been proposed, and reviews of existing methods are also provided. However, existing reviews are short of thorough investigations on a common benchmark. The present paper aims to review and benchmark existing methods in the new era of deep learning surface reconstruction. To this end, we contribute a large-scale benchmarking dataset consisting of both synthetic and real-scanned data; the benchmark includes object- and scene-level surfaces and takes into account various sensing imperfections that are commonly encountered in practical depth scanning. We conduct thorough empirical studies by comparing existing methods on the constructed benchmark, and pay special attention on robustness of existing methods against various scanning imperfections; we also study how different methods generalize in terms of reconstructing complex surface shapes. Our studies help identity the best conditions under which different methods work, and suggest some empirical findings. For example, while deep learning methods are increasingly popular in the research community, our systematic studies suggest that, surprisingly, a few classical methods perform even better in terms of both robustness and generalization; our studies also suggest that the practical challenges of misalignment of point sets from multi-view scanning, missing of surface points, and point outliers remain unsolved by all the existing surface reconstruction methods. We expect that the benchmark and our studies would be valuable both for practitioners and as a guidance for new innovations in future research. Zhangjin Huang, Yuxin Wen, Jinjuan Ren, Kui Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | STYX: Adaptive Poisoning Attacks Against Byzantine-Robust Defenses in Federated LearningabstractDecentralized training of machine learning models, for instance with federated learning protocols, continues to diffuse from theory toward practical applications and use cases. In federated learning (FL), a central server trains a model collaboratively with a group of users by communicating model updates, without the exchange of private user information. However, these systems can be influenced during training by malicious users who send poisoned updates. Because the training is decentralized and each user controls their own device, these users are free to poison the training protocol. In turn, this has lead to a number of proposals to incorporate aggregation strategies from byzantine-robust learning into the FL paradigm. Byzantine strategies are provably secure for simple model classes, and these robustness properties are often assumed to extend to neural models as well. In this work, we argue that a range of popular robust aggregation strategies, when applied to neural networks, can be trivially circumvented through simple adaptive attacks. We discuss the intuitions behind these adaptive attacks, and show that, despite their simplicity, they provide strong baselines that lead to significant decreases in model performance in FL systems. Yuxin Wen, Jonas Geiping, Micah Goldblum, Tom Goldstein |
ICASSP | 1 |
| 2023 | Decepticons: Corrupted Transformers Breach Privacy in Federated Learning for Language Models
Liam Fowl, Jonas Geiping, Steven Reich, Yuxin Wen, Wojciech Czaja, Micah Goldblum, Tom Goldstein |
ICLR | 4 |
| 2023 | Canary in a Coalmine: Better Membership Inference with Ensembled Adversarial Queries
Yuxin Wen, Arpit Bansal, Hamid Kazemi, Eitan Borgnia, Micah Goldblum, Jonas Geiping, Tom Goldstein |
ICLR | 1 |
| 2023 | A Watermark for Large Language ModelsabstractPotential harms of large language models can be mitigated by watermarking model output, i.e., embedding signals into generated text that are invisible to humans but algorithmically detectable from a short span of tokens. We propose a watermarking framework for proprietary language models. The watermark can be embedded with negligible impact on text quality, and can be detected using an efficient open-source algorithm without access to the language model API or parameters. The watermark works by selecting a randomized set of "green" tokens before a word is generated, and then softly promoting use of green tokens during sampling. We propose a statistical test for detecting the watermark with interpretable p-values, and derive an information-theoretic framework for analyzing the sensitivity of the watermark. We test the watermark using a multi-billion parameter model from the Open Pretrained Transformer (OPT) family, and discuss robustness and security. John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein |
ICML | 3 |
| 2023 | Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryabstractThe strength of modern generative models lies in their ability to be controlled through prompts. Hard prompts comprise interpretable words and tokens, and are typically hand-crafted by humans. Soft prompts, on the other hand, consist of continuous feature vectors. These can be discovered using powerful optimization methods, but they cannot be easily edited, re-used across models, or plugged into a text-based interface. We describe an easy-to-use approach to automatically optimize hard text prompts through efficient gradient-based optimization. Our approach can be readily applied to text-to-image and text-only applications alike. This method allows API users to easily generate, discover, and mix and match image concepts without prior knowledge of how to prompt the model. Furthermore, using our method, we can bypass token-level content filters imposed by Midjourney by optimizing through the open-sourced text encoder. Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, Tom Goldstein |
NeurIPS | 1 |
| 2023 | Tree-Rings Watermarks: Invisible Fingerprints for Diffusion ImagesabstractWatermarking the outputs of generative models is a crucial technique for tracing copyright and preventing potential harm from AI-generated content. In this paper, we introduce a novel technique called Tree-Ring Watermarking that robustly fingerprints diffusion model outputs. Unlike existing methods that perform post-hoc modifications to images after sampling, Tree-Ring Watermarking subtly influences the entire sampling process, resulting in a model fingerprint that is invisible to humans. The watermark embeds a pattern into the initial noise vector used for sampling. These patterns are structured in Fourier space so that they are invariant to convolutions, crops, dilations, flips, and rotations. After image generation, the watermark signal is detected by inverting the diffusion process to retrieve the noise vector, which is then checked for the embedded signal. We demonstrate that this technique can be easily applied to arbitrary diffusion models, including text-conditioned Stable Diffusion, as a plug-in with negligible loss in FID. Our watermark is semantically hidden in the image space and is far more robust than watermarking alternatives that are currently deployed. Yuxin Wen, John Kirchenbauer, Jonas Geiping, Tom Goldstein |
NeurIPS | 1 |
| 2023 | Machine learning techniques for stock price prediction and graphic signal recognition
Junde Chen, Yuxin Wen, Yaser Ahangari Nanehkaran, Md Suzauddola |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | A deep learning approach for inpatient length of stay and mortality prediction
Junde Chen, Trudi Di Qi, Jacqueline Vu, Yuxin Wen |
J. Biomed. Informatics | 4 |
| 2022 | Fishing for User Data in Large-Batch Federated Learning via Gradient MagnificationabstractFederated learning (FL) has rapidly risen in popularity due to its promise of privacy and efficiency. Previous works have exposed privacy vulnerabilities in the FL pipeline by recovering user data from gradient updates. However, existing attacks fail to address realistic settings because they either 1) require toy settings with very small batch sizes, or 2) require unrealistic and conspicuous architecture modifications. We introduce a new strategy that dramatically elevates existing attacks to operate on batches of arbitrarily large size, and without architectural modifications. Our model-agnostic strategy only requires modifications to the model parameters sent to the user, which is a realistic threat model in many scenarios. We demonstrate the strategy in challenging large-scale settings, obtaining high-fidelity data extraction in both cross-device and cross-silo federated learning. Code is available at https://github.com/JonasGeiping/breaching. Yuxin Wen, Jonas Geiping, Liam Fowl, Micah Goldblum, Tom Goldstein |
ICML | 1 |
| 2022 | MtCut: A Multi-Task Framework for Ranked List TruncationabstractRanked list truncation aims to cut the ranked results in short considering user-defined objectives, which balances the overall utility and user efforts over retrieval results. The exact selection of an optimal cut-off position brings potential benefits in various real-world applications, such as patent search and legal search. However, there is significant retrieval bias in the ranked list. The result scores and the disorder of document sequences cause difficulties in judging the relevance between the queries and documents -- alleviating the existing methods' performance improvement. In this work, we investigate the characteristics of retrieval bias on altering truncation and propose a multi-task truncation model, MtCut. It employs two auxiliary tasks to make complementary for the retrieval bias. As a practical evaluation, we explore its performance on two datasets, and the results show that MtCut outperforms the state-of-the-art methods on both F1-score and DCG metrics. Jianxin Li 0002, Tianchen Zhu, Haoyi Zhou, Qishan Zhu, Yuxin Wen, Hongming Piao |
WSDM | 6 |
| 2022 | Geometry-Aware Generation of Adversarial Point CloudsabstractMachine learning models have been shown to be vulnerable to adversarial examples. While most of the existing methods for adversarial attack and defense work on the 2D image domain, a few recent attempts have been made to extend them to 3D point cloud data. However, adversarial results obtained by these methods typically contain point outliers, which are both noticeable and easy to defend against using the simple techniques of outlier removal. Motivated by the different mechanisms by which humans perceive 2D images and 3D shapes, in this paper we propose the new design ofgeometry-aware objectives, whose solutions favor (the discrete versions of) the desired surface properties of smoothness and fairness. To generate adversarial point clouds, we use a targeted attack misclassification loss that supports continuous pursuit of increasingly malicious signals. Regularizing the targeted attack loss with our proposed geometry-aware objectives results in our proposed method,Geometry-Aware Adversarial Attack ($GeoA^3$GeoA3). The results of$GeoA^3$tend to be more harmful, arguably harder to defend against, and of the key adversarial characterization of being imperceptible to humans. While the main focus of this paper is to learn to generate adversarial point clouds, we also present a simple but effective algorithm termed$Geo_{+}A^3$-IterNormPro, with Iterative Normal Projection (IterNorPro) that solves a new objective function$Geo_{+}A^3$, towards surface-level adversarial attacks via generation of adversarial point clouds. We quantitatively evaluate our methods on both synthetic and physical objects in terms of attack success rate and geometric regularity. For a qualitative evaluation, we conduct subjective studies by collecting human preferences from Amazon Mechanical Turk. Comparative results in comprehensive experiments confirm the advantages of our proposed methods. Our source codes are publicly available athttps://github.com/Yuxin-Wen/GeoA3. Yuxin Wen, Jiehong Lin, Ke Chen 0004, C. L. Philip Chen, Kui Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Deep Optimized Priors for 3D Shape Modeling and ReconstructionabstractMany learning-based approaches have difficulty scaling to unseen data, as the generality of its learned prior is limited to the scale and variations of the training samples. This holds particularly true with 3D learning tasks, given the sparsity of 3D datasets available. We introduce a new learning framework for 3D modeling and reconstruction that greatly improves the generalization ability of a deep generator. Our approach strives to connect the good ends of both learning-based and optimization-based methods. In particular, unlike the common practice that fixes the pre-trained priors at test time, we propose to further optimize the learned prior and latent code according to the input physical measurements after the training. We show that the proposed strategy effectively breaks the barriers constrained by the pre-trained priors and could lead to high-quality adaptation to unseen data. We realize our framework using the implicit surface representation and validate the efficacy of our approach in a variety of challenging tasks that take highly sparse or collapsed observations as input. Experimental results show that our approach compares favorably with the state-of-the-art methods in terms of both generality and accuracy. Yuxin Wen, Weikai Chen 0001, Yongwei Chen, Kui Jia |
CVPR | 2 |
| 2021 | Sign-Agnostic Implicit Learning of Surface Self-Similarities for Shape Modeling and Reconstruction From Raw Point CloudsabstractShape modeling and reconstruction from raw point clouds of objects stand as a fundamental challenge in vision and graphics research. Classical methods consider analytic shape priors; however, their performance is degraded when the scanned points deviate from the ideal conditions of cleanness and completeness. Important progress has been recently made by data-driven approaches, which learn global and/or local models of implicit surface representations from auxiliary sets of training shapes. Motivated from a universal phenomenon that self-similar shape patterns of local surface patches repeat across the entire surface of an object, we aim to push forward the data-driven strategies and propose to learn a local implicit surface network for a shared, adaptive modeling of the entire surface for a direct surface reconstruction from raw point cloud; we also enhance the leveraging of surface self-similarities by improving correlations among the optimized latent codes of individual surface patches. Given that orientations of raw points could be unavailable or noisy, we extend signagnostic learning into our local implicit model, which enables our recovery of signed implicit fields of local surfaces from the unsigned inputs. We term our framework as Sign-Agnostic Implicit Learning of Surface Self-Similarities (SAIL-S3). With a global post-optimization of local sign flipping, SAIL-S3 is able to directly model raw, un-oriented point clouds and reconstruct high-quality object surfaces. Experiments show its superiority over existing methods. Wenbin Zhao, Jiabao Lei, Yuxin Wen, Jianguo Zhang 0001, Kui Jia |
CVPR | 3 |
| 2021 | Orthogonal Deep Neural NetworksabstractIn this paper, we introduce the algorithms of Orthogonal Deep Neural Networks (OrthDNNs) to connect with recent interest of spectrally regularized deep learning methods. OrthDNNs are theoretically motivated by generalization analysis of modern DNNs, with the aim to find solution properties of network weights that guarantee better generalization. To this end, we first prove that DNNs are of local isometry on data distributions of practical interest; by using a new covering of the sample space and introducing the local isometry property of DNNs into generalization analysis, we establish a new generalization error bound that is both scale- and range-sensitive to singular value spectrum of each of networks' weight matrices. We prove that the optimal bound w.r.t. the degree of isometry is attained when each weight matrix has a spectrum of equal singular values, among which orthogonal weight matrix or a non-square one with orthonormal rows or columns is the most straightforward choice, suggesting the algorithms of OrthDNNs. We present both algorithms of strict and approximate OrthDNNs, and for the later ones we propose a simple yet effective algorithm called Singular Value Bounding (SVB), which performs as well as strict OrthDNNs, but at a much lower computational cost. We also propose Bounded Batch Normalization (BBN) to make compatible use of batch normalization with OrthDNNs. We conduct extensive comparative studies by using modern architectures on benchmark image classification. Experiments show the efficacy of OrthDNNs. Kui Jia, Yuxin Wen, Tongliang Liu, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | A Neural Network-Based Joint Prognostic Model for Data Fusion and Remaining Useful Life PredictionabstractWith the rapid development of sensor and information technology, now multisensor data relating to the system degradation process are readily available for condition monitoring and remaining useful life (RUL) prediction. The traditional data fusion and RUL prediction methods are either not flexible enough to capture the highly nonlinear relationship between the health condition and the multisensor data or have not fully utilized the past observations to capture the degradation trajectory. In this article, we propose a joint prognostic model (JPM), where Bayesian linear models are developed for multisensor data, and an artificial neural network is proposed to model the nonlinear relationship between the residual life, the model parameters of each sensor data, and the observation epoch. A Bayesian updating scheme is developed to calculate the posterior distributions of the model parameters of each sensor data, which are further used to estimate the posterior predictive distributions of the residual life. The effectiveness and advantages of the proposed JPM are demonstrated using the commercial modular aero-propulsion system simulation data set. Yuxin Wen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Towards Understanding the Regularization of Adversarial Robustness on Neural NetworksabstractThe problem of adversarial examples has shown that modern Neural Network (NN) models could be rather fragile. Among the more established techniques to solve the problem, one is to require the model to be \emph{$\epsilon$-adversarially robust} (AR); that is, to require the model not to change predicted labels when any given input examples are perturbed within a certain range. However, it is observed that such methods would lead to standard performance degradation, i.e., the degradation on natural examples. In this work, we study the degradation through the regularization perspective. We identify quantities from generalization analysis of NNs; with the identified quantities we empirically find that AR is achieved by regularizing/biasing NNs towards less confident solutions by making the changes in the feature space (induced by changes in the instance space) of most layers smoother uniformly in all directions; so to a certain extent, it prevents sudden change in prediction w.r.t. perturbations. However, the end result of such smoothing concentrates samples around decision boundaries, resulting in less confident solutions, and leads to worse standard performance. Our studies suggest that one might consider ways that build AR into NNs in a gentler way to avoid the problematic regularization. Yuxin Wen, Kui Jia |
ICML | 1 |
| 2020 | Performance Evaluation of Probabilistic Methods Based on Bootstrap and Quantile Regression to Quantify PV Power Point Forecast UncertaintyabstractThis paper presents two probabilistic approaches based on bootstrap method and quantile regression (QR) method to estimate the uncertainty associated with solar photovoltaic (PV) power point forecasts. Solar PV output power forecasts are obtained using a hybrid intelligent model, which is composed of a data filtering technique based on wavelet transform (WT) and a soft computing model (SCM) based on radial basis function neural network (RBFNN) that is optimized by particle swarm optimization (PSO) algorithm. The point forecast capability of the proposed hybrid WT+RBFNN+PSO intelligent model is examined and compared with other hybrid models as well as individual SCM. The performance of the proposed bootstrap method in the form of probabilistic forecasts is compared with the QR method by generating different prediction intervals (PIs). Numerical tests using real data demonstrate that the point forecasts obtained from the proposed hybrid intelligent model can be effectively used to quantify PV power uncertainty. The performance of these two uncertainty quantification methods is assessed through reliability. Yuxin Wen, Donna AlHakeem, Paras Mandal, Shantanu Chakraborty, Yuan-Kang Wu, Tomonobu Senjyu, Sumit Paudyal, Tzu-Liang (Bill) Tseng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Learning to Discover Curbside Parking Spaces from Vehicle TrajectoriesabstractThe increasing number of vehicles makes it difficult to find parking spaces, especially for the citizens living in metropolises. Therefore, many people prefer parking at curbsides for convenience’s sake regardless of whether they are authorized or not for parking. It tends to cause a severe problem of traffic jams. In order to address the issue, urban planners are eager to mitigate the stress of the traffic by discovering more curbside parking spaces. The technique of image processing supports a modern solution that can only cover a small range of roads with the aid of CCTV surveillance systems. In this paper, we propose to leverage large-scale vehicle trajectories in Baidu Maps to help detect more popular curbside parking spaces, ensuring broad coverage of roads. However, the main challenge we encountered is that only a small proportion of curbside parking spaces are authorized/labeled. To tackle this challenge, we distill several effective features and develop an approach based on a weak-labeled active learning framework. The framework can utilize very few annotated training samples to label the rest unlabeled samples iteratively. In each round, our approach can adopt several pre-trained classifiers that are ensembled to find more positive or negative samples. We use two real-world datasets to measure the performance of several modern methods of curbside parking space detection. Extensive experimental results demonstrate that our method achieves significant improvements in Fl-score over the other approaches on the two datasets. Yuxin Wen, Jizhou Huang, Chongli Zhu, Ying Li 0123 |
IEEE BigData | 1 |
| 2019 | Multiple-Change-Point Modeling and Exact Bayesian Inference of Degradation Signal for Prognostic ImprovementabstractPrognostics play an increasingly important role in modern engineering systems for smart maintenance decision-making. In parametric regression-based approaches, the parametric models are often too rigid to model degradation signals in many applications. In this paper, we propose a Bayesian multiple-change-point (CP) modeling framework to better capture the degradation path and improve the prognostics. At the offline modeling stage, a novel stochastic process is proposed to model the joint prior of CPs and positions. All hyperparameters are estimated through an empirical two-stage process. At the online monitoring and remaining useful life (RUL) prediction stage, a recursive updating algorithm is developed to exactly calculate the posterior distribution and RUL prediction sequentially. To control the computational cost, a fixed-support-size strategy in the online model updating and a partial Monte Carlo strategy in the RUL prediction are proposed. The effectiveness and advantages of the proposed method are demonstrated through thorough simulation and real case studies. Yuxin Wen, Qiang Zhou 0002, Tzu-Liang (Bill) Tseng |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2017 | Multiple-Phase Modeling of Degradation Signal for Condition Monitoring and Remaining Useful Life PredictionabstractRemaining useful life prediction plays an important role in ensuring the safety, availability, and efficiency of various engineering systems. In this paper, we propose a flexible Bayesian multiple-phase modeling approach to characterize degradation signals for prognosis. The priors are specified with a novel stochastic process and the multiple-phase model is formulated to a novel state-space model to facilitate online monitoring and prediction. A particle filtering algorithm with stratified sampling and partial Gibbs resample-move strategy is developed for online model updating and residual life prediction. The advantages of the proposed method are demonstrated through extensive numerical studies and real case studies. Yuxin Wen, Yuan Yuan 0037 |
IEEE Trans. Reliab. | 1 |