Xingyu Xie

dblp:174/9633 · DBLP profile ↗
← Back
37ranked-venue papers
11as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 (Dis)Proving Spectre Security with Speculation-Passing Style
abstract
Constant-time (CT) verification tools are commonly used for detecting potential side-channel vulnerabilities in cryptographic libraries. Recently, a new class of tools, called speculative constant-time (SCT) tools, has also been used for detecting potential Spectre vulnerabilities. In many cases, these SCT tools have emerged as liftings of CT tools. However, these liftings are seldom defined precisely and are almost never analyzed formally. The goal of this paper is to address this gap, by developing formal foundations for these liftings, and to demonstrate that these foundations can yield practical benefits. Concretely, we introduce a program transformation, coined Speculation-Passing Style (SPS), for reducing SCT verification to CT verification. Essentially, the transformation instruments the program with a new input that corresponds to attacker-controlled predictions and modifies the program to follow them. This approach is sound and complete, in the sense that a program is SCT if and only if its SPS transform is CT. Thus, we can leverage existing CT verification tools to prove SCT; we illustrate this by combining SPS with three standard methodologies for CT verification, namely reducing it to noninterference, assertion safety, and dynamic taint analysis. We realize these combinations with three existing tools, EasyCrypt, Binsec/Rel , and CTGrind , and we evaluate them on Kocher’s benchmarks for Spectre-v1. Our results focus on Spectre-v1 in the standard CT leakage model; however, we also discuss applications of our method to other variants of Spectre and other leakage models.
Santiago Arranz-Olmos, Gilles Barthe, Lionel Blatter, Xingyu Xie, Zhiyuan Zhang 0005
Proc. ACM Program. Lang.4
2025 SEPARATE: A Simple Low-rank Projection for Gradient Compression in Modern Large-scale Model Training Process
abstract
Training Large Language Models (LLMs) presents a significant communication bottleneck, predominantly due to the growing scale of the gradient to communicate across multi-device clusters. However, how to mitigate communication overhead in practice remains a formidable challenge due to the weakness of the methodology of the existing compression methods, especially the neglect of the characteristics of the gradient. In this paper, we consider and demonstrate the low-rank properties of gradient and Hessian observed in LLMs training dynamic, and take advantage of such natural properties to design SEPARATE, a simple low-rank projection for gradient compression in modern large-scale model training processes. SEPARATE realizes dimensional reduction by common random Gaussian variables and an improved moving average error-feedback technique. We theoretically demonstrate that SEPARATE-based optimizers maintain the original convergence rate for SGD and Adam-Type optimizers for general non-convex objectives. Experimental results show that SEPARATE accelerates training speed by up to 2× for GPT-2-Medium pre-training, and improves performance on various benchmarks for LLAMA2-7B fine-tuning.
Hanzhen Zhao, Xingyu Xie, Cong Fang 0001, Zhouchen Lin
ICLR2
2025 GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
abstract
Speculative decoding accelerates inference in large language models (LLMs) by generating multiple draft tokens simultaneously. However, existing methods often struggle with token misalignment between the training and decoding phases, limiting their performance. To address this, we propose GRIFFIN, a novel framework that incorporates a token-alignable training strategy and a token-alignable draft model to mitigate misalignment. The training strategy employs a loss masking mechanism to exclude highly misaligned tokens during training, preventing them from negatively impacting the draft model's optimization. The token-alignable draft model introduces input tokens to correct inconsistencies in generated features. Experiments on LLaMA, Vicuna, Qwen and Mixtral models demonstrate that GRIFFIN achieves an average acceptance length improvement of over 8\% and a speedup ratio exceeding 7\%, outperforming current speculative decoding state-of-the-art methods. Our code and GRIFFIN's draft models will be released publicly in https://github.com/hsj576/GRIFFIN.
Shijing Hu 0001, Xingyu Xie, Zhihui Lu 0002, Kim-Chuan Toh, Pan Zhou 0002
NeurIPS3
2025 LoCo: Low-Bit Communication Adaptor for Large-Scale Model Training
abstract
To efficiently train large-scale models, low-bit gradient communication compresses full-precision gradients on local GPU nodes into low-precision ones for higher gradient synchronization efficiency among GPU nodes. However, it often degrades training quality due to compression information loss. To address this, we propose the Low-bit Communication Adaptor (LoCo), which compensates gradients on local GPU nodes before compression, ensuring efficient synchronization without compromising training quality. Specifically, LoCo designs a moving average of historical compensation errors to stably estimate concurrent compression error and then adopts it to compensate for the concurrent gradient compression, yielding a less lossless compression. This mechanism allows it to be compatible with general optimizers like Adam and sharding strategies like FSDP. Theoretical analysis shows that integrating LoCo into full-precision optimizers like Adam and SGD does not impair their convergence speed on non-convex problems. Experimental results show that across large-scale model training frameworks like Megatron-LM and PyTorch's FSDP, LoCo significantly improves communication efficiency, e.g., improving Adam's training speed by 14% to 40% without performance degradation on large language models like LLAMAs and MoEs.
Xingyu Xie, Zhijie Lin 0001, Kim-Chuan Toh, Pan Zhou 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Multistage Diffusion Model With Phase Error Correction for Fast PET Imaging
abstract
Fast PET imaging is clinically important for reducing motion artifacts and improving patient comfort. While recent diffusion-based deep learning methods have shown promise, they often fail to capture the true PET degradation process, suffer from accumulated inference errors, introduce artifacts, and require extensive reconstruction iterations. To address these challenges, we propose a novel multistage diffusion framework tailored for fast PET imaging. At the coarse level, we design a multistage structure to approximate the temporal non-linear PET degradation process in a data-driven manner, using paired PET images collected under different acquisition duration. A Phase Error Correction Network (PECNet) ensures consistency across stages by correcting accumulated deviations. At the fine level, we introduce a deterministic cold diffusion mechanism, which simulates intra-stage degradation through interpolation between known acquisition durations—significantly reducing reconstruction iterations to as few as 10. Evaluations on [68Ga]FAPI and [18F]FDG PET datasets demonstrate the superiority of our approach, achieving peak PSNRs of 36.2 dB and 39.0 dB, respectively, with average SSIMs over 0.97. Our framework offers high-fidelity PET imaging with fewer iterations, making it practical for accelerated clinical imaging.
Zhenxing Huang, Xingyu Xie, Qianyi Yang, Xinlan Yang, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Ruohua Chen, Zhanli Hu
IEEE J. Biomed. Health Informatics3
2025 Prompt-Agent-Driven Integration of Foundation Model Priors for Low-Count PET Reconstruction
abstract
Low-count Positron Emission Tomography reconstruction is critical for maintaining high imaging quality while minimizing tracer doses and radiation exposure. Although integrating structural information from CT and MR data has been shown to enhance PET reconstruction, this typically requires simultaneous PET and CT/MRI scans, complicating workflows and increasing radiation exposure. Recent advancements in foundation models offer a promising alternative to in-person CT/MRI imaging, potentially overcoming these limitations. However, the use of foundation models' segmentation masks as semantic guides has been observed to introduce erroneous structures in low-count PET reconstructions. To address this challenge, this work introduces an innovative prompting agent-based framework that dynamically interacts with the foundation model to retrieve and refine priors, minimizing undue influence on the reconstruction process. Specifically, a box agent is designed for single-instance local area information retrieval, while a point agent is introduced to progressively prompt broader semantic structures globally, utilizing history point prompts. Additionally, an MDP paradigm has been developed to address the challenges of utilizing historical point prompts while maintaining the independence required by MDPs. Evaluated on both simulated and real datasets, the proposed method demonstrates superior qualitative and quantitative performance compared to state-of-the-art methods, even those leveraging in-person CT/MRI priors.
Xingyu Xie, Mu Nan, Yaping Wu, Hairong Zheng, Dong Liang 0001, Zhanli Hu
IEEE Trans. Medical Imaging1
2024 GAuV: A Graph-Based Automated Verification Framework for Perfect Semi-Honest Security of Multiparty Computation Protocols
abstract
Proving the security of a Multiparty Computation (MPC) protocol is a difficult task. Under the current simulation-based definition of MPC, a security proof consists of a simulator, which is usually specific to the concrete protocol and requires to be manually constructed, together with a theoretical analysis of the output distribution of the simulator and corrupted parties’ views in the real world. This presents an obstacle in verifying the security of a given MPC protocol. Moreover, an instance of a secure MPC protocol can easily lose its security guarantee due to careless implementation, and such a security issue is hard to detect in practice.(p)(/p)In this work, we propose a general automated framework to verify the perfect security of instances of MPC protocols against the semi-honest adversary. Our framework has perfect soundness: any protocol that is proven secure under our framework is also secure under the simulation-based definition of MPC. We demonstrate the completeness of our framework by showing that for any instance of the well-known BGW protocol, our framework can prove its security for every corrupted party set with polynomial time. Unlike prior work that only focuses on black-box privacy which requires the outputs of corrupted parties to contain no information about the inputs of the honest parties, our framework may potentially be used to prove the security of arbitrary MPC protocols. (p)(/p)We implement our framework as a prototype. The evaluation shows that our prototype automatically proves the perfect semi-honest security of BGW protocols and B2A (binary to arithmetic) conversion protocols in reasonable durations.
Xingyu Xie, Tuowei Wang, Shizhen Xu
SP1
2024 Win: Weight-Decay-Integrated Nesterov Acceleration for Faster Network Training
abstract
Training deep networks on large-scale datasets is computationally challenging. This work explores the problem of “how to accelerate adaptive gradient algorithms in a general manner", and proposes an effective Weight-decay-Integrated Nesterov acceleration (Win) to accelerate adaptive algorithms. Taking AdamW and Adam as examples, per iteration, we construct a dynamical loss that combines the vanilla training loss and a dynamic regularizer inspired by proximal point method, and respectively minimize the first- and second-order Taylor approximations of dynamical loss to update variable. This yields our Win acceleration that uses a conservative step and an aggressive step to update, and linearly combines these two updates for acceleration. Next, we extend Win into Win2 which uses multiple aggressive update steps for faster convergence. Then we apply Win and Win2 to the popular LAMB and SGD optimizers. Our transparent derivation could provide insights for other accelerated methods and their integration into adaptive algorithms. Besides, we theoretically justify the faster convergence of Win- and Win2-accelerated AdamW, Adam and LAMB to their non-accelerated counterparts. Experimental results demonstrates the faster convergence speed and superior performance of our Win- and Win2-accelerated AdamW, Adam, LAMB and SGD over their vanilla counterparts on vision classification and language modeling tasks.
Pan Zhou 0002, Xingyu Xie, Zhouchen Lin, Kim-Chuan Toh, Shuicheng Yan
J. Mach. Learn. Res.2
2024 Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
abstract
In deep learning, different kinds of deep networks typically need different optimizers, which have to be chosen after multiple trials, making the training process inefficient. To relieve this issue and consistently improve the model training speed across deep networks, we propose the ADAptive Nesterov momentum algorithm, Adan for short. Adan first reformulates the vanilla Nesterov acceleration to develop a new Nesterov momentum estimation (NME) method, which avoids the extra overhead of computing gradient at the extrapolation point. Then Adan adopts NME to estimate the gradient's first- and second-order moments in adaptive gradient algorithms for convergence acceleration. Besides, we prove that Adan finds an$\epsilon$-approximate first-order stationary point within$\mathcal {O}(\epsilon ^{-3.5})$stochastic gradient complexity on the non-convex stochastic problems (e.g., deep learning problems), matching the best-known lower bound. Extensive experimental results show that Adan consistently surpasses the corresponding SoTA optimizers on vision, language, and RL tasks and sets new SoTAs for many popular networks and frameworks, e.g., ResNet, ConvNext, ViT, Swin, MAE, DETR, GPT-2, Transformer-XL, and BERT. More surprisingly, Adan can use half of the training cost (epochs) of SoTA optimizers to achieve higher or comparable performance on ViT, GPT-2, MAE,etc, and also shows great tolerance to a large range of minibatch size, e.g., from 1 k to 32 k.
Xingyu Xie, Pan Zhou 0002, Huan Li 0007, Zhouchen Lin, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Towards Understanding Convergence and Generalization of AdamW
abstract
AdamW modifies Adam by adding a decoupled weight decay to decay network weights per training iteration. For adaptive algorithms, this decoupled weight decay does not affect specific optimization steps, and differs from the widely used$\ell _{2}$-regularizer which changes optimization steps via changing the first- and second-order gradient moments. Despite its great practical success, for AdamW, its convergence behavior and generalization improvement over Adam and$\ell _{2}$-regularized Adam ($\ell _{2}$-Adam) remain absent yet. To solve this issue, we prove the convergence of AdamW and justify its generalization advantages over Adam and$\ell _{2}$-Adam. Specifically, AdamW provably converges but minimizes a dynamically regularized loss that combines vanilla loss and a dynamical regularization induced by decoupled weight decay, thus yielding different behaviors with Adam and$\ell _{2}$-Adam. Moreover, on both general nonconvex problems and PŁ-conditioned problems, we establish stochastic gradient complexity of AdamW to find a stationary point. Such complexity is also applicable to Adam and$\ell _{2}$-Adam, and improves their previously known complexity, especially for over-parametrized networks. Besides, we prove that AdamW enjoys smaller generalization errors than Adam and$\ell _{2}$-Adam from the Bayesian posterior aspect. This result, for the first time, explicitly reveals the benefits of decoupled weight decay in AdamW. Experimental results validate our theory.
Pan Zhou 0002, Xingyu Xie, Zhouchen Lin, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Global Convergence of Over-parameterized Deep Equilibrium Models
abstract
A deep equilibrium model (DEQ) is implicitly defined through an equilibrium point of an infinite-depth weight-tied model with an input-injection. Instead of infinite computations, it solves an equilibrium point directly with root-finding and computes gradients with implicit differentiation. In this paper, the training dynamics of over-parameterized DEQs are investigated, and we propose a novel probabilistic framework to overcome the challenge arising from the weight-sharing and the infinite depth. By supposing a condition on the initial equilibrium point, we prove that the gradient descent converges to a globally optimal solution at a linear convergence rate for the quadratic loss function. We further perform a fine-grained non-asymptotic analysis about random DEQs and the corresponding weight-untied models, and show that the required initial condition is satisfied via mild over-parameterization. Moreover, we show that the unique equilibrium point always exists during the training.
Zenan Ling, Xingyu Xie, Qiuhao Wang, Zongpeng Zhang, Zhouchen Lin
AISTATS2
2023 Win: Weight-Decay-Integrated Nesterov Acceleration for Adaptive Gradient Algorithms
Pan Zhou 0002, Xingyu Xie, Shuicheng Yan
ICLR2
2023 EditAnything: Empowering Unparalleled Flexibility in Image Editing and Generation
abstract
Image editing plays a vital role in computer vision field, aiming to realistically manipulate images while ensuring seamless integration. It finds numerous applications across various fields. In this work, we present EditAnything, a novel approach that empowers users with unparalleled flexibility in editing and generating image content. EditAnything introduces an array of advanced features, including cross-image dragging (e.g., try-on), region-interactive editing, controllable layout generation, and virtual character replacement. By harnessing these capabilities, users can engage in interactive and flexible editing, giving captivating outcomes that uphold the integrity of the original image. With its diverse range of tools, EditAnything caters to a wide spectrum of editing needs, pushing the boundaries of image editing and unlocking exciting new possibilities. The source code is released at https://github.com/sail-sg/EditAnything.
Shanghua Gao, Zhijie Lin 0001, Xingyu Xie, Pan Zhou 0002, Ming-Ming Cheng, Shuicheng Yan
ACM Multimedia3
2023 Task-Robust Pre-Training for Worst-Case Downstream Adaptation
abstract
Pre-training has achieved remarkable success when transferred to downstream tasks. In machine learning, we care about not only the good performance of a model but also its behavior under reasonable shifts of condition. The same philosophy holds when pre-training a foundation model. However, the foundation model may not uniformly behave well for a series of related downstream tasks. This happens, for example, when conducting mask recovery regression where the recovery ability or the training instances diverge like pattern features are extracted dominantly on pre-training, but semantic features are also required on a downstream task. This paper considers pre-training a model that guarantees a uniformly good performance over the downstream tasks. We call this goal as *downstream-task robustness*. Our method first separates the upstream task into several representative ones and applies a simple minimax loss for pre-training. We then design an efficient algorithm to solve the minimax loss and prove its convergence in the convex setting. In the experiments, we show both on large-scale natural language processing and computer vision datasets our method increases the metrics on worse-case downstream tasks. Additionally, some theoretical explanations for why our loss is beneficial are provided. Specifically, we show fewer samples are inherently required for the most challenging downstream task in some cases.
Jianghui Wang, Xingyu Xie, Cong Fang 0001, Zhouchen Lin
NeurIPS3
2023 Learning representation via indirect feature decorrelation with bi-vector-based contrastive learning for clustering
Xingyu Xie, Lei Zhang 0005, Yan Wang 0015, Zizhou Wang
Inf. Sci.1
2023 On the methodology of three-way structured merge in version control systems: Top-down, bottom-up, or both
Fengmin Zhu, Xingyu Xie, Dongyu Feng, Na Meng 0001, Fei He 0001
J. Syst. Archit.2
2023 Optimization Induced Equilibrium Networks: An Explicit Optimization Perspective for Understanding Equilibrium Models
abstract
To reveal the mystery behind deep neural networks (DNNs), optimization may offer a good perspective. There are already some clues showing the strong connection between DNNs and optimization problems, e.g., under a mild condition, DNN's activation function is indeed a proximal operator. In this paper, we are committed to providing a unified optimization induced interpretability for a special class of networks-equilibrium models, i.e., neural networks defined by fixed point equations, which have become increasingly attractive recently. To this end, we first decompose DNNs into a new class of unit layer that is the proximal operator of an implicit convex function while keeping its output unchanged. Then, the equilibrium model of the unit layer can be derived, we name it Optimization Induced Equilibrium Networks (OptEq). The equilibrium point of OptEq can be theoretically connected to the solution of a convex optimization problem with explicit objectives. Based on this, we can flexibly introduce prior properties to the equilibrium points: 1) modifying the underlying convex problems explicitly so as to change the architectures of OptEq; and 2) merging the information into the fixed point iteration, which guarantees to choose the desired equilibrium point when the fixed point set is non-singleton. We show that OptEq outperforms previous implicit models even with fewer parameters.
Xingyu Xie, Qiuhao Wang, Zenan Ling, Xia Li 0005, Guangcan Liu, Zhouchen Lin
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Convolutional Feature Descriptor Selection for Mammogram Classification
abstract
Breast cancer was the most commonly diagnosed cancer among women worldwide in 2020. Recently, several deep learning-based classification approaches have been proposed to screen breast cancer in mammograms. However, most of these approaches require additional detection or segmentation annotations. Meanwhile, some other image-level label-based methods often pay insufficient attention to lesion areas, which are critical for diagnosis. This study designs a novel deep-learning method for automatically diagnosing breast cancer in mammography, which focuses on the local lesion areas and only utilizes image-level classification labels. In this study, we propose to select discriminative feature descriptors from feature maps instead of identifying lesion areas using precise annotations. And we design a novel adaptive convolutional feature descriptor selection (AFDS) structure based on the distribution of the deep activation map. Specifically, we adopt the triangle threshold strategy to calculate a specific threshold for guiding the activation map to determine which feature descriptors (local areas) are discriminative. Ablation experiments and visualization analysis indicate that the AFDS structure makes the model easier to learn the difference between malignant and benign/normal lesions. Furthermore, since the AFDS structure can be regarded as a highly efficient pooling structure, it can be easily plugged into most existing convolutional neural networks with negligible effort and time consumption. Experimental results on two publicly available INbreast and CBIS-DDSM datasets indicate that the proposed method performs satisfactorily compared with state-of-the-art methods.
Dong Li 0051, Lei Zhang 0005, Jianwei Zhang 0016, Xingyu Xie
IEEE J. Biomed. Health Informatics4
2022 High Quality Segmentation for Ultra High-resolution Images
abstract
To segment 4K or 6K ultra high-resolution images needs extra computation consideration in image segmentation. Common strategies, such as downsampling, patch cropping, and cascade model, cannot address well the balance issue between accuracy and computation cost. Motivated by the fact that humans distinguish among objects continuously from coarse to precise levels, we propose the Continuous Refinement Model (CRM) for the ultra high-resolution segmentation refinement task. CRM continuously aligns the feature map with the refinement target and aggregates features to reconstruct these image details. Besides, our CRM shows its significant generalization ability to fill the resolution gap between low-resolution training images and ultra high-resolution testing ones. We present quantitative performance evaluation and visualization to show that our proposed method is fast and effective on image segmentation refinement. Code is available at https://github.com/dvlab-research/Entity/tree/main/CRM.
Tiancheng Shen, Yuechen Zhang, Lu Qi 0001, Jason Kuen, Xingyu Xie, Jianlong Wu, Zhe Lin 0001, Jiaya Jia
CVPR5
2022 Optimization inspired Multi-Branch Equilibrium Models
Mingjie Li 0007, Yisen Wang 0001, Xingyu Xie, Zhouchen Lin
ICLR3
2022 Mastery: Shifted-Code-Aware Structured Merging
Fengmin Zhu, Xingyu Xie, Dongyu Feng, Na Meng 0001, Fei He 0001
SETTA2
2021 Prob-CLR: A probabilistic approach to learn discriminative representation
Xingyu Xie, Minjuan Zhu, Yan Wang 0015, Lei Zhang 0005
Knowl. Based Syst.1
2020 Unified Graph and Low-Rank Tensor Learning for Multi-View Clustering
abstract
Multi-view clustering aims to take advantage of multiple views information to improve the performance of clustering. Many existing methods compute the affinity matrix by low-rank representation (LRR) and pairwise investigate the relationship between views. However, LRR suffers from the high computational cost in self-representation optimization. Besides, compared with pairwise views, tensor form of all views' representation is more suitable for capturing the high-order correlations among all views. Towards these two issues, in this paper, we propose the unified graph and low-rank tensor learning (UGLTL) for multi-view clustering. Specifically, on the one hand, we learn the view-specific affinity matrix based on projected graph learning. On the other hand, we reorganize the affinity matrices into tensor form and learn its intrinsic tensor based on low-rank tensor approximation. Finally, we unify these two terms together and jointly learn the optimal projection matrices, affinity matrices and intrinsic low-rank tensor. We also propose an efficient algorithm to iteratively optimize the proposed model. To evaluate the performance of the proposed method, we conduct extensive experiments on multiple benchmarks across different scenarios and sizes. Compared with the state-of-the-art approaches, our method achieves much better performance.
Jianlong Wu, Xingyu Xie, Liqiang Nie, Zhouchen Lin, Hongbin Zha
AAAI2
2020 Maximum-and-Concatenation Networks
abstract
While successful in many fields, deep neural networks (DNNs) still suffer from some open problems such as bad local minima and unsatisfactory generalization performance. In this work, we propose a novel architecture called Maximum-and-Concatenation Networks (MCN) to try eliminating bad local minima and improving generalization ability as well. Remarkably, we prove that MCN has a very nice property; that is, every local minimum of an (l+1)-layer MCN can be better than, at least as good as, the global minima of the network consisting of its first l layers. In other words, by increasing the network depth, MCN can autonomously improve its local minima’s goodness, what is more, it is easy to plug MCN into an existing deep model to make it also have this property. Finally, under mild conditions, we show that MCN can approximate certain continuous function arbitrarily well with high efficiency; that is, the covering number of MCN is much smaller than most existing DNNs such as deep ReLU. Based on this, we further provide a tight generalization bound to guarantee the inference ability of MCN when dealing with testing samples.
Xingyu Xie, Hao Kong 0002, Jianlong Wu, Wayne Zhang 0001, Guangcan Liu, Zhouchen Lin
ICML1
2020 NormalF-Net: Normal Filtering Neural Network for Feature-preserving Mesh Denoising
Zhiqi Li 0002, Yingkui Zhang, Yidan Feng, Xingyu Xie, Qiong Wang 0001, Mingqiang Wei, Pheng-Ann Heng
Comput. Aided Des.4
2020 Multi-Patch Collaborative Point Cloud Denoising via Low-Rank Recovery with Graph Constraint
abstract
Point cloud is the primary source from 3D scanners and depth cameras. It usually contains more raw geometric features, as well as higher levels of noise than the reconstructed mesh. Although many mesh denoising methods have proven to be effective in noise removal, they hardly work well on noisy point clouds. We propose a new multi-patch collaborative method for point cloud denoising, which is solved as a low-rank matrix recovery problem. Unlike the traditional single-patch based denoising approaches, our approach is inspired by the geometric statistics which indicate that a number of surface patches sharing approximate geometric properties always exist within a 3D model. Based on this observation, we define a rotation-invariant height-map patch (HMP) for each point by robust Bi-PCA encoding bilaterally filtered normal information, and group its non-local similar patches together. Within each group, all patches are geometrically similar, while suffering from noise. We pack the height maps of each group into an HMP matrix, whose initial rank is high, but can be significantly reduced. We design an improved low-rank recovery model, by imposing a graph constraint to filter noise. Experiments on synthetic and raw datasets demonstrate that our method outperforms state-of-the-art methods in both noise removal and feature preservation.
Honghua Chen, Mingqiang Wei, Yangxing Sun, Xingyu Xie, Jun Wang 0039
IEEE Trans. Vis. Comput. Graph.4
2019 Differentiable Linearized ADMM
abstract
Recently, a number of learning-based optimization methods that combine data-driven architectures with the classical optimization algorithms have been proposed and explored, showing superior empirical performance in solving various ill-posed inverse problems, but there is still a scarcity of rigorous analysis about the convergence behaviors of learning-based optimization. In particular, most existing analyses are specific to unconstrained problems but cannot apply to the more general cases where some variables of interest are subject to certain constraints. In this paper, we propose Differentiable Linearized ADMM (D-LADMM) for solving the problems with linear constraints. Specifically, D-LADMM is a K-layer LADMM inspired deep neural network, which is obtained by firstly introducing some learnable weights in the classical Linearized ADMM algorithm and then generalizing the proximal operator to some learnable activation function. Notably, we rigorously prove that there exist a set of learnable parameters for D-LADMM to generate globally converged solutions, and we show that those desired parameters can be attained by training D-LADMM in a proper way. To the best of our knowledge, we are the first to provide the convergence analysis for the learning-based optimization method on constrained problems.
Xingyu Xie, Jianlong Wu, Guangcan Liu, Zhisheng Zhong, Zhouchen Lin
ICML1
2019 Neural Ordinary Differential Equations with Envolutionary Weights
Lingshen He, Xingyu Xie, Zhouchen Lin
PRCV (1)2
2019 Matrix recovery with implicitly low-rank data
Xingyu Xie, Jianlong Wu, Guangcan Liu, Jun Wang 0039
Neurocomputing1
2019 Robust Low-rank subspace segmentation with finite mixture noise
Xianglin Guo, Xingyu Xie, Guangcan Liu, Mingqiang Wei, Jun Wang 0039
Pattern Recognit.2
2019 Mesh Denoising Guided by Patch Normal Co-Filtering via Kernel Low-Rank Recovery
abstract
Mesh denoising is a classical, yet not well-solved problem in digital geometry processing. The challenge arises from noise removal with the minimal disturbance of surface intrinsic properties (e.g., sharp features and shallow details). We propose a new patch normal co-filter (PcFilter) for mesh denoising. It is inspired by the geometry statistics which show that surface patches with similar intrinsic properties exist on the underlying surface of a noisy mesh. We model the PcFilter as a low-rank matrix recovery problem of similar-patch collaboration, aiming at removing different levels of noise, yet preserving various surface features. We generalize our model to pursue the low-rank matrix recovery in the kernel space for handling the nonlinear structure contained in the data. By making use of the block coordinate descent minimization and the specifics of a proximal based coordinate descent method, we optimize the nonlinear and nonconvex objective function efficiently. The detailed quantitative and qualitative results on synthetic and real data show that the PcFilter competes favorably with the state-of-the-art methods in surface accuracy and noise-robustness.
Mingqiang Wei, Xingyu Xie, Ligang Liu 0001, Jun Wang 0039, Harry Qin
IEEE Trans. Vis. Comput. Graph.3
2018 Redundancy-resistant Generative Hashing for Image Retrieval
abstract
By optimizing probability distributions over discrete latent codes, Stochastic Generative Hashing (SGH) bypasses the critical and intractable binary constraints on hash codes. While encouraging results were reported, SGH still suffers from the deficient usage of latent codes, i.e., there often exist many uninformative latent dimensions in the code space, a disadvantage inherited from its auto-encoding variational framework. Motivated by the fact that code redundancy usually is severer when more complex decoder network is used, in this paper, we propose a constrained deep generative architecture to simplify the decoder for data reconstruction. Specifically, our new framework forces the latent hashing codes to not only reconstruct data through the generative network but also retain minimal squared L2 difference to the last real-valued network hidden layer. Furthermore, during posterior inference, we propose to regularize the standard auto-encoding objective with an additional term that explicitly accounts for the negative redundancy degree of latent code dimensions. We interpret such modifications as Bayesian posterior regularization and design an adversarial strategy to optimize the generative, the variational, and the redundancy-resistanting parameters. Empirical results show that our new method can significantly boost the quality of learned codes and achieve state-of-the-art performance for image retrieval.
Changying Du, Xingyu Xie, Changde Du, Hao Wang 0005
IJCAI2
2018 Multi-view Adversarially Learned Inference for Cross-domain Joint Distribution Matching
abstract
Many important data mining problems can be modeled as learning a (bidirectional) multidimensional mapping between two data domains. Based on the generative adversarial networks (GANs), particularly conditional ones, cross-domain joint distribution matching is an increasingly popular kind of methods addressing such problems. Though significant advances have been achieved, there are still two main disadvantages of existing models, i.e., the requirement of large amount of paired training samples and the notorious instability of training. In this paper, we propose a multi-view adversarially learned inference (ALI) model, termed as MALI, to address these issues. Unlike the common practice of learning direct domain mappings, our model relies on shared latent representations of both domains and can generate arbitrary number of paired faking samples, benefiting from which usually very few paired samples (together with sufficient unpaired ones) is enough for learning good mappings. Extending the vanilla ALI model, we design novel discriminators to judge the quality of generated samples (both paired and unpaired), and provide theoretical analysis of our new formulation. Experiments on image-to-image translation, image-to-attribute generation (multi-label classification), attribute-to-image generation tasks demonstrate that our semi-supervised learning framework yields significant performance improvements over existing ones. Results on cross-modality retrieval show that our latent space based method can achieve competitive similarity search performance in relative fast speed, compared to those methods that compute similarities in the high-dimensional data space.
Changying Du, Changde Du, Xingyu Xie, Chen Zhang 0003, Hao Wang 0005
KDD3
2018 Implicit Block Diagonal Low-Rank Representation
abstract
While current block diagonal constrained subspace clustering methods are performed explicitly on the original data space, in practice, it is often more desirable to embed the block diagonal prior into the reproducing kernel Hilbert feature space by kernelization techniques, as the underlying data structure in reality is usually nonlinear. However, it is still unknown how to carry out the embedding and kernelization in the models with block diagonal constraints. In this paper, we shall take a step in this direction. First, we establish a novel model termed implicit block diagonal low-rank representation (IBDLR), by incorporating the implicit feature representation and block diagonal prior into the prevalent low-rank representation method. Second, mostly important, we show that the model in IBDLR could be kernelized by making use of a smoothed dual representation and the specifics of a proximal gradient-based optimization algorithm. Finally, we provide some theoretical analyses for the convergence of our optimization algorithm. Comprehensive experiments on synthetic and real-world data sets demonstrate the superiorities of our IBDLR over state-of-the-art methods.
Xingyu Xie, Xianglin Guo, Guangcan Liu, Jun Wang 0039
IEEE Trans. Image Process.1
2017 Surface reconstruction with data-driven exemplar priors
Oussama Remil, Qian Xie 0001, Xingyu Xie, Kai Xu 0004, Jun Wang 0039
Comput. Aided Des.3
2017 Data-Driven Sparse Priors of 3D Shapes
abstract
Abstract We present a sparse optimization framework for extracting sparse shape priors from a collection of 3D models. Shape priors are defined as point‐set neighborhoods sampled from shape surfaces which convey important information encompassing normals and local shape characterization. A 3D shape model can be considered to be formed with a set of 3D local shape priors, while most of them are likely to have similar geometry. Our key observation is that the local priors extracted from a family of 3D shapes lie in a very low‐dimensional manifold. Consequently, a compact and informative subset of priors can be learned to efficiently encode all shapes of the same family. A comprehensive library of local shape priors is first built with the given collection of 3D models of the same family. We then formulate a global, sparse optimization problem which enforces selecting representative priors while minimizing the reconstruction error. To solve the optimization problem, we design an efficient solver based on the Augmented Lagrangian Multipliers method (ALM). Extensive experiments exhibit the power of our data‐driven sparse priors in elegantly solving several high‐level shape analysis applications and geometry processing tasks, such as shape retrieval, style analysis and symmetry detection.
Oussama Remil, Qian Xie 0001, Xingyu Xie, Kai Xu 0004, Jun Wang 0039
Comput. Graph. Forum3
2016 Automatic Modeling of Urban Facades from Raw LiDAR Point Data
abstract
Abstract Modeling of urban facades from raw LiDAR point data remains active due to its challenging nature. In this paper, we propose an automatic yet robust 3D modeling approach for urban facades with raw LiDAR point clouds. The key observation is that building facades often exhibit repetitions and regularities. We hereby formulate repetition detection as an energy optimization problem with a global energy function balancing geometric errors, regularity and complexity of facade structures. As a result, repetitive structures are extracted robustly even in the presence of noise and missing data. By registering repetitive structures, missing regions are completed and thus the associated point data of structures are well consolidated. Subsequently, we detect the potential design intents (i.e., geometric constraints) within structures and perform constrained fitting to obtain the precise structure models. Furthermore, we apply structure alignment optimization to enforce position regularities and employ repetitions to infer missing structures. We demonstrate how the quality of raw LiDAR data can be improved by exploiting data redundancy, and discovering high level structural information (regularity and symmetry). We evaluate our modeling method on a variety of raw LiDAR scans to verify its robustness and effectiveness.
Jun Wang 0039, Yabin Xu, Oussama Remil, Xingyu Xie, Mingqiang Wei
Comput. Graph. Forum4