EDBT 2026 Demo / reviewers in the wild / expert
Yingzhen Yang
dblp:66/3838
· DBLP profile ↗
47ranked-venue papers
25as first author
19since 2021 · last 2025
0000-0003-0502-6122ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 19 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 10 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Declarative Privacy-Preserving Inference Queries
Ansh Tiwari, Summer Gautier, Rajan Hari Ambrish, Lixi Zhou, Yancheng Wang 0001, Deepti Gupta, Yingzhen Yang, Chaowei Xiao, Kanchan Chowdhury, Jia Zou 0001 |
DASFAA (6) | 8 |
| 2025 | Sketching for Convex and Nonconvex Regularized Least Squares with Sharp GuaranteesabstractRandomized algorithms play a crucial role in efficiently solving large-scale optimization problems. In this paper, we introduce Sketching for Regularized Optimization (SRO), a fast sketching algorithm designed for least squares problems with convex or nonconvex regularization. SRO operates by first creating a sketch of the original data matrix and then solving the sketched problem. We establish minimax optimal rates for sparse signal estimation by addressing the sketched sparse convex and nonconvex learning problems. Furthermore, we propose a novel Iterative SRO algorithm, which reduces the approximation error geometrically for sketched convex regularized problems. To the best of our knowledge, this work is among the first to provide a unified theoretical framework demonstrating minimax rates for convex and nonconvex sparse learning problems via sketching. Experimental results validate the efficiency and effectiveness of both the SRO and Iterative SRO algorithms. Yingzhen Yang, Ping Li 0001 |
ICLR | 1 |
| 2025 | Sharp Generalization for Nonparametric Regression by Over-Parameterized Neural Networks: A Distribution-Free Analysis in Spherical CovariateabstractSharp generalization bound for neural networks trained by gradient descent (GD) is of central interest in statistical learning theory and deep learning. In this paper, we consider nonparametric regression by an over-parameterized two-layer NN trained by GD. We show that, if the neural network is trained by GD with early stopping, then the trained network renders a sharp rate of the nonparametric regression risk of $\mathcal O(\epsilon_n^2)$, which is the same rate as that for the classical kernel regression trained by GD with early stopping, where $\epsilon_n$ is the critical population rate of the Neural Tangent Kernel (NTK) associated with the network and $n$ is the size of the training data. It is remarked that our result does not require distributional assumptions on the covariate as long as the covariate lies on the unit sphere, in a strong contrast with many existing results which rely on specific distributions such as the spherical uniform data distribution or distributions satisfying certain restrictive conditions. As a special case of our general result, when the eigenvalues of the associated NTK decay at a rate of $\lambda_j \asymp j^{-\frac{d}{d-1}}$ for $j \ge 1$ which happens under certain distributional assumption such as the training features follow the spherical uniform distribution, we immediately obtain the minimax optimal rate of $\mathcal O(n^{-\frac{d}{2d-1}})$, which is the major results of several existing works in this direction. The neural network width in our general result is lower bounded by a function of only $d$ and $\epsilon_n$, and such width does not depend on the minimum eigenvalue of the empirical NTK matrix whose lower bound usually requires additional assumptions on the training data. Our results are built upon two significant technical results which are of independent interest. First, uniform convergence to the NTK is established during the training process by GD, so that we can have a nice decomposition of the neural network function at any step of the GD into a function in the Reproducing Kernel Hilbert Space associated with the NTK and an error function with a small $L^{\infty}$-norm. Second, local Rademacher complexity is employed to tightly bound the Rademacher complexity of the function class comprising all the possible neural network functions obtained by GD. Our result formally fills the gap between training a classical kernel regression model and training an over-parameterized but finite-width neural network by GD for nonparametric regression without distributional assumptions about the spherical covariate. Yingzhen Yang |
ICML | 1 |
| 2025 | A New Concentration Inequality for Sampling Without Replacement and Its Application for Transductive LearningabstractWe introduce a new tool, Transductive Local Complexity (TLC), to analyze the generalization performance of transductive learning methods and motivate new transductive learning algorithms. Our work extends the idea of the popular Local Rademacher Complexity (LRC) to the transductive setting with considerable and novel changes compared to the analysis of typical LRC methods in the inductive setting. While LRC has been widely used as a powerful tool in the analysis of inductive models with sharp generalization bounds for classification and minimax rates for nonparametric regression, it remains an open problem whether a localized version of Rademacher complexity based tool can be designed and applied to transductive learning and gain sharp bound for transductive learning which is consistent with the inductive excess risk bound by (LRC). We give a confirmative answer to this open problem by TLC. Similar to the development of LRC, we build TLC by first establishing a novel and sharp concentration inequality for supremum of empirical processes for the gap between test and training loss in the setting of sampling uniformly without replacement. Then a peeling strategy and a new surrogate variance operator are used to derive the following excess risk bound in the transductive setting, which is consistent with that of the classical LRC based excess risk bound in the inductive setting. As an application of TLC, we use the new TLC tool to analyze the Transductive Kernel Learning (TKL) model, and derive sharper excess risk bound than that by the current state-of-the-art. As a result of independent interest, the concentration inequality for the test-train process is used to derive a sharp concentration inequality for the general supremum of empirical process involving random variables in the setting of sampling uniformly without replacement, with comparison to current concentration inequalities. Yingzhen Yang |
ICML | 1 |
| 2025 | Informative Synthetic Data Generation for Thorax Disease ClassificationabstractDeep Neural Networks (DNNs), including architectures such as Vision Transformers (ViTs), have achieved remarkable success in medical imaging tasks. However, their performance typically hinges on the availability of large-scale, high-quality labeled datasets-resources that are often scarce or infeasible to obtain in medical domains. Generative Data Augmentation (GDA) offers a promising remedy by supplementing training sets with synthetic data generated via generative models like Diffusion Models (DMs). Yet, this approach introduces a critical challenge: synthetic data often contains significant noise, which can degrade the performance of classifiers trained on such augmented datasets. Prior solutions, including data selection and re-weighting techniques, often rely on access to clean metadata or pretrained external classifiers. In this work, we propose \emph{Informative Data Selection} (IDS), a principled sample re-weighting framework grounded in the Information Bottleneck (IB) principle. IDS assigns higher weights to more informative synthetic samples, thereby improving classifier performance in GDA-enhanced training for thorax disease classification. Extensive experiments demonstrate that IDS significantly outperforms existing data selection and re-weighting baselines. Our code is publicly available at \url{https://github.com/Statistical-Deep-Learning/IDS}. Rajeev Goel, Marko Jojic, Alvin C. Silva, Teresa Wu, Yingzhen Yang |
UAI | 6 |
| 2025 | Deep Geometric Moments Promote Shape Consistency in Text-to-3D GenerationabstractTo address the data scarcity associated with 3D assets, 2D-lifting techniques such as Score Distillation Sampling (SDS) have become a widely adopted practice in text-to-3D generation pipelines. However, the diffusion models used in these techniques are prone to viewpoint bias and thus lead to geometric inconsistencies such as the Janus problem. To counter this, we introduce MT3D, a text-to-3D generative model that leverages a high-fidelity 3D object to overcome viewpoint bias and explicitly infuse geometric understanding into the generation pipeline. Firstly, we employ depth maps derived from a high-quality 3D model as control signals to guarantee that the generated 2D images preserve the funda-mental shape and structure, thereby reducing the inherent viewpoint bias. Next, we utilize deep geometric moments to ensure geometric consistency in the 3D representation explicitly. By incorporating geometric details from a 3D asset, MT3D enables the creation of diverse and geometri-cally consistent objects, thereby improving the quality and usability of our 3D representations. Project page and code: https://moment-3d.github.io/ Utkarsh Nath, Rajeev Goel, Eun Som Jeon, Changhoon Kim, Kyle Min 0001, Yezhou Yang, Yingzhen Yang, Pavan Turaga |
WACV | 7 |
| 2025 | Efficient Visual Transformer by Learnable Token MergingabstractSelf-attention and transformers have been widely used in deep learning. Recent efforts have been devoted to incorporating transformer blocks into different neural architectures, including those with convolutions, leading to various visual transformers for computer vision tasks. In this paper, we propose a novel and compact transformer block, Transformer with Learnable Token Merging (LTM), or LTM-Transformer. LTM-Transformer performs token merging in a learnable scheme. LTM-Transformer is compatible with many popular and compact transformer networks, and it reduces the FLOPs and the inference time of the visual transformers while maintaining or even improving the prediction accuracy. In the experiments, we replace all the transformer blocks in popular visual transformers, including MobileViT, EfficientViT, ViT, and Swin, with LTM-Transformer blocks, leading to LTM-Transformer networks with different backbones. The LTM-Transformer is motivated by reduction of Information Bottleneck, and a novel and separable variational upper bound for the IB loss is derived. The architecture of the mask module in our LTM blocks which generates the token merging mask is designed to reduce the derived upper bound for the IB loss. Extensive results on computer vision tasks evidence that LTM-Transformer renders compact and efficient visual transformers with comparable or much better prediction accuracy than the original visual transformers. Yancheng Wang 0001, Yingzhen Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | IDNet: A Novel Identity Document Dataset via Few-Shot and Quality-Driven Synthetic Data GenerationabstractEffective fraud detection and analysis of government-issued identity documents, such as passports, driver’s licenses, and identity cards, are essential in thwarting identity theft and bolstering security on online platforms. The accuracy of training fraud detection and analysis tools depends on the availability of extensive and diverse identity document datasets. However, current publicly available benchmark datasets for identity document analysis, including MIDV-500, MIDV-2020, and FMIDV, fall short in several aspects: they offer a limited number of samples of ten European country document types, cover insufficient varieties of fraud patterns, and seldom include alterations in critical personal identifying fields such as portrait images, limiting their utility in training models capable of detecting realistic frauds while preserving privacy. In response to these shortcomings, our research introduces a new benchmark dataset, IDNet, designed to advance privacy-preserving fraud detection efforts, synthesized by integrating the generative models and a Bayesian optimization approach. The IDNet dataset comprises 837, 060 images of synthetically generated identity documents, totaling approximately 490 gigabytes, categorized into 20 types from 10 U.S. states and 10 European countries, which is the largest identity document dataset publicly available today. We evaluated the fidelity and utility of IDNet to demonstrate the effectiveness of our unique synthetic data generation method. We also presented two use cases of the dataset, illustrating how it can aid in training privacy-preserving fraud detection methods, and facilitating the generation of camera and video capturing of identity documents. Lulu Xie, Yancheng Wang 0001, Soham Nag, Rajeev Goel, Niranjan Erappa Narayana Swamy, Yingzhen Yang, Chaowei Xiao, Jonathan Prisby, Ross Maciejewski, Jia Zou 0001 |
IEEE Big Data | 7 |
| 2024 | Visual Transformer with Differentiable Channel Selection: An Information Bottleneck Inspired ApproachabstractSelf-attention and transformers have been widely used in deep learning. Recent efforts have been devoted to incorporating transformer blocks into different types of neural architectures, including those with convolutions, leading to various visual transformers for computer vision tasks. In this paper, we propose a novel and compact transformer block, Transformer with Differentiable Channel Selection, or DCS-Transformer. DCS-Transformer features channel selection in the computation of the attention weights and the input/output features of the MLP in the transformer block. Our DCS-Transformer is compatible with many popular and compact transformer networks, such as MobileViT and EfficientViT, and it reduces the FLOPs of the visual transformers while maintaining or even improving the prediction accuracy. In the experiments, we replace all the transformer blocks in MobileViT and EfficientViT with DCS-Transformer blocks, leading to DCS-Transformer networks with different backbones. The DCS-Transformer is motivated by reduction of Information Bottleneck, and a novel variational upper bound for the IB loss which can be optimized by SGD is derived and incorporated into the training loss of the network with DCS-Transformer. Extensive results on image classification and object detection evidence that DCS-Transformer renders compact and efficient visual transformers with comparable or much better prediction accuracy than the original visual transformers. The code of DCS-Transformer is available at https://github.com/Statistical-Deep-Learning/DCS-Transformer. Ping Li 0001, Yingzhen Yang |
ICML | 3 |
| 2024 | Learning Low-Rank Feature for Thorax Disease ClassificationabstractDeep neural networks, including Convolutional Neural Networks (CNNs) and Visual Transformers (ViT), have achieved stunning success in the medical image domain. We study thorax disease classification in this paper. Effective extraction of features for the disease areas is crucial for disease classification on radiographic images. While various neural architectures and training techniques, such as self-supervised learning with contrastive/restorative learning, have been employed for disease classification on radiographic images, there are no principled methods that can effectively reduce the adverse effect of noise and background or non-disease areas on the radiographic images for disease classification. To address this challenge, we propose a novel Low-Rank Feature Learning (LRFL) method in this paper, which is universally applicable to the training of all neural networks. The LRFL method is both empirically motivated by a Low Frequency Property (LFP) and theoretically motivated by our sharp generalization bound for neural networks with low-rank features. LFP not only widely exists in deep neural networks for generic machine learning but also exists in all the thorax medical datasets studied in this paper. In the empirical study, using a neural network such as a ViT or a CNN pre-trained on unlabeled chest X-rays by Masked Autoencoders (MAE), our novel LRFL method is applied on the pre-trained neural network and demonstrates better classification results in terms of both multi-class area under the receiver operating curve (mAUC) and classification accuracy than the current state-of-the-art. The code of LRFL is available at \url{https://github.com/Statistical-Deep-Learning/LRFL}. Rajeev Goel, Utkarsh Nath, Alvin C. Silva, Teresa Wu, Yingzhen Yang |
NeurIPS | 6 |
| 2024 | Neural Architecture Search Finds Robust Models by Knowledge DistillationabstractDespite their superior performance, Deep Neural Networks (DNNs) are often vulnerable to adversarial attacks. Neural Architecture Search (NAS), a method for automatically designing the architectures of DNNs, has shown remarkable performance across various machine learning applications. However, the adversarial robustness of architectures learned by NAS against adversarial threats remains under-explored. By integrating a robust teacher, we examine whether NAS can yield a robust neural architecture by inheriting robustness from the teacher. In this paper, we propose Robust Neural Architecture Search by Cross-Layer Knowledge Distillation (RNAS-CL), a novel NAS algorithm that enhances the robustness of architectures learned by NAS through employing cross-layer knowledge distillation from a robust teacher. Distinct from previous knowledge distillation approaches that only align student-teacher outputs at the final layer, RNAS-CL dynamically searches for the optimal teacher layer to guide each student layer. Our experimental findings validate the effectiveness of RNAS-CL, demonstrating that it can generate both compact and adversarially robust neural architectures. Our results pave the way for developing new strategies for compact and robust neural architecture design applicable across various fields. The code of RNAS-CL is available at \url{https://github.com/Statistical-Deep-Learning/RNAS-CL}. Utkarsh Nath, Yingzhen Yang |
UAI | 3 |
| 2024 | RNAS-CL: Robust Neural Architecture Search by Cross-Layer Knowledge Distillation
Utkarsh Nath, Pavan Turaga, Yingzhen Yang |
Int. J. Comput. Vis. | 4 |
| 2023 | Eliciting Structural and Semantic Global Knowledge in Unsupervised Graph Contrastive LearningabstractGraph Contrastive Learning (GCL) has recently drawn much research interest for learning generalizable node representations in a self-supervised manner. In general, the contrastive learning process in GCL is performed on top of the representations learned by a graph neural network (GNN) backbone, which transforms and propagates the node contextual information based on its local neighborhoods. However, nodes sharing similar characteristics may not always be geographically close, which poses a great challenge for unsupervised GCL efforts due to their inherent limitations in capturing such global graph knowledge. In this work, we address their inherent limitations by proposing a simple yet effective framework -- Simple Neural Networks with Structural and Semantic Contrastive Learning} (S^3-CL). Notably, by virtue of the proposed structural and semantic contrastive learning algorithms, even a simple neural network can learn expressive node representations that preserve valuable global structural and semantic patterns. Our experiments demonstrate that the node representations learned by S^3-CL) achieve superior performance on different downstream tasks compared with the state-of-the-art unsupervised GCL methods. Implementation and more experimental details are publicly available at https://github.com/kaize0409/S-3-CL. Kaize Ding, Yancheng Wang 0001, Yingzhen Yang, Huan Liu 0001 |
AAAI | 3 |
| 2023 | Projective Proximal Gradient Descent for Nonconvex Nonsmooth Optimization: Fast Convergence Without Kurdyka-Lojasiewicz (KL) Property
Yingzhen Yang, Ping Li 0001 |
ICLR | 1 |
| 2023 | Locally Regularized Sparse Graph by Fast Proximal Gradient DescentabstractSparse graphs built by sparse representation has been demonstrated to be effective in clustering high-dimensional data. Albeit the compelling empirical performance, the vanilla sparse graph ignores the geometric information of the data by performing sparse representation for each datum separately. In order to obtain a sparse graph aligned with the local geometric structure of data, we propose a novel Support Regularized Sparse Graph, abbreviated as SRSG, for data clustering. SRSG encourages local smoothness on the neighborhoods of nearby data points by a well-defined support regularization term. We propose a fast proximal gradient descent method to solve the non-convex optimization problem of SRSG with the convergence matching the Nesterov’s optimal convergence rate of first-order methods on smooth and convex objective function with Lipschitz continuous gradient. Extensive experimental results on various real data sets demonstrate the superiority of SRSG over other competing clustering methods. Dongfang Sun, Yingzhen Yang |
UAI | 2 |
| 2022 | Discriminative Similarity for Data Clustering
Yingzhen Yang, Ping Li 0001 |
ICLR | 1 |
| 2022 | Benchmark of DNN Model Search at Deployment TimeabstractDeep learning has become the most popular direction in machine learning and artificial intelligence. However, the preparation of training data, as well as model training, are often time-consuming and become the bottleneck of the end-to-end machine learning lifecycle. Reusing models for inferring a dataset can avoid the costs of retraining. However, when there are multiple candidate models, it is challenging to discover the right model for reuse. Although there exist a number of model sharing platforms such as ModelDB, TensorFlow Hub, PyTorch Hub, and DLHub, most of these systems require model uploaders to manually specify the details of each model and model downloaders to screen keyword search results for selecting a model. We are lacking a highly productive model search tool that selects models for deployment without the need for any manual inspection and/or labeled data from the target domain. This paper proposes multiple model search strategies including various similarity-based approaches and non-similarity-based approaches. We design, implement and evaluate these approaches on multiple model inference scenarios, including activity recognition, image recognition, text classification, natural language processing, and entity matching. The experimental evaluation showed that our proposed asymmetric similarity-based measurement, adaptivity, outperformed symmetric similarity-based measurements and non-similarity-based measurements in most of the workloads. Lixi Zhou, Arindam Jain, Amitabh Das, Yingzhen Yang, Jia Zou 0001 |
SSDBM | 5 |
| 2022 | Noisy L0-sparse subspace clustering on dimensionality reduced dataabstractSparse subspace clustering methods with sparsity induced by L0-norm, such as L0-Sparse Subspace Clustering (L0-SSC), are demonstrated to be more effective than its L1 counterpart such as Sparse Subspace Clustering (SSC). However, the theoretical analysis of L0-SSC is restricted to clean data that lie exactly in subspaces. Real data often suffer from noise and they may lie close to subspaces. In this paper, we show that an optimal solution to the optimization problem of noisy L0-SSC achieves subspace detection property (SDP), a key element with which data from different subspaces are separated, under deterministic and semi-random model. Our results provide theoretical guarantee on the correctness of noisy L0-SSC in terms of SDP on noisy data for the first time, which reveals the advantage of noisy L0-SSC in terms of much less restrictive condition on subspace affinity. In order to improve the efficiency of noisy L0-SSC, we propose Noisy-DR-L0-SSC which provably recovers the subspaces on dimensionality reduced data. Noisy-DR-L0-SSC first projects the data onto a lower dimensional space by random projection, then performs noisy L0-SSC on the dimensionality reduced data for improved efficiency. Experimental results demonstrate the effectiveness of Noisy-DR-L0-SSC. Yingzhen Yang, Ping Li 0001 |
UAI | 1 |
| 2021 | FROS: Fast Regularized Optimization by SketchingabstractRandomized algorithms are important for solving large-scale optimization problems. In this paper, we propose Fast Regularized Optimization by Sketching (FROS) as an efficient solver for a general class of regularized optimization problems. FROS first generates a sketch of the original data matrix, then solves the sketched problem. Different from existing randomized algorithms, FROS handles general Frechet subdifferentiable regularization functions in an unified framework. It is proved that FROS achieves relative-error bounds for the approximation error between the optimization results of the sketched problem and that of the original problem for all convex and certain non-convex regularization. We further propose Iterative FROS which reduces the approximation error exponentially by iteratively invoking FROS. To our best knowledge, our results are among the few in approximation error of sketching algorithms for a broad class of optimization problems with general regularization. Experimental results demonstrate the effectiveness of the proposed FROS and Iterative FROS algorithms. Yingzhen Yang, Ping Li 0001 |
ISIT | 1 |
| 2020 | FSNet: Compression of Deep Convolutional Neural Networks by Filter Summary
Yingzhen Yang, Nebojsa Jojic, Jun Huan, Thomas S. Huang |
ICLR | 1 |
| 2019 | Fast Proximal Gradient Descent for A Class of Non-convex and Non-smooth Sparse Learning Problems
Yingzhen Yang |
UAI | 1 |
| 2018 | Dimensionality Reduced $\ell^{0}$-Sparse Subspace ClusteringabstractSubspace clustering partitions the data that lie on a union of subspaces. $\ell^{0}$-Sparse Subspace Clustering ($\ell^{0}$-SSC), which belongs to the subspace clustering methods with sparsity prior, guarantees the correctness of subspace clustering under less restrictive assumptions compared to its $\ell^{1}$ counterpart such as Sparse Subspace Clustering (SSC, Elhamifar et al., 2013) with demonstrated effectiveness in practice. In this paper, we present Dimensionality Reduced $\ell^{0}$-Sparse Subspace Clustering (DR-$\ell^{0}$-SSC). DR-$\ell^{0}$-SSC first projects the data onto a lower dimensional space by linear transformation, then performs $\ell^{0}$-SSC on the dimensionality reduced data. The correctness of DR-$\ell^{0}$-SSC in terms of the subspace detection property is proved, therefore DR-$\ell^{0}$-SSC recovers the underlying subspace structure in the original data from the dimensionality reduced data. Experimental results demonstrate the effectiveness of DR-$\ell^{0}$-SSC. Yingzhen Yang |
AISTATS | 1 |
| 2018 | WSNet: Compact and Efficient Networks Through Weight SamplingabstractWe present a new approach and a novel architecture, termed WSNet, for learning compact and efficient deep neural networks. Existing approaches conventionally learn full model parameters independently and then compress them via ad hoc processing such as model pruning or filter factorization. Alternatively, WSNet proposes learning model parameters by sampling from a compact set of learnable parameters, which naturally enforces parameter sharing throughout the learning process. We demonstrate that such a novel weight sampling approach (and induced WSNet) promotes both weights and computation sharing favorably. By employing this method, we can more efficiently learn much smaller networks with competitive performance compared to baseline networks with equal numbers of convolution filters. Specifically, we consider learning compact and efficient 1D convolutional neural networks for audio classification. Extensive experiments on multiple audio classification datasets verify the effectiveness of WSNet. Combined with weight quantization, the resulted models are up to 180x smaller and theoretically up to 16x faster than the well-established baselines, without noticeable performance drop. Xiaojie Jin 0004, Yingzhen Yang, Ning Xu 0001, Jianchao Yang, Nebojsa Jojic, Jiashi Feng, Shuicheng Yan |
ICML | 2 |
| 2018 | Subspace Learning by ℓ0-Induced Sparsity
Yingzhen Yang, Jiashi Feng, Nebojsa Jojic, Jianchao Yang, Thomas S. Huang |
Int. J. Comput. Vis. | 1 |
| 2017 | Support Regularized Sparse Coding and Its Fast Encoder
Yingzhen Yang, Pushmeet Kohli, Jianchao Yang, Thomas S. Huang |
ICLR (Poster) | 1 |
| 2017 | Neighborhood Regularized l^1-Graph
Yingzhen Yang, Jiashi Feng, Jianchao Yang, Thomas S. Huang |
UAI | 1 |
| 2016 | Epitomic Image Super-ResolutionabstractWe propose Epitomic Image Super-Resolution (ESR) to enhance the current internal SR methods that exploit the self-similarities in the input. Instead of local nearest neighbor patch matching used in most existing internal SR methods, ESR employs epitomic patch matching that features robustness to noise, and both local and non-local patch matching. Extensive objective and subjective evaluation demonstrate the effectiveness and advantage of ESR on various images. Yingzhen Yang, Zhangyang Wang, Shiyu Chang, Ding Liu 0001, Humphrey Shi, Thomas S. Huang |
AAAI | 1 |
| 2016 | On Order-Constrained Transitive Distance ClusteringabstractWe consider the problem of approximating order-constrained transitive distance (OCTD) and its clustering applications. Given any pairwise data, transitive distance (TD) is defined as the smallest possible "gap" on the set of paths connecting them. While such metric definition renders significant capability of addressing elongated clusters, it is sometimes also an over-simplified representation which loses necessary regularization on cluster structure and overfits to short links easily. As a result, conventional TD often suffers from degraded performance given clusters with "thick" structures. Our key intuition is that the maximum (path) order, which is the maximum number of nodes on a path, controls the level of flexibility. Reducing this order benefits the clustering performance by finding a trade-off between flexibility and regularization on cluster structure. Unlike TD, finding OCTD becomes an intractable problem even though the number of connecting paths is reduced. We therefore propose a fast approximation framework, using random samplings to generate multiple diversified TD matrices and a pooling to output the final approximated OCTD matrix. Comprehensive experiments on toy, image and speech datasets show the excellent performance of OCTD, surpassing TD with significant gains and giving state-of-the-art performance on several datasets. Zhiding Yu, Weiyang Liu, Wenbo Liu 0002, Yingzhen Yang, Ming Li 0026, B. V. K. Vijaya Kumar |
AAAI | 4 |
| 2016 | Studying Very Low Resolution Recognition Using Deep NetworksabstractVisual recognition research often assumes a sufficient resolution of the region of interest (ROI). That is usually violated in practice, inspiring us to explore the Very Low Resolution Recognition (VLRR) problem. Typically, the ROI in a VLRR problem can be smaller than 16 16 pixels, and is challenging to be recognized even by human experts. We attempt to solve the VLRR problem using deep learning methods. Taking advantage of techniques primarily in super resolution, domain adaptation and robust regression, we formulate a dedicated deep learning method and demonstrate how these techniques are incorporated step by step. Any extra complexity, when introduced, is fully justified by both analysis and simulation results. The resulting Robust Partially Coupled Networks achieves feature enhancement and recognition simultaneously. It allows for both the flexibility to combat the LR-HR domain mismatch, and the robustness to outliers. Finally, the effectiveness of the proposed models is evaluated on three different VLRR tasks, including face identification, digit recognition and font recognition, all of which obtain very impressive performances. Zhangyang Wang, Shiyu Chang, Yingzhen Yang, Ding Liu 0001, Thomas S. Huang |
CVPR | 3 |
| 2016 | D3: Deep Dual-Domain Based Fast Restoration of JPEG-Compressed ImagesabstractIn this paper, we design a Deep Dual-Domain (D3) based fast restoration model to remove artifacts of JPEG compressed images. It leverages the large learning capacity of deep networks, as well as the problem-specific expertise that was hardly incorporated in the past design of deep architectures. For the latter, we take into consideration both the prior knowledge of the JPEG compression scheme, and the successful practice of the sparsity-based dual-domain approach. We further design the One-Step Sparse Inference (1-SI) module, as an efficient and lightweighted feed-forward approximation of sparse coding. Extensive experiments verify the superiority of the proposed D3 model over several state-of-the-art methods. Specifically, our best model is capable of outperforming the latest deep model for around 1 dB in PSNR, and is 30 times faster. Zhangyang Wang, Ding Liu 0001, Shiyu Chang, Qing Ling 0001, Yingzhen Yang, Thomas S. Huang |
CVPR | 5 |
| 2016 | ℓ ^0 ℓ 0 -Sparse Subspace Clustering
Yingzhen Yang, Jiashi Feng, Nebojsa Jojic, Jianchao Yang, Thomas S. Huang |
ECCV (2) | 1 |
| 2016 | Learning A Deep ℓ∞ Encoder for Hashing
Zhangyang Wang, Yingzhen Yang, Shiyu Chang, Qing Ling 0001, Thomas S. Huang |
IJCAI | 2 |
| 2016 | Large-scale supervised similarity learning in networks
Shiyu Chang, Guo-Jun Qi, Yingzhen Yang, Charu C. Aggarwal, Meng Wang 0001, Thomas S. Huang |
Knowl. Inf. Syst. | 3 |
| 2015 | A Joint Optimization Framework of Sparse Coding and Discriminative Clustering
Zhangyang Wang, Yingzhen Yang, Shiyu Chang, Jinyan Li 0002, Simon Fong 0001, Thomas S. Huang |
IJCAI | 2 |
| 2015 | Designing a composite dictionary adaptively from joint examplesabstractWe study the complementary behaviors of external and internal examples in image restoration, and are motivated to formulate a composite dictionary design framework. The composite dictionary consists of the global part learned from external examples, and the sample-specific part learned from internal examples. The dictionary atoms in both parts are further adaptively weighted to emphasize their model statistics. Experiments demonstrate that the joint utilization of external and internal examples leads to substantial improvements, with successful applications in image denoising and super resolution. Zhangyang Wang, Yingzhen Yang, Jianchao Yang, Thomas S. Huang |
VCIP | 2 |
| 2015 | Learning Super-Resolution Jointly From External and Internal ExamplesabstractSingle image super-resolution (SR) aims to estimate a high-resolution (HR) image from a low-resolution (LR) input. Image priors are commonly learned to regularize the, otherwise, seriously ill-posed SR problem, either using external LR-HR pairs or internal similar patterns. We propose joint SR to adaptively combine the advantages of both external and internal SR methods. We define two loss functions using sparse coding-based external examples, and epitomic matching based on internal examples, as well as a corresponding adaptive weight to automatically balance their contributions according to their reconstruction errors. Extensive SR results demonstrate the effectiveness of the proposed method over the existing state-of-the-art methods, and is also verified by our subjective evaluation study. Zhangyang Wang, Yingzhen Yang, Shiyu Chang, Jianchao Yang, Thomas S. Huang |
IEEE Trans. Image Process. | 2 |
| 2014 | Data Clustering by Laplacian Regularized L1-GraphabstractL1-Graph has been proven to be effective in data clustering, which partitions the data space by using the sparse representation of the data as the similarity measure. However, the sparse representation is performed for each datum separately without taking into account the geometric structure of the data. Motivated by L1-Graph and manifold leaning, we propose Laplacian Regularized L1-Graph (LRℓ1-Graph) for data clustering. The sparse representations of LRℓ1-Graph are regularized by the geometric information of the data so that they vary smoothly along the geodesics of the data manifold by the graph Laplacian according to the manifold assumption. Moreover, we propose an iterative regularization scheme, where the sparse representation obtained from the previous iteration is used to build the graph Laplacian for the current iteration of regularization. The experimental results on real data sets demonstrate the superiority of our algorithm compared to L1-Graph and other competing clustering methods. Yingzhen Yang, Zhangyang Wang, Jianchao Yang, Jiangping Wang, Shiyu Chang, Thomas S. Huang |
AAAI | 1 |
| 2014 | Regularized l1-Graph for Data Clustering
Yingzhen Yang, Zhangyang Wang, Jianchao Yang, Jiawei Han 0001, Thomas S. Huang |
BMVC | 1 |
| 2014 | Epitomic image colorizationabstractImage colorization adds color to grayscale images. It not only increases the visual appeal of grayscale images, but also enriches the information conveyed by scientific images that lack color information. We develop a new image colorization method, epitomic image colorization, which automatically transfers color from the reference color image to the target grayscale image by a robust feature matching scheme using a new feature representation, namely the heterogeneous feature epitome. As a generative model, heterogeneous feature epitome is a condensed representation of image appearance which is employed for measuring the dissimilarity between reference patches and target patches in a way robust to noise in the reference image. We build a Markov Random Field (MRF) model with the learned heterogeneous feature epitome from the reference image, and inference in the MRF model achieves robust feature matching for transferring color. Our method renders better colorization results than the current state-of-the-art automatic colorization methods in our experiments. Yingzhen Yang, Xinqi Chu, Tian-Tsong Ng, Alex Yong Sang Chia, Jianchao Yang, Hailin Jin, Thomas S. Huang |
ICASSP | 1 |
| 2014 | Discriminative Exemplar clusteringabstractExemplar-based clustering methods partition the data space and identify the representative, or the exemplar, of each cluster. With the number of clusters adaptively determined, exemplar-based clustering methods are appealing since they avoid or alleviate the difficult task of estimating the latent parameters in case of complex models and high dimensionality of the data. Most exemplar-based clustering methods are based on generative models, where the exemplars serve as the parameters of the generative models. However, generative models do not consider the discriminative capability of the cluster boundaries explicitly described in discriminative models. In this paper, we present Discriminative Exemplar Clustering (DEC), that improves the discriminative power of exemplar-based clustering method by minimizing the misclassification error of the nonparametric unsupervised plug-in classifier while maintaining the appealing property of exemplar-based clustering. The optimization of DEC is performed in a pairwise Markov Random Field. Experimental results on synthetic and real data demonstrate the effectiveness of our method compared to other exemplar-based clustering methods. Yingzhen Yang, Feng Liang 0002, Thomas S. Huang |
ICASSP | 1 |
| 2014 | On a Theory of Nonparametric Pairwise Similarity for Clustering: Connecting Clustering to Classification
Yingzhen Yang, Feng Liang 0002, Shuicheng Yan, Zhangyang Wang, Thomas S. Huang |
NIPS | 1 |
| 2013 | Pairwise Clustering by Minimizing the Error of Unsupervised Nearest Neighbor ClassificationabstractPair wise clustering methods, including the popular graph cut based approaches such as normalized cut, partition the data space into clusters by the pair wise affinity between data points. The success of pair wise clustering largely depends on the pair wise affinity function defined over data points coming from different clusters. Interpreting the pair wise affinity in a probabilistic framework, we build the relationship between pair wise clustering and unsupervised classification by learning the soft Nearest Neighbor (NN) classifier from unlabeled data, and search for the optimal partition of the data points by minimizing the generalization error of the learned classifier associated with the data partitions. Modeling the underlying distribution of the data by non-parametric kernel density estimation, the asymptotic generalization error of the unsupervised soft NN classification involves only the pair wise affinity between data points. Moreover, such error rate reduces to the well-known kernel form of graph cut in case of uniform data distribution, which provides another understanding of the kernel similarity used in Laplacian Eigenmaps [1] which also assumes uniform distribution. By minimizing the generalization error bound, we propose a new clustering algorithm. Our algorithm efficiently partition the data by inference in a pair wise MRF model. Experimental results demonstrate the effectiveness of our method. Yingzhen Yang, Xinqi Chu, Thomas S. Huang |
ICMLA (2) | 1 |
| 2012 | Pairwise Exemplar ClusteringabstractExemplar-based clustering methods have been extensively shown to be effective in many clustering problems. They adaptively determine the number of clusters and hold the appealing advantage of not requiring the estimation of latent parameters, which is otherwise difficult in case of complicated parametric model and high dimensionality of the data. However, modeling arbitrary underlying distribution of the data is still difficult for existing exemplar-based clustering methods. We present Pairwise Exemplar Clustering (PEC) to alleviate this problem by modeling the underlying cluster distributions more accurately with non-parametric kernel density estimation. Interpreting the clusters as classes from a supervised learning perspective, we search for an optimal partition of the data that balances two quantities: 1 the misclassification rate of the data partition for separating the clusters; 2 the sum of within-cluster dissimilarities for controlling the cluster size. The broadly used kernel form of cut turns out to be a special case of our formulation. Moreover, we optimize the corresponding objective function by a new efficient algorithm for message computation in a pairwise MRF. Experimental results on synthetic and real data demonstrate the effectiveness of our method. Yingzhen Yang, Xinqi Chu, Feng Liang 0002, Thomas S. Huang |
AAAI | 1 |
| 2009 | Entertaining video warpingabstractWhile various techniques of image deformation have been developed and extensively applied in animation and morphing, there are few works to extend these techniques to handle videos, especially real-time warping of a meaningful moving part in the video like human face. An efficient online algorithm is proposed in this paper to implement real-time face warping for video sequence. We employ AdaBoost to detect the sixteen human facial feature points and implement a fast face warping frame by frame while maintaining both temporal and spatial continuity of the warped video. In order to reduce the shaking of the detected facial feature points due to the noise in the video, we develop a novel model called Frame Buffer. All these procedures are designed efficient to guarantee the real-time performance of our system (15 fps). In addition, many other types of warping functions are also compatible with our framework. It is shown that our algorithm can be applied in real-time special effect editing in video as well as other entertainment applications. Yingzhen Yang, Chengfang Song, Qunsheng Peng 0001 |
CAD/Graphics | 1 |
| 2009 | Image completion using structural priority belief propagationabstractA new image completion algorithm called Structural Priority Belief Propagation (SPBP) is presented to deal with LDV based image completion in this paper. LDV completion is a new form of image completion based on another large displacement view (LDV) of the same scene, no wonder, it has the potential of repairing large unknown region with salient structure information. In order to complete such unknown region, SPBP makes two important extensions over existing Priority-BP: dynamic weight of structural consistency and structural priority inheritance so as to propagate linear structure with correct priority, meanwhile it promotes texture propagation adhering to a global optimization scheme. Experimental results demonstrate that SPBP can obtain more satisfactory results than other LDV completion algorithms and it also performs well for traditional single image completion. Yingzhen Yang, Qunsheng Peng 0001 |
ACM Multimedia | 1 |
| 2009 | An improved belief propagation method for dynamic collage
Yingzhen Yang, Qunsheng Peng 0001, Yasuyuki Matsushita |
Vis. Comput. | 1 |
| 2008 | Distortion Optimization based Image Completion from a Large Displacement ViewabstractAbstract We present a new image completion method based on an additional large displacement view (LDV) of the same scene for faithfully repairing large missing regions on the target image in an automatic way. A coarse‐to‐fine distortion correction algorithm is proposed to minimize the perspective distortion in the corresponding parts for the common scene regions on the LDV image. First, under the assumption of a planar scene, the LDV image is warped according to a homography to generate the initial correction result. Second, the residual distortions in the common known scene regions are revealed by means of a mismatch detection mechanism and relaxed by energy optimization of overlap correspondences, with the expectations of color constancy and displacement field smoothness. The fundamental matrix for the two views is then computed based on the reliable correspondence set. Third, under the constraints of epipolar geometry, displacement field smoothness and color consistency of the neighboring pixels, the missing pixels are orderly restored according to a specially defined repairing priority function. We finally eliminate the ghost effect between the repaired region and its surroundings by Poisson image blending. Experimental results demonstrate that our method outperforms recent state‐of‐the‐art image completion methods for repairing large missing area with complex structure information. Yingzhen Yang, Qunsheng Peng 0001, Wei Chen 0001 |
Comput. Graph. Forum | 2 |