EDBT 2026 Demo / reviewers in the wild / expert
Ju Sun
dblp:31/6843
· DBLP profile ↗
26ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-2017-5903ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 7 since 2021Theory of computation · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal-Consistent Video Restoration with Pre-trained Diffusion ModelsabstractVideo restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they suffer from approximation errors during reverse diffusion and insufficient temporal consistency. Moreover, dealing with 3D video data, VR is inherently computationally intensive. In this paper, we advocate viewing the reverse process in DMs as a function and present a novel Maximum a Posterior (MAP) framework that directly parameterizes video frames in the seed space of DMs, eliminating approximation errors. We also introduce strategies to promote bilevel temporal consistency: semantic consistency by leveraging clustering structures in the seed space, and pixel-level consistency by progressive warping with optical flow refinements. Extensive experiments on multiple virtual reality tasks demonstrate superior visual quality and temporal consistency achieved by our method compared to the state-of-the-art. Hengkang Wang, Huidong Liu, Chien-Chih Wang, Hongdong Li, Bryan Wang, Ju Sun |
AAAI | 8 |
| 2024 | DMPlug: A Plug-in Method for Solving Inverse Problems with Diffusion ModelsabstractPretrained diffusion models (DMs) have recently been popularly used in solving inverse problems (IPs). The existing methods mostly interleave iterative steps in the reverse diffusion process and iterative steps to bring the iterates closer to satisfying the measurement constraint. However, such interleaving methods struggle to produce final results that look like natural objects of interest (i.e., manifold feasibility) and fit the measurement (i.e., measurement feasibility), especially for nonlinear IPs. Moreover, their capabilities to deal with noisy IPs with unknown types and levels of measurement noise are unknown. In this paper, we advocate viewing the reverse process in DMs as a function and propose a novel plug-in method for solving IPs using pretrained DMs, dubbed DMPlug. DMPlug addresses the issues of manifold feasibility and measurement feasibility in a principled manner, and also shows great potential for being robust to unknown types and levels of noise. Through extensive experiments across various IP tasks, including two linear and three nonlinear IPs, we demonstrate that DMPlug consistently outperforms state-of-the-art methods, often by large margins especially for nonlinear IPs. Hengkang Wang, Taihui Li, Yuxiang Wan, Tiancong Chen, Ju Sun |
NeurIPS | 6 |
| 2024 | Interpretable deep learning methods for multiview learningabstractBACKGROUND: Technological advances have enabled the generation of unique and complementary types of data or views (e.g. genomics, proteomics, metabolomics) and opened up a new era in multiview learning research with the potential to lead to new biomedical discoveries. RESULTS: We propose iDeepViewLearn (Interpretable Deep Learning Method for Multiview Learning) to learn nonlinear relationships in data from multiple views while achieving feature selection. iDeepViewLearn combines deep learning flexibility with the statistical benefits of data and knowledge-driven feature selection, giving interpretable results. Deep neural networks are used to learn view-independent low-dimensional embedding through an optimization problem that minimizes the difference between observed and reconstructed data, while imposing a regularization penalty on the reconstructed data. The normalized Laplacian of a graph is used to model bilateral relationships between variables in each view, therefore, encouraging selection of related variables. iDeepViewLearn is tested on simulated and three real-world data for classification, clustering, and reconstruction tasks. For the classification tasks, iDeepViewLearn had competitive classification results with state-of-the-art methods in various settings. For the clustering task, we detected molecular clusters that differed in their 10-year survival rates for breast cancer. For the reconstruction task, we were able to reconstruct handwritten images using a few pixels while achieving competitive classification accuracy. The results of our real data application and simulations with small to moderate sample sizes suggest that iDeepViewLearn may be a useful method for small-sample-size problems compared to other deep learning methods for multiview learning. CONCLUSION: iDeepViewLearn is an innovative deep learning model capable of capturing nonlinear relationships between data from multiple views while achieving feature selection. It is fully open source and is freely available at https://github.com/lasandrall/iDeepViewLearn . Hengkang Wang, Ju Sun, Sandra Safo |
BMC Bioinform. | 3 |
| 2024 | Blind Image Deblurring with Unknown Kernel Size and Substantial Noise
Zhong Zhuang, Taihui Li, Hengkang Wang, Ju Sun |
Int. J. Comput. Vis. | 4 |
| 2024 | A Unified Analysis of AdaGrad With Weighted Aggregation and Momentum AccelerationabstractIntegrating adaptive learning rate and momentum techniques into stochastic gradient descent (SGD) leads to a large class of efficiently accelerated adaptive stochastic algorithms, such as AdaGrad, RMSProp, Adam, AccAdaGrad, and so on. In spite of their effectiveness in practice, there is still a large gap in their theories of convergences, especially in the difficult nonconvex stochastic setting. To fill this gap, we propose weighted AdaGrad with unified momentum and dubbed AdaUSM, which has the main characteristics that: 1) it incorporates a unified momentum scheme that covers both the heavy ball (HB) momentum and the Nesterov accelerated gradient (NAG) momentum and 2) it adopts a novel weighted adaptive learning rate that can unify the learning rates of AdaGrad, AccAdaGrad, Adam, and RMSProp. Moreover, when we take polynomially growing weights in AdaUSM, we obtain its O(log(T)/√T) convergence rate in the nonconvex stochastic setting. We also show that the adaptive learning rates of Adam and RMSProp correspond to taking exponentially growing weights in AdaUSM, thereby providing a new perspective for understanding Adam and RMSProp. Finally, comparative experiments of AdaUSM against SGD with momentum, AdaGrad, AdaEMA, Adam, and AMSGrad on various deep learning models and datasets are also carried out. Li Shen 0008, Congliang Chen, Fangyu Zou, Zequn Jie, Ju Sun, Wei Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Rethinking Transfer Learning for Medical Image Classification
Le Peng, Hengyue Liang, Gaoxiang Luo, Taihui Li, Ju Sun |
BMVC | 5 |
| 2023 | Deep Random Projector: Accelerated Deep Image PriorabstractDeep image prior (DIP) has shown great promise in tackling a variety of image restoration (IR) and general visual inverse problems, needing no training data. However, the resulting optimization process is often very slow, inevitably hindering DIP's practical usage for time-sensitive scenarios. In this paper, we focus on IR, and propose two crucial modifications to DIP that help achieve substantial speedup: 1) optimizing the DIP seed while freezing randomly-initialized network weights, and 2) reducing the network depth. In addition, we reintroduce explicit priors, such as sparse gradient prior-encoded by total-variation regularization, to preserve the DIP peak performance. We evaluate the proposed method on three IR tasks, including image denoising, image super-resolution, and image inpainting, against the original DIP and variants, as well as the competing metaDIP that uses meta-learning to learn good initializers with extra data. Our method is a clear winner in obtaining competitive restoration quality in a minimal amount of time. Our code is available at https://github.com/sun-umn/Deep-Random-Projector. Taihui Li, Hengkang Wang, Zhong Zhuang, Ju Sun |
CVPR | 4 |
| 2023 | Robust Autoencoders for Collective Corruption RemovalabstractRobust PCA is a standard tool for learning a linear subspace in the presence of sparse corruption or rare outliers. What about robustly learning manifolds that are more realistic models for natural data, such as images? There have been several recent attempts to generalize robust PCA to manifold settings. In this paper, we propose ℓ1- and scaling-invariant ℓ1/ℓ2-robust autoencoders based on a surprisingly compact formulation built on the intuition that deep autoencoders perform manifold learning. We demonstrate on several standard image datasets that the proposed formulation significantly outperforms all previous methods in collectively removing sparse corruption, without clean images for training. Moreover, we also show that the learned manifold structures can be generalized to unseen data samples effectively. Taihui Li, Hengkang Wang, Le Peng, Xian'e Tang, Ju Sun |
ICASSP | 5 |
| 2023 | Random Projector: Efficient Deep Image PriorabstractDeep image prior (DIP) has shown great promise in tackling a range of image restoration problems. However, its optimization is extremely sluggish, which inevitability hinders its practical usage when there are hard time constraints. To mitigate this issue, we propose a more compact and efficient model, dubbed random projector (RP), and freeze the convolutional layers of the neural network to prevent slow learning. We further make use of an explicit prior—total variation— to regularize the reconstructed natural images and promote pleasure-looking images. We evaluate our proposed method on different image restoration tasks such as image denoising and image inpainting, and conduct comparisons with DIP and its prevalent variants. The experimental results suggest that our proposed random projector achieves competitive restoration quality in terms of PSNR while it significantly reduces the optimization (OPT) time. Taihui Li, Zhong Zhuang, Hengkang Wang, Ju Sun |
ICASSP | 4 |
| 2023 | Optimization for Robustness Evaluation Beyond ℓp MetricsabstractEmpirical evaluation of the adversarial robustness of deep learning models involves solving non-trivial constrained optimization problems. Popular numerical algorithms to solve these constrained problems rely predominantly on projected gradient descent (PGD) and mostly handle adversarial perturbations modeled by the ℓ1, ℓ2, and ℓ∞metrics. In this paper, we introduce a novel algorithmic framework that blends a general-purpose constrained-optimization solver PyGRANSO, With Constraint-Folding (PWCF), to add reliability and generality to robustness evaluation. PWCF 1) finds good-quality solutions without the need of delicate hyperparameter tuning and 2) can handle more general perturbation types, e.g., modeled by general ℓp(where p > 0) and perceptual (nonℓp) distances, which are inaccessible to existing PGD-based algorithms. Hengyue Liang, Buyun Liang 0001, Tim Mitchell, Ju Sun |
ICASSP | 5 |
| 2022 | Evaluation of federated learning variations for COVID-19 diagnosis using chest radiographs from 42 US and European hospitalsabstractOBJECTIVE: Federated learning (FL) allows multiple distributed data holders to collaboratively learn a shared model without data sharing. However, individual health system data are heterogeneous. "Personalized" FL variations have been developed to counter data heterogeneity, but few have been evaluated using real-world healthcare data. The purpose of this study is to investigate the performance of a single-site versus a 3-client federated model using a previously described Coronavirus Disease 19 (COVID-19) diagnostic model. Additionally, to investigate the effect of system heterogeneity, we evaluate the performance of 4 FL variations. MATERIALS AND METHODS: We leverage a FL healthcare collaborative including data from 5 international healthcare systems (US and Europe) encompassing 42 hospitals. We implemented a COVID-19 computer vision diagnosis system using the Federated Averaging (FedAvg) algorithm implemented on Clara Train SDK 4.0. To study the effect of data heterogeneity, training data was pooled from 3 systems locally and federation was simulated. We compared a centralized/pooled model, versus FedAvg, and 3 personalized FL variations (FedProx, FedBN, and FedAMP). RESULTS: We observed comparable model performance with respect to internal validation (local model: AUROC 0.94 vs FedAvg: 0.95, P = .5) and improved model generalizability with the FedAvg model (P < .05). When investigating the effects of model heterogeneity, we observed poor performance with FedAvg on internal validation as compared to personalized FL algorithms. FedAvg did have improved generalizability compared to personalized FL algorithms. On average, FedBN had the best rank performance on internal and external validation. CONCLUSION: FedAvg can significantly improve the generalization of the model compared to other personalization FL algorithms; however, at the cost of poor internal validity. Personalized FL may offer an opportunity to develop both internal and externally validated algorithms. Le Peng, Gaoxiang Luo, Andrew Walker, Zach Zaiman, Emma K. Jones, Hemant Gupta, Kristopher Kersten, John L. Burns, Christopher A. Harle, Tanja Magoc, Benjamin Shickel, Scott D. Steenburg, Tyler J. Loftus, Genevieve B. Melton, Judy Gichoya, Ju Sun, Christopher J. Tignanelli |
J. Am. Medical Informatics Assoc. | 16 |
| 2021 | Self-Validation: Early Stopping for Single-Instance Deep Generative Priors
Taihui Li, Zhong Zhuang, Hengyue Liang, Le Peng, Hengkang Wang, Ju Sun |
BMVC | 6 |
| 2019 | Subgradient Descent Learns Orthogonal Dictionaries
Yu Bai 0017, Qijia Jiang, Ju Sun |
ICLR (Poster) | 3 |
| 2017 | Complete Dictionary Recovery Over the Sphere I: Overview and the Geometric PictureabstractWe consider the problem of recovering a complete (i.e., square and invertible) matrix A0, from Y ∈ Rn×pwith Y = A0X0, provided X0is sufficiently sparse. This recovery problem is central to theoretical understanding of dictionary learning, which seeks a sparse representation for a collection of input signals and finds numerous applications in modern signal processing and machine learning. We give the first efficient algorithm that provably recovers A0when X0has O (n) nonzeros per column, under suitable probability model for X0. In contrast, prior results based on efficient algorithms either only guarantee recovery when X0has O(√n) zeros per column, or require multiple rounds of semidefinite programming relaxation to work when X0has O(n) nonzeros per column. Our algorithmic pipeline centers around solving a certain nonconvex optimization problem with a spherical constraint. In this paper, we provide a geometric characterization of the objective landscape. In particular, we show that the problem is highly structured with high probability: 1) there are no “spurious” local minimizers and 2) around all saddle points the objective has a negative directional curvature. This distinctive structure makes the problem amenable to efficient optimization algorithms. In a companion paper, we design a second-order trust-region algorithm over the sphere that provably converges to a local minimizer from arbitrary initializations, despite the presence of saddle points. Ju Sun, Qing Qu 0001, John Wright 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2017 | Complete Dictionary Recovery Over the Sphere II: Recovery by Riemannian Trust-Region MethodabstractWe consider the problem of recovering a complete (i.e., square and invertible) matrix A0, from Y ∈ Rn×pwith Y = A0X0, provided X0is sufficiently sparse. This recovery problem is central to theoretical understanding of dictionary learning, which seeks a sparse representation for a collection of input signals and finds numerous applications in modern signal processing and machine learning. We give the first efficient algorithm that provably recovers A0when X0has O (n) nonzeros per column, under suitable probability model for X0. Our algorithmic pipeline centers around solving a certain nonconvex optimization problem with a spherical constraint, and hence is naturally phrased in the language of manifold optimization. In a companion paper, we have showed that with high probability, our nonconvex formulation has no “spurious” local minimizers and around any saddle point, the objective function has a negative directional curvature. In this paper, we take advantage of the particular geometric structure and describe a Riemannian trust region algorithm that provably converges to a local minimizer with from arbitrary initializations. Such minimizers give excellent approximations to the rows of X0. The rows are then recovered by a linear programming rounding and deflation. Ju Sun, Qing Qu 0001, John Wright 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2016 | A geometric analysis of phase retrievalabstractGiven nonlinear measurements yk= |〈ak, x〉| for k = 1,...,m, is it possible to recover x ∈ ℂn? This generalized phase retrieval (GPR) problem is a fundamental task in various disciplines. Natural nonconvex methods often work remarkably well for GPR in practice, but lack clear theoretical explanations. In this paper, we take a step towards bridging this gap. We show that when the sensing vectors ak's are generic (i.i.d. complex Gaussian) and the number of measurements is large enough (m ≥ Cn log3n), with high probability (w.h.p.), a natural least-squares formulation for GPR has the following benign geometric structure: (1) all local minimizers are global-they are the target signal x and its equivalent copies; and (2) the objective function has a negative directional curvature around each saddle point. Such structure allows a number of algorithmic possibilities for efficient global optimization. We describe a second-order trust-region algorithm that provably finds a global minimizer in polynomial time, from arbitrary initializations. Ju Sun, Qing Qu 0001, John Wright 0001 |
ISIT | 1 |
| 2016 | Finding a Sparse Vector in a Subspace: Linear Sparsity Using Alternating DirectionsabstractIs it possible to find the sparsest vector (direction) in a generic subspace S ⊆ ℝpwith dim(S) = n <; p? This problem can be considered a homogeneous variant of the sparse recovery problem and finds connections to sparse dictionary learning, sparse PCA, and many other problems in signal processing and machine learning. In this paper, we focus on a planted sparse model for the subspace: the target sparse vector is embedded in an otherwise random subspace. Simple convex heuristics for this planted recovery problem provably break down when the fraction of nonzero entries in the target sparse vector substantially exceeds O(1/√n). In contrast, we exhibit a relatively simple nonconvex approach based on alternating directions, which provably succeeds even when the fraction of nonzero entries is Ω(1). To the best of our knowledge, this is the first practical algorithm to achieve linear scaling under the planted sparse model. Empirically, our proposed algorithm also succeeds in more challenging data models, e.g., sparse dictionary learning. Qing Qu 0001, Ju Sun, John Wright 0001 |
IEEE Trans. Inf. Theory | 2 |
| 2015 | Complete Dictionary Recovery Using Nonconvex OptimizationabstractWe consider the problem of recovering a complete (i.e., square and invertible) dictionary mb A_0, from mb Y = mb A_0 mb X_0 with mb Y ∈\mathbb R^n \times p. This recovery setting is central to the theoretical understanding of dictionary learning. We give the first efficient algorithm that provably recovers mb A_0 when mb X_0 has O(n) nonzeros per column, under suitable probability model for mb X_0. Prior results provide recovery guarantees when mb X_0 has only O(\sqrtn) nonzeros per column. Our algorithm is based on nonconvex optimization with a spherical constraint, and hence is naturally phrased in the language of manifold optimization. Our proofs give a geometric characterization of the high-dimensional objective landscape, which shows that with high probability there are no spurious local minima. Experiments with synthetic data corroborate our theory. Full version of this paper is available online: \urlhttp://arxiv.org/abs/1504.06785. Ju Sun, Qing Qu 0001, John Wright 0001 |
ICML | 1 |
| 2014 | Finding a sparse vector in a subspace: Linear sparsity using alternating directions
Qing Qu 0001, Ju Sun, John Wright 0001 |
NIPS | 2 |
| 2014 | Efficient Point-to-Subspace Query in ℓ1 with Application to Robust Object Instance RecognitionabstractMotivated by vision tasks such as robust face and object recognition, we consider the following general problem: given a collection of low-dimensional linear subspaces in a high-dimensional ambient (image) space, and a query point (image), efficiently determine the nearest subspace to the query in $\ell^1$ distance. In contrast to the naive exhaustive search which entails large-scale linear programs, we show that the computational burden can be cut down significantly by a simple two-stage algorithm: (1) projecting the query and database subspaces into lower-dimensional space by random Cauchy matrix and solving small-scale distance evaluations (linear programs) in the projection space to locate the nearest candidates; (2) with few candidates upon independent repetition of (1), getting back to the high-dimensional space and performing exhaustive search. To preserve the identity of the nearest subspace with nontrivial probability, the projection dimension typically is a low-order polynomial of the subspace dimension multiplied by a logarithm of the number of the subspaces (Theorem 2.1). The reduced dimensionality and hence complexity render the proposed algorithm particularly relevant to vision applications such as robust face and object instance recognition that we investigate empirically. Ju Sun, John Wright 0001 |
SIAM J. Imaging Sci. | 1 |
| 2013 | Robust Recovery of Subspace Structures by Low-Rank RepresentationabstractIn this paper, we address the subspace clustering problem. Given a set of data samples (vectors) approximately drawn from a union of multiple subspaces, our goal is to cluster the samples into their respective subspaces and remove possible outliers as well. To this end, we propose a novel objective function named Low-Rank Representation (LRR), which seeks the lowest rank representation among all the candidates that can represent the data samples as linear combinations of the bases in a given dictionary. It is shown that the convex program associated with LRR solves the subspace clustering problem in the following sense: When the data is clean, we prove that LRR exactly recovers the true subspace structures; when the data are contaminated by outliers, we prove that under certain conditions LRR can exactly recover the row space of the original data and detect the outlier as well; for data corrupted by arbitrary sparse errors, LRR can also approximately recover the row space with theoretical guarantees. Since the subspace membership is provably determined by the row space, these further imply that LRR can perform robust subspace clustering and error correction in an efficient and effective way. Guangcan Liu, Zhouchen Lin, Shuicheng Yan, Ju Sun, Yong Yu 0001, Yi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Efficient Point-to-Subspace Query in ℓ1 with Application to Robust Face Recognition
Ju Sun, John Wright 0001 |
ECCV (4) | 1 |
| 2010 | Randomized Locality Sensitive Vocabularies for Bag-of-Features Model
Yadong Mu, Ju Sun, Tony X. Han, Loong Fah Cheong, Shuicheng Yan |
ECCV (3) | 2 |
| 2010 | Activity recognition using dense long-duration trajectoriesabstractCurrent research on visual action/activity analysis has mostly exploited appearance-based static feature descriptions, plus statistics of short-range motion fields. The deliberate ignorance of dense, long-duration motion trajectories as features is largely due to the lack of mature mechanism for efficient extraction and quantitative representation of visual trajectories. In this paper, we propose a novel scheme for extraction and representation of dense, long-duration trajectories from video sequences, and demonstrate its ability to handle video sequences containing occlusions, camera motions, and nonrigid deformations. Moreover, we test the scheme on the KTH action recognition dataset, and show its promise as a scheme for general purpose long-duration motion description in realistic video sequences. Ju Sun, Yadong Mu, Shuicheng Yan, Loong Fah Cheong |
ICME | 1 |
| 2009 | Hierarchical spatio-temporal context modeling for action recognitionabstractThe problem of recognizing actions in realistic videos is challenging yet absorbing owing to its great potentials in many practical applications. Most previous research is limited due to the use of simplified action databases under controlled environments or focus on excessively localized features without sufficiently encapsulating the spatio-temporal context. In this paper, we propose to model the spatio-temporal context information in a hierarchical way, where three levels of context are exploited in ascending order of abstraction: 1) point-level context (SIFT average descriptor), 2) intra-trajectory context (trajectory transition descriptor), and 3) inter-trajectory context (trajectory proximity descriptor). To obtain efficient and compact representations for the latter two levels, we encode the spatiotemporal context information into the transition matrix of a Markov process, and then extract its stationary distribution as the final context descriptor. Building on the multichannel nonlinear SVMs, we validate this proposed hierarchical framework on the realistic action (HOHA) and event (LSCOM) recognition databases, and achieve 27% and 66% relative performance improvements over the state-of-the-art results, respectively. We further propose to employ the Multiple Kernel Learning (MKL) technique to prune the kernels towards speedup in algorithm evaluation. Ju Sun, Xiao Wu 0004, Shuicheng Yan, Loong Fah Cheong, Tat-Seng Chua, Jintao Li 0001 |
CVPR | 1 |
| 2008 | 3D ordinal constraint in spatial configuration for robust scene recognitionabstractThis paper proposes a scene recognition strategy that integrates the appearance based local SURF features and the geometry based 3D ordinal constraint. Firstly, we show that spatial ordinal ranks of 3D landmarks are well correlated across large camera viewpoint and view direction changes and thus serve as a powerful tool for scene recognition. Secondly, ordinal depth information is acquired in a simple and robust manner when the camera undergoes a bio-mimic ‘Turn-back-and-Look’ TBL) motion. Thirdly, a scene recognition strategy is proposed by combining local SURF feature matches and global 3D rank correlation coefficient into the scene recognition decision process. The performance is validated and evaluated over four indoor and outdoor databases. Ching Lik Teo, Shimiao Li, Loong Fah Cheong, Ju Sun |
ICPR | 4 |