Jia Li 0002

dblp:23/6950-2 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-2956-2846ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
Ziming Yu, Pan Zhou 0002, Sike Wang, Jia Li 0002, Mi Tian 0008, Hua Huang 0001
ICCV4
2025 Mixture of Group Experts for Learning Invariant Representations
abstract
Sparsely activated Mixture-of-Experts (MoE) models effectively increase the number of parameters while maintaining consistent computational costs per token. However, vanilla MoE models often suffer from limited diversity and specialization among experts, constraining their performance and scalability, especially as the number of experts increases. In this paper, we present a novel perspective on vanilla MoE with top-k routing inspired by sparse representation. This allows us to bridge established theoretical insights from sparse representation into MoE models. Building on this foundation, we propose a group sparse regularization approach for the input of top-k routing, termed Mixture of Group Experts (MoGE). MoGE indirectly regularizes experts by imposing structural constraints on the routing inputs, while preserving the original MoE architecture. Furthermore, we organize the routing input into a 2D topographic map, spatially grouping neighboring elements. This structure enables MoGE to capture representations invariant to minor transformations, thereby significantly enhancing expert diversity and specialization. Comprehensive evaluations across various Transformer models for image classification and language modeling tasks demonstrate that MoGE substantially outperforms its MoE counterpart, with minimal additional memory and computation overhead. Our approach provides a simple yet effective solution to scale the number of experts and reduce redundancy among them. Our code is available at: https://github.com/wangyuankl123/MoGE.
Lei Kang 0007, Jia Li 0002, Mi Tian 0008, Hua Huang 0001
MMAsia2
2025 Mixture of Group Experts for Multi-task Dense Prediction
Lei Kang 0007, Jia Li 0002, Hua Huang 0001
PRCV (3)2
2025 Joint learning of RGBW color filter arrays and demosaicking
Chen-Yan Bai, Faqi Liu, Jia Li 0002
Pattern Recognit.3
2024 4-bit Shampoo for Memory-Efficient Network Training
abstract
Second-order optimizers, maintaining a matrix termed a preconditioner, are superior to first-order optimizers in both theory and practice. The states forming the preconditioner and its inverse root restrict the maximum size of models trained by second-order optimizers. To address this, compressing 32-bit optimizer states to lower bitwidths has shown promise in reducing memory usage. However, current approaches only pertain to first-order optimizers. In this paper, we propose the first 4-bit second-order optimizers, exemplified by 4-bit Shampoo, maintaining performance similar to that of 32-bit ones. We show that quantizing the eigenvector matrix of the preconditioner in 4-bit Shampoo is remarkably better than quantizing the preconditioner itself both theoretically and experimentally. By rectifying the orthogonality of the quantized eigenvector matrix, we enhance the approximation of the preconditioner's eigenvector matrix, which also benefits the computation of its inverse 4-th root. Besides, we find that linear square quantization slightly outperforms dynamic tree quantization when quantizing second-order optimizer states. Evaluation on various networks for image classification and natural language modeling demonstrates that our 4-bit Shampoo achieves comparable performance to its 32-bit counterpart while being more memory-efficient.
Sike Wang, Pan Zhou 0002, Jia Li 0002, Hua Huang 0001
NeurIPS3
2024 Universal deep demosaicking for sparse color filter arrays
Chen-Yan Bai, Wenxing Qiao, Jia Li 0002
Signal Process. Image Commun.3
2024 Spatial-Frequency Fusion for Bayer Demosaicking
abstract
Deep learning-based demosaicking for the Bayer color filter array (CFA) has great advances. However, most existing demosaicking methods only work in the spatial domain and rarely explore the spectral characteristic of Bayer CFA. In this letter, we first attempt to address Bayer demosaicking in both spatial and frequency domains and propose a spatial-frequency fusion network. It consists of three key modules: spatial-domain branch, frequency-domain branch, and dual-domain fusion. Spatial-domain branch employs the standard convolution to extract local features in the spatial domain, while frequency-domain branch adopts frequency selection to maintain CFA's periodicity and achieve the image-wide receptive field for obtaining global features. Dual-domain fusion integrates the complementary representation of the two types of features. Extensive experiments validate the effectiveness of the proposed network.
Chen-Yan Bai, Jia Li 0002, Jinbiao Wang
IEEE Signal Process. Lett.2
2024 Fine-Tuning for Bayer Demosaicking Through Periodic-Consistent Self-Supervised Learning
abstract
Deep learning-based Bayer demosaicking methods have achieved superior performance. They require a large amount of paired images, which are challenging to collect. So most of these methods resort to using simulated images, where the raw images are sampled from the RGB images with the Bayer color filter array (CFA). However, there is a domain gap between simulated images and real images, which makes the learned networks less practical in the real world. To address this problem, we propose a simple and effective self-supervised learning framework for Bayer demosaicking. We observe that the Bayer CFA is periodic, and its four equivalent patterns can be transformed into each other through simple translation. Accordingly, we propose the concept of periodic-consistent demosaicking, which means that a demosaicking network should produce an identical demosaicked image for the same scene captured by the four equivalent patterns. Our framework can fine-tune existing demosaicking networks using only single raw images. We demonstrate the effectiveness of our framework on both clean and noisy images. Source code will be released.
Songze He, Jia Li 0002
IEEE Signal Process. Lett.4
2024 IMU-Assisted Accurate Blur Kernel Re-Estimation in Non-Uniform Camera Shake Deblurring
abstract
Image deblurring for camera shake is a highly regarded problem in the field of computer vision. A promising solution is the patch-wise non-uniform image deblurring algorithms, where a linear transformation model is typically established between different blur kernels to re-estimate poorly estimated blur kernels. However, the linear model struggles to effectively describe the nonlinear transformation relationships between blur kernels. A key observation is that the inertial measurement unit (IMU) provides motion data of the camera, which is helpful in describing the landmarks of the blur kernel. This paper presents a new IMU-assisted method for the re-estimation of poorly estimated blur kernels. This method establishes a nonlinear transformation relationship model between blur kernels of different patches using IMU motion data. Subsequently, an optimization problem is applied to re-estimate poorly estimated blur kernels by incorporating this relationship model with neighboring well-estimated kernels. Experimental results demonstrate that this blur kernel re-estimation method outperforms existing methods.
Jianxiang Rong, Hua Huang 0001, Jia Li 0002
IEEE Trans. Image Process.3
2023 Universal Demosaicking for Interpolation-Friendly RGBW Color Filter Arrays
abstract
Interpolation-friendly RGBW color filter arrays (CFAs) and the popular sequential demosaicking contain the idea of computational photography, where the CFA and the demosaicking method are co-designed. Due to the advantages, interpolation-friendly RGBW CFAs have been extensively used in commercial color cameras. However, most associated demosaicking methods rely on strict assumptions or are limited to a few specific CFAs with a given camera. In this paper, we propose a universal demosaicking method for interpolation-friendly RGBW CFAs, which enables the comparison of different CFAs. Our new method belongs to sequential demosaicking, i.e., W channel is interpolated first and then RGB channels are reconstructed with guidance from the interpolated W channel. Specifically, it first interpolates the W channel using only available W pixels followed by an aliasing reduction technique to remove aliasing artifacts. Then it employs an image decomposition model to built relations between W channel and each of RGB channels with known RGB values, which can be easily generalized to the full-size demosaicked image. We apply the linearized alternating direction method (LADM) to solve it with convergence guarantee. Our demosaicking method can be applied to all interpolation-friendly RGBW CFAs with varying color cameras and lighting conditions. Extensive experiments confirm the universal property and advantage of our proposed method with both simulated and real raw images.
Jia Li 0002, Chen-Yan Bai, Hua Huang 0001
IEEE Trans. Image Process.1
2022 Training Neural Networks by Lifted Proximal Operator Machines
abstract
We present the lifted proximal operator machine (LPOM) to train fully-connected feed-forward neural networks. LPOM represents the activation function as an equivalent proximal operator and adds the proximal operators to the objective function of a network as penalties. LPOM is block multi-convex in all layer-wise weights and activations. This allows us to develop a new block coordinate descent (BCD) method with convergence guarantee to solve it. Due to the novel formulation and solving method, LPOM only uses the activation function itself and does not require any gradient steps. Thus it avoids the gradient vanishing or exploding issues, which are often blamed in gradient-based methods. Also, it can handle various non-decreasing Lipschitz continuous activation functions. Additionally, LPOM is almost as memory-efficient as stochastic gradient descent and its parameter tuning is relatively easy. We further implement and analyze the parallel solution of LPOM. We first propose a general asynchronous-parallel BCD method with convergence guarantee. Then we use it to solve LPOM, resulting in asynchronous-parallel LPOM. For faster speed, we develop the synchronous-parallel LPOM. We validate the advantages of LPOM on various network architectures and datasets. We also apply synchronous-parallel LPOM to autoencoder training and demonstrate its fast convergence and superior performance.
Jia Li 0002, Mingqing Xiao 0002, Cong Fang 0001, Yue Dai 0003, Chao Xu 0006, Zhouchen Lin
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Lifted Proximal Operator Machines
abstract
We propose a new optimization method for training feedforward neural networks. By rewriting the activation function as an equivalent proximal operator, we approximate a feedforward neural network by adding the proximal operators to the objective function as penalties, hence we call the lifted proximal operator machine (LPOM). LPOM is block multiconvex in all layer-wise weights and activations. This allows us to use block coordinate descent to update the layer-wise weights and activations. Most notably, we only use the mapping of the activation function itself, rather than its derivative, thus avoiding the gradient vanishing or blow-up issues in gradient based training methods. So our method is applicable to various non-decreasing Lipschitz continuous activation functions, which can be saturating and non-differentiable. LPOM does not require more auxiliary variables than the layer-wise activations, thus using roughly the same amount of memory as stochastic gradient descent (SGD) does. Its parameter tuning is also much simpler. We further prove the convergence of updating the layer-wise weights and activations and point out that the optimization could be made parallel by asynchronous update. Experiments on MNIST and CIFAR-10 datasets testify to the advantages of LPOM.
Jia Li 0002, Cong Fang 0001, Zhouchen Lin
AAAI1
2019 Group Collaborative Representation for Image Set Classification
Bo Liu 0050, Liping Jing, Jia Li 0002, Jian Yu 0001, Alex Gittens, Michael W. Mahoney
Int. J. Comput. Vis.3
2019 Convolutional sparse coding for demosaicking with panchromatic pixels
Chen-Yan Bai, Jia Li 0002
Signal Process. Image Commun.2
2017 Automatic Design of High-Sensitivity Color Filter Arrays With Panchromatic Pixels
abstract
In most of existing digital cameras, color images have to be reconstructed from raw images which only have one color sensed at each pixel, as their imaging sensors are covered by color filter arrays (CFAs). At each pixel a CFA usually allows only a portion of the light spectrum to pass through and thereby reduces the light sensitivity of pixels. To address this issue, previous works have explored adding panchromatic pixels into CFAs. However, almost all existing methods assign panchromatic pixels empirically, making the designed CFAs prone to aliasing artifacts. In this paper, based on a mathematical model we propose a fully automatic approach to designing high-sensitivity CFAs using panchromatic pixels. By the frequency structure representation of CFAs, we formulate high-sensitivity CFA design as a continuous multi-objective optimization problem, where robustness to aliasing artifacts and percentage of panchromatic pixels are simultaneously maximized. We analyze the characteristics of our new formulation. According to the analysis, we develop a new method to propose frequency structure candidates, which can produce CFAs that reach a desired percentage of panchromatic pixels. Then for each candidate, we optimize parameters to obtain the final CFA, which is an appropriately balanced solution to the multi-objective optimization problem. We formulate the two design procedures as constrained optimization problems and solve them using the alternating direction method (ADM). Extensive experiments confirm the advantage of the proposed method in both low-light and normal-light conditions.
Jia Li 0002, Chen-Yan Bai, Zhouchen Lin, Jian Yu 0001
IEEE Trans. Image Process.1
2017 Optimized Color Filter Arrays for Sparse Representation-Based Demosaicking
abstract
Demosaicking is the problem of reconstructing a color image from the raw image captured by a digital color camera that covers its only imaging sensor with a color filter array (CFA). Sparse representation-based demosaicking has been shown to produce superior reconstruction quality. However, almost all existing algorithms in this category use the CFAs, which are not specifically optimized for the algorithms. In this paper, we consider optimally designing CFAs for sparse representation-based demosaicking, where the dictionary is well-chosen. The fact that CFAs correspond to the projection matrices used in compressed sensing inspires us to optimize CFAs via minimizing the mutual coherence. This is more challenging than that for traditional projection matrices because CFAs have physical realizability constraints. However, most of the existing methods for minimizing the mutual coherence require that the projection matrices should be unconstrained, making them inapplicable for designing CFAs. We consider directly minimizing the mutual coherence with the CFA's physical realizability constraints as a generalized fractional programming problem, which needs to find sufficiently accurate solutions to a sequence of nonconvex nonsmooth minimization problems. We adapt the redistributed proximal bundle method to address this issue. Experiments on benchmark images testify to the superiority of the proposed method. In particular, we show that a simple sparse representation-based demosaicking algorithm with our specifically optimized CFA can outperform LSSC [1]. To the best of our knowledge, it is the first sparse representation-based demosaicking algorithm that beats LSSC in terms of CPSNR.
Jia Li 0002, Chen-Yan Bai, Zhouchen Lin, Jian Yu 0001
IEEE Trans. Image Process.1
2016 Robust graph learning via constrained elastic-net regularization
Bo Liu 0050, Liping Jing, Jian Yu 0001, Jia Li 0002
Neurocomputing4
2016 Automatic Design of Color Filter Arrays in the Frequency Domain
abstract
In digital color imaging, the raw image is typically obtained through a single sensor covered by a color filter array (CFA), which allows only one color component to be measured at each pixel. The procedure to reconstruct a full color image from the raw image is known as demosaicking. Since the CFA may cause irreversible visual artifacts, the CFA and the demosaicking algorithm are crucial to the quality of demosaicked images. Fortunately, the design of CFAs in the frequency domain provides a theoretical approach to handling this issue. However, almost all the existing design methods in the frequency domain involve considerable human effort. In this paper, we present a new method to automatically design CFAs in the frequency domain. Our method is based on the frequency structure representation of mosaicked images. We utilize a multi-objective optimization approach to propose frequency structure candidates, in which the overlap among the frequency components of images mosaicked with the CFA is minimized. Then, we optimize parameters for each candidate, which is formulated as a constrained optimization problem. We use the alternating direction method to solve it. Our parameter optimization method is applicable to arbitrary frequency structures, including those with conjugate replicas of chrominance components. Experiments on benchmark images confirm the advantage of the proposed method.
Chen-Yan Bai, Jia Li 0002, Zhouchen Lin, Jian Yu 0001
IEEE Trans. Image Process.2
2015 Penrose Demosaicking
abstract
The Penrose pixel layout, an aperiodic pixel layout in rhombus Penrose tiling, has been shown to substantially outperform the existing square pixel layout in super-resolution. However, it was tested only on grayscale images. To study its performance on color images, we have to reconstruct regular color images from Penrose raw images, i.e., images with only one color component at each Penrose pixel, resulting in the problem of demosaicking from Penrose pixels. Penrose demosaicking is more difficult than regular demosaicking, because none of the color components of the reconstructed regular color images are available. Therefore, most of the traditional demosaicking methods do not apply. We develop a sparse representation-based method for Penrose demosaicking. Extensive experiments show that Penrose pixel layout outperforms regular pixel layouts in terms of both perceptual evaluation and S-CIELAB. The Penrose pixel layout is unique among all irregular layouts because it is uniformly three-colorable and it has only two pixel shapes, thick and thin rhombi, making its manufacturing relatively easy.
Chen-Yan Bai, Jia Li 0002, Zhouchen Lin, Jian Yu 0001, Yen-Wei Chen 0001
IEEE Trans. Image Process.2
2014 Constrained Least Squares Regression for Semi-Supervised Learning
Bo Liu 0050, Liping Jing, Jian Yu 0001, Jia Li 0002
PAKDD (2)4
2012 Multi-scale Convolutional Neural Networks for Natural Scene License Plate Detection
Jia Li 0002, Changyong Niu
ISNN (2)1