EDBT 2026 Demo / reviewers in the wild / expert
Jinshan Zeng
dblp:57/10341
· DBLP profile ↗
64ranked-venue papers
16as first author
47since 2021 · last 2026
0000-0003-1719-3358ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 11 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoRA-E2: Effective and Efficient Low-rank AdaptationabstractLow-rank adaptation (LoRA) has emerged as an efficient fine-tuning technique for large language models, enabling parameter-efficient updates while maintaining task performance. However, LoRA suffers from two key issues: 1) inefficient feature learning when the width n (embedding dimension) is large, and 2) ineffective updates to the adapter matrix A due to the initialization of B as zero. We propose LoRA-E2, which utilizes a Gaussian initialization with variance Θ(n-3/4) for A, and employs the Gauss-Seidel iteration to train B and A. We theoretically show that LoRA-E2 enables more stable and efficient feature learning with effective parameter updates over standard LoRA. Empirically, LoRA-E2 achieves consistent gains in both natural language understanding and generation tasks. On the GLUE benchmark with T5-base, it improves performance by 1-10% over LoRA. When fine-tuning LLaMA 2-7B on MetaMathQA with GSM8K as validation, LoRA-E2 surpasses LoRA by 1-2% and converges up to ∼1/43× faster. Code is available at https://github.com/whu-totemdb/LoRA-E2. Shengkun Zhu, Jinshan Zeng, Sheng Wang 0007, Yuan Sun 0003, Shangfeng Chen, Yuan Yao 0011, Qiang Yang 0001 |
WWW | 2 |
| 2026 | Distinct Polyp Generator Network for polyp segmentation
Huan Wan, Jing Ai, Xin Wei 0002, Jinshan Zeng, Jianyi Wan |
Image Vis. Comput. | 5 |
| 2026 | DeRe-Net: details restoration networks for polyp segmentation
Huan Wan, Qinqin Wang, Jinshan Zeng, Xin Wei 0002 |
Multim. Syst. | 3 |
| 2026 | Cross-lingual font generation via patch-level style contrastive learning and relative position awareness
Jinshan Zeng, Yiyang Yuan, Xijia Wang, Yefei Wang |
Pattern Recognit. | 1 |
| 2026 | Highly-Efficient Large-Scale k-means with Individual Fairness
Shengkun Zhu, Jinshan Zeng, Yuan Sun 0003, Sheng Wang 0007, Yushuai Ji, Feiping Nie 0001, Xiaodong Li 0001, Zhiyong Peng 0001 |
Proc. VLDB Endow. | 2 |
| 2026 | VTMedSeg: Geometry-Guided Vision-Text Model With Concept-Aware Fusion for Medical Ultrasound Image SegmentationabstractMedical ultrasound image segmentation is a vital diagnostic technique for identifying abnormalities, including tumors, cysts, and inflammation. To obtain accurate segmentation efficiently, many methods have been proposed. Amongst, vision–text models have shown promising progress; however, their performance is limited by the lack of high-quality paired text descriptions in ultrasound datasets. To address this issue, we propose VTMedSeg. In VTMedSeg, a geometry-guided text generator is developed to automatically synthesize medical descriptions by using connected component analysis and multi-dimensional geometric analysis. Then, the text features of these descriptions and the visual features of the corresponding ultrasound images are effectively integrated into a concept-aware cross-modal fusion module through multi-scale cross-modal alignment and medical knowledge graph modeling. Extensive experiments on four public medical ultrasound datasets demonstrate that VTMedSeg outperforms state-of-the-art vision–text models across multiple metrics. Huan Wan, Wujian Xu, Yiwen Zou, Jinshan Zeng, Xin Wei 0002 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Learning a Self-Supervised Low-Rank Decomposition Network for Hyperspectral Image Super-ResolutionabstractHyperspectral image (HSI) super-resolution, which reconstructs a high-resolution HSI (HR-HSI) through hyperspectral and multispectral image fusion (HMIF) tasks that integrate a low-resolution HSI (LR-HSI) with a high-resolution multispectral image (HR-MSI), has emerged as a promising technique for enhancing spatial–spectral quality. Recently, low-rank representations have demonstrated significant advances in various hyperspectral-related applications, providing an effective solution to HMIF tasks. However, most existing methods rely on model priors to learn the low-rank representation of HSIs, which restricts their adaptability to low-rank variations across different datasets. To address this issue, this paper introduces a self-supervised low-rank decomposition network (SSLRDN) framework specifically designed for HMIF, inspired by the observation that the HR-MSI and HR-HSI of the same scene share highly similar spatial features, whereas different hyperspectral scenes exhibit variations in both spectral and spatial features. In SSLRDN, we develop a self-supervised network to adaptively learn the low-rank decomposition (spectral subspace and spatial coefficients) across different HR-HSIs, overcoming the inefficiency of conventional alternating optimization methods where factor updates fail to mutually promote each other. Given the spatial feature consistency between HR-MSI and HR-HSI, we leverage the rich spatial information from HR-MSI to guide the learning of spatial coefficients in HR-HSI. To enhance the self-supervised learning of spatial coefficient images, we further integrate an externally pre-trained denoiser to improve their estimation accuracy, effectively fusing and mutually promoting both self-supervised and pre-trained learning paradigms. Experimental results show that the proposed method achieves superior performance in both visual quality and quantitative metrics, without requiring pretraining on external datasets. Yong Chen 0013, Xinfeng Gui, Feiwang Yuan, Wei He 0003, Jinshan Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | BraveANN: Robust Approximate Nearest Neighbor Search for Billion-Scale Vectors
Shengkun Zhu, Jinshan Zeng, Sheng Wang 0007, Yuan Sun 0003, Yuhui Lai, Zhiyong Peng 0001 |
World Wide Web (WWW) | 4 |
| 2025 | Self-Supervised Collaborative Information Bottleneck for Text Readability AssessmentabstractText readability assessment involves categorizing texts based on readers' comprehension levels. Hybrid automatic readability assessment (ARA) models, combining deep and linguistic features, have recently attracted rising attention due to their impressive performance. However, existing hybrid ARA models generally ignore the specific-intrinsic information of deep and linguistic representations, and cannot fully explore their common-intrinsic information. In this paper, we introduce a self-supervised collaborative information bottleneck (SCIB) module for ARA to address these issues. Specifically, we collaboratively consider both specific-intrinsic and common-intrinsic information of the linguistic representation and various levels of deep representations including the document-, sentence- and word-level deep representations, and yield their refined representations via a self-supervised information bottleneck scheme. Extensive experiments are conducted on four English and two Chinese corpora to demonstrate the effectiveness of the proposed model. Experimental results show that the proposed model outperforms state-of-the-art models in terms of four important evaluation metrics, and the suggested SCIB module can effectively capture the specific- and common-intrinsic information. Jinshan Zeng, Xianglong Yu, Xianchao Tong, Wenyan Xiao |
AAAI | 1 |
| 2025 | TriDE-Net: Triple-Densely Extraction Network for Precise Skin Lesion SegmentationabstractAccurate skin lesion segmentation is crucial for the quantitative analysis of skin cancer. Despite the significant advancements achieved by the deep-learning methods, the segmentation of skin lesions with irregular shapes and significant size variations is still challenging. To address the problem, we propose a Triple-Densely Extraction Network (TriDE-Net) for skin lesion segmentation, aiming to heavily extract multi-scale features in the inter- and intra-feature layers. In the TriDE-Net, a Feature-Intensive Capture Module (FICM) is designed to essentially extract multi-scale features from the intra-feature layers in a dually dense manner, and FICM is densely deployed in each skip-connection path to exploit features from the inter-feature layers. Moreover, we developed a Feature Adaptive Fusion Module (FAFM) to aggregate the decoding features to obtain accurate segmentation results. Comprehensive experiments on four widely-used skin lesion datasets consistently demonstrate that our TriDE-Net outperforms the state-of-the-art methods, with the Dice coefficient improving to 93.27%. Huan Wan, Taona Deng, Wujian Xu, Xin Wei 0002, Jinshan Zeng |
ICASSP | 5 |
| 2025 | Learning Stroke-Order Dynamics in Few-Shot Font Generation via Sequential AwarenessabstractFew-shot font generation has garnered significant attention due to its wide range of applications. The mainstream methods are based on the idea of the style and content disentangled representation learning and can be mainly categorized into two kinds of methods according to the prior used, i.e., the deep prior and glyph prior. However, the prior information used in existing methods mainly focuses on static spatial information and ignores dynamic temporal information symbolizing the internal correlation of characters, which results in stroke misalignment and poor performance on the generation of glyph articulations. To address these issues, we propose a novel few-shot font generation model by learning stroke-order dynamics via sequential awareness, where both the static spatial stroke information and dynamic temporal stroke-order information are incorporated into the generation. By leveraging these kinds of stroke information, the issues of stroke misalignment and poor articulation generation can be significantly alleviated. We conduct extensive experiments over 150 fonts, which show the superiority of the proposed model compared to state-of-the-art, and good generalization performance for the cross-lingual font generation. Jinshan Zeng, Yiyang Yuan, Yefei Wang, Xijia Wang |
ICASSP | 1 |
| 2025 | FedAPM: Federated Learning via ADMM with Partial Model PersonalizationabstractIn federated learning (FL), the assumption that datasets from different devices are independent and identically distributed (i.i.d.) often does not hold due to user differences, and the presence of various data modalities across clients makes using a single model impractical. Personalizing certain parts of the model can effectively address these issues by allowing those parts to differ across clients, while the remaining parts serve as a shared model. However, we found that partial model personalization may exacerbate client drift (each client's local model diverges from the shared model), thereby reducing the effectiveness and efficiency of FL algorithms. We propose an FL framework based on the alternating direction method of multipliers (ADMM), referred to as FedAPM, to mitigate client drift. We construct the augmented Lagrangian function by incorporating first-order and second-order proximal terms into the objective, with the second-order term providing fixed correction and the first-order term offering compensatory correction between the local and shared models. Our analysis demonstrates that FedAPM, by using explicit estimates of the Lagrange multiplier, is more stable and efficient in terms of convergence compared to other FL frameworks. We establish the global convergence of FedAPM training from arbitrary initial points to a stationary point, achieving three types of rates: constant, linear, and sublinear, under mild assumptions. We conduct experiments using four heterogeneous and multimodal datasets with different metrics to validate the performance of FedAPM. Specifically, FedAPM achieves faster and more accurate convergence, outperforming the SOTA methods with average improvements of 12.3% in test accuracy, 16.4% in F1 score, and 18.0% in AUC while requiring fewer communication rounds. Shengkun Zhu, Feiteng Nie, Jinshan Zeng, Sheng Wang 0007, Yuan Sun 0003, Yuan Yao 0011, Shangfeng Chen, Quanqing Xu, Chuanhui Yang |
KDD (2) | 3 |
| 2025 | EdgeFont: Enhancing style and content representations in few-shot font generation with multi-scale edge self-supervision
Yefei Wang, Kangyue Xiong, Yiyang Yuan, Jinshan Zeng |
Expert Syst. Appl. | 4 |
| 2025 | Few-shot font generation via stroke prompt and hierarchical representation learning
Jinshan Zeng, Yiyang Yuan, Ling Tu, Yefei Wang |
Expert Syst. Appl. | 1 |
| 2025 | EDG-CDM: A New Encoder-Guided Conditional Diffusion Model-Based Image Synthesis Method for Limited DataabstractABSTRACT The Diffusion Probabilistic Model (DM) has emerged as a powerful generative model in the field of image synthesis, capable of producing high‐quality and realistic images. However, training DM requires a large and diverse dataset, which can be challenging to obtain. This limitation weakens the model's generalisation and robustness when training data is limited. To address this issue, EDG‐CDM, an innovative encoder‐guided conditional diffusion model was proposed for image synthesis with limited data. Firstly, the authors pre‐train the encoder by introducing noise to capture the distribution of image features and generate the condition vector through contrastive learning and KL divergence. Next, the encoder undergoes further training with classification to integrate image class information, providing more favourable and versatile conditions for the diffusion model. Subsequently, the encoder is connected to the diffusion model, which is trained using all available data with encoder‐provided conditions. Finally, the authors evaluate EDG‐CDM on various public datasets with limited data, conducting extensive experiments and comparing our results with state‐of‐the‐art methods using metrics such as Fréchet Inception Distance and Inception Score. Our experiments demonstrate that EDG‐CDM outperforms existing models by consistently achieving the lowest FID scores and the highest IS scores, highlighting its effectiveness in generating high‐quality and diverse images with limited training data. These results underscore the significance of EDG‐CDM in advancing image synthesis techniques under data‐constrained scenarios. Haopeng Lei, Kaijun Liang, Mingwen Wang 0001, Jinshan Zeng, Guoliang Luo |
IET Comput. Vis. | 5 |
| 2025 | Low-Rank Tensor Meets Deep Prior: Coupling Model-Driven and Data-Driven Methods for Hyperspectral Image ReconstructionabstractSnapshot compressive imaging (SCI) captures a 3D hyperspectral image (HSI) using a 2D compressive measurement and reconstructs the desired 3D HSI from that 2D measurement. The effective reconstruction method thus is crucial in SCI. Despite recent successes of deep learning (DL)-based methods over traditional approaches, they often ignore the intrinsic characteristics of HSI and are trained for a specific imaging system using sufficient paired datasets. To address this, we propose a novel self-supervised HSI reconstruction framework called low-rank tensor meets deep prior (LDMeet), which couples model-driven and data-driven methods. The design of LDMeet is inspired by the traditional model-driven low-rank tensor prior constructed based on domain knowledge, which can explore the intrinsic global spatial-spectral correlation of HSI and make the reconstruction method interpretable. To further utilize the powerful learning ability of DL-based approaches, we introduce a self-supervised spatial-spectral guided network (SSG-Net) into LDMeet to learn the implicit deep spatial-spectral prior of HSI without requiring training data, making it adaptable to various imaging systems. An efficient alternating direction method of multiplier (ADMM) is designed to solve the LDMeet model. Comprehensive experiments confirm that our LDMeet achieves superior results compared to self-supervised HSI reconstruction methods, while also yielding competitive results with supervised learning methods. Yong Chen 0013, Feiwang Yuan, Wenzhen Lai, Jinshan Zeng, Wei He 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Fusing Global Structural and Local Deep Features for Thick Cloud Removal in Multitemporal Remote Sensing ImagesabstractOptical remote sensing images are inevitably affected by thick cloud cover, leading to information missing, which seriously hinders subsequent Earth observation tasks. Existing cloud removal methods typically focus on extracting either global or local features, but lack effective fusion of these two aspects, resulting in deficiencies in recovering global structure and fine details. However, fusing multiple features is challenging for single-driven methods due to the diversity of images and clouds. To this end, this paper proposes a novel joint-driven cloud removal framework that can simultaneously exploit the global structural and local deep features of both the image and cloud components. Specifically, in our joint-driven scheme, we devise a flexible model-driven low-rank group sparse decomposition method, in which the global temporal-spectral correlations and shared sparse features of multitemporal remote sensing images and clouds from different scenes are effectively captured. To fully account for the differences in the local features of images and clouds from different scenes, we devise effective data-driven dual self-supervised networks to adaptively learn their local deep features. Additionally, we design an efficient half-quadratic splitting algorithm to iteratively optimize the joint-driven model, fostering mutual enhancement between the two and achieving further performance improvements. Experimental results show that the proposed method outperforms existing thick cloud removal approaches in both simulated and real-world datasets, under varying factors such as spectral bands, cloud coverage, time nodes, and spatial resolutions. Yeqi Xu, Yong Chen 0013, Wei He 0003, Min Huang 0005, Jinshan Zeng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Feature Fusion-Guided Network With Sparse Prior Constraints for Unsupervised Hyperspectral Image Quality ImprovementabstractDue to imaging hardware limitations and atmospheric interference, hyperspectral image (HSI) often suffers from low spatial resolution or mixed noise degradation. HSI fusion and denoising are two key strategies to improve HSI quality. Traditional model-based methods rely on data-specific manual priors, while supervised deep learning methods are typically developed specifically for a single task and require a large number of training datasets. To address these limitations, we propose a novel unsupervised feature fusion-guided network (UFFGNet) as a general prior that effectively leverages multi-scale semantic features from a guidance image while incorporating sparse prior constraints to suppress outliers. Specifically, UFFGNet comprises a deep feature extraction network to capture multiscale semantic features from a guidance image, and an attentionbased feature generation network that generates an output image from random noise. These two networks are connected by a feature refinement module to embed the refined features from the feature aggregation module into the generation network. Furthermore, the sparse prior constraint is incorporated to model sparse noise (including impulse noise, stripe artifacts, and deadlines) in HSI, thereby improving the robustness of UFFGNet. The proposed network optimizes network parameters in an unsupervised manner without external additional training data, using the fidelity term in the degradation model as a loss function to learn the prior information of the original image. Extensive experiments demonstrate that the proposed UFFGNet outperforms state-of-the-art methods in both HSI fusion and denoising tasks, significantly improving HSI quality. Feiwang Yuan, Yong Chen 0013, Wei He 0003, Jinshan Zeng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | GuessGas: Tell Me Fine-Grained Gas Consumption of My Smart Contract and WhyabstractSmart contracts with excessive gas consumption can cause economic losses, such as black hole contracts. Actual gas consumption depends on runtime information and has a probability distribution under different runtime situations. However, existing static analysis tools (e.g., Solc) cannot define runtime information and only provide an approximate upper bound on gas consumption without explanation. To address the challenge, we propose a label named GCL, which describes the probability distribution of gas consumption, a code representation method containing domain features and a graph neural network (GNN) named attention-based graph isomorphism network (AGIN) oriented to domain feature, and SubgraphGas, a domain-oriented subgraph-level GNN explanation model. By combining AGIN and SubgraphGas, we have created a new explainable gas consumption prediction model (EGE). Our evaluations show that EGE outperforms prediction schemes based on Bi-LSTM. And EGE has similar explainability accuracy to general methods, but it is more efficient. Renxiong Chen, Zhenchang Xing, Jinshan Zeng, Qinghua Lu 0001, Xiwei Xu 0001 |
IEEE Trans. Reliab. | 4 |
| 2024 | InterpretARA: Enhancing Hybrid Automatic Readability Assessment with Linguistic Feature Interpreter and Contrastive LearningabstractThe hybrid automatic readability assessment (ARA) models that combine deep and linguistic features have recently received rising attention due to their impressive performance. However, the utilization of linguistic features is not fully realized, as ARA models frequently concentrate excessively on numerical values of these features, neglecting valuable structural information embedded within them. This leads to limited contribution of linguistic features in these hybrid ARA models, and in some cases, it may even result in counterproductive outcomes. In this paper, we propose a novel hybrid ARA model named InterpretARA through introducing a linguistic interpreter to better comprehend the structural information contained in linguistic features, and leveraging the contrastive learning that enables the model to understand relative difficulty relationships among texts and thus enhances deep representations. Both document-level and segment-level deep representations are extracted and used for the readability assessment. A series of experiments are conducted over four English corpora and one Chinese corpus to demonstrate the effectiveness of the proposed model. Experimental results show that InterpretARA outperforms state-of-the-art models in most corpora, and the introduced linguistic interpreter can provide more useful information than existing ways for ARA. Jinshan Zeng, Xianchao Tong, Xianglong Yu, Wenyan Xiao |
AAAI | 1 |
| 2024 | SCI-Font: Enhancing Content-Style Representation for Chinese Calligraphy Generation with Skeleton, Contour and Inexact Paired Data
Yefei Wang, Jialu Xiong, Jinshan Zeng |
ICANN (3) | 5 |
| 2024 | CLIP-MSA: Incorporating Inter-Modal Dynamics and Common Knowledge to Multimodal Sentiment Analysis With ClipabstractMultimodal Sentiment Analysis (MSA) aims to yield the sentiment polarities of speakers in video streams based on multiple modal features such as textual, acoustic and visual features, and has attracted amounts of attention in recent years. Existing MSA models often yield unimodal embeddings from the associated modal features individually, while overlooking the importance of inter-modal dynamics and common knowledge in the extraction of unimodal embeddings, resulting in the limited performance. In this paper, we suggest a novel MSA model called CLIP-MSA through incorporating the inter-modal dynamics and common knowledge into the generation of unimodal representations with the Contrastive Language-Image Pre-training (CLIP), and fusing the textual, acoustic and visual representations with a hierarchical co-attention mechanism. Numerous experimental results over two benchmark datasets show that the proposed model outperforms existing state-of-the-art models on CMU-MOSI, and provides competitive performance on CMU-MOSEI, in terms of four commonly used evaluation metrics. Pingting Cai, Tanyue Nie, Jinshan Zeng |
ICASSP | 4 |
| 2024 | CLIP-Font: Sementic Self-Supervised Few-Shot Font Generation with ClipabstractFont design is a very resource-intensive endeavor, especially for intricate fonts. The task of few-shot font generation (FFG) has attracted great interest recently. This method captures style from a limited set of reference glyphs and then transfers it to other characters to generate diverse style fonts. Existing FFG methods mainly revolve around learning font content or style. However, these methods often only learn content or style or lack the ability to represent style and content, resulting in poor font quality. To address these issues, we introduce CLIP-Font—a novel few-shot font generation model. CLIP-Font uses font text semantics for self-supervision to guide font generation at the content level, and uses attention-based contrast learning at the style level to capture the representation capabilities of the font fine-grained style enhancement model. Experimental results on various datasets demonstrate the effectiveness of our method, surpassing the performance of existing FFG techniques. Jialu Xiong, Yefei Wang, Jinshan Zeng |
ICASSP | 3 |
| 2024 | SCA-Font: Enhancing Few-Shot Generation with Style-Content AggregationabstractThe few-shot font generation (FFG) task aims to create a new font library using only a small number of reference samples. The predominated methods for this task are mainly based on the style-content disentangled representation learning. Existing style-content disentangling based few-shot font generation models are mainly devoted to the extraction of better content and style features by leveraging extra prior information such as strokes and skeletons, or introducing auxiliary networks while ignoring the aggregation scheme of style and content features. To address this issue, we propose a novel few-shot font generation method called SCA-Font by introducing an effective style-content feature aggregation module (SCAM), where the content features from the source characters and the style features from the target reference characters are effectively aggregated by a novel neural network. Experimental results on a dataset of 35 font styles collected by ourselves demonstrate that the proposed SCA-Font model outperforms state-of-the-art models in both quantitative results and the quality of generated characters. We also verify the effect of the number of shots for the proposed model. Numerical experiment results show that six shots of reference characters are preferred to achieve the best performance of the proposed model. Yefei Wang, Sunzhe Yang, Kangyue Xiong, Jinshan Zeng |
IJCNN | 4 |
| 2024 | Unidirectional Spatial and Spectral Smoothed Tensor Ring Decomposition for Hyperspectral Image Denoising and DestripingabstractIn this letter, we propose a novel unidirectional spatial and spectral smoothed tensor ring (U3STR) decomposition for hyperspectral image (HSI) denoising and destriping. The powerful tensor ring (TR) decomposition is introduced to explore the global spatial-spectral correlation of HSI, which transforms the restoration of HSI into estimating three TR factors. To address the local spatial-spectral smoothness of HSI and the directional characteristic of stripe noise, unidirectional spatial and spectral smoothed constraints are applied to the horizontal spatial and spectral TR factors, respectively. Moreover, considering that the stripe noise shares spatial correlation and local smoothness with the image component, we strategically utilize band-by-band low rank and unidirectional total variation (TV) regularization, effectively disentangling stripe noise from the image content without conflicting the image regularization. The proposed U3STR model is solved by the alternating direction method of multipliers (ADMM) algorithm effectively. Experimental results demonstrate that our method outperforms other HSI restoration methods in denoising and destriping, notably enhancing the quality of the restored image by an average of 3 dB over existing methods. Yong Chen 0013, Jinshan Zeng, Wei He 0003, Min Huang 0005 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Sequential Manipulation Against Rank Aggregation: Theory and AlgorithmabstractRank aggregation with pairwise comparisons is widely encountered in sociology, politics, economics, psychology, sports, etc. Given the enormous social impact and the consequent incentives, the potential adversary has a strong motivation to manipulate the ranking list. However, the ideal attack opportunity and the excessive adversarial capability cause the existing methods to be impractical. To fully explore the potential risks, we leverage an online attack on the vulnerable data collection process. Since it is independent of rank aggregation and lacks effective protection mechanisms, we disrupt the data collection process by fabricating pairwise comparisons without knowledge of the future data or the true distribution. From the game-theoretic perspective, the confrontation scenario between the online manipulator and the ranker who takes control of the original data source is formulated as a distributionally robust game that deals with the uncertainty of knowledge. Then we demonstrate that the equilibrium in the above game is potentially favorable to the adversary by analyzing the vulnerability of the sampling algorithms such as Bernoulli and reservoir methods. According to the above theoretical analysis, different sequential manipulation policies are proposed under a Bayesian decision framework and a large class of parametric pairwise comparison models. For attackers with complete knowledge, we establish the asymptotic optimality of the proposed policies. To increase the success rate of the sequential manipulation with incomplete knowledge, a distributionally robust estimator, which replaces the maximum likelihood estimation in a saddle point problem, provides a conservative data generation solution. Finally, the corroborating empirical evidence shows that the proposed method manipulates the results of rank aggregation methods in a sequential manner. Ke Ma 0001, Qianqian Xu 0001, Jinshan Zeng, Wei Liu 0005, Xiaochun Cao, Yingfei Sun, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | A guidable nonlocal low-rank approximation model for hyperspectral image denoising
Yong Chen 0013, Jinshan Zeng, Wenzhen Lai, Xinfeng Gui, Tai-Xiang Jiang |
Signal Process. | 3 |
| 2024 | Thick Cloud Removal in Multitemporal Remote Sensing Images via Low-Rank Regularized Self-Supervised NetworkabstractThe existence of thick clouds covers the comprehensive Earth observation of optical remote sensing images (RSIs). Cloud removal is an effective and economical preprocessing step to improve the subsequent applications of RSIs. Deep learning (DL)-based methods have attracted much attention and achieved state-of-the-art results. However, most of these methods suffer from the following issues: 1) ignore the physical characteristics of RSIs; 2) require paired images with/without cloud or extra auxiliary images (such as SAR); and 3) demand the cloud mask. These issues might have limited the flexibility of existing networks. In this paper, we propose a novel low-rank regularized self-supervised network (LRRSSN) that couples model-driven and data-driven methods to remove the thick cloud from multitemporal remote sensing images (MRSIs). First, motivated by the equal importance of image and cloud components as well as their intrinsic characteristics, we decompose the observed image into low-rank image and structural sparse cloud components. In this way, we obtain a model-driven thick cloud removal method where the spectral-temporal low-rank correlation of the image component and the spectral structural sparsity of the cloud component are effectively exploited. Second, to capture the complex nonlinear features of different scenarios, the data-driven self-supervised network that does not require external training datasets is designed to explore the deep prior of the image component. Third, the coupled model-driven and data-driven LRRSSN is optimized by an efficient half-quadratic splitting algorithm. Finally, without knowing the exact cloud mask, we estimate the cloud mask to preserve information in cloud-free areas as much as possible. Experiments conducted in synthetic and real-world scenarios demonstrate the effectiveness of the proposed approach. Yong Chen 0013, Wei He 0003, Jinshan Zeng, Min Huang 0005, Yu-Bang Zheng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Fast Large-Scale Hyperspectral Image Denoising via Noniterative Low-Rank Subspace RepresentationabstractDenoising of hyperspectral image (HSI) is challenging, especially when dealing with large-scale data. Model-based methods show promise in HSI denoising due to their good generalization, but they suffer from computational complexity due to complex priors [like nonlocal self-similarity (NSS)] and iterations, resulting in low efficiency for large-scale HSI processing. To address these challenges, we propose a fast large-scale HSI denoising (FallHyDe) method based on noniterative low-rank (LR) subspace representation to enjoy high denoising efficiency, effectiveness, and flexibility simultaneously. By leveraging the global spectral property of HSI, FallHyDe efficiently estimates spectral subspace and spatial representation coefficients (SRCs) from the observed noisy HSI, reducing computation complexity caused by the high spectral dimension during processing. In addition, we innovatively explore the presence of high signal-to-noise ratio bands (HSNRBs) in real HSI, enabling fast SRC estimation through a least squares problem without relying on complex priors and iterations. FallHyDe requires neither iteration nor parameter tuning, enabling our method to process large-scale HSI denoising quickly and flexibly. Experimental results on both simulated and real HSI datasets demonstrate that our proposed method not only achieves competitive results in quality but also speeds up the restoration by more than ten times than the representative fast HSI denoising methods. The code is available athttps://chenyong1993.github.io/yongchen.github.io/. Yong Chen 0013, Jinshan Zeng, Wei He 0003, Xi-Le Zhao, Tai-Xiang Jiang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Hyperspectral Compressive Snapshot Reconstruction via Coupled Low-Rank Subspace Representation and Self-Supervised Deep NetworkabstractCoded aperture snapshot spectral imaging (CASSI) is an important technique for capturing three-dimensional (3D) hyperspectral images (HSIs), and involves an inverse problem of reconstructing the 3D HSI from its corresponding coded 2D measurements. Existing model-based and learning-based methods either could not explore the implicit feature of different HSIs or require a large amount of paired data for training, resulting in low reconstruction accuracy or poor generalization performance as well as interpretability. To remedy these deficiencies, this paper proposes a novel HSI reconstruction method, which exploits the global spectral correlation from the HSI itself through a formulation of model-driven low-rank subspace representation and learns the deep prior by a data-driven self-supervised deep learning scheme. Specifically, we firstly develop a model-driven low-rank subspace representation to decompose the HSI as the product of an orthogonal basis and a spatial representation coefficient, then propose a data-driven deep guided spatial-attention network (called DGSAN) to adaptively reconstruct the implicit spatial feature of HSI by learning the deep coefficient prior (DCP), and finally embed these implicit priors into an iterative optimization framework through a self-supervised training way without requiring any training data. Thus, the proposed method shall enhance the reconstruction accuracy, generalization ability, and interpretability. Extensive experiments on several datasets and imaging systems validate the superiority of our method. The source code and data of this article will be made publicly available at https://github.com/ChenYong1993/LRSDN. Yong Chen 0013, Wenzhen Lai, Wei He 0003, Xi-Le Zhao, Jinshan Zeng |
IEEE Trans. Image Process. | 5 |
| 2024 | Combining Low-Rank and Deep Plug-and-Play Priors for Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) is a promising technique that captures a 3-D hyperspectral image (HSI) by a 2-D detector in a compressed manner. The ill-posed inverse process of reconstructing the HSI from their corresponding 2-D measurements is challenging. However, current approaches either neglect the underlying characteristics, such as high spectral correlation, or demand abundant training datasets, resulting in an inadequate balance among performance, generalizability, and interpretability. To address these challenges, in this article, we propose a novel approach called LR2DP that integrates the model-driven low-rank prior and data-driven deep priors for SCI reconstruction. This approach not only captures the spectral correlation and deep spatial features of HSI but also takes advantage of both model-based and learning-based methods without requiring any extra training datasets. Specifically, to preserve the strong spectral correlation of the HSI effectively, we propose that the HSI lies in a low-rank subspace, thereby transforming the problem of reconstructing the HSI into estimating the spectral basis and spatial representation coefficient. Inspired by the mutual promotion of unsupervised deep image prior (DIP) and trained deep denoising prior (DDP), we integrate the unsupervised network and pre-trained deep denoiser into the plug-and-play (PnP) regime to estimate the representation coefficient together, aiming to explore the internal target image prior (learned by DIP) and the external training image prior (depicted by pre-trained DDP) of the HSI. An effective half-quadratic splitting (HQS) technique is employed to optimize the proposed HSI reconstruction model. Extensive experiments on both simulated and real datasets demonstrate the superiority of the proposed method over the state-of-the-art approaches. Yong Chen 0013, Xinfeng Gui, Jinshan Zeng, Xi-Le Zhao, Wei He 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Revealing the Unseen: AI Chain on LLMs for Predicting Implicit Dataflows to Generate Dataflow Graphs in Dynamically Typed CodeabstractDataflow graphs (DFGs) capture definitions (defs) and uses across program blocks, which is a fundamental program representation for program analysis, testing and maintenance. However, dynamically typed programming languages like Python present implicit dataflow issues that make it challenging to determine def-use flow information at compile time. Static analysis methods like Soot and WALA are inadequate for handling these issues, and manually enumerating comprehensive heuristic rules is impractical. Large pre-trained language models (LLMs) offer a potential solution, as they have powerful language understanding and pattern matching abilities, allowing them to predict implicit dataflow by analyzing code context and relationships between variables, functions, and statements in code. We propose leveraging LLMs’ in-context learning ability to learn implicit rules and patterns from code representation and contextual information to solve implicit dataflow problems. To further enhance the accuracy of LLMs, we design a five-step chain of thought (CoT) and break it down into an Artificial Intelligence (AI) chain, with each step corresponding to a separate AI unit to generate accurate DFGs for Python code. Our approach’s performance is thoroughly assessed, demonstrating the effectiveness of each AI unit in the AI Chain. Compared to static analysis, our method achieves 82% higher def coverage and 58% higher use coverage in DFG generation on implicit dataflow. We also prove the indispensability of each unit in the AI Chain. Overall, our approach offers a promising direction for building software engineering tools by utilizing foundation models, eliminating significant engineering and maintenance effort, but focusing on identifying problems for AI to solve. Zhiwen Luo, Zhenchang Xing, Jinshan Zeng, Jieshan Chen, Xiwei Xu 0001, Yong Chen 0013 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Enhancing Chinese Calligraphy Generation with Contour and Region-aware AttentionabstractChinese calligraphy generation is an important problem involved in many applications. Existing methods generally regard it as an image-to-image translation problem. However, existing models meet the challenge of poor generation performance for the Chinese calligraphy generation due to the lack of effective guided information. In this paper, we propose a novel model called CRA-GAN for the Chinese calligraphy generation through utilizing the contours of calligraphy characters as certain guided information and introducing a region-aware attention module to capture their content regions, motivated by the observation that the contour and content region provide certain delicate characteristics on the calligraphy style and content, respectively. Noticing that there is usually a large glyph difference between the source and target fonts, resulting in the degradation of model performance, we borrow the idea of adaptive pre-deformation from the literature to address this issue. A series of experiments are conducted to show the effectiveness of the suggested contour and region-aware attention, as well as the used adaptive pre-deformation operation. The outperformance of the proposed model over the state-of-the-art models is also demonstrated through extensive quantitative and qualitative comparisons over nine Chinese calligraphic font datasets. Jinshan Zeng, Ling Tu, Yefei Wang, Jiguo Zeng |
IJCNN | 1 |
| 2023 | Zero-Shot Chinese Character Recognition with Stroke- and Radical-Level DecompositionsabstractZero-shot Chinese character recognition has attracted rising attention in recent years. Existing methods for this problem are mainly based on either certain low-level stroke-based decomposition or medium-level radical-based decomposition. Considering that the stroke- and radical-level decompositions can provide different levels of information, we propose an effective zero-shot Chinese character recognition method by combining them. The proposed method consists of a training stage and an inference stage. In the training stage, we adopt two similar encoder-decoder models to yield the estimates of stroke and radical encodings, which together with the true encodings are then used to formalize the associated stroke and radical losses for training. A similarity loss is introduced to regularize stroke and radical encoders to yield features of the same characters with high correlation. In the inference stage, two key modules, i.e., the stroke screening module (SSM) and feature matching module (FMM) are introduced to tackle the deterministic and confusing cases respectively. In particular, we introduce an effective stroke rectification scheme in FMM to enlarge the candidate set of characters for final inference. Numerous experiments over three benchmark datasets covering the handwritten, printed artistic and street view scenarios are conducted to demonstrate the effectiveness of the proposed method. Numerical results show that the proposed method outperforms the state-of-the-art methods in both character and radical zero-shot settings, and maintains competitive performance in the traditional seen character setting. Jinshan Zeng, Ruiying Xu, Hongwei Li 0017 |
IJCNN | 1 |
| 2023 | F3KM: Federated, Fair, and Fast k-meansabstractThis paper proposes a federated, fair, and fast k-means algorithm (F3KM) to solve the fair clustering problem efficiently in scenarios where data cannot be shared among different parties. The proposed algorithm decomposes the fair k-means problem into multiple subproblems and assigns each subproblem to a client for local computation. Our algorithm allows each client to possess multiple sensitive attributes (or have no sensitive attributes). We propose an in-processing method that employs the alternating direction method of multipliers (ADMM) to solve each subproblem. During the procedure of solving subproblems, only the computation results are exchanged between the server and the clients, without exchanging the raw data. Our theoretical analysis shows that F3KM is efficient in terms of both communication and computation complexities. Specifically, it achieves a better trade-off between utility and communication complexity, and reduces the computation complexity to linear with respect to the dataset size. Our experiments show that F3KM achieves a better trade-off between utility and fairness than other methods. Moreover, F3KM is able to cluster five million points in one hour, highlighting its impressive efficiency. Shengkun Zhu, Quanqing Xu, Jinshan Zeng, Sheng Wang 0007, Yuan Sun 0003, Zhifeng Yang, Chuanhui Yang, Zhiyong Peng 0001 |
Proc. ACM Manag. Data | 3 |
| 2023 | Exploring Structural Sparsity of Deep Networks Via Inverse Scale SpacesabstractThe great success of deep neural networks is built upon their over-parameterization, which smooths the optimization landscape without degrading the generalization ability. Despite the benefits of over-parameterization, a huge amount of parameters makes deep networks cumbersome in daily life applications. On the other hand, training neural networks without over-parameterization faces many practical problems, e.g., being trapped in the local optimal. Though techniques such as pruning and distillation are developed, they are expensive in fully training a dense network as backward selection methods; and there is still a void on systematically exploring forward selection methods for learning structural sparsity in deep networks. To fill in this gap, this paper proposes a new approach based on differential inclusions of inverse scale spaces. Specifically, our method can generate a family of models from simple to complex ones along the dynamics via coupling a pair of parameters, such that over-parameterized deep models and their structural sparsity can be explored simultaneously. This kind of differential inclusion scheme has a simple discretization, dubbed Deep structure splitting Linearized Bregman Iteration (DessiLBI), whose global convergence in learning deep networks could be established under the Kurdyka-Łojasiewicz framework. Particularly, we explore several applications of DessiLBI, including finding sparse structures of networks directly via the coupled structure parameter and growing networks from simple to complex ones progressively. Experimental evidence shows that our method achieves comparable and even better performance than the competitive optimizers in exploring the sparse structure of several widely used backbones on the benchmark datasets. Remarkably, with early stopping, our method unveils "winning tickets" in early epochs: the effective sparse network structures with comparable test accuracy to fully trained over-parameterized models, that are further transferable to similar alternative tasks. Furthermore, our method is able to grow networks efficiently with adaptive filter configurations, demonstrating the good performance with much less computational cost. Codes and models can be downloaded at https://github.com/DessiLBI2020/DessiLBI. Yanwei Fu 0001, Chen Liu 0030, Zuyuan Zhong, Xinwei Sun 0001, Jinshan Zeng, Yuan Yao 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | A Tale of HodgeRank and Spectral Method: Target Attack Against Rank Aggregation is the Fixed Point of Adversarial GameabstractRank aggregation with pairwise comparisons has shown promising results in elections, sports competitions, recommendations, and information retrieval. However, little attention has been paid to the security issue of such algorithms, in contrast to numerous research work on the computational and statistical characteristics. Driven by huge profit, the potential adversary has strong motivation and incentives to manipulate the ranking list. Meanwhile, the intrinsic vulnerability of the rank aggregation methods is not well studied in the literature. To fully understand the possible risks, we focus on the purposeful adversary who desires to designate the aggregated results by modifying the pairwise data in this paper. From the perspective of the dynamical system, the attack behavior with a target ranking list is a fixed point belonging to the composition of the adversary and the victim. To perform the targeted attack, we formulate the interaction between the adversary and the victim as a game-theoretic framework consisting of two continuous operators while Nash equilibrium is established. Then two procedures against HodgeRank and RankCentrality are constructed to produce the modification of the original data. Furthermore, we prove that the victims will produce the target ranking list once the adversary masters the complete information. It is noteworthy that the proposed methods allow the adversary only to hold incomplete information or imperfect feedback and perform the purposeful attack. The effectiveness of the suggested target attack strategies is demonstrated by a series of toy simulations and several real-world data experiments. These experimental results show that the proposed methods could achieve the attacker's goal in the sense that the leading candidate of the perturbed ranking list is the designated one by the adversary. Ke Ma 0001, Qianqian Xu 0001, Jinshan Zeng, Guorong Li, Xiaochun Cao, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Image Processing and Control of Tracking Intelligent Vehicle Based on Grayscale CameraabstractIn order to realize the rapid and stable recognition and automatic tracking of various complex roads by the intelligent vehicles, this paper proposes image processing and cascade Proportion Integration Differentiation (PID) steering and speed control algorithms based on CMOS grayscale cameras in the context of the national college student intelligent vehicle competition. First, the grayscale image of the track is acquired by the grayscale camera. Then, the Otsu method is used to binarize the image, and the information of black boundary guide line is extracted. In order to improve the speed of the race, various track elements in the image are identified and classified, and the deviation between the actual centerline position and the ideal centerline position of the intelligent vehicle is calculated. Third, the discrete incremental cascade PID control algorithm is used to calculate the pulse width modulation (PWM) signal corresponding to the deviation. And the PWM signal is acted on the steering motor through the driving circuit, driving the intelligent vehicle to always drive along the middle road, so as to achieve the purpose of automatic tracking guidance. Experiments prove that the intelligent vehicle of this design can identify complex roads quickly and in a stable way, accurately complete automatic tracking, and obtain higher speed performance. Jinshan Zeng, Hongtu Xie |
IPAS | 4 |
| 2022 | Fully corrective gradient boosting with squared hinge: Fast learning rates and early stopping
Jinshan Zeng, Shaobo Lin |
Neural Networks | 1 |
| 2022 | Poisoning Attack Against Estimating From Pairwise ComparisonsabstractAs pairwise ranking becomes broadly employed for elections, sports competitions, recommendation, information retrieval and so on, attackers have strong motivation and incentives to manipulate or disrupt the ranking list. They could inject malicious comparisons into the training data to fool the target ranking algorithm. Such a technique is called "poisoning attack" in regression and classification tasks. In this paper, to the best of our knowledge, we initiate the first systematic investigation of data poisoning attack on the pairwise ranking algorithms, which can be generally formalized as the dynamic and static games between the ranker and the attacker, and can be modeled as certain kinds of integer programming problems mathematically. To break the computational hurdle of the underlying integer programming problems, we reformulate them into the distributionally robust optimization (DRO) problems, which are computational tractable. Based on such DRO formulations, we propose two efficient poisoning attack algorithms and establish the associated theoretical guarantees including the existence of Nash equilibrium and the generalization ability bounds. The effectiveness of the suggested poisoning attack strategies is demonstrated by a series of toy simulations and several real data experiments. These experimental results show that the proposed methods can significantly reduce the performance of the ranker in the sense that the correlation between the true ranking list and the aggregated results with toxic data can be decreased dramatically. Ke Ma 0001, Qianqian Xu 0001, Jinshan Zeng, Xiaochun Cao, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Hyperspectral Image Denoising Using Factor Group Sparsity-Regularized Nonconvex Low-Rank ApproximationabstractHyperspectral image (HSI) mixed noise removal is a fundamental problem and an important preprocessing step in remote sensing fields. The low-rank approximation-based methods have been verified effective to encode the global spectral correlation for HSI denoising. However, due to the large scale and complexity of real HSI, previous low-rank HSI denoising techniques encounter several problems, including coarse rank approximation (such as nuclear norm), the high computational cost of singular value decomposition (SVD) (such as Schatten$p$-norm), and adaptive rank selection (such as low-rank factorization). In this article, two novel factor group sparsity-regularized nonconvex low-rank approximation (FGSLR) methods are introduced for HSI denoising, which can simultaneously overcome the mentioned issues of previous works. The FGSLR methods capture the spectral correlation via low-rank factorization, meanwhile utilizing factor group sparsity regularization to further enhance the low-rank property. It is SVD-free and robust to rank selection. Moreover, FGSLR is equivalent to Schatten$p$-norm approximation (Theorem 1), and thus FGSLR is tighter than the nuclear norm in terms of rank approximation. To preserve the spatial information of HSI in the denoising process, the total variation regularization is also incorporated into the proposed FGSLR models. Specifically, the proximal alternating minimization is designed to solve the proposed FGSLR models. Experimental results have demonstrated that the proposed FGSLR methods significantly outperform existing low-rank approximation-based HSI denoising methods. Yong Chen 0013, Ting-Zhu Huang, Wei He 0003, Xi-Le Zhao, Hongyan Zhang 0001, Jinshan Zeng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Exploring Nonlocal Group Sparsity Under Transform Learning for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising has been regarded as an effective and economical preprocessing step in data subsequent applications. Recent nonlocal low-rank approximation on each full band patch group has demonstrated their superiority for HSI denoising. These methods, however, directly design the low-rank regularization to the grouped patch image itself (i.e., original domain), which ignores the spatial information of the grouped patch image and cannot explores the potential structure. To address these issues, this paper proposes a nonlocal group sparsifying transform learning method (dubbed TLNLGS) for HSI denoising. Motivated by the global spectral correlation in the HSI, we firstly impose a certain low-dimensional subspace hypothesis over the HSI to prevent the heavy computation burden with the spectral band increases, and then explore a discriminatively intrinsic nonlocal group sparse prior of the reduced image by transform model. The learned group sparse prior can not only excavate the nonlocal self-similarity as recent nonlocal low-rank approximation methods but also preserve the local spatial smooth structure of the image. Moreover, compared with the fixed transform domain (e.g., gradient and discrete cosine transformation domains), the transform learning scheme can improve the sparse representation ability. An efficient block coordinate descent (BCD) algorithm is developed to solve the proposed model. Extensive experiments, including simulated and real HSI datasets, indicate the superiority of the proposed TLNLGS method over the state-of-the-art HSI denoising approaches. Yong Chen 0013, Wei He 0003, Xi-Le Zhao, Ting-Zhu Huang, Jinshan Zeng, Hui Lin 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Hyperspectral and Multispectral Image Fusion Using Factor Smoothed Tensor Ring DecompositionabstractFusing a pair of low-spatial-resolution hyperspectral image (LR-HSI) and high-spatial-resolution multispectral image (HR-MSI) has been regarded as an effective and economical strategy to achieve HR-HSI, which is essential to many applications. Among existing fusion models, the tensor ring (TR) decomposition-based model has attracted rising attention due to its superiority in approximating high-dimensional data compared to other traditional matrix/tensor decomposition models. Unlike directly estimating HR-HSI in traditional models, the TR fusion model translates the fusion procedure into an estimate of the TR factor of HR-HSI, which can efficiently capture the spatial–spectral correlation of HR-HSI. Although the spatial–spectral correlation has been preserved well by TR decomposition, the spatial–spectral continuity of HR-HSI is ignored in existing TR decomposition models, sometimes resulting in poor quality of reconstructed images. In this article, we introduce a factor smoothed regularization for TR decomposition to capture the spatial–spectral continuity of HR-HSI. As a result, our proposed model is calledfactor smoothed TR decompositionmodel, dubbedFSTRD. In order to solve the suggested model, we develop an efficient proximal alternating minimization algorithm. A series of experiments on four synthetic datasets and one real-world dataset show that the quality of reconstructed images can be significantly improved by the introduced factor smoothed regularization, and thus, the suggested method yields the best performance by comparing it to state-of-the-art methods. Yong Chen 0013, Jinshan Zeng, Wei He 0003, Xi-Le Zhao, Ting-Zhu Huang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | StrokeGAN: Reducing Mode Collapse in Chinese Font Generation via Stroke EncodingabstractThe generation of stylish Chinese fonts is an important problem involved in many applications. Most of existing generation methods are based on the deep generative models, particularly, the generative adversarial networks (GAN) based models. However, these deep generative models may suffer from the mode collapse issue, which significantly degrades the diversity and quality of generated results. In this paper, we introduce a one-bit stroke encoding to capture the key mode information of Chinese characters and then incorporate it into CycleGAN, a popular deep generative model for Chinese font generation. As a result we propose an efficient method called StrokeGAN, mainly motivated by the observation that the stroke encoding contains amount of mode information of Chinese characters. In order to reconstruct the one-bit stroke encoding of the associated generated characters, we introduce a stroke-encoding reconstruction loss imposed on the discriminator. Equipped with such one-bit stroke encoding and stroke-encoding reconstruction loss, the mode collapse issue of CycleGAN can be significantly alleviated, with an improved preservation of strokes and diversity of generated characters. The effectiveness of StrokeGAN is demonstrated by a series of generation tasks over nine datasets with different fonts. The numerical results demonstrate that StrokeGAN generally outperforms the state-of-the-art methods in terms of content and recognition accuracies, as well as certain stroke error, and also generates more realistic characters. Jinshan Zeng, Mingwen Wang 0001, Yuan Yao 0011 |
AAAI | 1 |
| 2021 | On ADMM in Deep Learning: Convergence and Saturation-AvoidanceabstractIn this paper, we develop an alternating direction method of multipliers (ADMM) for deep neural networks training with sigmoid-type activation functions (called sigmoid-ADMM pair), mainly motivated by the gradient-free nature of ADMM in avoiding the saturation of sigmoid-type activations and the advantages of deep neural networks with sigmoid-type activations (called deep sigmoid nets) over their rectified linear unit (ReLU) counterparts (called deep ReLU nets) in terms of approximation. In particular, we prove that the approximation capability of deep sigmoid nets is not worse than that of deep ReLU nets by showing that ReLU activation fucntion can be well approximated by deep sigmoid nets with two hidden layers and finitely many free parameters but not vice-verse. We also establish the global convergence of the proposed ADMM for the nonlinearly constrained formulation of the deep sigmoid nets training from arbitrary initial points to a Karush-Kuhn-Tucker (KKT) point at a rate of order O(1/k). Besides sigmoid activation, such a convergence theorem holds for a general class of smooth activations. Compared with the widely used stochastic gradient descent (SGD) algorithm for the deep ReLU nets training (called ReLU-SGD pair), the proposed sigmoid-ADMM pair is practically stable with respect to the algorithmic hyperparameters including the learning rate, initial schemes and the pro-processing of the input data. Moreover, we find that to approximate and learn simple but important functions the proposed sigmoid-ADMM pair numerically outperforms the ReLU-SGD pair. Jinshan Zeng, Shaobo Lin, Yuan Yao 0011, Ding-Xuan Zhou |
J. Mach. Learn. Res. | 1 |
| 2021 | Fast Stochastic Ordinal Embedding With Variance Reduction and Adaptive Step SizeabstractLearning representation from relative similarity comparisons, often called ordinal embedding, gains rising attention in recent years. Most of the existing methods are based on semi-definite programming (SDP), which is generally time-consuming and degrades the scalability, especially confronting large-scale data. To overcome this challenge, we propose a stochastic algorithm called SVRG-SBB, which has the following features: i) achieving good scalability via dropping positive semi-definite (PSD) constraints as serving a fast algorithm, i.e., stochastic variance reduced gradient (SVRG) method, and ii) adaptive learning via introducing a new, adaptive step size called the stabilized Barzilai-Borwein (SBB) step size. Theoretically, under some natural assumptions, we show theO(1/T) O(1T) rate of convergence to a stationary point of the proposed algorithm, where T T is the number of total iterations. Under the further Polyak-Łojasiewicz assumption, we can show the global linear convergence (i.e., exponentially fast converging to a global optimum) of the proposed algorithm. Numerous simulations and real-world data experiments are conducted to show the effectiveness of the proposed algorithm by comparing with the state-of-the-art methods, notably, much lower computational cost with good prediction performance. Ke Ma 0001, Jinshan Zeng, Jiechao Xiong, Qianqian Xu 0001, Xiaochun Cao, Wei Liu 0005, Yuan Yao 0011 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Random Sketching for Neural Networks With ReLUabstractTraining neural networks is recently a hot topic in machine learning due to its great success in many applications. Since the neural networks' training usually involves a highly nonconvex optimization problem, it is difficult to design optimization algorithms with perfect convergence guarantees to derive a neural network estimator of high quality. In this article, we borrow the well-known random sketching strategy from kernel methods to transform the training of shallow rectified linear unit (ReLU) nets into a linear least-squares problem. Using the localized approximation property of shallow ReLU nets and a recently developed dimensionality-leveraging scheme, we succeed in equipping shallow ReLU nets with a specific random sketching scheme. The efficiency of the suggested random sketching strategy is guaranteed by theoretical analysis and also verified via a series of numerical experiments. Theoretically, we show that the proposed random sketching is almost optimal in terms of both approximation capability and learning performance. This implies that random sketching does not degenerate the performance of shallow ReLU nets. Numerically, we show that random sketching can significantly reduce the computational burden of numerous backpropagation (BP) algorithms while maintaining their learning performance. Di Wang 0008, Jinshan Zeng, Shaobo Lin |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion PathsabstractOver-parameterization is ubiquitous nowadays in training neural networks to benefit both optimization in seeking global optima and generalization in reducing prediction error. However, compressive networks are desired in many real world applications and direct training of small networks may be trapped in local optima. In this paper, instead of pruning or distilling over-parameterized models to compressive ones, we propose a new approach based on differential inclusions of inverse scale spaces. Specifically, it generates a family of models from simple to complex ones that couples a pair of parameters to simultaneously train over-parameterized deep models and structural sparsity on weights of fully connected and convolutional layers. Such a differential inclusion scheme has a simple discretization, proposed as Deep structurally splitting Linearized Bregman Iteration (DessiLBI), whose global convergence analysis in deep learning is established that from any initializations, algorithmic iterations converge to a critical point of empirical risks. Experimental evidence shows that DessiLBI achieve comparable and even better performance than the competitive optimizers in exploring the structural sparsity of several widely used backbones on the benchmark datasets. Remarkably, with early stopping, DessiLBI unveils “winning tickets” in early epochs: the effective sparse structure with comparable test accuracy to fully trained over-parameterized models. Yanwei Fu 0001, Chen Liu 0030, Xinwei Sun 0001, Jinshan Zeng, Yuan Yao 0011 |
ICML | 5 |
| 2019 | Image Smoothing Via Gradient Sparsity and Surface Area MinimizationabstractImage smoothing is a very important topic in image processing. Among these image smoothing methods, the L0gradient minimization method is one of the most popular ones. However, the L0gradient minimization method suffers from the staircasing effect and over-sharpening issue, which highly degrade the quality of the smoothed image. To overcome these issues, we use not only the L0gradient term for finding edges, but also a surface area based term for the purpose of smoothing the inside of each region. An alternating minimization algorithm is suggested to efficiently solve the proposed model, where each subproblem has a closed-form solution. Leveraging the introduced surface area term, the proposed method can effectively alleviate the staircasing effect and the over-sharpening issue. The superiority of our method over the state-of-the-art methods is demonstrated by a series of experiments. Jun Liu 0012, Ming Yan 0006, Jinshan Zeng, Tieyong Zeng |
ICIP | 3 |
| 2019 | Global Convergence of Block Coordinate Descent in Deep LearningabstractDeep learning has aroused extensive attention due to its great empirical success. The efficiency of the block coordinate descent (BCD) methods has been recently demonstrated in deep neural network (DNN) training. However, theoretical studies on their convergence properties are limited due to the highly nonconvex nature of DNN training. In this paper, we aim at providing a general methodology for provable convergence guarantees for this type of methods. In particular, for most of the commonly used DNN training models involving both two- and three-splitting schemes, we establish the global convergence to a critical point at a rate of ${\cal O}(1/k)$, where $k$ is the number of iterations. The results extend to general loss functions which have Lipschitz continuous gradients and deep residual networks (ResNets). Our key development adds several new elements to the Kurdyka-Lojasiewicz inequality framework that enables us to carry out the global convergence analysis of BCD in the general scenario of deep learning. Jinshan Zeng, Tsz Kit Lau, Shaobo Lin, Yuan Yao 0011 |
ICML | 1 |
| 2019 | Fast Learning With Polynomial KernelsabstractThis paper proposes a new learning system of low computational cost, called fast polynomial kernel learning (FPL), based on regularized least squares with polynomial kernel and subsampling. The almost optimal learning rate as well as the feasibility verifications including the subsampling mechanism and solvability of FPL are provided in the framework of learning theory. Our theoretical assertions are verified by numerous toy simulations and real data applications. The studies in this paper show that FPL can reduce the computational burden of kernel methods without sacrificing its generalization ability very much. Shaobo Lin, Jinshan Zeng |
IEEE Trans. Cybern. | 2 |
| 2019 | Constructive Neural Network LearningabstractIn this paper, we aim at developing scalable neural network-type learning systems. Motivated by the idea of constructive neural networks in approximation theory, we focus on constructing rather than training feed-forward neural networks (FNNs) for learning, and propose a novel FNNs learning system called the constructive FNN (CFN). Theoretically, we prove that the proposed method not only overcomes the classical saturation problem for constructive FNN approximation, but also reaches the optimal learning rate when the regression function is smooth, while the state-of-the-art learning rates established for traditional FNNs are only near optimal (up to a logarithmic factor). A series of numerical simulations are provided to show the efficiency and feasibility of CFN. Shaobo Lin, Jinshan Zeng, Xiaoqin Zhang 0002 |
IEEE Trans. Cybern. | 2 |
| 2018 | Stochastic Non-Convex Ordinal Embedding With Stabilized Barzilai-Borwein Step SizeabstractLearning representation from relative similarity comparisons, often called ordinal embedding, gains rising attention in recent years. Most of the existing methods are batch methods designed mainly based on the convex optimization, say, the projected gradient descent method. However, they are generally time-consuming due to that the singular value decomposition (SVD) is commonly adopted during the update, especially when the data size is very large. To overcome this challenge, we propose a stochastic algorithm called SVRG-SBB, which has the following features: (a) SVD-free via dropping convexity, with good scalability by the use of stochastic algorithm, i.e., stochastic variance reduced gradient (SVRG), and (b) adaptive step size choice via introducing a new stabilized Barzilai-Borwein (SBB) method as the original version for convex problems might fail for the considered stochastic non-convex optimization problem. Moreover, we show that the proposed algorithm converges to a stationary point at a rate O(1/T) in our setting, where T is the number of total iterations. Numerous simulations and real-world data experiments are conducted to show the effectiveness of the proposed algorithm via comparing with the state-of-the-art methods, particularly, much lower computational cost with good prediction performance. Ke Ma 0001, Jinshan Zeng, Jiechao Xiong, Qianqian Xu 0001, Xiaochun Cao, Wei Liu 0005, Yuan Yao 0011 |
AAAI | 2 |
| 2018 | Finding Global Optima in Nonconvex Stochastic Semidefinite Optimization with Variance ReductionabstractThere is a recent surge of interest in nonconvex reformulations via low-rank factorization for stochastic convex semidefinite optimization problem in the purpose of efficiency and scalability. Compared with the original convex formulations, the nonconvex ones typically involve much fewer variables, allowing them to scale to scenarios with millions of variables. However, it opens a new challenge that under what conditions the nonconvex stochastic algorithms may find the global optima effectively despite their empirical success in applications. In this paper, we provide an answer that a stochastic gradient descent method with variance reduction, can be adapted to solve the nonconvex reformulation of the original convex problem, with a global linear convergence, i.e., converging to a global optimum exponentially fast, at a proper initial choice in the restricted strongly convex case. Experimental studies on both simulation and real-world applications on ordinal embedding are provided to show the effectiveness of the proposed algorithms. Jinshan Zeng, Ke Ma 0001, Yuan Yao 0011 |
AISTATS | 1 |
| 2018 | Greedy Criterion in Orthogonal Greedy LearningabstractOrthogonal greedy learning (OGL) is a stepwise learning scheme that starts with selecting a new atom from a specified dictionary via the steepest gradient descent (SGD) and then builds the estimator through orthogonal projection. In this paper, we found that SGD is not the unique greedy criterion and introduced a new greedy criterion, called as " -greedy threshold" for learning. Based on this new greedy criterion, we derived a straightforward termination rule for OGL. Our theoretical study shows that the new learning scheme can achieve the existing (almost) optimal learning rate of OGL. Numerical experiments are also provided to support that this new scheme can achieve almost optimal generalization performance while requiring less computation than OGL. Lin Xu 0001, Shaobo Lin, Jinshan Zeng, Yi Fang 0006, Zongben Xu |
IEEE Trans. Cybern. | 3 |
| 2017 | Learning Rates for Classification with Gaussian KernelsabstractThis letter aims at refined error analysis for binary classification using support vector machine (SVM) with gaussian kernel and convex loss. Our first result shows that for some loss functions, such as the truncated quadratic loss and quadratic loss, SVM with gaussian kernel can reach the almost optimal learning rate provided the regression function is smooth. Our second result shows that for a large number of loss functions, under some Tsybakov noise assumption, if the regression function is infinitely smooth, then SVM with gaussian kernel can achieve the learning rate of order [Formula: see text], where [Formula: see text] is the number of samples. Shaobo Lin, Jinshan Zeng, Xiangyu Chang |
Neural Comput. | 2 |
| 2015 | Jackson-type inequalities for spherical neural networks with doubling weights
Shaobo Lin, Jinshan Zeng, Lin Xu 0001, Zongben Xu |
Neural Networks | 2 |
| 2015 | Error Estimate for Spherical Neural Networks Interpolation
Shaobo Lin, Jinshan Zeng, Zongben Xu |
Neural Process. Lett. | 2 |
| 2014 | Learning Rates of lq Coefficient Regularization Learning with Gaussian KernelabstractRegularization is a well-recognized powerful strategy to improve the performance of a learning machine and l(q) regularization schemes with 0 < q < ∞ are central in use. It is known that different q leads to different properties of the deduced estimators, say, l(2) regularization leads to a smooth estimator, while l(1) regularization leads to a sparse estimator. Then how the generalization capability of l(q) regularization learning varies with q is worthy of investigation. In this letter, we study this problem in the framework of statistical learning theory. Our main results show that implementing l(q) coefficient regularization schemes in the sample-dependent hypothesis space associated with a gaussian kernel can attain the same almost optimal learning rates for all 0 < q < ∞. That is, the upper and lower bounds of learning rates for l(q) regularization learning are asymptotically identical for all 0 < q < ∞. Our finding tentatively reveals that in some modeling contexts, the choice of q might not have a strong impact on the generalization capability. From this perspective, q can be arbitrarily specified, or specified merely by other nongeneralization criteria like smoothness, computational complexity or sparsity. Shaobo Lin, Jinshan Zeng, Jian Fang 0001, Zongben Xu |
Neural Comput. | 2 |
| 2014 | Sparse solution of underdetermined linear equations via adaptively iterative thresholding
Jinshan Zeng, Shaobo Lin, Zongben Xu |
Signal Process. | 1 |
| 2013 | Accelerated L1/2 regularization based SAR imaging via BCR and reduced Newton skills
Jinshan Zeng, Zongben Xu, Bingchen Zhang, Wen Hong, Yirong Wu |
Signal Process. | 1 |
| 2012 | Efficient DPCA SAR imaging with fast iterative spectrum reconstruction method
Jian Fang 0001, Jinshan Zeng, Zongben Xu |
Sci. China Inf. Sci. | 2 |
| 2012 | Sparse SAR imaging based on L 1/2 regularization
Jinshan Zeng, Jian Fang 0001, Zongben Xu |
Sci. China Inf. Sci. | 1 |
| 2011 | SAR imaging from compressed measurements based on L1/2 regularizationabstractIn this paper, a novel synthetic aperture radar (SAR) imaging method based on L1/2regularization is proposed. Our method implements SAR imaging from compressed measurements with high resolution, enhanced features, reduced sidelobes and suppressed artifacts. Real SAR data experiments are implemented to demonstrate the outperformance of our method. The experiment results demonstrate that our method needs far below the traditional Nyquist rate to guarantee successful imaging. Compared to the prevalent L1regularization-based methods, there is a significant reduction of the sampling rate for SAR imaging. The sampling rate used by our method is about half of the L1regularization-based methods in the real SAR data experiments. Jinshan Zeng, Zongben Xu, Bingchen Zhang, Wen Hong, Yirong Wu |
IGARSS | 1 |