Lin Zuo

dblp:75/5553 · DBLP profile ↗
← Back
53ranked-venue papers
10as first author
39since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 11 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language Models
abstract
Dropout is a widely used regularization technique which improves the generalization ability of a model by randomly dropping neurons. In light of this, we propose Dropout Prompt Learning, which aims for applying dropout to improve the robustness of the vision-language models. Different from the vanilla dropout, we apply dropout on the tokens of the textual and visual branches, where we evaluate the token significance considering both intra-modal context and inter-modal alignment, enabling flexible dropout probabilities for each token. Moreover, to maintain semantic alignment for general knowledge transfer while encouraging the diverse representations that dropout introduces, we further propose residual entropy regularization. Experiments on 11 benchmarks show our method's effectiveness in challenging scenarios like low-shot learning, long-tail classification, and out-of-distribution generalization. Notably, our method surpasses regularization-based methods including KgCoOp by 5.10% and PromptSRC by 2.13% in performance on base-to-novel generalization.
Lin Zuo, Mengmeng Jing, Kunbin He
AAAI2
2026 Instruction-Guided Cross-Modal Clustering for Training-Free Visual Token Pruning in Vision-Language Models
abstract
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data such as images and text. However, the number of visual tokens in these models often far exceeds that of textual tokens, resulting in substantial redundancy and high inference costs. Existing pruning methods primarily rely on either unimodal information or cross-modal attention mechanisms. The former often overlooks the semantic alignment between instructions and visual representations in the multimodal space, while the latter is prone to attention drift and dispersion, leading to significant performance degradation under high pruning ratios. All the above issues stem from the lack of effective textual guidance during the pruning process. To identify effective informational cues for guiding pruning, we conduct an in-depth analysis of the interaction between language instructions and visual features based on the cross-modal information bottleneck attribution (CIBA) method, revealing the presence of noun anchors. Based on this analysis, we propose the Instruction-Guided Cross-Modal Clustering Token Pruning (ICCTP) method, a plug-and-play, training-free pruning paradigm. Specifically, ICCTP first leverages global attention to retain a small set of visual tokens that preserve global context. It then extracts nouns from the instruction as clustering centers to perform cross-modal clustering over the remaining visual tokens. To balance semantic diversity and global relevance while reducing intra-cluster redundancy, we design an importance scoring mechanism. Finally, visual tokens within each cluster are pruned according to a specified pruning ratio. We evaluate ICCTP on multiple VLM architectures, including LLaVA-1.5-7B, LLaVA-1.5-13B, and LLaVA-NeXT-7B. Experimental results show that ICCTP maintains strong performance across various pruning rates without requiring retraining. Notably, even under an extreme setting where 94.4% of visual tokens are removed, ICCTP retains 90.02% of the original accuracy while reducing TFLOPs by 82.36%.
Yunqian Yu, Yunya Zhang, Tonglan Xie, Mengmeng Jing, Lin Zuo
AAAI6
2026 SQL-CoT: Teaching Compact LLMs Structured Reasoning for Text-to-SQL
Kunbin He, Zhikun Zheng, Yuguo Hu, Lin Zuo
ICIC (24)5
2026 Flexible CLIP: Decoupled Vision and Auxiliary Semantics for Flexible Object Recognition
Yunya Zhang, Tonglan Xie, Jiafu Yan, Kunshan Yang, Lin Zuo
ICIC (18)6
2026 SARAG: Sufficiency-Aware Retrieval Augmented Generation
Zhikun Zheng, Wenwei Luo, Lin Zuo
ICIC (22)5
2026 Data poisoning-based backdoor attacks against supervised learning rules of Spiking Neural Networks
Lingxin Jin, Wei Jiang 0016, Jinyu Zhan, Meiyu Lin, Letian Chen, Boran Quan, Lin Zuo, Xingzhi Zhou 0001, Maregu Assefa, Naoufel Werghi
J. Syst. Archit.7
2026 Learning spatio-temporal consistency in spiking neural networks by self-distillation
Lin Zuo, Yongqi Ding, Mengmeng Jing, Kunshan Yang, Hanpu Deng
Pattern Recognit.1
2026 MarkingVLM: Vision-Language Model for Few-Shot IC Marking Detection
abstract
Integrated circuit (IC) marking detection faces critical “Triple C” challenges: compact marking scales, complex multidirectional layouts, and costly annotation requirements. Traditional models struggle with data scarcity in dynamic manufacturing environments where fine-grained labeled data are expensive and time-consuming to obtain. This work presents MarkingVLM, a specialized vision-language model designed for IC marking detection with exceptional few-shot learning capabilities. MarkingVLM adapts cross-modal understanding from Contrastive Language-Image Pre-training (CLIP) through unique innovations: 1) dual-granularity detection combining patch-level semantic alignment with pixel-level visual decoding; 2) layout mirror module enabling efficient multidirectional text flow recognition through feature-level augmentation; and 3) enhanced prompt engineering with learnable contexts for effective domain adaptation. Extensive experiments across two IC marking datasets with distinct characteristics substantiate superior performance: 94.2% precision and 96.5% recall on Dataset-1, and 89.6% precision with 92.6% recall on the more challenging Dataset-2. Most significantly, MarkingVLM achieves 92.7% precision with 32 training samples and maintains over 80% recall with merely 8 samples, demonstrating notable data efficiency improvement over conventional methods. Results establish a new paradigm for industrial text detection by bridging open-domain vision-language knowledge with specialized manufacturing requirements.
Zhongshu Chen, Zhenghua Chen, Lin Zuo, Yu Liu 0006
IEEE Trans. Ind. Informatics3
2025 Prelude echoes Finale: Video Domain Adaptation with Fine-grained Temporal Consistency
abstract
The prelude and finale of one video usually share similar themes and content, with the finale often echoing or deepening the concepts introduced in the prelude. This temporal correlation enhances the understanding of video content. Existing video domain adaptation methods, however, ignore this fine-grained temporal structure in the clip level, focusing instead on coarser temporal consistency at the video level. To make the prelude echo the finale, we propose to optimize fine-grained temporal consistency for video domain adaptation. Specifically, we first extract the prelude (the first n clips) and the finale (the last n clips) from the same video. Then, we maximize the feature similarity of the prelude and the finale, which enhances the connections between clips and achieves the temporal consistency of each video across source and target domains. On the other hand, we generate the cross-domain fusion features and optimize their discriminability to make the source domain features echo the target. As a result, the source and target domains are aligned. Extensive experiments on video domain adaptation benchmarks demonstrate the effectiveness of our method.
Mengmeng Jing, Xianlong Tian, Yuguo Hu, Lin Zuo
ICASSP5
2025 Learning Latent Representations with Codebook Priors for Single-Image Flare Removal
Shimin Luo, Yunya Zhang, Hanpu Deng, Mengmeng Jing, Lin Zuo
ICIC (3)6
2025 Multi-bit Mechanism: Towards Ultra-Low Time Steps for Spiking Neural Networks
Yongjun Xiao, Pei He, Hanpu Deng, Tonglan Xie, Mengmeng Jing, Lin Zuo
ICIC (21)6
2025 Rethinking Spiking Neural Networks from an Ensemble Learning Perspective
abstract
Spiking neural networks (SNNs) exhibit superior energy efficiency but suffer from limited performance. In this paper, we consider SNNs as ensembles of temporal subnetworks that share architectures and weights, and highlight a crucial issue that affects their performance: excessive differences in initial states (neuronal membrane potentials) across timesteps lead to unstable subnetwork outputs, resulting in degraded performance. To mitigate this, we promote the consistency of the initial membrane potential distribution and output through membrane potential smoothing and temporally adjacent subnetwork guidance, respectively, to improve overall stability and performance. Moreover, membrane potential smoothing facilitates forward propagation of information and backward propagation of gradients, mitigating the notorious temporal gradient vanishing problem. Our method requires only minimal modification of the spiking neurons without adapting the network structure, making our method generalizable and showing consistent performance gains in 1D speech, 2D object, and 3D point cloud recognition tasks. In particular, on the challenging CIFAR10-DVS dataset, we achieved 83.20\% accuracy with only four timesteps. This provides valuable insights into unleashing the potential of SNNs.
Yongqi Ding, Lin Zuo, Mengmeng Jing, Pei He, Hanpu Deng
ICLR2
2025 Diffusion-Driven Source Consistency for Gradual Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to directly transfer models from source domain to target domain. When the domain shifts are large, UDA performs poorly. Gradual Domain Adaptation (GDA) alleviates this problem in a mild way by gradually adapting from source to target domain using multiple intermediate domains. In this paper, we propose Source Adaptive Diffusion Model (SADiff) to leverage intermediate domains to get more linear and effective performance across samples. SADiff employs diffusion model to simulate the feature distribution shift between domains, ensuring smoothness and continuity in latent space through regularization constraints on cross-domain mixup features. Additionally, we introduce an energy-based labeling strategy to improve label consistency. Finally, we gradually improved the adaptive ability of the model, achieving smooth and effective domain adaptation. Extensive experiments on four GDA benchmarks verified the effectiveness of the proposed method.
Wenwei Luo, Yuguo Hu, Jiafu Yan, Mengmeng Jing, Lin Zuo
ICME5
2025 Bayesian-Inspired Cross-Spectral Fusion Network for Robust Depth Estimation
abstract
Depth estimation is a critical task in multimedia, with widespread applications in autonomous driving, robotics, and 3D reconstruction. Traditional methods relying on single-spectral data suffer from low-light and nighttime conditions. In view of this, we propose BCSFDepth, a bayesian-inspired cross-spectral fusion network fuses infrared and visible images to enhance depth estimation. Our method incorporates a frequency-domain interaction enhancement module and a Bayesian uncertainty-weighted fusion mechanism to effectively integrate complementary information across modalities. The proposed framework dynamically adjusts modality weights, mitigating the impact of low-quality features. We collected and calibrated a real-world cross-spectral dataset, IVLD, to evaluate our method under diverse weather conditions. Experimental results demonstrate substantial improvements in depth estimation accuracy, particularly in noisy and complex environments.
Jiafu Yan, Wenwei Luo, Yuguo Hu, Changhua Zhang, Mengmeng Jing, Lin Zuo
ICME6
2025 A Semantic-Enhanced Heterogeneous Graph Learning Method for Flexible Objects Recognition
abstract
Flexible objects recognition remains a significant challenge due to its inherently diverse shapes and sizes, translucent attributes, and subtle inter-class differences. Graph-based models, such as graph convolution networks and graph vision models, are promising in flexible objects recognition due to their ability of capturing variable relations within the flexible objects. These methods, however, often focus on global visual relationships or fail to align semantic and visual information. To alleviate these limitations, we propose a semantic-enhanced heterogeneous graph learning method. First, an adaptive scanning module is employed to extract discriminative semantic context, facilitating the matching of flexible objects with varying shapes and sizes while aligning semantic and visual nodes to enhance cross-modal feature correlation. Second, a heterogeneous graph generation module aggregates global visual and local semantic node features, improving the recognition of flexible objects. Additionally, We introduce the FSCW, a large-scale flexible dataset curated from existing sources. We validate our method through extensive experiments on flexible datasets (FDA and FSCW), and challenge benchmarks (CIFAR-100 and ImageNet-Hard), demonstrating competitive performance.
Kunshan Yang, Wenwei Luo, Yuguo Hu, Jiafu Yan, Mengmeng Jing, Lin Zuo
ICME6
2025 Toward End-to-End Bearing Fault Diagnosis for Industrial Scenarios with Spiking Neural Networks
abstract
This paper explores the application of spiking neural networks (SNNs), known for their low-power binary spikes, to bearing fault diagnosis, bridging the gap between high-performance AI algorithms and real-world industrial scenarios.In particular, we identify two key limitations of existing SNN fault diagnosis methods: inadequate encoding capacity that necessitates cumbersome data preprocessing, and non-spike-oriented architectures that constrain the performance of SNNs.To alleviate these problems, we propose a Multi-scale Residual Attention SNN (MRA-SNN) to simultaneously improve the efficiency, performance, and robustness of SNN methods.By incorporating a lightweight attention mechanism, we have designed a multi-scale attention encoding module to extract multiscale fault features from vibration signals and encode them as spatio-temporal spikes, eliminating the need for complicated preprocessing.Then, the spike residual attention block extracts high-dimensional fault features and enhances the expressiveness of sparse spikes with the attention mechanism for end-to-end diagnosis.In addition, the performance and robustness of MRA-SNN is further enhanced by introducing the lightweight attention mechanism within the spiking neurons to simulate the biological dendritic filtering effect.Extensive experiments on MFPT, JNU, Bearing, and Gearbox benchmark datasets demonstrate that MRA-SNN significantly outperforms existing methods in terms of accuracy, energy
Lin Zuo, Yongqi Ding, Mengmeng Jing, Kunshan Yang, Yunqian Yu
KDD (2)1
2025 Chain-of-Thought Guided Semantic Debiasing for Low-Shot Vision-Language Tasks
Kunbin He, Zhikun Zheng, Mengmeng Jing, Lin Zuo
ACM Multimedia5
2025 PA-HOI: A Physics-Aware Human and Object Interaction Dataset
Ruiyan Wang, Lin Zuo, Zonghao Lin, Qiang Wang 0061, Zhengxue Cheng, Rong Xie 0004, Jun Ling, Li Song 0001
ACM Multimedia2
2025 Bridging Inter-Class Ambiguity and Spatial Variability in Flexible Object Recognition via Graph Distillation
abstract
Flexible object recognition remains challenging in multimedia scenarios due to inherently diverse shapes and sizes, and subtle inter-class differences. Graph-based vision models show promise in flexible objects recognition by capturing variable relationships. However, they suffer from two problems: (1) inter-class ambiguity hinders model discrimination and (2) frequent scale changes degrade model generalization. To address these limitations, we propose a unified graph distillation framework that enhances inter-class discrimination and spatial generalization while maintaining computational efficiency. For inter-class ambiguity problem, we introduce a virtual prototype module that dynamically generates learnable class prototypes via clustering intermediate features. These prototypes are incorporated into the distillation loss to sharpen decision boundaries. A global-local distillation mechanism further capture both image-level global semantics and patch-level local details, enhancing inter-class discrimination. For frequent scale changes problem, we design a patch-aware distillation strategy that transfers knowledge across multiple patch scales, strengthening the student model's spatial generalization to match various shapes and sizes of flexible objects, thus alleviate generalization degradation. Extensive experiments on flexible-object datasets (FDA, FSCW, CCSN) and challenging benchmarks (CIFAR-100, Mini-ImageNet) confirm effectiveness and efficiency of our method.
Lin Zuo, Kunshan Yang, Mengmeng Jing, Xiangxu Zhao, Jiaqiao Chen
ACM Multimedia1
2025 Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-Distillers
abstract
Brain-inspired spiking neural networks (SNNs) promise to be a low-power alternative to computationally intensive artificial neural networks (ANNs), although performance gaps persist. Recent studies have improved the performance of SNNs through knowledge distillation, but rely on large teacher models or introduce additional training overhead. In this paper, we show that SNNs can be naturally deconstructed into multiple submodels for efficient self-distillation. We treat each timestep instance of the SNN as a submodel and evaluate its output confidence, thus efficiently identifying the strong and the weak. Based on this strong and weak relationship, we propose two efficient self-distillation schemes: (1) Strong2Weak: During training, the stronger "teacher" guides the weaker "student", effectively improving overall performance. (2) Weak2Strong: The weak serve as the "teacher", distilling the strong in reverse with underlying dark knowledge, again yielding significant performance gains. For both distillation schemes, we offer flexible implementations such as ensemble, simultaneous, and cascade distillation. Experiments show that our method effectively improves the discriminability and overall performance of the SNN, while its adversarial robustness is also enhanced, benefiting from the stability brought by self-distillation. This ingeniously exploits the temporal properties of SNNs and provides insight into how to efficiently train high-performance SNNs.
Yongqi Ding, Lin Zuo, Mengmeng Jing, Kunshan Yang, Pei He, Tonglan Xie
NeurIPS2
2025 Visual residual aggregation network for visual-language prompt tuning
Yunqian Yu, Xianlong Tian, Mengmeng Jing, Lin Zuo
Appl. Intell.6
2025 Adaptive recognition of flexible objects via hierarchical deformable graph network
Kunshan Yang, Lin Zuo, Mengmeng Jing, Yongqi Ding, Xianlong Tian
Inf. Sci.2
2025 DviT: Debiased variational inference for multi-modal mutual prompt tuning
Zhongshu Chen, Zhikun Zheng, Lin Zuo
Knowl. Based Syst.5
2025 A Computation-Quantized Training Framework to Generate Accuracy Lossless QNNs for One-Shot Deployment in Embedded Systems
abstract
Quantized Neural Networks (QNNs) have received increasing attention, since they can enrich intelligent applications deployed on embedded devices with limited resources, such as mobile devices and AIoT systems. Unfortunately, the numerical and computational discrepancies between training systems (i.e., servers) and deployment systems (e.g., embedded ends) may lead to large accuracy drop for QNNs in real deployments. We propose a Computation-Quantized Training Framework (CQTF), which simulates deployment-time fixed-point computation during training to enable one-shot, lossless deployment. The training procedure of CQTF is built upon a well-formulated quantization-specific numerical representation that quantifies both numerical and computational discrepancies between training and deployment. Leveraging this representation, forward propagation executes all computations in quantization mode to simulate deployment-time inference, while backward propagation identifies and mitigates gradient vanishing through an efficient floating-point gradient update scheme. Benchmark-based experiments demonstrate the efficiency of our approach, which can achieve no accuracy loss from training to deployment. Compared with existing five frameworks, the deployed accuracy of CQTF can be improved by up to 18.41%.
Xingzhi Zhou 0001, Wei Jiang 0016, Jinyu Zhan, Lingxin Jin, Lin Zuo
IEEE Trans. Computers5
2025 Flexible ViG: Learning the Self-Saliency for Flexible Object Recognition
abstract
Existing computer vision methods mainly focus on the recognition of rigid objects, whereas the recognition of flexible objects remains unexplored. Recognizing flexible objects poses significant challenges due to their inherently diverse shapes and sizes, translucent attributes, ambiguous boundaries, and subtle inter-class differences. In this paper, we claim that these problems primarily arise from the lack of object saliency. To this end, we propose the Flexible Vision Graph Neural Network (FViG) to optimize the self-saliency and thereby improve the discrimination of the representations for flexible objects. Specifically, on one hand, we propose to maximize the channel-aware saliency by extracting the weight of neighboring graph nodes, which is employed to identify flexible objects with minimal inter-class differences. On the other hand, we maximize the spatial-aware saliency based on clustering to aggregate neighborhood information for the centroid graph nodes. This introduces local context information and enables extracting of consistent representation, effectively adapting to the shape and size variations in flexible objects. To verify the performance of flexible objects recognition thoroughly, for the first time we propose the Flexible Dataset (FDA), which consists of various images of flexible objects collected from real-world scenarios or online. Extensive experiments evaluated on our FDA, FireNet, CIFAR-100 and ImageNet-Hard datasets demonstrate the effectiveness of our method on enhancing the discrimination of flexible objects.
Kunshan Yang, Lin Zuo, Mengmeng Jing, Xianlong Tian, Kunbin He, Yongqi Ding
IEEE Trans. Circuits Syst. Video Technol.2
2025 ICMarkingNet: An Ultrafast and Streamlined Deep Model for IC Marking Inspection
abstract
This study presents ICMarkingNet, an end-to-end model for integrated circuit (IC) the marking inspection task. The model pinpoints markings against an IC image utilizing a saliency-guided regime under weakly supervised learning. Through the introduction of a novel direction representation alongside a transform-based rotation method, the model achieves improved accuracy in recognizing marking directions. Furthermore, the incorporation of a newly introduced sampling method, namely LinkSampling, enables the model to extract character features with high consistency, and empowers the model to excel in word-level marking recognition tasks. Notably, ICMarkingNet is uniquely engineered to function within a compact and streamlined pipeline, facilitating execution entirely on graphics process units. Experiments on a real-line IC marking dataset exhibit an f1-score of 96.88% and the inspection speed surpassing 100 samples per second, validating the superior performance and efficiency of the proposed model over both general end-to-end text recognition models and existing IC marking inspection frameworks.
Zhongshu Chen, Lin Zuo, Yu Liu 0006
IEEE Trans. Ind. Informatics3
2024 Shrinking Your TimeStep: Towards Low-Latency Neuromorphic Object Recognition with Spiking Neural Networks
abstract
Neuromorphic object recognition with spiking neural networks (SNNs) is the cornerstone of low-power neuromorphic computing. However, existing SNNs suffer from significant latency, utilizing 10 to 40 timesteps or more, to recognize neuromorphic objects. At low latencies, the performance of existing SNNs is drastically degraded. In this work, we propose the Shrinking SNN (SSNN) to achieve low-latency neuromorphic object recognition without reducing performance. Concretely, we alleviate the temporal redundancy in SNNs by dividing SNNs into multiple stages with progressively shrinking timesteps, which significantly reduces the inference latency. During timestep shrinkage, the temporal transformer smoothly transforms the temporal scale and preserves the information maximally. Moreover, we add multiple early classifiers to the SNN during training to mitigate the mismatch between the surrogate gradient and the true gradient, as well as the gradient vanishing/exploding, thus eliminating the performance degradation at low latency. Extensive experiments on neuromorphic datasets, CIFAR10-DVS, N-Caltech101, and DVS-Gesture have revealed that SSNN is able to improve the baseline accuracy by 6.55% ~ 21.41%. With only 5 average timesteps and without any data augmentation, SSNN is able to achieve an accuracy of 73.63% on CIFAR10-DVS. This work presents a heterogeneous temporal scale SNN and provides valuable insights into the development of high-performance, low-latency SNNs.
Yongqi Ding, Lin Zuo, Mengmeng Jing, Pei He, Yongjun Xiao
AAAI2
2024 Graph Neural Network-Based Structured Scene Graph Generation for Efficient Wildfire Detection
Yanning Ye, Shimin Luo, Mengmeng Jing, Yongqi Ding, Kunbin He, Lin Zuo
ICIC (3)6
2024 Weakly Supervised End-to-End Learning for Inspection on Multidirectional Integrated Circuit Markings in Surface Mount Technology
abstract
Integrated circuit (IC) marking inspection is a crucial task to ensure product quality in electronics manufacturing. Due to the diversity of marking appearance, high environmental complexity, and massive annotation costs, it is, however, still a great challenge to accurately recognize IC markings in a real-time fashion at some production stages, such as surface mount technology (SMT). In this article, an end-to-end deep learning model with three branches is put forth for IC marking inspection. The saliency activation branch provides powerful shared feature representation, and by incorporating with the weakly supervised mechanism, it can generate precise character localization information with coarse-grained annotation. The direction recognition and character recognition branches utilize shared saliency maps to sample word-level and character-level features, respectively, such that the network can properly recognize markings in different orientations, and especially perform well on the chips with multidirectional markings. The proposed character box refinement method allows the network to adapt to tiny size and tight-layout IC markings, and a new loss function called ED-Loss is designed for error estimation between the unaligned sequences. Experiments on a real SMT chip dataset with highly diverse IC images show that the model reaches a recall rate of 96.34%, with an inspection speed close to 30 fps. The comparative experiments with the state-of-the-art models demonstrate that our model has superior performances in terms of accuracy, efficiency, and adaptability.
Zhongshu Chen, Lin Zuo, Changhua Zhang, Yu Liu 0006
IEEE Trans. Ind. Informatics2
2024 An End-to-End Bilateral Network for Multidefect Detection of Solid Propellants
abstract
Defect detection tasks of solid propellants (SPs), involving size, shape, and surface defects, are essential for ensuring the quality of many industrial products. Developing separate models for three tasks, however, is complicated and inefficient due to the redundant deployment. Multitask learning (MTL), with its potential for knowledge sharing, may greatly reduce the space and power consumption, but still faces the challenges of destructive interference and empirical tradeoffs between tasks. To this end, a novel end-to-end network for multidefect detection of SPs is put forth: 1) a new setting for MTL without any empirical tradeoffs is introduced, in which the knowledge is shared while the models are not visible to each other among diverse tasks; 2) in this setting, a bilateral feature extractor is constructed to extract both low- and high-level features, and a feature fusion module is further exploited to encourage each task to adaptively learn the task-specific knowledge; 3) an end-to-end training manner with a dynamic balance strategy and a gradient stop-flow strategy is designed to ensure that different tasks can benefit from, but do not interfere with, each other; 4) the introduction of semantic knowledge from the size detection branch enables the surface detection branch to learn semantic features beyond only pixel-to-pixel mapping. A smoothness construction loss is further designed to boost the performance of the surface detection task. Experimental results on an image dataset from a real-world manufacturing line show that the setting for MTL has the superiority in terms of the model size, inference speed, and detection accuracy.
Zhongshu Chen, Lin Zuo, Tangfan Xiahou, Yu Liu 0006
IEEE Trans. Ind. Informatics4
2023 Multi-gate Mixture-of-Contrastive-Experts with Graph-based Gating Mechanism for TV Recommendation
abstract
With the rapid development of smart TV, TV recommendation is attracting more and more users. TV users usually distribute in multiple regions with different cultures and hence have diverse TV program preferences. From the perspective of engineering practice and performance improvement, it's very essential to model users from multiple regions with one single model. In previous work, Multi-gate Mixture-of-Expert (MMoE) has been widely adopted in multi-task and multi-domain recommendation scenarios. In practice, however, we first observe the embeddings generated by experts tend to be homogeneous which may result in high semantic similarities among embeddings that reduce the capability of Multi-gate Mixture-of-Expert (MMoE) model. Secondly, we also find there are lots of commonalities and differences between multiple regions regarding user preferences. Therefore, it's meaningful to model the complicated relationships between regions. In this paper, we first introduce contrastive learning to overcome the expert representation degeneration problem. The embeddings of two augmented samples generated by the same experts are pushed closer to enhance the alignment, and the embeddings of the same samples generated by different experts are pushed away in vector space to improve uniformity. Then we propose a Graph-based Gating Mechanism to empower typical Multi-gate Mixture-of-Experts. Graph-based MMoE is able to recognize the commonalities and differences among multiple regions by introducing a Graph Neural Network (GNN) with region similarity prior. We name our model Multi-gate Mixture-of-Contrastive-Experts model with Graph-based Gating Mechanism (MMoCEG). Extensive offline experiments and online A/B tests on a commercial TV service provider over 100 million users and 2.3 million items demonstrate the efficacy of MMoCEG compared to the existing models.
Cong Zhang 0016, Lin Zuo, Junlan Feng, Chao Deng 0002, Haitao Zeng, Yaohong Zhao
CIKM3
2023 AAKD-Net: Attention-Based Adversarial Knowledge Distillation Network for Image Classification
Fukang Zheng, Lin Zuo, Wenwei Luo, Yuguo Hu
ICONIP (7)2
2023 An improved probabilistic spiking neural network with enhanced discriminative ability
Yongqi Ding, Lin Zuo, Kunshan Yang, Zhongshu Chen, Tangfan Xiahou
Knowl. Based Syst.2
2022 Building Height Restoration Method of Remote Sensing Images based on Faster RCNN
abstract
To accurately obtain building height information from a single remote sensing image, we propose a height restoration method, which mainly is composed of two parts, building shadow rotation detection and building height calculation. The first part adds a skip connection structure and rotated branches based on Faster RCNN and achieves rotation shadow detection. The latter uses imaging date and geographic latitude to restore building height based on the geometric relationship between the building and its shadow. The experiment shows that the accuracy of height restoration is 95.04%. Compared with the state-of-the-art method, our method has the superiority of simple implementation, less data, fast speed, and high accuracy.
Xucan Chen, Lin Zuo
ICTAI3
2022 Energy-Based Domain Generalization for Face Anti-Spoofing
abstract
With various unforeseeable face presentation attacks (PA) springing up, face anti-spoofing (FAS) urgently needs to generalize to unseen scenarios. Research on generalizable FAS has lately attracted growing attention. Existing methods cast FAS as a vanilla binary classification problem and address it by a standard discriminative classifier p(y|x) under a domain generalization framework. However, discriminative models are unreliable for samples far away from the training distribution. In this paper, we resort to an energy-based model (EBM) to tackle FAS in a generative perspective. Our motivation is to model the joint density p(x,y), which allows to compute not only p(y|x) but also p(x). Due to the intractability of direct modeling, we use EBMs as an alternative to probabilistic estimation. With energy-based training, real faces are encouraged to get low free energy associated with the marginal probability p(x) of real faces, and all samples with high free energy are regarded as fake faces, thus rejecting any kind of PA out of the distribution of real faces. To learn to generalize to unseen domains, we generate diverse and novel populations in feature space under the guidance of energy model. Our model is updated in a meta-learning schema, where the original source samples are utilized for meta-training and the generated ones for meta-testing. We validate our method on four widely used FAS datasets. Comprehensive experimental results demonstrate the effectiveness of our method compared with state-of-the-arts.
Zhekai Du, Jingjing Li 0001, Lin Zuo, Lei Zhu 0002, Ke Lu 0001
ACM Multimedia3
2022 An Adaptive Deep Learning Framework for Fast Recognition of Integrated Circuit Markings
abstract
Fast recognition of integrated circuit (IC) markings is an essential but challenging task in electronic device manufacturing lines. This article develops an adaptive deep learning framework to facilitate the fast marking recognition of IC chips. The proposed framework contains four deep learning components, namely, chip segmentation, orientation correction, character extraction, and character recognition. The four components utilize different convolutional neural network structures to guarantee excellent adaptivity to a wide range of IC types and mitigate the influence of the low-quality chip images. In particular, the character extraction model is comprised of two improved label generation strategies and a proposed border correction method, so as to accommodate tiny scale chips and compactly printed markings. Experiments from the chip image dataset of a real laptop manufacturing line reached a recognition Precision of 91.73% and the Recall of 92.93%. The results demonstrate the superiority of the proposed framework to the state-of-the-art models and the effectiveness of handling a great diversity of chips with different scales, shapes, text fonts, marking colors, and layouts.
Zhongshu Chen, Changhua Zhang, Lin Zuo, Tangfan Xiahou, Yu Liu 0006
IEEE Trans. Ind. Informatics3
2021 CMRD-Net: An Improved Method for Underwater Image Enhancement
abstract
Underwater image enhancement is a challenging task due to the degradation of image quality in underwater complicated lighting conditions and scenes. In recent years, most methods improve the visual quality of underwater images by using deep Convolutional Neural Networks and Generative Adversarial Networks. However, the majority of existing methods do not consider that the attenuation degrees of R, G, B channels of the underwater image are different, leading to a sub-optimal performance. Based on this observation, we propose a Channel-wise Multi-scale Residual Dense Network called CMRD-Net, which learns the weights of different color channels instead of treating all the channels equally. More specifically, the Channel-wise Multi-scale Fusion Residual Attention Block (CMFRAB) is involved in the CMRD-Net to obtain a better ability of feature extraction and representation. Notably, we evaluate the effectiveness of our model by comparing it with recent state-of-the-art methods. Extensive experimental results show that our method can achieve a satisfactory performance on a popular public dataset.
Fengjie Xu, Changhua Zhang, Zhongshu Chen, Zhekai Du, Lin Zuo
MMAsia6
2021 Dynamic Charging Scheme Problem With Actor-Critic Reinforcement Learning
abstract
The energy problem is one of the most important challenges in the application of sensor networks. With the development of wireless charging technology and intelligent mobile charger (MC), the energy problem can be solved by the wireless charging strategy. In the practical application of wireless rechargeable sensor networks (WRSNs), the energy consumption rate of nodes is dynamically changed due to many uncertainties, such as the death and different transmission tasks of sensor nodes. However, existing works focus on on-demand schemes, which not fully consider real-time global charging scheduling. In this article, a novel dynamic charging scheme (DCS) in WRSN based on the actor-critic reinforcement learning (ACRL) algorithm is proposed. In the ACRL, we introduce gated recurrent units (GRUs) to capture the relationships of charging actions in time sequence. Using the actor network with one GRU layer, we can pick up an optimal or near-optimal sensor node from candidates as the next charging target more quickly and speed up the training of the model. Meanwhile, we take the tour length and the number of dead nodes as the reward signal. Actor and critic networks are updated by the error criterion function of R and V. Compared with current on-demand charging scheduling algorithms, extensive simulations show that the proposed ACRL algorithm surpasses heuristic algorithms, such as the Greedy, DP, nearest job next with preemption, and TSCA in the average lifetime and tour length, especially against the size and complexity increasing of WRSNs.
Nianbo Liu, Lin Zuo, Yong Feng 0004, Minghui Liu 0002, Hai-gang Gong, Ming Liu 0002
IEEE Internet Things J.3
2021 Challenging tough samples in unsupervised domain adaptation
Lin Zuo, Mengmeng Jing, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001, Yang Yang 0002
Pattern Recognit.1
2020 Mapping Electric Transmission Line Infrastructure from Aerial Imagery with Deep Learning
abstract
Access to electricity positively correlates with many beneficial socioeconomic outcomes in the developing world including improvements in education, health, and poverty. Efficient planning for electricity access requires information on the location of existing electric transmission and distribution infrastructure; however, the data on existing infrastructure is often unavailable or expensive. We propose a deep learning based method to automatically detect electric transmission infrastructure from aerial imagery and quantify those results with traditional object detection performance metrics. In addition, we explore two challenges to applying these techniques at scale: (1) how models trained on particular geographies generalize to other locations and (2) how the spatial resolution of imagery impacts infrastructure detection accuracy. Our approach results in object detection performance with an F1 score of 0.53 (0.47 precision and 0.60 recall). Using training data that includes more diverse geographies improves performance across the 4 geographies that we examined. Image resolution significantly impacts object detection performance and decreases precipitously as the image resolution decreases.
Ben Alexander, Wendell Cathcart, Atsushi Hu, Varun Nair, Lin Zuo, Jordan M. Malof, Leslie M. Collins, Kyle Bradbury
IGARSS6
2020 Mobile parking incentives for vehicular networks: a deep reinforcement learning approach
Nianbo Liu, Lin Zuo, Hai-gang Gong, Minghui Liu 0002, Ming Liu 0002
CCF Trans. Pervasive Comput. Interact.3
2020 A spiking neural network with probability information transmission
Lin Zuo, Yi Chen 0034, Lei Zhang 0005, Changle Chen
Neurocomputing1
2020 A robust approach to reading recognition of pointer meters based on improved mask-RCNN
Lin Zuo, Changhua Zhang, Zhehan Zhang
Neurocomputing1
2020 Leveraging unpaired out-of-domain data for image captioning
Xinghan Chen, Zheng Wang 0044, Lin Zuo, Bo Li 0072, Yang Yang 0002
Pattern Recognit. Lett.4
2020 Cross-Modal Attention With Semantic Consistence for Image-Text Matching
abstract
The task of image-text matching refers to measuring the visual-semantic similarity between an image and a sentence. Recently, the fine-grained matching methods that explore the local alignment between the image regions and the sentence words have shown advance in inferring the image-text correspondence by aggregating pairwise region-word similarity. However, the local alignment is hard to achieve as some important image regions may be inaccurately detected or even missing. Meanwhile, some words with high-level semantics cannot be strictly corresponding to a single-image region. To tackle these problems, we address the importance of exploiting the global semantic consistence between image regions and sentence words as complementary for the local alignment. In this article, we propose a novel hybrid matching approach named Cross-modal Attention with Semantic Consistency (CASC) for image-text matching. The proposed CASC is a joint framework that performs cross-modal attention for local alignment and multilabel prediction for global semantic consistence. It directly extracts semantic labels from available sentence corpus without additional labor cost, which further provides a global similarity constraint for the aggregated region-word similarity obtained by the local alignment. Extensive experiments on Flickr30k and Microsoft COCO (MSCOCO) data sets demonstrate the effectiveness of the proposed CASC on preserving global semantic consistence along with the local alignment and further show its superior image-text matching performance compared with more than 15 state-of-the-art methods.
Xing Xu 0001, Yang Yang 0002, Lin Zuo, Fumin Shen, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.4
2020 Exploring nonnegative and low-rank correlation for noise-resistant spectral clustering
Zheng Wang 0044, Lin Zuo, Jing Ma 0004, Jingjing Li 0001, Zhao Kang 0001, Lei Zhang 0038
World Wide Web2
2019 Residual Graph Convolutional Networks for Zero-Shot Learning
abstract
Most existing Zero-Shot Learning (ZSL) approaches adopt the semantic space as a bridge to classify unseen categories. However, it is difficult to transfer knowledge from seen categories to unseen categories through semantic space, since the correlations among categories are uncertain and ambiguous in the semantic space. In this paper, we formulated zero-shot learning as a classifier weight regression problem. Specifically, we propose a novel Residual Graph Convolution Network (ResGCN) which takes word embeddings and knowledge graph as inputs and outputs a visual classifier for each category. ResGCN can effectively alleviate the problem of over-smoothing and over-fitting. During the test, an unseen image can be classified by ranking the inner product of its visual feature and predictive visual classifiers. Moreover, we provide a new method to build a better knowledge graph. Our approach not only further enhances the correlations among categories, but also makes it easy to add new categories to the knowledge graph. Experiments conducted on the large-scale ImageNet 2011 21K dataset demonstrate that our method significantly outperforms existing state-of-the-art approaches.
Jiwei Wei, Yang Yang 0002, Jingjing Li 0001, Lei Zhu 0002, Lin Zuo, Heng Tao Shen
MMAsia5
2017 A Fast Precise-Spike and Weight-Comparison Based Learning Approach for Evolving Spiking Neural Networks
Lin Zuo, Hong Qu 0002, Malu Zhang
ICONIP (3)1
2017 A Dynamic Region Generation Algorithm for Image Segmentation Based on Spiking Neural Network
Lin Zuo, Linyao Ma, Yanqing Xiao, Malu Zhang, Hong Qu 0002
ICONIP (3)1
2013 Convergence and Chaos of a Class of Discrete-Time Background Neural Networks with Uniform Firing Rate
abstract
The dynamical properties of a class of discrete-time background network with uniform firing rate are investigated. The conditions for stability are derived. To guaranteed the boundness of all trajectories of the discrete-time background network, several invariant sets are obtained. It's then proved that any trajectories of the network starting from each of the invariant sets will converge. In addition to the stability and convergence analysis, bifurcation and chaos are also discussed. It's shown that the network can engender bifurcation and chaos with the increase of background input. The Lyapunov exponents are finally computed to confirm the existence of chaos. Since the background networks originate from the study of the activities of brain and chaotic activities are ubiquitous in the human brain, the chaos analysis of the background networks is significant.
Lin Zuo, Jinrong Hu
DASC2
2007 Sequential Pattern-Based Cache Replacement in Servlet Container
Lin Zuo, Jun Wei 0001, Hua Zhong 0001, Tao Huang 0001
ICWE2
2006 A Fault-Tolerant Scheme for Complex Transaction Patterns in J2EE
abstract
End-to-end reliability is an important issue in building large-scale distributed enterprise applications based on multi-tier architecture, but the support of reliability as adopted in conventional replication or transactions mechanisms is not enough due to their distinct objectives - replication guarantees the liveness of computational operations by using forward error recovery, while transactions guarantee the safety of application data by using backward error recovery. Combining the two mechanisms for stronger reliability is a challenging task Current solutions, however, are typically on the assumption of simple transaction pattern where a request from a single client executes in the context of exactly one transaction at the middle-tier application server, and seldom think about some complex patterns, such as client transaction enclosing multiple client requests or nested transactions. In this paper, we first identify four transaction pattern classes, and then propose a fault-tolerant scheme that can uniformly provide exactly-once semantic reliability support for these patterns. In this scheme, application servers are passively replicated to endow business logics with high reliability and high availability. In addition, by replicating transaction coordinator, the blocking problem of 2PC protocol during distributed transactions processing is eliminated. We have implemented this approach and integrated it into our own J2EE application server, OnceAS. Also, its effectiveness is discussed in different transaction patterns and the corresponding performance is evaluated
Lin Zuo, Shaohua Liu 0002, Jun Wei 0001
EDOC1
2006 Combining Replication with Transaction Processing for Enhanced Reliability in J2EE
abstract
The multi-tier architecture of J2EE provides good modularity and scalability by partitioning an application into several tiers, and becomes the mainstream for distributed applications development on Internet/Intranet. Current reliability solutions in this architecture are typically dependent on either replication, which provides at-least-once guarantee, or transaction processing, which guarantees at-most-once semantics. In practice, the end-to-end reliability guarantee of exactly-once semantics is necessary, especially for some complex transaction scenarios, such as client transaction or nested transaction. In this paper, we describe a fault-tolerant algorithm that can provide this enhanced reliability support through combining replication and transaction processing. We use passive replication to protect business processing at middle-tier application server. A client stub transparently intercepts client request and automatically resubmits it in the case of failure. In addition, transaction coordinator is passively replicated to prevent the blocking problem of distributed transaction. Also, different application scenarios are discussed to illustrate the effectiveness of this algorithm, and a performance study based on our implementation in J2EE application server, OnceAS, shows the overhead of it is acceptable
Lin Zuo, Shaohua Liu 0002, Jun Wei 0001
ISSRE1