EDBT 2026 Demo / reviewers in the wild / expert
Shujing Lyu
dblp:229/8425
· DBLP profile ↗
27ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0003-2623-1379ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LithoMamba: High-fidelity lithography simulation with State Space ModelsabstractLithography simulation is a critical technology in modern semiconductor manufacturing, yet existing deep learning models often fail to accurately model the complex, long-range optical physics due to the inherent locality of convolution. This limitation results in insufficient simulation fidelity and poses significant challenges for optimization tasks. To overcome this challenge, we introduce LithoMamba, the first generative framework to leverage Mamba for high-fidelity lithography simulation. Our architecture uses a Mamba Generator to model global and long-range optical interactions, while a local, MLP-free Discriminator provides precise, spatial feedback to ensure fine-grained pattern fidelity. This global-local design enables our model to achieve both physical realism and exceptional detail. Our experiments show that LithoMamba outperforms existing methods, both in quantitative and qualitative results. These findings demonstrate the promise of State Space Models for improving lithography simulation and suggest new possibilities for combining physics with generative AI in chip manufacturing. Daohui Wang, Shujing Lyu, Pourya Shamsolmoali, Jiwei Shen, Yue Lu 0001 |
DATE | 3 |
| 2026 | Feature Enhancement Module Based on Class-Centric Loss for Fine-Grained Visual ClassificationabstractWe propose a novel feature enhancement module designed for fine-grained visual classification tasks, which can be seamlessly integrated into various backbone architectures, including both convolutional neural network (CNN)-based and Transformer-based networks. The plug-and-play module outputs pixel-level feature maps and performs a weighted fusion of filtered features to enhance fine-grained feature representation. We introduce a class-centric loss function that optimizes the alignment of samples with their target class centers by pulling them toward the center of the target class while simultaneously pushing them away from the center of the most visually similar nontarget classes. Soft labels are employed to mitigate overfitting, ensuring the model generalizes well to unseen examples. Our approach consistently delivers significant improvements in accuracy across various mainstream backbone architectures, underscoring its versatility and robustness. Furthermore, we achieved the highest accuracy on the NABirds (NAB) and our proprietary lock cylinder datasets. We have released our source code and pretrained model on GitHub: https://github.com/Richard5413/FEM-CC.git. Daohui Wang, Shujing Lyu, Tian Wei, Yue Lu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | MAED: Mask assignment encoder decoder solver for multiple patterning layout decomposition
Jiwei Shen, Shujing Lyu, Yue Lu 0001 |
Expert Syst. Appl. | 3 |
| 2025 | GCENet: A geometric correspondence estimation network for tracking and loop detection in visual-inertial SLAM
Jichao Zhou, Jiwei Shen, Shujing Lyu, Yue Lu 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Leveraging Predictions of Task-Related Latents for Interactive Visual NavigationabstractInteractive visual navigation (IVN) involves tasks where embodied agents learn to interact with the objects in the environment to reach the goals. Current approaches exploit visual features to train a reinforcement learning (RL) navigation control policy network. However, RL-based methods continue to struggle at the IVN tasks as they are inefficient in learning a good representation of the unknown environment in partially observable settings. In this work, we introduce predictions of task-related latents (PTRLs), a flexible self-supervised RL framework for IVN tasks. PTRL learns the latent structured information about environment dynamics and leverages multistep representations of the sequential observations. Specifically, PTRL trains its representation by explicitly predicting the next pose of the agent conditioned on the actions. Moreover, an attention and memory module is employed to associate the learned representation to each action and exploit spatiotemporal dependencies. Furthermore, a state value boost module is introduced to adapt the model to previously unseen environments by leveraging input perturbations and regularizing the value function. Sample efficiency in the training of RL networks is enhanced by modular training and hierarchical decomposition. Extensive evaluations have proved the superiority of the proposed method in increasing the accuracy and generalization capacity. Jiwei Shen, Yue Lu 0001, Shujing Lyu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Enhanced Deep Reinforcement Learning for Parcel Singulation in Non-Stationary EnvironmentsabstractIn the rapidly expanding logistics sector, parcel singulation has emerged as a significant bottleneck. To address this, we propose an automated parcel singulator utilizing a sparse actuator array, which presents an optimal balance between cost and efficiency, albeit requiring a sophisticated control policy. In this study, we frame the parcel singulation issue as a Markov Decision Process with a variable state space dimension, addressed through a deep reinforcement learning (RL) algorithm complemented by a State Space Standardization Module (S3). Distinct from previous RL approaches, our methodology initially considers the non-stationary environment during the problem modeling phase. To counter this challenge, the S3 module standardizes the dynamic input state, thereby stabilizing the RL training process. We validate our method through simulation experiments in complex environments, comparing it with several baseline algorithms. Results indicate that our algorithm excels in parcel singulation tasks, achieving a higher success rate and enhanced efficiency. Jiwei Shen, Hu Lu, Shujing Lyu, Yue Lu 0001 |
ICASSP | 4 |
| 2024 | Enhancing Reinforcement Learning via Causally Correct Input Identification and Targeted InterventionabstractCausal confusion, characterized by the learning of spurious correlations, detrimentally affects the generalization and effectiveness of reinforcement learning (RL) algorithms, especially in environments without latent confounders often encountered in robot autonomous navigation tasks. This study addresses this gap by developing a causal structure within a Partially Observable Markov Decision Process (POMDP). Subsequently, we introduce a targeted intervention that mitigates the influence of spurious correlations by isolating causally significant state variables and discarding irrelevant inputs. Testing in three real-world scenarios confirms the approach’s feasibility and superiority in enhancing the RL algorithms’ performance and generalization ability, signifying a promising step towards more robust online RL frameworks. Jiwei Shen, Hu Lu, Shujing Lyu, Yue Lu 0001 |
ICASSP | 4 |
| 2024 | RSTAN: Residual Spatio-Temporal Attention Network for End-to-End Human Fall Detection
Yaru Jiang, Shujing Lyu, Hongjian Zhan, Yue Lu 0001 |
ICPR (15) | 2 |
| 2024 | Learning to Detect Lithography Defects in SEM Images
Hu Lu, Botong Zhao, Jiwei Shen, Hongjian Zhan, Shujing Lyu, Yue Lu 0001 |
ICPR (5) | 5 |
| 2024 | Enhancing parcel singulation efficiency through transformer-based position attention and state space augmentation
Jiwei Shen, Hu Lu, Shujing Lyu, Yue Lu 0001 |
Expert Syst. Appl. | 3 |
| 2024 | NDOrder: Exploring a novel decoding order for scene text recognition
Dajian Zhong, Hongjian Zhan, Shujing Lyu, Cong Liu 0006, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Expert Syst. Appl. | 3 |
| 2024 | LithoPW: Leveraging Visual Memory Encoding and Defect-Aware Optimization for Precise Determination of the Lithography Process WindowsabstractLithography stands as a critical step in the manufacturing of integrated circuits, where the precise control of focus and exposure dose parameters is vital for optimal results. The conventional methodologies for defining lithography process windows often face difficulties with managing measurement errors, detecting printed defects, and exploiting visual features from Scanning Electron Microscope (SEM) images. This paper proposes LithoPW, a novel framework that utilizes visual features of SEM images for the determination of process windows. This approach is comprised of a denoising module, a Transformer-based visual memory encoder, and a defect-aware process window optimization module. The denoising module incorporates a Transformer architecture to mitigate the impact of noise, thereby enhancing the efficiency of downstream tasks in leveraging information embedded within SEM images. The transformer-based visual memory encoder discerns each SEM image as a Query, maintaining neighbouring SEM images in memory as Key and Value elements, thereby facilitating precise lithography quality classification associated with the query image. The defect-aware process window optimization module heightens the reliability of the results by adjusting the process window according to the defects identified within the SEM images. Experimental results confirm the efficacy of our framework, highlighting its promising application in lithography production for accurate process window determination. Jiwei Shen, Shujing Lyu, Yue Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | A New Few-Shot Learning-Based Model for Prohibited Objects Detection in Cluttered Baggage X-Ray Images Through Edge Detection and Reverse ValidationabstractDetecting prohibited items via X-ray screening at airports and sensitive venues is essential for preventing smuggling and breaches of security. The difficulty in prohibited items inspection lies in accurately detecting prohibited items in complex X-ray images and limited access to X-ray images containing prohibited items. Few-shot detection aims at learning with limited examples and assigning a category label to each object. However, most few-shot learning methods do not focus on the edge information of the occluded object in X-ray images, which is crucial for the model to detect prohibited items in the X-ray images. In this paper, we presents a method (RVViT) for few-shot prohibited items detection tasks which fully acknowledges the significance of X-ray penetrability and increases the stability of few-shot learning model. Specifically, a Transformer encoder is firstly adopted for generating high-level semantic features that contain global information. At the same time, an edge detection module is devised for enhancing the edge information of prohibited items. Moreover, to further improve the stability of the few-shot learning model and ensure prototype consistency between the support and query samples, a reverse validation strategy is proposed to assist training. Extensive experiments demonstrate our method outperforms state-of-the-art approaches in terms of detection with a small number of samples. Shujing Lyu, Palaiahnakote Shivakumara, Michael Blumenstein, Yue Lu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Few-Shot Segmentation for Prohibited Items Inspection With Patch-Based Self-Supervised Learning and Prototype Reverse ValidationabstractProhibited items inspection using X-ray screening is essential for reducing the risk of crime and terrorist attacks. The difficulty in prohibited items inspection lies in accurately detecting prohibited items in complex X-ray images and limited access to X-ray images containing prohibited items. Few-shot segmentation aims at learning with limited examples and assigning a category label to each image pixel. However, current few-shot methods are mostly full-supervised and less robust to the prohibited items categories that did not appear during training process. In this paper, we propose a method for few-shot prohibited items segmentation tasks which utilize unlabeled data and better leverage the representation of input samples during model training process. Specifically, a patch-based self-supervised embedding network is firstly devised as the base learner to learn an abstract representation of the observation from unlabeled samples. Then we apply few-shot learning and generate abstract representation related to prohibited items from support sample within the embedding space, which is followed by obtaining the corresponding class-specific prototype representations via masked average pooling. The distance between each pixel of query sample and prototypes are calculated to predict the label of each pixel. Moreover, prototype reverse validation strategy (PRV) is proposed to further exploit the support representation to assist training. Extensive experiments show that our proposed method outperforms the state-of-the-art by delivering a higher accuracy on automated prohibited items inspection and requiring less labeled samples. Shujing Lyu, Yue Lu 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | SGBANet: Semantic GAN and Balanced Attention Network for Arbitrarily Oriented Scene Text Recognition
Dajian Zhong, Shujing Lyu, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
ECCV (28) | 2 |
| 2022 | Text proposals with location-awareness-attention network for arbitrarily shaped scene text detection and recognition
Dajian Zhong, Shujing Lyu, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Expert Syst. Appl. | 2 |
| 2021 | DenseNet-CTC: An end-to-end RNN-free architecture for context-free string recognition
Hongjian Zhan, Shujing Lyu, Yue Lu 0001, Umapada Pal 0001 |
Comput. Vis. Image Underst. | 2 |
| 2020 | FACLSTM: ConvLSTM with focused attention for scene text recognition
Wenjing Jia, Xiangjian He, Michael Blumenstein, Shujing Lyu, Yue Lu 0001 |
Sci. China Inf. Sci. | 6 |
| 2019 | DeepText: Detecting Text from the Wild with Multi-ASPP-Assembled DeepLababstractIn this paper, we address the issue of scene text detection in the way of direct regression and successfully adapt an effective semantic segmentation model, DeepLab v3+ [1], for this application. In order to handle texts with arbitrary orientations and sizes and improve the recall of small texts, we propose to extract features of multiple scales by inserting multiple Atrous Spatial Pyramid Pooling (ASPP) layers to the DeepLab after the feature maps with different resolutions. Then, we set multiple auxiliary IoU losses at the decoding stage and make auxiliary connections from the intermediate encoding layers to the decoder to assist network training and enhance the discrimination ability of lower encoding layers. Experiments conducted on the benchmark scene text dataset ICDAR2015 demonstrate the superior performance of our proposed network, named as DeepText, over the state-of-the-art approaches. Wenjing Jia, Xiangjian He, Yue Lu 0001, Michael Blumenstein, Shujing Lyu |
ICDAR | 7 |
| 2019 | Change Detection via Graph Matching and Multi-View Geometric ConstraintsabstractChange detection is a critical preprocessing step of visual perception with broad prospects. Its primary challenge is to identify all the meaningful changes from a target image to the source image, which is observed of the same scene and has a different perspective as well. A robust change detection method involving graph matching and geometric constraints is proposed in this paper. Maximum common sub-graph matching is applied for alleviating the risk of suboptimal results and geometric constraints are used to remove the possible mistaken results. Detection results in different real-world scenes with respect to considerable textural moved objects show that the proposed method is more robust than the state-of-the-art methods. Jiwei Shen, Shujing Lyu, Yue Lu 0001 |
ICIP | 2 |
| 2019 | Writing Style Adversarial Network for Handwritten Chinese Character Recognition
Shujing Lyu, Hongjian Zhan, Yue Lu 0001 |
ICONIP (4) | 2 |
| 2019 | Modified Adaptive Implicit Shape Model for Object Detection
Ziyan Xu, Shujing Lyu, Weiping Jin, Yue Lu 0001 |
ICONIP (5) | 2 |
| 2019 | Residual CRNN and Its Application to Handwritten Digit String Recognition
Hongjian Zhan, Shujing Lyu, Xiao Tu, Yue Lu 0001 |
ICONIP (5) | 2 |
| 2019 | Improving Text-Independent Chinese Writer Identification with the Aid of Character PairsabstractText-independent Chinese writer identification does not depend on the text content of the query and reference handwritings. In order to deal with the uncertainty of the text content, text-independent approaches usually give special attention to the global writing style of handwriting, rather than the properties of each individual character or word. Thanks to the existence of high-frequency characters, some characters probably appear in both the query and reference handwritings in most cases. If character images in the query handwriting are similar to those in the reference handwriting, this query handwriting and the corresponding reference handwriting are very likely to be written by the identical writer. In this paper, we exploit the above characteristic to improve the performance of Chinese writer identification. We first present an identification scheme using edge co-occurrence feature (ECF). Then, we detect the character pairs in the query and reference handwritings using a two-step framework and propose the displacement field-based similarity (DFS) to determine whether a character pair is written by the identical writer. The character pairs help to re-rank the candidate list obtained by text-independent ECF-based similarity and finally decide the writer of the query handwriting. The proposed method is evaluated on the HIT-MW and CASIA-2.1 datasets. Experimental results demonstrate that our proposed method outperforms the existing ones, and its Top-1 accuracy on the two datasets reaches 97.1% and 98.3%, respectively. Yujie Xiong, Li Liu 0010, Shujing Lyu, Patrick Shen-Pei Wang, Yue Lu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2018 | Improving Off-Line Handwritten Chinese Character Recognition with Semantic Information
Hongjian Zhan, Shujing Lyu, Yue Lu 0001 |
ICONIP (5) | 2 |
| 2018 | A Fusion Strategy for the Single Shot Text DetectorabstractIn this paper, we propose a new fusion strategy for scene text detection. The system is based on a single fully convolution network, which outputs the coordinates of text bounding boxes at multiple scales. We improve the performance of text detection by combining a fusion strategy. This strategy obtains precise text bounding boxes according to the confidence of candidate text boxes. It exhibits promising robustness and discriminative power by fusing text boxes. Experimental results on ICDAR2011 and ICDAR2013 datasets indicate the effectiveness and robustness of the proposed fusion strategy with an F-measure of 87%, which outperforms the base network 2%. Shujing Lyu, Yue Lu 0001, Patrick Shen-Pei Wang |
ICPR | 2 |
| 2018 | Handwritten Digit String Recognition using Convolutional Neural NetworkabstractString recognition is one of the most important tasks in computer vision applications. Recently the combinations of convolutional neural network (CNN) and recurrent neural network (RNN) have been widely applied to deal with the issue of string recognition. However RNNs are not only hard to train but also time-consuming. In this paper, we propose a new architecture which is based on CNN only, and apply it to handwritten digit string recognition (HDSR). This network is composed of three parts from bottom to top: feature extraction layers, feature dimension transposition layers and an output layer. Motivated by its super performance of DenseNet, we utilize dense blocks to conduct feature extraction. At the top of the network, a CTC (connectionist temporal classification) output layer is used to calculate the loss and decode the feature sequence, while some feature dimension transposition layers are applied to connect feature extraction and output layer. The experiments have demonstrated that, compared to other methods, the proposed method obtains significant improvements on ORAND-CAR-A and ORAND-CAR-B datasets with recognition rates 92.2% and 94.02%, respectively. Hongjian Zhan, Shujing Lyu, Yue Lu 0001 |
ICPR | 2 |