EDBT 2026 Demo / reviewers in the wild / expert
Yirui Wu
dblp:71/8497
· DBLP profile ↗
70ranked-venue papers
33as first author
42since 2021 · last 2026
0000-0003-3022-3718ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 13 first-author · 13 since 2021Artificial intelligence and machine learning · 31 · 15 first-author · 19 since 2021Computer networks · 5 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zero-shot Recommendation: Towards Class Semantic Relation Learning for Inferring Labels of Unseen Micro-videosabstractMicro-video label prediction plays a pivotal role on contemporary video-sharing platforms, such as Kwai and Tiktok. The emergence of video content lacking labels presents a formidable challenge for conventional user interest prediction methods. This paper addresses the challenge of micro-video label prediction, particularly for unseen videos, by proposing a zero-shot method called Class Semantic Relation Learning (CSRL). Unlike traditional user interest prediction models, CSRL leverages the pre-trained Large Language Model (LLM) to enhance prediction accuracy for unlabeled videos. The novelty of CSRL lies in its integration of three key components: a raw feature autoencoder, LLM-enhanced features, and a decomposed graph network. The decomposed graph network is specifically designed to disentangle the relationships between labeled and unlabeled videos, offering a significant improvement over previous methods. By fusing hidden topics with LLM-enhanced text, CSRL effectively handles sparse video features. Experiments on large-scale datasets from the Kwai platform show that CSRL achieves state-of-the-art results, with up to 44.64% improvement in Hit Ratio (HR), highlighting its superiority over existing zero-shot recommendation models in predicting user interests within the user-video network. Junyang Chen 0001, Huan Wang 0005, Yirui Wu, Qiuzhen Lin, Yunfeng Diao, Junkai Ji |
AAAI | 3 |
| 2026 | BTMC: Hyper-relation extraction via BiTime-LSTM and multi-relation contrastive learning
Tingting Hang, Haichao Ding, Jun Feng 0001, Yirui Wu |
Expert Syst. Appl. | 5 |
| 2026 | Edge-Cloud Collaborated Prototype Graph Network for Efficient Few-Shot Object DetectionabstractWith the rapid development of industrial automation, few-shot object detection has emerged as a promising solution for recognizing novel categories using only limited annotated data. However, existing approaches often suffer from high computational complexity and limited adaptability when deployed in resource-constrained industrial environments. To achieve precise detection, efficiency, and security, this paper proposes a collaborative computing framework based on an Edge-Cloud Dual-Prototype Graph Convolutional Network (EC-DP-GCN) for few-shot object detection with hierarchical knowledge embedding. The framework comprises three key components: a device–edge–cloud architecture, a Positive-Negative Prototype (PNP) module, and a Class-Prototype-Sample Hierarchical Graph (CPS-HG) module. Specifically, the PNP module explicitly models intra-class diversity by constructing discriminative positive and negative prototypes from limited support samples, thereby enhancing prototype representativeness. In addition, we further introduce the CPS-HG module, which treats the dual prototypes as class-based prior knowledge and models the relationships among samples through a hierarchical graph structure encompassing class, prototype, and sample levels. This design effectively expands the semantic margins in the embedding space to improve knowledge-guided detection. Extensive experiments on the PASCAL VOC and MS COCO benchmarks demonstrate that EC-DP-GCN significantly outperforms strong baselines and previous state-of-the-art methods, achieving an average improvement of 1.1% in 10-shot detection scenarios. Yirui Wu, Xinfu Liu 0001, Shaohua Wan 0001, Guohua Lv, Jiehan Zhou, Joel J. P. C. Rodrigues |
IEEE Internet Things J. | 1 |
| 2026 | Alignment-aware fine-tuning of vision-language models for out-of-distribution generalization
Yirui Wu, Mohammed A.-M. Salem, Lixin Yuan, Junyang Chen 0001, Huan Wang 0005, Shaohua Wan 0001 |
Multim. Syst. | 2 |
| 2026 | Continual relation extraction with wake-sleep memory consolidation
Tingting Hang, Jun Huang 0003, Yirui Wu, Umapada Pal 0001, Palaiahnakote Shivakumara |
Pattern Recognit. | 4 |
| 2026 | Diffusion models with spatial control and attention fusion for incremental few-shot semantic segmentation
Guangchen Shi, Yirui Wu, Palaiahnakote Shivakumara, Shirong Zou, Tong Lu 0002 |
Pattern Recognit. | 2 |
| 2026 | Representative instance selection strategy for discriminative features
Lixin Yuan, Ningyu Du, Yirui Wu, Palaiahnakote Shivakumara, Umapada Pal 0001 |
Pattern Recognit. Lett. | 3 |
| 2026 | Plausible Deniable Medical Image Encryption by Large Language Models and Reversible Content-Aware StrategyabstractThere is a rising concern about healthcare system security, where data loss could bring lots of damages to patients and hospitals. As a promising encryption method for medical images, DNA encoding own characteristics of high speed, parallelism computation, minimal storage, and unbreakable cryptosystems. Inspired by the idea of involving Large Language Models(LLMs) to improve DNA encoding, we propose a medical image encryption method with LLM-enhanced DNA encoding, which consists of LLM enhancing module and content-aware permutation&diffusion module. Regarding medical images generally have plain backgrounds with low-entropy pixels, the first module compresses pixels into highly compact signals with features of probabilistic varying and plausibly deniability, serving as another LLM-based layer of defense against privacy breaches before DNA encoding. The second module not only adds permutation by randomly sampling from a redundant correlation between adjacent pixels to break the internal links between pixels but also performs a DNA-based diffusion process to greatly increase the complexity of cracking. Experiments on ChestXray-14, COVID-CT and fcon-1000 datasets show that the proposed method outperforms all comparative methods in sensitivity, correlation and entropy. Yirui Wu, Xinfu Liu 0001, Lucia Cascone, Michele Nappi, Shaohua Wan 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Deconfound Semantic Shift and Incompleteness in Incremental Few-shot Semantic SegmentationabstractIncremental few-shot semantic segmentation (IFSS) expands segmentation capacity of the trained model to segment new-class images with few samples. However, semantic meanings may shift from background to object class or vice versa during incremental learning. Moreover, new-class samples often lack representative attribute features when the new class greatly differs from the pre-learned old class. In this paper, we propose a causal framework to discuss the cause of semantic shift and incompleteness in IFSS, and we deconfound the revealed causal effects from two aspects. First, we propose a Causal Intervention Module (CIM) to resist semantic shift. CIM progressively and adaptively updates prototypes of old class, and removes the confounder in an intervention manner. Second, a Prototype Refinement Module (PRM) is proposed to complete the missing semantics. In PRM, knowledge gained from the episode learning scheme assists in fusing features of new-class and old-class prototypes. Experiments on both PASCAL-VOC 2012 and ADE20k benchmarks demonstrate the outstanding performance of our method. Yirui Wu, Yuhang Xia, Lixin Yuan, Junyang Chen 0001, Jun Liu 0036, Shaohua Wan 0001 |
AAAI | 1 |
| 2025 | SeqBench: Benchmarking Sequential Narrative Generation in Text-to-Video ModelsabstractText-to-video (T2V) generation models have made significant progress in creating visually appealing videos. However, they struggle with generating coherent sequential narratives that require logical progression through multiple events. Existing T2V benchmarks primarily focus on visual quality metrics but fail to evaluate narrative coherence over extended sequences. To bridge this gap, we present SeqBench, a comprehensive benchmark for evaluating sequential narrative coherence in T2V generation. SeqBench includes a carefully designed dataset of 320 prompts spanning various narrative complexities, with 2,560 human-annotated videos generated from 8 state-of-the-art T2V models. Additionally, we design a Dynamic Temporal Graphs (DTG)-based automatic evaluation metric, which can efficiently capture long-range dependencies and temporal ordering while maintaining computational efficiency. Our DTG-based metric demonstrates a strong correlation with human annotations. Through systematic evaluation using SeqBench, we reveal critical limitations in current T2V models: failure to maintain consistent object states across multi-action sequences, physically implausible results in multi-object scenarios, and difficulties in preserving realistic timing and ordering relationships between sequential actions. SeqBench provides the first systematic framework for evaluating narrative coherence in T2V generation and offers concrete insights for improving sequential reasoning capabilities in future models. Please refer to https://videobench.github.io/SeqBench.github.io/ for more details. Zhengxu Tang, Zizheng Wang, Luning Wang, Zitao Shuai, Siyu Qian, Yirui Wu, Haosong Rao, Chenwei Wu 0006 |
CBMI | 7 |
| 2025 | Stray Intrusive Outliers-Based Feature Selection on Intra-Class Asymmetric Instance Distribution or Multiple High-Density ClustersabstractFor data with intra-class Asymmetric instance Distribution or Multiple High-density Clusters (ADMHC), outliers are real and have specific patterns for data classification, where the class body is necessary and difficult to identify. Previous Feature Selection (FS) methods score features based on all training instances or rarely target intra-class ADMHC. In this paper, we propose a supervised FS method, Stray Intrusive Outliers-based FS (SIOFS), for data classification with intra-class ADMHC. By focusing on Stray Intrusive Outliers (SIOs), SIOFS modifies the skewness coefficient and fuses the threshold in the 3$\sigma$ principle to identify the class body, scoring features based on the intrusion degree of SIOs. In addition, the refined density-mean center is proposed to represent the general characteristics of the class body reasonably. Mathematical formulations, proofs, and logical exposition ensure the rationality and universality of the settings in the proposed SIOFS method. Extensive experiments on 16 diverse benchmark datasets demonstrate the superiority of SIOFS over 12 state-of-the-art FS methods in terms of classification accuracy, normalized mutual information, and confusion matrix. SIOFS source codes is available at https://github.com/XXXly/2025-ICML-SIOFS Lixin Yuan, Yirui Wu, Minglei Yuan, Jun Liu 0036 |
ICML | 2 |
| 2025 | Diffuse&Refine: Intrinsic Knowledge Generation and Aggregation for Incremental Object DetectionabstractIncremental Object Detection(IOD) targets at progressively extending capability of object detectors to recognize new classes. However, representation confusion between old and new classes leads to catastrophic forgetting. To alleviate this problem, we propose DiffKA, with intrinsic knowledge generated and aggregated by forward and backward diffusion, gradually establishing rigid class boundary. With incremental streaming data, forward diffusion spreads information to generate potential inter-class associations among new- and old-class prototypes within a hierarchical tree, named as Intrinsic Correlation Tree(ICTree), to store intrinsic knowledge. Afterwards, backward diffusion refines and aggregates the generated knowledge in ICTree, explicitly establishing rigid class boundary to mitigate representation confusion. To keep semantic consistency with extreme IOD settings, we reorganize semantic relevance of old- and new-class prototypes in paradigms to adaptively and effectively update DiffKA. Experiments on MS COCO dataset show DiffKA achieves state-of-the-art performance on IOD tasks with significant advantages. Yirui Wu, Lixin Yuan, Jun Liu 0036, Junyang Chen 0001, Huan Wang 0005, Wenhai Wang |
IJCAI | 2 |
| 2025 | Geo-CF2Net: Geometry-Prior Cross-Frequency Interactive Fusion Network for 3D Human Action RecognitionabstractDynamic point cloud-based human action recognition has garnered increasing attention due to its inherent advantages in privacy preservation and structural completeness. Current methods typically rely on nested point spatio-temporal convolutions to understand motion semantics in a bottom-up manner, which is intractable for capturing high-fidelity human dynamics disentangled from spatio-temporal interference. Motivated by this, designing a practical spatio-temporal factorization backbone is essential. However, the repeated coarsening of aggregated features along the spatial dimension often leads to the degradation of intrinsic geometric texture relations within point cloud data. Moreover, discretizing continuous visual data into isolated temporal hyperpoints significantly diminishes temporal continuity, resulting in the fragmentation of human action. To circumvent above limitations, we propose a novel Geometry-Prior Cross-Frequency Interactive Fusion Network (Geo-CF2Net). Specifically, we investigate a Spatial-Geometry Pose Prior (SGPP) module, which compensates for pose information loss during spatial downsampling by explicitly modeling geometric constraints among neighboring points. In addition, we elaborate on a Temporal Motion Unit Interactive Coordination (TMIC) module to track the interactive composite semantics of low-frequency steady-state venations and high-frequency transient-state details within a high-dimensional pose evolution flow. Extensive experiments on three public benchmarks substantiate the superiority of Geo-CF2Net over state-of-the-art methods. Qian Huang 0008, Xing Li 0005, Shihao Han, Yirui Wu, Xin Li 0090, Ziyang Yin |
ACM Multimedia | 7 |
| 2025 | Cross-level Distillation Based Machine Unlearning with Contrastive Enhanced Knowledge
Shijia Qiao, Xinfu Liu 0001, Lixin Yuan, Yirui Wu |
PRCV (1) | 6 |
| 2025 | Edge-Computing-Driven Active-Reference Fusion for Few-Shot Semantic Segmentation
Yirui Wu, Xinfu Liu 0001, Guangchen Shi, Shaohua Wan 0001 |
IEEE Internet Things J. | 1 |
| 2025 | EPM: Evolutionary Perception Method for Anomaly Detection in Noisy Dynamic GraphsabstractWith the rapid expansion of interactions across various domains such as knowledge graphs and social networks, anomaly detection in dynamic graphs has become increasingly critical for mitigating potential risks. However, existing anomaly detection methods often assume noise-free dynamic graphs, overlooking the prevalence of noisy dynamic graphs in real-world applications. Specifically, noisy dynamic graphs affected by structural noises-such as spurious and missing nodes and edges-struggle to consistently provide reliable structural evidence for anomaly detection. To tackle this challenge, we propose an Evolutionary Perception Method (EPM) for identifying anomalous nodes in noisy dynamic graphs by resisting the interference of structural noises. EPM primarily consists of two components: a dynamic fitter and a filtering reviser. The dynamic fitter characterizes the interaction dynamics of nodes that removes and generates links at each period as a multiple superposition state, utilizing various link prediction algorithms to fit evolutionary mechanisms. Additionally, the filtering reviser designs evolutional entropies to quantify the evolutional uncertainty in multiple superposition states, further designing the Kalman filter to optimize these entropies. Extensive experiments show that the proposed EPM method surpasses state-of-the-art approaches in detecting anomalous nodes in noisy dynamic graphs. Huan Wang 0005, Junyang Chen 0001, Yirui Wu, Victor C. M. Leung, Di Wang 0015 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Sparse Enhanced Network: An Adversarial Generation Method for Robust Augmentation in Sequential RecommendationabstractSequential Recommendation plays a significant role in daily recommendation systems, such as e-commerce platforms like Amazon and Taobao. However, even with the advent of large models, these platforms often face sparse issues in the historical browsing records of individual users due to new users joining or the introduction of new products. As a result, existing sequence recommendation algorithms may not perform well. To address this, sequence-based data augmentation methods have garnered attention. Existing sequence enhancement methods typically rely on augmenting existing data, employing techniques like cropping, masking prediction, random reordering, and random replacement of the original sequence. While these methods have shown improvements, they often overlook the exploration of the deep embedding space of the sequence. To tackle these challenges, we propose a Sparse Enhanced Network (SparseEnNet), which is a robust adversarial generation method. SparseEnNet aims to fully explore the hidden space in sequence recommendation, generating more robust enhanced items. Additionally, we adopt an adversarial generation method, allowing the model to differentiate between data augmentation categories and achieve better prediction performance for the next item in the sequence. Experiments have demonstrated that our method achieves a remarkable 4-14% improvement over existing methods when evaluated on the real-world datasets. (https://github.com/junyachen/SparseEnNet) Junyang Chen 0001, Guoxuan Zou, Pan Zhou 0001, Yirui Wu, Zhenghan Chen, Houcheng Su, Huan Wang 0005, Zhiguo Gong |
AAAI | 4 |
| 2024 | Few-shot Semantic Segmentation via Perceptual Attention and Spatial ControlabstractFew-shot semantic segmentation (FSS) aims to locate pixels of unseen classes with clues from a few labeled samples. Recently, thanks to profound prior knowledge, diffusion models have been expanded to achieve FSS tasks. However, due to probabilistic noising and denoising processes, it is difficult for them to maintain spatial relationships between inputs and outputs, leading to inaccurate segmentation masks. To address this issue, we propose a Diffusion-based Segmentation network (DiffSeg), which decouples probabilistic denoising and segmentation processes. Specifically, DiffSeg leverages attention maps extracted from a pretrained diffusion model as support-query interaction information to guide segmentation, which mitigates the impact of probabilistic processes while benefiting from rich prior knowledge of diffusion models. In the segmentation stage, we present a Perceptual Attention Module (PAM), where two cross-attention mechanisms capture semantic information of support-query interaction and spatial information produced by the pretrained diffusion model. Furthermore, a self-attention mechanism within PAM ensures a balanced dependence for segmentation, thus preventing inconsistencies between the aforementioned semantic and spatial information. Additionally, considering the uncertainty inherent in the generation process of diffusion models, we equip DiffSeg with a Spatial Control Module (SCM), which models spatial structural information of query images to control boundaries of attention maps, thus aligning the spatial location between knowledge representation and query images. Experiments on PASCAL-5i and COCO datasets show that DiffSeg achieves new state-of-the-art performance with remarkable advantages. Guangchen Shi, Yirui Wu, Danhuai Zhao, Tong Lu 0002 |
ACM Multimedia | 3 |
| 2024 | Feature Fusion Pyramid Network for End-to-End Scene Text DetectionabstractHow to properly involve text characteristics like multi-scale, arbitrary direction, length aspect ratio, into detection network design has become a hot topic in computer vision. Feature Pyramid Network (FPN) is a typical method to achieve robust text detection, where its low-level and high-level feature map retains spatial structure and global semantic information, respectively. However, its strict hierarchical structure fails to fuse low-level and high-level information to improve the distinguish ability of feature map. To address this problem, we propose a novel feature fusion pyramid network for end-to-end scene text detection by fusing multi-modal information. By dividing pyramid structure into high-level and low-level layers, channel and spatial attention modules are adopted to enhance high-level and low-level feature representation by encoding channel and spatial-wise context information, respectively. In order to reduce information loss by layer transmission, a special residual network is designed to achieve short-cut between high-level and low-level features, so as to realize multi-modal feature fusion. Experiments show the precision and recall of the proposed method on ICDAR2015, ICDAR2017-MLT, and MSRA-TD500 datasets reach 88.7%/82.1%, 77.0%/60.3%, and 85.3%/74.8%, respectively. Yirui Wu, Lilai Zhang, Hao Li 0089, Shaohua Wan 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2024 | CDT-CAD: Context-Aware Deformable Transformers for End-to-End Chest Abnormality Detection on X-Ray ImagesabstractDeep learning methods have achieved great success in medical image analysis domain. However, most of them suffer from slow convergency and high computing cost, which prevents their further widely usage in practical scenarios. Moreover, it has been proved that exploring and embedding context knowledge in deep network can significantly improve accuracy. To emphasize these tips, we present CDT-CAD, i.e., context-aware deformable transformers for end-to-end chest abnormality detection on X-Ray images. CDT-CAD firstly constructs an iterative context-aware feature extractor, which not only enlarges receptive fields to encode multi-scale context information via dilated context encoding blocks, but also captures unique and scalable feature variation patterns in wavelet frequency domain via frequency pooling blocks. Afterwards, a deformable transformer detector on the extracted context features is built to accurately classify disease categories and locate regions, where a small set of key points are sampled, thus leading the detector to focus on informative feature subspace and accelerate convergence speed. Through comparative experiments on Vinbig Chest and Chest Det 10 Datasets, CDT-CAD demonstrates its effectiveness in recognizing chest abnormities and outperforms 1.4% and 6.0% than the existing methods in$AP_{5}0$and$AR$on VinBig dateset, and 0.9% and 2.1% on Chest Det-10 dataset, respectively. Yirui Wu, Qiran Kong, Lilai Zhang, Aniello Castiglione, Michele Nappi, Shaohua Wan 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | Edge Computing and Few-Shot Learning Featured Intelligent Framework in Digital Twin Empowered Mobile NetworksabstractDigital twins (DT) and mobile networks have evolved forms of intelligence in Internet of Things (IoT). In this work, we consider a Digital Twin Mobile Network (DTMN) scenario with few multimedia samples. Facing challenges of knowledge extraction with few samples, stable interaction with dynamic changes of multimedia data, time and privacy saving in low-resource mobile network, we propose an edge computing and few-shot learning featured intelligent framework. Considering time-sensitive property of transmission and privacy risks of directly uploads in mobile network, we deploy edge computing to locally run networks for analysis, thus saving time to offload computing request and enhancing privacy by encrypting original data. Inspired by remarkable relationship representation of graphs, we build Graph Neural Network (GNN) in cloud to map physical mobile systems to virtual entities with DT, thus performing semantic inferences in cloud with few samples uploaded by edges. Occasionally, node features in GNN could converge to similar, non-discriminative embeddings, causing catastrophic unstable phenomena. An iterative reweight and drop structure (IRDS) is thus constructed in cloud, which nonetheless contributes stability with respect to edge uncertainty. As part of IRDS, a drop Edge&Node scheme is proposed to randomly remove certain nodes and edges, which not only enhances distinguished capability of graph neighbor patterns, but also offers data encryption with random strategy. We show one implementation case of image classification in social network, where experiments on public datasets show that our framework is effective with user-friendly advantages and significant intelligence. Yirui Wu, Yong Lai 0001, Liang Zhao 0004, Xiaoheng Deng, Shaohua Wan 0001 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2023 | CDText: Scene text detector based on context-aware deformable transformer
Yirui Wu, Qiran Kong, Yong Lai 0001, Fabio Narducci, Shaohua Wan 0001 |
Pattern Recognit. Lett. | 1 |
| 2023 | GDRL: An interpretable framework for thoracic pathologic prediction
Yirui Wu, Hao Li 0089, Andrea Casanova, Andrea F. Abate, Shaohua Wan 0001 |
Pattern Recognit. Lett. | 1 |
| 2023 | A novel method of data and feature enhancement for few-shot image classification
Yirui Wu, Benze Wu, Shaohua Wan 0001 |
Soft Comput. | 1 |
| 2023 | Joint Intent Detection Model for Task-oriented Human-Computer Dialogue System using Asynchronous TrainingabstractHow to accurately understand low-resource languages is the core of the task-oriented human-computer dialogue system. Language understanding consists of two sub-tasks, i.e., intent detection and slot filling. Intent detection still faces challenges due to semantic ambiguity and implicit intentions with users’ input. Moreover, separately modeling intent detection and slot filling significantly decrease the correctness and relevance between questions and answers. To address these issues, we propose a joint intent detection method using asynchronous training strategy. The proposed method firstly encodes local text information extracted by CNN and relationship information among words emphasized by attention structure. Later, a joint intent detection model with asynchronous training strategy is proposed by either fusing hidden states of intent detection and slot filling layers, or adopting the key information to fine-tune the whole network, greatly increasing the relevance of intent detection and slot filling subtasks. The accuracy achieved by the proposed method tested on an open-source airline travel dataset and a self-collected electricity service dataset, i.e., ATIS and ECSF, are 97.49% and 89.68%, respectively, which proves the effectiveness of joint learning and asynchronous training. Yirui Wu, Hao Li 0089, Lilai Zhang, Qian Huang 0008, Shaohua Wan 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2023 | Edge-AI-Driven Framework with Efficient Mobile Network Design for Facial Expression RecognitionabstractFacial Expression Recognition (FER) in the wild poses significant challenges due to realistic occlusions, illumination, scale, and head pose variations of the facial images. In this article, we propose an Edge-AI-driven framework for FER. On the algorithms aspect, we propose two attention modules, Arbitrary-oriented Spatial Pooling (ASP) and Scalable Frequency Pooling (SFP), for effective feature extraction to improve classification accuracy. On the systems aspect, we propose an edge-cloud joint inference architecture for FER to achieve low-latency inference, consisting of a lightweight backbone network running on the edge device, and two optional attention modules partially offloaded to the cloud. Performance evaluation demonstrates that our approach achieves a good balance between classification accuracy and inference latency. Yirui Wu, Lilai Zhang, Zonghua Gu 0001, Hu Lu, Shaohua Wan 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2023 | Medical Image Encryption by Content-Aware DNA Computing for Secure HealthcareabstractThere exists a rising concern on security of healthcare data and service. Even small lost, stolen, displaced, hacked, or communicated in personal health data could bring huge damage to patients. Therefore, we propose a novel content-aware deoxyribonucleic acid (DNA) computing system to encrypt medical images, thus guaranteeing privacy and promoting secure healthcare environment. The proposed system consists of sender and receiver to perform tasks of encryption and decryption, respectively, where both contain the same structure design, but perform opposite operations. In either sender or receiver, we design a randomly DNA encoding and a content-aware permutation and diffusion module. Considering introducing random mechanism to increase difficulty of cracking, the former module builds a random encryption rule selector in DNA encoding process by randomly mapping quantity of medical image pixels to outputs. Meanwhile, the latter module constructs a permutation sequence, which not only encodes information of pixel values, but also involves redundant correlation between adjacent pixels located in a patch. Such design brings awareness property of medical image content to greatly increase complexity in cracking by embedding semantical information for encryption. We demonstrate that the proposed system successfully improve cybersecurity of medical images against various attacks in robustness and effectiveness when transmitting data in wireless broadcasting scenarios. Yirui Wu, Lilai Zhang, Stefano Berretti, Shaohua Wan 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Digital Twin of Intelligent Small Surface Defect Detection with Cyber-manufacturing SystemsabstractWith the remarkable technological development in cyber-physical systems, industry 4.0 has evolved by use of a significant concept named digital twin (DT). However, it is still difficult to construct a relationship between twin simulation and a real scenario considering dynamic variations, especially when dealing with small surface defect detection tasks with high performance and computation resource requirements. In this article, we aim to construct cyber-manufacturing systems to achieve a DT solution for small surface defect detection task. Focusing on DT-based solution, the proposed system consists of an Edge–Cloud architecture and a surface defect detection algorithm. Considering dynamic characteristics and real-time response requirement, Edge–Cloud architecture is built to achieve smart manufacturing by efficiently collecting, processing, analyzing, and storing data produced by factory. A deep learning–based algorithm is then constructed to detect surface defeats based on multi-modal data, i.e., imaging and depth data. Experiments show the proposed algorithm could achieve high accuracy and recall in small defeat detection task, thus constructing DT in cyber-manufacturing. Yirui Wu, Guoqiang Yang, Tong Lu 0002, Shaohua Wan 0001 |
ACM Trans. Internet Techn. | 1 |
| 2022 | Learning Group-Disentangled Representation for Interpretable Thoracic Pathologic PredictionabstractDeep learning methods have shown significant performance in medical image analysis tasks. However, they generally act like ”black box” without explanations in both feature extraction and decision processes, leading to lack of clinical insights and high risk assessments. To aid deep learning in envisioning diseases with visual clues, we propose Representation Group-Disentangling Network (RGD-Net), which can completely disentangle feature space of input X-ray images into several independent feature groups, each corresponding to a specific disease. Taking several semantically related and labeled X-ray images as input, RGD-Net firstly extracts completely group-disentangled representations of diseases through Group-Disentangle Module, which applies group-swap and linking operations to construct latent space by enforcing semantic consistency of attributes. To prevent learning degenerate representations defined as shortcut problem, we further introduce adversarial constricts on mapping from features to diseases, thus avoiding model collapse with former free-form disentanglement. Experiments on chestxray-14 and ChestXpert datasets demonstrate that RGD-Net are effective in predicting diseases with remarkable advantages, which leverage potential factors contributing to different diseases, thus enhancing interpretability in working patterns of deep learning methods. Hao Li 0089, Yirui Wu, Hexuan Hu 0001, Hu Lu, Yong Lai 0001, Shaohua Wan 0001 |
BIBM | 2 |
| 2022 | Incremental Few-Shot Semantic Segmentation via Embedding Adaptive-Update and Hyper-class RepresentationabstractIncremental few-shot semantic segmentation (IFSS) targets at incrementally expanding model's capacity to segment new class of images supervised by only a few samples. However, features learned on old classes could significantly drift, causing catastrophic forgetting. Moreover, few samples for pixel-level segmentation on new classes lead to notorious overfitting issues in each learning session. In this paper, we explicitly represent class-based knowledge for semantic segmentation as a category embedding and a hyper-class embedding, where the former describes exclusive semantical properties, and the latter expresses hyper-class knowledge as class-shared semantic properties. Aiming to solve IFSS problems, we present EHNet, i.e., Embedding adaptive-update and Hyper-class representation Network from two aspects. First, we propose an embedding adaptive-update strategy to avoid feature drift, which maintains old knowledge by hyper-class representation, and adaptively update category embeddings with a class-attention scheme to involve new classes learned in individual sessions. Second, to resist overfitting issues caused by few training samples, a hyper-class embedding is learned by clustering all category embeddings for initialization and aligned with category embedding of the new class for enhancement, where learned knowledge assists to learn new knowledge, thus alleviating performance dependence on training data scale. Significantly, these two designs provide representation capability for classes with sufficient semantics and limited biases, enabling to perform segmentation tasks requiring high semantic dependence. Experiments on PASCAL-5i and COCO datasets show that EHNet achieves new state-of-the-art performance with remarkable advantages. Guangchen Shi, Yirui Wu, Jun Liu 0036, Shaohua Wan 0001, Wenhai Wang, Tong Lu 0002 |
ACM Multimedia | 2 |
| 2022 | AI for Online Customer Service: Intent Recognition and Slot Filling Based on Deep Learning Technology
Yirui Wu, Wenqin Mao, Jun Feng 0001 |
Mob. Networks Appl. | 1 |
| 2022 | A novel forget-update module for few-shot domain generalization
Minglei Yuan, Chunhao Cai, Tong Lu 0002, Yirui Wu |
Pattern Recognit. | 4 |
| 2022 | CE-text: A context-Aware and embedded text detector in natural scene images
Yirui Wu, Shaohua Wan 0001 |
Pattern Recognit. Lett. | 1 |
| 2021 | CT-CAD: Context-Aware Transformers for End-to-End Chest Abnormality Detection on X-RaysabstractSupervised based deep learning methods have achieved great success in medical image analysis domain. Essentially, most of them could be further improved by exploring and embedding context knowledge for accuracy boosting. Moreover, they generally suffer from slow convergency and high computing cost, which prevents their usage in a practical scenario. To tackle these problems, we present CT-CAD, context-aware transformers for end-to-end chest abnormality detection on X-Ray images. The proposed method firstly constructs a context-aware feature extractor, which enlarges receptive fields to encode multi-scale context information via an iterative feature fusion scheme and dilated context encoding blocks. Afterwards, deformable transformer detector are built for category classification and location regression, where their deformable attention block attend to a small set of key sampling points, thus allowing the transformer to focus on feature subspace and accelerate convergence speed. Through comparative experiments on Vinbig Chest and Chest Det10 Datasets, the proposed CT-CAD demonstrates its effectiveness and outperforms the existing methods in mAP and training epoches. Qiran Kong, Yirui Wu, Chi Yuan |
BIBM | 2 |
| 2021 | PolarText: Single-stage Scene Text Detection with Polar RepresentationabstractAlthough deep learning has achieved great success in object detection recently, scene text detection is still a challenging task, due to inherent difficulties of locating texts in complex scenes. Many approaches adopt inspirations from segmentation to detect arbitrary shaped scene text. However, most segmentation based methods have high computation cost and generally needs a lot of refinements to get accurate results. To ease this problem, we propose a novel single-stage method, i.e., PolarText network, which detects text regions by generating contour points in polar coordinates. PolarText not only relieves the burden of high computation cost by directly regressing contour points instead of pixels, but also fits with intrinsic characteristics of text instances by centers and contours, thus suppressing mislabeling boundary pixels caused by pixel-level labeling. To cope with polar representation, PolarText utilizes Polar IoU loss and polar centerness to generalize effective paradigms from box representation for polar representation. In addition, we add a dedicated bounding box branch to work with text detection since most text instances are approximately rectangular in shape. Compared with the existing methods, the proposed method achieves superior results in both accuracy and efficiency by testing on CTW 1500 and ICDAR 2015 datasets. Qiran Kong, Yirui Wu, Shaohua Wan 0001 |
EUC | 2 |
| 2021 | An Image Enhancement Method for Few-shot ClassificationabstractIn order to predict the unknown image categories, few-shot image classification has recently become a very hot field. However, many methods need a large number of samples to support in order to achieve enough functions. This makes the whole network de amplification to meet a large number of effective feature extraction, and reduces the efficiency of few-shot classification to a certain extent. To solve these problems, we propose a dilate convolutional network with data enhancement. This network can not only meet the necessary features of image classification without increasing the number of samples, but also has a structure that utilizes a large number of effective features without sacrificing efficiency. The cutout structure can enhance the data by adding a fixed area 0 mask matrix in the process of image input. The structure of FAU uses dilate convolution and uses the characteristics of a sequence to improve the efficiency of the network. Benze Wu, Yirui Wu, Shaohua Wan 0001 |
EUC | 2 |
| 2021 | A Image Enhancement Method for Few-shot Classificationabstractn order to predict the unknown image categories, few-shot image classification has recently become a very hot field. However, many methods need a large number of samples to support in order to achieve enough functions. This makes the whole network de amplification to meet a large number of effective feature extraction, and reduces the efficiency of few-shot classification to a certain extent. To solve these problems, we propose a dilate convolutional network with data enhancement. This network can not only meet the necessary features of image classification without increasing the number of samples, but also has a structure that utilizes a large number of effective features without sacrificing efficiency. The cutout structure can enhance the data by adding a fixed area 0 mask matrix in the process of image input. The structure of FAU uses dilate convolution and uses the characteristics of a sequence to improve the efficiency of the network. Benze Wu, Yirui Wu, Shaohua Wan 0001 |
EUC | 2 |
| 2021 | ARNet: Active-Reference Network for Few-Shot Image Semantic SegmentationabstractTo make predictions on unseen classes, few-shot segmentation becomes a research focus recently. However, most methods build on pixel-level annotation requiring quantity of manual work. Moreover, inherent information on same-category objects to guide segmentation could have large diversity in feature representation due to differences in size, appearance, layout, and so on. To tackle these problems, we present an active-reference network (ARNet) for few-shot segmentation. The proposed active-reference mechanism not only supports accurately cooccurrent objects in either support or query images, but also relaxes high constraint on pixel-level labeling, allowing for weakly boundary labeling. To extract more intrinsic feature representation, a category-modulation module (CMM) is further applied to fuse features extracted from multiple support images, thus forgetting useless and enhancing contributive information. Experiments on PASCAL-5idataset show the proposed method achieves a m-IOU score of 56.5% for 1-shot and 59.8% for 5-shot segmentation, being 0.5% and 1.3% higher than current state-of-the-art method. Guangchen Shi, Yirui Wu, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002 |
ICME | 2 |
| 2021 | Spatial and Temporal Aware Graph Convolutional Network for Flood ForecastingabstractIntelligent flood forecasting systems provide an effective means to forecast flood disaster. Accurate flood flow value prediction is a huge challenge since it's influenced by both spatial and temporal relationship among flood factors. Popular deep learning structures like Long Short-Term Memory (LSTM) network lacks abilities of modeling the spatial correlations of hydrological data, thus cannot yield satisfactory prediction results. Moreover, not all the temporal information is always valuable for flood forecasting. In this paper, we proposed a novel spatial and temporal aware Graph Convolution Network (ST-GCN) for flood prediction, which is capable to extract spatial-temporal information from raw flood data. Moreover, a temporal attention mechanism is introduced to weight the importance of different time steps, thus involving global temporal information to improve flood prediction accuracy. Compared with the existing methods, results on two self-collected datasets show that ST-GCN greatly improves the prediction performance. Jun Feng 0001, Yirui Wu, Yuqi Xi |
IJCNN | 3 |
| 2021 | Joint extraction of entities and overlapping relations using source-target entity labeling
Tingting Hang, Jun Feng 0001, Yirui Wu, Le Yan 0002 |
Expert Syst. Appl. | 3 |
| 2021 | Multiple attention encoded cascade R-CNN for scene text detection
Yirui Wu, Shaohua Wan 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Multi-scale relation reasoning for multi-modal Visual Question Answering
Yirui Wu, Shaohua Wan 0001 |
Signal Process. Image Commun. | 1 |
| 2020 | Dynamic Low-Light Image Enhancement for Object Detection via End-to-End TrainingabstractObject detection based on convolutional neural networks is a hot research topic in computer vision. The illumination component in the image has a great impact on object detection, and it will cause a sharp decline in detection performance under low-light conditions. Using low-light image enhancement technique as a pre-processing mechanism can improve image quality and obtain better detection results. However, due to the complexity of low-light environments, the existing enhancement methods may have negative effects on some samples. Therefore, it is difficult to improve the overall detection performance in low-light conditions. In this paper, our goal is to use image enhancement to improve object detection performance rather than perceptual quality for humans. We propose a novel framework that combines low-light enhancement and object detection for end-to-end training. The framework can dynamically select different enhancement subnetworks for each sample to improve the performance of the detector. Our proposed method consists of two stage: the enhancement stage and the detection stage. The enhancement stage dynamically enhances the low-light images under the supervision of several enhancement methods and output corresponding weights. During the detection stage, the weights offers information on object classification to generate high-quality region proposals and in turn result in accurate detection. Our experiments present promising results, which show that the proposed method can significantly improve the detection performance in low-light environment. Tong Lu 0002, Yirui Wu |
ICPR | 3 |
| 2020 | Multi-scale Relational Reasoning with Regional Attention for Visual Question AnsweringabstractOne of the main challenges of visual question answering (VQA) lies in properly reasoning relations among visual regions involved in the question. In this paper, we propose a novel neural network to perform question-guided relational reasoning in multi-scales for visual question answering, in which each region of image is enhanced by regional attention. Specifically, we present regional attention module, which consists of a soft attention module and a hard attention module, to select informative regions of the image according to informative evaluations implemented by question-guided soft attention. Combinations of different informative regions are then concatenated with question embedding in different scales to capture relational information. Relational reasoning module can extract question-based relational information among regions, in which multi-scale mechanism gives it the ability to model scaled relationships with diversity making it sensitive to numbers. We conduct experiments to show that our proposed architecture is effective and achieves a new state-of-the-art on VQA v2. Yun-Tao Ma, Tong Lu 0002, Yirui Wu |
ICPR | 3 |
| 2020 | Context-Aware Residual Network with Promotion Gates for Single Image Super-Resolution
Xiaozhong Ji, Yirui Wu, Tong Lu 0002 |
MMM (2) | 2 |
| 2020 | TK-Text: Multi-shaped Scene Text Detection via Instance Segmentation
Xiaoge Song, Yirui Wu, Wenhai Wang, Tong Lu 0002 |
MMM (2) | 2 |
| 2020 | CASR: a context-aware residual network for single-image super-resolution
Yirui Wu, Xiaozhong Ji, Wanting Ji, Helen Zhou |
Neural Comput. Appl. | 1 |
| 2020 | Network Attacks Detection Methods Based on Deep Learning Techniques: A SurveyabstractWith the development of the fifth-generation networks and artificial intelligence technologies, new threats and challenges have emerged to wireless communication system, especially in cybersecurity. In this paper, we offer a review on attack detection methods involving strength of deep learning techniques. Specifically, we firstly summarize fundamental problems of network security and attack detection and introduce several successful related applications using deep learning structure. On the basis of categorization on deep learning methods, we pay special attention to attack detection methods built on different kinds of architectures, such as autoencoders, generative adversarial network, recurrent neural network, and convolutional neural network. Afterwards, we present some benchmark datasets with descriptions and compare the performance of representing approaches to show the current working state of attack detection methods with deep learning structures. Finally, we summarize this paper and discuss some ways to improve the performance of attack detection under thoughts of utilizing deep learning structures. Yirui Wu, Dabao Wei, Jun Feng 0001 |
Secur. Commun. Networks | 1 |
| 2019 | A Text-Context-Aware CNN Network for Multi-oriented and Multi-language Scene Text DetectionabstractThe existing deep learning based state-of-theart scene text detection methods treat scene texts a type of general objects, or segment text regions directly. The latter category achieves remarkable detection results on arbitrary orientation and large aspect ratios of scene texts based on instance segmentation algorithms. However, due to the lack of context information with consideration of scene text unique characteristics, directly applying instance segmentation to text detection task is prone to result in low accuracy, especially producing false positive detection results. To ease this problem, we propose a novel text-context-aware scene text detection CNN structure, which appropriately encodes channel and spatial attention information to construct context-aware and discriminative feature map for multi-oriented and multi-language text detection tasks. With high representation ability of text context-aware feature map, the proposed instance segmentation based method can not only robustly detect multi-oriented and multi-language text from natural scene images, but also produce better text detection results by greatly reducing false positives. Experiments on ICDAR2015 and ICDAR2017-MLT datasets show that the proposed method has achieved superior performances in precision, recall and F-measure than most of the existing studies. Minglong Xue, Tong Lu 0002, Yirui Wu, Palaiahnakote Shivakumara |
ICDAR | 4 |
| 2019 | Hierarchical Bayesian Network Based Incremental Model for Flood Prediction
Yirui Wu, Weigang Xu, Qinghan Yu, Jun Feng 0001, Tong Lu 0002 |
MMM (1) | 1 |
| 2019 | An Automatic System for Generating Artificial Fake Character Images
Yisheng Yue, Palaiahnakote Shivakumara, Yirui Wu, Tong Lu 0002, Umapada Pal 0001 |
MMM (2) | 3 |
| 2019 | Deep spatiotemporal LSTM network with temporal pattern feature for 3D human action recognitionabstractAbstract With the rapid development of RGB‐D cameras and pose estimation techniques, action recognition based on three‐dimensional skeleton data has gained significant attention in the artificial intelligence community. In this paper, we incorporate temporal pattern descriptors of joint positions with the currently popular long short‐term memory (LSTM)–based learning scheme to obtain accurate and robust action recognition. Considering that actions are essentially formed by small subactions, we first utilize a two‐dimensional wavelet transform to extract temporal pattern descriptors in the frequency domain for each subaction. Afterward, we design a novel LSTM structure to extract deep features, which model a long‐term spatiotemporal correlation between body parts. Since temporal pattern descriptors and LSTM deep features can be regarded as multimodal representations for actions, we fuse them with an autoencoder network to achieve a more effective feature descriptor for action recognition. Experimental results on three challenging data sets with several comparative methods demonstrate the effectiveness of the proposed method for three‐dimensional action recognition. Yirui Wu, Lianglei Wei, Yucong Duan |
Comput. Intell. | 1 |
| 2018 | New COLD Feature Based Handwriting Analysis for Enthnicity/Nationality IdentificationabstractIdentifying crime for forensic investigating teams when crimes involve people of different nationals is challenging. This paper proposes a new method for ethnicity (nationality) identification based on Cloud of Line Distribution (COLD) features of handwriting components. The proposed method, at first, uses tangent angle of the contour pixels in each row and the mean of intensity values of each row for segmenting text lines. For segmented text lines, we use tangent angle and direction of base lines to remove rule lines in the image. We use polygonal approximation for finding dominant points for contours of edge components. Then the proposed method connects the nearest dominant points of every dominant point, which results in line segments of dominant point pairs. For each line segment, the proposed method estimates angle and length, which gives a point in polar domain. For all the line segments, the proposed method generates dense points in polar domain, which results in COLD distribution. As character component shapes change, according to nationals, the shape of the distribution changes. This observation is extracted based on distance from pixels of distribution to Principal Axis of the distribution. Then the features are subjected to an SVM classifier for identifying nationals. Experiments are conducted on a complex dataset, which show the proposed method is effective and outperforms the existing method. Sauradip Nag, Palaiahnakote Shivakumara, Yirui Wu, Umapada Pal 0001, Tong Lu 0002 |
ICFHR | 3 |
| 2018 | End-To-End Chromosome Karyotyping with Data Augmentation Using GANabstractClassifying human chromosomes from input cell images, i.e., karyotyping, requires domain expertise and quantity of manual effort to perform. In this paper, we propose an end-to-end chromosome karyotyping method, which can automatically detect, segment and classify chromosomes from cell images. During detection, we explore Extremal Regions (ER) to obtain chromosome candidates in input images. During segmentation, we segment overlapping chromosome candidates by approximating chromosome shapes with eclipses. In classification, we first propose Multiple Distribution Generative Advertising Network (MD-GAN) to effectively cover diverse data modes and generate more labeled samples for data augmentation. Then, we finetune pre-trained convolutional neural network (CNN) to classify chromosomes with samples generated by MD-GAN. We demonstrate the accuracy of the proposed end-to-end method in detecting, segmenting and classifying by experiments on a self-collected dataset. Experiments also prove data augmentation with MD-GAN could improve classification performance of CNN. Yirui Wu, Yisheng Yue, Tong Lu 0002 |
ICIP | 1 |
| 2018 | Em-SLAM: a Fast and Robust Monocular SLAM Method for Embedded SystemsabstractSimultaneous Localization and Mapping (SLAM) is difficult to deploy in the embedded systems due to its high computation cost and stable input requirements. Building on excellent algorithms of recent years, we present Em-SLAM, a monocular SLAM method which is fast and robust in the embedded system. We present Em-SLAM in three stages comprising initial pose estimation, iterative pose optimization and correspondences, and mapping with nearest frame queue. During the first stage, we perform stable initial pose estimation based on the matched ORB features extracted around the selected key points. Regarding initial pose and corresponding key points as input, the second stage of Em-SLAM iteratively optimizes these inputs values by tracking key points in the new frames. At the last stage, we firstly determine keyframes with the help of the proposed nearest frame queue and then design a greedy search algorithm to find matched ORB features between keyframes, which are adopted for compact and robust map reconstruction. Due to the special designs for the embedded systems, Em-SLAM demonstrates a high accurate and fast performance on the embedded system for all SLAM tasks: tracking, mapping and loop closing. We evaluate Em-SLAM on he most popular datasets by comparing with one latest SLAM method. Yirui Wu, Zhikai Li, Palaiahnakote Shivakumara, Tong Lu 0002 |
ICPR | 1 |
| 2018 | Context-Aware Attention LSTM Network for Flood PredictionabstractTo minimize the negative impacts brought by floods, researchers from pattern recognition community utilize artificial intelligence based methods to solve the problem of flood prediction. Inspired by the significant power of Long Short-Term Memory (LSTM) networks in modeling the dynamics and dependencies of sequential data, we intend to utilize LSTM networks to predict sequential flow rate values based on a set of collected flood factors. Since not all factors are informative for flood prediction and the irrelevant factors often bring a lot of noise, we need to pay more attention to the informative ones. However, original LSTM doesn't have strong attention capability. Hence we propose an context-aware attention LSTM (CA-LSTM) network for flood prediction, which is capable to selectively focus on informative factors. During training, the local context-aware attention model is constructed by learning probability distributions between flow rate and hidden output of each LSTM cell. During testing, the learned local attention model assign weights to adjust relations between input factors and predictions at all steps of LSTM network. We conduct experiments on a flood dataset with several comparative methods to demonstrate high accuracy of the proposed method and the effectiveness of the proposed context-aware attention model. Yirui Wu, Zhaoyang Liu 0001, Weigang Xu, Jun Feng 0001, Palaiahnakote Shivakumara, Tong Lu 0002 |
ICPR | 1 |
| 2018 | Local and Global Bayesian Network based Model for Flood PredictionabstractTo minimize the negative impacts brought by floods, researchers from pattern recognition community pay special attention to the problem of flood prediction by involving technologies of machine learning. In this paper, we propose to construct hierarchical Bayesian network to predict floods for small rivers, which appropriately embed hydrology expert knowledge for high rationality and robustness. We present the construction of the hierarchical Bayesian network in two stages comprising local and global network construction. During the local network construction, we firstly divide the river watershed into small local regions. Following the idea of a famous hydrology model - the Xinanjiang model, we establish the entities and connections of the local Bayesian network to represent the variables and physical processes of the Xinanjiang model, respectively. During the global network construction, intermediate variables for local regions, computed by the local Bayesian network, are coupled to offer an estimation for time-varying values of flow rate by proper inferences of the global network. At last, we propose to improve the output of Bayesian network by utilizing former flow rate values. We demonstrate the accuracy and robustness of the proposed method by conducting experiments on a collected dataset with several comparative methods. Yirui Wu, Weigang Xu, Jun Feng 0001, Palaiahnakote Shivakumara, Tong Lu 0002 |
ICPR | 1 |
| 2018 | Cloud of Line Distribution and Random Forest Based Text Detection from Natural/Video Scene Images
Wenhai Wang, Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002 |
MMM (2) | 2 |
| 2018 | A Novel 3D Human Action Recognition Framework for Video Content Analysis
Lianglei Wei, Yirui Wu, Wenhai Wang, Tong Lu 0002 |
MMM (1) | 2 |
| 2017 | A Robust Symmetry-Based Method for Scene/Video Text Detection through Neural NetworkabstractText detection in video/scene images has gained a significant attention in the field of image processing and document analysis due to the inherent challenges caused by variations in contrast, orientation, background, text type, font type, non-uniform illumination and so on. In this paper, we propose a novel text detection method to explore symmetry property and appearance features of text for improved accuracy and robustness. First, the proposed method explores Extremal Regions (ER) for detecting text candidates in images. Then we propose a novel feature named as Multi-domain Strokes Symmetry Histogram (MSSH) for each text candidate, which describes the inherent symmetry property of stroke pixel pairs in gray, gradient and frequency domains. Furthermore, deep convolutional features are extracted to describe the appearance for each text candidate. We further fuse them by Auto-Encoder network to define a more discriminative text descriptor for classification. Finally, the proposed method constructs text lines based on the classification results. We demonstrate the effectiveness and robustness detection results of our proposed method by testing on four different benchmark databases. Yirui Wu, Wenhai Wang, Palaiahnakote Shivakumara, Tong Lu 0002 |
ICDAR | 1 |
| 2017 | Multi-glimpse LSTM with color-depth feature fusion for human detectionabstractWith the development of depth cameras such as Kinect and Intel Realsense, RGB-D based human detection receives continuous research attention due to its usage in a variety of applications. In this paper, we propose a new Multi-Glimpse LSTM (MG-LSTM) network, in which multi-scale contextual information is sequentially integrated to promote the human detection performance. Furthermore, we propose a feature fusion strategy based on our MG-LSTM network to better incorporate the RGB and depth information. To the best of our knowledge, this is the first attempt to utilize LSTM structure for RGB-D based human detection. Our method achieves superior performance on two publicly available datasets. Hengduo Li, Jun Liu 0036, Guyue Zhang, Yirui Wu |
ICIP | 5 |
| 2017 | Deep-dense Conditional Random Fields for Object Co-segmentationabstractWe address the problem of object co-segmentation in images. Object co-segmentation aims to segment common objects in images and has promising applications in AI agents. We solve it by proposing a co-occurrence map, which measures how likely an image region belongs to an object and also appears in other images. The co-occurrence map of an image is calculated by combining two parts: objectness scores of image regions and similarity evidences from object proposals across images. We introduce a deep-dense conditional random field framework to infer co-occurrence maps. Both similarity metric and objectness measure are learned end-to-end in a single deep network. We evaluate our method on two benchmarks and achieve competitive performance. Ze-Huan Yuan, Tong Lu 0002, Yirui Wu |
IJCAI | 3 |
| 2017 | Visual Robotic Object Grasping Through Combining RGB-D Data and 3D Meshes
Yiyang Zhou, Wenhai Wang, Wenjie Guan, Yirui Wu, Heng Lai, Tong Lu 0002, Min Cai |
MMM (1) | 4 |
| 2017 | FreeScup: A Novel Platform for Assisting Sculpture Pose DesignabstractSculpture design is challenging due to its inherent difficulty in characterizing artworks quantitatively; thus, few works have been done to assist sculpture design in the past decades in the multimedia community. We have cooperated with several sculptors on analyzing styles of different artists consisting of Giacometti, Augeuste Rodin, Henry Moore, and Marino Marini from which we find pose editing plays an important role in sculpture design. Motivated by this, we present a novel platform that allows sculptors to edit virtual three-dimensional (3-D) sculptures by a free way. The proposed platform consists of three modules, namely,sculpture initialization,sculptor-sculpture mapping, andinteractive pose editing. In sculpture initialization, a virtual 3-D sculpture is first incrementally reconstructed from multiview images. Then, we define Laplace operator and its corresponding spectrum to describe the geometry information of the reconstructed sculpture. During sculptor–sculpture mapping, we apply spectral analysis on the low-frequency parts of the spectrum to search for candidate editing points on the surface of the sculpture. Next, body actions of the sculptor are captured by Kinect and further mapped onto editing points as a predefined configuration set. Finally, during interactive pose editing, a real-time Kinect-driven sculpture pose editing scheme is presented, which not only preserves geometry features of the sculpture but also allows instant changes of sculpture poses. We demonstrate that our platform successfully assists sculptors on real-time pose editing by comparing its performance with those of the existing sculpture assisting methods. Yirui Wu, Tong Lu 0002, Ze-Huan Yuan |
IEEE Trans. Multim. | 1 |
| 2016 | EvaToon: A novel graph matching system for evaluating cartoon drawingsabstractImitation cartoon drawing is an important skill for cartoonists, requiring quantity of efforts on practising and guidance. In this paper, we propose EvaToon, an imitated drawing evaluate system, which automatically assigns judging scores and marks improper drawing regions. With our system, cartoonists can practise and get guidance by themselves. We have cooperated with several experts on developing such an evaluation system. Based on their guide, we present EvaToon in two stages comprising cartoon drawings analyzing and similarity evaluating. During analyzing, we first locate contour pixels with high curvature as interest points and then extract multi-scale features around interest points to hierarchically describe shape. During evaluating, we first match interest points between original and imitated drawing based on distance of features. After matching, we construct a regression tree to map high dimensional difference of matching features to scores and marks based on quantity of manually evaluated training examples. Finally, our system matches an input imitated drawing with the original one and predicts its scores automatically. We demonstrate the accuracy of our EvaToon system in matching and predicting and prove the capability of describing shape of our proposed features by experiments on a collected dataset of imitated drawings. Yirui Wu, Xianli Zhou, Tong Lu 0002, Guo Mei, Linbi Sun |
ICPR | 1 |
| 2016 | Contour Restoration of Text Components for Recognition in Video/Scene ImagesabstractText recognition in video/natural scene images has gained significant attention in the field of image processing in many computer vision applications, which is much more challenging than recognition in plain background images. In this paper, we aim to restore complete character contours in video/scene images from gray values, in contrast to the conventional techniques that consider edge images/binary information as inputs for text detection and recognition. We explore and utilize the strengths of zero crossing points given by the Laplacian to identify stroke candidate pixels (SPC). For each SPC pair, we propose new symmetry features based on gradient magnitude and Fourier phase angles to identify probable stroke candidate pairs (PSCP). The same symmetry properties are proposed at the PSCP level to choose seed stroke candidate pairs (SSCP). Finally, an iterative algorithm is proposed for SSCP to restore complete character contours. Experimental results on benchmark databases, namely, the ICDAR family of video and natural scenes, Street View Data, and MSRA data sets, show that the proposed technique outperforms the existing techniques in terms of both quality measures and recognition rate. We also show that character contour restoration is effective for text detection in video and natural scene images. Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan, Michael Blumenstein, G. Hemantha Kumar 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | FreeScup: A novel platform for assisting sculpture pose designabstractSculpture design is challenging due to its inherent difficulty in characterizing an artwork quantitatively, and few works have been done to assist sculpture design. We present a novel platform to help sculptors in two stages, comprising automatic sculpture reconstruction and free spectral-based sculpture pose editing. During sculpture reconstruction, we co-segment a sculpture from real scene images of different views through a two-label MRF framework, aiming at performing sculpture reconstruction efficiently. During sculpture pose editing, we automatically extract candidate editing points on the sculpture by searching in the spectrums of Laplacian operator. After manually mapping body joints of a sculptor to particular editing points, we further construct a global Laplacian-based linear system by adopting the spectrums of Laplacian operator and using Kinect captured body motions for real time pose editing. The constructed system thus allows the sculptor to freely edit different kinds of sculpture artworks through Kinect. Experimental results demonstrate that our platform successfully assists sculptors in real-time pose editing. Yirui Wu, Tong Lu 0002, Ze-Huan Yuan |
ICME | 1 |
| 2015 | HIRM: A handle-independent reduced model for incremental mesh editing
Yirui Wu, Oscar Kin-Chung Au, Chiew-Lan Tai, Tong Lu 0002 |
Comput. Aided Geom. Des. | 1 |
| 2015 | A new ring radius transform-based thinning method for multi-oriented video characters
Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001 |
Int. J. Document Anal. Recognit. | 1 |
| 2011 | Multiclass object detection by combining local appearances and contextabstractIn this paper, we present a novel approach for multiclass object detection by combining local appearances and contextual constraints. We first construct a multiclass Hough forest of local patches, which can well deal with multiclass object deformations and local appearance variations, due to randomization and discrimination of the forest. Then, in the object hypothesis space, a new multiclass context model is proposed to capture relative location constraints, disambiguating appearance inputs in multiclass object detection. Finally, multiclass objects are detected with a greedy search algorithm efficiently. Experimental evaluations on two image data sets show that the combination of local appearances and context achieves state-of-the-art performance in multiclass object detection. Limin Wang 0002, Yirui Wu, Tong Lu 0002 |
ACM Multimedia | 2 |