VLDB 2026 Research / reviewers in the wild / expert
Zhi-Ping Shi 0002
dblp:143/4573-2 · also Zhiping Shi 0002
· DBLP profile ↗
91ranked-venue papers
3as first author
44since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 23 · 10 since 2021Software engineering, systems software and programming languages · 16 · 7 since 2021Systems, architecture and hardware · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Theory of computation · 5 · 1 first-author · 3 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Formal reasoning about Bernstein-Vazirani algorithm
Zhi-Ping Shi 0002, Shanyan Chen, Ximeng Li 0003 |
J. Log. Algebraic Methods Program. | 2 |
| 2026 | Region-Based Collaborative Caching With Joint Latency and Lifetime Optimization for Hybrid SMR-Flash StorageabstractShingled Magnetic Recording (SMR) disks have been improved significantly over traditional Hard Disk Drives (HDDs) in storage capacity and cost by a shingled-like structure that overlaps with adjacent tracks. However, the overall system performance is severely sacrificed due to the large portion of Read-Merge-Write (RMW) operations triggered by the non-sequential nature of SMR disks. To handle such performance degradation, Persistent Cache (PC) and built-in NAND flash cache are applied to absorb non-sequential writes. However, when the cache is full, triggering write-back operations extends the I/O response time. Additionally, the erase of NAND flash also compromises its lifetime.This paper aims to manage the NAND flash region and the original SMR band region uniformly instead of only adopting NAND flash as a cache layer. To achieve this, we propose a Region-based Co-optimized strategy named Multi-Regional Collaborative Management (MCM). Our approach segments I/O requests using a fast and efficient Bloom filter alongside a Band-based partitioning module. Additionally, the NAND flash memory region is managed based on the characteristics of the segmented data, enabling maximum control over frequent data updates in the NAND flash memory. The NAND flash cache will not serve as the first-level cache as before. Meanwhile, through Region-aware Wear-leveling (WL) and garbage collection (GC) strategies, the lifetime of NAND flash memory is extended, and cleaning efficiency is improved. In addition, further analysis of the impact of the wear-leveling parameters on RMWs and NAND flash lifetime under different workloads. According to the experimental results, compared with the typical Skylight (baseline), our scheme reduces the average response time and write amplification by 74.56% and 98.99%, respectively. At the same time, our approach also outperforms some state-of-the-art solutions, such as MU-RMW, RMW-F, and FC. Zhengang Chen, Zhi-Ping Shi 0002, Tianyu Wang 0009 |
IEEE Trans. Computers | 3 |
| 2025 | Strategy-Aware Liquidity for Account-Based Blockchains
Ximeng Li 0003, Sensen Chen, Qianying Zhang, Zhi-Ping Shi 0002 |
SETTA | 6 |
| 2025 | Formalization of robot collision detection method based on conformal geometric algebra
Shanyan Chen, Zhi-Ping Shi 0002, Ximeng Li 0003 |
Formal Methods Syst. Des. | 4 |
| 2025 | Joint 3-D Human Reconstruction and Hybrid Pose Self-Supervision for Action RecognitionabstractAction recognition is a promising task of identifying human activities in videos or images. Human movement is often accompanied by occlusion and blurring, resulting in many approximate behaviors that cannot be correctly recognized. Most methods directly employ local or motion cues to improve global features. Due to insufficient exploration of 3-D depth information, similar actions from the 2-D perspective still cannot be particularity distinguished. To tackle this challenge, this article proposes a novel action recognition approach termed RPS-Net, integrating 3-D human body reconstruction and hybrid pose self-supervision (HPSS). The designed 3-D reconstruction network is named CFFormer, which leveraging a context-fusion structure to generate 3-D meshes as auxiliary input. Among them, the mesh rotated 90° (M9) is exploited to provide multiperspective associative information, which employing spatial residual learning to enhance extra 3-D cues under various interpose variations. Meanwhile, the original images and meshes from the same perspective will perform HPSS with temporal encoding. It is responsible for capturing subtle differences in key point positions and structures across multiple dimensions. Extensive experiments demonstrate that our proposed method significantly surpasses most existing algorithms. And the Top-1 recognition accuracy on SSV2, HMDB51, and Olympic Sports datasets can reach 73.4%, 87.7%, and 96.2%, respectively. Hexin Wang, Luwei Li, Yuxuan Qiu, Zhi-Ping Shi 0002 |
IEEE Internet Things J. | 5 |
| 2025 | Bullet-Screen-Emoji Attack With Temporal Difference Noise for Video Action RecognitionabstractRecent studies have shown that video action recognition models are also vulnerable to fooling by adversarial samples. However, currently existing video attack methods usually require high computational overhead (e.g., they generate adversarial perturbations for all frames by default), and most of them are difficult to implement printable attacks in the physical world. To address the above issues, we devise a novel efficient and effective framework for video action recognition attack: Bullet-Screen-Emoji Attack with Temporal Difference Noise (BSE), a reinforcement learning-based black-box attack method that fools the model by simply generating adversarial bullet screens for key frame and scrolling them on clean video. The agent is optimized to make the optimal actions, i.e., searching key frame. Moreover, we introduce a simple and effective temporal difference noise to enhance the attack capability of the adversarial bullet screen and accelerate the convergence speed. Most importantly, BSE enables printable physical attacks. Extensive experiments show that our proposed BSE achieves promising attack performance on mainstream datasets (HMDB51, UCF101 and Kinetics-400) and in the physical world with high efficiency. Yongkang Zhang 0001, Jun Li 0072, Zhi-Ping Shi 0002, Jian Yang 0030, Kaixin Yang, Qiuyan Liang, Xianglong Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Advancing Visible-Infrared Person Re-Identification: Synergizing Visual-Textual Reasoning and Cross-Modal Feature AlignmentabstractVisible-infrared person re-identification (VI-ReID) is a critical cross-modality fine-grained classification task with significant implications for public safety and security applications. Existing VI-ReID methods primarily focus on extracting modality-invariant features for person retrieval. However, due to the inherent lack of texture information in infrared images, these modality-invariant features tend to emphasize global contexts. Consequently, individuals with similar silhouettes are often misidentified, posing potential risks to security systems and forensic investigations. To address this problem, this paper innovatively introduces natural language descriptions to learn the global-local contexts for VI-ReID. Specifically, we design a framework that jointly optimizes visible-infrared alignment plus (VIAP) and visual-textual reasoning (VTR), and introduces local-global joint measure (LJM) to enhance the metric, while proposing a human-LLM collaborative approach to incorporate textual descriptions into existing cross-modal person re-identification datasets. VIAP achieves cross-modal alignment between RGB and IR. It can explicitly utilize designed frequency-aware modality alignment and relationship-reinforced fusion to explore the potential of local cues in global features and modality-invariant information. VTR proposes pooling selection and dual-level reasoning mechanisms to force the image encoder to pay attention to significant regions based on textual descriptions. LJM proposes introducing local feature distances into the measure stage metric to enhance the relevance of matching using fine-grained information. Extensive experimental results on the popular SYSU-MM01 and RegDB datasets show that the proposed method significantly outperforms state-of-the-art approaches. The dataset is publicly available athttps://github.com/qyx596/vireid-caption. Yuxuan Qiu, Wei Song 0010, Jiawei Liu 0001, Zhi-Ping Shi 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Efficient Algorithms for Influence Maximization in Hypergraphs by Stratified SamplingabstractInfluence maximization (IM) aims to identify$k$vertices that maximize influence spread across a network. While well-studied in regular graphs, IM in hypergraphs presents unique challenges: conventional graph-based IM methods fail to capture hypergraph-specific structural properties, and existing hypergraph IM algorithms lack theoretical guarantees for time complexity and approximation quality. We address these gaps withHyperIM, a novel algorithm leveraging stratified sampling to generate random reversible reachable sets for efficient seed selection. Our key innovation lies in dual-perspective stratified sampling: assigning sampling probabilities based on vertex structural properties while applying size-adaptive sampling strategies. This approach optimizes seed selection, reduces computational costs, and provides rigorous theoretical guarantees. We further proposeHyperIM_BRR, which optimizes the required number of reversible reachable sets, achieving substantial cost reduction without sacrificing accuracy. Extensive experiments on real-world hypergraphs demonstrate that our algorithms significantly outperform state-of-the-art methods, delivering faster execution times and superior influence spread. Lingling Zhang 0006, Tiancheng Lu, Zhi-Ping Shi 0002, Zhiwei Zhang 0002, Ye Yuan 0001, Guoren Wang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Video Motion Blur Attack via Grad-Weighted and Discrete-Fusion Based Perturbation GenerationabstractRecent research has shown that deep learning networks are vulnerable to adversarial samples. Although there has been great progress in the study of adversarial attacks on images, there is relatively little research on adversarial attacks in the video domain, especially on intrinsic factors of videos, such as motion blur. In this paper, we devise a novel Grad-Weighted based One-step Motion Blur Attack (GWO-MBA) and a Discrete-Fusion based Progressive Motion Blur Attack (DFP-MBA) for video recognition, starting from the idea of integrating global adversarial attacks and adversarial patch attacks. Concretely, we use gradient maps to filter and weighted fusion motion blur (termed GWO-MBA) to achieve the attack that matches the motion information in the context of the video. In order to make the generated motion blur attack perturbations more natural and improve the attack success rate, we further introduce a progressive decomposition motion blur strategy (termed DFP-MBA) to progressively fuse more realistic discrete motion blurs. Besides, we propose an Aggressive Motion Blur Generation (AMBG), which generates natural motion blur based on the video context and has a better attack effect. The extensive experiments, on the HMDB-51 and UCF-101 datasets, demonstrate the effectiveness and superiority of our proposed attack method. In addition, the attack effectiveness of the mainstream denoising defense model and the deblur model further validates the robustness of our attack method. Guoming Wu, Jun Li 0072, Yangfan Xu, Zhi-Ping Shi 0002, Xianglong Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Hard-Sample Style Guided Patch Attack With RL-Enhanced Motion Pattern for Video RecognitionabstractAdversarial attacks have been extensively studied in the image field. In recent years, research has shown that video recognition models are also vulnerable to adversarial examples. However, most studies about adversarial attacks for video models have focused on perturbation-based methods, while patch-based black-box attacks have received less attention. Despite the excellent performance of perturbation-based attacks, these attacks are impractical for real-world implementation. Most existing patch-based black-box attacks require occluding larger areas and performing more queries to the target model. In this paper, we propose a hard-sample style guided patch attack with reinforcement learning (RL) enhanced motion patterns for video recognition (HSPA). Specifically, we utilize the style features of video hard samples and transfer their multi-dimensional style features to images to obtain a texture patch set. Then we use reinforcement learning to locate the patch coordinates and obtain a specific adversarial motion pattern of the patch to successfully perform an effective attack on a video recognition model in both the spatial and temporal dimensions. Our experiments on three widely-used video action recognition models (C3D, LRCN, and TDN) and two mainstream datasets (UCF-101 and HMDB-51) demonstrate the superior performance of our method compared to other state-of-the-art approaches. Jian Yang 0030, Jun Li 0072, Yunong Cai, Guoming Wu, Zhi-Ping Shi 0002, Chaodong Tan, Xianglong Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | SQLPass: A Semantic Effective Fuzzing Method for DBMSabstractFuzzing, as an effective method for software system defect testing and bug mining, utilizes random data generated by mutation to execute programs to trigger potential bugs in the tested program. However, due to the strict syntax and semantic checks in DBMSs, existing mutation-based testing methods are difficult to generate test cases with correct syntax and semantics. To ensure the syntax and semantic correctness of the test case generated by mutation, this paper proposes the concepts of weak semantic correlation nodes and semantic relationship tables. At the same time, a set of mutation operators for weak semantic correlation nodes in the syntax tree is designed, and each time the “optimal” mutation operator is selected to replace, delete, and insert the target node subtree in the test case syntax tree. In response to semantic errors during the mutation, this paper checks and corrects them through a pre-extracted semantic relationship table, further ensuring the correctness of the semantics of the test cases generated by the mutation. This paper utilizes this semantic effective fuzzing method for DBMSs to implement a new fuzzing framework for DBMSs, SQLPass. We evaluated SQLPass on four popular DBMSs: SQLite3, MySQL, MariaDB and PostgreSQL. In our experiment, SQLPass achieves 5.7%-94.2% higher semantic correctness than state-of-the-art tools, and explores 1.3%-52% more code coverage than other tools. Notably, it discovered four unknown bugs on SQLite3 and MariaDB and has submitted them to the vendor for confirmation. Yixiao Yang, Zhi-Ping Shi 0002, Rui Wang 0024 |
COMPSAC | 4 |
| 2024 | Refinement Verification of OS Services based on a Verified Preemptive MicrokernelabstractAbstract An OS microkernel can be extended by implementing services upon it. A service could introduce an object that references a kernel object, and implement a group of functions that invokes the functions for manipulating the kernel object. We consider the scenario where the microkernel has been verified with machine-checkable proofs, while the services remain to be verified. Moreover, the verification of the microkernel is not performed with the verification of subsequent extension in mind. We address the problem of how to build sufficiently on the verification results for the microkernel, in achieving the verification of the services. Our methodology consists of enhancements to the verification framework for the microkernel, and the design of invariants for establishing the connection between the service-level objects and the kernel-level objects. Using the methodology, we have conducted a substantial formal verification of a group of services extending the inter-task communication functionalities of the preemptive microkernel $$\mu \!\!\text{ C }\!\!\text{/ }\!\!\!\text{ OS-II }$$ μ C / OS-II . Our verification uncovers dormant bugs and provides a level of correctness assurance for the services that is above what is achievable through extensive testing. Ximeng Li 0003, Shanyan Chen, Qianying Zhang, Zhi-Ping Shi 0002 |
FASE | 6 |
| 2024 | Augur: Predictive and Adaptive Data Management in SSD-SMR Hybrid Storage System Using Reinforcement LearningabstractIn SSD-SMR hybrid storage systems, the Persistent Cache (PC) cleaning process in Shingled Magnetic Recording (SMR) disks involves frequent Read-Modify-Write (RMW) operations, significantly increasing system tail latency. Although traditional Q-learning methods alleviate this issue to some extent, the expansion of state space results in an increase in the dimensions of the Q-Table, leading to significant rises in computational and storage costs, thus limiting its ability to handle complex problems. In this paper, we propose an intelligent data prediction management strategy named Augur, based on Deep Reinforcement Learning (DRL), that predicts nodes where high latency is imminent and directs the agent named Pythia to timely redirect the flow of suspect data, effectively reducing RMW operations in SMR. We implement our technique on a real SSD-SMR hybrid storage system. Experimental results show that, compared to Q-learning and Skylight, our method reduces the system tail latency at the 99.9thpercentile by 23.61% and 58.35%, and reduces average request latency by 22.54% and 47.56%, respectively. Zhengang Chen, Zhi-Ping Shi 0002 |
HPCC | 4 |
| 2024 | Joint Visual-Textual Reasoning and Visible-Infrared Modality Alignment for Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a cross-modality fine-grained classification task. Existing approaches for VI-ReID mainly explore modality-invariant features for person retrieval. However, modality-invariant features pay more attention to global contexts, due to the lack of texture information in infrared images. This leads to a person with similar silhouette often being misidentified. Targeting this problem, this paper innovatively introduces natural language specification to learn global-local contexts for VI-ReID. Specifically, our framework jointly optimizes visible-infrared alignment (VIA) and visual-textual reasoning (VTR). VIA achieves cross-modal between RGB and IR. It can explicitly utilize designed modality-guided alignment and relationship-reinforced fusion to explore the potential of local cues in global features. VTR proposes the pooling selection and dual-level reasoning mechanisms to force the image encoder to pay attention to significant regions based on textual descriptions. Extensive experimental results on the popular SYSU-MM01 and RegDB datasets show that the proposed method significantly outperforms state-of-the-art approaches. Yuxuan Qiu, Wei Song 0010, Jiawei Liu 0001, Zhi-Ping Shi 0002 |
ICME | 5 |
| 2024 | A survey on person and vehicle re-identificationabstractAbstract Person/vehicle re‐identification aims to use technologies such as cross‐camera retrieval to associate the same person (same vehicle) in the surveillance videos at different locations, different times, and images captured by different cameras so as to achieve cross‐surveillance image matching, person retrieval and trajectory tracking. It plays an extremely important role in the fields of intelligent security, criminal investigation etc. In recent years, the rapid development of deep learning technology has significantly propelled the advancement of re‐identification (Re‐ID) technology. An increasing number of technical methods have emerged, aiming to enhance Re‐ID performance. This paper summarises four popular research areas in the current field of re‐identification, focusing on the current research hotspots. These areas include the multi‐task learning domain, the generalisation learning domain, the cross‐modality domain, and the optimisation learning domain. Specifically, the paper analyses various challenges faced within these domains and elaborates on different deep learning frameworks and networks that address these challenges. A comparative analysis of re‐identification tasks from various classification perspectives is provided, introducing mainstream research directions and current achievements. Finally, insights into future development trends are presented. Zhaofa Wang, Zhi-Ping Shi 0002, Qichuan Geng |
IET Comput. Vis. | 3 |
| 2024 | Cross-project concurrency bug prediction using domain-adversarial neural network
Fangyun Qin, Zheng Zheng 0001, Yulei Sui, Siqian Gong, Zhi-Ping Shi 0002, Kishor S. Trivedi |
J. Syst. Softw. | 5 |
| 2024 | Diffusion Patch Attack With Spatial-Temporal Cross-Evolution for Video RecognitionabstractDeep neural networks (DNNs) have demonstrated excellent performance across various domains. However, recent studies have shown that deep neural networks are vulnerable to adversarial examples, including DNN-based video action recognition models. While much of the existing research on adversarial attacks against video models focuses on perturbation-based attacks, there is limited research on patch-based black-box attacks. Existing patch-based attack algorithms suffer from the problem of a large search space of optimization algorithms and use patches with simple content, leading to suboptimal attack performance or requiring a large number of queries. To address these challenges, we propose the “Diffusion Patch Attack (DPA) with Spatial-Temporal Cross-Evolution (STCE) for Video Recognition,” a novel approach that integrates the excellent properties of the diffusion model into video black-box adversarial attacks for the first time. This integration significantly narrows the parameter search space while enhancing the adversarial content of patches. Moreover, we introduce the spatial-temporal cross-evolutionary algorithm to adapt to the narrowed search space. Specifically, we separate the spatial and temporal parameters and then employ an alternate evolutionary strategy for each parameter type. Extensive experiments conducted on three widely used video action recognition models (C3D, NL, and TPN) and two benchmark datasets (UCF-101 and HMDB-51) demonstrate the superior performance of our approach compared to other state-of-the-art black-box patch attack algorithms. Jian Yang 0030, Zhiyu Guan, Jun Li 0072, Zhi-Ping Shi 0002, Xianglong Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Adaptive Impulsive Consensus of Nonlinear Multiagent Systems With Limited Bandwidth Under Uncertain Deception AttacksabstractAt present, the dynamic encoding–decoding scheme is utilized in the impulsive consensus control of multiagent systems (MASs) to solve the limited bandwidth problem. However, the unknown nonlinear dynamics and deception attacks will generate some uncertainties inevitably in the encoding–decoding, which may cause the quantizer saturation and then influence the consensus performance. Therefore, the impulsive consensus control problem of uncertain nonlinear MASs based on encoding–decoding under deception attacks is investigated in this article. To address the system uncertainty, an adaptive algorithm with neural networks is designed. Under the proposed estimator for each follower, the designed adaptive law for every follower only needs the information from its own sensor and estimator instead of the quantized information from its neighbors. Then a more general scenario of uncertain deception attack is considered, in which the uncertain deception attacks can occur in the both parts of hybrid impulsive control protocol. An attack observer is introduced to handle this more complex attack by compensating impacts of deception attacks. Next the sufficient conditions for secure consensus of MASs with limited bandwidth are derived. Moreover, although some uncertainties exist in the encoding–decoding, the quantizer saturation can be eliminated by adjusting the parameters of controller. Finally, the validity of given theorems is demonstrated by the simulation experiments. Chang-E Ren, Zhi-Ping Shi 0002, C. L. Philip Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Efficiently estimating node influence through group sampling over large graphs
Lingling Zhang 0006, Zhi-Ping Shi 0002, Zhiwei Zhang 0002, Ye Yuan 0001, Guoren Wang |
World Wide Web (WWW) | 2 |
| 2023 | Formal Verification of Interrupt Isolation for the TrustZone-based TEEabstractARM TrustZone is a hardware security technology commonly used to implement the Trusted Execution Environment (TEE), which is a hardware-based isolated execution environment in which security-critical software can execute without interference and is widely applied to embedded systems. Isolation of interrupts in the two security domains of TrustZone is an important security mechanism that plays a vital role in protecting the execution of the TEE. This paper introduces a formal modeling and verification approach for the interrupt isolation mechanism of the TEE based on TrustZone. We propose a model formalizing the TrustZone-based TEE system at the instruction level, which focuses on modeling system behaviors of interrupt handling and captures crucial instructions related to interrupt isolation. We formally define three types of properties for the model: correctness, safety, and information flow security. The verification results in the theorem prover Isabelle/HOL show that the model satisfies the above properties, which indicates that the interrupt isolation mechanism of TrustZone correctly enforces the isolation of interrupts and protects the TEE from being interfered with non-secure interrupts. Qianying Zhang, Ximeng Li 0003, Zhi-Ping Shi 0002 |
APSEC | 6 |
| 2023 | Region-based Flash Caching with Joint Latency and Lifetime Optimization in Hybrid SMR Storage SystemsabstractThe frequent Read-Modify-Write operations (RMWs) in Shingled Magnetic Recording (SMR) disks severely degrade the random write performance of the system. Although the adoption of persistent cache (PC) and built-in NAND flash cache alleviates some of the RMWs, when the cache is full, the triggered write-back operations still prolong I/O response time and the erasure of NAND flash also sacrifices its lifetime. In this paper, we propose a Region-based Co-optimized strategy named Multi-Regional Collaborative Management (MCM) to optimize the average response time by separately managing sequential/random and hot/cold data and extend the NAND flash lifetime by a region-aware wear-leveling strategy. The experimental results show that our MCM reduces 71 % of the average response time and 96% of RMWs on average compared with the Skylight (baseline). For the comparison with the state-of-art flash-based cache (FC) approach, we can still save the average response time and flash erase operations by 17.2 % and 33.32 %, respectively. Zhengang Chen, Zhi-Ping Shi 0002, Tianyu Wang 0009 |
DATE | 3 |
| 2023 | Exploiting 3D Human Recovery for Action Recognition with Spatio-Temporal Bifurcation FusionabstractAction recognition utilizes information in images or videos to analyze and classify human behaviors. The existing methods usually exploit 2D pose to improve classification features. Due to the lack of 3D cues, some approximate behaviors in 2D perspective cannot be recognized. In this paper, we propose a novel action recognition method with 3D human recovery and spatio-temporal bifurcations fusion. It consists of 3D spatial branch and 2D temporal branch. The 3D spatial branch exploits overlapping human models from 3D recovery to learn 3D recognition feature. The 2D temporal branch utilizes channel attention mechanism to enhance dynamic associated feature. Two types of features are fused by adaptive weights module, in order to improve the recognition of approximate behavior in 2D perspective. Extensive experiments show that this proposed method outperforms most of the state-of-the-art methods on the Olympic Sport, Diving48 and Human3.6M datasets. Qichuan Geng, Zhi-Ping Shi 0002 |
ICASSP | 4 |
| 2023 | RDEPD: Re-Exploring Depth Estimation for Pedestrian DetectionabstractPedestrian detection is a fundamental task in computer vision field. It remains challenging due to perspective affine and object occlusion. To alleviate these problems, this paper re-thinks the assistance of depth estimation, and then proposes a novel pedestrian detection method named RDEPD. The method consists of a data augmentation module based on depth (DamD) and a detection framework with learnable attention module based on depth(LAmD) and self-suitable NMS(S2NMS). DamD provides hard examples with mosaics at the depth gradient, which can improve the generalization ability. The detection framework pays more attention on these instances with occlusion or perspective affine by LamD and S2NMS. Where LAmD is responsible for integrating depth and RGB cues to guide localization, and S2NMS exploits every possible predicting box to improve the detecting precision. Extensive experimental results demonstrate that the proposed RDEPD significantly outperforms most state-of-the-art methods on the authoritative MOT17det, CrowdHuman, and Citypersons datasets, especially for occlusion situations. Yifei Pei, Zhi-Ping Shi 0002, Qichuan Geng, Zhaofa Wang, Yongkang Zhang 0001 |
ICIP | 2 |
| 2023 | Parallelizable Simple Recurrent Units with Hierarchical Memory
Yu Qiao 0001, Hengyi Zhang, Yuan Tian 0017, Zhenzhou Shao, Zhi-Ping Shi 0002 |
ICONIP (15) | 7 |
| 2023 | 6D Object Pose Estimation with Attention Aware Bi-gated Fusion
Laichao Wang, Weiding Lu, Yuan Tian 0017, Zhenzhou Shao, Zhi-Ping Shi 0002 |
ICONIP (2) | 6 |
| 2023 | Combat data shift in few-shot learning with knowledge graph
Yongchun Zhu, Fuzhen Zhuang, Xiangliang Zhang 0001, Zhi-Ping Shi 0002, Juan Cao 0001, Qing He 0003 |
Frontiers Comput. Sci. | 5 |
| 2023 | Imperceptible Adversarial Attack With Multigranular Spatiotemporal Attention for Video Action RecognitionabstractIn recent years, the application of video Internet of Things (IoT) in various cities and public places has brought unprecedented opportunities to the security field and achieved great success. However, the latest research shows that video recognition models are also vulnerable to adversarial examples, but adversarial examples based on physical attacks are easily detected by humans, making it difficult to pass human review. To address this problem, in this article, we propose to introduce a novel multigranular spatiotemporal attention network (MSANet), which can attack the video action recognition models imperceptibly. Specifically, to exploit video motion information more effectively and to reduce the detectability of attack perturbations, we design a multiplexed spatiotemporal attention module to select and enhance spatial regions and temporal frames at coarse-grained and fine-grained levels, respectively, thus maintaining a certain degree of smoothness while reducing the perturbation size and avoiding attacking overfitting. In addition, our proposed MSANet achieves imperceptible perturbations to video sequences through alternate iterative optimization combined with the PGD attack mechanism. extended experimental results on two different models (e.g., TDN and TSM) and two widely used data sets [HMDB-51 (Kuehne et al., 2011) and UCF-101 (Soomro et al., 2012)], compared to the state-of-the-art model, demonstrate the effectiveness of our devised video action recognition attack approach. Guoming Wu, Yangfan Xu, Jun Li 0072, Zhi-Ping Shi 0002, Xianglong Liu 0001 |
IEEE Internet Things J. | 4 |
| 2023 | Temporal Transformer Networks With Self-Supervision for Action RecognitionabstractIn recent years, Internet of Things (IoT) has made rapid development, and IoT devices are developing towards intelligence. IoT terminal devices represented by surveillance cameras play an irreplaceable role in modern society, most of them are integrated with video action recognition and other intelligent functions. However, their performance is somewhat affected by the limitation of computing resources of IoT terminal devices and the lack of long-range non-linear temporal relation modeling and reverse motion information modeling. To address this urgent problem, we introduce a startling Temporal Transformer Network with Self-supervision (TTSN). Our high-performance TTSN mainly consists of a temporal transformer module and a temporal sequence self-supervision module. Concisely speaking, we utilize the efficient temporal transformer module to model the non-linear temporal dependencies among non-local frames, which significantly enhances complex motion feature representations. The temporal sequence self-supervision module we employ unprecedentedly adopts the streamlined strategy of “random batch random channel” to reverse the sequence of video frames, allowing robust extractions of motion information representation from inversed temporal dimensions and improving the generalization capability of the model. Extensive experiments on three widely used datasets (HMDB51, UCF101, and Something-something V1) have conclusively demonstrated that our proposed TTSN is promising as it successfully achieves state-of-the-art performance for video action recognition. Our TTSN provides the possibility for its application in IoT scenarios due to its computational complexity and high performance. With the rapid development of the Internet of Things (IoT), more and more data is being disseminated in the form of video, which also puts new requirements on the understanding and modeling of video data. In recent years, 2D Convolutional Networks-based video action recognition has encouragingly gained wide popularity; However, constrained by the lack of long-range non-linear temporal relation modeling and reverse motion information modeling, the performance of existing models is, therefore, undercut seriously. To address this urgent problem, we introduce a startling Temporal Transformer Network with Self-supervision (TTSN). Our high-performance TTSN mainly consists of a temporal transformer module and a temporal sequence self-supervision module. Concisely speaking, we utilize the efficient temporal transformer module to model the non-linear temporal dependencies among non-local frames, which significantly enhances complex motion feature representations. The temporal sequence self-supervision module we employ unprecedentedly adopts the streamlined strategy of “random batch random channel” to reverse the sequence of video frames, allowing robust extractions of motion information representation from inversed temporal dimensions and improving the generalization capability of the model. Extensive experiments on three widely used datasets (HMDB51, UCF101, and Something-something V1) have conclusively demonstrated that our proposed TTSN is promising as it successfully achieves state-of-the-art performance for video action recognition. As a result, our work provides new attention and self-supervised algorithm for processing video data in IoT. Yongkang Zhang 0001, Jun Li 0072, Guoming Wu, Zhi-Ping Shi 0002, Zhaoxun Liu, Zizhang Wu, Xianglong Liu 0001 |
IEEE Internet Things J. | 6 |
| 2023 | Formalization of the inverse kinematics of three-fingered dexterous hand
Shanyan Chen, Zhi-Ping Shi 0002, Ximeng Li 0003, Jingzhi Zhang |
J. Log. Algebraic Methods Program. | 4 |
| 2023 | A unified proof technique for verifying program correctness with big-step semantics
Ximeng Li 0003, Qianying Zhang, Zhi-Ping Shi 0002 |
J. Syst. Archit. | 4 |
| 2023 | Spatio-Temporal Adaptive Network With Bidirectional Temporal Difference for Action RecognitionabstractAction Recognition is a fundamental task in computer vision field, with a wide range of applications in autonomous driving, security monitoring, etc. However, previous action recognition approaches usually suffer from the inappropriate spatio-temporal modeling or high computational consumption (e.g., 3D CNN). In this paper, we propose a novel Spatio-Temporal Adaptive Network (STANet) with bidirectional temporal difference, consisting of a Temporal Adaptive module (TA) and a Spatial Adaptive (SA) module, to sufficiently extract the crucial motion information and model the spatial pivotal appearance information from both forward and backward perspectives, respectively. Specifically, the Temporal Adaptive module uses bidirectional temporal differences to learn valuable motion trends and balance the static semantics and dynamic motion for a certain action during information fusion; while the Spatial Adaptive module uses the bidirectional temporal difference to obtain the spatio-channel attention to stress the discriminative position-relevant and semantic-relevant appearance features. Extensive experiments conducted on widely-used action recognition benchmarks UCF-101, HMDB-51, Something-Something V1, and Kinetics-400 prove the effectiveness of the proposed methods compared to other state-of-the-art approaches. Zhilei Li, Jun Li 0072, Yuqing Ma, Rui Wang 0024, Zhi-Ping Shi 0002, Yifu Ding 0001, Xianglong Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | MuLHiTA: A Novel Multiclass Classification Framework With Multibranch LSTM and Hierarchical Temporal Attention for Early Detection of Mental StressabstractMental stress is an increasingly common psychological issue leading to diseases such as depression, addiction, and heart attack. In this study, an early detection framework based on electroencephalogram (EEG) data is developed for reducing the risk of these diseases. In existing frameworks, signals are often segmented into smaller sections prior to being input to a deep neural network. However, this approach ignores the fundamental nature of EEG signals as a carrier of valuable information (e.g., the integrity of frequency and phase, and temporal fluctuations of EEG components). As such, this type of segmenting may lead to information loss and a failure to effectively identify mental stress levels. Thus, we propose a novel multiclass classification framework termed multibranch LSTM and hierarchical temporal attention (MuLHiTA) for the early identification of mental stress levels. It specifically focuses on not only intraslice (within each slice) but also interslice (between different slices) samples in parallel. This was achieved by including two complementary branches, each of which integrated a specifically designed attention module into a bidirectional long short-term memory (BLSTM) network, enabling extraction of the most discriminative features from interslice and intraslice EEG signals simultaneously. The outputs of attention modules were then summed to obtain a feature representation that contributes to reduce overfitting and more effective multiclass classification. In addition, electrode positions were optimized using neural activity areas under high-stress conditions, thereby reducing computational costs by minimizing the number of critical electrodes. MuLHiTA was evaluated across one private [Montreal imaging stress task (MIST)] and two publicly available EEG datasets [EEG during mental arithmetic tasks (DMAT) and Simultaneous task EEG workload (STEW)]. These were divided into training and test sets using an 8:2 ratio, and the training data were further divided into training and validation sets using a fivefold cross-validation (CV) method, in which the model with the highest accuracy among the five was selected. The model was trained once more with the full training set, and the test data were then used to evaluate its performance. This approach achieved average classification accuracies of 93.58%, 91.80%, and 99.71% for the MIST, STEW, and DMAT datasets, respectively. Experimental results showed MuLHiTA was superior to state-of-the-art algorithms, including EEGNet, BLSTM, EEGLearn, convolutional neural network (CNN)-long short-term memory (LSTM), and convolutional recurrent attention model (CRAM), for multiclass classification. This demonstrates the viability of MuLHiTA for the early detection of mental stress. Likun Xia, Ziheng Guo, Jinhong Ding, Ming Ma 0004, Guoxi Gan, Yehan Xu, Jingyu Luo, Zhi-Ping Shi 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 11 |
| 2022 | Cross-Guided Feature Fusion with Intra-Modality Reweighting for Multi-Spectral Pedestrian DetectionabstractMulti-spectral pedestrian detection has gained extensive attention over the past decade. To alleviate the problem of modality imbalance in the multi-spectral tasks, a novel cross-guided feature fusion network based on the auto-encoder framework is proposed using RGB-thermal image pairs as inputs. To obtain the complementary features, a cross-guided loss is designed, so that the output images are balanced with both modalities in an unsupervised manner. An intra-modality reweighting module is implemented to filter the redundant features before the fusion. Finally, YOLOv3 is chosen as the detector fed by the fused features. The proposed method is verified using the public KAIST and VOT-RGBT datasets. Experimental results demonstrate that the proposed method can outperform the state-of-the-art methods, the miss rate of pedestrian detection reaches 48.57% and 4.52% using KAIST and VOT-RGBT datasets, respectively. Zhenzhou Shao, Ying Qu 0001, Jun Zhang 0031, Zhi-Ping Shi 0002 |
ICPR | 8 |
| 2022 | TMN: Temporal-guided Multiattention Network for Action Recognitionabstract2D convolutional neural network, due to its low computational complexity and fast recognition speed, has attracted more and more attention from researchers in the field of video action recognition. Temporal shift and temporal differential, have made tremendous progress, but the lack of crucial spatiotemporal attention mechanism has led to huge performance loss. To address this issue, we propose a Temporal-guided Multiattention Network (TMN), which fully excavate and fuse spatio-temporal attention information for effective video action recognition. Concretely, the multi-attention module squeezes and expands spatio-temporal features to achieve weighting of corresponding regions for video in spatio-temporal dimensions, while the adaptive temporal guidance module imports temporal guiding signal to the spatial attention and re-weight the global temporal attention to accomplish the accurate temporal modeling. Extensive experiments and analyses show that our proposed temporal-guided multiattention network can achieve state-of-the-art promising video action recognition performance on the widely used benchmarks (HMDB51, UCF101 and Something-Something V1). Yongkang Zhang 0001, Guoming Wu, Yangfan Xu, Zhi-Ping Shi 0002, Jun Li 0072 |
ICPR | 5 |
| 2022 | Making Information Hiding Effective AgainabstractInformation hiding (IH) is an important building block for many defenses against code reuse attacks, such as code-pointer integrity (CPI), control-flow integrity (CFI) and fine-grained code (re-)randomization, because of its effectiveness and performance. It employs randomization to probabilistically “hide” sensitive memory areas, called safe areas, from attackers and ensures their addresses are not leaked by any pointers directly. These defenses used safe areas to protect their critical data, such as jump targets and randomization secrets. However, recent works have shown that IH is vulnerable to various attacks. In this article, we propose a new IH technique called SafeHidden. It continuously re-randomizes the locations of safe areas and thus prevents the attackers from probing and inferring the memory layout to find its location. A new thread-private memory mechanism is proposed to isolate the thread-local safe areas and prevent adversaries from reducing the randomization entropy. It also randomizes the safe areas after the TLB misses to prevent attackers from inferring the address of safe areas using cache side-channels. Existing IH-based defenses can utilize SafeHidden directly without any change. Our experiments show that SafeHidden not only prevents existing attacks effectively but also incurs low performance overhead. Zhe Wang 0017, Chenggang Wu 0002, Yinqian Zhang, Bowen Tang 0001, Pen-Chung Yew, Mengyao Xie, Yuanming Lai, Yan Kang 0002, Yueqiang Cheng, Zhi-Ping Shi 0002 |
IEEE Trans. Dependable Secur. Comput. | 10 |
| 2022 | Two-Branch Attention Network via Efficient Semantic Coupling for One-Shot LearningabstractOver the past few years, Convolutional Neural Networks (CNNs) have achieved remarkable advancement for the tasks of one-shot image classification. However, the lack of effective attention modeling has limited its performance. In this paper, we propose a Two-branch (Content-aware and Position-aware) Attention (CPA) Network via an Efficient Semantic Coupling module for attention modeling. Specifically, we harness content-aware attention to model the characteristic features (e.g., color, shape, texture) as well as position-aware attention to model the spatial position weights. In addition, we exploit support images to improve the learning of attention for the query images. Similarly, we also use query images to enhance the attention model of the support set. Furthermore, we design a local-global optimizing framework that further improves the recognition accuracy. The extensive experiments on four common datasets (miniImageNet, tieredImageNet, CUB-200-2011, CIFAR-FS) with three popular networks (DPGN, RelationNet and IFSL) demonstrate that our devised CPA module equipped with local-global Two-stream framework (CPAT) can achieve state-of-the-art performance, with a significant improvement in accuracy of 3.16% on CUB-200-2011 in particular. Jun Li 0072, Duorui Wang, Xianglong Liu 0001, Zhi-Ping Shi 0002, Meng Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | DDSAS: Dynamic and Differentiable Space-Architecture SearchabstractNeural Architecture Search (NAS) has made remarkable progress in automatically designing neural networks. However, existing differentiable NAS and stochastic NAS methods are either biased towards exploitation and thus may converge to a local minimum, or biased towards exploration and thus converge slowly. In this work, we propose a Dynamic and Differentiable Space-Architecture Search (DDSAS) method to address the exploration-exploitation dilemma. DDSAS dynamically samples space, searches architectures in the sampled subspace with gradient descent, and leverages the Upper Confidence Bound (UCB) to balance exploitation and exploration. The whole search space is elastic, offering flexibility to evolve and to consider resource constraints. Experiments on image classification datasets demonstrate that with only 4GB memory and 3 hours for searching, DDSAS achieves 2.39% test error on CIFAR10, 16.26% test error on CIFAR100, and 23.9% test error when transferring to ImageNet. When directly searching on ImageNet, DDSAS achieves comparable accuracy with more than 6.5 times speedup over state-of-the-art methods. The source codes are available at https://github.com/xingxing-123/DDSAS. Longxing Yang, Yu Hu 0001, Shun Lu 0001, Jilin Mei, Yiming Zeng 0003, Zhi-Ping Shi 0002, Yinhe Han 0001, Xiaowei Li 0001 |
ACML | 7 |
| 2021 | Learning Attention Map for 3D Human Recovery from a Single RGB Image
Jun Li 0072, Zhi-Ping Shi 0002 |
BMVC | 4 |
| 2021 | Fast and Unsupervised Non-Local Feature Learning for Direct Volume Rendering of 3D Medical ImagesabstractTo improve the efficiency of medical visualization for computer aided surgery, we propose a fast and unsupervised 3D-CNN based non-local feature learning network. The proposed network consists of an encoder structure and a decoder structure. The encoder of the network projects the cube into a high-dimensional feature space, and the decoder of the network reconstructs the cube from the feature space. The decoder of the network serves as a dictionary shared by the cube to enforce the features for similar parts to be similar although they may distribute at disjointed locations. With such structures, the network is able to extract non-local features of the entire data. Moreover, a sparse constraint is incorporated into the network to increase the discriminative of the non-local features. Then the extracted non-local features of each voxel are fused with the corresponding position matrix and Hessian matrix for the voxel classification using Random Forest. Finally, a multidimensional transfer function is designed to enable the volume rendering. Experimental results demonstrate that the proposed method outperforms the state-of-the-art methods with much less training time. Xinmei Fu, Zhenzhou Shao, Ying Qu 0001, Yibo Zou, Zhi-Ping Shi 0002, Jindong Tan |
IROS | 6 |
| 2021 | Reasoning About Iteration and Recursion Uniformly Based on Big-Step Semantics
Ximeng Li 0003, Qianying Zhang, Zhi-Ping Shi 0002 |
SETTA | 4 |
| 2021 | Adaptive Event-Triggered Control for Nonlinear Multi-Agent Systems with State Time Delay and Unknown External DisturbanceabstractThis paper solves the adaptive event-triggered consensus problem for nonlinear multi-agent systems with state time delay and unknown external disturbance. First, the disturbance is estimated by designing a novel adaptive disturbance observer. Next, Lyapunov-Krasovskii functional is utilized to deal with the state delay. Finally, in accordance with the designed event-triggered condition, an adaptive controller is proposed, in which neural networks are applied to approximate the unknown nonlinearity. According to Lyapunov stability theory, it can be proved that nonlinear multi-agent systems can reach consensus under the proposed controller. The simulation validates the theoretical results. Quanxin Fu, Chang-E Ren, Jiaang Zhang, Zhi-Ping Shi 0002 |
SMC | 4 |
| 2021 | Semi-Supervised Domain Adaption Classifier via Broad Learning SystemabstractBroad Learning System (BLS) is effective and efficient in dealing with various machine learning problems. When constructing a BLS based classifier with domain adaption capability, the mapped features and enhancement nodes are determined by random parameters. Thus, there will be discrepancy between the hidden layer features of the source domain and target domain. Therefore, we propose a new semi-supervised classifier with domain adaption capability based on BLS. Firstly, in order to reduce the discrepancy between the hidden layer features of domains, our method aligns the second-order statistics of mapped features due to the fact that enhancement nodes are generated by mapped features. Further, when learning the classifier through the aligned features, we embed balanced distribution adaptation to improve domain adaption capability of the classifier. Experiments on benchmark datasets demonstrate our method has better classification accuracy than some existences. Zehua Xuan, Chang-E Ren, Zhi-Ping Shi 0002 |
SMC | 3 |
| 2021 | Formalization of Euler-Lagrange Equation Set Based on Variational Calculus in HOL Light
Jingzhi Zhang, Ximeng Li 0003, Zhi-Ping Shi 0002, Yongdong Li |
J. Autom. Reason. | 5 |
| 2021 | A Semi-supervised Learning Approach Based on Adaptive Weighted Fusion for Automatic Image AnnotationabstractTo learn a well-performed image annotation model, a large number of labeled samples are usually required. Although the unlabeled samples are readily available and abundant, it is a difficult task for humans to annotate large numbers of images manually. In this article, we propose a novel semi-supervised approach based on adaptive weighted fusion for automatic image annotation that can simultaneously utilize the labeled data and unlabeled data to improve the annotation performance. At first, two different classifiers, constructed based on support vector machine and covolutional neural network, respectively, are trained by different features extracted from the labeled data. Therefore, these two classifiers are independently represented as different feature views. Then, the corresponding features of unlabeled images are extracted and input into these two classifiers, and the semantic annotation of images can be obtained respectively. At the same time, the confidence of corresponding image annotation can be measured by an adaptive weighted fusion strategy. After that, the images and its semantic annotations with high confidence are submitted to the classifiers for retraining until a certain stop condition is reached. As a result, we can obtain a strong classifier that can make full use of unlabeled data. Finally, we conduct experiments on four datasets, namely, Corel 5K, IAPR TC12, ESP Game, and NUS-WIDE. In addition, we measure the performance of our approach with standard criteria, including precision, recall, F-measure, N+, and mAP. The experimental results show that our approach has superior performance and outperforms many state-of-the-art approaches. Zhixin Li 0001, Canlong Zhang, Huifang Ma, Weizhong Zhao, Zhi-Ping Shi 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2020 | Relational Graph Neural Network with Hierarchical Attention for Knowledge Graph CompletionabstractThe rapid proliferation of knowledge graphs (KGs) has changed the paradigm for various AI-related applications. Despite their large sizes, modern KGs are far from complete and comprehensive. This has motivated the research in knowledge graph completion (KGC), which aims to infer missing values in incomplete knowledge triples. However, most existing KGC models treat the triples in KGs independently without leveraging the inherent and valuable information from the local neighborhood surrounding an entity. To this end, we propose a Relational Graph neural network with Hierarchical ATtention (RGHAT) for the KGC task. The proposed model is equipped with a two-level attention mechanism: (i) the first level is the relation-level attention, which is inspired by the intuition that different relations have different weights for indicating an entity; (ii) the second level is the entity-level attention, which enables our model to highlight the importance of different neighboring entities under the same relation. The hierarchical attention mechanism makes our model more effective to utilize the neighborhood information of an entity. Finally, we extensively validate the superiority of RGHAT against various state-of-the-art baselines. Zhao Zhang 0011, Fuzhen Zhuang, Hengshu Zhu, Zhi-Ping Shi 0002, Hui Xiong 0001, Qing He 0003 |
AAAI | 4 |
| 2020 | Formal Verification of Atomicity Requirements for Smart Contracts
Ximeng Li 0003, Zhi-Ping Shi 0002 |
APLAS | 4 |
| 2020 | Formal Verification of Memory Isolation for the TrustZone-based TEEabstractThe trusted execution environment (TEE) is the security basis of embedded systems, which can provide a hardware-based isolated execution environment for security-sensitive components. Isolation of memory is a critical mechanism of TEE, the security of which plays a very important role in TEE's construction. In this paper, we present a formal verification of security properties about the memory isolation mechanism of TEE systems based on the ARM TrustZone, which is a hardware security technology commonly used on billions of ARM processors to create TEE. We establish a formal model of memory isolation, which consists of the formalization of ARMv8 architecture hardware components related to memory isolation and the formalization of a TrustZone monitor supporting world switch. We formally explicit and verify the correctness properties of memory management along with the information flow security properties of the memory isolation mechanism. The formalizations and verifications are all performed in the interactive theorem prover Isabelle/HOL. Yuwei Ma, Qianying Zhang, Shijun Zhao, Ximeng Li 0003, Zhi-Ping Shi 0002 |
APSEC | 6 |
| 2020 | Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video PredictionabstractVideo prediction is a pixel-wise dense prediction task to infer future frames based on past frames. Missing appearance details and motion blur are still two major problems for current models, leading to image distortion and temporal inconsistency. We point out the necessity of exploring multi-frequency analysis to deal with the two problems. Inspired by the frequency band decomposition characteristic of Human Vision System (HVS), we propose a video prediction network based on multi-level wavelet analysis to uniformly deal with spatial and temporal information. Specifically, multi-level spatial discrete wavelet transform decomposes each video frame into anisotropic sub-bands with multiple frequencies, helping to enrich structural information and reserve fine details. On the other hand, multilevel temporal discrete wavelet transform which operates on time axis decomposes the frame sequence into sub-band groups of different frequencies to accurately capture multifrequency motions under a fixed frame rate. Extensive experiments on diverse datasets demonstrate that our model shows significant improvements on fidelity and temporal consistency over the state-of-the-art works. Source code and videos are available at https://github.com/Bei-Jin/STMFANet. Beibei Jin, Yu Hu 0001, Qiankun Tang, Jingyu Niu, Zhi-Ping Shi 0002, Yinhe Han 0001, Xiaowei Li 0001 |
CVPR | 5 |
| 2020 | Lightdet: A Lightweight and Accurate Object Detection NetworkabstractThe extensive computational burden limits the usage of accurate but complex object detectors in resource-bounded scenarios. In this paper, we present a lightweight object detector, named LightDet, to address this dilemma. We design a lightweight backbone that is able to capture rich low-level features by the proposed Detail-Preserving Module. To effectively aggregate bottom and top-down features, we introduce an efficient Feature-Preserving and Refinement Module. A lightweight prediction head is employed to further reduce the entire network complexity. Experimental results show that our LightDet achieves 75.5% mAP on PASCAL VOC 2007 at the speed of 250 FPS and 24.0% mAP on MS COCO dataset. Qiankun Tang, Zhi-Ping Shi 0002, Yu Hu 0001 |
ICASSP | 3 |
| 2020 | Formalizing the Transaction Flow Process of Hyperledger Fabric
Ximeng Li 0003, Qianying Zhang, Zhi-Ping Shi 0002 |
ICFEM | 4 |
| 2020 | DAPC: Domain Adaptation People Counting via Style-level Transfer Learning and Scene-aware Estimation
Xingsen Wen, Zhi-Ping Shi 0002 |
ICPR | 3 |
| 2020 | Batch Normalization Masked Sparse Autoencoder for Robotic Grasping DetectionabstractTo improve the accuracy of the grasping detection, this paper proposes a novel detector with batch normalization masked evaluation model. It is designed with a two-layer sparse autoencoder, and a Batch Normalization based mask is incorporated into the second layer of the model to effectively reduce the features with weak correlation. The extracted features from such model are more distinctive, which guarantees the higher accuracy of the grasping detection. Extensive experiments show that the proposed evaluation model outperforms the state-of- the-art, and the recognition accuracy can reach 95.51% for robotic grasping detection. Zhenzhou Shao, Ying Qu 0001, Guangli Ren, Zhi-Ping Shi 0002, Jindong Tan |
IROS | 6 |
| 2020 | Deep Visible and Thermal Image Fusion with Cross-Modality Feature Selection for Pedestrian Detection
Zhenzhou Shao, Zhi-Ping Shi 0002 |
NPC | 3 |
| 2020 | Formalization of Camera Pose Estimation Algorithm based on Rodrigues FormulaabstractAbstract Camera pose estimation is key to the proper functioning of robotic systems, supporting critical tasks such as robot navigation, target tracking, camera calibration, etc.Whilemultiple algorithms solving this problem have been proposed, their correctness has rarely been validated using formal techniques. This is true despite the fact that the adoption of formal verification is essential for the reliability of safety-critical systems, and for their certification to high assurance levels. In this article, we present an effort in formally verifying an algorithm for camera pose estimation in an interactive theorem prover. The algorithm leverages the power of Rodrigues formula to solve the pose estimation problem under conditions for which existing solutions cannot be applied. The technical ingredients include (but are not limited to) mechanized proofs of the Rodrigues formula (along with its Cayley decomposition form) and the least squares method for fitting data. Based on the formalization of the algorithm, we formally derive and verify its general solution and unique solution. Shanyan Chen, Ximeng Li 0003, Qianying Zhang, Zhi-Ping Shi 0002 |
Formal Aspects Comput. | 5 |
| 2020 | Adaptive iterative learning consensus control for second-order multi-agent systems with unknown control gains
Guilu Li, Chang-E Ren, C. L. Philip Chen, Zhi-Ping Shi 0002 |
Neurocomputing | 4 |
| 2020 | Formalization of continuous Fourier transform in verifying applications for dependable cyber-physical systems
Jie Zhang 0074, Zhi-Ping Shi 0002, Yi Wang 0003, Yongdong Li |
J. Syst. Archit. | 3 |
| 2020 | Online Rail Surface Inspection Utilizing Spatial Consistency and ContinuityabstractRail surface inspection using visual inspection system is an important part of railway maintenance. However, accurate and efficient identification of possible defects remains challenging. This paper proposes a background-oriented defect inspector (BODI) to improve defect detection by considering specified characteristics of the track during inspection. Reformulating the inspection task in this manner offers a new way to model rail surface images. More specifically, BODI features a random sampling stage to obtain a compact background representation without any prior information. A sufficient number of random selections generates adequate and diverse background statistics, and defect-determination and a fusion of procedures then determine whether current pixel belongs to the background. Finally, a background update mechanism and parallelism ensure real-time applicability. The proposed BODI is evaluated on a working railway line. The experimental results demonstrate that it outperforms state-of-the-art methods. Jinrui Gan, Jianzhu Wang, Haomin Yu, Qingyong Li, Zhi-Ping Shi 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2019 | Formal Modelling and Verification of Spinlocks at Instruction LevelabstractSpinlocks have been widely used as a solution for synchronous accesses to shared resources, and their correctness is critical to guarantee the consistency of concurrent processes. This paper presents formal models and machine-checked verification of the correctness of spinlocks at instruction level. We present the formal verification of two spinlocks, which are spinlocks implemented based on the ARM instructions and the x86 instructions, respectively. Our model formalizes the lowlevel instructions that are necessary to capture the execution of spinlocks, characterizes the processor hardware mechanisms related to each instruction, and considers the context switches on processors and two-level scheduling of processors and processes. We specify the correctness property of our models, that is, accesses of a critical section satisfy mutual exclusion, and verify that the models satisfy the property using the theorem prover Isabelle/HOL. With the verification experience, we give some suggestions on how to implement spinlock leveraging the ARM ISA. Qianying Zhang, Zhi-Ping Shi 0002, Minhua Wu |
APSEC | 4 |
| 2019 | Towards Verifying Ethereum Smart Contracts at Intermediate Language Level
Ximeng Li 0003, Zhi-Ping Shi 0002, Qianying Zhang |
ICFEM | 2 |
| 2019 | P-Minder: A CNN Based Sidewalk Segmentation Approach for Phubber Safety ApplicationsabstractObstacle detection on smartphones can support the obstacle avoidance for phubbers, but detecting a driveway phubbing walking is difficult for it. However, driveway and sidewalk with different textures can be divided by semantic segmentation. Therefore, this paper proposes a phubber safety approach, named P-Minder (Phubber Minder), to predict the transition between sidewalk and driveway based on sidewalk semantic segmentation. It determines the range of the sidewalk and driveway respectively by the results of sidewalk segmentation, then reminds the phubbers through their smart phones before they enter the driveway. P-Minder includes a walking detector to provide low-consumption and an adjacent detection to assure the detection accuracy. Additionally, P-Minder can detect the existence of manhole covers or stairs on sidewalk. During the experiment, we establish a sidewalk dataset includes 6 classes: driveway, sidewalk, blind road, stairs, manhole covers and car barrier. The experimental result proves that the segmentation model can reach 77.9% miou. Besides, the method finally achieves 74.19% detection accuracy in the experiment with a low computing resources cost. Zhi-Ping Shi 0002 |
ICIP | 3 |
| 2019 | An IBP-CNN Based Fast Block Partition For Intra PredictionabstractThe increase of block size to 64×64 in HEVC leads to the increase of computational complexity of intra prediction. The convolution neural network (CNN) shows advantages in the extraction and application of image features than the traditional intra prediction optimization algorithm which is developed manually. For reducing the computational complexity of intra prediction, a CNN-based algorithm, intra block partition CNN (IBP-CNN) is proposed in this paper to get the block partition. First, a database which is consisted of coding tree unit (CTU) images and label images is established. The position of pixels in the label images are consistent with those in CTU images. Second, the texture features are analyzed by IBP-CNN to get the block partition. Then the output of the network is adjusted according to the quadtree structure of HEVC to facilitate the calculation of rate distortion (RD) cost. The method proposed in this paper reduces the average coding time of about 59.07% and the average BD-rate is about 1.55%. Wenpeng Ren, Zhi-Ping Shi 0002 |
PCS | 4 |
| 2019 | A HOL Theory of the Differential for Matrix FunctionsabstractThe differential of matrix functions(DMF) plays an important role in mathematics and engineering. Common applications of it are found in optimization analysis, computer vision, robotics, etc. In this paper, a formal method based on HOL is used to construct the DMF based on Fréchet differential in matrix space. In order to illustrate the practical effectiveness of our work, we use our formalization to verify a property of matrix exponential. Yuhan Nie, Zhi-Ping Shi 0002, Aixuan Wu, Ximeng Li 0003 |
TASE | 2 |
| 2019 | SafeHidden: An Efficient and Secure Information Hiding Technique Using Re-randomization
Zhe Wang 0017, Chenggang Wu 0002, Yinqian Zhang, Bowen Tang 0001, Pen-Chung Yew, Mengyao Xie, Yuanming Lai, Yan Kang 0002, Yueqiang Cheng, Zhi-Ping Shi 0002 |
USENIX Security Symposium | 10 |
| 2019 | Formalization of Geometric Algebra in HOL Light
Li-Ming Li, Zhi-Ping Shi 0002, Qianying Zhang, Yong-Dong Li |
J. Autom. Reason. | 2 |
| 2019 | Multi-representation adaptation network for cross-domain image classification
Yongchun Zhu, Fuzhen Zhuang, Jindong Wang 0001, Jingwu Chen, Zhi-Ping Shi 0002, Wenjuan Wu, Qing He 0003 |
Neural Networks | 5 |
| 2019 | SoftME: A Software-Based Memory Protection Approach for TEE System to Resist Physical AttacksabstractThe development of the Internet of Things has made embedded devices widely used. Embedded devices are often used to process sensitive data, making them the target of attackers. ARM TrustZone technology is used to protect embedded device data from compromised operating systems and applications. But as the value of the data stored in embedded devices increases, more and more effective physical attacks have emerged. However, TrustZone cannot resist physical attacks. We propose SoftME, an approach that utilizes the on-chip memory space to provide a trusted execution environment for sensitive applications. We protect the confidentiality and integrity of the data stored on the off-chip memory. In addition, we design task scheduling in the encryption process. We implement a prototype system of our approach on the development board supporting TrustZone and evaluate the overhead of our approach. The experimental results show that our approach improves the security of the system, and there is no significant increase in system overhead. Qianying Zhang, Shijun Zhao, Zhi-Ping Shi 0002 |
Secur. Commun. Networks | 4 |
| 2019 | SemRec: a personalized semantic recommendation method based on weighted heterogeneous information networks
Chuan Shi 0001, Zhiqiang Zhang 0012, Yugang Ji, Philip S. Yu, Zhi-Ping Shi 0002 |
World Wide Web | 6 |
| 2018 | Runtime Verification of Robots Collision Avoidance Case StudyabstractThe robot has attracted much attention to anticipate improved quality of human life. Real-time obstacle avoidance is one of hot spots of the research. Runtime verification is a real-time and lightweight verification technology to verify the properties in many fields. In this case study, we use the JavaMOP, a runtime verification tool to verify the implementations' correctness of the safety strategies for avoiding collision as a complement to the design. The design of the safety strategies can be classified as the pre-contact safety strategy and the post-contact safety strategy. The former can avoid obstacles and the latter can reduce the physical damage after a collision. Additionally, this case study also proposes a new method of dynamic parameter selection. It can automatically update the parameters during the operation of the robot without having to get familiar with and to modify robot programs. Because some special parameters may change as mutative factors or be updated by engineers' experience in the uncertain environment, they cannot be fixed. We follow the JavaMOP specification to describe informal requirements using the FSM and ptLTL languages. Finally, the experimental results verify the correctness of the safety strategies and the effectiveness of dynamic parameter selection. Chenxia Luo, Rui Wang 0024, Yu Jiang 0001, Zhi-Ping Shi 0002 |
COMPSAC (1) | 7 |
| 2018 | Formalization of Symplectic Geometry in HOL-Light
Zhi-Ping Shi 0002, Qianying Zhang, Yongdong Li |
ICFEM | 3 |
| 2018 | Formal analysis of the kinematic Jacobian in screw theoryabstractAbstract As robotic systems flourish, reliability has become a topic of paramount importance in the human–robot relationship. The Jacobian matrix in screw theory underpins the design and optimization of robotic manipulators. Kernel properties of robotic manipulators, including dexterity and singularity, are characterized with the Jacobian matrix. The accurate specification and the rigorous analysis of the Jacobian matrix are indispensable in guaranteeing correct evaluation of the kinematics performance of manipulators. In this paper, a formal method for analyzing the Jacobian matrix in screw theory is presented using the higher-order logic theorem prover HOL4. Formalizations of twists and the forward kinematics are performed using the product of exponentials formula and the theory of functional matrices. To the best of our knowledge, this work is the first to formally analyze the kinematic Jacobian using theorem proving. The formal modeling and analysis of the Stanford manipulator demonstrate the effectiveness and applicability of the proposed approach to the formal verification of the kinematic properties of robotic manipulators. Zhi-Ping Shi 0002, Aixuan Wu, Xiumei Yang, Yongdong Li |
Formal Aspects Comput. | 1 |
| 2018 | A cyber-enabled visual inspection system for rail corrugation
Qingyong Li, Zhi-Ping Shi 0002, Huayan Zhang, Yunqiang Tan, Shengwei Ren |
Future Gener. Comput. Syst. | 2 |
| 2018 | Sensor attack detection using history based pairwise inconsistency
Rui Wang 0024, Yu Jiang 0001, Houbing Song, Chenxia Luo, Zhi-Ping Shi 0002 |
Future Gener. Comput. Syst. | 8 |
| 2017 | Formalization and analysis of jacobian matrix in screw theory and its application in kinematic singularityabstractAccurate specification and rigorous analysis of Jacobian matrix are indispensable to guarantee correct evaluation on the manipulator kinematics performance. In this paper, a formal analysis method of the Jacobian matrix in the screw theory is presented by using the higher-order logic theorem prover HOL4. Formalizations of twists and the forward kinematics are characterized with the product of exponential formula and the theory of functional matrices. To the best of our knowledge, this work is the first to formally reason about the spatial Jacobian using theorem proving. The formal modeling and analysis of a 3-DOF planar manipulator substantiate the effectiveness and applicability of the proposed approach to formally verify the kinematics properties of manipulator. Aixuan Wu, Zhi-Ping Shi 0002, Xiumei Yang, Yongdong Li |
IROS | 2 |
| 2015 | Embedded System Design with Reliability-Centric OptimizationabstractEmbedded systems are becoming increasingly popular due to their widespread applications. Hardware/software partitioning with reliability in consideration is becoming one of the most crucial steps in the design of complex embedded systems, especially for safety critical applications. In this paper, a reliability-centric approach is proposed to increase the reliability and decrease the potential errors of the embedded system during the partition stage. We use task graph to model the system function blocks, and treat reliability as the first-class metric during the mathematical formalization and optimization procedure. Experimental evaluations demonstrate that the proposed approach obtains a better solution with the trade off of time, cost and reliability compared to existing partitioning methods, with higher and more accurate reliability characterization for safety-critical systems. Yuanyuan Hou, Rui Wang 0024, Yu Jiang 0001, Zhi-Ping Shi 0002, Jie Zhang 0074 |
COMPSAC | 6 |
| 2014 | Formal verification of a collision-free algorithm of dual-arm robot in HOL4abstractPossessing two manipulators heightens the ability of dual-arm robots (DAR) to conduct complex tasks, while raising hazard that the two manipulators might collide with each other or with other objects. DARs are usually equipped with a collision-free motion planning algorithms (CFMPA) to prevent the two manipulators from colliding. The CFMPA searches the motion paths of robot manipulators, which are expected to be as short and smooth as possible under the premise of ensuring safety. It is important to ensure that the algorithm is correct and efficient. It is not enough to apply traditional test methods to determine whether DARs can work in safety-critical applications. In this paper, theorem proving technology is employed to analyze the correctness and efficiency of a classical CFMPA. The CFMPA is outlined, and then formalized in high order logic with the theorem prover HOL4. An inconsistency in the range of motions of the robot manipulators in the algorithm is discovered. An improved algorithm is therefore proposed. Formal verification with HOL4 proves the correctness and efficiency of the proposed algorithm that has already run on a real DAR as well, in conformity with our expectation. Zhi-Ping Shi 0002, Chunna Zhao, Jie Zhang 0074, Hongxing Wei |
ICRA | 2 |
| 2014 | A co-boost framework for learning object categories from Google Images with 1st and 2nd order features
Xi Liu 0009, Zhi-Ping Shi 0002, Zhongzhi Shi |
Vis. Comput. | 2 |
| 2012 | Extracting discriminative features for CBIR
Zhi-Ping Shi 0002, Xi Liu 0009, Qingyong Li, Qing He 0003, Zhongzhi Shi |
Multim. Tools Appl. | 1 |
| 2011 | Modeling continuous visual features for semantic image annotation and retrieval
Zhixin Li 0001, Zhi-Ping Shi 0002, Xi Liu 0009, Zhongzhi Shi |
Pattern Recognit. Lett. | 2 |
| 2010 | Automatic image annotation with continuous PLSAabstractAutomatic image annotation has become an important and challenging problem due to the existence of semantic gap. In this paper, we firstly extend probabilistic latent semantic analysis (PLSA) to model continuous quantity. In addition, corresponding Expectation-Maximization (EM) algorithm is derived to determine the model parameters. Furthermore, in order to deal with the data of different modalities in terms of their characteristics, we present a semantic annotation model which employs continuous PLSA and standard PLSA to model visual features and textual words respectively. The model learns the correlation between these two modalities by an asymmetric learning approach and then it can predict semantic annotation for unseen images. We compare our approach with several state-of-the-art approaches on a standard Corel dataset. The experiment results show that our approach performs more effectively and accurately. Zhixin Li 0001, Zhi-Ping Shi 0002, Xi Liu 0009, Zhongzhi Shi |
ICASSP | 2 |
| 2010 | A novel sparse coding model based on structural similarityabstractUnderstanding and modeling the function of the neurons and neural systems are primary goal of systems neuroscience. Sparse coding theory demonstrates that the neurons in primary visual cortex form a sparse representation of natural scenes in the viewpoint of statistics. In this paper, we propose a novel sparse coding model based on structural similarity (SS_SC) for natural image feature extraction. The advantage for our model is to be able to preserve structural information from a scene, which human visual perception is highly adapted for. Using the proposed sparse coding model, the validity of image feature extraction is testified. Furthermore, compared with standard sparse coding (SC) model, the experimental results show that the quality of reconstructed images obtained by our method outperforms the SC method. Zhi-Ping Shi 0002, Xi Liu 0009, Zhongzhi Shi |
ICASSP | 2 |
| 2010 | Sorted label classifier chains for learning images with multi-labelabstractIn the real world, images always have several visual objects instead of only one, which makes it difficult for conventional object recognition methods to deal with them. In this paper, we present a topologically sorted classifier chain method for learning images with multi-label. We first provide a means of generating a topo-logically sorted label chain ordering by employing a topological sort algorithm and then apply the chain ordering to the classifier chain model proposed by [1] to classify multi-label images. Our method can capture the correlations between labels very effectively due to the sorted label chain ordering and the advantages brought by classifier chain method. We evaluate the proposed method on Corel dataset and demonstrate the micro and macro F1 measures superior to the state-of-the-art methods. Xi Liu 0009, Zhi-Ping Shi 0002, Zhixin Li 0001, Xishun Wang, Zhongzhi Shi |
ACM Multimedia | 2 |
| 2010 | Fusing semantic aspects for image annotation and retrieval
Zhixin Li 0001, Zhi-Ping Shi 0002, Xi Liu 0009, Zhongzhi Shi |
J. Vis. Commun. Image Represent. | 2 |
| 2009 | Modeling latent aspects for automatic image annotationabstractIn this paper, we present an approach based on probabilistic latent semantic analysis (PLSA) to accomplish the tasks of automatic image annotation. In order to model training data precisely, we represent an image as a bag of visual words and employ two PLSA models to capture semantic information from visual and textual modalities respectively. Furthermore, an adaptive learning approach is proposed to combine the aspects learned from both modalities. For each image document, distribution over aspects is fused by different weight in terms of the entropy of its feature distribution. Consequently, the two models are linked with the same distribution over aspects. This structure can predict semantic annotation for an unseen image because it associates visual and textual modalities properly. We compare our approach with several previous approaches on a standard Corel dataset. The experiment results show that our approach performs more effectively and accurately. Zhixin Li 0001, Zhi-Ping Shi 0002, Zhongzhi Shi |
ICIP | 2 |
| 2009 | Filter object categories using CoBoost with 1ST and 2ND order featuresabstractWe describe a method for filtering object category from a large number of noisy images. This problem is particularly difficult due to the greater variation within object categories and lack of labeled object images. Our method deals with it by combining a co-training algorithm CoBoost with two features - 1stand 2ndorder features, which define bag of words representation and spatial relationship between local features respectively. We iteratively train two boosting classifiers based on the 1stand 2ndorder features, during which each classifier provides labeled data for the other classifier. It is effective because the 1stand 2ndorder features make up an independent and redundant feature split. We evaluate our method on Berg dataset and demonstrate the precision comparative to the state-of-the-art. Xi Liu 0009, Zhi-Ping Shi 0002, Zhongzhi Shi |
ICIP | 2 |
| 2009 | Learning image semantics with latent aspect modelabstractAutomatic image annotation has become an important and challenging problem due to the existence of semantic gap. In this paper, we present an approach based on probabilistic latent semantic analysis (PLSA) to accomplish the tasks of semantic image annotation and retrieval. In order to model training images precisely, we employ two PLSA models to capture semantic information from visual and textual modalities respectively. Then an adaptive asymmetric learning approach is proposed to fuse aspects which are learned from both modalities. For each image document, the weight of each modality is determined by its contribution to the content of the image. Consequently, the two models are linked with the same distribution over aspects. This structure can predict semantic annotation for an unseen image because it associates visual and textual modalities properly. Finally, we compare our approach with several previous approaches on a standard Corel dataset. The experiment results show that our approach performs more effective and accurate. Zhixin Li 0001, Xi Liu 0009, Zhi-Ping Shi 0002, Zhongzhi Shi |
ICME | 3 |
| 2009 | Filter object categories: employing visual consistency and semisupervised approachabstractWe describe a method for filtering object category from a large number of noisy images. This problem is particularly difficult due to the greater variation within object categories and only a few labeled object images available. Our method deals with it by using visual consistency and semi-supervised approach. The images of one category often share some visual consistency so that the most irrelevant images can be first removed. Among the left images, a voting method is used to obtain more object exemplars with the initial object exemplars manually selected by users. Finally with all the obtained exemplars and those unlabeled images, we create a semi-supervised classifier to rank all the images. We evaluate our method on Berg dataset and demonstrate the precision comparative to the state-of-the-art. Besides, we collect five more categories from Google images to show the effectiveness of the method. Xi Liu 0009, Zhixin Li 0001, Zhi-Ping Shi 0002, Zhongzhi Shi |
ICME | 3 |
| 2009 | An Efficient Coding Model for Image Representation
Zhi-Ping Shi 0002, Zhixin Li 0001, Zhongzhi Shi |
ICONIP (1) | 2 |
| 2009 | Coboost learning of visual categories with 1st and 2nd order features from Google imagesabstractConventional object recognition techniques rely heavily on manually annotated image datasets to achieve good performances. However, collecting high quality datasets is really laborious. In this paper, we propose a semi-supervised framework for learning visual categories from Google Images. The 1st and 2nd order features, which define bag of words representation and spatial relationship between local features respectively, make up an independent and redundant feature split. We then integrate a cotraining algorithm CoBoost with these two features. We create two boosting classifiers based on the 1st and 2nd order features respectively in the training, during which one classifier provides labels for the other. Besides, the 2nd order features are generated dynamically rather than extracted exhaustively to avoid high computation. An active learning technique is also introduced to further improve the performance. We evaluate our method on the benchmark datasets, showing results competitive with the state-of-the-art unsupervised approaches and some supervised techniques. Xi Liu 0009, Zhi-Ping Shi 0002, Zhixin Li 0001, Zhongzhi Shi |
ACM Multimedia | 2 |
| 2009 | An index and retrieval framework integrating perceptive features and semantics for multimedia databases
Zhi-Ping Shi 0002, Qing He 0003, Zhongzhi Shi |
Multim. Tools Appl. | 1 |
| 2007 | Image Retrieval Based on Fuzzy Color SemanticsabstractIn order to improve the performance of content-based image retrieval (CBIR) systems, the 'semantic gap' between the low-level visual features and the high-level semantic features attracts more and more research interest. We propose an approach to describe and to extract the fuzzy color semantics. According to human color perception model, we utilize the linguistic variable to describe the image color semantics, so it becomes possible to depict the image in linguistic expression such as mostly red. Furthermore, we apply the feedforward neural network to model the vagueness of human color perception and to extract the fuzzy semantic feature vector. Our experiments show that the color semantic features have good accordance with the human perception, and also have good retrieval performance. In some extent, our approach shows the potential to reduce the semantic gap in CBIR. Qingyong Li, Zhi-Ping Shi 0002, Siwei Luo |
FUZZ-IEEE | 2 |
| 2007 | A Neural Network Approach for Bridging the Semantic Gap in Texture Image RetrievalabstractOne of the big challenges faced by content-based image retrieval (CBIR) is the 'semantic gap' between the visual features and the richness of human semantics for image content. We put forward a neural network approach to extract the image fuzzy semantics ground on linguistic expression based image description framework (LEBID). We utilize the linguistic variable to depict the texture semantics according to Tamura texture model, so we can describe the image in linguistic expression such as coarse, very line-like. Moreover, we use feedforward neural network (NN) to model the vagueness of human visual perception and to extract the fuzzy semantic feature. Our experiments demonstrate that NN outperforms other method such as genetic algorithm on the complexity of model, and it also achieves good retrieval performance. Qingyong Li, Zhi-Ping Shi 0002, Siwei Luo |
IJCNN | 2 |