EDBT 2026 Demo / reviewers in the wild / expert
Shiqiang Zhu
dblp:122/3795
· DBLP profile ↗
32ranked-venue papers
2as first author
27since 2021 · last 2026
0000-0002-5687-4001ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 16 since 2021Systems, architecture and hardware · 10 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MSBPD: multi-scale contextual semantics enhancement with bidirectional parallel prompt decoding for document-level event extraction
Shiqiang Zhu, Qianxi Hou |
Knowl. Inf. Syst. | 1 |
| 2025 | FCRF: Flexible Constructivism Reflection for Long-Horizon Robotic Task Planning with Large Language ModelsabstractAutonomous error correction is critical for domestic robots to achieve reliable execution of complex long-horizon tasks. Prior work has explored self-reflection in Large Language Models (LLMs) for task planning error correction; however, existing methods are constrained by inflexible self-reflection mechanisms that limit their effectiveness. Motivated by these limitations and inspired by human cognitive adaptation, we propose the Flexible Constructivism Reflection Framework (FCRF), a novel Mentor-Actor architecture that enables LLMs to perform flexible self-reflection based on task difficulty, while constructively integrating historical valuable experience with failure lessons. We evaluated FCRF on diverse domestic tasks through simulation in AlfWorld and physical deployment in the real-world environment. Experimental results demonstrate that FCRF significantly improves overall performance and self-reflection flexibility in complex long-horizon robotic tasks. Website at https://mongoosesyf.github.io/FCRF.github.io/ Jiatao Zhang, Zeng Gu, Qingmiao Liang, Tuocheng Hu, Wei Song 0008, Shiqiang Zhu |
IROS | 7 |
| 2025 | A Crab-Inspired Soft Gripper with Single-Finger Dexterous Grasping CapabilitiesabstractSoft grippers conform to the shape and surface properties of the objects to be grasped, effectively avoiding damage to soft and fragile items. Despite the variety of existing soft gripper designs, their structures lack sufficient flexibility for effectively grasping slender objects or operating in narrow spaces. To address these challenges, we propose a soft gripper with single-finger grasping capabilities, inspired by the structure of crab claws. The structural design and the fabrication method of the gripper are introduced, and the analytical bending model is derived. Experiments are conducted under typical operating conditions to validate the model, and the results indicate that the measured data are in good accordance with the predicted responses. Furthermore, a series of grasping experiments are carried out to test the single-finger grasping capabilities of the proposed soft gripper. The results indicate that the proposed soft gripper can efficiently and stably grasp slender or irregular objects with a single finger. In particular, it demonstrates suitability for operations in narrow spaces and shows potential for handling complex tasks. This innovative design effectively reduces the complexity of the system, while exhibiting promising capabilities in grasping slender or irregular objects and operating within restricted spaces. Yunce Zhang, Haobin Lv, Yixiang Liu, Zhe Min, Shizhao Zhou, Tao Wang 0072, Shiqiang Zhu, Rui Song 0002 |
IROS | 7 |
| 2025 | Efficient FPGA Implementation of Multi-Channel Pipelined Large FFT Architectures Based on SA-MDF AlgorithmabstractFPGA implementation of a multi-channel pipelined large FFT architecture is challenging due to its complex inter-channel data scheduling, high-throughput requirement, and resource-constrained hardware. By transforming to 2D-FFT implementation, investigating different binary tree schemes, and exploring various radices, butterflies, as well as data path structures, many hardware architectures have been designed to enhance single-channel large FFT or multi-channel medium-small size FFT performance. These designs fall short in addressing the demands of multi-channel pipelined and large FFT applications. In this article, a self-attention multipath delay feedback (SA-MDF) algorithm is proposed to analyze and identify the most critical bottleneck, then automatically pay attention to improve it, and finally generate the optimal FFT framework by exhaustively exploring the design space. The proposed algorithm alleviates the design difficulties and speeds up the FPGA implementation. Furthermore, an approximate roofline model and a novel binary tree scheme are introduced to further minimize the utilization of on-chip memory. A comprehensive comparison in terms of principles, implementation methods, and optimization effects is conducted when compared with other multi-channel FFT architectures. Experimental results show that the proposed FFT architectures are superior to other FFT implementations in terms of channel count, FFT length, high-throughput data arrangement, and adaptability to diverse hardware platforms. Tang Hu, Chunling Hao, Xier Wang, Songnan Ren, Zhiwei Xu 0003, Shiqiang Zhu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | Leveraging the efficiency of multi-task robot manipulation via task-evoked planner and reinforcement learningabstractMulti-task learning has expanded the boundaries of robotic manipulation, enabling the execution of increasingly complex tasks. However, policies learned through reinforcement learning exhibit limited generalization and narrow distributions, which restrict their effectiveness in multi-task training. Addressing the challenge of obtaining policies with generalization and stability represents a non-trivial problem. To tackle this issue, we propose a planning-guided reinforcement learning method. It leverages a task-evoked planner(TEP) and a reinforcement learning approach with planner’s guidance. TEP utilizes reusable samples as the source, with the aim of learning reachability information across different task scenarios. Then in reinforcement learning, TEP assesses and guides the Actor towards better outputs and smoothly enhances the performance in multi-task benchmarks. We evaluate this approach within the Meta-World framework and compare it with prior works in terms of learning efficiency and effectiveness. Depending on experimental results, our method has more efficiency, higher success rates, and demonstrates more realistic behavior. Haofu Qian, Jiatao Zhang, Jason Gu, Wei Song 0008, Shiqiang Zhu |
ICRA | 7 |
| 2024 | FLTRNN: Faithful Long-Horizon Task Planning for Robotics with Large Language ModelsabstractRecent planning methods based on Large Language Models typically employ the In-Context Learning paradigm. Complex long-horizon planning tasks require more context(including instructions and demonstrations) to guarantee that the generated plan can be executed correctly. However, in such conditions, LLMs may overlook(unfaithful) the rules in the given context, resulting in the generated plans being invalid or even leading to dangerous actions. In this paper, we investigate the faithfulness of LLMs for complex long-horizon tasks. Inspired by human intelligence, we introduce a novel framework named FLTRNN. FLTRNN employs a language-based RNN structure to integrate task decomposition and memory management into LLM planning inference, which could effectively improve the faithfulness of LLMs and make the planner more reliable. We conducted experiments in VirtualHome household tasks. Results show that our model significantly improves faithfulness and success rates for complex long-horizon tasks. Website at https://tannl.github.io/FLTRNN.github.io/ Jiatao Zhang, Lanling Tang, Qiwei Meng, Haofu Qian, Wei Song 0008, Shiqiang Zhu, Jason Gu |
ICRA | 8 |
| 2024 | Whole-Body Inverse Kinematics and Operation-Oriented Motion Planning for Robot Mobile ManipulationabstractHigh DoF mobile manipulation of robots is a nonlinear, nonchain redundant problem. In this article, we focus on two subissues of robot mobile manipulation: whole-body inverse kinematics (whole-body IK) and operation-oriented motion planning (OOMP). Whole-body IK solves the robot arm joint configuration and the mobile base position configuration according to the target pose. OOMP generates a feasible trajectory from the current pose to the target pose. The trajectory can avoid obstacles and touch operated objects. We introduce neural network optimization (NNO) methods with two variations to solve whole-body IK and OOMP, respectively. For whole-body IK, we design a fully connected network (FCN) to predict ten DoF of position and joint configurations based on the target pose. We use these ten DoF configurations to derive the predicted pose for online optimization. For OOMP, we design a GRU-based network to generate trajectories based on the initial and goal states. We mainly adopt sphere masks to modify the point cloud properties of the target object dynamically. During optimization, the trajectory keeps away from point clouds but approaches sphere masks. Finally, we conduct extensive experiments both on a Franka Panda robot and a mobile dual-arm robot. The results demonstrate the superior performance of our NNO method on whole body IK and OOMP, and implement mobile manipulation in different environments successfully. Tianlei Jin, Jiakai Zhu, Shiqiang Zhu, Zaixing He, Shuyou Zhang 0001, Wei Song 0008, Jason Gu |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | HEU-Net: hybrid attention residual block-based network with external skip connections for metal corrosion semantic segmentation
Tiancheng Zhu, Shiqiang Zhu, Hongliang Ding, Wei Song 0008, Cunjun Li |
Vis. Comput. | 2 |
| 2023 | Fast Contextual Scene Graph Generation with Unbiased Context AugmentationabstractScene graph generation (SGG) methods have historically suffered from long-tail bias and slow inference speed. In this paper, we notice that humans can analyze relationships between objects relying solely on context descriptions, and this abstract cognitive process may be guided by experience. For example, given descriptions of cup and table with their spatial locations, humans can speculate possible relationshipsor. Even without visual appearance information, some impossible predicates like flying in and looking at can be empirically excluded. Accordingly, we propose a contextual scene graph generation (C-SGG) method without using visual information and introduce a context augmentation method. We propose that slight perturbations in the position and size of objects do not essentially affect the relationship between objects. Therefore, at the context level, we can produce diverse context descriptions by using a context augmentation method based on the original dataset. These diverse context descriptions can be used for unbiased training of C-SGG to alleviate long-tail bias. In addition, we also introduce a context guided visual scene graph generation (CV-SGG) method, which leverages the C-SGG experience to guide vision to focus on possible predicates. Through extensive experiments on the publicly available dataset, C-SGG alleviates long-tail bias and omits the huge computation of visual feature extraction to realize real-time SGG. CV-SGG achieves a great trade-off between common predicates and tail predicates. Tianlei Jin, Fangtai Guo, Qiwei Meng, Shiqiang Zhu, Xiangming Xi, Wen Wang 0017, Zonghao Mu, Wei Song 0008 |
CVPR | 4 |
| 2023 | Semi-Supervised Domain Generalization with Graph-Based ClassifierabstractSemi-supervised domain generalization (SSDG) has recently emerged as a potential research topic. Compared to domain generalization, SSDG represents a realistic and challenging goal, which only requires a few labels from source domains. To tackle this problem, this work presents a novel pseudo-labeling method that facilitates incremental learning on a large amount of unlabeled data. With edge weighting optimization, the proposed method utilizes the graph Laplacian regularizer (GLR) in a multi-class setting that relies on the generated similarity graph. The proposed overall SSDG scheme mitigates the overfitting problem by an adaptive threshold module based on a two-stage GLR denoiser. Our experiments on PACS and OfficeHome verify that the proposed method effectively improves the quality of pseudo-labeling and domain generalization, achieving top performance in terms of accuracy. Minxiang Ye, Shiqiang Zhu, Anhuan Xie, Senwei Xiang |
ICASSP | 3 |
| 2023 | KGNet: Knowledge-Guided Networks for Category-Level 6D Object Pose and Size EstimationabstractDespite the giant leap made in object 6D pose estimation and robotic grasping under structured scenarios, most approaches depend heavily on the exact CAD models of target objects beforehand, thereby limiting their wide applications. To address this, we propose a novel knowledge-guided network - KGNet to estimate the pose and size of category-level unseen objects. This network includes three primary innovations: knowledge-guided categorical model generation, pointwise deformation probability matrix and synergetic RGBD feature fusion, with the former two leveraging categorical object knowledge for unseen object reconstruction and the latter one facilitating pose-sensitive feature extraction. Exten-sive experiments on CAMERA25 and REAL275 verify their effectiveness, and KGNet achieves the SOTA performance on these two acknowledged benchmarks. Additionally, a real-world robotic grasping experiment is conducted, and its results further qualitatively prove the practicability and robustness of KGNet. Qiwei Meng, Jason Gu, Shiqiang Zhu, Jianfeng Liao, Tianlei Jin, Fangtai Guo, Wen Wang 0017, Wei Song 0008 |
ICRA | 3 |
| 2023 | RFFCE: Residual Feature Fusion and Confidence Evaluation Network for 6DoF Pose EstimationabstractIn this paper, we propose a novel RGBD-based object 6DoF pose estimation network - RFFCE. It is a two-stage method that firstly leverages deep neural networks for feature extraction and object points matching, and then the geometric principles are utilized for final pose computation. Our approach consists of three primary innovations: residual feature fusion for representative RGBD feature extraction; confidence evaluation and confidence-based paired points offsets regression for self-evaluation and self-optimization respectively. Their effectiveness is verified through an ablation study, and our RFFCE achieves the SOTA performance on LineMOD, Occlusion-LineMOD and YCB-Video datasets. Additionally, we also conduct a real-world object grasping experiment for visualization and qualitative evaluation of the RFFCE. Qiwei Meng, Shanshan Ji, Shiqiang Zhu, Tianlei Jin, Jason Gu, Wei Song 0008 |
ICRA | 3 |
| 2023 | Towards Safe and Aggressive Motion Generation for Dynamic Targets Pick-and-PlaceabstractIn this paper, we present a framework to generate time-optimal trajectories for dynamic target pick-and-place tasks. We develop an optimization-based trajectory generation method for manipulators, which can conduct spatial-temporal deformation under user-defined requirements. We formulate the problem of dynamic target pick-and-place, in which the trajectory duration and jerk are optimized and terminal states are adjusted instead of being fixed. The motions are constrained within the mechanical limits and to avoid collisions. Constraints transcription is adopted to convert constraints to weighted penalties. Then the problem can be solved based on the trajectory generation method with a high-level optimizer. We integrate the proposed method with online perception into a robot arm platform, in which a conveyor belt is used to transport the objects. Simulations and real-world experiments are conducted under a range of object speeds. Results show that the proposed method achieves online grasping under the object velocity up to 0.5m/s with an average computing time of 190ms. Jianfeng Liao, Shiqiang Zhu, Wei Song 0008, Yinchun Huang |
IROS | 5 |
| 2023 | B2C-AFM: Bi-Directional Co-Temporal and Cross-Spatial Attention Fusion Model for Human Action RecognitionabstractHuman Action Recognition plays a driving engine of many human-computer interaction applications. Most current researches focus on improving the model generalization by integrating multiple homogeneous modalities, including RGB images, human poses, and optical flows. Furthermore, contextual interactions and out-of-context sign languages have been validated to depend on scene category and human per se. Those attempts to integrate appearance features and human poses have shown positive results. However, with human poses' spatial errors and temporal ambiguities, existing methods are subject to poor scalability, limited robustness, and sub-optimal models. In this paper, inspired by the assumption that different modalities may maintain temporal consistency and spatial complementarity, we present a novel Bi-directional Co-temporal and Cross-spatial Attention Fusion Model (B2C-AFM). Our model is characterized by the asynchronous fusion strategy of multi-modal features along temporal and spatial dimensions. Besides, the novel explicit motion-oriented pose representations called Limb Flow Fields (Lff) are explored to alleviate the temporal ambiguity regarding human poses. Experiments on publicly available datasets validate our contributions. Abundant ablation studies experimentally show that B2C-AFM achieves robust performance across seen and unseen human actions. The codes are available at https://github.com/gftww/B2C.git. Fangtai Guo, Tianlei Jin, Shiqiang Zhu, Xiangming Xi, Wen Wang 0017, Qiwei Meng, Wei Song 0008, Jiakai Zhu |
IEEE Trans. Image Process. | 3 |
| 2023 | Genetic Programming for Dynamic Workflow Scheduling in Fog ComputingabstractDynamicWorkflowScheduling inFogComputing (DWSFC) is an important optimisation problem with many real-world applications. The current workflow scheduling problems only consider cloud servers but ignore the roles of mobile devices and edge servers. Some applications need to consider the mobile devices, edge, and cloud servers simultaneously, making them work together to generate an effective schedule. In this article, a new problem model for DWSFC is considered and a new simulator is designed for the new DWSFC problem model. The designed simulator takes the mobile devices, edge, and cloud servers as a whole system, where they all can execute tasks. In the designed simulator, two kinds of decision points are considered, which are the routing decision points and the sequencing decision points. To solve this problem, a newMulti-TreeGeneticProgramming (MTGP) method is developed to automatically evolve scheduling heuristics that can make effective real-time decisions on these decision points. The proposed MTGP method with a multi-tree representation can handle the routing decision points and sequencing decision points simultaneously. The experimental results show that the proposed MTGP can achieve significantly better test performance (reduce the makespan by up to 50%) on all the tested scenarios than existing state-of-the-art methods. Meng Xu 0008, Yi Mei 0001, Shiqiang Zhu, Beibei Zhang 0007, Fangfang Zhang 0003, Mengjie Zhang 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | BCOT: A Markerless High-Precision 3D Object Tracking BenchmarkabstractTemplate-based 3D object tracking still lacks a high-precision benchmark of real scenes due to the difficulty of annotating the accurate 3D poses of real moving video objects without using markers. In this paper, we present a multi-view approach to estimate the accurate 3D poses of real moving objects, and then use binocular data to construct a new benchmark for monocular textureless 3D object tracking. The proposed method requires no markers, and the cameras only need to be synchronous, relatively fixed as cross-view and calibrated. Based on our object-centered model, we jointly optimize the object pose by minimizing shape reprojection constraints in all views, which greatly improves the accuracy compared with the single-view approach, and is even more accurate than the depth-based method. Our new benchmark dataset contains 20 textureless objects, 22 scenes, 404 video sequences and 126K images captured in real scenes. The annotation error is guaranteed to be less than 2mm, according to both theoretical analysis and validation experiments. We reevaluate the state-of-the-art 3D object tracking methods with our dataset, reporting their performance ranking in real scenes. Our BCOT benchmark and code can be found at https://ar3dv.github.io/BCOT-Benchmark/. Bin Wang 0035, Shiqiang Zhu, Xin Cao 0010, Fan Zhong 0001, Wenxuan Chen, Jason Gu, Xueying Qin |
CVPR | 3 |
| 2022 | Deep Markov Clustering for Panoptic SegmentationabstractPanoptic segmentation is a challenging scene understanding task that unifies semantic segmentation and instance segmentation. Namely, each pixel of an image is assigned a semantic label and an instance id. Existing works have elaborated end-to-end panoptic segmentation networks and made great progress in non-proposal-based methods. In this work, we adopt a box-free strategy and incorporate a graph-based clustering method to merge repetitive kernel weights for object instances. An alternative graph-based clustering algorithm like Markov clustering performs effective random walks for unsupervised clustering without pre-defined cluster numbers. Our proposed deep Markov clustering scheme provides an efficient alternative to guarantee instance-aware label prediction in both training and inference stages. On the COCO dataset, our method achieves promising accuracy (PQ=42.1), which is comparable with state-of-the-art methods. Minxiang Ye, Shiqiang Zhu, Anhuan Xie, Dan Zhang 0006 |
ICASSP | 3 |
| 2022 | Independent Relationship Detection for Real-Time Scene Graph Generation
Tianlei Jin, Wen Wang 0017, Shiqiang Zhu, Xiangming Xi, Qiwei Meng, Zonghao Mu, Wei Song 0008 |
ICONIP (4) | 3 |
| 2022 | DMM: Dual-Modal Model for Person Re-IdentificationabstractThis paper explores how to boost the performance of current person re-identification (ReID) models by incorporating auxiliary information such as contour sketch. Most current ReID methods consider only RGB images as input, with little attention on extra yet important information contained in other modal images. We propose a dual-modal model (DMM), consisting of a main stream that inputs RGB images, and an auxiliary stream that inputs other modal images, to explore how the auxiliary information will help to promote the performance of existing ReID models. To fuse these two streams, a novel dual-modal attention (DMA) mechanism is proposed. Specifically, we apply spatial attention to auxiliary feature maps to take full advantage of the informative spatial locations contained in this stream. Then channel attention is applied to the spatially refined main feature maps, resulting in further refined representations. Moreover, we adopt DMA at multiple scales to exploit different semantics from low to high levels, which finally generates more discriminative feature representations. Comprehensive experiments on publicly available datasets, Market1501, DukeMTMC, MSMT17, and Black ReID, show that our proposal achieves SOTA results. Wen Wang 0017, Shunda Hu, Shiqiang Zhu, Zhiyong Huang 0005, Zheyuan Lin, Tianlei Jin |
IJCNN | 3 |
| 2022 | Depth-aware gaze-following via auxiliary networks for roboticsabstractGaze-Following aims to predict the gaze target of a subject within an image, and information on orientation and depth greatly improves this task. However, previous methods require additional datasets to obtain depth or orientation information, leading to cumbersome training or inference processes. To this end, we propose an end-to-end depth-aware gaze-following approach that incorporates depth and orientation information without additional datasets. Our approach identifies a primary task, gaze-following, supervised by true labels from the gaze-following dataset and two auxiliary tasks, scene depth estimation and 3D orientation estimation, supervised by generated pseudo labels. Intermediate auxiliary features are integrated into the primary task network as implicit information. We propose a residual filter module for screening useful information that can enhance gaze-following prediction performance. Extensive experiments on GazeFollow and VideoAttentionTarget show that our approach achieves state-of-the-art results (0.120 Ave. Dist. achieved on GazeFollow and 0.104 L2 Dist. achieved on VideoAttentionTarget). Finally, we apply our approach to a real robot for understanding human attention and intention. Compared to the previous depth considered gaze-following method, our method saves half of the computation time. Tianlei Jin, Qizhi Yu, Shiqiang Zhu, Zheyuan Lin, Yuanhai Zhou, Wei Song 0008 |
Eng. Appl. Artif. Intell. | 3 |
| 2021 | Multi-Person Gaze-Following with Numerical Coordinate RegressionabstractGaze-Following is a complex task that needs to combine the gaze with the scene. Previous works performed well on predicting single-person gaze-following but expensive computations are impractical to the real-world project. Moreover, when there are multiple people appearing at the same time, previous works will excecute repeated scene feature extraction. In addition, obtaining gaze target point through the heatmap argmax method seems to be a convention for gaze-following while the quantization error of the heatmap is ignored. In this paper, a simple but efficient network structure is proposed to provide shared scene features for the multi-person gaze-following, and a numerical coordinate regression is firstly introduced to calculate the gaze target point and regression loss. Our experiments show that the accuracy of our method can achieve SOTA on both GazeFollow dataset and VideoAttentionTarget dataset. At the same time, by using the ghostnet, the FLOPs of our method is only about 1/18 of other methods with the same accuracy. Further, sharing scene features saves nearly 40% of inference time in multi-person gaze-following task when more than 6 people in the frame. Tianlei Jin, Zheyuan Lin, Shiqiang Zhu, Wen Wang 0017, Shunda Hu |
FG | 3 |
| 2021 | Interesting Receptive Region and Feature Excitation for Partial Person Re-identification
Qiwei Meng, Shanshan Ji, Shiqiang Zhu, Jianjun Gu 0004 |
ICANN (4) | 4 |
| 2021 | Multi-branch Fusion Fully Convolutional Network for Person Re-Identification
Shanshan Ji, Shiqiang Zhu, Qiwei Meng, Jianjun Gu 0004 |
ICONIP (3) | 3 |
| 2021 | A Capturability-based Control Framework for the Underactuated Bipedal Walking *abstractThis work considers the control of underactuated bipedal walking, and a novel capturability-based control framework is presented. Compared with traditional approaches, the presented control method does not rely on the use of the Poincaré map, which may take significant computational cost. Firstly, a new definition of stable walking is presented, and a novel foot-placement based control method is proposed. Then, a controller design method is presented based on this control method. For the controller design, the foot placement adjustment is achieved by updating the virtual constraints using a heuristic method, and an improved virtual constraint control method is proposed to enforce the virtual constraints. Finally, the effectiveness of the presented control framework is illustrated on a five-link underactuated planar biped by numerical simulations. Haihui Yuan, Sumian Song, Ruilong Du, Shiqiang Zhu, Jason Gu, Mingguo Zhao, Jianxin Pang |
ICRA | 4 |
| 2021 | Automatic Pancreas Segmentation in CT Images With Distance-Based Saliency-Aware DenseASPP NetworkabstractPancreas identification and segmentation is an essential task in the diagnosis and prognosis of pancreas disease. Although deep neural networks have been widely applied in abdominal organ segmentation, it is still challenging for small organs (e.g. pancreas) that present low contrast, highly flexible anatomical structure and relatively small region. In recent years, coarse-to-fine methods have improved pancreas segmentation accuracy by using coarse predictions in the fine stage, but only object location is utilized and rich image context is neglected. In this paper, we propose a novel distance-based saliency-aware model, namely DSD-ASPP-Net, to fully use coarse segmentation to highlight the pancreas feature and boost accuracy in the fine segmentation stage. Specifically, a DenseASPP (Dense Atrous Spatial Pyramid Pooling) model is trained to learn the pancreas location and probability map, which is then transformed into saliency map through geodesic distance-based saliency transformation. In the fine stage, saliency-aware modules that combine saliency map and image context are introduced into DenseASPP to develop the DSD-ASPP-Net. The architecture of DenseASPP brings multi-scale feature representation and achieves larger receptive field in a denser way, which overcome the difficulties brought by variable object sizes and locations. Our method was evaluated on both public NIH pancreas dataset and local hospital dataset, and achieved an average Dice-Sørensen Coefficient (DSC) value of 85.49±4.77% on the NIH dataset, outperforming former coarse-to-fine methods. Peijun Hu, Xiang Li 0182, Yu Tian 0002, Tianyu Tang, Xueli Bai, Shiqiang Zhu, Tingbo Liang, Jingsong Li 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Multicenter Privacy-Preserving Cox Analysis Based on Homomorphic EncryptionabstractThe Cox proportional hazards model is one of the most widely used methods for analyzing survival data. Data from multiple data providers are required to improve the generalizability and confidence of the results of Cox analysis; however, such data sharing may result in leakage of sensitive information, leading to financial fraud, social discrimination or unauthorized data abuse. Some privacy-preserving Cox regression protocols have been proposed in past years, but they lack either security or functionality. In this paper, we propose a privacy-preserving Cox regression protocol for multiple data providers and researchers. The proposed protocol allows researchers to train models on horizontally or vertically partitioned datasets while providing privacy protection for both the sensitive data and the trained models. Our protocol utilizes threshold homomorphic encryption to guarantee security. Experimental results demonstrate that with the proposed protocol, Cox regression model training over 9 variables in a dataset of 113,035 samples takes approximately 44 min, and the trained model is almost the same as that obtained with the original nonsecure Cox regression protocol; therefore, our protocol is a potential candidate for practical real-world applications in multicenter medical research. Yao Lu 0009, Yu Tian 0002, Shiqiang Zhu, Jingsong Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | EHR-Oriented Knowledge Graph System: Toward Efficient Utilization of Non-Used Information Buried in Routine Clinical PracticeabstractNon-used clinical information has negative implications on healthcare quality. Clinicians pay priority attention to clinical information relevant to their specialties during routine clinical practices but may be insensitive or less concerned about information showing disease risks beyond their specialties, resulting in delayed and missed diagnoses or improper management. In this study, we introduced an electronic health record (EHR)-oriented knowledge graph system to efficiently utilize non-used information buried in EHRs. EHR data were transformed into a semantic patient-centralized information model under the ontology structure of a knowledge graph. The knowledge graph then creates an EHR data trajectory and performs reasoning through semantic rules to identify important clinical findings within EHR data. A graphical reasoning pathway illustrates the reasoning footage and explains the clinical significance for clinicians to better understand the neglected information. An application study was performed to evaluate unconsidered chronic kidney disease (CKD) reminding for non-nephrology clinicians to identify important neglected information. The study covered 71,679 patients in non-nephrology departments. The system identified 2,774 patients meeting CKD diagnosis criteria and 10,377 patients requiring high attention. A follow-up study of 5,439 patients showed that 82.1% of patients who met the diagnosis criteria and 61.4% of patients requiring high attention were confirmed to be CKD positive during follow-up research. The application demonstrated that the proposed approach is feasible and effective in clinical information utilization. Additionally, it's valuable as an explainable artificial intelligence to provide interpretable recommendations for specialist physicians to understand the importance of non-used data and make comprehensive decisions. Yong Shang, Yu Tian 0002, Kewei Lyu, Ran Xin, Tingbo Liang, Shiqiang Zhu, Jingsong Li 0001 |
IEEE J. Biomed. Health Informatics | 9 |
| 2020 | RBFNN-Based Adaptive Sliding Mode Control Design for Delayed Nonlinear Multilateral Telerobotic System With Cooperative ManipulationabstractMultilateral telerobotic system has potential applications in the industry environments with the advantages of cooperative manipulation for the remote and hazardous tasks, and its control design is quite challenging due to several coupling issues such as stability, position tracking, force feedback, and cooperative manipulation under time delays, various uncertainties, and external disturbance. In this paper, a novel radial basis function neural network (RBFNN) based adaptive sliding mode control design is proposed for nonlinear multilateral telerobotic system with n-master-n-slave manipulators. The environment force is modeled with a general form via the RBFNN-based environment parameters estimation in the slave side. The estimated environment parameters (nonpower signals) are transmitted to rebuild the environment dynamics in the master side and provide the good force feedback for the human operators. The RBFNN-based adaptive sliding mode controllers are designed separately for master and slave manipulators to achieve good position tracking under parameter variations and external disturbance. The coordinated force distribution algorithm is designed to achieve cooperative manipulation with the balance of force acting on the target object. The theoretical analysis is given and the comparative experiment for a nonlinear multilateral telerobotic system with 2-master-2-slave manipulators is implemented. The results show the good performance of our design. Zheng Chen 0004, Fanghao Huang, Weichao Sun, Jason Gu, Shiqiang Zhu |
IEEE Trans. Ind. Informatics | 8 |
| 2016 | Cascade force control of lower limb hydraulic exoskeleton for human performance augmentationabstractRecently the research on hydraulically actuated exoskeleton becomes an attractive topic for those application requirements of human performance augmentation. The control goal of this type of exoskeleton system is to minimize the human machine interaction force. And it becomes more challenging for hydraulically actuated lower limb exoskeleton where the multi-variable nonlinear dynamics is quite complicated and multiple walking phases are existing. Furthermore, since the exoskeleton is driven by the hydraulic actuators, the accurate output force tracking can not be easily realized due to the large compressibility of hydraulic oil. This paper focuses on the human machine interaction force control and the walking phase partition of the hydraulically actuated lower limb exoskeleton. Firstly, a cascade interaction force control strategy is proposed for a 3-DOF support leg which is the basic partitioned module of the lower limb exoskeleton. The spring model is built for the dynamics of human-machine interface, and a high level controller minimizing the integral of human-machine interaction force is designed to generate the desired joint trajectories of the exoskeleton which can be considered as the human motion intent. Subsequently, an independent joint based PID controller is developed in the low level to achieve the good tracking of the above generated human motion intent. Secondly, the exoskeleton system in different walking phases is partitioned into three serial chain manipulator modules. For each serial chain manipulator module, the proposed cascade interaction force controller is applied to minimize the human machine interaction force at the end effector. Finally, the walking experiments with 20Kg load on a practical hydraulically actuated lower limb exoskeleton are carried out to validate the effectiveness of the proposed approach. Shan Chen 0003, Zheng Chen 0004, Bin Yao 0001, Xiaocong Zhu, Shiqiang Zhu, Qingfeng Wang 0001 |
IECON | 5 |
| 2016 | Convolutional neural network based deep conditional random fields for stereo matching
Shiqiang Zhu, Zhengzhe Cui |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Edge-preserving guided filtering based cost aggregation for stereo matching
Shiqiang Zhu, Xuequn Zhang |
J. Vis. Commun. Image Represent. | 1 |
| 2010 | Saturated output feedback tracking control for robot manipulators via fuzzy self-tuningabstractThis paper concerns the problem of output feedback tracking (OFT) control with bounded torque inputs of robot manipulators, and proposes a novel saturated OFT controller based on fuzzy self-tuning proportional and derivative (PD) gains. First, aiming to accomplish the whole closed-loop control with only position measurements, a linear filter is involved to generate a pseudo velocity error signal. Second, different from previous strategies, the arctangent function with error-gain is applied to ensure the boundedness of the torque control input, and an explicit system stability proof is made by using the theory of singularly perturbed systems. Moreover, a fuzzy self-tuning PD regulator, which guarantees the continuous stability of the overall closed-loop system, is added to obtain an adaptive performance in tackling the disturbances during tracking control. Simulation showed that the proposed controller gains more satisfactory tracking results than the others, with a better dynamic response performance and stronger anti-disturbance capability. Huashan Liu, Shiqiang Zhu, Zhang-wei Chen |
J. Zhejiang Univ. Sci. C | 2 |