EDBT 2026 Demo / reviewers in the wild / expert
Jiyong Zhang 0001
dblp:95/6425-1
· DBLP profile ↗
74ranked-venue papers
4as first author
48since 2021 · last 2026
0000-0001-9600-8477ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 28 since 2021Artificial intelligence and machine learning · 27 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 4 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 3Human-computer interaction and ubiquitous computing · 3 · 1 first-authorTheory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object DetectionabstractRecently segment anything model (SAM) has attracted widespread concerns, and it is often treated as a vision foundation model for universal segmentation. Some researchers have attempted to directly apply the foundation model to the RGB-D video salient object detection (RGB-D VSOD) task, which often encounters three challenges, including the dependence on manual prompts, the high memory consumption of sequential adapters, and the computational burden of memory attention. To address the limitations, we propose a novel method, namely Segment Anything Model with Depth-guided Adaptive Queries (SAM-DAQ), which adapts SAM2 to pop-out salient objects from videos by seamlessly integrating depth and temporal cues within a unified framework. Firstly, we deploy a parallel adapter-based multi-modal image encoder (PAMIE), which incorporates several depth-guided parallel adapters (DPAs) in a skip-connection way. Remarkably, we fine-tune the frozen SAM encoder under prompt-free conditions, where the DPA utilizes depth cues to facilitate the fusion of multi-modal features. Secondly, we deploy a query-driven temporal memory (QTM) module, which unifies the memory bank and prompt embeddings into a learnable pipeline. Concretely, by leveraging both frame-level queries and video-level queries simultaneously, the QTM module can not only selectively extract temporal consistency features but also iteratively update the temporal representations of the queries. Extensive experiments are conducted on three RGB-D VSOD datasets, and the results show that the proposed SAM-DAQ consistently outperforms state-of-the-art methods in terms of all evaluation metrics. Xiaofei Zhou 0003, Runmin Cong, Guodao Zhang, Zhi Liu 0003, Jiyong Zhang 0001 |
AAAI | 7 |
| 2026 | Dynamic Channel Collaboration Framework for Panoramic Image Enhancement: A Neurobiologically-Inspired ApproachabstractPanoramic images are critical for immersive VR/AR and 6DoF yet degraded by compression artifacts, projection distortion, and uneven sampling, with existing hybrid CNN-Transformer models struggling to reconcile fine details and structural consistency in panoramas; to address this, we propose Dynamic Channel Collaboration (DCC-Former) for panoramic enhancement, inspired by primate vision's hierarchical processing and three strategies: strengthening local feature representation via reparameterization and gating, enhancing global context with adaptive self-attention, and enabling cross-scale aggregation through cascaded multi-scale fusion, aligned with biological vision's ventral-dorsal stream division and fovea-periphery resource allocation to balance detail preservation, global consistency, and computational efficiency-extensive experiments on benchmark datasets demonstrate DCC-Former outperforms SOTA in restoration quality and inference efficiency, providing a practical-efficient paradigm for high-resolution panoramic enhancement. Ziyi Cao, Hongkui Wang, Haibing Yin, Tiansong Li, Jiyong Zhang 0001, Xiaofeng Huang, Xia Wang 0006, Ruiyang Fu |
DCC | 5 |
| 2026 | AND-GS: Adaptive supervision of normal and depth in Gaussian splatting for accurate and efficient surface reconstruction
Xiang Le, Qiang Zhao 0005, Haofan Ren, Zhongtian Zheng, Tingyu Wang 0002, Jiyong Zhang 0001, Chenggang Yan 0001 |
Neurocomputing | 6 |
| 2026 | Multi-scale sampling and feature fusion for dynamic human rendering
Kainan Yu, Bolun Zheng, Qianyu Zhang 0002, Fangni Chen, Jiyong Zhang 0001, Canjin Wang |
J. Vis. Commun. Image Represent. | 6 |
| 2026 | Video Demoiréing With Spatial-Temporal Filtering in Frequency DomainabstractWhen acquiring images or videos of electronic displays, moiré patterns often arise due to aliasing between overlapping pixel grids, substantially compromising the perceptual quality of the captured content. Although frequency domain techniques have demonstrated high efficacy in image demoiréing, existing video approaches often overlook inter-frame frequency contextual relationships. This limitation restricts their capacity to achieve consistent temporal coherence and reconstruction fidelity. To overcome these challenges, we introduce a novel network (STFNet) with spatial-temporal filtering in frequency domain for video demoiréing. The proposed architecture comprises two dedicated stages: (1) Temporal-Guided Filtering (TGF), aims to adaptively incorporate temporal cues into learnable bandpass filters; and (2) Joint Filtering with Partially Shared Passbands (JFPS), which enhances representation learning of low-frequency moiré textures through strategic parameter sharing. Comprehensive evaluations on multiple public benchmarks confirm the superiority of our method. STFNet consistently outperforms state-of-the-art alternatives across both quantitative metrics and perceptual quality assessments, demonstrating robust performance in dynamic moiré suppression and detail preservation. Zhongqi Liu, Bolun Zheng, Qianyu Zhang 0002, Heng Jin, Qiankun Li 0005, Xu Jia 0012, Jiyong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | Reparameterization-Driven Depthwise Separable Large-Kernel Network for Lightweight Salient Object Detection of Strip Steel Surface DefectsabstractWith the rapid development of neural networks, strip steel surface defect detection, as an important task in computer vision, has achieved remarkable progress. However, state-of-the-art methods still face a tradeoff between accuracy and efficiency. High-performing models are usually large and computationally expensive, whereas lightweight models often suffer from limited detection accuracy. To address this issue, we first propose a spatial channel enhancement (SCE) module, which consists of a reparameterizable depthwise large-kernel convolution and a reparameterizable pointwise (RepPw) convolution. The proposed SCE module enlarges the receptive field and strengthens long-range spatial and channel interactions while preserving computational efficiency. Based on the SCE module, we propose a novel lightweight saliency model for strip steel surface defects, namely, reparameterization-driven depthwise separable large-kernel network (RepDSLKNet). RepDSLKNet employs SCE modules to build an encoder and a decoder, and utilizes cascaded channel attention (CCA) modules for the feature fusion. The lightweight architecture can effectively extract and fuse the semantic information and detailed features of strip steel surface defects, thereby improving the accuracy and speed of detection with a small model size. With an input size of $224 \times 224$ , our RepDSLKNet has only 0.47 M parameters and 0.42 G FLOPs during inference. Compared to the current state-of-the-art methods, our approach achieves a 19-fold improvement in throughput and a twofold reduction in latency. Experiments on two public strip steel defect datasets demonstrate that RepDSLKNet delivers competitive performance against state-of-the-art methods. Xiaofei Zhou 0003, Zhenkun Mo, Gongyang Li, Liuxin Bao, Xiaobin Xu 0002, Jiyong Zhang 0001 |
IEEE Trans. Cybern. | 6 |
| 2026 | Scale-Invariant Feature Matching Network for V-D-T Few-Shot Semantic SegmentationabstractMulti-modal few-shot semantic segmentation (FSS) aims to perform dense prediction from multiple modality images including visible image, depth image, and thermal image with a few annotated samples. However, some efforts treat the three modality information equally, where they don't incorporate the inherent differences among multiple modalities. Besides, the objects vary in size greatly, and the cutting-edge matching paradigms fail to establish an effective support-query connection. Therefore, we propose a novel scale-invariant feature matching network (i.e., SFM-Net), which consists of an encoder, a feature matching block, a feature elevation block, and a decoder, to conduct visible-depth-thermal (V-D-T) few-shot semantic segmentation. Firstly, in the encoder part, after the extraction of multi-level initial features, we fuse each level's RGB feature and thermal feature, yielding the support features and the query features. Secondly, in the feature matching block, a pixel-to-patch cross-attention (PTPCA) module is deployed to explore the correlation between each level's support feature and the query feature, where the pixel-to-patch pooling (PTP-pool) units are designed to build scale-invariant relationships, generating the coarse mask for the query image. Thirdly, in the feature elevation block, we employ the prior-related fusion (PF) module to integrate the depth image with a coarse mask via the cross-attention mechanism, yielding the enhanced coarse prediction result, which is further aggregated in a bottom-up way. Finally, in the decoder, we deploy a reverse attention (RA) unit to gradually explore the complementarity between object internal regions and spatial details, and further generate the final segmentation results via conventional convolution layers. Extensive experiments are conducted on the VDT-2048- $5^{i}$ dataset, and the results show that our model outperforms the state-of-the-art methods with a large margin. Xiaofei Zhou 0003, Deyang Liu, Jiyong Zhang 0001, Runmin Cong |
IEEE Trans. Image Process. | 5 |
| 2025 | PLGMNet: Parallel Local-Global Mamba Network for Real-Time Steel Surface Defect Detection
Chenlei Li, Xiaofei Zhou 0003, Yong Wu 0007, Deyang Liu, Jiyong Zhang 0001, Zhi Liu 0003 |
PRCV (17) | 6 |
| 2025 | Multi-modal feature integration network for Visible-Depth-Thermal salient object detection
Fengyv Cui, Xiaofei Zhou 0003, Liuxin Bao, Bin Wan, Jiyong Zhang 0001 |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Bidirectionally Guided Multi-Scale Feature Decoding Network for High-Resolution Salient Object DetectionabstractABSTRACT With the advancement of modern camera technology, the resolution and quality of images have been significantly improved. High‐resolution images can provide more detailed and clearer information, but they may also introduce more noise, which interferes with the detection and localization of salient objects. To address this issue, existing high‐resolution salient object detection methods either design complex network structures or adopt multi‐modal fusion. However, these approaches often consume significant computing and storage resources. This leads to redundancy of irrelevant features and loss of critical details. In this paper, we propose a network called bidirectionally guided multi‐scale feature decoding network for high‐resolution salient object detection. The model incorporates a bidirectional guidance method to explore the complementarity between encoding and decoding features, thereby achieving a comprehensive combination and enhancement of features. Additionally, in the decoder, multi‐scale encoding features are obtained and utilized sequentially to enhance feature learning and improve the accuracy of salient object detection. Specifically, our model consists of an encoder, a guided multi‐scale feature enhancement (GMFE) module, a guided feature fusion (GFF) module, and a multi‐scale feature decoder (MFD) module. First, multi‐scale encoding features are extracted through the encoder. These features are then fed into the GMFE module to enhance the multi‐scale encoding features under the guidance of saliency map derived from the decoding features of the previous layer. Subsequently, in the GFF module, the enhanced encoding features are fused with the decoding features from the previous layer. Finally, in the MFD module, the bidirectionally guided multi‐scale encoding features is integrated to generate an accurate saliency map. Experiments on two high‐resolution and two low‐resolution datasets demonstrate that our model outperforms on high‐resolution datasets while maintaining competitive performance on low‐resolution datasets, underscoring its effectiveness across varying image qualities. Jiangping Tang, Shuyao Guo, Xiaofei Zhou 0003, Liuxin Bao, Jiyong Zhang 0001 |
IET Image Process. | 5 |
| 2025 | Consistency perception network for 360° omnidirectional salient object detection
Hongfa Wen, Zunjie Zhu, Xiaofei Zhou 0003, Jiyong Zhang 0001, Chenggang Yan 0001 |
Neurocomputing | 4 |
| 2025 | Pyramid Learnable Bandpass Filters for Ultra-High-Definition Image DemoiréingabstractMoiré patterns usually depend on the style of display grids and the position of shooting camera, appearing in the form of stripes, meshes or ripples, with various and irregular colors. Compared with low-resolution moiré images, high-definition (HD) and ultra-high-definition (UHD) moiré images exhibit more complex moiré patterns, e.g., wider distribution of moiré frequencies and higher coupling degree of moirés of different scales, which poses a greater challenge to the modeling capabilities of the model. To address these challenges, we propose a novel Pyramid Learnable Bandpass Filtering Network (PBNet) for demoiréing UHD images. Specifically, we propose a pyramid learnable bandpass filter (P-LBF) to perform multi-scale filtering in the same semantic context to obtain richer frequency domain information. The P-LBF contains three stages: aligning, filtering and fusing. First, we introduce a pyramid alignment (DA) to align neighbor pixels for eliminating the deviations raised by different styles of display grids and relative position of the shooting camera. Then, a pyramid filtering (PF) is conducted to model the complex and variable moiré patterns with aligned neighbor pixels. Finally, the frequency domain responses of these different scales are fused with a multi-dimensional feature fusion (MFF). The PBNet is constructed based on the P-LBF, incorporating a cross-layer feature fusion (CLF) module to facilitate more effective information interaction between features at different depths. Extensive experiments on four public datasets show that our model achieves state-of-the-art performance for both high- and low-resolution moiré images. The code is publicly available at:https://github.com/liuzhongqi1/PBNet. Zhongqi Liu, Bolun Zheng, Qianyu Zhang 0002, Xu Jia 0012, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | IFENet: Interaction, Fusion, and Enhancement Network for V-D-T Salient Object DetectionabstractVisible-depth-thermal (VDT) salient object detection (SOD) aims to highlight the most visually attractive object by utilizing the triple-modal cues. However, existing models don't give sufficient exploration of the multi-modal correlations and differentiation, which leads to unsatisfactory detection performance. In this paper, we propose an interaction, fusion, and enhancement network (IFENet) to conduct the VDT SOD task, which contains three key steps including the multi-modal interaction, the multi-modal fusion, and the spatial enhancement. Specifically, embarking on the Transformer backbone, our IFENet can acquire multi-scale multi-modal features. Firstly, the inter-modal and intra-modal graph-based interaction (IIGI) module is deployed to explore inter-modal channel correlation and intra-modal long-term spatial dependency. Secondly, the gated attention-based fusion (GAF) module is employed to purify and aggregate the triple-modal features, where multi-modal features are filtered along spatial, channel, and modality dimensions, respectively. Lastly, the frequency split-based enhancement (FSE) module separates the fused feature into high-frequency and low-frequency components to enhance spatial information (i.e., boundary details and object location) of the salient object. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art models. Our code and results are available at https://github.com/Lx-Bao/IFENet. Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Runmin Cong, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | AGFNet: Adaptive Gated Fusion Network for RGB-T Semantic SegmentationabstractRGB-T semantic segmentation can effectively pop-out objects from challenging scenarios (e.g., low illumination and low contrast environments) by combining RGB and thermal infrared images. However, the existing cutting-edge RGB-T semantic segmentation methods often present insufficient exploration of multi-modal feature fusion, where they overlook the differences between the two modalities. In this paper, we propose an adaptive gated fusion network (AGFNet) to conduct RGB-T semantic segmentation, where the multi-modal features are combined via the gating mechanisms and the spatial details are enhanced via the introduction of edge information. Specifically, the AGFNet employs a cross-modal adaptive gated-attention fusion (CAGF) module to aggregate the RGB and thermal features, where we give a sufficient exploration of the complementarity between the two-modal features via the gated attention unit (GAU). Particularly, in GAU, the gates can be used to purify the features, and the channel and spatial attention mechanisms are further employed to enhance the two-modal features interactively. Then, we design an edge detection (ED) module to learn the object-related edge cues, which simultaneously incorporates local detail information from low-level features and global location information from high-level features. After that, we deploy the edge guidance (EG) module to emphasize the spatial details of the fused features. Next, we deploy the contextual elevation (CE) module to enrich the contextual information of features by iteratively introducing the sine and cosine functions. Finally, considering that the quality of thermal images is usually lower than that of RGB images, we progressively integrate the multi-level RGB encoder features with multi-level decoder features, thereby focusing more on appearance information. Following this way, we can acquire the final high-quality segmentation result. Extensive experiments are performed on three public datasets including MFNet, PST900 and FMB datasets, and the experimental results show that our method achieves competitive performance when compared with the 22 state-of-the-art methods. Xiaofei Zhou 0003, Liuxin Bao, Haibing Yin, Qiuping Jiang, Jiyong Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | It Takes Two: Accurate Gait Recognition in the Wild via Cross-granularity AlignmentabstractExisting studies for gait recognition primarily utilized sequences of either binary silhouette or human parsing to encode the shapes and dynamics of persons during walking. Silhouettes exhibit accurate segmentation quality and robustness to environmental variations, but their low information entropy may result in sub-optimal performance. In contrast, human parsing provides fine-grained part segmentation with higher information entropy, but the segmentation quality may deteriorate due to the complex environments. To discover the advantages of silhouette and parsing and overcome their limitations, this paper proposes a novel cross-granularity alignment gait recognition method, named XGait, to unleash the power of gait representations of different granularity. To achieve this goal, the XGait first contains two branches of backbone encoders to map the silhouette sequences and the parsing sequences into two latent spaces, respectively. Moreover, to explore the complementary knowledge across the features of two representations, we design the Global Cross-granularity Module (GCM) and the Part Cross-granularity Module (PCM) after the two encoders. In particular, the GCM aims to enhance the quality of parsing features by leveraging global features from silhouettes, while the PCM aligns the dynamics of human parts between silhouette and parsing features using the high information entropy in parsing sequences. In addition, to effectively guide the alignment of two representations with different granularity at the part level, an elaborate-designed learnable division mechanism is proposed for the parsing features. Finally, comprehensive experiments on two large-scale gait datasets not only show the superior performance of XGait with the Rank-1 accuracy of 80.5% on Gait3D and 88.3% CCPG but also reflect the robustness of the learned features even under challenging conditions like occlusions and cloth changes Jinkai Zheng, Xinchen Liu, Boyue Zhang 0004, Chenggang Yan 0001, Jiyong Zhang 0001, Wu Liu 0005, Yongdong Zhang 0001 |
ACM Multimedia | 5 |
| 2024 | GINet:Graph interactive network with semantic-guided spatial refinement for salient object detection in optical remote sensing images
Chenwei Zhu, Xiaofei Zhou 0003, Liuxin Bao, Hongkui Wang, Shuai Wang 0003, Zunjie Zhu, Chenggang Yan 0001, Jiyong Zhang 0001 |
J. Vis. Commun. Image Represent. | 8 |
| 2024 | Interactive Fusion and Correlation Network for Three-Modal Images Few-Shot Semantic SegmentationabstractThis letter presents a novel method for three-modal images few-shot semantic segmentation. Some previous efforts fuse multiple modalities before feature correlation, while this changes the original visual information that is useful to subsequent feature matching. Others are built based on early correlation learning, which can cause details loss and thereby defects multi-modal integration. To address these challenges, we build a novel interactive fusion and correlation network (IFCNet). Specifically, the proposed fusing and correlating (FC) module performs feature correlating and attention-based multi-modal fusing interactively, which establishes effective inter-modal complementarity and benefits intra-modal query-support correlation. Furthermore, we add a multi-modal correlation (MC) module, which leverages multi-layer cosine similarity maps to enrich multi-modal visual correspondence. Experiments on the VDT-2048-5$^{i}$dataset demonstrate the network's superior performance, which outperforms existing state-of-the-art methods in both 1-shot and 5-shot settings. The study also includes an ablation analysis to validate the contributions of the FC module and the MC module to the overall segmentation accuracy. Haolan He, Xianguo Dong, Xiaofei Zhou 0003, Bo Wang 0031, Jiyong Zhang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Quality-Aware Selective Fusion Network for V-D-T Salient Object DetectionabstractDepth images and thermal images contain the spatial geometry information and surface temperature information, which can act as complementary information for the RGB modality. However, the quality of the depth and thermal images is often unreliable in some challenging scenarios, which will result in the performance degradation of the two-modal based salient object detection (SOD). Meanwhile, some researchers pay attention to the triple-modal SOD task, namely the visible-depth-thermal (VDT) SOD, where they attempt to explore the complementarity of the RGB image, the depth image, and the thermal image. However, existing triple-modal SOD methods fail to perceive the quality of depth maps and thermal images, which leads to performance degradation when dealing with scenes with low-quality depth and thermal images. Therefore, in this paper, we propose a quality-aware selective fusion network (QSF-Net) to conduct VDT salient object detection, which contains three subnets including the initial feature extraction subnet, the quality-aware region selection subnet, and the region-guided selective fusion subnet. Firstly, except for extracting features, the initial feature extraction subnet can generate a preliminary prediction map from each modality via a shrinkage pyramid architecture, which is equipped with the multi-scale fusion (MSF) module. Then, we design the weakly-supervised quality-aware region selection subnet to generate the quality-aware maps. Concretely, we first find the high-quality and low-quality regions by using the preliminary predictions, which further constitute the pseudo label that can be used to train this subnet. Finally, the region-guided selective fusion subnet purifies the initial features under the guidance of the quality-aware maps, and then fuses the triple-modal features and refines the edge details of prediction maps through the intra-modality and inter-modality attention (IIA) module and the edge refinement (ER) module, respectively. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art methods with a large margin. Our code and results are available at https://github.com/Lx-Bao/QSFNet. Liuxin Bao, Xiaofei Zhou 0003, Xiankai Lu, Yaoqi Sun, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Image Process. | 7 |
| 2023 | 360$^{\circ }$ Omnidirectional Salient Object Detection with Multi-scale Interaction and Densely-Connected Prediction
Haowei Dai, Liuxin Bao, Kunye Shen, Xiaofei Zhou 0003, Jiyong Zhang 0001 |
ICIG (1) | 5 |
| 2023 | GFNet: gated fusion network for video saliency prediction
Songhe Wu, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001 |
Appl. Intell. | 6 |
| 2023 | SMINet: Semantics-aware multi-level feature interaction network for surface defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Haibing Yin, Ji Hu 0002, Jiyong Zhang 0001, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 7 |
| 2023 | Aggregating transformers and CNNs for salient object detection in optical remote sensing images
Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Haibing Yin, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001 |
Neurocomputing | 6 |
| 2023 | STI-Net: Spatiotemporal integration network for video saliency detection
Xiaofei Zhou 0003, Weipeng Cao, Hanxiao Gao, Zhong Ming 0001, Jiyong Zhang 0001 |
Inf. Sci. | 5 |
| 2023 | SRI-Net: Similarity retrieval-based inference network for light field salient object detection
Chengtao Lv, Xiaofei Zhou 0003, Deyang Liu, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2023 | CANet: Context-aware Aggregation Network for Salient Object Detection of Surface Defects
Bin Wan, Xiaofei Zhou 0003, Mang Xiao, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 7 |
| 2023 | Caps-SSENet: An Improved Estimation Method for SAR Ship SizeabstractAccurate estimation of the sizes of ship targets plays a critical role in the task of ship classification in synthetic aperture radar (SAR) images. Existing deep neural networks (DNNs)-based methods for SAR ship size estimation (SSE) often adopt a fully connected structure that has limited capability in accurately modeling the relationships of features extracted from SAR images, leading to degraded performance of size estimation. It has been demonstrated that capsule networks provide new guidelines to capture relationships of image features by replacing traditional neurons with capsules, where the dynamic routing strategy is used to calculate correlations among capsules. In this letter, we propose an improved method for SAR SSE based on the capsule network named Caps-SSE network (SSENet). In our Caps-SSENet, a capsule-neural-mixing size mapping module is designed to transform the extracted image features into capsules and complete the estimation of ship sizes using informative feature correlations from dynamic routing. In addition, an average scaled mean square error (ASMSE) loss is proposed to improve the size estimation performance of small ships. Experimental results based on measured SAR data show that the proposed method reduces the estimation error of ship sizes in SAR images in comparison with the existing state-of-the-art method. Yu Liu 0005, Xueqian Wang 0002, Zhizhuo Jiang, Gang Li 0008, Bolun Zheng, Jiyong Zhang 0001, You He 0003 |
IEEE Geosci. Remote. Sens. Lett. | 8 |
| 2023 | Transformer-Based Multi-Scale Feature Integration Network for Video Saliency PredictionabstractMost cutting-edge video saliency prediction models rely on spatiotemporal features extracted by 3D convolutions due to its local contextual cues acquirement ability. However, the shortage of 3D convolutions is that it cannot effectively capture long-term spatiotemporal dependencies in videos. To address this limitation, we propose a novel Transformer-based Multi-scale Feature Integration Network (TMFI-Net) for video saliency prediction, where the proposed TMFI-Net consists of a semantic-guided encoder and a hierarchical decoder. Firstly, embarking on the Transformer-based multi-level spatiotemporal features, the semantic-guided encoder enhances the features by inserting the high-level feature into each level feature via a top-down pathway and a longitudinal connection, which endows the multi-level spatiotemporal features with rich contextual information. In this way, the features are steered to give more concerns to saliency regions. Secondly, the hierarchical decoder employs a multi-dimensional attention (MA) module to elevate features along channel, temporal, and spatial dimensions jointly. Successively, the hierarchical decoder deploys a progressive decoding block to conduct an initial saliency prediction, which provides a coarse localization of saliency regions. Lastly, considering the complementarity of different saliency predictions, we integrate all initial saliency prediction results into the final saliency map. Comprehensive experimental results on four video saliency datasets firmly demonstrate that our model achieves superior performance when compared with the state-of-the-art video saliency models. The code is available athttps://github.com/wusonghe/TMFI-Net. Xiaofei Zhou 0003, Songhe Wu, Bolun Zheng, Shuai Wang 0003, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Edge-Guided Recurrent Positioning Network for Salient Object Detection in Optical Remote Sensing ImagesabstractOptical remote sensing images (RSIs) have been widely used in many applications, and one of the interesting issues about optical RSIs is the salient object detection (SOD). However, due to diverse object types, various object scales, numerous object orientations, and cluttered backgrounds in optical RSIs, the performance of the existing SOD models often degrade largely. Meanwhile, cutting-edge SOD models targeting optical RSIs typically focus on suppressing cluttered backgrounds, while they neglect the importance of edge information which is crucial for obtaining precise saliency maps. To address this dilemma, this article proposes an edge-guided recurrent positioning network (ERPNet) to pop-out salient objects in optical RSIs, where the key point lies in the edge-aware position attention unit (EPAU). First, the encoder is used to give salient objects a good representation, that is, multilevel deep features, which are then delivered into two parallel decoders, including: 1) an edge extraction part and 2) a feature fusion part. The edge extraction module and the encoder form a U-shape architecture, which not only provides accurate salient edge clues but also ensures the integrality of edge information by extra deploying the intraconnection. That is to say, edge features can be generated and reinforced by incorporating object features from the encoder. Meanwhile, each decoding step of the feature fusion module provides the position attention about salient objects, where position cues are sharpened by the effective edge information and are used to recurrently calibrate the misaligned decoding process. After that, we can obtain the final saliency map by fusing all position attention cues. Extensive experiments are conducted on two public optical RSIs datasets, and the results show that the proposed ERPNet can accurately and completely pop-out salient objects, which consistently outperforms the state-of-the-art SOD models. Xiaofei Zhou 0003, Kunye Shen, Li Weng, Runmin Cong, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Cybern. | 6 |
| 2022 | A Novel RVFL-Based Algorithm Selection Approach for Software Model Checking
Weipeng Cao, Yuhao Wu 0001, Qiang Wang 0020, Jiyong Zhang 0001, Meikang Qiu |
KSEM (3) | 4 |
| 2022 | Evolutionary Multi-objective Architecture Search Framework: Application to COVID-19 3D CT Classification
Xin He 0019, Guohao Ying, Jiyong Zhang 0001, Xiaowen Chu 0001 |
MICCAI (1) | 3 |
| 2022 | Gait Recognition in the Wild with Multi-hop Temporal SwitchabstractExisting studies for gait recognition are dominated by in-the-lab scenarios. Since people live in real-world senses, gait recognition in the wild is a more practical problem that has recently attracted the attention of the community of multimedia and computer vision. Current methods that obtain state-of-the-art performance on in-the-lab benchmarks achieve much worse accuracy on the recently proposed in-the-wild datasets because these methods can hardly model the varied temporal dynamics of gait sequences in unconstrained scenes. Therefore, this paper presents a novel multi-hop temporal switch method to achieve effective temporal modeling of gait patterns in real-world scenes. Concretely, we design a novel gait recognition network, named Multi-hop Temporal Switch Network (MTSGait), to learn spatial features and multi-scale temporal features simultaneously. Different from existing methods that use 3D convolutions for temporal modeling, our MTSGait models the temporal dynamics of gait sequences by 2D convolutions. By this means, it achieves high efficiency with fewer model parameters and reduces the difficulty in optimization compared with 3D convolution-based models. Based on the specific design of the 2D convolution kernels, our method can eliminate the misalignment of features among adjacent frames. In addition, a new sampling strategy, i.e., non-cyclic continuous sampling, is proposed to make the model learn more robust temporal features. Finally, the proposed method achieves superior performance on two public gait in-the-wild datasets, i.e., GREW and Gait3D, compared with state-of-the-art methods. Jinkai Zheng, Xinchen Liu, Xiaoyan Gu 0001, Yaoqi Sun, Chuang Gan 0001, Jiyong Zhang 0001, Wu Liu 0005, Chenggang Yan 0001 |
ACM Multimedia | 6 |
| 2022 | A hybrid controller for safe and efficient longitudinal collision avoidance control
Qiang Wang 0020, Xinlei Zheng, Jiyong Zhang 0001, Joseph Sifakis |
J. Syst. Archit. | 3 |
| 2022 | Fully Squeezed Multiscale Inference Network for Fast and Accurate Saliency Detection in Optical Remote-Sensing ImagesabstractRecently, salient object detection in optical remote-sensing images (RSIs) has received more and more attention. To tackle the challenges of RSIs including large-scale variation of objects, cluttered background, irregular shape of objects, and big difference in illumination, the cutting-edge convolutional neural network (CNN)-based models are proposed and have achieved an encouraging performance. However, the performance of the top-level models usually depends on the large model size and high computational cost, which limits their practical applications. To remedy the issue, we introduce a fully squeezed multiscale (FSM) module to equip the entire network. Specifically, the FSM module squeezes the feature maps from high dimension to low dimension and introduces the multiscale strategy to endow the capability of feature characterization with different receptive fields and different contexts. Based on the FSM module, we build the FSM inference network (FSMI-Net) to pop-out salient objects from optical RSIs, which is with fewer parameters and fast inference speed. Particularly, the proposed FSMI-Net only contains 3.6M parameters, and its GPU running speed is about 28 fps for$384 \times 384$inputs, which is superior to the existing saliency models targeting optical RSIs. Extensive comparisons are performed on two public optical RSIs datasets, and our FSMI-Net achieves comparable detection accuracy when compared with the state-of-the-art models, where our model realizes a balance between the computational cost and detection performance. Kunye Shen, Xiaofei Zhou 0003, Bin Wan, Jiyong Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | An Adaptive IoT Network Security Situation Prediction Model
Hongyu Yang 0003, Xugao Zhang, Jiyong Zhang 0001 |
Mob. Networks Appl. | 4 |
| 2022 | Learning Frequency Domain Priors for Image DemoireingabstractImage demoireing is a multi-faceted image restoration task involving both moire pattern removal and color restoration. In this paper, we raise a general degradation model to describe an image contaminated by moire patterns, and propose a novel multi-scale bandpass convolutional neural network (MBCNN) for single image demoireing. For moire pattern removal, we propose a multi-block-size learnable bandpass filters (M-LBFs), based on a block-wise frequency domain transform, to learn the frequency domain priors of moire patterns. We also introduce a new loss function named Dilated Advanced Sobel loss (D-ASL) to better sense the frequency information. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, and then performs local fine tuning of the color per pixel. To determine the most appropriate frequency domain transform, we investigate several transforms including DCT, DFT, DWT, learnable non-linear transform and learnable orthogonal transform. We finally adopt the DCT. Our basic model won the AIM2019 demoireing challenge. Experimental results on three public datasets show that our method outperforms state-of-the-art methods by a large margin. Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Xiang Tian 0002, Jiyong Zhang 0001, Yaoqi Sun, Lin Liu 0016, Ales Leonardis, Gregory Slabaugh |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | FANet: Feature aggregation network for RGBD saliency detection
Xiaofei Zhou 0003, Hongfa Wen, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
Signal Process. Image Commun. | 5 |
| 2022 | Each Part Matters: Local Patterns Facilitate Cross-View Geo-LocalizationabstractCross-view geo-localization is to spot images of the same geographic target from different platforms,e.g., drone-view cameras and satellites. It is challenging in the large visual appearance changes caused by extreme viewpoint variations. Existing methods usually concentrate on mining the fine-grained feature of the geographic target in the image center, but underestimate the contextual information in neighbor areas. In this work, we argue that neighbor areas can be leveraged as auxiliary information, enriching discriminative clues for geo-localization. Specifically, we introduce a simple and effective deep neural network, called Local Pattern Network (LPN), to take advantage of contextual information in an end-to-end manner. Without using extra part estimators, LPN adopts a square-ring feature partition strategy, which provides the attention according to the distance to the image center. It eases the part matching and enables the part-wise representation learning. Owing to the square-ring partition design, the proposed LPN has good scalability to rotation variations and achieves competitive results on three prevailing benchmarks,i.e., University-1652, CVUSA and CVACT. Besides, we also show the proposed LPN can be easily embedded into other frameworks to further boost performance. Tingyu Wang 0002, Zhedong Zheng, Chenggang Yan 0001, Jiyong Zhang 0001, Yaoqi Sun, Bolun Zheng, Yi Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Edge-Aware Multiscale Feature Integration Network for Salient Object Detection in Optical Remote Sensing ImagesabstractThe optical remote sensing images (RSIs) show various spatial resolutions and cluttered background, where salient objects with different scales, types, and orientations are presented in diverse RSI scenes. Therefore, it is inappropriate to directly extend cutting-edge saliency detection methods for conventional RGB images to optical RSIs. Besides, the existing saliency models targeting RSIs often render imperfect saliency maps, where some of them are with coarse boundary details. To solve this problem, this article attempts to introduce the edge information to precisely detect salient objects in RSIs. Accordingly, we propose an edge-aware multiscale feature integration network (EMFI-Net) for salient object detection by conducting multiscale feature integration under the explicit and implicit assistance of salient edge cues. Specifically, our network contains two parts including the encoder and decoder. First, the encoder extracts multiscale deep features from three RSIs with different resolutions, where the high-level deep semantic features from three RSIs are integrated using a cascaded feature fusion module. Second, the encoder explicitly enriches the multiscale deep features by integrating the salient edge cues extracted by a salient edge extraction module. Meanwhile, we also implicitly deploy an edge-aware constraint to the supervision of the saliency map prediction by introducing a hybrid loss function. Finally, the decoder integrates the enriched multiscale deep features in a coarse-to-fine way, yielding a high-quality saliency map. The experiments conducted on two public optical RSI datasets clearly prove the effectiveness and superiority of the proposed EMFI-Net against the state-of-the-art saliency models. Xiaofei Zhou 0003, Kunye Shen, Zhi Liu 0003, Chen Gong 0002, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Age-Invariant Face Recognition by Multi-Feature Fusionand Decomposition with Self-attentionabstractDifferent from general face recognition, age-invariant face recognition (AIFR) aims at matching faces with a big age gap. Previous discriminative methods usually focus on decomposing facial feature into age-related and age-invariant components, which suffer from the loss of facial identity information. In this article, we propose a novel Multi-feature Fusion and Decomposition (MFD) framework for age-invariant face recognition, which learns more discriminative and robust features and reduces the intra-class variants. Specifically, we first sample multiple face images of different ages with the same identity as a face time sequence. Then, the multi-head attention is employed to capture contextual information from facial feature series, extracted by the backbone network. Next, we combine feature decomposition with fusion based on the face time sequence to ensure that the final age-independent features effectively represent the identity information of the face and have stronger robustness against the aging process. Besides, we also mitigate imbalanced age distribution in the training data by a re-weighted age loss. We experimented with the proposed MFD over the popular CACD and CACD-VS datasets, where we show that our approach improves the AIFR performance than previous state-of-the-art methods. We simultaneously show the performance of MFD on LFW dataset. Chenggang Yan 0001, Lixuan Meng, Liang Li 0003, Jian Yin 0003, Jiyong Zhang 0001, Yaoqi Sun, Bolun Zheng |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2021 | Automated Model Design and Benchmarking of Deep Learning Models for COVID-19 Detection with Chest CT ScansabstractThe COVID-19 pandemic has spread globally for several months. Because its transmissibility and high pathogenicity seriously threaten people's lives, it is crucial to accurately and quickly detect COVID-19 infection. Many recent studies have shown that deep learning (DL) based solutions can help detect COVID-19 based on chest CT scans. However, most existing work focuses on 2D datasets, which may result in low quality models as the real CT scans are 3D images. Besides, the reported results span a broad spectrum on different datasets with a relatively unfair comparison. In this paper, we first use three state-of-the-art 3D models (ResNet3D101, DenseNet3D121, and MC3\_18) to establish the baseline performance on three publicly available chest CT scan datasets. Then we propose a differentiable neural architecture search (DNAS) framework to automatically search the 3D DL models for 3D chest CT scans classification and use the Gumbel Softmax technique to improve the search efficiency. We further exploit the Class Activation Mapping (CAM) technique on our models to provide the interpretability of the results. The experimental results show that our searched models (CovidNet3D) outperform the baseline human-designed models on three datasets with tens of times smaller model size and higher accuracy. Furthermore, the results also verify that CAM can be well applied in CovidNet3D for COVID-19 datasets to provide interpretability for medical diagnosis. Code: https://github.com/HKBU-HPML/CovidNet3D. Xin He 0019, Xiaowen Chu 0001, Shaohuai Shi, Jiangping Tang, Xin Liu 0027, Chenggang Yan 0001, Jiyong Zhang 0001, Guiguang Ding |
AAAI | 8 |
| 2021 | TraND: Transferable Neighborhood Discovery for Unsupervised Cross-Domain Gait RecognitionabstractGait, i.e., the movement pattern of human limbs during locomotion, is a promising biometrie for identification of persons. Despite significant improvement in gait recognition with deep learning, existing studies still neglect a more practical but challenging scenario - unsupervised cross-domain gait recognition which aims to learn a model on a labeled dataset then adapt it to an unlabeled dataset. Due to the domain shift and class gap, directly applying a model trained on one source dataset to other target datasets usually obtains very poor results. Therefore, this paper proposes a Transferable Neighborhood Discovery (TraND) framework to bridge the domain gap for unsupervised cross-domain gait recognition. To learn effective prior knowledge for gait representation, we first adopt a backbone network pre- trained on the labeled source data in a supervised manner. Then we design an end-to-end trainable approach to automatically discover the confident neighborhoods of unlabeled samples in the latent space. During training, the class consistency indicator is adopted to select confident neighborhoods of samples based on their entropy measurements. Moreover, we explore a high- entropy-first neighbor selection strategy, which can effectively transfer prior knowledge to the target domain. Our method achieves the state-of-the-art results on two public datasets, i.e., CASIA-B and OU-LP. Jinkai Zheng, Xinchen Liu, Chenggang Yan 0001, Jiyong Zhang 0001, Wu Liu 0005, Xiao-Ping Zhang 0002, Tao Mei 0001 |
ISCAS | 4 |
| 2021 | Heuristic Depth Estimation with Progressive Depth Reconstruction and Confidence-Aware LossabstractRecently deep learning-based depth estimation has shown the promising result, especially with the help of sparse depth reference samples. Existing works focus on directly inferring the depth information from sparse samples with high confidence. In this paper, we propose a Heuristic Depth Estimation Network (HDEN) with progressive depth reconstruction and confidence-aware loss. The HDEN leverages the reference samples with low confidence to distill the spatial geometric and local semantic information for dense depth prediction. Specifically, we first train a U-NET network to generate a coarse-level dense reference map. Second, the progressive depth reconstruction module successively reconstructs the fine-level dense depth map from different scales, where a multi-level upsampling block is designed to recover the local structure of object. Finally, the confidence-aware loss is proposed to trigger the reference samples with low confidence, which enforces the model focusing on estimating the depth of the tiny structure. Extensive experiments on the NYU-Depth-v2 and KITTI-Odometry dataset show the effectiveness of our method. Visualization results demonstrate that the dense depth maps generated by HDEN have better consistency at the entity edge with RGB image. Liang Li 0003, Chenggang Yan 0001, Yaoqi Sun, Tao Shen 0004, Jiyong Zhang 0001 |
ACM Multimedia | 6 |
| 2021 | Cross-modal semantic correlation learning by Bi-CNN networkabstractAbstract Cross modal retrieval can retrieve images through a text query and vice versa. In recent years, cross modal retrieval has attracted extensive attention. The purpose of most now available cross modal retrieval methods is to find a common subspace and maximize the different modal correlation. To generate specific representations consistent with cross modal tasks, this paper proposes a novel cross modal retrieval framework, which integrates feature learning and latent space embedding. In detail, we proposed a deep CNN and a shallow CNN to extract the feature of the samples. The deep CNN is used to extract the representation of images, and the shallow CNN uses a multi‐dimensional kernel to extract multi‐level semantic representation of text. Meanwhile, we enhance the semantic manifold by constructing cross modal ranking and within‐modal discriminant loss to improve the division of semantic representation. Moreover, the most representative samples are selected by using online sampling strategy, so that the approach can be implemented on a large‐scale data. This approach not only increases the discriminative ability among different categories, but also maximizes the relativity between different modalities. Experiments on three real word datasets show that the proposed method is superior to the popular methods. Liang Li 0003, Chenggang Yan 0001, Yaoqi Sun, Jiyong Zhang 0001 |
IET Image Process. | 6 |
| 2021 | Geometric attentional dynamic graph convolutional neural networks for point cloud analysis
Yiming Cui 0002, Xin Liu 0027, Hongmin Liu 0001, Jiyong Zhang 0001, Alina Zare, Bin Fan 0001 |
Neurocomputing | 4 |
| 2021 | Leveraging graph neural networks for point-of-interest recommendations
Jiyong Zhang 0001, Xin Liu 0027, Xiaofei Zhou 0003, Xiaowen Chu 0001 |
Neurocomputing | 1 |
| 2021 | An integrated classification model for incremental learning
Ji Hu 0002, Chenggang Yan 0001, Xin Liu 0027, Chengwei Ren, Jiyong Zhang 0001, Dongliang Peng 0001, Yi Yang 0001 |
Multim. Tools Appl. | 6 |
| 2021 | Dynamic Selective Network for RGB-D Salient Object DetectionabstractRGB-D saliency detection is receiving more and more attention in recent years. There are many efforts have been devoted to this area, where most of them try to integrate the multi-modal information, i.e. RGB images and depth maps, via various fusion strategies. However, some of them ignore the inherent difference between the two modalities, which leads to the performance degradation when handling some challenging scenes. Therefore, in this paper, we propose a novel RGB-D saliency model, namely Dynamic Selective Network (DSNet), to perform salient object detection (SOD) in RGB-D images by taking full advantage of the complementarity between the two modalities. Specifically, we first deploy a cross-modal global context module (CGCM) to acquire the high-level semantic information, which can be used to roughly locate salient objects. Then, we design a dynamic selective module (DSM) to dynamically mine the cross-modal complementary information between RGB images and depth maps, and to further optimize the multi-level and multi-scale information by executing the gated and pooling based selection, respectively. Moreover, we conduct the boundary refinement to obtain high-quality saliency maps with clear boundary details. Extensive experiments on eight public RGB-D datasets show that the proposed DSNet achieves a competitive and excellent performance against the current 17 state-of-the-art RGB-D SOD models. Hongfa Wen, Chenggang Yan 0001, Xiaofei Zhou 0003, Runmin Cong, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Yongjun Bao, Guiguang Ding |
IEEE Trans. Image Process. | 7 |
| 2021 | Deep Unsupervised Binary Descriptor Learning Through Locality Consistency and Self DistinctivenessabstractDeep learning has been successfully applied to learn local feature descriptors in recent years. However, most of existing methods are supervised methods relying on a large number of labeled training patches, which are also proposed for learning real valued descriptors. In this paper, we propose a novel unsupervised deep learning method for binary descriptor learning. The binary descriptors are much more compact and efficient than the real valued descriptors and unsupervised leaning is highly required in many applications due to its label-free characteristic as the annotations are sometimes expensive to obtain. The core idea of our method is to explore the locality consistency in the descriptor space as well as to distinguish different patches while maintaining the ability to match a patch with its geometric transformed ones. We also give a theorical analysis about the role of batch normalization in learning effective binary descriptors. Benefited from this analysis, there is no need to append two additional losses on minimizing the quantization error and maximizing the entropy to the final learning objective like previous works did, thus simplifying our network training. Experiments on four benchmarks demonstrate that the proposed method is able to learn binary descriptors significantly outperforming previous unsupervised binary descriptors, even superior to most supervised ones. Especially, it obtains 21.2% of improvement on the UBC Phototour dataset, and 19.8%, 26.7%, 26.0% of improvements for patch verification, matching, retrieval tasks respectively on the HPatches dataset compared to the previous best unsupervised method. Bin Fan 0001, Hongmin Liu 0001, Hui Zeng 0003, Jiyong Zhang 0001, Xin Liu 0027, Junwei Han 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Heterogeneous Transfer Learning with Weighted Instance-Correspondence DataabstractInstance-correspondence (IC) data are potent resources for heterogeneous transfer learning (HeTL) due to the capability of bridging the source and the target domains at the instance-level. To this end, people tend to use machine-generated IC data, because manually establishing IC data is expensive and primitive. However, existing IC data machine generators are not perfect and always produce the data that are not of high quality, thus hampering the performance of domain adaption. In this paper, instead of improving the IC data generator, which might not be an optimal way, we accept the fact that data quality variation does exist but find a better way to use the data. Specifically, we propose a novel heterogeneous transfer learning method named Transfer Learning with Weighted Correspondence (TLWC), which utilizes IC data to adapt the source domain to the target domain. Rather than treating IC data equally, TLWC can assign solid weights to each IC data pair depending on the quality of the data. We conduct extensive experiments on HeTL datasets and the state-of-the-art results verify the effectiveness of TLWC. Xiaoming Jin, Guiguang Ding, Jungong Han, Jiyong Zhang 0001, Sicheng Zhao |
AAAI | 6 |
| 2020 | Research Progress of Zero-Shot Learning Beyond Computer Vision
Weipeng Cao, Yuhao Wu 0001, Zhong Ming 0001, Zhiwu Xu 0001, Jiyong Zhang 0001 |
ICA3PP (2) | 6 |
| 2020 | Distributed and Parallel Ensemble Classification for Big Data Based on Kullback-Leibler Random Sample Partition
Chenghao Wei, Jiyong Zhang 0001, Timur Valiullin, Weipeng Cao, Qiang Wang 0020 |
ICA3PP (1) | 2 |
| 2020 | A Variational Generative Network Based Network Threat Situation Assessment
Hongyu Yang 0003, Renyun Zeng, Fengyan Wang, Guangquan Xu, Jiyong Zhang 0001 |
ICICS | 5 |
| 2020 | A General Re-Ranking Method Based On Metric Learning For Person Re-IdentificationabstractWhen Person Re-identification is considered as a retrieval task, re-ranking becomes a critical part of improving the re-identification accuracy. Most of the existing re-ranking methods focus on k -nearest neighbors, which requires a lot of queries and memory. In this paper, we propose a Feature Relation Map based Similarity Evaluation (FRM-SE) model to tackle this problem. The Feature Relation Map is utilized to automatically mine the latent relation between the k -neighbors through convolution operation. The re-ranking distance is learned through the FRM-SE model with metric learning. Further, we optimize the existing re-ranking method to utilize the advantage of the FRM-SE model for maintaining a balance between accuracy and complexity. The proposed approach is validated on two benchmark datasets, Market1501 and CUHK03. Results show that our re-ranking method is superior to the state-of-the-art re-ranking methods. Furthermore, in the transfer learning setting, the model trained on either Market1501 or CUHK03 can achieve a comparable accuracy improvement on the DuekMTMC dataset, which validates the generalization of our SE model. Tongkun Xu, Jiamin Hou, Jiyong Zhang 0001, Xinhong Hao, Jian Yin 0003 |
ICME | 4 |
| 2020 | Deep Space Probing for Point Cloud Analysisabstract3D points distribute in a continuous 3D space irregularly, thus directly adapting 2D image convolution to 3D points is not an easy job. Previous works often artificially divide the space into regular grids, yet it could be suboptimal to learn geometry. In this paper, we propose SPCNN, namely, Space Probing Convolutional Neural Network, which naturally generalizes image CNN to deal with point clouds. The key idea of SPCNN is learning to probe the 3D space in an adaptive manner. Specifically, we define a pool of learnable convolutional weights, and let each point in the local region learn to choose a suitable convolutional weight from the pool. This is achieved by constructing a geometry guided index-mapping function that implicitly establishes a correspondence between convolutional weights and some local regions in the neighborhood (Fig. 1). In this way, the index-mapping function learns to adaptively partition nearby space for local geometry pattern recognition. With this convolution as a basic operator, SPCNN, a hierarchical architecture can be developed for effective point cloud analysis. Extensive experiments on challenging benchmarks across three tasks demonstrate that SPCNN achieves the state-of-the-art or has competitive performance. Yirong Yang, Bin Fan 0001, Yongcheng Liu, Jiyong Zhang 0001, Xin Liu 0027, Xinyu Cai, Shiming Xiang, Chunhong Pan |
ICPR | 5 |
| 2020 | Real-World Automatic Makeup via Identity Preservation Makeup NetabstractThis paper focuses on the real-world automatic makeup problem. Given one non-makeup target image and one reference image, the automatic makeup is to generate one face image, which maintains the original identity with the makeup style in the reference image. In the real-world scenario, face makeup task demands a robust system against the environmental variants. The two main challenges in real-world face makeup could be summarized as follow: first, the background in real-world images is complicated. The previous methods are prone to change the style of background as well; second, the foreground faces are also easy to be affected. For instance, the ``heavy'' makeup may lose the discriminative information of the original identity. To address these two challenges, we introduce a new makeup model, called Identity Preservation Makeup Net (IPM-Net), which preserves not only the background but the critical patterns of the original identity. Specifically, we disentangle the face images to two different information codes, i.e., identity content code and makeup style code. When inference, we only need to change the makeup style code to generate various makeup images of the target person. In the experiment, we show the proposed method achieves not only better accuracy in both realism (FID) and diversity (LPIPS) in the test set, but also works well on the real-world images collected from the Internet. Zhikun Huang, Zhedong Zheng, Chenggang Yan 0001, Hongtao Xie 0001, Yaoqi Sun, Jiyong Zhang 0001 |
IJCAI | 7 |
| 2020 | Video object detection for autonomous driving: Motion-aid feature calibration
Dongfang Liu, Yiming Cui 0002, Victor Y. Chen, Jiyong Zhang 0001, Bin Fan 0001 |
Neurocomputing | 4 |
| 2020 | Cross-modal feature extraction and integration based RGBD saliency detection
Liang Pan, Xiaofei Zhou 0003, Jiyong Zhang 0001, Chenggang Yan 0001 |
Image Vis. Comput. | 4 |
| 2020 | Attention-guided RGBD saliency detection using appearance information
Xiaofei Zhou 0003, Gongyang Li, Chen Gong 0002, Zhi Liu 0003, Jiyong Zhang 0001 |
Image Vis. Comput. | 5 |
| 2020 | Towards context-aware collaborative filtering by learning context-aware latent representations
Xin Liu 0027, Jiyong Zhang 0001, Chenggang Yan 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Improved covariant local feature detector
Zhanqiang Huo, Hongmin Liu 0001, Jing Wang 0093, Xin Liu 0027, Jiyong Zhang 0001 |
Pattern Recognit. Lett. | 6 |
| 2020 | A Key Business Node Identification Model for Internet of Things SecurityabstractBased on the research of business continuity and information security of the Internet of Things (IoT), a key business node identification model for the Internet of Things security is proposed. First, the business nodes are obtained based on the business process, and the importance decision matrix of business nodes is constructed by quantifying the evaluation attributes of nodes. Second, the attribute weights are improved by the analytic hierarchy process (AHP) and entropy weighting method from subjective and objective dimensions to form the combination weight decision matrix, and the analytic hierarchy process and entropy weighting VIKOR (AE-VIKOR) method are used to calculate the business node importance coefficient to identify the key nodes. Finally, according to the NSL-KDD dataset, the network security events of IoT network intrusion detection based on machine learning are monitored purposefully, and after the information security event occurs in the smart mobile phone, which impacts through IoT on the business system, the impact of the key business node on business continuity is analyzed, and the business continuity risk value is calculated to evaluate the business risk to prove the effectiveness of the model. The experimental results of the civil aviation departure business show that the AE-VIKOR method can effectively identify key business node, and the impact of the key business node on business continuity is analyzed, which further proves the efficiency and accuracy of the model in identifying the key business node. Lixia Xie, Huiyu Ni, Hongyu Yang 0003, Jiyong Zhang 0001 |
Secur. Commun. Networks | 4 |
| 2020 | An Unsupervised Learning-Based Network Threat Situation Assessment Model for Internet of ThingsabstractWith the wide application of network technology, the Internet of Things (IoT) systems are facing the increasingly serious situation of network threats; the network threat situation assessment becomes an important approach to solve these problems. Aiming at the traditional methods based on data category tag that has high modeling cost and low efficiency in the network threat situation assessment, this paper proposes a network threat situation assessment model based on unsupervised learning for IoT. Firstly, we combine the encoder of variational autoencoder (VAE) and the discriminator of generative adversarial networks (GAN) to form the V-G network. Then, we obtain the reconstruction error of each layer network by training the network collection layer of the V-G network with normal network traffic. Besides, we conduct the reconstruction error learning by the 3-layer variational autoencoder of the output layer and calculate the abnormal threshold of the training. Moreover, we carry out the group threat testing with the test dataset containing abnormal network traffic and calculate the threat probability of each test group. Finally, we obtain the threat situation value (TSV) according to the threat probability and the threat impact. The simulation results show that, compared with the other methods, this proposed method can evaluate the overall situation of network security threat more intuitively and has a stronger characterization ability for network threats. Hongyu Yang 0003, Renyun Zeng, Fengyan Wang, Guangquan Xu, Jiyong Zhang 0001 |
Secur. Commun. Networks | 5 |
| 2020 | Enabling 5G: sentimental image dominant graph topic model for cross-modality topic detection
Liang Li 0003, Wenchao Li 0004, Jiyong Zhang 0001, Chenggang Yan 0001 |
Wirel. Networks | 4 |
| 2019 | Truncated Gradient Confidence-Weighted Based Online Learning for Imbalance Streaming DataabstractOnline learning for imbalanced streaming data is an important and challenging problem for many classification tasks in the machine learning research field. Traditional online learning algorithms are mainly focused on classification tasks with balanced data, and with little consideration about the characteristics of imbalanced streaming data. In this paper, we propose a novel online learning algorithm called Truncated Gradient Confidence-Weighted (TGCW), which integrate the truncated gradient algorithm with the confidence weighted algorithm together to improve the feature selection ability while reducing the dimensions of imbalanced streaming data effectively. We study a number of classification tasks with various imbalance data ratio including the pedestrian detection application and compare the performance of the TGCW algorithm with traditional online learning algorithms, and empirical results show that the TGCW algorithm can achieve better performance consistently than other baseline approaches. Ji Hu 0002, Chenggang Yan 0001, Xin Liu 0027, Jiyong Zhang 0001, Dongliang Peng 0001, Yi Yang 0001 |
ICME | 4 |
| 2019 | Landmark Selection for Zero-shot LearningabstractZero-shot learning (ZSL) is an emerging research topic whose goal is to build recognition models for previously unseen classes. The basic idea of ZSL is based on heterogeneous feature matching which learns a compatibility function between image and class features using seen classes. The function is constructed based on one-vs-all training in which each class has only one class feature and many image features. Existing ZSL works mostly treat all image features equivalently. However, in this paper we argue that it is more reasonable to use some representative cross-domain data instead of all. Motivated by this idea, we propose a novel approach, termed as Landmark Selection(LAST) for ZSL. LAST is able to identify representative cross-domain features which further lead to better image-class compatibility function. Experiments on several ZSL datasets including ImageNet demonstrate the superiority of LAST to the state-of-the-arts. Guiguang Ding, Jungong Han, Chenggang Yan 0001, Jiyong Zhang 0001, Qionghai Dai |
IJCAI | 5 |
| 2019 | Meta-Learning for Low-resource Natural Language Generation in Task-oriented Dialogue SystemsabstractNatural language generation (NLG) is an essential component of task-oriented dialogue systems. Despite the recent success of neural approaches for NLG, they are typically developed for particular domains with rich annotated training examples. In this paper, we study NLG in a low-resource setting to generate sentences in new scenarios with handful training examples. We formulate the problem from a meta-learning perspective, and propose a generalized optimization-based approach (Meta-NLG) based on the well-recognized model-agnostic meta-learning (MAML) algorithm. Meta-NLG defines a set of meta tasks, and directly incorporates the objective of adapting to new low-resource NLG tasks into the meta-learning optimization process. Extensive experiments are conducted on a large multi-domain dataset (MultiWoz) with diverse linguistic variations. We show that Meta-NLG significantly outperforms other training procedures in various low-resource configurations. We analyze the results, and demonstrate that Meta-NLG adapts extremely fast and well to low-resource situations. Fei Mi, Minlie Huang, Jiyong Zhang 0001, Boi Faltings |
IJCAI | 3 |
| 2019 | Deep fusion based video saliency detection
Hongfa Wen, Xiaofei Zhou 0003, Yaoqi Sun, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Decomposing Activities of Daily Living to Discover Routine ClustersabstractThe modern sensor technology helps us collect time series data for activities of daily living (ADLs), which in turn can be used to infer broad patterns, such as common daily routines. Most of the existing approaches either rely on a model trained by a preselected and manually labeled set of activities, or perform micro-pattern analysis with manually selected length and number of micro-patterns. Since real life ADL datasets are massive, such approaches would be too costly to apply. Thus, there is a need to formulate unsupervised methods that can be applied to different time scales.We propose a novel approach to discover clusters of daily activity routines.We use a matrix decomposition method to isolate routines and deviations to obtain two different sets of clusters. We obtain the final memberships via the cross product of these sets. We validate our approach using two real-life ADL datasets and a well-known artificial dataset. Based on average silhouette width scores, our approach can capture strong structures in the underlying data. Furthermore, results show that our approach improves on the accuracy of the baseline algorithms by 12% with a statistical significance (p < 0.05) using the Wilcoxon signed-rank comparison test. Onur Yürüten, Jiyong Zhang 0001, Pearl Pu |
AAAI | 2 |
| 2014 | Predictors of life satisfaction based on daily activities from mobile sensor dataabstractIn recent years much research work has been dedicated to detecting user activity patterns from sensor data such as location, movement and proximity. However, how daily activities are correlated to people's happiness (such as their satisfaction from work and social lives) is not well explored. In this work, we propose an approach to investigate the relationship between users' daily activity patterns and their life satisfaction level. From a well-known longitudinal dataset collected by mobile devices, we extract various activity features through location and proximity information, and compute the entropies of these data to capture the regularities of the behavioral patterns of the participants. We then perform component analysis and structural equation modeling to identify key behavior contributors to self-reported satisfaction scores. Our results show that our analytical procedure can identify meaningful assumptions of causality between activities and satisfaction. Particularly, keeping regularity in daily activities can significantly improve the life satisfaction. Onur Yürüten, Jiyong Zhang 0001, Pearl Pu |
CHI | 2 |
| 2008 | A visual interface for critiquing-based recommender systemsabstractCritiquing-based recommender systems provide an efficient way for users to navigate through complex product spaces even if they are not familiar with the domain details in e-commerce environments. While recent researchers have mainly concentrated on methods for generating high quality compound critiques, to date there has been a lack of comprehensive investigation on the interface design issues. Traditionally the interface is textual, which shows compound critiques in plain text and may not be easily understood. In this paper we propose a new visual interface which represents various critiques by a set of meaningful icons. Results from our real-user evaluation show that the visual interface can improve the performance of critique-based recommenders by attracting users to apply the compound critiques more frequently and reducing users' interaction effort substantially when the product domain is complex. Users' subjective feedback also shows that the visual interface is highly promising in enhancing users' shopping experience. Jiyong Zhang 0001, Nicolas Jones, Pearl Pu |
EC | 1 |
| 2007 | A comparison of two compound critiquing systemsabstractCompound critiques allow users to simultaneously express directional preferences over several product attributes. Presenting the user with compound critiques is not a new idea. The original Find-Me Systems (e.g., Car Navigator) showed static compound critiques; they didn't change irrespective of user preferences or the product availability. Recently, a number of techniques for dynamically generating compound critiques have been proposed. While these techniques have been evaluated in isolation, to date no direct comparison of these (in terms of their interfacing characteristics and recommendation performance) has been reported. Motivated by this, our research groups have come together to carry out this comparison for the approaches we each take. The user study platform that we have developed facilitates the comparison of various critiquing based recommenders. In this paper we report the first set of results from a comprehensive real-user evaluation of two dynamic compound critique systems using this evaluation platform. James Reilly 0001, Jiyong Zhang 0001, Lorraine McGinty, Pearl Pu, Barry Smyth |
IUI | 2 |
| 2007 | Refining preference-based search results through Bayesian filteringabstractPreference-based search (PBS) is a popular approach for helping consumers find their desired items from online catalogs. Currently most PBS tools generate search results by a certain set of criteria based on preferences elicited from the current user during the interaction session. Due to the incompleteness and uncertainty of the user's preferences, the search results are often inaccurate and may contain items that the user has no desire to select. In this paper we develop an efficient Bayesian filter based on a group of users' past choice behavior and use it to refine the search results by filtering out items which are unlikely to be selected by the user. Our preliminary experiment shows that our approach is highly promising in generating more accurate search results and saving user's interaction effort. Jiyong Zhang 0001, Pearl Pu |
IUI | 1 |
| 2007 | A recursive prediction algorithm for collaborative filtering recommender systemsabstractCollaborative filtering (CF) is a successful approach for building online recommender systems. The fundamental process of the CF approach is to predict how a user would like to rate a given item based on the ratings of some nearest-neighbor users (user-based CF) or nearest-neighbor items (item-based CF). In the user-based CF approach, for example, the conventional prediction procedure is to find some nearest-neighbor users of the active user who have rated the given item, and then aggregate their rating information to predict the rating for the given item. In reality, due to the data sparseness, we have observed that a large proportion of users are filtered out because they don't rate the given item, even though they are very close to the active user. In this paper we present a recursive prediction algorithm, which allows those nearest-neighbor users to join the prediction process even if they have not rated the given item. In our approach, if a required rating value is not provided explicitly by the user, we predict it recursively and then integrate it into the prediction process. We study various strategies of selecting nearest-neighbor users for this recursive process. Our experiments show that the recursive prediction algorithm is a promising technique for improving the prediction accuracy for collaborative filtering recommender systems. Jiyong Zhang 0001, Pearl Pu |
RecSys | 1 |
| 2007 | Evaluating compound critiquing recommenders: a real-user studyabstractConversational recommender systems are designed to help users to more efficiently navigate complex product spaces by alternatively making recommendations and inviting users' feedback. Compound critiquing techniques provide an efficient way for users to feed back their preferences (in terms of several simultaneous product attributes) when interfacing with conversational recommender systems. For example, in the laptop domain a user might wish to express a preference for a laptop that is "Cheaper, Lighter, with a Larger Screen". While recently a number of techniques for dynamically generating compound critiques have been proposed, to date there has been a lack of direct comparison of these approaches in a real-user study. In this paper we will compare two alternative approaches to the dynamic generation of compound critiques based on ideas from data mining and multi-attribute utility theory. We will demonstrate how both approaches support users to more efficiently navigate complex product spaces highlighting, in particular, the influence of product complexity and interface strategy on recommendation performance and user satisfaction. James Reilly 0001, Jiyong Zhang 0001, Lorraine McGinty, Pearl Pu, Barry Smyth |
EC | 2 |