VLDB 2026 Research / reviewers in the wild / expert
Zhiyong Li 0001
dblp:59/5464-1
· DBLP profile ↗
83ranked-venue papers
14as first author
47since 2021 · last 2026
0000-0001-9720-5915ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 7 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-author · 17 since 2021Systems, architecture and hardware · 9 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Computer networks · 3 · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World ModelsabstractBirds' Eye View (BEV) semantic segmentation is an indispensable perception task in end-to-end autonomous driving systems. Unsupervised and semi-supervised learning for BEV tasks, as pivotal for real-world applications, underperform due to the homogeneous distribution of the labeled data. In this work, we explore the potential of synthetic data from driving world models to enhance the diversity of labeled data for robustifying BEV segmentation. Yet, our preliminary findings reveal that generation noise in synthetic data compromises efficient BEV model learning. To fully harness the potential of synthetic data from world models, this article proposes NRSeg, a noise-resilient learning framework for BEV semantic segmentation. Specifically, a Perspective-Geometry Consistency Metric (PGCM) is proposed to quantitatively evaluate the guidance capability of generated data for model learning. This metric originates from the alignment measure between the perspective road mask of generated data and the mask projected from the BEV labels. Moreover, a Bi-Distribution Parallel Prediction (BiDPP) is designed to enhance the inherent robustness of the model, where the learning process is constrained through parallel prediction of multinomial and Dirichlet distributions. The former efficiently predicts semantic probabilities, whereas the latter adopts evidential deep learning to realize uncertainty quantification. Furthermore, a Hierarchical Local Semantic Exclusion (HLSE) module is designed to address the non-mutual exclusivity inherent in BEV semantic segmentation tasks. The proposed framework is evaluated on BEV semantic segmentation using data generated by multiple world models, with comprehensive testing conducted on the public nuScenes dataset under unsupervised and semi-supervised settings. Experimental results demonstrate that NRSeg achieves state-of-the-art performance, yielding the highest improvements in mIoU of 13.8% and 11.4% in unsupervised and semi-supervised BEV segmentation tasks, respectively. The source code will be made publicly available at https://github.com/lynn-yu/NRSeg. Siyu Li 0002, Yihong Cao, Kailun Yang 0001, Zhiyong Li 0001, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Scene Graph-Guided SegCaptioning Transformer With Fine-Grained Alignment for Controllable Video Segmentation and CaptioningabstractRecent advancements in multimodal large models have significantly bridged the representation gap between diverse modalities, catalyzing the evolution of video multimodal interpretation, which enhances users' understanding of video content by generating correlated modalities. However, most existing video multimodal interpretation methods primarily concentrate on global comprehension with limited user interaction. To address this, we propose a novel task, Controllable Video Segmentation and Captioning (SegCaptioning), which empowers users to provide specific prompts, such as a bounding box around an object of interest, to simultaneously generate correlated masks and captions that precisely embody user intent. An innovative framework, Scene Graph-guided Fine-grained SegCaptioning Transformer (SG-FSCFormer), is designed to integrate a Prompt-guided Temporal Graph Former to effectively capture and represent user intent through an adaptive prompt adaptor, ensuring that the generated content aligns well with the user's requirements. Furthermore, our model introduces a Fine-grained Mask-linguistic Decoder to collaboratively predict high-quality caption-mask pairs using a Multi-entity Contrastive loss, while providing fine-grained alignment between each mask and its corresponding caption tokens, thereby enhancing the user's comprehension of videos. Comprehensive experiments conducted on two benchmark datasets demonstrate that SG-FSCFormer achieves remarkable performance, effectively capturing user intent and generating precise multimodal outputs tailored to user specifications. Our code is available at https://github.com/XuZhang1211/SG-FSCFormer. Xu Zhang 0025, Jin Yuan 0002, BinHong Yang, Xuan Liu 0001, Qianjun Zhang, Yuyi Wang 0001, Zhiyong Li 0001, Hanwang Zhang |
IEEE Trans. Image Process. | 7 |
| 2025 | SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioningabstractControllable image semantic understanding tasks, such as captioning or segmentation, necessitate users to input a prompt (e.g., text or bounding boxes) to predict a unique outcome, presenting challenges such as high-cost prompt input or limited information output. This paper introduces a new task ``Image Collaborative Segmentation and Captioning'' (SegCaptioning), which aims to translate a straightforward prompt, like a bounding box around an object, into diverse semantic interpretations represented by (caption, masks) pairs, allowing flexible result selection by users. This task poses significant challenges, including accurately capturing a user's intention from a minimal prompt while simultaneously predicting multiple semantically aligned caption words and masks. Technically, we propose a novel Scene Graph Guided Diffusion Model that leverages structured scene graph features for correlated mask-caption prediction. Initially, we introduce a Prompt-Centric Scene Graph Adaptor to map a user's prompt to a scene graph, effectively capturing his intention. Subsequently, we employ a diffusion process incorporating a Scene Graph Guided Bimodal Transformer to predict correlated caption-mask pairs by uncovering intricate correlations between them. To ensure accurate alignment, we design a Multi-Entities Contrastive Learning loss to explicitly align visual and textual entities by considering inter-modal similarity, resulting in well-aligned caption-mask pairs. Extensive experiments conducted on two datasets demonstrate that SGDiff achieves superior performance in SegCaptioning, yielding promising results for both captioning and segmentation tasks with minimal prompt input. Xu Zhang 0025, Jin Yuan 0002, Hanwang Zhang, Guojin Zhong, Yongsheng Zang, Jiacheng Lin, Zhiyong Li 0001 |
AAAI | 7 |
| 2025 | AVAM: A Universal Training-Free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-Image Question Answering
Kang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan 0002, Zhiyong Li 0001 |
ICCV | 5 |
| 2025 | Multi-Resolution Decomposable Diffusion Model for Non-Stationary Time Series Anomaly DetectionabstractRecently, generative models have shown considerable promise in unsupervised time series anomaly detection. Nonetheless, the task of effectively capturing complex temporal patterns and minimizing false alarms becomes increasingly challenging when dealing with non-stationary time series, characterized by continuously fluctuating statistical attributes and joint distributions. To confront these challenges, we underscore the benefits of multi-resolution modeling, which improves the ability to distinguish between anomalies and non-stationary behaviors by leveraging correlations across various resolution scales. In response, we introduce a **M**ulti-Res**o**lution **De**composable Diffusion **M**odel (MODEM), which integrates a coarse-to-fine diffusion paradigm with a frequency-enhanced decomposable network to adeptly navigate the intricacies of non-stationarity. Technically, the coarse-to-fine diffusion model embeds cross-resolution correlations into the forward process to optimize diffusion transitions mathematically. It then innovatively employs low-resolution recovery to guide the reverse trajectories of high-resolution series in a coarse-to-fine manner, enhancing the model's ability to learn and elucidate underlying temporal patterns. Furthermore, the frequency-enhanced decomposable network operates in the frequency domain to extract globally shared time-invariant information and time-variant temporal dynamics for accurate series reconstruction. Extensive experiments conducted across five real-world datasets demonstrate that our proposed MODEM achieves state-of-the-art performance and can be generalized to other time series tasks. Guojin Zhong, Pan Wang 0011, Jin Yuan 0002, Zhiyong Li 0001, Long Chen 0016 |
ICLR | 4 |
| 2025 | Resource-Efficient Affordance Grounding with Complementary Depth and Semantic PromptsabstractAffordance refers to the functional properties that an agent perceives and utilizes from its environment, and is key perceptual information required for robots to perform actions. This information is rich and multimodal in nature. Existing multimodal affordance methods face limitations in extracting useful information, mainly due to simple structural designs, basic fusion methods, and large model parameters, making it difficult to meet the performance requirements for practical deployment. To address these issues, this paper proposes the BiT-Align image-depth-text affordance mapping framework. The framework includes a Bypass Prompt Module (BPM) and a Text Feature Guidance (TFG) attention selection mechanism. BPM integrates the auxiliary modality depth image directly as a prompt to the primary modality RGB image, embedding it into the primary modality encoder without introducing additional encoders. This reduces the model’s parameter count and effectively improves functional region localization accuracy. The TFG mechanism guides the selection and enhancement of attention heads in the image encoder using textual features, improving the understanding of affordance characteristics. Experimental results demonstrate that the proposed method achieves significant performance improvements on public AGD20K and HICO-IIF datasets. On the AGD20K dataset, compared with the current state-of-the-art method, we achieve a 6.0% improvement in the KLD metric, while reducing model parameters by 88.8%, demonstrating practical application values. The source code will be made publicly available at https://github.com/DAWDSE/BiT-Align. Fan Yang 0063, Guoliang Zhu, Hao Shi 0004, Yukun Zuo, Wenrui Chen, Zhiyong Li 0001, Kailun Yang 0001 |
IROS | 8 |
| 2025 | One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing ScenesabstractDeformable object manipulation in robotics presents significant challenges due to uncertainties in component properties, diverse configurations, visual interference, and ambiguous prompts. These factors complicate both perception and control tasks. To address these challenges, we propose a novel method for One-Shot Affordance Grounding of Deformable Objects (OS-AGDO) in egocentric organizing scenes, enabling robots to recognize previously unseen deformable objects with varying colors and shapes using minimal samples. Specifically, we first introduce the Deformable Object Semantic Enhancement Module (DefoSEM), which enhances hierarchical understanding of the internal structure and improves the ability to accurately identify local features, even under conditions of weak component information. Next, we propose the ORB-Enhanced Keypoint Fusion Module (OEKFM), which optimizes feature extraction of key components by leveraging geometric constraints and improves adaptability to diversity and visual interference. Additionally, we propose an instance-conditional prompt based on image data and task context, which effectively mitigates the issue of region ambiguity caused by prompt words. To validate these methods, we construct a diverse real-world dataset, AGDDO15, which includes 15 common types of deformable objects and their associated organizational actions. Experimental results demonstrate that our approach significantly outperforms state-of-the-art methods, achieving improvements of 6.2%, 3.2%, and 2.9% in KLD, SIM, and NSS metrics, respectively, while exhibiting high generalization performance. Source code and benchmark dataset are made publicly available at https://github.com/Dikay1/OS-AGDO. Wanjun Jia, Fan Yang 0063, Mengfei Duan, Xianchi Chen, Yinxi Wang, Yiming Jiang 0001, Wenrui Chen, Kailun Yang 0001, Zhiyong Li 0001 |
IROS | 9 |
| 2025 | Unveiling the Potential of Segment Anything Model 2 for RGB-Thermal Semantic Segmentation with Language GuidanceabstractThe perception capability of robotic systems relies on the richness of the dataset. Although Segment Anything Model 2 (SAM2), trained on large datasets, demonstrates strong perception potential in perception tasks, its inherent training paradigm prevents it from being suitable for RGB-T tasks. To address these challenges, we propose SHIFNet, a novel SAM2-driven Hybrid Interaction Paradigm that unlocks the potential of SAM2 with linguistic guidance for efficient RGB-Thermal perception. Our framework consists of two key components: (1) Semantic-Aware Cross-modal Fusion (SACF) module that dynamically balances modality contributions through text-guided affinity learning, overcoming SAM2’s inherent RGB bias; (2) Heterogeneous Prompting Decoder (HPD) that enhances global semantic information through a semantic enhancement module and then combined with category embeddings to amplify cross-modal semantic consistency. With 32.27M trainable parameters, SHIFNet achieves state-of-the-art segmentation performance on public benchmarks, reaching 89.8% on PST900 and 67.8% on FMB, respectively. The framework facilitates the adaptation of pre-trained large models to RGB-T segmentation tasks, effectively mitigating the high costs associated with data collection while endowing robotic systems with comprehensive perception capabilities. The source code will be made publicly available at https://github.com/iAsakiT3T/SHIFNet. Zhiyong Li 0001, Xu Zheng 0002, Kailun Yang 0001 |
IROS | 5 |
| 2025 | Expression Prompt Collaboration Transformer for universal referring video object segmentation
Jiacheng Lin, Guojin Zhong, Haolong Fu, Ke Nai, Kailun Yang 0001, Zhiyong Li 0001 |
Knowl. Based Syst. | 7 |
| 2025 | Task-Oriented Tool Manipulation With Robotic Dexterous Hands: A Knowledge Graph Approach From Fingers to FunctionalityabstractA primary challenge in robotic tool use is achieving precise manipulation with dexterous robotic hands to mimic human actions. It requires understanding human tool use and allocating specific functions to each robotic finger for fine control. Existing work has primarily focused on the overall grasping capabilities of robotic hands, often neglecting the functional allocation among individual fingers during object interaction. In response to this, we introduce a semantic knowledge-driven approach to distribute functions among fingers for tool manipulation. Central to this approach is the finger-to-function (F2F) knowledge graph, which captures human expertise in tool use and establishes relationships between tool attributes, tasks, and manipulation elements, including functional fingers, components, required force, and gestures. We also develop a manipulation element-oriented prediction algorithm using knowledge graph semantic embedding, enhancing the prediction of manipulation elements' speed and accuracy. Additionally, we propose the functionality-integrated adaptive force feedback manipulation (FAFM) module, which integrates manipulation elements with adaptive force feedback to achieve precise finger-level control. Our framework does not rely on extensive annotated data for supervision but utilizes semantic constraints from F2F to guide tool manipulation. The proposed method demonstrates superior performance and generalizability in real-world scenarios, achieving an 8% higher success rate in grasping and manipulation of representative tool instances compared to the existing state-of-the-art methods. The dataset and code are available at https://github.com/yangfan293/F2F. Fan Yang 0063, Wenrui Chen, Sijie Wu, Xin Li 0082, Zhiyong Li 0001, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 6 |
| 2025 | AdaptiveClick: Click-Aware Transformer With Adaptive Focal Loss for Interactive Image SegmentationabstractInteractive image segmentation (IIS) has emerged as a promising technique for decreasing annotation time. Substantial progress has been made in pre- and post-processing for IIS, but the critical issue of interaction ambiguity, notably hindering segmentation quality, has been under-researched. To address this, we introduce ADAPTIVE CLICK - a click-aware transformer incorporating an adaptive focal loss (AFL) that tackles annotation inconsistencies with tools for mask- and pixel-level ambiguity resolution. To the best of our knowledge, AdaptiveClick is the first transformer-based, mask-adaptive segmentation framework for IIS. The key ingredient of our method is the click-aware mask-adaptive transformer decoder (CAMD), which enhances the interaction between click and image features. Additionally, AdaptiveClick enables pixel-adaptive differentiation of hard and easy samples in the decision space, independent of their varying distributions. This is primarily achieved by optimizing a generalized AFL with a theoretical guarantee, where two adaptive coefficients control the ratio of gradient values for hard and easy pixels. Our analysis reveals that the commonly used Focal and BCE losses can be considered special cases of the proposed AFL. With a plain ViT backbone, extensive experimental results on nine datasets demonstrate the superiority of AdaptiveClick compared to state-of-the-art methods. The source code is publicly available at https://github.com/lab206/AdaptiveClick. Jiacheng Lin, Kailun Yang 0001, Alina Roitberg, Siyu Li 0002, Zhiyong Li 0001, Shutao Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Learning Granularity-Aware Affordances From Human-Object Interaction for Tool-Based Functional Dexterous GraspingabstractTo enable robots to use tools, the initial step is teaching robots to employ dexterous gestures for touching specific areas precisely where tasks are performed. Affordance features of objects serve as a bridge in the functional interaction between agents and objects. However, leveraging these affordance cues to help robots achieve functional tool grasping remains unresolved. To address this, we propose a granularity-aware affordance feature extraction method for locating functional affordance areas and predicting dexterous coarse gestures. We study the intrinsic mechanisms of human tool use. On the one hand, we use fine-grained affordance features of object-functional finger contact areas to locate functional affordance regions. On the other hand, we use highly activated coarse-grained affordance features in hand-object interaction regions to predict grasp gestures. Additionally, we introduce a model-based postprocessing module that transforms affordance localization and gesture prediction into executable robotic actions. This forms GAAF-Dex, a complete framework that learns granularity-aware affordances from human-object interaction to enable tool-based functional grasping with dexterous hands. Unlike fully supervised methods that require extensive data annotation, we employ a weakly supervised approach to extract relevant cues from exocentric (Exo) images of hand-object interactions to supervise feature extraction in egocentric (Ego) images. To support this approach, we have constructed a small-scale dataset, functional affordance hand (FAH)-object interaction dataset, which includes nearly 6k images of functional hand-object interaction Exo images and Ego images of 18 commonly used tools performing six tasks. Extensive experiments on the dataset demonstrate that our method outperforms state-of-the-art methods, and real-world localization and grasping experiments validate the practical applicability of our approach. The source code and the established dataset are available at https://github.com/yangfan293/GAAF-DEX. Fan Yang 0063, Wenrui Chen, Kailun Yang 0001, Conghui Tang, Zhiyong Li 0001, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | LF Tracy: A Unified Single-Pipeline Paradigm for Salient Object Detection in Light Field Cameras
Jiaming Zhang 0001, Kunyu Peng, Xina Cheng, Zhiyong Li 0001, Kailun Yang 0001 |
ICPR (17) | 6 |
| 2024 | CF-Deformable DETR: An End-to-End Alignment-Free Model for Weakly Aligned Visible-Infrared Object Detection
Haolong Fu, Jin Yuan 0002, Guojin Zhong, Jiacheng Lin, Zhiyong Li 0001 |
IJCAI | 6 |
| 2024 | MambaMOS: LiDAR-based 3D Moving Object Segmentation with Motion-aware State Space ModelabstractLiDAR-based Moving Object Segmentation (MOS) aims to locate and segment moving objects in point clouds of the current scan using motion information from previous scans. Despite the promising results achieved by previous MOS methods, several key issues, such as the weak coupling of temporal and spatial information, still need further study. In this paper, we propose a novel LiDAR-based 3D Moving Object Segmentation with Motion-aware State Space Model, termed MambaMOS. Firstly, we develop a novel embedding module, the Time Clue Bootstrapping Embedding (TCBE), to enhance the coupling of temporal and spatial information in point clouds and alleviate the issue of overlooked temporal clues. Secondly, we introduce the Motion-aware State Space Model (MSSM) to endow the model with the capacity to understand the temporal correlations of the same object across different time steps. Specifically, MSSM emphasizes the motion states of the same object at different time steps through two distinct temporal modeling and correlation steps. We utilize an improved state space model to represent these motion differences, significantly modeling the motion states. Finally, extensive experiments on the SemanticKITTI-MOS and KITTI-Road benchmarks demonstrate that the proposed MambaMOS achieves state-of-the-art performance. The source code is publicly available at https://github.com/Terminal-K/MambaMOS Kang Zeng, Hao Shi 0004, Jiacheng Lin, Siyu Li 0002, Jintao Cheng, Kaiwei Wang, Zhiyong Li 0001, Kailun Yang 0001 |
ACM Multimedia | 7 |
| 2024 | Privacy preservation network with global-aware focal loss for Interactive Personal Visual Privacy Preservation
Jiacheng Lin, Haolong Fu, Yifan Li 0005, Jin Yuan 0002, Zhiyong Li 0001 |
Neurocomputing | 7 |
| 2024 | Attribute discrimination combined with selected sample dropout for unsupervised domain adaptive person re-identification
Guanzhong Yang, Zhiyong Li 0001 |
Image Vis. Comput. | 5 |
| 2024 | Two3-AnoECG: ECG anomaly detection with two-stream networks and two-stage training using two double-throw switchesabstractThe electrocardiogram (ECG) is a highly cost-effective and convenient diagnostic tool that can aid in the diagnosis of a wide range of cardiovascular conditions, such as arrhythmias , myocardial ischemia , myocardial infarction, and heart failure. Its usefulness lies in its ability to provide information about the electrical activity and rhythm of the heart, making it an essential component of the diagnostic process for many cardiac conditions. However, the interpretation of ECG signals often requires a high level of expertise from medical professionals. When reading complex ECGs, noncardiologists often need to consult with cardiologists. Even cardiologists may make errors in their judgments after prolonged reading of ECGs. Therefore, the accurate and timely diagnosis of cardiovascular diseases using ECG signals is a complex task. In this paper, we categorize this problem as a multilabel, multidimensional time-series data anomaly detection task. We propose a method called T w o 3 -AnoECG for ECG anomaly detection, which improves upon existing ECG feature extraction and data segmentation methods and introduces a two-stream network that is trained in two stages using two double-throw switches to control the input and model structure of each stage. We further propose T w o 3 -EnsECG, which combines the anomaly score generated by T w o 3 -AnoECG and multiple baseline methods, to improve the overall performance of anomaly detection in ECG signals. The experimental results demonstrate the effectiveness of our proposed method. Yifan Li 0005, Weixun Cai, Jiacheng Lin, Zhiyong Li 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Multiscale deep feature selection fusion network for referring image segmentation
Xianwen Dai, Jiacheng Lin, Ke Nai, Qingpeng Li, Zhiyong Li 0001 |
Multim. Tools Appl. | 5 |
| 2024 | Click-Pixel Cognition Fusion Network With Balanced Cut for Interactive Image SegmentationabstractInteractive image segmentation (IIS) has been widely used in various fields, such as medicine, industry, etc. However, some core issues, such as pixel imbalance, remain unresolved so far. Different from existing methods based on pre-processing or post-processing, we analyze the cause of pixel imbalance in depth from the two perspectives of pixel number and pixel difficulty. Based on this, a novel and unified Click-pixel Cognition Fusion network with Balanced Cut (CCF-BC) is proposed in this paper. On the one hand, the Click-pixel Cognition Fusion (CCF) module, inspired by the human cognition mechanism, is designed to increase the number of click-related pixels (namely, positive pixels) being correctly segmented, where the click and visual information are fully fused by using a progressive three-tier interaction strategy. On the other hand, a general loss, Balanced Normalized Focal Loss (BNFL), is proposed. Its core is to use a group of control coefficients related to sample gradients and forces the network to pay more attention to positive and hard-to-segment pixels during training. As a result, BNFL always tends to obtain a balanced cut of positive and negative samples in the decision space. Theoretical analysis shows that the commonly used Focal and BCE losses can be regarded as special cases of BNFL. Experiment results of five well-recognized datasets have shown the superiority of the proposed CCF-BC method compared to other state-of-the-art methods. The source code is publicly available at https://github.com/lab206/CCF-BC. Jiacheng Lin, Xiaohui Wei 0001, Puhong Duan, Renwei Dian, Zhiyong Li 0001, Shutao Li 0001 |
IEEE Trans. Image Process. | 7 |
| 2024 | PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image SegmentationabstractIntegration of diverse visual prompts like clicks, scribbles, and boxes in interactive image segmentation significantly facilitates users' interaction as well as improves interaction efficiency. However, existing studies primarily encode the position or pixel regions of prompts without considering the contextual areas around them, resulting in insufficient prompt feedback, which is not conducive to performance acceleration. To tackle this problem, this paper proposes a simple yet effective Probabilistic Visual Prompt Unified Transformer (PVPUFormer) for interactive image segmentation, which allows users to flexibly input diverse visual prompts with the probabilistic prompt encoding and feature post-processing to excavate sufficient and robust prompt features for performance boosting. Specifically, we first propose a Probabilistic Prompt-unified Encoder (PPuE) to generate a unified one-dimensional vector by exploring both prompt and non-prompt contextual information, offering richer feedback cues to accelerate performance improvement. On this basis, we further present a Prompt-to-Pixel Contrastive (P2C) loss to accurately align both prompt and pixel features, bridging the representation gap between them to offer consistent feature representations for mask prediction. Moreover, our approach designs a Dual-cross Merging Attention (DMA) module to implement bidirectional feature interaction between image and prompt features, generating notable features for performance improvement. A comprehensive variety of experiments on several challenging datasets demonstrates that the proposed components achieve consistent improvements, yielding state-of-the-art interactive segmentation performance. Our code is available at https://github.com/XuZhang1211/PVPUFormer. Xu Zhang 0025, Kailun Yang 0001, Jiacheng Lin, Jin Yuan 0002, Zhiyong Li 0001, Shutao Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | DTCLMapper: Dual Temporal Consistent Learning for Vectorized HD Map ConstructionabstractTemporal information plays a pivotal role in Bird’s-Eye-View (BEV) driving scene understanding, which can alleviate the visual information sparsity. However, the indiscriminate temporal fusion method will cause the barrier of feature redundancy when constructing vectorized High-Definition (HD) maps. In this paper, we revisit the temporal fusion of vectorized HD maps, focusing on temporal instance consistency and temporal map consistency learning. To improve the representation of instances in single-frame maps, we introduce a novel method, DTCLMapper. This approach uses a dual-stream temporal consistency learning module that combines instance embedding with geometry maps. In the instance embedding component, our approach integrates temporal Instance Consistency Learning (ICL), ensuring consistency from vector points and instance features aggregated from points. A vectorized points pre-selection module is employed to enhance the regression efficiency of vector points from each instance. Then aggregated instance features obtained from the vectorized points preselection module are grounded in contrastive learning to realize temporal consistency, where positive and negative samples are selected based on position and semantic information. The geometry mapping component introduces Map Consistency Learning (MCL) designed with self-supervised learning. The MCL enhances the generalization capability of our consistent learning approach by concentrating on the global location and distribution constraints of the instances. Extensive experiments on well-recognized benchmarks indicate that the proposed DTCLMapper achieves state-of-the-art performance in vectorized mapping tasks, reaching 61.9% and 65.1% mAP scores on the nuScenes and Argoverse datasets, respectively. The source code is available athttps://github.com/lynn-yu/DTCLMapper. Siyu Li 0002, Jiacheng Lin, Hao Shi 0004, Jiaming Zhang 0001, Song Wang 0019, You Yao, Zhiyong Li 0001, Kailun Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | EchoTrack: Auditory Referring Multi-Object Tracking for Autonomous DrivingabstractThis paper introduces the task of Auditory Referring Multi-Object Tracking (AR-MOT), which dynamically tracks specific objects in a video sequence based on audio expressions and appears as a challenging problem in autonomous driving. Due to the lack of semantic modeling capacity in audio and video, existing works have mainly focused on text-based multi-object tracking, which often comes at the cost of tracking quality, interaction efficiency, and even the safety of assistance systems, limiting the application of such methods in autonomous driving. In this paper, we delve into the problem of AR-MOT from the perspective of audio-video fusion and audio-video tracking. We put forward EchoTrack, an end-to-end AR-MOT framework with dual-stream vision transformers. The dual streams are intertwined with our Bidirectional Frequency-domain Cross-attention Fusion Module (Bi-FCFM), which bidirectionally fuses audio and video features from both frequency- and spatiotemporal domains. Moreover, we propose the Audio-visual Contrastive Tracking Learning (ACTL) regime to extract homogeneous semantic features between expressions and visual objects by learning homogeneous features between different audio and video objects effectively. Aside from the architectural design, we establish the first set of large-scale AR-MOT benchmarks, including Echo-KITTI, Echo-KITTI+, and Echo-BDD. Extensive experiments on the established benchmarks demonstrate the effectiveness of the proposed EchoTrack and its components. The source code and datasets are available athttps://github.com/lab206/EchoTrack. Jiacheng Lin, Kunyu Peng, Zhiyong Li 0001, Rainer Stiefelhagen, Kailun Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | SA2E-AD: A Stacked Attention Autoencoder for Anomaly Detection in Multivariate Time SeriesabstractAnomaly detection for multivariate time series is an essential task in the modern industrial field. Although several methods have been developed for anomaly detection, they usually fail to effectively exploit the metrical-temporal correlation and the other dependencies among multiple variables. To address this problem, we propose a stacked attention autoencoder for anomaly detection in multivariate time series (SA2E-AD); it focuses on fully utilizing the metrical and temporal relationships among multivariate time series. We design a multiattention block, alternately containing the temporal attention and metrical attention components in a hierarchical structure to better reconstruct normal time series, which is helpful in distinguishing the anomalies from the normal time series. Meanwhile, a two-stage training strategy is designed to further separate the anomalies from the normal data. Experiments on three publicly available datasets show that SA2E-AD outperforms the advanced baseline methods in detection performance and demonstrate the effectiveness of each part of the process in our method. Zhiyong Li 0001, Zhibang Yang, Xu Zhou 0001, Yifan Li 0005, Ziyan Wu 0006, Lingzhao Kong, Ke Nai |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | LRAF-Net: Long-Range Attention Fusion Network for Visible-Infrared Object DetectionabstractVisible-infrared object detection aims to improve the detector performance by fusing the complementarity of visible and infrared images. However, most existing methods only use local intramodality information to enhance the feature representation while ignoring the efficient latent interaction of long-range dependence between different modalities, which leads to unsatisfactory detection performance under complex scenes. To solve these problems, we propose a feature-enhanced long-range attention fusion network (LRAF-Net), which improves detection performance by fusing the long-range dependence of the enhanced visible and infrared features. First, a two-stream CSPDarknet53 network is used to extract the deep features from visible and infrared images, in which a novel data augmentation (DA) method is designed to reduce the bias toward a single modality through asymmetric complementary masks. Then, we propose a cross-feature enhancement (CFE) module to improve the intramodality feature representation by exploiting the discrepancy between visible and infrared images. Next, we propose a long-range dependence fusion (LDF) module to fuse the enhanced features by associating the positional encoding of multimodality features. Finally, the fused features are fed into a detection head to obtain the final detection results. Experiments on several public datasets, i.e., VEDAI, FLIR, and LLVIP, show that the proposed method obtains state-of-the-art performance compared with other methods. Haolong Fu, Shixun Wang, Puhong Duan, Changyan Xiao, Renwei Dian, Shutao Li 0001, Zhiyong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image GenerationabstractThe recently rising markup-to-image generation poses greater challenges as compared to natural image generation, due to its low tolerance for errors as well as the complex sequence and context correlations between markup and rendered image. This paper proposes a novel model named "Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment'' (FSA-CDM), which introduces contrastive positive/negative samples into the diffusion model to boost performance for markup-to-image generation. Technically, we design a fine-grained cross-modal alignment module to well explore the sequence similarity between the two modalities for learning robust feature representations. To improve the generalization ability, we propose a contrast-augmented diffusion model to explicitly explore positive and negative samples by maximizing a novel contrastive variational objective, which is mathematically inferred to provide a tighter bound for the model's optimization. Moreover, the context-aware cross attention module is developed to capture the contextual information within markup language during the denoising process, yielding better noise prediction results. Extensive experiments are conducted on four benchmark datasets from different domains, and the experimental results demonstrate the effectiveness of the proposed components in FSA-CDM, significantly exceeding state-of-the-art performance by about 2% ~ 12% DTW improvements. Guojin Zhong, Jin Yuan 0002, Pan Wang 0011, Kailun Yang 0001, Weili Guan, Zhiyong Li 0001 |
ACM Multimedia | 6 |
| 2023 | A Text-Specific Domain Adaptive Network for Scene Text Detection in the Wild
Jin Yuan 0002, Zhiyong Li 0001 |
Appl. Intell. | 6 |
| 2023 | Domain adaptive multigranularity proposal network for text detection under extreme traffic scenes
Zhiyong Li 0001, Jiacheng Lin, Ke Nai, Jin Yuan 0002, Yifan Li 0005 |
Comput. Vis. Image Underst. | 2 |
| 2023 | BRPPNet: Balanced privacy protection network for referring personal image privacy protection
Jiacheng Lin, Xianwen Dai, Ke Nai, Jin Yuan 0002, Zhiyong Li 0001, Xu Zhang 0025, Shutao Li 0001 |
Expert Syst. Appl. | 5 |
| 2023 | M3GAN: A masking strategy with a mutable filter for multidimensional anomaly detection
Yifan Li 0005, Ziyan Wu 0006, Fan Yang 0063, Zhiyong Li 0001 |
Knowl. Based Syst. | 6 |
| 2023 | DO-SA&R: Distant Object Augmented Set Abstraction and Regression for Point-Based 3D Object DetectionabstractPoint-based 3D detection approaches usually suffer from the severe point sampling imbalance problem between foreground and background. We observe that prior works have attempted to alleviate this imbalance by emphasizing foreground sampling. However, even adequate foreground sampling may be extremely unbalanced between nearby and distant objects, yielding unsatisfactory performance in detecting distant objects. To tackle this issue, this paper first proposes a novel method named Distant Object Augmented Set Abstraction and Regression (DO-SA&R) to enhance distant object detection, which is vital for the timely response of decision-making systems like autonomous driving. Technically, our approach first designs DO-SA with novel distant object augmented farthest point sampling (DO-FPS) to emphasize sampling on distant objects by leveraging both object-dependent and depth-dependent information. Then, we propose distant object augmented regression to reweight all the instance boxes for strengthening regression training on distant objects. In practice, the proposed DO-SA&R can be easily embedded into the existing modules, yielding consistent performance improvements, especially on detecting distant objects. Extensive experiments are conducted on the popular KITTI, nuScenes and Waymo datasets, and DO-SA&R demonstrates superior performance, especially for distant object detection. Our code is available at https://github.com/mikasa3lili/DO-SAR. Jiacheng Lin, Ke Nai, Jin Yuan 0002, Zhiyong Li 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | DCT-GAN: Dilated Convolutional Transformer-Based GAN for Time Series Anomaly DetectionabstractTime series anomaly detection (TSAD) is an essential problem faced in several fields, e.g., fault detection, fraud detection, and intrusion detection, etc. Although TSAD is a crucial problem in anomaly detection, few solutions in anomaly detection are suitable for it at present. Recently, some researchers use GAN-based methods such as TAnoGAN and TadGAN to solve TSAD problem. However, problems such as model collapse, low generalization capability and poor accuracy still exist. In this article, we proposed a Dilated Convolutional Transformer-based GAN (DCT-GAN) to enhance accuracy and improve generalization capability of the model. Specifically, DCT-GAN utilize several generators and a single discriminator to alleviate the mode collapse problem. Each generator consists of a dilated convolutional neural network and a Transformer block to obtain fine-grained and coarse-grained information of the time series, which is a useful component to improve generalization capability. We also use weight-based mechanism to balance these generators. Experiments verify the effectiveness of our method and each part of DCT-GAN. Yifan Li 0005, Jia Zhang 0005, Zhiyong Li 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Robust Visual Tracking via Multitask Sparse Correlation Filters LearningabstractIn this article, a novel multitask sparse correlation filters (MTSCF) model, which introduces multitask sparse learning into the CFs framework, is proposed for visual tracking. Specifically, the proposed MTSCF method exploits multitask learning to take the interdependencies among different visual features (e.g., histogram of oriented gradient (HOG), color names, and CNN features) into account to simultaneously learn the CFs and make the learned filters enhance and complement each other to boost the tracking performance. Moreover, it also performs feature selection to dynamically select discriminative spatial features from the target region to distinguish the target object from the background. A$l_{2,1}$regularization term is considered to realize multitask sparse learning. In order to solve the objective model, alternating direction method of multipliers is utilized for learning the CFs. By considering multitask sparse learning, the proposed MTSCF model can fully utilize the strength of different visual features and select effective spatial features to better model the appearance of the target object. Extensive experiment results on multiple tracking benchmarks demonstrate that our MTSCF tracker achieves competitive tracking performance in comparison with several state-of-the-art trackers. Ke Nai, Zhiyong Li 0001, Yihui Gan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | TP-FER: An Effective Three-phase Noise-tolerant Recognizer for Facial Expression RecognitionabstractSingle-label facial expression recognition (FER), which aims to classify single expression for facial images, usually suffers from the label noisy and incomplete problem, where manual annotations for partial training images exist wrong or incomplete labels, resulting in performance decline. Although prior work has attempted to leverage external sources or manual annotations to handle this problem, it usually requires extra costs. This article explores a simple yet effective three-phase paradigm (“warm-up,” “selection,” and “relabeling”) for FER task. First, the warm-up phase attempts to build an initial recognition network based on noisy samples for discriminative feature extractions and facial expression predictions. Then, the second selection phase defines several rules to choose high confident samples according to prediction scores, and the third relabeling phase assigns two potential labels to those samples for network updating according to a composite two-label loss. Compared with the previous studies, the three-phase learning could effectively correct noisy labels in the ground truth without extra information and automatically assign two potential labels to single-label samples without manual annotations. As a result, the label information is purified and supplemented with few cost, yielding significant performance improvement. Extensive experiments are conducted on three datasets, and the experimental results demonstrate that our approach is robust to noisy training samples and outperforms several state-of-the-art methods. Jin Yuan 0002, Zhiyong Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | JDAN: Joint Detection and Association Network for Real-Time Online Multi-Object TrackingabstractIn the last few years, enormous strides have been made for object detection and data association, which are vital subtasks for one-stage online multi-object tracking (MOT). However, the two separated submodules involved in the whole MOT pipeline are processed or optimized separately, resulting in a complex method design and requiring manual settings. In addition, few works integrate the two subtasks into a single end-to-end network to optimize the overall task. In this study, we propose an end-to-end MOT network called joint detection and association network (JDAN) that is trained and inferred in a single network. All layers in JDAN are differentiable, and can be optimized jointly to detect targets and output an association matrix for robust multi-object tracking. What’s more, we generate suitable pseudo-labels to address the data inconsistency between object detection and association. The detection and association submodules could be optimized by the composite loss function that is derived from the detection results and the generated pseudo association labels, respectively. The proposed approach is evaluated on two MOT challenge datasets, and achieves promising performance compared with classic and latest methods. Zhiyong Li 0001, Jin Yuan 0002, Shutao Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Negative Stiffness Analysis and Regulation of In-Hand Manipulation with Underactuated Compliant HandsabstractThis paper addresses the generation mechanism and avoidance method of negative stiffness during in-Hand manipulation with underactuated compliant hands. Firstly, a planar hand with two three-jointed fingers manipulating a rectangular is set, and a quasi-static underactuated operation model is established. Secondly, based on this simulation model, we investigated the stiffness evolution during in-hand manipulation, and analyze the influence factors of system stiffness. Finally, a stiffness regulation method is developed to avoid negative stiffness during in-hand manipulation. The method is validated by simulation. The research results are beneficial to improve the performance of underactuated in-hand manipulation. Wenrui Chen, Qiang Diao, Yaonan Wang 0001, Cuo Yan, Zhiyong Li 0001 |
ICRA | 7 |
| 2022 | STURE: Spatial-Temporal Mutual Representation Learning for robust data association in online multi-object tracking
Zhiyong Li 0001, Ke Nai |
Comput. Vis. Image Underst. | 2 |
| 2022 | BTN: Neuroanatomical aligning between visual object tracking in deep neural network and smooth pursuit in brain
Zhiyong Li 0001, Ke Nai, Jin Yuan 0002, Shutao Li 0001, Xianghua Li |
Neurocomputing | 2 |
| 2022 | Dynamic feature fusion with spatial-temporal context for robust object tracking
Ke Nai, Zhiyong Li 0001 |
Pattern Recognit. | 2 |
| 2022 | Learning Channel-Aware Correlation Filters for Robust Object TrackingabstractCorrelation filters with Convolutional Neural Networks (CNNs) features have obtained tremendous attention and success in visual tracking. However, redundant and noisy feature channels existed in CNN features may cause severe over-fitting and greatly limit the discriminative power of the tracking model. To tackle the issue, in this paper, we develop a new and effective channel-aware correlation filters (CACF) method for boosting the tracking performance. Our CACF method aims to dynamically select representative and discriminative feature channels from high-dimensional CNN features to reduce the model complexity and better distinguish the target object from the background. Moreover, the CACF model is solved by the alternating direction method of multipliers (ADMM) to learn correlation filters. By retaining reliable feature channels, our CACF tracking method can reach better generalization ability and discriminative ability to accurately localize the target object. Comprehensive experiments are conducted on challenging tracking datasets, and the experiment results prove that our CACF method obtains favorable tracking accuracy compared to several popular tracking methods. Ke Nai, Zhiyong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Discriminative Style Learning for Cross-Domain Image CaptioningabstractThe cross-domain image captioning, which is trained on a source domain and generalized to other domains, usually faces the large domain shift problem. Although prior work has attempted to leverage both paired source and unpaired target data to minimize this shift, the performance is still unsatisfactory. One main reason lies in the large discrepancy in language expression between two domains, where diverse language styles are adopted to describe an image from different views, resulting in different semantic descriptions for an image. To tackle this problem, this paper proposes a Style-based Cross-domain Image Captioner (SCIC) which incorporates the discriminative style information into the encoder-decoder framework, and interprets an image as a special sentence according to external style instructions. Technically, we design a novel "Instruction-based LSTM", which adds the instruct gate to collect a style instruction, and then outputs a specified format according to that instruction. Two objectives are designed to train I-LSTM: 1) generating correct image descriptions and 2) generating correct styles, thus the model is expected to accurately capture the semantic meanings of an image by the special caption as well as understand the syntactic structure of the caption. We use MS-COCO as the source domain, and Oxford-102, CUB-200, Flickr30k as the target domains. Experimental results demonstrate that our model consistently outperforms the previous methods, and the style information incorporating with I-LSTM significantly improves the performance, with 5% CIDEr improvements at least on all datasets. Jin Yuan 0002, Shuai Zhu, Shuyin Huang, Hanwang Zhang, Yaoqiang Xiao, Zhiyong Li 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2022 | Learning a Dynamic Feature Fusion Tracker for Object TrackingabstractObject tracking is a key component of self-driving systems and has important meanings to alleviate traffic accidents. Therefore, it is meaningful to design a high performance and real-time tracker for improving the stability and safety of self-driving systems. In this paper, an effective and efficient feature fusion tracker, which dynamically fuses gradient and color features to model the appearance of the target object, is designed with the correlation filters framework for fast tracking. To be specific, two complementary correlation filters for gradient (e.g. HOG) and color (e.g. ColorNames) features are maintained during tracking, and the proposed feature fusion method adaptively adjusts the weights of them to deal with large appearance changes of the target object in challenging tracking scenes. The weights are decided by the consistency of the final tracking result and the predicted results obtained by two correlation filters. Moreover, a failure detection scheme is designed to alleviate the model drift issue caused by undesirable model updates to improve the tracking accuracy. If a tracking result is identified as a failed case, re-detection operations are performed to accurately localize the target object. The experimental results prove that the proposed tracker can achieve competitive tracking performance and a satisfactory tracking speed of 25.3 FPS in comparison with several state-of-the-art trackers on challenging tracking benchmarks. Zhiyong Li 0001, Ke Nai, Guiji Li, Shilong Jiang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Improving Short Text Classification Using Context-Sensitive Representations and Content-Aware Extended Topic Knowledge
Zhihao Ye, Rui Wen 0001, Xi Chen 0003, Zhiyong Li 0001, Ke Nai, Yefeng Zheng 0001 |
PAKDD (2) | 6 |
| 2021 | Person re-identification with part prediction alignment
Zhiyong Li 0001, Jingyi Lv, Jin Yuan 0002 |
Comput. Vis. Image Underst. | 1 |
| 2021 | Dynamic resource allocation for jointing vehicle-edge deep neural network inference
Zhiyong Li 0001, Ke Nai |
J. Syst. Archit. | 2 |
| 2021 | Siamese target estimation network with AIoU loss for real-time visual tracking
Zhiyong Li 0001, Chenming Hu, Ke Nai, Jin Yuan 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | History-based attention in Seq2Seq model for multi-label text classification
Yaoqiang Xiao, Jin Yuan 0002, Songrui Guo, Yi Xiao 0004, Zhiyong Li 0001 |
Knowl. Based Syst. | 6 |
| 2020 | A Variant of Recurrent Entity Networks for Targeted Aspect-Based Sentiment AnalysisabstractDeep neural network models have achieved promising results on targeted aspect-based sentiment analysis. However, previous models did not effectively match long-distance fine-grained sentiment polarity with the associated target and aspects, and the interdependence among the specific target, corresponding aspects, and the context is always ignored. This work proposes a novel recurrent entity memory network that employs word-level information and sentence-level hidden memory to entity state tracking. In addition, the entity state is utilized to fine-tune the target embedding and aspect embedding. The experimental results showed that the proposed model outperformed previous models. Zhihao Ye, Zhiyong Li 0001 |
ECAI | 2 |
| 2020 | Document and Word Representations Generated by Graph Convolutional Network and BERT for Short Text ClassificationabstractIn many studies, the graph convolution neural networks were used to solve different natural language processing (NLP) problems. However, few researches employ graph convolutional network for text classification, especially for short text classification. In this work, a special text graph of the short-text corpus is created, and then a short-text graph convolutional network (STGCN) is developed. Specifically, different topic models for short text are employed, and a short text short-text graph based on the word co-occurrence, document word relations, and text topic information, is developed. The word and sentence representations generated by the STGCN are considered as the classification feature. In addition, a pre-trained word vector obtained by the BERTs hidden layer is employed, which greatly improves the classification effect of our model. The experimental results show that our model outperforms the state-of-the-art models on multiple short text datasets. Zhihao Ye, Gongyao Jiang, Zhiyong Li 0001, Jin Yuan 0002 |
ECAI | 4 |
| 2020 | A Stackelberg game approach to multiple resources allocation and pricing in mobile edge computing
Zhiyong Li 0001, Bo Yang 0021, Ke Nai, Keqin Li 0001 |
Future Gener. Comput. Syst. | 2 |
| 2020 | Deep feature learning for gender classification with covered/camouflaged facesabstractThe great attention to gender classification is increasing recently as genders carry rich information related to male and female social activities. Extracting discriminating visual representations for gender classification is challenging especially with covered or camouflaged faces. In this work, the authors propose a network that uses a combination of inceptions with variational feature learning (VFL) loss function. The proposed network recognises the gender of normal or covered/camouflaged faces through the middle face part. This network trained on the middle part of the faces that contain both eyes with a small margin from the top‐left corner to the bottom‐right corner of the area of the eyes. Experimental results showed that the proposed network achieved state‐of‐art performance on five public data sets: FEI, SCIEN, AR FACES, LFW, and ADIENCE. They also evaluated the authors’ network on another new collected data set for covered and camouflaged faces and obtained encouraging outcomes. Mohammed Alghaili, Zhiyong Li 0001, Hamdi A. R. Ali |
IET Image Process. | 2 |
| 2020 | Reliable correlation tracking via dual-memory selection model
Guiji Li, Manman Peng, Ke Nai, Zhiyong Li 0001, Keqin Li 0001 |
Inf. Sci. | 4 |
| 2020 | Person re-identification with expanded neighborhoods distance re-ranking
Jingyi Lv, Zhiyong Li 0001, Ke Nai, Jin Yuan 0002 |
Image Vis. Comput. | 2 |
| 2020 | Non-local attention association scheme for online multi-object tracking
Saizhou Wang, Jingyi Lv, Chenming Hu, Zhiyong Li 0001 |
Image Vis. Comput. | 5 |
| 2020 | Real-time traffic sign detection and classification towards real traffic scene
Yiqiang Wu, Zhiyong Li 0001, Ke Nai, Jin Yuan 0002 |
Multim. Tools Appl. | 2 |
| 2020 | Multi-view correlation tracking with adaptive memory-improved update model
Guiji Li, Manman Peng, Ke Nai, Zhiyong Li 0001, Keqin Li 0001 |
Neural Comput. Appl. | 4 |
| 2020 | Gated CNN: Integrating multi-scale feature layers for object detection
Jin Yuan 0002, Heng-Chang Xiong, Yi Xiao 0004, Weili Guan, Meng Wang 0001, Richang Hong, Zhiyong Li 0001 |
Pattern Recognit. | 7 |
| 2020 | Image Captioning with a Joint Attention Mechanism by Visual Concept SamplesabstractThe attention mechanism has been established as an effective method for generating caption words in image captioning; it explores one noticed subregion in an image to predict a related caption word. However, even though the attention mechanism could offer accurate subregions to train a model, the learned captioner may predict wrong, especially for visual concept words, which are the most important parts to understand an image. To tackle the preceding problem, in this article we propose Visual Concept Enhanced Captioner, which employs a joint attention mechanism with visual concept samples to strengthen prediction abilities for visual concepts in image captioning. Different from traditional attention approaches that adopt one LSTM to explore one noticed subregion each time, Visual Concept Enhanced Captioner introduces multiple virtual LSTMs in parallel to simultaneously receive multiple subregions from visual concept samples. Then, the model could update parameters by jointly exploring these subregions according to a composite loss function. Technically, this joint learning is helpful in finding the common characters of a visual concept, and thus it enhances the prediction accuracy for visual concepts. Moreover, by integrating diverse visual concept samples from different domains, our model can be extended to bridge visual bias in cross-domain learning for image captioning, which saves the cost for labeling captions. Extensive experiments have been conducted on two image datasets (MSCOCO and Flickr30K), and superior results are reported when comparing to state-of-the-art approaches. It is impressive that our approach could significantly increase BLUE-1 and F1 scores, which demonstrates an accuracy improvement for visual concepts in image captioning. Jin Yuan 0002, Songrui Guo, Yi Xiao 0004, Zhiyong Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2019 | A Spatial-Aware TrackerabstractIn this paper, a novel spatial-aware tracker (SAT), which utilizes the Siamese network and multiple correlation filters, is proposed to deal with fast motion and model drift problem in visual tracking. Specifically, the Siamese network is first used by an adaptive spatial search strategy to detect the target object in larger search areas. An extended search patch is generated if the target position obtained by the Siamese network is far away from the previous target position. Then, multiple correlation filters perform detection operations on both the extended search patch and the original search patch. With the proposed spatial selection scheme, SAT can accurately track the target object in challenging tracking scenes. By taking advantage of the Siamese network and multiple correlation filters, the proposed SAT tracker can effectively deal with fast motion and model drift problems to achieve better tracking performance. Extensive experimental results demonstrate that the proposed SAT tracker performs superiorly against several state-of-the-art trackers on OTB-2015 tracking benchmark. Zhiyong Li 0001, Ximing Xiang, Ke Nai, Shilong Jiang |
ICIP | 1 |
| 2019 | Chemical reaction optimization for virtual machine placement in cloud computing
Zhiyong Li 0001, Tingkun Yuan, Shilong Jiang |
Appl. Intell. | 1 |
| 2019 | FSB-EA: Fuzzy search bias guided constraint handling technique for evolutionary algorithm
Zhiyong Li 0001, Shiwen Zhang 0004, Shilong Jiang, Yu Gu 0018, Mourad Nouioua |
Expert Syst. Appl. | 1 |
| 2019 | DCDG-EA: Dynamic convergence-diversity guided evolutionary algorithm for many-objective optimization
Zhiyong Li 0001, Mourad Nouioua, Shilong Jiang, Yu Gu 0018 |
Expert Syst. Appl. | 1 |
| 2019 | Person re-identification based on re-ranking with expanded k-reciprocal nearest neighbors
Jin Yuan 0002, Zhiyong Li 0001, Yiqiang Wu, Mourad Nouioua, Guoqi Xie |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Multi-pattern correlation tracking
Ke Nai, Degui Xiao, Zhiyong Li 0001, Shilong Jiang, Yu Gu 0018 |
Knowl. Based Syst. | 3 |
| 2019 | DELR: A double-level ensemble learning method for unsupervised anomaly detection
Jia Zhang 0005, Zhiyong Li 0001, Ke Nai, Yu Gu 0018, Ahmed Sallam |
Knowl. Based Syst. | 2 |
| 2018 | Envy-free auction mechanism for VM pricing and allocation in clouds
Bo Yang 0021, Zhiyong Li 0001, Shilong Jiang, Keqin Li 0001 |
Future Gener. Comput. Syst. | 2 |
| 2018 | Visual tracking via context-aware local sparse appearance model
Guiji Li, Manman Peng, Ke Nai, Zhiyong Li 0001, Keqin Li 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Visual Tracking With Weighted Adaptive Local Sparse Appearance Model via Spatio-Temporal Context LearningabstractSparse representation has been widely exploited to develop an effective appearance model for object tracking due to its well discriminative capability in distinguishing the target from its surrounding background. However, most of these methods only consider either the holistic representation or the local one for each patch with equal importance, and hence may fail when the target suffers from severe occlusion or large-scale pose variation. In this paper, we propose a simple yet effective approach that exploits rich feature information from reliable patches based on weighted local sparse representation that takes into account the importance of each patch. Specifically, we design a reconstruction-error based weight function with the reconstruction error of each patch via sparse coding to measure the patch reliability. Moreover, we explore spatio-temporal context information to enhance the robustness of the appearance model, in which the global temporal context is learned via incremental subspace and sparse representation learning with a novel dynamic template update strategy to update the dictionary, while the local spatial context considers the correlation between the target and its surrounding background via measuring the similarity among their sparse coefficients. Extensive experimental evaluations on two large tracking benchmarks demonstrate favorable performance of the proposed method over some state-of-the-art trackers. Zhetao Li, Jie Zhang 0136, Kaihua Zhang 0001, Zhiyong Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Robust Object Tracking via Local Sparse Appearance ModelabstractIn this paper, we propose a novel local sparse representation-based tracking framework for visual tracking. To deeply mine the appearance characteristics of different local patches, the proposed method divides all local patches of a candidate target into three categories, which are stable patches, valid patches, and invalid patches. All these patches are assigned different weights to consider the different importance of the local patches. For stable patches, we introduce a local sparse score to identify them, and discriminative local sparse coding is developed to decrease the weights of background patches among the stable patches. For valid patches and invalid patches, we adopt local linear regression to distinguish the former from the latter. Furthermore, we propose a weight shrinkage method to determine weights for different valid patches to make our patch weight computation more reasonable. Experimental results on public tracking benchmarks with challenging sequences demonstrate that the proposed method performs favorably against other state-of-the-art tracking methods. Ke Nai, Zhiyong Li 0001, Guiji Li, Shanquan Wang |
IEEE Trans. Image Process. | 2 |
| 2017 | Using differential evolution strategies in chemical reaction optimization for global numerical optimization
Mourad Nouioua, Zhiyong Li 0001 |
Appl. Intell. | 2 |
| 2017 | Robust object tracking based on adaptive templates matching via the fusion of multiple features
Zhiyong Li 0001, Ke Nai |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Quantum-Inspired Hyper-Heuristics for Energy-Aware Scheduling on Heterogeneous Computing SystemsabstractPower and performance tradeoff optimization is one of the most significant issues on heterogeneous multiprocessor or multicomputer systems (HMCSs) with dynamically variable voltage. In this paper, the problem is defined as energy-constrained performance optimization and performance-constrained energy optimization. Task scheduling for precedence-constrained parallel applications represented by a directed acyclic graph (DAG) in HMCSs is an NP-HARD problem. Over the last three decades, several task scheduling techniques have been developed for energy-aware scheduling. However, it is impossible for a single task scheduling technique to outperform all other techniques for all types of applications and situations. Motivated by these observations, hyperheuristic framework is introduced. Moreover, a quantum-inspired high-level learning strategy is proposed to improve the performance of this framework. Meanwhile, a fast solution evaluation technique is designed to reduce the computational burden for each iteration step. Experimental results show that the fast solution evaluation technique can improve average algorithm search speed by 38 percent and that the proposed algorithm generally exhibits outstanding convergence performance. Zhiyong Li 0001, Bo Yang 0021, Günter Rudolph |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Stackelberg Game Approach for Energy-Aware Resource Allocation in Data CentersabstractData centers hosting distributed computing systems consume huge amounts of electrical energy, contributing to high operational costs, whereas the utilization of data centers continues to be very low. Moreover, a data center generally consists of heterogeneous servers with different performance and energy. Failure to fully consider the heterogeneity of servers will lead to both sub-optimal energy saving and performance. In this study, we employ game theoretic approaches to model the problem of minimizing energy consumption as a Stackelberg game. In our model, the system monitor, who plays the role of the leader, can maximize profit by adjusting resource provisioning, whereas scheduler agents, who act as followers, can select resources to obtain optimal performance. In addition, we model the problem of minimizing average response time of tasks as a noncooperative game among decentralized scheduler agents as they compete with one another in the sharing resources. Several algorithms are presented to implement the game models. Simulation results demonstrate that the proposed technique has immense potential to improve energy efficiency under dynamic work scenarios without compromising service level agreements. Bo Yang 0021, Zhiyong Li 0001, Tao Wang 0016, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Orthogonal chemical reaction optimization algorithm for global numerical optimization problems
Zhiyong Li 0001, Tien Trong Nguyen |
Expert Syst. Appl. | 1 |
| 2015 | PS-ABC: A hybrid algorithm based on particle swarm and artificial bee colony for high-dimensional optimization problems
Zhiyong Li 0001, Weiyou Wang, Yanyan Yan |
Expert Syst. Appl. | 1 |
| 2015 | Robust object tracking via multi-feature adaptive fusion based on stability: contrast analysis
Zhiyong Li 0001, Shuang He, Mervat Hashem |
Vis. Comput. | 1 |
| 2015 | Moving object tracking based on multi-independent features distribution fields with comprehensive spatial feature similarity
Zhiyong Li 0001, Xiaoping Yu, Mervat Hashem |
Vis. Comput. | 1 |
| 2014 | A hybrid algorithm based on particle swarm and chemical reaction optimization
Tien Trong Nguyen, Zhiyong Li 0001, Shiwen Zhang 0004, Tung Khac Truong |
Expert Syst. Appl. | 2 |
| 2014 | Proactive workload management in dynamic virtualized environments
Ahmed Sallam, Kenli Li 0001, Aijia Ouyang, Zhiyong Li 0001 |
J. Comput. Syst. Sci. | 4 |
| 2009 | Fourier Volume Rendering on GPGPU
Degui Xiao, Zhiyong Li 0001, Kenli Li 0001 |
ISNN (3) | 4 |
| 2009 | An Efficient Large-Scale Volume Data Compression Algorithm
Degui Xiao, Liping Zhao 0005, Zhiyong Li 0001, Kenli Li 0001 |
ISNN (3) | 4 |
| 2007 | A framework of quantum-inspired multi-objective evolutionary algorithms and its convergence conditionabstractA general framework of quantum-inspired multi-objective evolutionary algorithms as well as one of its sufficient convergence conditions to Pareto optimal set is proposed. Zhiyong Li 0001, Günter Rudolph |
GECCO | 1 |
| 2007 | On the Convergence Properties of Quantum-Inspired Multi-Objective Evolutionary Algorithms
Zhiyong Li 0001, Günter Rudolph |
ICIC (3) | 1 |