Pinxue Guo

dblp:333/7534 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
26since 2021 · last 2026
0000-0002-4388-9757ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Seeing Is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
abstract
Multimodal Large Language Models (MLLMs) have unlocked powerful cross-modal capabilities, but still significantly suffer from hallucinations. As such, accurate detection of hallucinations in MLLMs is imperative for ensuring their reliability in practical applications. To this end, guided by the principle of “Seeing is Believing”, we introduce VBackChecker, a novel reference-free hallucination detection framework that verifies the consistency of MLLM-generated responses with visual inputs, by leveraging a pixel-level Grounding LLM equipped with reasoning and referring segmentation capabilities. This referencefree framework not only effectively handles rich-context scenarios, but also offers interpretability. To facilitate this, an innovative pipeline is accordingly designed for generating instruction-tuning data (R-Instruct), featuring richcontext descriptions, grounding masks, and hard negative samples. We further establish R 2 -HalBench, a new hallucination benchmark for MLLMs, which, unlike previous benchmarks, encompasses real-world, rich-context descriptions from 18 MLLMs with high-quality annotations, spanning diverse object-, attribute-, and relationship-level details. VBackChecker outperforms prior complex frameworks and achieves state-of-the-art performance on R^2 -HalBench, even rivaling GPT-4o’s capabilities in hallucination detection. It also surpasses prior methods in the pixel-level grounding task, achieving over a 10% improvement.
Pinxue Guo, Chongruo Wu, Xinyu Zhou 0006, Lingyi Hong, Zhaoyu Chen 0001, Kaixun Jiang, Sen-Ching S. Cheung, Wei Zhang 0016
AAAI1
2026 LVOS: A Benchmark for Large-Scale Long-Term Video Object Segmentation
abstract
Video object segmentation (VOS) aims to distinguish and track target objects in a video. Despite the excellent performance achieved by off-the-shelf VOS models, part of the existing VOS benchmarks mainly focuses on short-term videos, where objects remain visible most of the time. However, these benchmarks may not fully capture challenges encountered in practical applications, and the absence of long-term datasets restricts further investigation of VOS in realistic scenarios. Thus, we propose a novel benchmark named LVOS, comprising 720 videos with 296,401 frames and 407,945 high-quality annotations. Videos in LVOS last 1.14 minutes on average. Each video includes various attributes, especially challenges encountered in the wild, such as long-term reappearing and cross-temporal similar objects. Compared to previous benchmarks, our LVOS better reflects VOS models' performance in real scenarios. Based on LVOS, we evaluate 15 existing VOS models under 3 different settings and conduct a comprehensive analysis. On LVOS, these models suffer a large performance drop, highlighting the challenge of achieving precise tracking and segmentation in real-world scenarios. Attribute-based analysis indicates that one of the significant factors contributing to accuracy decline is the increased video length, interacting with complex challenges such as long-term reappearance, cross-temporal confusion, and occlusion, which emphasize LVOS's crucial role. We hope our LVOS can advance development of VOS in real scenes.
Lingyi Hong, Zhongying Liu, Chenzhi Tan, Yuang Feng, Xinyu Zhou 0006, Pinxue Guo, Zhaoyu Chen 0001, Shuyong Gao, Wei Zhang 0016
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 ClickVOS: Click Video Object Segmentation
abstract
Video Object Segmentation (VOS) task aims to segment objects in videos. However, previous settings either require time-consuming manual masks of target objects at the first frame during inference or lack the flexibility to specify arbitrary objects of interest. To address these limitations, we propose the setting named Click Video Object Segmentation (ClickVOS) which segments objects of interest across the whole video according to a single click per object in the first frame. And we provide the extended datasets DAVIS-P and YouTubeVOS-P that with point annotations to support this task. ClickVOS is of significant practical applications and research implications due to its only 1-2 seconds interaction time for indicating an object, comparing annotating the mask of an object needs several minutes. However, ClickVOS also presents increased challenges. To address this task, we propose an end-to-end baseline approach named called Attention Before Segmentation (ABS), motivated by the attention process of humans. ABS utilizes the given point in the first frame to perceive the target object through a concise yet effective segmentation attention. Although the initial object mask is possibly inaccurate, in our ABS, as the video goes on, the initially imprecise object mask can self-heal instead of deteriorating due to error accumulation, which is attributed to our designed improvement memory that continuously records stable global object memory and updates detailed dense memory. In addition, we conduct various baseline explorations utilizing off-the-shelf algorithms from related fields, which could provide insights for the further exploration of ClickVOS. The experimental results demonstrate the superiority of the proposed ABS approach. Extended datasets and codes will be available at https://github.com/PinxueGuo/ClickVOS.
Pinxue Guo, Lingyi Hong, Xinyu Zhou 0006, Shuyong Gao, Wanyun Li, Zhaoyu Chen 0001, Xiaoqiang Li 0002, Wei Zhang 0016
IEEE Trans. Circuits Syst. Video Technol.1
2025 OpenVIS: Open-vocabulary Video Instance Segmentation
abstract
Open-vocabulary Video Instance Segmentation (OpenVIS) can simultaneously detect, segment, and track arbitrary object categories in a video, without being constrained to categories seen during training. In this work, we propose InstFormer, a carefully designed framework for the OpenVIS task that achieves powerful open-vocabulary capabilities through lightweight fine-tuning with limited-category data. InstFormer begins with the open-world mask proposal network, encouraged to propose all potential instance class-agnostic masks by the contrastive instance margin loss. Next, we introduce InstCLIP, adapted from pre-trained CLIP with Instance Guidance Attention, which encodes open-vocabulary instance tokens efficiently. These instance tokens not only enable open-vocabulary classification but also offer strong universal tracking capabilities. Furthermore, to prevent the tracking module from being constrained by the training data with limited categories, we propose the universal rollout association, which transforms the tracking problem into predicting the next frame’s instance tracking token. The experimental results demonstrate the proposed InstFormer achieve state-of-the-art capabilities on a comprehensive OpenVIS evaluation benchmark, while also achieves competitive performance in fully supervised VIS task.
Pinxue Guo, Peiyang He, Tianjun Xiao
AAAI1
2025 General Compression Framework for Efficient Transformer Object Tracking
abstract
Previous works have attempted to improve tracking efficiency through lightweight architecture design or knowledge distillation from teacher models to compact student trackers. However, these solutions often sacrifice accuracy for speed to a great extent, and also have the problems of complex training process and structural limitations. Thus, we propose a general model compression framework for efficient transformer object tracking, named CompressTracker, to reduce model size while preserving tracking accuracy. Our approach features a novel stage division strategy that segments the transformer layers of the teacher model into distinct stages to break the limitation of model structure. Additionally, we also design a unique replacement training technique that randomly substitutes specific stages in the student model with those from the teacher model, as opposed to training the student model in isolation. Replacement training enhances the student model's ability to replicate the teacher model's behavior and simplifies the training process. To further forcing student model to emulate teacher model, we incorporate prediction guidance and stage-wise feature mimicking to provide additional supervision during the teacher model's compression process. CompressTracker is structurally agnostic, making it compatible with any transformer architecture. We conduct a series of experiment to verify the effectiveness and generalizability of our CompressTracker. Our CompressTracker-SUTrack, compressed from SUTrack, retains about 99 performance on LaSOT (72.2 AUC) while achieves 2.42x speed up. Code is available at https://github.com/LingyiHongfd/CompressTracker.
Lingyi Hong, Xinyu Zhou 0006, Shilin Yan, Pinxue Guo, Kaixun Jiang, Zhaoyu Chen 0001, Shuyong Gao, Xingdong Sheng, Wei Zhang 0016, Hong Lu 0001
ICCV5
2025 VideoSAM: Open-World Video Segmentation
abstract
Video segmentation is essential for advancing robotics and autonomous driving, particularly in open-world settings where continuous perception and object association across video frames are critical. While the Segment Anything Model (SAM) has excelled in static image segmentation, extending its capabilities to video segmentation poses significant challenges. We tackle two major hurdles: a) SAM's embedding limitations in associating objects across frames, and b) granularity inconsistencies in object segmentation. To this end, we introduce VideoSAM, an end-to-end framework designed to address these challenges by improving object tracking and segmentation consistency in dynamic environments. VideoSAM integrates an agglomerated backbone, RADIO, enabling object association through similarity metrics and introduces Cycle-ack-Pairs Propagation with a memory mechanism for stable object tracking. Additionally, we incorporate an autoregressive object-token mechanism within the SAM decoder to maintain consistent granularity across frames. Our method is extensively evaluated on the UVO and BURST benchmarks, and robotic videos from RoboTAP, demonstrating its effectiveness and robustness in real-world scenarios. All codes will be available.
Pinxue Guo, Jianxiong Gao, Chongruo Wu, Tong He 0002, Zheng Zhang 0001, Tianjun Xiao
ICRA1
2025 Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment
abstract
Preference alignment in diffusion models has primarily focused on benign human preferences (e.g., aesthetic). In this paper, we propose a novel perspective: framing unrestricted adversarial example generation as a problem of aligning with adversary preferences. Unlike benign alignment, adversarial alignment involves two inherently conflicting preferences: visual consistency and attack effectiveness, which often lead to unstable optimization and reward hacking (e.g., reducing visual quality to improve attack success). To address this, we propose APA (Adversary Preferences Alignment), a two-stage framework that decouples conflicting preferences and optimizes each with differentiable rewards. In the first stage, APA fine-tunes LoRA to improve visual consistency using rule-based similarity reward. In the second stage, APA updates either the image latent or prompt embedding based on feedback from a substitute classifier, guided by trajectory-level and step-wise rewards. To enhance black-box transferability, we further incorporate a diffusion augmentation strategy. Experiments demonstrate that APA achieves significantly better attack transferability while maintaining high visual consistency, inspiring further research to approach adversarial attacks from an alignment perspective.
Kaixun Jiang, Zhaoyu Chen 0001, Haijing Guo, Jiyuan Fu, Pinxue Guo, Hao Tang 0005, Bo Li 0115
NeurIPS6
2025 Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
abstract
UAV tracking can be widely applied in scenarios such as disaster rescue, environmental monitoring, and logistics transportation. However, existing UAV tracking methods predominantly emphasize speed and lack exploration in semantic awareness, which hinders the search region from extracting accurate localization information from the template. The limitation results in suboptimal performance under typical UAV tracking challenges such as camera motion, fast motion, and low resolution, etc. To address this issue, we propose a dynamic semantic aware correlation modeling tracking framework. The core of our framework is a Dynamic Semantic Relevance Generator, which, in combination with the correlation map from the Transformer, explore semantic relevance. The approach enhances the search region's ability to extract important information from the template, improving accuracy and robustness under the aforementioned challenges. Additionally, to enhance the tracking speed, we design a pruning method for the proposed framework. Therefore, we present multiple model variants that achieve trade-offs between speed and accuracy, enabling flexible deployment according to the available computational resources. Experimental results validate the effectiveness of our method, achieving competitive performance on multiple UAV tracking datasets.
Xinyu Zhou 0006, Tongxin Pan, Lingyi Hong, Pinxue Guo, Haijing Guo, Zhaoyu Chen 0001, Kaixun Jiang
NeurIPS4
2025 Self-supervised video object segmentation via pseudo label rectification
Pinxue Guo, Wei Zhang 0016, Xiaoqiang Li 0002, Jianping Fan 0007
Pattern Recognit.1
2024 Out of Thin Air: Exploring Data-Free Adversarial Robustness Distillation
abstract
Adversarial Robustness Distillation (ARD) is a promising task to solve the issue of limited adversarial robustness of small capacity models while optimizing the expensive computational costs of Adversarial Training (AT). Despite the good robust performance, the existing ARD methods are still impractical to deploy in natural high-security scenes due to these methods rely entirely on original or publicly available data with a similar distribution. In fact, these data are almost always private, specific, and distinctive for scenes that require high robustness. To tackle these issues, we propose a challenging but significant task called Data-Free Adversarial Robustness Distillation (DFARD), which aims to train small, easily deployable, robust models without relying on data. We demonstrate that the challenge lies in the lower upper bound of knowledge transfer information, making it crucial to mining and transferring knowledge more efficiently. Inspired by human education, we design a plug-and-play Interactive Temperature Adjustment (ITA) strategy to improve the efficiency of knowledge transfer and propose an Adaptive Generator Balance (AGB) module to retain more data information. Our method uses adaptive hyperparameters to avoid a large number of parameter tuning, which significantly outperforms the combination of existing techniques. Meanwhile, our method achieves stable and reliable performance on multiple benchmarks.
Zhaoyu Chen 0001, Dingkang Yang, Pinxue Guo, Kaixun Jiang, Lizhe Qi
AAAI4
2024 OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning
abstract
Visual object tracking aims to localize the target object of each frame based on its initial appearance in the first frame. Depending on the input modility, tracking tasks can be divided into RGB tracking and RGB+X (e.g. RGB+N, and RGB+D) tracking. Despite the different input modalities, the core aspect of tracking is the temporal matching. Based on this common ground, we present a general framework to unify various tracking tasks, termed as One Tracker. One- Tracker first performs a large-scale pre-training on a RGB tracker called Foundation Tracker. This pretraining phase equips the Foundation Tracker with a stable ability to estimate the location of the target object. Then we regard other modality information as prompt and build Prompt Tracker upon Foundation Tracker. Through freezing the Foundation Tracker and only adjusting some additional trainable parameters, Prompt Tracker inhibits the strong localization ability from Foundation Tracker and achieves parameter- efficient finetuning on downstream RGB+X tracking tasks. To evaluate the effectiveness of our general framework OneTracker, which is consisted of Foundation Tracker and Prompt Tracker, we conduct extensive experiments on 6 popular tracking tasks across 11 benchmarks and our One- Tracker outperforms other models and achieves state-of-the-art performance.
Lingyi Hong, Shilin Yan, Renrui Zhang, Wanyun Li, Xinyu Zhou 0006, Pinxue Guo, Kaixun Jiang, Zhaoyu Chen 0001
CVPR6
2024 OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework
Wanyun Li, Pinxue Guo, Xinyu Zhou 0006, Lingyi Hong, Yangji He, Wei Zhang 0016
ECCV (58)2
2024 X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
Pinxue Guo, Wanyun Li, Lingyi Hong, Xinyu Zhou 0006, Zhaoyu Chen 0001, Kaixun Jiang, Wei Zhang 0016
ACM Multimedia1
2024 TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
Xinyu Zhou 0006, Kaixun Jiang, Lingyi Hong, Pinxue Guo, Zhaoyu Chen 0001, Weifeng Ge
ACM Multimedia5
2024 DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
abstract
Previous visual object tracking methods employ image-feature regression models or coordinate autoregression models for bounding box prediction. Image-feature regression methods heavily depend on matching results and do not utilize positional prior, while the autoregressive approach can only be trained using bounding boxes available in the training set, potentially resulting in suboptimal performance during testing with unseen data. Inspired by the diffusion model, denoising learning enhances the model’s robustness to unseen data. Therefore, We introduce noise to bounding boxes, generating noisy boxes for training, thus enhancing model robustness on testing data. We propose a new paradigm to formulate the visual object tracking problem as a denoising learning process. However, tracking algorithms are usually asked to run in real-time, directly applying the diffusion model to object tracking would severely impair tracking speed. Therefore, we decompose the denoising learning process into every denoising block within a model, not by running the model multiple times, and thus we summarize the proposed paradigm as an in-model latent denoising learning process. Specifically, we propose a denoising Vision Transformer (ViT), which is composed of multiple denoising blocks. In the denoising block, template and search embeddings are projected into every denoising block as conditions. A denoising block is responsible for removing the noise in a predicted bounding box, and multiple stacked denoising blocks cooperate to accomplish the whole denoising process. Subsequently, we utilize image features and trajectory information to refine the denoised bounding box. Besides, we also utilize trajectory memory and visual memory to improve tracking stability. Experimental results validate the effectiveness of our approach, achieving competitive performance on several challenging datasets. The proposed in-model latent denoising tracker achieve real-time speed, rendering denoising learning applicable in the visual object tracking community.
Xinyu Zhou 0006, Lingyi Hong, Kaixun Jiang, Pinxue Guo, Weifeng Ge
NeurIPS5
2024 Boosting the transferability of adversarial attacks with global momentum initialization
Zhaoyu Chen 0001, Kaixun Jiang, Dingkang Yang, Lingyi Hong, Pinxue Guo, Haijing Guo
Expert Syst. Appl.6
2024 HFVOS: History-Future Integrated Dynamic Memory for Video Object Segmentation
abstract
Memory-based methods have substantially enhanced the precision of video object segmentation (VOS) by storing features in an expanding memory bank. However, this comes at the cost of increased computational demands and storage overhead. While recent methods have sought to alleviate this issue via compression or selection strategies, their reliance solely on history cues and simple memory structures result in precision degradation and intrinsic limitations, such as error accumulation and poor robustness. In this paper, we introduce HFVOS, an efficient yet effective framework to bolster VOS performance in both speed and precision by meticulously considering the memory design with low redundancy, high accuracy, and adaptability. First, we construct a novel hierarchical memory update pipeline with the proposed Buffered Memory Mechanism, which incorporates both future and history cues to reduce redundancy and improve the utility of memory. Second, we propose an Adaptive Dual-stream Selection Network (ADSN) to carry out the adaptive selection and drop operations of the memory update, and integrate an ADSN based long-term memory to enhance the robustness, especially for long videos. Furthermore, to further boost HFVOS, a progressive selection loss is designed to facilitate ADSN gradually adapt to fewer features while preserving high precision. Experiments show that HFVOS achieves the state-of-the-art segmentation precision and speed on both short-term datasets (DAVIS-17 val: 86.8%J&Fand 33.0 FPS, DAVIS-16 val: 92.0%J&Fand 42.0 FPS) and long-term datasets (LVOS val: 58.0%J&Fand 37.4 FPS). Code will be available at https://github.com/L599wy/HFVOS.
Wanyun Li, Jack Fan, Pinxue Guo, Lingyi Hong, Wei Zhang 0016
IEEE Trans. Circuits Syst. Video Technol.3
2023 Gated Enhanced RPN and Hybrid-View for Few-Shot Object Detection
abstract
Few-Shot Object Detection (FSOD) is designed to detect unseen objects using a few examples. Matching between query image and support instance via metric learning has been shown to be an effective FSOD method. In previous work, the quality of proposals generated by the RPN is not yet high due to the lack of a fine-grained matching strategy, and the detector performs feature integration only at the spatial location level is limited. Therefore, in this paper, we propose a new method for few-shot object detection which obtains high-quality proposals by the Gated Enhanced RPN (GRPN). To further improve the detection performance from different views, we propose a Hybrid-View Detector to categorize proposals more comprehensively. We perform extensive experiments on the MS-COCO benchmark and the experimental results demonstrate the effectiveness of the proposed method and outperform the state-of-the-art on most shot cases.
Xujun Wei, Zechu Zhou, Pinxue Guo
ICASSP3
2023 Robust Video Object Segmentation with Restricted Attention
abstract
This paper focuses on the two problems of the similar objects distraction and the lack of robustness for unseen object categories in semi-supervised video object segmentation task. Existing methods have achieved great results on the benchmark dataset, but these two problems still have not been completely solved. We propose the Robust Video Object Segmentation With Restricted Attention (RVOSR), which can suppress the effects caused by similar objects and filter out noise confusion from other irrelevant regions. Meanwhile augmenting the semantic information of features, which makes the features more suitable for video object segmentation task. Extensive experiments demonstrate the effectiveness of our approach and achieve the state-of-the-art performance on the widely-used VOS benchmarks including DAVIS-2016 (92.1% $\mathcal{J}{{\& }}\mathcal{F}$), DAVIS-2017 (86.8% $\mathcal{J}{{\& }}\mathcal{F}$) and YouTubeVOS-2019 (84.8%).
Huaizheng Zhang, Pinxue Guo, Zhongwen Le
ICASSP2
2023 LVOS: A Benchmark for Long-term Video Object Segmentation
abstract
Existing video object segmentation (VOS) benchmarks focus on short-term videos which just last about 3-5 seconds and where objects are visible most of the time. These videos are poorly representative of practical applications, and the absence of long-term datasets restricts further investigation of VOS on the application in realistic scenarios. So, in this paper, we present a new benchmark dataset named LVOS, which consists of 220 videos with a total duration of 421 minutes. To the best of our knowledge, LVOS is the first densely annotated long-term VOS dataset. The videos in our LVOS last 1.59 minutes on average, which is 20 times longer than videos in existing VOS datasets. Each video includes various attributes, especially challenges deriving from the wild, such as long-term reappearing and cross-temporal similar objeccts. Based on LVOS, we assess existing video object segmentation algorithms and propose a Diverse Dynamic Memory network (DDMemory) that consists of three complementary memory banks to exploit temporal information adequately. The experimental results demonstrate the strength and weaknesses of prior methods, pointing promising directions for further study. Data and code are available at https://lingyihongfd.github.io/lvos.github.io/.
Lingyi Hong, Zhongying Liu, Wei Zhang 0016, Pinxue Guo, Zhaoyu Chen 0001
ICCV5
2023 Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution Detection
abstract
Detecting out-of-distribution inputs for visual recognition models has become critical in safe deep learning. This paper proposes a novel hierarchical visual category modeling scheme to separate out-of-distribution data from in-distribution data through joint representation learning and statistical modeling. We learn a mixture of Gaussian models for each in-distribution category. There are many Gaussian mixture models to model different visual categories. With these Gaussian models, we design an in-distribution score function by aggregating multiple Mahalanobis-based metrics. We don’t use any auxiliary outlier data as training samples, which may hurt the generalization ability of out-of-distribution detection algorithms. We split the ImageNet-1k dataset into ten folds randomly. We use one fold as the in-distribution dataset and the others as out-of-distribution datasets to evaluate the proposed method. We also conduct experiments on seven popular benchmarks, including CIFAR, iNaturalist, SUN, Places, Textures, ImageNet-O, and OpenImage-O. Extensive experiments indicate that the proposed method outperforms state-of-the-art algorithms clearly. Meanwhile, we find that our visual representation has a competitive performance when compared with features learned by classical methods. These results demonstrate that the proposed method hasn’t weakened the discriminative ability of visual recognition models and keeps high efficiency in detecting out-of-distribution samples.
Xinyu Zhou 0006, Pinxue Guo, Yixuan Sun, Weifeng Ge
ICCV3
2023 Exploring the Adversarial Robustness of Video Object Segmentation via One-shot Adversarial Attacks
abstract
Video object segmentation (VOS) is a fundamental task for computer vision and multimedia. Despite significant progress of VOS models in recent works, there has been little research on the VOS models' adversarial robustness, posing serious security risks in the VOS models' practical applications (e.g., autonomous driving and video surveillance). Adversarial robustness refers to the ability of the model to resist malicious attacks on adversarial examples. To address this gap, we propose a one-shot adversarial robustness evaluation framework (i.e., the adversary only perturbs the first frame) for VOS models, including white-box and black-box attacks. For white-box attacks, we introduce Objective Attention (OA) and Boundary Attention (BA) mechanisms to enhance the attention of attack on objects from both pixel and object levels while mitigating issues such as multi-objects attack imbalance, attack bias towards the background, and boundary reservation. For black-box attacks, we propose the Video Diverse Input (VDI) module, which utilizes data augmentation to simulate historical information, improving our method's black-box transferability. We conduct extensive experiments to evaluate the adversarial robustness of VOS models with different structures. Our experimental results reveal that existing VOS models are more vulnerable to our attacks (both white-box and black-box) compared to other state-of-the-art attacks. We further analyze the influence of different designs (e.g., memory and matching mechanisms) on adversarial robustness. Finally, we provide insights for designing more secure VOS models in the future.
Kaixun Jiang, Lingyi Hong, Zhaoyu Chen 0001, Pinxue Guo, Zeng Tao, Yan Wang 0068
ACM Multimedia4
2023 A Capture to Registration Framework for Realistic Image Super-Resolution in the Industry Environment
abstract
The acquisition and processing of visual data in industrial environments are of paramount importance. High-resolution (HR) images offer superior clarity and richer textural detail compared to low-resolution (LR) images. On the one hand, owing to the incorporation of richer information, HR images demonstrate substantially enhanced performance compared to LR images in downstream applications, such as anomaly detection. On the other hand, they provide valuable insights to designers and quality inspectors who require a detailed understanding of the images. Currently, the majority of research on super-resolution focuses on natural scenes such as cities and fields, however, the development of datasets for industrial scenes is still in its infancy. To address the image distortion in building realistic LR-HR image pairs in the industry environment, we design a capture to registration framework. It consists of the standard imaging system, physical calibration of the imaging system, as well as the rigid to elastic registration of the LR-HR image pairs. Thus, we build the first realistic industrial sence super-resolution dataset (IndSR), comprises of 50 sets of calibrated images with three scale factors and five typical defects. To benchmark IndSR, we employ quantitative, qualitative, and task-oriented studies to evaluate the representative super-resolution and anomaly detection methods. Besides, we systematically investigate and discuss the performances and results of the existing SISR methods to advance research in the field of super-resolution in industry environment. The IndSR dataset can be available from https://byw4ng.github.io/IndSR/.
Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Junxiong Lin, Zeng Tao, Pinxue Guo, Zhaoyu Chen 0001, Kaixun Jiang, Shaoqi Yan, Shuyong Gao
ACM Multimedia6
2023 Reading Relevant Feature from Global Representation Memory for Visual Object Tracking
abstract
Reference features from a template or historical frames are crucial for visual object tracking. Prior works utilize all features from a fixed template or memory for visual object tracking. However, due to the dynamic nature of videos, the required reference historical information for different search regions at different time steps is also inconsistent. Therefore, using all features in the template and memory can lead to redundancy and impair tracking performance. To alleviate this issue, we propose a novel tracking paradigm, consisting of a relevance attention mechanism and a global representation memory, which can adaptively assist the search region in selecting the most relevant historical information from reference features. Specifically, the proposed relevance attention mechanism in this work differs from previous approaches in that it can dynamically choose and build the optimal global representation memory for the current frame by accessing cross- frame information globally. Moreover, it can flexibly read the relevant historical information from the constructed memory to reduce redundancy and counteract the negative effects of harmful information. Extensive experiments validate the effectiveness of the proposed method, achieving competitive performance on five challenging datasets with 71 FPS.
Xinyu Zhou 0006, Pinxue Guo, Lingyi Hong, Wei Zhang 0016, Weifeng Ge
NeurIPS2
2023 Memory Network With Pixel-Level Spatio-Temporal Learning for Visual Object Tracking
abstract
Making full use of temporal and spatial information is critical to cope with the appearance changes of objects in visual object tracking. However, existing methods in the tracking field, which employ a memory network at frame level to learn this information, bring redundancy and cannot build long-term relationships among historical frames due to the limited memory size. In this paper, we propose a novel memory network, Pixel-level Spatio-Temporal Memory (PSTM), which organizes object features in an efficient way to leverage temporal and spatial context information. Specifically, PSTM is constructed and updated by a memory writer, which includes a pixel-level updating strategy to maintain the temporal consistency and dynamically memorize the noteworthy variations. Furthermore, in order to exploit relationships between the object and search region and precisely estimate the state of the object, we propose a memory reader, Pixel-wise Matching and Refinement module (PMR), and model spatial context without a complex manual-designed mechanism. Comprehensive experiments and comparisons on challenging large-scale benchmarks, including GOT-10k, TrackingNet, LaSOT, OTB2015, VOT2020, and NfS, have demonstrated the effectiveness of our proposed method, which performs favorably against state-of-the-art trackers.
Zechu Zhou, Xinyu Zhou 0006, Zhaoyu Chen 0001, Pinxue Guo, Qian-Yu Liu
IEEE Trans. Circuits Syst. Video Technol.4
2022 Adaptive Online Mutual Learning Bi-Decoders for Video Object Segmentation
abstract
One of the major challenges facing video object segmentation (VOS) is the gap between the training and test datasets due to unseen category in test set, as well as object appearance change over time in the video sequence. To overcome such challenges, an adaptive online framework for VOS is developed with bi-decoders mutual learning. We learn object representation per pixel with bi-level attention features in addition to CNN features, and then feed them into mutual learning bi-decoders whose outputs are further fused to obtain the final segmentation result. We design an adaptive online learning mechanism via a deviation correcting trigger such that bi-decoders online mutual learning will be activated when the previous frame is segmented well meanwhile the current frame is segmented relatively worse. Knowledge distillation from the well segmented previous frames, along with mutual learning between bi-decoders, improves generalization ability and robustness of VOS model. Thus, the proposed model adapts to the challenging scenarios including unseen categories, object deformation, and appearance variation during inference. We extensively evaluate our model on widely-used VOS benchmarks including DAVIS-2016, DAVIS-2017, YouTubeVOS-2018, YouTubeVOS-2019, and UVO. Experimental results demonstrate the superiority of the proposed model over state-of-the-art methods.
Pinxue Guo, Wei Zhang 0016, Xiaoqiang Li 0002
IEEE Trans. Image Process.1