Libo Zhang 0001

dblp:78/33-1 · DBLP profile ↗
← Back
78ranked-venue papers
11as first author
57since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 7 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 2 first-author · 29 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorComputer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Structured Context Learning for Generic Event Boundary Detection
abstract
Generic Event Boundary Detection (GEBD) aims to identify moments in videos that humans perceive as event boundaries. This paper proposes a novel method for addressing this task, called Structured Context Learning, which introduces the Structured Partition of Sequence (SPoS) to provide a structured context for learning temporal information. Our approach is end-to-end trainable and flexible, not restricted to specific temporal models like GRU, LSTM, and Transformers. This flexibility enables our method to achieve a better speed-accuracy trade-off. Specifically, we apply SPoS to partition the input frame sequence and provide a structured context for the subsequent temporal model. Notably, SPoS’s overall computational complexity is linear with respect to the video length. We next calculate group similarities to capture differences between frames, and a lightweight fully convolutional network is utilized to determine the event boundaries based on the grouped similarity maps. To remedy the ambiguities of boundary annotations, we adapt the Gaussian kernel to preprocess the ground-truth event boundaries. Our proposed method has been extensively evaluated on the challenging Kinetics-GEBD, TAPOS, and shot transition detection datasets, demonstrating its superiority over existing state-of-the-art methods.
Xin Gu 0003, Dexiang Hong, Libo Zhang 0001, Tiejian Luo, Longyin Wen, Heng Fan 0001
WACV5
2026 RepoCBench: A benchmark for C-oriented repository-level code generation with large language models and agents
Kaichun Yao, Libo Zhang 0001, Chen Zhao 0024
Expert Syst. Appl.3
2025 GSOT3D: Towards Generic 3D Single Object Tracking in the Wild
abstract
In this paper, we present a novel benchmark, GSOT3D, that aims at facilitating development of generic 3D single object tracking (SOT) in the wild. Specifically, GSOT3D offers 620 sequences with 123K frames, and covers a wide selection of 54 object categories. Each sequence is offered with multiple modalities, including the point cloud (PC), RGB image, and depth. This allows GSOT3D to support various 3D tracking tasks, such as single-modal 3D SOT on PC and multi-modal 3D SOT on RGB-PC or RGB-D, and thus greatly broadens research directions for 3D object tracking. To provide highquality per-frame 3D annotations, all sequences are labeled manually with multiple rounds of meticulous inspection and refinement. To our best knowledge, GSOT3D is the largest benchmark dedicated to various generic 3D object tracking tasks. To understand how existing 3D trackers perform and to provide comparisons for future research on GSOT3D, we assess eight representative point cloud-based tracking models. Our evaluation results exhibit that these models heavily degrade on GSOT3D, and more efforts are required for robust and generic 3D object tracking. Besides, to encourage future research, we present a simple yet effective generic 3D tracker, named PROT3D, that localizes the target object via a progressive spatial-temporal network and outperforms all current solutions by a large margin. By releasing GSOT3D, we expect to advance further 3D tracking in future research and applications. Our benchmark and model as well as the evaluation results will be publicly released at our webpage https://github.com/ailovejinx/GSOT3D.
Yifan Jiao, Junhua Ding 0001, Qing Yang 0003, Song Fu, Heng Fan 0001, Libo Zhang 0001
ICCV7
2025 Attention to Trajectory: Trajectory-Aware Open-Vocabulary Tracking
abstract
Open-Vocabulary Multi-Object Tracking (OV-MOT) aims to enable approaches to track objects without being limited to a predefined set of categories. Current OV-MOT methods typically rely primarily on instance-level detection and association, often overlooking trajectory information that is unique and essential for object tracking tasks. Utilizing trajectory information can enhance association stability and classification accuracy, especially in cases of occlusion and category ambiguity, thereby improving adaptability to novel classes. Thus motivated, in this paper we propose \textbf{TRACT}, an open-vocabulary tracker that leverages trajectory information to improve both object association and classification in OV-MOT. Specifically, we introduce a \textit{Trajectory Consistency Reinforcement} (\textbf{TCR}) strategy, that benefits tracking performance by improving target identity and category consistency. In addition, we present \textbf{TraCLIP}, a plug-and-play trajectory classification module. It integrates \textit{Trajectory Feature Aggregation} (\textbf{TFA}) and \textit{Trajectory Semantic Enrichment} (\textbf{TSE}) strategies to fully leverage trajectory information from visual and language perspectives for enhancing the classification results. Extensive experiments on OV-TAO show that our TRACT significantly improves tracking performance, highlighting trajectory information as a valuable asset for OV-MOT. Code will be released.
Yifan Jiao, Dan Meng 0001, Heng Fan 0001, Libo Zhang 0001
ICCV5
2025 Multi-Reward as Condition for Instruction-based Image Editing
abstract
High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable Diffusion, DALL-E) which are not trained for image editing. Accordingly, these datasets suffer from inaccurate instruction following, poor detail preserving, and generation artifacts. In this paper, we propose to address the training data quality issue with multi-perspective reward data instead of refining the ground-truth image quality. 1) we first design a quantitative metric system based on best-in-class LVLM (Large Vision Language Model), i.e., GPT-4o in our case, to evaluate the generation quality from 3 perspectives, namely, instruction following, detail preserving, and generation quality. For each perspective, we collected quantitative score in $0\sim 5$ and text descriptive feedback on the specific failure points in ground-truth edited images, resulting in a high-quality editing reward dataset, i.e., RewardEdit20K. 2) We further proposed a novel training framework to seamlessly integrate the metric output, regarded as multi-reward, into editing models to learn from the imperfect training triplets. During training, the reward scores and text descriptions are encoded as embeddings and fed into both the latent space and the U-Net of the editing models as auxiliary conditions. During inference, we set these additional conditions to the highest score with no text description for failure points, to aim at the best generation outcome. 3) We also build a challenging evaluation benchmark with real-world images/photos and diverse editing instructions, named as Real-Edit. Experiments indicate that our multi-reward conditioned model outperforms its no-reward counterpart on two popular editing pipelines, i.e., InsPix2Pix and SmartEdit. Code is released at https://github.com/bytedance/Multi-Reward-Editing.
Xin Gu 0003, Libo Zhang 0001, Longyin Wen, Tiejian Luo, Sijie Zhu
ICLR3
2025 Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
abstract
Transformer has attracted increasing interest in spatio-temporal video grounding, or STVG, owing to its end-to-end pipeline and promising result. Existing Transformer-based STVG approaches often leverage a set of object queries, which are initialized simply using zeros and then gradually learn target position information via iterative interactions with multimodal features, for spatial and temporal localization. Despite simplicity, these zero object queries, due to lacking target-specific cues, are hard to learn discriminative target information from interactions with multimodal features in complicated scenarios (e.g., with distractors or occlusion), resulting in degradation. Addressing this, we introduce a novel $\textbf{T}$arget-$\textbf{A}$ware Transformer for $\textbf{STVG}$ ($\textbf{TA-STVG}$), which seeks to adaptively generate object queries via exploring target-specific cues from the given video-text pair, for improving STVG. The key lies in two simple yet effective modules, comprising text-guided temporal sampling (TTS) and attribute-aware spatial activation (ASA), working in a cascade. The former focuses on selecting target-relevant temporal cues from a video utilizing holistic text information, while the latter aims at further exploiting the fine-grained visual attribute information of the object from previous target-aware temporal cues, which is applied for object query initialization. Compared to existing methods leveraging zero-initialized queries, object queries in our TA-STVG, directly generated from a given video-text pair, naturally carry target-specific cues, making them adaptive and better interact with multimodal features for learning more discriminative information to improve STVG. In our experiments on three benchmarks, including HCSTVG-v1/-v2 and VidSTG, TA-STVG achieves state-of-the-art performance and significantly outperforms the baseline, validating its efficacy. Moreover, TTS and ASA are designed for general purpose. When applied to existing methods such as TubeDETR and STCAT, we show substantial performance gains, verifying its generality. Code is released at https://github.com/HengLan/TA-STVG.
Xin Gu 0003, Yaojie Shen, Chenxi Luo, Tiejian Luo, Yan Huang 0002, Yuewei Lin, Heng Fan 0001, Libo Zhang 0001
ICLR8
2025 CGTrack: Cascade Gating Network with Hierarchical Feature Aggregation for UAV Tracking
abstract
Recent advancements in visual object tracking have markedly improved the capabilities of unmanned aerial vehicle (UAV) tracking, which is a critical component in real-world robotics applications. While the integration of hierarchical lightweight networks has become a prevalent strategy for enhancing efficiency in UAV tracking, it often results in a significant drop in network capacity, which further exacerbates challenges in UAV scenarios, such as frequent occlusions and extreme changes in viewing angles. To address these issues, we introduce a novel family of UAV trackers, termed CGTrack, which combines explicit and implicit techniques to expand network capacity within a coarse-to-fine framework. Specifically, we first introduce a Hierarchical Feature Cascade (HFC) module that leverages the spirit of feature reuse to increase network capacity by integrating the deep semantic cues with the rich spatial information, incurring minimal computational costs while enhancing feature representation. Based on this, we design a novel Lightweight Gated Center Head (LGCH) that utilizes gating mechanisms to decouple target-oriented coordinates from previously expanded features, which contain dense local discriminative information. Extensive experiments on three challenging UAV tracking benchmarks demonstrate that CGTrack achieves state-of-the-art performance while running fast. Code will be available at https://github.com/Nightwatch-Fox11/CGTrack.
Weihong Li 0002, Xiaoqiong Liu, Heng Fan 0001, Libo Zhang 0001
ICRA4
2025 LaMOT: Language-Guided Multi-Object Tracking
abstract
Vision-Language MOT is a critical tracking problem that has recently garnered increasing attention. It aims to track objects based on human language commands, displacing the traditional use of templates or pre-set information from training sets in conventional tracking tasks. However, a key challenge remains in understanding why language is used for tracking, hindering further development. In this paper, we introduce Language-Guided MOT, a unified task framework, and LaMOT, a corresponding large-scale benchmark, which encompasses diverse scenarios and language descriptions and comprises 1,660 sequences from 4 different datasets. The purpose of LaMOT is to unify various Vision-Language MOT tasks while providing a standardized evaluation platform. To ensure high-quality annotations, we manually assign appropriate descriptive texts to each target in every video and conduct careful inspection and correction. To our knowledge, LaMOt is the first benchmark dedicated to Language-Guided MOT. Additionally, we propose a simple yet effective tracker, termed LaMOTer. By establishing a unified task framework, providing challenging benchmarks, and offering insights for future algorithm design and evaluation, we expect to contribute to the advancement of research in Vision-Language MOT. We will release the data at https://github.com/Nathan-Li123/LaMOT.
Xiaoqiong Liu, Luke Liu, Heng Fan 0001, Libo Zhang 0001
ICRA5
2025 G3 CN: Gaussian Topology Refinement Gated Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
Graph Convolutional Networks (GCNs) have proven to be highly effective for skeleton-based action recognition, primarily due to their ability to leverage graph topology for feature aggregation, a key factor in extracting meaningful representations. However, despite their success, GCNs often struggle to effectively distinguish between ambiguous actions, revealing limitations in the representation of learned topological and spatial features. To address this challenge, we propose a novel approach, Gaussian Topology Refinement Gated Graph Convolution (G3CN), to address the challenge of distinguishing ambiguous actions in skeleton-based action recognition. G3CN incorporates a Gaussian filter to refine the skeleton topology graph, improving the representation of ambiguous actions. Additionally, Gated Recurrent Units (GRUs) are integrated into the GCN framework to enhance information propagation between skeleton points. Our method shows strong generalization across various GCN backbones. Extensive experiments on NTU RGB+D, NTU RGB+D 120, and NW-UCLA benchmarks demonstrate that G3CN effectively improves action recognition, particularly for ambiguous samples.
Haiqing Ren, Zhongkai Luo, Heng Fan 0001, Xiaohui Yuan 0001, Guanchen Wang, Libo Zhang 0001
IROS6
2025 Robust Ego-Exo Correspondence with Long-Term Memory
abstract
Establishing object-level correspondence between egocentric and exocentric views is essential for intelligent assistants to deliver precise and intuitive visual guidance. However, this task faces numerous challenges, including extreme viewpoint variations, occlusions, and the presence of small objects. Existing approaches usually borrow solutions from video object segmentation models, but still suffer from the aforementioned challenges. Recently, the Segment Anything Model 2 (SAM 2) has shown strong generalization capabilities and excellent performance in video object segmentation. Yet, when simply applied to the ego-exo correspondence (EEC) task, SAM 2 encounters severe difficulties due to ineffective ego-exo feature fusion and limited long-term memory capacity, especially for long videos. Addressing these problems, we propose a novel EEC framework based on SAM 2 with long-term memories by presenting a dual-memory architecture and an adaptive feature routing module inspired by Mixture-of-Experts (MoE). Compared to SAM 2, our approach features **(i)** a Memory-View MoE module which consists of a dual-branch routing mechanism to adaptively assign contribution weights to each expert feature along both channel and spatial dimensions, and **(ii)** a dual-memory bank system with a simple yet effective compression strategy to retain critical long-term information while eliminating redundancy. In the extensive experiments on the challenging EgoExo4D benchmark, our method, dubbed ***LM-EEC***, achieves new state-of-the-art results and significantly outperforms existing methods and the SAM 2 baseline, showcasing its strong generalization across diverse scenarios. Our code and model are available at https://github.com/juneyeeHu/LM-EEC.
Bing Fan, Xin Gu 0003, Haiqing Ren, Dongfang Liu, Heng Fan 0001, Libo Zhang 0001
NeurIPS7
2025 Compiler-R1: Towards Agentic Compiler Auto-tuning with Reinforcement Learning
abstract
Compiler auto-tuning optimizes pass sequences to improve performance metrics such as Intermediate Representation (IR) instruction count. Although recent advances leveraging Large Language Models (LLMs) have shown promise in automating compiler tuning, two significant challenges still remain: the absence of high-quality reasoning datasets for agents training, and limited effective interactions with the compilation environment. In this work, we introduce Compiler-R1, the first reinforcement learning (RL)-driven framework specifically augmenting LLM capabilities for compiler auto-tuning. Compiler-R1 features a curated, high-quality reasoning dataset and a novel two-stage end-to-end RL training pipeline, enabling efficient environment exploration and learning through an outcome-based reward. Extensive experiments across seven datasets demonstrate Compiler-R1 achieving an average 8.46\% IR instruction count reduction compared to opt -Oz, showcasing the strong potential of RL-trained LLMs for compiler optimization. Our code and datasets are publicly available at https://github.com/Panhaolin2001/Compiler-R1.
Haolin Pan, Kaichun Yao, Libo Zhang 0001, Mingjie Xing
NeurIPS6
2025 PlanarTrack: A high-quality and challenging benchmark for large-scale planar object tracking
Yifan Jiao, Xiaoqiong Liu, Xiaohui Yuan 0001, Heng Fan 0001, Libo Zhang 0001
Comput. Vis. Image Underst.6
2025 Learning to zoom: Exploiting mixed-scale contextual information for object detection
Boying Wang, Ruyi Ji, Libo Zhang 0001, Jing Liu 0001
Expert Syst. Appl.3
2025 High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
Libo Zhang 0001, Jiali Yao, Heng Fan 0001
Int. J. Comput. Vis.1
2025 AttMOT: Improving Multiple-Object Tracking by Introducing Auxiliary Pedestrian Attributes
abstract
Multiobject tracking (MOT) is a fundamental problem in computer vision with numerous applications, such as intelligent surveillance and automated driving. Despite the significant progress made in MOT, pedestrian attributes, such as gender, hairstyle, body shape, and clothing features, which contain rich and high-level information, have been less explored. To address this gap, we propose a simple, effective, and generic method to predict pedestrian attributes to support general reidentification (Re-ID) embedding. We first introduce attribute multi-object tracking (AttMOT), a large, highly enriched synthetic dataset for pedestrian tracking, containing over 80k frames and six million pedestrian identity switches (IDs) with different times, weather conditions, and scenarios. To the best of authors' knowledge, AttMOT is the first MOT dataset with semantic attributes. Subsequently, we explore different approaches to fuse Re-ID embedding and pedestrian attributes, including attention mechanisms, which we hope will stimulate the development of attribute-assisted MOT. The proposed method attribute-assisted method (AAM) demonstrates its effectiveness and generality on several representative pedestrian MOT benchmarks, including MOT17 and MOT20, through experiments on the AttMOT dataset. When applied to the state-of-the-art trackers, AAM achieves consistent improvements in multi-object tracking accuracy (MOTA), higher order tracking accuracy (HOTA), association accuracy (AssA), IDs, and IDF1 scores. For instance, on MOT17, the proposed method yields a +1.1 MOTA, +1.7 HOTA, and +1.8 IDF1 improvement when used with FairMOT. To further encourage related research, we release the data and code at https://github.com/HengLan/AttMOT.
Dan Meng 0001, Heng Fan 0001, Libo Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2025 CARL: Unsupervised Code-Based Adversarial Attacks for Programming Language Models via Reinforcement Learning
abstract
Code based adversarial attacks play a crucial role in revealing vulnerabilities of software system. Recently, pre-trained programming language models (PLMs) have demonstrated remarkable success in various significant software engineering tasks, progressively transforming the paradigm of software development. Despite their impressive capabilities, these powerful models are vulnerable to adversarial attacks. Therefore, it is necessary to carefully investigate the robustness and vulnerabilities of the PLMs by means of adversarial attacks. Adversarial attacks entail imperceptible input modifications that cause target models to make incorrect predictions. Existing approaches for attacking PLMs often employ either identifier renaming or the greedy algorithm, which may yield sub-optimal performance or lead to high inference times. In response to these limitations, we propose CARL, an unsupervised black-box attack model that leverages reinforcement learning to generate imperceptible adversarial examples. Specifically, CARL comprises a programming language encoder and a perturbation prediction layer. In order to achieve more effective and efficient attack, we cast the task as a sequence decision-making process, optimizing through policy gradient with a suite of reward functions. We conduct extensive experiments to validate the effectiveness of CARL on code summarization, code translation, and code refinement tasks, covering various programming languages and PLMs. The experimental results demonstrate that CARL surpasses state-of-the-art code attack models, achieving the highest attack success rate across multiple tasks and PLMs while maintaining high attack efficiency, imperceptibility, consistency, and fluency.
Kaichun Yao, Hao Wang 0093, Chuan Qin 0002, Hengshu Zhu, Libo Zhang 0001
ACM Trans. Softw. Eng. Methodol.6
2024 Context-Guided Spatio-Temporal Video Grounding
abstract
Spatio-temporal video grounding (or STVG) task aims at locating a spatio-temporal tube for a specific instance given a text query. Despite advancements, current methods easily suffer the distractors or heavy object appearance variations in videos due to insufficient object information from the text, leading to degradation. Addressing this, we propose a novel framework, context-guided STVG (CG-STVG), which mines discriminative instance context for object in videos and applies it as a supplementary guidance for target localization. The key of CG-STVG lies in two specially designed modules, including instance context generation (ICG), which focuses on discovering visual context information (in both appearance and motion) of the instance, and instance context refinement (ICR), which aims to improve the instance context from ICG by eliminating irrelevant or even harmful information from the context. During grounding, ICG, together with ICR, are deployed at each decoding stage of a transformer architecture for instance context learning. Particularly, instance context learned from one decoding stage is fed to the next stage, and leveraged as a guidance containing rich and discriminative object feature to enhance the target-awareness in decoding feature, which conversely benefits generating better new instance context to improve localization finally. Compared to existing methods, CG-STVG enjoys object information in text query and guidance from mined instance visual context for more accurate target localization. In experiments on HCSTVG-v1/-v2 and VidSTG, CG-STVG sets new state-of-the-arts in m_tIoU and m_vIoU on all of them, showing efficacy. Code is released at https://github.com/HengLan/CGSTVG.
Xin Gu 0003, Heng Fan 0001, Yan Huang 0002, Tiejian Luo, Libo Zhang 0001
CVPR5
2024 Kernel Adaptive Convolution for Scene Text Detection via Distance Map Prediction
abstract
Segmentation-based scene text detection algorithms that are accurate to the pixel level can satisfy the detection of arbitrary shape scene text and have received widespread attention. On the one hand, due to the complexity and di-versity of the scene text, the convolution with a fixed kernel size has some limitations in extracting the visual features of the scene text. On the other hand, most of the existing segmentation-based algorithms only segment the center of the text, losing information such as the edges and directions of the text, with limited detection accuracy. There are also some improved algorithms that use iterative cor-rections or introduce other multiple information to improve text detection accuracy but at the expense of efficiency. To address these issues, this paper proposes a simple and effective scene text detection method, the Kernel Adaptive Con-volution, which is designed with a Kernel Adaptive Con-volution Module for scene text detection via predicting the distance map. Specifically, first, we design an extensible kernel adaptive convolution module (KACM) to extract vi-sual features from multiple convolutions with different ker-nel sizes in an adaptive manner. Secondly, our method pre-dicts the text distance map under the supervision of a pri-ori information (including direction map, and foreground segmentation map) and completes the text detection from the predicted distance map. Experiments on four publicly available datasets prove the effectiveness of our algorithm, in which the accuracy and efficiency of both the Total- Text and TD500 outperform the state-of-the-art algorithm. The algorithm efficiency is improved while the accuracy is com-petitive on ArT and CTW1500.
Jinzhi Zheng, Heng Fan 0001, Libo Zhang 0001
CVPR3
2024 Beyond MOT: Semantic Multi-object Tracking
Hao Wang 0093, Jiali Yao, Shaohua Dong, Heng Fan 0001, Libo Zhang 0001
ECCV (35)8
2024 BPDO: Boundary Points Dynamic Optimization for Arbitrary Shape Scene Text Detection
abstract
Arbitrary shape scene text detection is of great importance in scene understanding tasks. Due to the complexity and diversity of text in natural scenes, existing scene text algorithms have limited accuracy for detecting arbitrary shape text. In this paper, we propose a novel arbitrary shape scene text detector through boundary points dynamic optimization(BPDO). The proposed model is designed with a text aware module (TAM) and a boundary point dynamic optimization module (DOM). Specifically, the model designs a text aware module based on segmentation to obtain boundary points describing the central region of the text by extracting a priori information about the text region. Then, based on the idea of deformable attention, it proposes a dynamic optimization model for boundary points, which gradually optimizes the exact position of the boundary points based on the information of the adjacent region of each boundary point. Experiments on CTW-1500, Total-Text, and MSRATD500 datasets show that the model proposed in this paper achieves a performance that is better than or comparable to the state-of-the-art algorithm, proving the effectiveness of the model.
Jinzhi Zheng, Libo Zhang 0001, Chen Zhao 0024
ICASSP2
2024 Text Region Multiple Information Perception Network for Scene Text Detection
abstract
Segmentation-based scene text detection algorithms can handle arbitrary shape scene texts and have strong robustness and adaptability, so it has attracted wide attention. Existing segmentation-based scene text detection algorithms usually only segment the pixels in the center region of the text, while ignoring other information of the text region, such as edge information, distance information, etc., thus limiting the detection accuracy of the algorithm for scene text. This paper proposes a plug-and-play module called the Region Multiple Information Perception Module (RMIPM) to enhance the detection performance of segmentation-based algorithms. Specifically, we design an improved module that can perceive various types of information about scene text regions, such as text foreground classification maps, distance maps, direction maps, etc. Experiments on MSRA-TD500 and TotalText datasets show that our method achieves comparable performance with current state-of-the-art algorithms.
Jinzhi Zheng, Libo Zhang 0001, Chen Zhao 0024
ICASSP2
2024 MaGIC: Multi-modality Guided Image Completion
abstract
Vanilla image completion approaches exhibit sensitivity to large missing regions, attributed to the limited availability of reference information for plausible generation. To mitigate this, existing methods incorporate the extra cue as guidance for image completion. Despite improvements, these approaches are often restricted to employing a *single modality* (e.g., *segmentation* or *sketch* maps), which lacks scalability in leveraging multi-modality for more plausible completion. In this paper, we propose a novel, simple yet effective method for **M**ulti-mod**a**l **G**uided **I**mage **C**ompletion, dubbed **MaGIC**, which not only supports a wide range of single modality as the guidance (e.g., *text*, *canny edge*, *sketch*, *segmentation*, *depth*, and *pose*), but also adapts to arbitrarily customized combinations of these modalities (i.e., *arbitrary multi-modality*) for image completion. For building MaGIC, we first introduce a modality-specific conditional U-Net (MCU-Net) that injects single-modal signal into a U-Net denoiser for single-modal guided image completion. Then, we devise a consistent modality blending (CMB) method to leverage modality signals encoded in multiple learned MCU-Nets through gradient guidance in latent space. Our CMB is *training-free*, thereby avoiding the cumbersome joint re-training of different modalities, which is the secret of MaGIC to achieve exceptional flexibility in accommodating new modalities for completion. Experiments show the superiority of MaGIC over state-of-the-art methods and its generalization to various completion tasks.
Hao Wang 0093, Tiejian Luo, Heng Fan 0001, Libo Zhang 0001
ICLR5
2024 VastTrack: Vast Category Visual Object Tracking
abstract
In this paper, we propose a novel benchmark, named VastTrack, aiming to facilitate the development of general visual tracking via encompassing abundant classes and videos. VastTrack consists of a few attractive properties: (1) Vast Object Category. In particular, it covers targets from 2,115 categories, significantly surpassing object classes of existing popular benchmarks (e.g., GOT-10k with 563 classes and LaSOT with 70 categories). Through providing such vast object classes, we expect to learn more general object tracking. (2) Larger scale. Compared with current benchmarks, VastTrack provides 50,610 videos with 4.2 million frames, which makes it to date the largest dataset in term of the number of videos, and hence could benefit training even more powerful visual trackers in the deep learning era. (3) Rich Annotation. Besides conventional bounding box annotations, VastTrack also provides linguistic descriptions with more than 50K sentences for the videos. Such rich annotations of VastTrack enable the development of both vision-only and vision-language tracking. In order to ensure precise annotation, each frame in the videos is manually labeled with multi-stage of careful inspections and refinements. To understand performance of existing trackers and to provide baselines for future comparison, we extensively evaluate 25 representative trackers. The results, not surprisingly, display significant drops compared to those on current datasets due to lack of abundant categories and videos from diverse scenarios for training, and more efforts are urgently required to improve general visual tracking. Our VastTrack, the toolkit, and evaluation results are publicly available at https://github.com/HengLan/VastTrack.
Junyuan Gao, Weihong Li 0002, Shaohua Dong, Heng Fan 0001, Libo Zhang 0001
NeurIPS8
2024 Local Compressed Video Stream Learning for Generic Event Boundary Detection
Libo Zhang 0001, Xin Gu 0003, Tiejian Luo, Heng Fan 0001
Int. J. Comput. Vis.1
2024 Correction: PIDray: A Large-Scale X-ray Benchmark for Real-World Prohibited Item Detection
Libo Zhang 0001, Lutao Jiang, Ruyi Ji, Heng Fan 0001
Int. J. Comput. Vis.1
2024 Robust Domain Adaptive Object Detection With Unified Multi-Granularity Alignment
abstract
Domain adaptive detection aims to improve the generalization of detectors on target domain. To reduce discrepancy in feature distributions between two domains, recent approaches achieve domain adaption through feature alignment in different granularities via adversarial learning. However, they neglect the relationship between multiple granularities and different features in alignment, degrading detection. Addressing this, we introduce a unified multi-granularity alignment (MGA)-based detection framework for domain-invariant feature learning. The key is to encode the dependencies across different granularities including pixel-, instance-, and category-levels simultaneously to align two domains. Specifically, based on pixel-level features, we first develop an omni-scale gated fusion (OSGF) module to aggregate discriminative representations of instances with scale-aware convolutions, leading to robust multi-scale detection. Besides, we introduce multi-granularity discriminators to identify where, either source or target domains, different granularities of samples come from. Note that, MGA not only leverages instance discriminability in different categories but also exploits category consistency between two domains for detection. Furthermore, we present an adaptive exponential moving average (AEMA) strategy that explores model assessments for model update to improve pseudo labels and alleviate local misalignment problem, boosting detection robustness. Extensive experiments on multiple domain adaption scenarios validate the superiority of MGA over other approaches on FCOS and Faster R-CNN detectors.
Libo Zhang 0001, Wenzhang Zhou, Heng Fan 0001, Tiejian Luo, Haibin Ling
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Text with Knowledge Graph Augmented Transformer for Video Captioning
abstract
Video captioning aims to describe the content of videos using natural language. Although significant progress has been made, there is still much room to improve the performance for real-world applications, mainly due to the long-tail words challenge. In this paper, we propose a text with knowledge graph augmented transformer (TextKG)for video captioning. Notably, TextKG is a two-stream transformer, formed by the external stream and internal stream. The external stream is designed to absorb additional knowledge, which models the interactions between the additional knowledge, e.g., pre-built knowledge graph, and the built-in information of videos, e.g., the salient object regions, speech transcripts, and video captions, to mitigate the long-tail words challenge. Meanwhile, the internal stream is designed to exploit the multi-modality information in videos (e.g., the appearance of video frames, speech transcripts, and video captions) to ensure the quality of caption results. In addition, the cross attention mechanism is also used in between the two streams for sharing information. In this way, the two streams can help each other for more accurate results. Extensive experiments conducted on four challenging video captioning datasets, i.e., YouCookII, ActivityNet Captions, MSR-VTT, and MSVD, demonstrate that the proposed method performs favorably against the state-of-the-art methods. Specifically, the proposed TextKG method out-performs the best published results by improving 18.7% absolute CIDEr scores on the YouCookII dataset.
Xin Gu 0003, Libo Zhang 0001, Tiejian Luo, Longyin Wen
CVPR4
2023 Two Birds, One Stone: A Unified Framework for Joint Learning of Image and Video Style Transfers
abstract
Current arbitrary style transfer models are limited to either image or video domains. In order to achieve satisfying image and video style transfers, two different models are inevitably required with separate training processes on image and video domains, respectively. In this paper, we show that this can be precluded by introducing UniST, a Unified Style Transfer framework for both images and videos. At the core of UniST is a domain interaction transformer (DIT ), which first explores context information within the specific domain and then interacts contextualized domain information for joint learning. In particular, DIT enables exploration of temporal information from videos for the image style transfer task and meanwhile allows rich appearance texture from images for video style transfer, thus leading to mutual benefits. Considering heavy computation of traditional multi-head self-attention, we present a simple yet effective axial multi-head self-attention (AMSA) for DIT , which improves computational efficiency while maintains style transfer performance. To verify the effectiveness of UniST, we conduct extensive experiments on both image and video style transfer tasks and show that UniST performs favorably against state-of-the-art approaches on both tasks. Code is available at https://github.com/NevSNev/UniST.
Bohai Gu, Heng Fan 0001, Libo Zhang 0001
ICCV3
2023 PlanarTrack: A Large-scale Challenging Benchmark for Planar Object Tracking
abstract
Planar object tracking is a critical computer vision problem and has drawn increasing interest owing to its key roles in robotics, augmented reality, etc. Despite rapid progress, its further development, especially in the deep learning era, is largely hindered due to the lack of large-scale challenging benchmarks. Addressing this, we introduce PlanarTrack, a large-scale challenging planar tracking benchmark. Specifically, PlanarTrack consists of 1,000 videos with more than 490K images. All these sequences are collected in complex unconstrained scenarios from the wild, which makes PlanarTrack, compared with existing benchmarks, more challenging but realistic for real-world applications. To ensure the high-quality annotation, each frame in PlanarTrack is manually labeled using four corners with multiple-round careful inspection and refinement. To our best knowledge, PlanarTrack, to date, is the largest and the most challenging dataset dedicated to planar object tracking. In order to analyze the proposed PlanarTrack, we evaluate 10 planar trackers and conduct comprehensive comparisons and in-depth analysis. Our results, not surprisingly, demonstrate that current top-performing planar trackers degenerate significantly on the challenging PlanarTrack and more efforts are needed to improve planar tracking in the future. In addition, we further derive a variant named PlanarTrackBBfor generic object tracking from our PlanarTrack. Our evaluation of 10 excellent generic trackers on PlanarTrackBBmanifests that, surprisingly, PlanarTrackBBis even more challenging than several popular generic tracking benchmarks and more attention should be paid to handle such planar objects, though they are rigid. All benchmarks and evaluations are released at https://hengfan2010.github.io/projects/PlanarTrack/.
Xiaoqiong Liu, Ziruo Yi, Libo Zhang 0001, Yan Huang 0002, Qing Yang 0003, Heng Fan 0001
ICCV6
2023 Accurate and Fast Compressed Video Captioning
abstract
Existing video captioning approaches typically require to first sample video frames from a decoded video and then conduct a subsequent process (e.g., feature extraction and/or captioning model learning). In this pipeline, manual frame sampling may ignore key information in videos and thus degrade performance. Additionally, redundant information in the sampled frames may result in low efficiency in the inference of video captioning. Addressing this, we study video captioning from a different perspective in compressed domain, which brings multi-fold advantages over the existing pipeline: 1) Compared to raw images from the decoded video, the compressed video, consisting of I-frames, motion vectors and residuals, is highly distinguishable, which allows us to leverage the entire video for learning without manual sampling through a specialized model design; 2) The captioning model is more efficient in inference as smaller and less redundant information is processed. We propose a simple yet effective end-to-end transformer in the compressed domain for video captioning that enables learning from the compressed video for captioning. We show that even with a simple design, our method can achieve state-of-the-art performance on different benchmarks while running almost 2× faster than existing approaches. Code is available at https://github.com/acherstyx/CoCap.
Yaojie Shen, Xin Gu 0003, Heng Fan 0001, Longyin Wen, Libo Zhang 0001
ICCV6
2023 Unsupervised Domain Adaptive Detection with Network Stability Analysis
abstract
Domain adaptive detection aims to improve the generality of a detector, learned from the labeled source domain, on the unlabeled target domain. In this work, drawing inspiration from the concept of stability from the control theory that a robust system requires to remain consistent both externally and internally regardless of disturbances, we propose a novel framework that achieves unsupervised domain adaptive detection through stability analysis. In specific, we treat discrepancies between images and regions from different domains as disturbances, and introduce a novel simple but effective Network Stability Analysis (NSA) framework that considers various disturbances for domain adaptation. Particularly, we explore three types of perturbations including heavy and light image-level disturbances and instance-level disturbance. For each type, NSA performs external consistency analysis on the outputs from raw and perturbed images and/or internal consistency analysis on their features, using teacher-student models. By integrating NSA into Faster R-CNN, we immediately achieve state-of-the-art results. In particular, we set a new record of 52.7% mAP on Cityscapes-to-FoggyCityscapes, showing the potential of NSA for domain adaptive detection. It is worth noticing, our NSA is designed for general purpose, and thus applicable to one-stage detection model (e.g., FCOS) besides the adopted one, as shown by experiments. Code is released at https://github.com/tiankongzhang/NSA.
Wenzhang Zhou, Heng Fan 0001, Tiejian Luo, Libo Zhang 0001
ICCV4
2023 CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition
Jinzhi Zheng, Ruyi Ji, Libo Zhang 0001, Chen Zhao 0024
ICONIP (6)3
2023 A Semantic and Structural Transformer for Code Summarization Generation
abstract
Currently most methods cast code summarization generation as a machine translation task. Wherein the Transformer framework is a representative among them. Thanks to the attention mechanism in the Transformer, such a framework has achieved the state-of-the-art performance. Unfortunately, the Transformer encounters a series of challenges when generalizing to code summarization generation domain. Compared with natural language, code sequence is characterized by more complex multi-modal features, and difficult to extract these features only by the original Transformer structure. To further improve the performance, we make full use of code semantic and structural information in abstract syntax tree to build a simple yet effective framework, which consists of self-attention and graph based module to integrate code semantic information and syntax tree structure information. Besides, to compensate for the insufficiency of Transformer in encoding local features, we present a well-designed local RNN module. Extensive experiments show that the proposed method performs on par with the state-of-the-art methods on two public benchmarks, including Java and Python datasets. The comprehensive ablation studies further demonstrate the effectiveness of architecture design choices. The source code is released at https://github.com/tzv314159/SSTrans.git.
Ruyi Ji, Zhenyu Tong, Tiejian Luo, Jing Liu 0001, Libo Zhang 0001
IJCNN5
2023 Siamese self-supervised learning for fine-grained visual classification
Ruyi Ji, Libo Zhang 0001
Comput. Vis. Image Underst.3
2023 Collaborative three-stream transformers for video captioning
Hao Wang 0093, Libo Zhang 0001, Heng Fan 0001, Tiejian Luo
Comput. Vis. Image Underst.2
2023 AnimalTrack: A Benchmark for Multi-Animal Tracking in the Wild
Libo Zhang 0001, Junyuan Gao, Heng Fan 0001
Int. J. Comput. Vis.1
2023 PIDray: A Large-Scale X-ray Benchmark for Real-World Prohibited Item Detection
Libo Zhang 0001, Lutao Jiang, Ruyi Ji, Heng Fan 0001
Int. J. Comput. Vis.1
2023 Coorp: Satisfying Low-Latency and High-Throughput Requirements of Wireless Network for Coordinated Robotic Learning
abstract
In coordinated robotic learning, multiple robots share the same wireless channel for communication, and bring together latency-sensitive (LS) network flows for control and bandwidth-hungry (BH) flows for distributed learning. Unfortunately, existing wireless network supporting systems cannot coordinate these two network flows to meet their own requirements: 1) prioritized contention systems (e.g., EDCA) prevent LS messages from timely acquiring the wireless channel because multiple wireless network interface cards (WNICs) with BH messages are contending for the channel 2) global planning systems (e.g., SchedWiFi) have to reserve a notable time window in the shared channel for each LS flow, suffering from severe bandwidth degradation (up to 42%). We present the coordinated preemption method to meet both requirements for LS flows and BH flows. Globally (among multiple robots), coordinated preemption eliminates unnecessary contention of BH flows by making them transmit in a round-robin manner, such that LS flows have the highest chance to win the contention against BH flows, without sacrificing overall bandwidth from the perspective of coordinated robotic learning applications. Locally (within the same robot), coordinated preemption in real time predicts the periodic transmission of LS flows from the upper application and conservatively limits packets of BH flows buffered in the WNIC only before LS packets arriving, reducing the bandwidth devoted to preemption. COORP, our implementation of coordinated preemption, reduced the violation of latency requirements from 53.9% (EDCA) to 8.8% (comparable to SchedWiFi). Regarding learning quality, COORP achieved a comparable (at times the same) learning reward with EDCA, which grew up to 76% faster than SchedWiFi.
Shengliang Deng, Xiuxian Guan, Zekai Sun, Shixiong Zhao, Tianxiang Shen, Xusheng Chen, Tianyang Duan, Jia Pan 0001, Libo Zhang 0001, Heming Cui
IEEE Internet Things J.11
2023 Dual Transformer With Multi-Grained Assembly for Fine-Grained Visual Classification
abstract
Fine-grained visual classification requires distinguishing sub-categories within the same super-category, which suffers from small inter-class and large intra-class variances. This paper aims to improve the FGVC task towards better performance, for which we deliver a novel dual Transformer framework (coined Dual-TR) with multi-grained assembly. The Dual-TR is well-designed to encode fine-grained objects by two parallel hierarchies, which is amenable to capturing the subtle yet discriminative cues via the self-attention mechanism in ViT. Specifically, we perform orthogonal multi-grained assembly within the Transformer structure for a more robust representation, i.e., intra-layer and inter-layer assembly. The former aims to explore the informative feature in various self-attention heads within the Transformer layer. The latter pays attention to the token assembly across Transformer layers. Meanwhile, we introduce the constraint of center loss to pull intra-class samples’ compactness and push that of inter-class samples. Extensive experiments show that Dual-TR performs on par with the state-of-the-art methods on four public benchmarks, including CUB-200-2011, NABirds, iNaturalist2017, and Stanford Dogs. The comprehensive ablation studies further demonstrate the effectiveness of architectural design choices.
Ruyi Ji, Libo Zhang 0001, Jing Liu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Bridging Multi-Scale Context-Aware Representation for Object Detection
abstract
Feature Pyramid Network (FPN) exploits multi-scale fusion representation to deal with scale variances in object detection. However, it ignores the context information gap across different levels. In this paper, we develop a plug-and-play detector, the multi-scale context-aware feature pyramid network to unleash the power of feature pyramid representation. Based on the dilated feature map at the highest level of the backbone, we propose the cross-scale context aggregation block to make full use of context information in the feature pyramid. Moreover, we extract discriminative features among different levels by the adaptive context aggregation block for robust object detection. Comprehensive experiments on MS-COCO demonstrate the effectiveness and efficiency of the proposed network, where about 1.0~3.0 AP improvements are achieved compared with existing FPN-based methods. In addition, we also conduct extensive experiments on pixel-level prediction tasks, i.e., instance segmentation, semantic segmentation, and panoptic segmentation, which further verify the effectiveness of the proposed method.
Boying Wang, Ruyi Ji, Libo Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 End-to-End Compressed Video Representation Learning for Generic Event Boundary Detection
abstract
Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be decoded before feeding into the network, which demands considerable computational power and storage space. To that end, we propose a new end-to-end compressed video representation learning for event boundary detection that leverages the rich information in the compressed domain, i.e., RGB, motion vectors, residuals, and the internal group of pictures (GOP) structure, withoutfully decoding the video. Specifically, we first use the Con-vNets to extract features of the I-frames in the Gaps. After that, a light-weight spatial-channel compressed encoder is designed to compute the feature representations of the P-frames based on the motion vectors, residuals and representations of their dependent I-frames. A temporal contrastive module is proposed to determine the event boundaries of video sequences. To remedy the ambiguities of annotations and speed up the training process, we use the Gaussian kernel to preprocess the ground-truth event boundaries. Extensive experiments conducted on the Kinetics-GEBD dataset demonstrate that the proposed method achieves comparable results to the state-of-the-art methods with 4.5 x faster running speed.
Longyin Wen, Dexiang Hong, Tiejian Luo, Libo Zhang 0001
CVPR6
2022 LD-ConGR: A Large RGB-D Video Dataset for Long-Distance Continuous Gesture Recognition
abstract
Gesture recognition plays an important role in natural human-computer interaction and sign language recognition. Existing research on gesture recognition is limited to close-range interaction such as vehicle gesture control and face-to-face communication. To apply gesture recognition to long-distance interactive scenes such as meetings and smart homes, a large RGB-D video dataset LD-ConGR is established in this paper. LD-ConGR is distinguished from existing gesture datasets by its long-distance gesture collection, fine-grained annotations, and high video qual-ity. Specifically, 1) the farthest gesture provided by the LD-ConGR is captured 4m away from the camera while existing gesture datasets collect gestures within 1m from the camera; 2) besides the gesture category, the temporal segmentation of gestures and hand location are also anno-tated in LD-ConGR; 3) videos are captured at high reso-lution (1280 x 720 for color streams and 640 x 576 for depth streams) and high frame rate (30 fps). On top of the LD-ConGR, a series of experimental and studies are conducted, and the proposed gesture region estimation and key frame sampling strategies are demonstrated to be effective in dealing with long-distance gesture recognition and the uncertainty of gesture duration. The dataset and experimen-tal results presented in this paper are expected to boost the research of long-distance gesture recognition. The dataset is available at https://github.com/Diananini/LD-ConGR-CVPR2022.
Libo Zhang 0001
CVPR2
2022 Multi-Granularity Alignment Domain Adaptation for Object Detection
abstract
Domain adaptive object detection is challenging due to distinctive data distribution between source domain and target domain. In this paper, we propose a unified multi-granularity alignment based object detection framework towards domain-invariant feature learning. To this end, we encode the dependencies across different granularity perspectives including pixel-, instance-, and category-levels simultaneously to align two domains. Based on pixel-level feature maps from the backbone network, we first develop the omniscale gated fusion module to aggregate discriminative representations of instances by scale-aware convolutions, leading to robust multi-scale object detection. Meanwhile, the multi-granularity discriminators are proposed to identify which domain different granularities of samples (i.e., pixels, instances, and categories) come from. Notably, we leverage not only the instance discriminability in different categories but also the category consistency between two domains. Extensive experiments are carried out on multiple domain adaptation scenarios, demonstrating the effectiveness of our framework over state-of-the-art algorithms on top of anchor-free FCOS and anchor-based Faster R-CNN detectors with different backbones.
Wenzhang Zhou, Dawei Du, Libo Zhang 0001, Tiejian Luo
CVPR3
2022 AutoTransition: Learning to Recommend Video Transition Effects
Yaojie Shen, Libo Zhang 0001, Xiaojie Jin 0004
ECCV (38)2
2022 Unbiased Multi-modality Guidance for Image Inpainting
Dawei Du, Libo Zhang 0001, Tiejian Luo
ECCV (16)3
2022 High-Fidelity Image Inpainting with GAN Inversion
Libo Zhang 0001, Heng Fan 0001, Tiejian Luo
ECCV (16)2
2022 ROG: A High Performance and Robust Distributed Training System for Robotic IoT
abstract
Critical robotic tasks such as rescue and disaster response are more prevalently leveraging ML (Machine Learning) models deployed on a team of wireless robots, on which data parallel (DP) training over Internet of Things of these robots (robotic IoT) can harness the distributed hardware resources to adapt their models to changing environments as soon as possible. Unfortunately, due to the need for DP synchronization across all robots, the instability in wireless networks (i.e., fluctuating bandwidth due to occlusion and varying communication distance) often leads to severe stall of robots, which affects the training accuracy within a tight time budget and wastes energy stalling. Existing methods to cope with the instability of datacenter networks are incapable of handling such straggler effect. That is because they are conducting model-granulated transmission scheduling, which is much more coarse-grained than the granularity of transient network instability in real-world robotic IoT networks, making a previously reached schedule mismatch with the varying bandwidth during transmission. We present ROG, the first ROw-Granulated distributed training system optimized for ML training over unstable wireless networks. ROG confines the granularity of transmission and synchronization to each row of a layer’s parameters and schedules the transmission of each row adaptively to the fluctuating bandwidth. In this way the ML training process can update partial and the most important gradients of a stale robot to avoid triggering stalls, while provably guaranteeing convergence. The evaluation shows that, given the same training time, ROG achieved about 4.9%~6.5% training accuracy gain compared with the baselines and saved 20.4%~50.7% of the energy to achieve the same training accuracy.
Xiuxian Guan, Zekai Sun, Shengliang Deng, Xusheng Chen, Shixiong Zhao, Zongyuan Zhang, Tianyang Duan, Chenshu Wu, Yong Cui 0001, Libo Zhang 0001, Rui Wang 0007, Heming Cui
MICRO11
2021 Rethinking Object Detection in Retail Stores
abstract
The conventional standard for object detection uses a bounding box to represent each individual object instance. However, it is not practical in the industry-relevant applications in the context of warehouses due to severe occlusions among groups of instances of the same categories. In this paper, we propose a new task, i.e., simultaneously object localization and counting, abbreviated as Locount, which requires algorithms to localize groups of objects of interest with the number of instances. However, there does not exist a dataset or benchmark designed for such a task. To this end, we collect a large-scale object localization and counting dataset with rich annotations in retail stores, which consists of 50,394 images with more than 1.9 million object instances in 140 categories. Together with this dataset, we provide a new evaluation protocol and divide the training and testing subsets to fairly evaluate the performance of algorithms for Locount, developing a new benchmark for the Locount task. Moreover, we present a cascaded localization and counting network as a strong baseline, which gradually classifies and regresses the bounding boxes of objects with the predicted numbers of instances enclosed in the bounding boxes, trained in an end-to-end manner. Extensive experiments are conducted on the proposed dataset to demonstrate its significance and the analysis is provided to indicate future directions. Dataset is available at https://isrc.iscas.ac.cn/gitlab/research/locount-dataset.
Yuanqiang Cai, Longyin Wen, Libo Zhang 0001, Dawei Du, Weiqiang Wang 0001
AAAI3
2021 Towards Real-world X-ray Security Inspection: A High-Quality Benchmark And Lateral Inhibition Module For Prohibited Items Detection
abstract
Prohibited items detection in X-ray images often plays an important role in protecting public safety, which often deals with color-monotonous and luster-insufficient objects, resulting in unsatisfactory performance. Till now, there have been rare studies touching this topic due to the lack of specialized high-quality datasets. In this work, we first present a High-quality X-ray (HiXray) security inspection image dataset, which contains 102,928 common prohibited items of 8 categories. It is the largest dataset of high quality for prohibited items detection, gathered from the real-world airport security inspection and annotated by professional security inspectors. Besides, for accurate prohibited item detection, we further propose the Lateral Inhibition Module (LIM) inspired by the fact that humans recognize these items by ignoring irrelevant information and focusing on identifiable characteristics, especially when objects are overlapped with each other. Specifically, LIM, the elaborately designed flexible additional module, suppresses the noisy information flowing maximumly by the Bidirectional Propagation (BP) module and activates the most identifiable charismatic, boundary, from four directions by Boundary Activation (BA) module. We evaluate our method extensively on HiXray and OPIXray and the results demonstrate that it outperforms SOTA detection methods.1
Renshuai Tao, Yanlu Wei, Xiangjian Jiang, Hainan Li, Haotong Qin, Jiakai Wang, Yuqing Ma, Libo Zhang 0001, Xianglong Liu 0001
ICCV8
2021 Towards Real-World Prohibited Item Detection: A Large-Scale X-ray Benchmark
abstract
Automatic security inspection using computer vision technology is a challenging task in real-world scenarios due to various factors, including intra-class variance, class imbalance, and occlusion. Most of the previous methods rarely solve the cases that the prohibited items are deliberately hidden in messy objects due to the lack of large-scale datasets, restricted their applications in real-world scenarios. Towards real-world prohibited item detection, we collect a large-scale dataset, named as PIDray, which covers various cases in real-world scenarios for prohibited item detection, especially for deliberately hidden items. With an intensive amount of effort, our dataset contains 12 categories of prohibited items in 47, 677 X-ray images with high-quality annotated segmentation masks and bounding boxes. To the best of our knowledge, it is the largest prohibited items detection dataset to date. Meanwhile, we design the selective dense attention network (SDANet) to construct a strong baseline, which consists of the dense attention module and the dependency refinement module. The dense attention module formed by the spatial and channel-wise dense attentions, is designed to learn the discriminative features to boost the performance. The dependency refinement module is used to exploit the dependencies of multi-scale features. Extensive experiments conducted on the collected PIDray dataset demonstrate that the proposed method performs favorably against the state-of-the-art methods, especially for detecting the deliberately hidden items.
Boying Wang, Libo Zhang 0001, Longyin Wen, Xianglong Liu 0001
ICCV2
2021 Non-deterministic and emotional chatting machine: learning emotional conversation generation using conditional variational autoencoders
Kaichun Yao, Libo Zhang 0001, Tiejian Luo, Dawei Du
Neural Comput. Appl.2
2021 Scale-Residual Learning Network for Scene Text Detection
abstract
Detecting incidentally captured text in the wild remains an open problem due to challenging factors including unconstrained scenarios and large scale variation. In this paper, we establish a large-scale scene text detection dataset (LS-Text), containing 36, 000 images and 270, 783 text instances with various scales and complex scenarios, to promote the research of text detection. We propose a Scale-residual Learning Network (SLN) to deal with the scale variation problem in a progressive optimization manner. Specifically, we integrate both learnable feature concatenation and feature up-sampling operator. It can effectively eliminate the residuals between the outputs of SLN and ground-truth text instances by processing both the Feature Fusion Residuals (FFR) and the Scale Transformation Residuals (STR), simultaneously. By stacking multi-scale feature maps in a deep-to-shallow manner, SLN continuously optimizes feature representation by accumulating strong semantic information and rich texture details in a scale-residual learning way. Extensive experimental results on five challenging datasets demonstrate the state-of-the-art performance of the proposed SLN model, and the challenging aspects related to real-world scenarios of the proposed LS-Text dataset. Both the source code of SLN and the LS-Text dataset are available athttps://github.com/SLN-Text-Detection.
Yuanqiang Cai, Chang Liu 0047, Peirui Cheng, Dawei Du, Libo Zhang 0001, Weiqiang Wang 0001, Qixiang Ye
IEEE Trans. Circuits Syst. Video Technol.5
2021 SiamCAN: Real-Time Visual Tracking Based on Siamese Center-Aware Network
abstract
In this article, we present a novel Siamese center-aware network (SiamCAN) for visual tracking, which consists of the Siamese feature extraction subnetwork, followed by the classification, regression, and localization branches in parallel. The classification branch is used to distinguish the target from background, and the regression branch is introduced to regress the bounding box of the target. To reduce the impact of manually designed anchor boxes to adapt to different target motion patterns, we design the localization branch to localize the target center directly to assist the regression branch generating accurate results. Meanwhile, we introduce the global context module into the localization branch to capture long-range dependencies for more robustness to large displacements of the target. A multi-scale learnable attention module is used to guide these three branches to exploit discriminative features for better performance. Extensive experiments on 9 challenging benchmarks, namely VOT2016, VOT2018, VOT2019, OTB100, LTB35, LaSOT, TC128, UAV123 and VisDrone-SOT2019 demonstrate that SiamCAN achieves leading accuracy with high efficiency. Our source code is available at https://isrc.iscas.ac.cn/gitlab/research/siamcan.
Wenzhang Zhou, Longyin Wen, Libo Zhang 0001, Dawei Du, Tiejian Luo
IEEE Trans. Image Process.3
2021 Graph Regularized Flow Attention Network for Video Animal Counting From Drones
abstract
In this paper, we propose a large-scale video based animal counting dataset collected by drones (AnimalDrone) for agriculture and wildlife protection. The dataset consists of two subsets, i.e., PartA captured on site by drones and PartB collected from the Internet, with rich annotations of more than 4 million objects in 53, 644 frames and corresponding attributes in terms of density, altitude and view. Moreover, we develop a new graph regularized flow attention network (GFAN) to perform density map estimation in dense crowds of video clips with arbitrary crowd density, perspective, and flight altitude. Specifically, our GFAN method leverages optical flow to warp the multi-scale feature maps in sequential frames to exploit the temporal relations, and then combines the enhanced features to predict the density maps. Moreover, we introduce the multi-granularity loss function including pixel-wise density loss and region-wise count loss to enforce the network to concentrate on discriminative features for different scales of objects. Meanwhile, the graph regularizer is imposed on the density maps of multiple consecutive frames to maintain temporal coherency. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method, compared with several state-of-the-art counting algorithms. The AnimalDrone dataset is available at https://github.com/VisDrone/AnimalDrone.
Pengfei Zhu 0001, Dawei Du, Libo Zhang 0001, Qinghua Hu
IEEE Trans. Image Process.5
2021 FastUDP: a highly scalable user-level UDP framework in multi-core systems for fast packet I/O
Heng Zhang 0005, Libo Zhang 0001
J. Supercomput.3
2021 Iterative Knowledge Distillation for Automatic Check-Out
abstract
Automatic Check-Out (ACO) provides an object detection based mechanism for retailers to process the purchases of customers automatically. However, it suffers a lot from the domain shift problem because of different data distribution between the single item in training exemplar images and mixed items in testing checkout images. In this paper, we propose a new iterative knowledge distillation method to solve the domain adaptation problem for this task. First, we develop a new augmentation data strategy to generate synthesized checkout images. It can extract segmented items from the training images by the coarse-to-fine strategy and filter items with unrealistic poses by pose pruning. Second, we propose a dual pyramid scale network (DPSNet) to exploit the multi-scale feature representation in joint detection and counting views. Third, the iterative knowledge distillation training strategy is developed to make full use of both image-level and instance-level samples to narrow the semantic gap between source domain and target domain. Extensive experiments on the large-scale Retail Product Checkout (RPC) dataset show the proposed DPSNet can achieve state-of-the-art performance compared with existing methods. The source codes can be found athttps://isrc.iscas.ac.cn/gitlab/research/dpsnet.
Libo Zhang 0001, Dawei Du, Tiejian Luo
IEEE Trans. Multim.1
2021 Multi-peak Graph-based Multi-instance Learning for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD), aiming to detect objects with only image-level annotations, has become one of the research hotspots over the past few years. Recently, much effort has been devoted to WSOD for the simple yet effective architecture and remarkable improvements have been achieved. Existing approaches using multiple-instance learning usually pay more attention to the proposals individually, ignoring relation information between proposals. Besides, to obtain pseudo-ground-truth boxes for WSOD, MIL-based methods tend to select the region with the highest confidence score and regard those with small overlap as background category, which leads to mislabeled instances. As a result, these methods suffer from mislabeling instances and lacking relations between proposals, degrading the performance of WSOD. To tackle these issues, this article introduces a multi-peak graph-based model for WSOD. Specifically, we use the instance graph to model the relations between proposals, which reinforces multiple-instance learning process. In addition, a multi-peak discovery strategy is designed to avert mislabeling instances. The proposed model is trained by stochastic gradients decent optimizer using back-propagation in an end-to-end manner. Extensive quantitative and qualitative evaluations on two publicly challenging benchmarks, PASCAL VOC 2007 and PASCAL VOC 2012, demonstrate the superiority and effectiveness of the proposed approach.
Ruyi Ji, Ze-Yu Liu 0011, Libo Zhang 0001, Jianwei Liu 0006, Chen Zhao 0024
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization
abstract
Fine-grained visual categorization (FGVC) is an important but challenging task due to high intra-class variances and low inter-class variances caused by deformation, occlusion, illumination, etc. An attention convolutional binary neural tree architecture is presented to address those problems for weakly supervised FGVC. Specifically, we incorporate convolutional operations along edges of the tree structure, and use the routing functions in each node to determine the root-to-leaf computational paths within the tree. The final decision is computed as the summation of the predictions from leaf nodes. The deep convolutional operations learn to capture the representations of objects, and the tree structure characterizes the coarse-to-fine hierarchical feature learning process. In addition, we use the attention transformer module to enforce the network to capture discriminative features. The negative log-likelihood loss is used to train the entire network in an end-to-end fashion by SGD with back-propagation. Several experiments on the CUB-200-2011, Stanford Cars and Aircraft datasets demonstrate that the proposed method performs favorably against the state-of-the-arts.
Ruyi Ji, Longyin Wen, Libo Zhang 0001, Dawei Du, Chen Zhao 0024, Xianglong Liu 0001, Feiyue Huang
CVPR3
2020 Learning Semantic Neural Tree for Human Parsing
Ruyi Ji, Dawei Du, Libo Zhang 0001, Longyin Wen, Chen Zhao 0024, Feiyue Huang, Siwei Lyu
ECCV (13)3
2020 Spatial Attention Pyramid Network for Unsupervised Domain Adaptation
Dawei Du, Libo Zhang 0001, Longyin Wen, Tiejian Luo, Pengfei Zhu 0001
ECCV (13)3
2020 Guided Attention Network for Object Detection and Counting on Drones
abstract
Object detection and counting are related but challenging problems, especially for drone based scenes with small objects and cluttered background. In this paper, we propose a new Guided Attention network (GAnet) to deal with both object detection and counting tasks based on the feature pyramid. Different from the previous methods relying on unsupervised attention modules, we fuse different scales of feature maps by using the proposed weakly-supervised Background Attention (BA) between the background and objects for more semantic feature representation. Then, the Foreground Attention (FA) module is developed to consider both global and local appearance of the object to facilitate accurate localization. Moreover, the new data argumentation strategy is designed to train a robust model in the drone based scenes with various illumination conditions. Extensive experiments on three challenging benchmarks (i.e., UAVDT, CARPK and PUCPR+) show the state-of-the-art detection and counting performance of the proposed method compared with existing methods. Code can be found at https://isrc.iscas.ac.cn/gitlab/research/ganet.
Yuanqiang Cai, Dawei Du, Libo Zhang 0001, Longyin Wen, Weiqiang Wang 0001, Siwei Lyu
ACM Multimedia3
2020 Occluded Prohibited Items Detection: An X-ray Security Inspection Benchmark and De-occlusion Attention Module
abstract
Security inspection often deals with a piece of baggage or suitcase where objects are heavily overlapped with each other, resulting in an unsatisfactory performance for prohibited items detection in X-ray images. In the literature, there have been rare studies and datasets touching this important topic. In this work, we contribute the first high-quality object detection dataset for security inspection, named Occluded Prohibited Items X-ray (OPIXray) image benchmark. OPIXray focused on the widely-occurred prohibited item "cutter", annotated manually by professional inspectors from the international airport. The test set is further divided into three occlusion levels to better understand the performance of detectors. Furthermore, to deal with the occlusion in X-ray images detection, we propose the De-occlusion Attention Module (DOAM), a plug-and-play module that can be easily inserted into and thus promote most popular detectors. Despite the heavy occlusion in X-ray imaging, shape appearance of objects can be preserved well, and meanwhile different materials visually appear with different colors and textures. Motivated by these observations, our DOAM simultaneously leverages the different appearance information of the prohibited item to generate the attention map, which helps refine feature maps for the general detectors. We comprehensively evaluate our module on the OPIXray dataset, and demonstrate that our module can consistently improve the performance of the state-of-the-art detection methods such as SSD, FCOS, etc, and significantly outperforms several widely-used attention mechanisms. In particular, the advantages of DOAM are more significant in the scenarios with higher levels of occlusion, which demonstrates its potential application in real-world inspections. The OPIXray benchmark and our model are released at https://github.com/OPIXray-author/OPIXray.
Yanlu Wei, Renshuai Tao, Zhangjie Wu, Yuqing Ma, Libo Zhang 0001, Xianglong Liu 0001
ACM Multimedia5
2020 Towards interpretable and robust hand detection via pixel-wise prediction
Libo Zhang 0001, Tiejian Luo, Lili Tao
Pattern Recognit.2
2020 Dual Encoding for Abstractive Text Summarization
abstract
Recurrent neural network-based sequence-to-sequence attentional models have proven effective in abstractive text summarization. In this paper, we model abstractive text summarization using a dual encoding model. Different from the previous works only using a single encoder, the proposed method employs a dual encoder including the primary and the secondary encoders. Specifically, the primary encoder conducts coarse encoding in a regular way, while the secondary encoder models the importance of words and generates more fine encoding based on the input raw text and the previously generated output text summarization. The two level encodings are combined and fed into the decoder to generate more diverse summary that can decrease repetition phenomenon for long sequence generation. The experimental results on two challenging datasets (i.e., CNN/DailyMail and DUC 2004) demonstrate that our dual encoding model performs against existing methods.
Kaichun Yao, Libo Zhang 0001, Dawei Du, Tiejian Luo, Lili Tao
IEEE Trans. Cybern.2
2019 Scale Invariant Fully Convolutional Network: Detecting Hands Efficiently
abstract
Existing hand detection methods usually follow the pipeline of multiple stages with high computation cost, i.e., feature extraction, region proposal, bounding box regression, and additional layers for rotated region detection. In this paper, we propose a new Scale Invariant Fully Convolutional Network (SIFCN) trained in an end-to-end fashion to detect hands efficiently. Specifically, we merge the feature maps from high to low layers in an iterative way, which handles different scales of hands better with less time overhead comparing to concatenating them simply. Moreover, we develop the Complementary Weighted Fusion (CWF) block to make full use of the distinctive features among multiple layers to achieve scale invariance. To deal with rotated hand detection, we present the rotation map to get rid of complex rotation and derotation layers. Besides, we design the multi-scale loss scheme to accelerate the training process significantly by adding supervision to the intermediate layers of the network. Compared with the state-of-the-art methods, our algorithm shows comparable accuracy and runs a 4.23 times faster speed on the VIVA dataset and achieves better average precision on Oxford hand detection dataset at a speed of 62.5 fps.
Dawei Du, Libo Zhang 0001, Tiejian Luo, Feiyue Huang, Siwei Lyu
AAAI3
2019 Data Priming Network for Automatic Check-Out
abstract
Automatic Check-Out (ACO) receives increased interests in recent years. An important component of the ACO system is the visual item counting, which recognizes the categories and counts of the items chosen by the customers. However, the training of such a system is challenged by the domain adaptation problem, in which the training data are images from isolated items while the testing images are for collections of items. Existing methods solve this problem with data augmentation using synthesized images, but the image synthesis leads to unreal images that affect the training process. In this paper, we propose a new data priming method to solve the domain adaptation problem. Specifically, we first use pre-augmentation data priming, in which we remove distracting background from the training images using the coarse-to-fine strategy and select images with realistic view angles by the pose pruning method. In the post-augmentation step, we train a data priming network using detection and counting collaborative learning, and select more reliable images from testing data to fine-tune the final visual item tallying network. Experiments on the large scale Retail Product Checkout (RPC) dataset demonstrate the superiority of the proposed method, i.e., we achieve 80.51% checkout accuracy compared with 56.68% of the baseline methods. The source codes can be found in https://isrc.iscas.ac.cn/gitlab/research/acm-mm-2019-ACO.
Dawei Du, Libo Zhang 0001, Tiejian Luo, Qi Tian 0001, Longyin Wen, Siwei Lyu
ACM Multimedia3
2018 Learning to Communicate via Supervised Attentional Message Processing
abstract
Many tasks in AI require the collaboration of multiple agents. Generally, these agents cooperate with each other by message-passing communication. However, agents may suffer from being overwhelmed by massive received messages and have difficulties in obtaining useful information. To this end, we use an attention-based message processing (AMP) method to model agents' interactions by considering the relevance of each received message. To improve the efficiency of learning correct interactions, a supervised variant SAMP is then proposed to directly optimize the attentional weights in AMP with a target auxiliary interaction matrix from the environment. The empirical results demonstrate our proposal outperforms other competing multi-agent methods in "predator-prey-toxin" domain, and prove the superiority of SAMP in correctly guiding the optimization of attentional weights in AMP.
Zhaoqing Peng, Libo Zhang 0001, Tiejian Luo
CASA2
2018 Teaching Machines to Ask Questions
abstract
We propose a novel neural network model that aims to generate diverse and human-like natural language questions. Our model not only directly captures the variability in possible questions by using a latent variable, but also generates certain types of questions by introducing an additional observed variable. We deploy our model in the generative adversarial network (GAN) framework and modify the discriminator which not only allows evaluating the question authenticity, but predicts the question type. Our model is trained and evaluated on a question-answering dataset SQuAD, and the experimental results shown the proposed model is able to generate diverse and readable questions with the specific attribute.
Kaichun Yao, Libo Zhang 0001, Tiejian Luo, Lili Tao
IJCAI2
2018 Multi-agent Communication with Attentional and Recurrent Message Integration
abstract
Effective communication is significant for solving cooperative tasks in multi-agent domain. Agents coordinate their behaviors by appropriately modeling the communication signals or messages sent from others. To this end, agents are required to filter noise and obtain useful information from received messages, and learn to adapt to the dynamics of messages number. In this paper, we propose an attentional and recurrent message integration method (ARMI) that handles the dynamics by recurrently decoding messages, and performs attentional integration based on the relevance of each message. We evaluate our proposal on a new “predator-prey-toxin” environment where the number of agents changes, and the results outperform other competing multi-agent methods. Further investigations are also done to prove the superiority of ARMI in collaborating agents' behaviors for complex tasks and establishing interpretable communication protocol.
Zhaoqing Peng, Libo Zhang 0001, Tiejian Luo
ISCC2
2018 Deep reinforcement learning for extractive document summarization
Kaichun Yao, Libo Zhang 0001, Tiejian Luo
Neurocomputing2
2017 EpCom: A parallel community detection approach for epidemic diffusion over social networks
abstract
Detecting community structure in epidemics networks is crucial for the assessment of epidemic dynamics and effective control of disease spread by targeting at the individuals bridging communities. Common community detection models (e.g., cut-criteria and modularity-criteria based model) are efficient in optimal quality of network partitions. However, most of the approaches fail to consider the dynamic infected possibility in person-to-person interactions. In addition, they present high computational complexity, which was limited by the scale of networks and the performance of hardware platform. In this paper, we propose a Jaccard distance based community detection model by considering both the quality of network partitions and the dynamics of infected interacts (i.e., edges) between two individuals in epidemic diffusion. Then, we design a novel parallel approach based on the high parallism of GPU, called EpCom, for boosting the performance and scalability of parallel community detection over large-scale epidemic networks. From the evaluation results, the proposed GPU-based implementation EpCom exhibits great performance and achieves maximum 604 million TEPS (traversed edges per second), which corresponds to up to 54.2 times and 15.6 times than CPU-based NCut and Louvain approaches separately.
Heng Zhang 0005, Libo Zhang 0001, Da Cheng, Chen Zhao 0024
BIBM2
2017 Accelerating Core Decomposition in Large Temporal Networks Using GPUs
Heng Zhang 0005, Haibo Hou, Libo Zhang 0001
ICONIP (1)3
2017 Learning Path Generation Method Based on Migration Between Concepts
Libo Zhang 0001, Tiejian Luo
KSEM2
2016 A fast filter tracker against serious occlusion
abstract
Many tracking algorithms applied in medical image processing, such as observing the movement of cells, have a great improvement in accuracy and robustness. However, it is difficult to deal with the large area occlusion and complete occlusion. In this paper, we propose a fast scale adaptive tracking algorithm based on correlation filtering. Except tracking the change of the target scale quickly, our method can also deal with the problem of large area occlusion and the complete disappearance of the target. Compared with the outstanding scale adaptive tracking method, the proposed method demonstrates higher performances in terms of the accuracy of tracking the target and the real-time performance.
Libo Zhang 0001, Tiejian Luo, Yihan Sun 0002
BIBM1
2016 Semantic analysis based on human thought pattern
abstract
Semantic analysis is an important component of recommendation systems and information retrieval in computer aided detection. Previous researches have made certain breakthroughs in disease diagnosis and drugs recommended by semantic analysis. We propose a bilateral shortest paths method for computing semantic relatedness based on the human thought patterns for making sufficient use of the hyperlink structure. The proposed novel method exploits bilateral shortest paths method to calculate word similarity, and employs the method of matrix partition to calculate text similarity. Finally, an evaluation based on WS353-Ex and Lee datasets is carried out and the result shows that we obtain effective performance.
Libo Zhang 0001, Tiejian Luo, Yihan Sun 0002
BIBM1
2016 A novel saliency detection method via manifold ranking and compactness prior
abstract
For improving the performance of saliency detection, several algorithms used graph construction have achieved excellent results. This paper proposes a novel bottom-up approach of saliency detection, which takes the advantages of both prior background and compactness. At first, we optimize the image boundary selections, by removing erroneous sections with a fixed threshold, to achieve more accurate saliency estimation results. The saliency map obtained by ranking with background queries can be optimized with compactness prior. The objects of salient are connected regions which are group together, with a compact form which are spatial distributed. Compared to the 8 state-of-the-art saliency detection approaches, our experimental results which test on the three public datasets show that the proposed algorithm improves accuracy and robustness significantly. This algorithm can find its potential applications in many different areas, but it is best suit for medical science and technologies because of high accuracy requirements. It can be used in the medical imaging processing to accurately differentiate tumor from bones, muscles and fats.
Libo Zhang 0001, Zakir Ullah, Yihan Sun 0002, Tiejian Luo
BIBM1
2016 MLPF algorithm for tracking fast moving target against light interference
abstract
In order to deal with the difficulty of tracking the fast moving aerial targets with light interference, we propose an improved particle tracking algorithm named multi-layers particle filter (MLPF). In MLPF, the particles are divided into three categories: the main particles (M-particles), the subordinate particles (S-particles) and the regenerate particles (R-particles). In the phase of resampling and state estimating, only M-particles are involved, then the R-particles are generated and considered as new S-particles in the next cycle. To a certain extent, our algorithm maintains the diversity of particles and reduces the computation time. Besides, MLPF has significant improvements on overcoming the tracing error after the sudden disappearance of the target and solving the degradation of particles. We demonstrate effectiveness of our proposed algorithm through systematic experiments. Experimental results show MLPF has better tracking effect compared to the traditional particle filter (PF) when the target is moving fast and affected by light interference. In the first experiment, the running time has been reduced from 47s to 21s while the precision increased from 64% to 96%. And for the second experiment, the running time has been reduced from 237s to 121s while precision increased from 46% to 89%.
Libo Zhang 0001, Yuanqiang Cai, Zakir Ullah, Tiejian Luo
ICPR1
2016 TACE: A Toolkit for Analyzing Concept Evolution in Computing Curricula
abstract
Effective teaching and learning requires to knowing whether some concepts are needed to grasping.Our investigation for ACM Computing Curricula from 1991 to 2013 shows that the numbers of concepts are increased by 5 times.That phenomenon makes learners harder to distinguish which concept is update or not.In this paper, we develop a solution to explore the body of knowledge of computing.The proposed toolkit uses a graph model to represent the disciplinary knowledge structure.The analytic results by TACE give us insight for the various subjects in computing discipline.Our findings show that 61.3% concepts in CC1991 are obsolete.In CC2001, the proportion of obsolete concepts drops to 11.5%, and in CS2008 it is 16.8%.The OS, IM, DS, AL's knowledge areas are more stable than CN, NC, GV, AR.The TACE's framework is highly modular, adaptive and extendible for analyzing other discipline's curricula.
Tiejian Luo, Libo Zhang 0001
SEKE2