EDBT 2026 Demo / reviewers in the wild / expert
Jianbin Jiao
dblp:02/6888
· DBLP profile ↗
109ranked-venue papers
2as first author
56since 2021 · last 2026
0000-0003-0454-3929ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 68 · 2 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 67 · 25 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsabstractBackdoor attacks embed malicious behaviors into Large Language Models (LLMs), enabling adversaries to trigger harmful outputs or bypass safety controls. However, the persistence of the implanted backdoors under user-driven post-deployment continual fine-tuning has been rarely examined. Most prior works evaluate the effectiveness and generalization of implanted backdoors only at releasing and empirical evidence shows that naively injected backdoor persistence degrades after updates. In this work, we study whether and how implanted backdoors persist through a multi‑stage post-deployment fine‑tuning. We propose P‑Trojan, a trigger‑based attack algorithm that explicitly optimizes for backdoor persistence across repeated updates. By aligning poisoned gradients with those of clean tasks on token embeddings, the implanted backdoor mapping is less likely to be suppressed or forgotten during subsequent updates. Theoretical analysis shows the feasibility of such persistent backdoor attacks after continual fine-tuning. And experiments conducted on the Qwen2.5 and LLaMA3 families of LLMs, as well as diverse task sequences, demonstrate that P‑Trojan achieves over \textbf{99\%} persistence while preserving clean‑task accuracy. Our findings highlight the need for persistence-aware evaluation and stronger defenses in realistic model adaptation pipelines. Jianbin Jiao, Junge Zhang |
AAAI | 3 |
| 2026 | MAGIC: Deep Geometric Evolution with Structural Consensus for Temporal Knowledge Graph ReasoningabstractTemporal Knowledge Graph (TKG) reasoning remains challenging to characterize with conventional flat representations due to its intrinsic heterogeneous structure.Existing multigeometry approaches face two key bottlenecks: 1) the Riemannian depth barrier driven by numerical instability, which restricts models to shallow architectures; and 2) gate collapse, where adaptive fusion mechanisms suffer from gradient starvation and degenerate into singlegeometry solutions.To this end, we propose MAGIC (Multi-geometry Annealing Graph Interaction with Consensus).Our framework introduces a Tangent-Residual Engine in multigeometric spaces, which enables the first stable 8-layer geometric evolution and reveals a phenomenon termed Geometric Annealing, where manifold curvature spontaneously evolves from semantic flatness in shallow layers to structural complexity in deeper layers.We further design an explicit reasoning module with structural consensus, leveraging geometric invariants and structural priors to regulate gradient flow, prevent collapse, and ensure robust synergy across Hyperbolic, Spherical, and Euclidean spaces.Experiments show that MAGIC achieves stateof-the-art performance in TKG reasoning, improving MRR by up to 2.9 points. Chengao Liu, Yuan Li 0055, Yingze Wang, Jianbin Jiao |
ACL (1) | 4 |
| 2026 | Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Hanghang Ma, Xiaoqi Ma, Xiaoming Wei, Jianbin Jiao, Enhua Wu, Jie Hu 0019 |
Int. J. Comput. Vis. | 7 |
| 2026 | FedPSAWA: Federated personalization with state aware weighting aggregation for cross subject seizure prediction
Peipei Gu, Jibin Shou, Yuping Zhao, Meiyan Xu, Jiayang Guo, Yan Zhang 0109, Jianbin Jiao, Jingzhu Li |
Neurocomputing | 9 |
| 2026 | CAGS: Open-vocabulary 3D scene understanding with context-aware Gaussian splatting
Wei Sun 0056, Yuan Li 0055, Jianbin Jiao |
Image Vis. Comput. | 3 |
| 2026 | KAN-YOLO: a nonlinear-enhanced and state space duality-driven YOLO network for robust ship detection in complex maritime scenes
Jingzhu Li, Lintao Xu, Jianbin Jiao, Qimeng Huang, Shunqing Yang, Pupu Wang, Ao Jiao |
Multim. Syst. | 3 |
| 2026 | SAPNet++: Evolving Point-Prompted Instance Segmentation With Semantic and Spatial AwarenessabstractSingle-point annotation is increasingly prominent in visual tasks for labeling cost reduction. However, it challenges tasks requiring high precision, such as the point-prompted instance segmentation (PPIS) task, which aims to estimate precise masks using single-point prompts to train a segmentation network. Due to the constraints of point annotations, granularity ambiguity and boundary uncertainty arise i.e., the difficulty distinguishing between different levels of detail (e.g., whole object vs. parts) and the challenge of precisely delineating object boundaries. Previous works have usually inherited the paradigm of mask generation along with proposal selection to achieve PPIS. However, proposal selection relies solely on category information, failing to resolve the ambiguity of different granularity. Furthermore, mask generators offer only finite discrete solutions that often deviate from actual masks, particularly at boundaries. To address these issues, we propose the Semantic-Aware Point-Prompted Instance Segmentation Network (SAPNet). It integrates Point Distance Guidance and Box Mining Strategy to tackle group and local issues caused by the point's granularity ambiguity. Additionally, we incorporate completeness scores within proposals to add spatial granularity awareness, enhancing multiple instance learning (MIL) in proposal selection termed S-MIL. The Multi-level Affinity Refinement conveys pixel and semantic clues, narrowing boundary uncertainty during mask refinement. These modules culminate in SAPNet++, mitigating point prompt's granularity ambiguity and boundary uncertainty and significantly improving segmentation performance. Extensive experiments on four challenging datasets validate the effectiveness of our methods, highlighting the potential to advance PPIS. Zhaoyang Wei, Xumeng Han, Xuehui Yu, Xue Yang 0005, Guorong Li, Zhenjun Han, Jianbin Jiao |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Audio denoising and audio-visual localization for unmanned aerial vehicles
Zhaojian Li 0002, Dengdian Huang, Jianbin Jiao, Jingzhu Li |
Pattern Recognit. | 4 |
| 2026 | Quantized DiT with hadamard transformation: A technical report
Wenxi Yang, Jianbin Jiao |
Pattern Recognit. Lett. | 3 |
| 2025 | EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement LearningabstractLarge Language Models (LLMs) have shown impressive reasoning capabilities in well-defined problems with clear solutions, such as mathematics and coding. However, they still struggle with complex real-world scenarios like business negotiations, which require strategic reasoning—an ability to navigate dynamic environments and align long-term goals amidst uncertainty.Existing methods for strategic reasoning face challenges in adaptability, scalability, and transferring strategies to new contexts.To address these issues, we propose explicit policy optimization (EPO) for strategic reasoning, featuring an LLM that provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior.To improve adaptability and policy transferability, we train the strategic reasoning model via multi-turn reinforcement learning (RL), utilizing process rewards and iterative self-play.Experiments across social and physical domains demonstrate EPO’s ability of long-term goal alignment through enhanced strategic reasoning, achieving state-of-the-art performance on social dialogue and web navigation tasks. Our findings reveal various collaborative reasoning mechanisms emergent in EPO and its effectiveness in generating novel strategies, underscoring its potential for strategic reasoning in real-world applications. Code and data are available at https://github.com/lxqpku/EPO. Yongbin Li 0001, Yuchuan Wu, Aobo Kong, Fei Huang 0002, Jianbin Jiao, Junge Zhang |
ACL (1) | 8 |
| 2025 | Adaptive Keyframe Sampling for Long Video UnderstandingabstractMultimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. However, when the visual input changes from a single image to a long video, the above paradigm encounters difficulty because the vast amount of video tokens has significantly exceeded the maximal capacity of MLLMs. Therefore, existing video-based MLLMs are mostly established upon sampling a small portion of tokens from input data, which can cause key information to be lost and thus produce incorrect answers. This paper presents a simple yet effective algorithm named Adaptive Keyframe Sampling (AKS). It inserts a plug-and-play module known as keyframe selection, which aims to maximize the useful information with a fixed number of video tokens. We formulate keyframe selection as an optimization involving (1) the relevance between the keyframes and the prompt, and (2) the coverage of the keyframes over the video, and present an adaptive algorithm to approximate the best solution. Experiments on two long video understanding benchmarks validate that AKS improves video QA accuracy (beyond strong baselines) upon selecting informative keyframes. Our study reveals the importance of information pre-filtering in video-based MLLMs. Our codes are available at https://github.com/ncTimTang/AKS Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, Qixiang Ye |
CVPR | 5 |
| 2025 | Night-to-Day Translation for Nighttime SurveillanceabstractNighttime surveillance suffers from degradation due to poor illumination and arduous human annotations. It is challengable and remains a security risk at night. Existing methods rely on multi-spectral images to perceive objects in the dark, which are troubled by color absence. We argue that the ultimate solution for nighttime surveillance is night-to-day translation, or Night2Day, which aims to translate a surveillance scene from nighttime to the daytime while maintaining semantic consistency. To achieve this, this paper presents a Disentangled Contrastive (DiCo) learning method. Specifically, to address the poor and complex illumination in the nighttime scenes, we propose a learnable physical prior, i.e., the color invariant, which provides a stable perception of a highly dynamic night environment and can be incorporated into the learning pipeline of neural networks. Targeting the surveillance scenes, we develop a disentangled representation, which is an auxiliary pretext task that separates surveillance scenes into the foreground and background with contrastive learning. Such a strategy can extract the semantics without supervision and boost our model to achieve instanceaware translation. Finally, we incorporate all the modules above into generative adversarial networks and achieve high-fidelity translation. This paper also contributes a new surveillance dataset called NightSuR. It includes six scenes to support the study on nighttime surveillance. This dataset collects nighttime images with different properties of nighttime environments, such as flare and extreme darkness. Extensive experiments demonstrate that our method outperforms existing works significantly. Jingzhu Li, Guanzhou Lan, Jianbin Jiao |
ICPADS | 4 |
| 2025 | VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual EvidenceabstractWith the rapid development of MLLMs, evaluating their visual capabilities has become increasingly crucial. Current benchmarks primarily fall into two main types: basic perception benchmarks,which focus on local details but lack deep reasoning (e.g., ''what is in the image?''), and mainstream reasoning benchmarks, which concentrate on prominent image elements but may fail to assess subtle clues requiring intricate analysis. However, profound visual understanding and complex reasoning depend more on interpreting subtle, inconspicuous local details than on perceiving salient, macro-level objects. These details, though occupying minimal image area, often contain richer, more critical information for robust analysis. To bridge this gap, we introduce the VER-Bench, a novel framework to evaluate MLLMs' ability to: 1) identify fine-grained visual clues, often occupying, on average, just 0.25% of the image area; 2) integrate these clues with world knowledge for complex reasoning. Comprising 374 carefully designed questions across Geospatial, Temporal, Situational, Intent, System State, and Symbolic reasoning, each question in VER-Bench is accompanied by structured evidence: visual clues and question-related reasoning derived from them. VER-Bench reveals current models' limitations in extracting subtle visual evidence and constructing evidence-based reasoning chains, highlighting the need to enhance models' capabilities in fine-grained visual evidence extraction, integration, and reasoning for genuine visual understanding and human-like analysis. The dataset is available at https://github.com/verbta/ACMMM-25-Materials. Chenhui Qiang, Zhaoyang Wei, Xumeng Han, Siyao Li, Xiangyuan Lan, Jianbin Jiao, Zhenjun Han |
ACM Multimedia | 7 |
| 2025 | Dual Discrepancy-Based Continuation Learning for Hybrid Supervised Multi-View 3D Object Detection
Tianyu Wang 0028, Feng Liu 0050, Jianbin Jiao, Fang Wan 0001 |
PRCV (17) | 4 |
| 2025 | P2Object: Single Point Supervised Object Detection and Instance Segmentation
Pengfei Chen 0004, Xuehui Yu, Xumeng Han, Kuiran Wang, Guorong Li, Lingxi Xie, Zhenjun Han, Jianbin Jiao |
Int. J. Comput. Vis. | 8 |
| 2025 | DiffusionLoc: A diffusion model-based framework for crowd localization
Yuan Li 0055, Yanzhao Zhou, Jianbin Jiao |
Image Vis. Comput. | 5 |
| 2025 | PoinKI: A Poincaré-Embedded knowledge-integrated model for dynamic link prediction
Chengao Liu, Yuan Li 0055, Siheng Ning, Jianming Zhu 0001, Jianbin Jiao |
Knowl. Based Syst. | 5 |
| 2025 | Beyond masking: Demystifying token-based pre-training for vision transformers
Yunjie Tian, Lingxi Xie, Jiemin Fang, Jianbin Jiao, Qi Tian 0001 |
Pattern Recognit. | 4 |
| 2025 | ClickTrack: Towards real-time interactive single object tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu, Guorong Li, Xiangyuan Lan, Qixiang Ye, Jianbin Jiao, Zhenjun Han |
Pattern Recognit. | 7 |
| 2025 | Depth-Guided Texture Diffusion for Image Semantic SegmentationabstractDepth information provides valuable insights into the 3D structure, especially the outline of objects, which can be utilized to enhance semantic segmentation. However, a naive fusion of depth information can disrupt features and compromise accuracy due to the gap between depth and RGB modalities. In this work, we introduce a depth-guided texture diffusion approach that effectively tackles the outlined challenge. Our method extracts low-level features from edges and textures to create a texture image. This image is then selectively diffused across the depth map, enhancing structural information vital for precisely extracting object outlines. By integrating this enhanced depth map with the original RGB image into a joint feature embedding, our method effectively bridges the disparity between depth and RGB modalities, enabling more accurate semantic segmentation. We conduct comprehensive experiments on diverse, widely-used datasets covering various semantic segmentation tasks, including Camouflaged Object Detection (COD), Salient Object Detection (SOD), and indoor semantic segmentation. With source-free estimated depth or depth captured by depth cameras, our method consistently outperforms existing baselines and achieves new state-of-the-art results, demonstrating the effectiveness of our depth-guided texture diffusion for image semantic segmentation. The source code and datasets are publicly available athttps://github.com/Wistzz/Texture-Diffusion.git. Wei Sun 0056, Yuan Li 0055, Qixiang Ye, Jianbin Jiao, Yanzhao Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Exploring Complicated Search Spaces With Interleaving-Free SamplingabstractConventional neural architecture search (NAS) algorithms typically work on search spaces with short-distance node connections. We argue that such designs, though safe and stable, are obstacles to exploring more effective network architectures. In this brief, we explore the search algorithm upon a complicated search space with long-distance connections and show that existing weight-sharing search algorithms fail due to the existence of interleaved connections (ICs). Based on the observation, we present a simple-yet-effective algorithm, termed interleaving-free neural architecture search (IF-NAS). We further design a periodic sampling strategy to construct subnetworks during the search procedure, avoiding the ICs to emerge in any of them. In the proposed search space, IF-NAS outperforms both random sampling and previous weight-sharing search algorithms by significant margins. It can also be well-generalized to the microcell-based spaces. This study emphasizes the importance of macrostructure and we look forward to further efforts in this direction. The code is available at github.com/sunsmarterjie/IFNAS. Yunjie Tian, Lingxi Xie, Jiemin Fang, Jianbin Jiao, Qixiang Ye, Qi Tian 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | BadRL: Sparse Targeted Backdoor Attack against Reinforcement LearningabstractBackdoor attacks in reinforcement learning (RL) have previously employed intense attack strategies to ensure attack success. However, these methods suffer from high attack costs and increased detectability. In this work, we propose a novel approach, BadRL, which focuses on conducting highly sparse backdoor poisoning efforts during training and testing while maintaining successful attacks. Our algorithm, BadRL, strategically chooses state observations with high attack values to inject triggers during training and testing, thereby reducing the chances of detection. In contrast to the previous methods that utilize sample-agnostic trigger patterns, BadRL dynamically generates distinct trigger patterns based on targeted state observations, thereby enhancing its effectiveness. Theoretical analysis shows that the targeted backdoor attack is always viable and remains stealthy under specific assumptions. Empirical results on various classic RL tasks illustrate that BadRL can substantially degrade the performance of a victim agent with minimal poisoning efforts (0.003% of total training steps) during training and infrequent attacks during testing. Code is available at: https://github.com/7777777cc/code. Yuzhe Ma, Jianbin Jiao, Junge Zhang |
AAAI | 4 |
| 2024 | Semantic-aware SAM for Point-Prompted Instance SegmentationabstractSingle-point annotation in visual tasks, with the goal of minimizing labelling costs, is becoming increasingly prominent in research. Recently, visual foundation models, such as Segment Anything (SAM), have gained widespread usage due to their robust zero-shot capabilities and exceptional annotation performance. However, SAM's class-agnostic output and high confidence in local segmentation introduce semantic ambiguity, posing a challenge for precise category-specific segmentation. In this paper, we introduce a cost-effective category-specific segmenter using SAM. To tackle this challenge, we have devised a Semantic-Aware Instance Segmentation Network (SAPNet) that integrates Multiple Instance Learning (MIL) with matching capability and SAM with point prompts. SAPNet strategically selects the most representative mask proposals generated by SAM to supervise segmentation, with a specific focus on object category information. Moreover, we introduce the Point Distance Guidance and Box Mining Strategy to mitigate inherent challenges: group and local issues in weakly supervised segmentation. These strategies serve to further enhance the overall segmentation performance. The experimental results on Pascal VOC and COCO demonstrate the promising performance of our proposed SAPNet, emphasizing its semantic matching capabilities and its potential to advance point-prompted instance segmentation. The code is available at https://github.com/zhaoyangwei123/SAPNet. Zhaoyang Wei, Pengfei Chen 0004, Xuehui Yu, Guorong Li, Jianbin Jiao, Zhenjun Han |
CVPR | 5 |
| 2024 | P2Seg: Pointly-supervised Segmentation via Mutual DistillationabstractPoint-level Supervised Instance Segmentation (PSIS) aims to enhance the applicability and scalability of instance segmentation by utilizing low-cost yet instance-informative annotations. Existing PSIS methods usually rely on positional information to distinguish objects, but predicting precise boundaries remains challenging due to the lack of contour annotations. Nevertheless, weakly supervised semantic segmentation methods are proficient in utilizing intra-class feature consistency to capture the boundary contours of the same semantic regions. In this paper, we design a Mutual Distillation Module (MDM) to leverage the complementary strengths of both instance position and semantic information and achieve accurate instance-level object perception. The MDM consists of Semantic to Instance (S2I) and Istance to Semantic (I2S). S2I is guided by the precise boundaries of semantic regions to learn the association between annotated points and instance contours. I2S leverages discriminative relationships between instances to facilitate the differentiation of various objects within the semantic map. Extensive experiments substantiate the efficacy of MDM in fostering the synergy between instance and semantic information, consequently improving the quality of instance-level object representations. Our method achieves 55.7 mAP50 and 17.6 mAP on the PASCAL VOC and MS COCO datasets, significantly outperforming recent PSIS methods and several box-supervised instance segmentation competitors. Xuehui Yu, Xumeng Han, Wenwen Yu, Zhixun Huang, Jianbin Jiao, Zhenjun Han |
ICLR | 6 |
| 2024 | Position: Foundation Agents as the Paradigm Shift for Decision MakingabstractDecision making demands intricate interplay between perception, memory, and reasoning to discern optimal policies. Conventional approaches to decision making face challenges related to low sample efficiency and poor generalization. In contrast, foundation models in language and vision have showcased rapid adaptation to diverse new tasks. Therefore, we advocate for the construction of foundation agents as a transformative shift in the learning paradigm of agents. This proposal is underpinned by the formulation of foundation agents with their fundamental characteristics and challenges motivated by the success of large language models (LLMs). Moreover, we specify the roadmap of foundation agents from large interactive data collection or generation, to self-supervised pretraining and adaptation, and knowledge and value alignment with LLMs. Lastly, we pinpoint critical research questions derived from the formulation and delineate trends for foundation agents supported by real-world use cases, addressing both technical and theoretical aspects to propel the field towards a more comprehensive and impactful future. Xingzhou Lou, Jianbin Jiao, Junge Zhang |
ICML | 3 |
| 2024 | Towards Precise 3D Human Pose Estimation with Multi-Perspective Spatial-Temporal Relational Transformersabstract3D human pose estimation captures the human joint points in three-dimensional space while keeping the depth information and physical structure. That is essential for applications that require precise pose information, such as humancomputer interaction, scene understanding, and rehabilitation training. Due to the challenges in data collection, mainstream datasets of 3D human pose estimation are primarily composed of multi-view video data collected in laboratory environments, which contains rich spatial-temporal correlation information besides the image frame content. Given the remarkable selfattention mechanism of transformers, capable of capturing the spatial-temporal correlation from multi-view video datasets, we propose a multi-stage framework for 3D sequence-to-sequence (seq2seq) human pose detection. Firstly, the spatial module represents the human pose feature by intra-image content, while the frame-image relation module extracts temporal relationships and 3D spatial positional relationship features between the multiperspective images. Secondly, the self-attention mechanism is adopted to eliminate the interference from non-human body parts and reduce computing resources. Our method is evaluated on Human3.6M, a popular 3D human pose detection dataset. Experimental results demonstrate that our approach achieves stateof-the-art performance on this dataset. The source code will be available at https://github.com/WUJINHUAN/3D-human-pose. Jianbin Jiao, Xina Cheng, Xiaoting Yin, Hao Shi 0004, Kailun Yang 0001 |
IJCNN | 1 |
| 2024 | VMamba: Visual State Space ModelabstractDesigning computationally efficient network architectures remains an ongoing necessity in computer vision. In this paper, we adapt Mamba, a state-space language model, into VMamba, a vision backbone with linear time complexity. At the core of VMamba is a stack of Visual State-Space (VSS) blocks with the 2D Selective Scan (SS2D) module. By traversing along four scanning routes, SS2D bridges the gap between the ordered nature of 1D selective scan and the non-sequential structure of 2D vision data, which facilitates the collection of contextual information from various sources and perspectives. Based on the VSS blocks, we develop a family of VMamba architectures and accelerate them through a succession of architectural and implementation enhancements. Extensive experiments demonstrate VMamba’s
promising performance across diverse visual perception tasks, highlighting its superior input scaling efficiency compared to existing benchmark models. Source code is available at https://github.com/MzeroMiko/VMamba Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang 0001, Qixiang Ye, Jianbin Jiao, Yunfan Liu 0001 |
NeurIPS | 8 |
| 2024 | SRX Net: A point cloud segmentation framework based on surface representation and X-Net
Wei Zhou 0027, Jianbin Jiao, Mingan Wei, Wang Nei, Hai-Xia Xu 0001 |
Comput. Graph. | 2 |
| 2024 | Unifying multimodal interactions for rumor diffusion prediction with global hypergraph modeling
Yuan Li 0055, Jialing Zou, Jianming Zhu 0001, Dingning Liu, Jianbin Jiao |
Knowl. Based Syst. | 6 |
| 2024 | PointBiMssc: Bidirectional Multiscale Attention-Based Point Cloud Semantic Segmentation for Water Conservancy EnvironmentabstractPoint cloud semantic segmentation is a key technique for the digital twin construction of water conservancy projects, which can realize the identification and change detection of terrain features. However, constructing full-range point cloud data of the water conservancy environment remains one of the critical challenges in achieving comprehensive digital twin construction. Meanwhile, the openness of the water conservancy environment also makes its point cloud data structure highly complex, posing challenges to the accuracy and robustness of point cloud semantic segmentation algorithms. Therefore, we adopt unmanned aerial vehicle (UAV)-borne lidar to scan water conservancy scenes and construct a large-scale point cloud dataset, Water Conservancy Segment 3-D (WCS3D), with approximately 265 million points. On this basis, we propose a point cloud segmentation model named PointBiMssc based on a bidirectional multiscale attention mechanism for point cloud semantic segmentation in the water conservancy environment. A series of experiments conducted on the WCS3D dataset demonstrate that the PointBiMssc model can accurately complete point cloud semantic segmentation tasks and generate high-precision segmentation boundaries, outperforming the latest transformer-based models in mean intersection over union (mIoU) and overall pointwise accuracy (OA) evaluation metrics to achieve state-of-the-art performance. Code:https://github.com/JJBUP/PointBiMssc, dataset:https://github.com/XTU-SCS-HappyCV/WCS3D. Wei Zhou 0027, Jianbin Jiao, Hai-Xia Xu 0001, Mingan Wei, Xueqiang Zhao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network With Token MigrationabstractWe propose integrally pre-trained transformer pyramid network (iTPN), towards jointly optimizing the network backbone and the neck, so that transfer gap between representation models and downstream tasks is minimal. iTPN is born with two elaborated designs: 1) The first pre-trained feature pyramid upon vision transformer (ViT). 2) Multi-stage supervision to the feature pyramid using masked feature modeling (MFM). iTPN is updated to Fast-iTPN, reducing computational memory overhead and accelerating inference through two flexible designs. 1) Token migration: dropping redundant tokens of the backbone while replenishing them in the feature pyramid without attention operations. 2) Token gathering: reducing computation cost caused by global attention by introducing few gathering tokens. The base/large-level Fast-iTPN achieve 88.75%/89.5% top-1 accuracy on ImageNet-1 K. With 1× training schedule using DINO, the base/large-level Fast-iTPN achieves 58.4%/58.8% box AP on COCO object detection, and a 57.5%/58.7% mIoU on ADE20 K semantic segmentation using MaskDINO. Fast-iTPN can accelerate the inference procedure by up to 70%, with negligible performance loss, demonstrating the potential to be a powerful backbone for downstream vision tasks. Yunjie Tian, Lingxi Xie, Jihao Qiu, Jianbin Jiao, Yaowei Wang 0001, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | CPR++: Object Localization via Single Coarse Point SupervisionabstractPoint-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance due to the inconsistency of annotated points. Existing POL heavily rely on strict annotation rules, which are difficult to define and apply, to handle the problem. In this study, we propose coarse point refinement (CPR), which to our best knowledge is the first attempt to alleviate semantic variance from an algorithmic perspective. CPR reduces the semantic variance by selecting a semantic centre point in a neighbourhood region to replace the initial annotated point. Furthermore, We design a sampling region estimation module to dynamically compute a sampling region for each object and use a cascaded structure to achieve end-to-end optimization. We further integrate a variance regularization into the structure to concentrate the predicted scores, yielding CPR++. We observe that CPR++ can obtain scale information and further reduce the semantic variance in a global region, thus guaranteeing high-performance object localization. Extensive experiments on four challenging datasets validate the effectiveness of both CPR and CPR++. We hope our work can inspire more research on designing algorithms rather than annotation rules to address the semantic variance problem in POL. Xuehui Yu, Pengfei Chen 0004, Kuiran Wang, Xumeng Han, Guorong Li, Zhenjun Han, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Self-supervised feature-gate coupling for dynamic network pruning
Chang Liu 0047, Jianbin Jiao, Qixiang Ye |
Pattern Recognit. | 3 |
| 2024 | Generic-to-Specific Distillation of Masked AutoencodersabstractTo transfer the representation capacity of large pre-trained models to lightweight models, knowledge distillation has been widely explored. However, conventional single-stage distillation methods are prone to getting stuck in the transfer of task-specific knowledge, making it difficult to retain task-agnostic knowledge which is crucial for model generalization. In this study, we propose generic-to-specific distillation (G2SD), to boost lightweight models under the assistance of large models pre-trained by masked image modeling. In generic distillation, the decoder of a small model is encouraged to align feature predictions with that of a large model, so that task-agnostic knowledge can be transferred. In specific distillation, predictions of the small model are encouraged to be consistent with those of the large model, to guarantee task performance. G2SD is also applicable for heterogeneous settings(i.e., distilling from ViT to CNN). With G2SD, the ViT-Small model respectively achieves 98.9%, 98.4%, 99.3% and 98.9% accuracies when compared with its teachers (ViT-Base) for image classification, object detection, semantic segmentation and video recognition tasks. The lightweight ResNet models are improved to a new height on image classification task. The code is available at github.com/pengzhiliang/G2SD. Zhiliang Peng, Li Dong 0004, Furu Wei, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Generic-to-Specific Distillation of Masked AutoencodersabstractLarge vision Transformers (ViTs) driven by self-supervised pre-training mechanisms achieved unprecedented progress. Lightweight ViT models limited by the model capacity, however, benefit little from those pre-training mechanisms. Knowledge distillation defines a paradigm to transfer representations from large (teacher) models to small (student) ones. However, the conventional single-stage distillation easily gets stuck on task-specific transfer, failing to retain the task-agnostic knowledge crucial for model generalization. In this study, we propose generic-to-specific distillation (G2SD), to tap the potential of small ViT models under the supervision of large models pretrained by masked autoencoders. In generic distillation, decoder of the small model is encouraged to align feature predictions with hidden representations of the large model, so that task-agnostic knowledge can be transferred. In specific distillation, predictions of the small model are constrained to be consistent with those of the large model, to transfer task-specific features which guarantee task performance. With G2SD, the vanilla ViT-Small model respectively achieves 98.7%, 98.1% and 99.3% the performance of its teacher (ViT-Base) for image classification, object detection, and semantic segmentation, setting a solid baseline for two-stage vision distillation. Code will be available at https://github.com/pengzhiliang/G2SD Zhiliang Peng, Li Dong 0004, Furu Wei, Jianbin Jiao, Qixiang Ye |
CVPR | 5 |
| 2023 | Integrally Pre-Trained Transformer Pyramid NetworksabstractIn this paper, we present an integral pre-training framework based on masked image modeling (MIM). We advocate for pre-training the backbone and neck jointly so that the transfer gap between MIM and downstream recognition tasks is minimal. We make two technical contributions. First, we unify the reconstruction and recognition necks by inserting a feature pyramid into the pre-training stage. Second, we complement mask image modeling (MIM) with masked feature modeling (MFM) that offers multi-stage supervision to the feature pyramid. The pre-trained models, termed integrally pre-trained transformer pyramid networks (iTPNs), serve as powerful foundation models for visual recognition. In particular, the base/large-level iTPN achieves an 86.2%/87.8% top-1 accuracy on ImageNet-1K, a 53.2%/55.6% box AP on COCO object detection with 1× training schedule using Mask-RCNN, and a 54.7%/57.7% mIoU on ADE20K semantic segmentation using UPerHead – all these results set new records. Our work inspires the community to work on unifying upstream pre-training and downstream fine-tuning tasks. Code is available at github.com/sunsmarterjie/iTPN. Yunjie Tian, Lingxi Xie, Zhaozhi Wang, Longhui Wei, Xiaopeng Zhang 0008, Jianbin Jiao, Yaowei Wang 0001, Qi Tian 0001, Qixiang Ye |
CVPR | 6 |
| 2023 | Spatial Self-Distillation for Object Detection with Inaccurate Bounding BoxesabstractObject detection via inaccurate bounding boxes supervision has boosted a broad interest due to the expensive high-quality annotation data or the occasional inevitability of low annotation quality (e.g. tiny objects). The previous works usually utilize multiple instance learning (MIL), which highly depends on category information, to select and refine a low-quality box. Those methods suffer from object drift, group prediction and part domination problems without exploring spatial information. In this paper, we heuristically propose a Spatial Self-Distillation based Object Detector (SSD-Det) to mine spatial information to refine the inaccurate box in a self-distillation fashion. SSD-Det utilizes a Spatial Position Self-Distillation (SPSD) module to exploit spatial information and an interactive structure to combine spatial information and category information, thus constructing a high-quality proposal bag. To further improve the selection procedure, a Spatial Identity Self-Distillation (SISD) module is introduced in SSD-Det to obtain spatial confidence to help select the best proposals. Experiments on MS-COCO and VOC datasets with noisy box annotation verify our method’s effectiveness and achieve state-of-the-art performance. The code is available at https://github.com/ucas-vg/PointTinyBenchmark/tree/SSD-Det. Pengfei Chen 0004, Xuehui Yu, Guorong Li, Zhenjun Han, Jianbin Jiao |
ICCV | 6 |
| 2023 | Conformer: Local Features Coupling Global Representations for Recognition and DetectionabstractWith convolution operations, Convolutional Neural Networks (CNNs) are good at extracting local features but experience difficulty to capture global representations. With cascaded self-attention modules, vision transformers can capture long-distance feature dependencies but unfortunately deteriorate local feature details. In this paper, we propose a hybrid network structure, termed Conformer, to take both advantages of convolution operations and self-attention mechanisms for enhanced representation learning. Conformer roots in feature coupling of CNN local features and transformer global representations under different resolutions in an interactive fashion. Conformer adopts a dual structure so that local details and global dependencies are retained to the maximum extent. We also propose a Conformer-based detector (ConformerDet), which learns to predict and refine object proposals, by performing region-level feature coupling in an augmented cross-attention fashion. Experiments on ImageNet and MS COCO datasets validate Conformer's superiority for visual recognition and object detection, demonstrating its potential to be a general backbone network. Zhiliang Peng, Zonghao Guo, Yaowei Wang 0001, Lingxi Xie, Jianbin Jiao, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Rethinking Sampling Strategies for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that sampling strategy plays an equally important role. We analyze the reasons for the performance differences between various sampling strategies under the same framework and loss function. We suggest that deteriorated over-fitting is an important factor causing poor performance, and enhancing statistical stability can rectify this problem. Inspired by that, a simple yet effective approach is proposed, termed group sampling, which gathers samples from the same class into groups. The model is thereby trained using normalized group samples, which helps alleviate the negative impact of individual samples. Group sampling updates the pipeline of pseudo-label generation by guaranteeing that samples are more efficiently classified into the correct classes. It regulates the representation learning process, enhancing statistical stability for feature representation in a progressive fashion. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 show that group sampling achieves performance comparable to state-of-the-art methods and outperforms the current techniques under purely camera-agnostic settings. Code has been available at https://github.com/ucas-vg/GroupSampling. Xumeng Han, Xuehui Yu, Guorong Li, Jian Zhao 0006, Gang Pan 0002, Qixiang Ye, Jianbin Jiao, Zhenjun Han |
IEEE Trans. Image Process. | 7 |
| 2023 | Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV TrackingabstractUnmanned Aerial Vehicles (UAV) have many applications in both commerce and recreation. However, irresponsibly operated UAVs will pose a threat to public safety. Therefore, developing our understanding of UAVs and their uses is of particular interest. This paper considers tracking UAVs, which provide multifaceted information around location, paths and trajectories. To facilitate research on this topic, we introduce a new benchmark, herein referred to as Anti-UAV, which provides a novel direction for UAV tracking with more than 300 video pairs containing over 580 k manually annotated bounding boxes. Addressing anti-UAV research challenges could help to design anti-UAV systems, which in turn may improve surveillance. Accordingly, we have proposed a simple yet effective approach, called dual-flow semantic consistency (DFSC) is proposed for UAV tracking. Modulated by the semantic flow across video sequences, tracker learns more robust class-level semantic information and obtains more discriminative instance-level features. Experiments highlight significant performance gain with the proposed approach over state-of-the-art trackers and the challenging aspects of Anti-UAV. The Anti-UAV benchmark and the code for the proposed approach have been made publicly available athttps://github.com/ucas-vg/Anti-UAVandhttps://github.com/ZhaoJ9014/Anti-UAV. Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang 0051, Junliang Xing, Guorong Li, Guodong Guo, Qixiang Ye, Jianbin Jiao, Jian Zhao 0006, Zhenjun Han |
IEEE Trans. Multim. | 10 |
| 2022 | Dynamic Perception Framework for Fine-Grained RecognitionabstractFine-grained recognition poses the challenge of discriminating categories with only small subtle visual differences, which can be easily overwhelmed by diverse appearance within categories. Conventional approaches generally locate discriminative parts and then recognize the part-based features. However, we find that tuning the effective receptive field (ERF) of the network to the task plays the key role, which enables significant regions to contribute more to the output. Inspired by the receptive field stimulation mechanism of the visual cortex, we propose a Dynamic Perception framework as a solution. Our framework adapts the ERF by considering the image space and the kernel space simultaneously. In the image space, the Spatial Selective Sampling module is adopted to enlarge informative regions locally. In the kernel space, Spatial Selective Kernel convolution is introduced to adapt different kernel sizes for regions of interest and backgrounds by embedding spatial attention in the multi-path convolution. Extensive experiments on challenging benchmarks, including CUB-200-2011, FGVC-Aircraft, and Stanford Cars, demonstrate that our method yields a performance boost over the state-of-the-art methods. Yao Ding 0006, Zhenjun Han, Yanzhao Zhou, Yi Zhu 0004, Jie Chen 0001, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | Convex-Hull Feature Adaptation for Oriented and Densely Packed Object DetectionabstractDetecting oriented and densely packed objects is a challenging problem considering that the receptive field intersection between objects causes spatial feature aliasing. In this paper, we propose a convex-hull feature adaptation (CFA) approach, with the aim to configure convolutional features in accordance with irregular object layouts. CFA roots in the convex-hull feature representation, which defines a set of dynamically sampled feature points guided by the convex intersection over union (CIoU) to bound object extent. CFA pursues optimal feature assignment by constructing convex-hull sets and iteratively splitting positive or negative convex-hulls. By simultaneously considering overlapping convex-hulls and objects and penalizing convex-hulls shared by multiple objects, CFA defines a systematic way to adapt convolutional features on regular grids to objects of irregular shapes. Experiments on DOTA and SKU110K-R datasets show that CFA achieved new state-of-the-art performance for detecting oriented and densely packed objects. CFA also sets a solid baseline for convex polygon prediction on the MS COCO dataset defined for general object detection. Code is available athttps://github.com/SDL-GuoZonghao/BeyondBoundingBox. Zonghao Guo, Xiaosong Zhang 0004, Chang Liu 0047, Xiangyang Ji, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Configurable Graph Reasoning for Visual Relationship DetectionabstractVisual commonsense knowledge has received growing attention in the reasoning of long-tailed visual relationships biased in terms of object and relation labels. Most current methods typically collect and utilize external knowledge for visual relationships by following the fixed reasoning path of {subject, object → predicate} to facilitate the recognition of infrequent relationships. However, the knowledge incorporation for such fixed multidependent path suffers from the data set biased and exponentially grown combinations of object and relation labels and ignores the semantic gap between commonsense knowledge and real scenes. To alleviate this, we propose configurable graph reasoning (CGR) to decompose the reasoning path of visual relationships and the incorporation of external knowledge, achieving configurable knowledge selection and personalized graph reasoning for each relation type in each image. Given a commonsense knowledge graph, CGR learns to match and retrieve knowledge for different subpaths and selectively compose the knowledge routed path. CGR adaptively configures the reasoning path based on the knowledge graph, bridges the semantic gap between the commonsense knowledge, and the real-world scenes and achieves better knowledge generalization. Extensive experiments show that CGR consistently outperforms previous state-of-the-art methods on several popular benchmarks and works well with different knowledge graphs. Detailed analyses demonstrated that CGR learned explainable and compelling configurations of reasoning paths. Yi Zhu 0004, Xiwen Liang, Bingqian Lin, Qixiang Ye, Jianbin Jiao, Liang Lin 0004, Xiaodan Liang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Beyond Bounding-Box: Convex-Hull Feature Adaptation for Oriented and Densely Packed Object DetectionabstractDetecting oriented and densely packed objects remains challenging for spatial feature aliasing caused by the intersection of reception fields between objects. In this paper, we propose a convex-hull feature adaptation (CFA) approach for configuring convolutional features in accordance with oriented and densely packed object layouts. CFA is rooted in convex-hull feature representation, which defines a set of dynamically predicted feature points guided by the convex intersection over union (CIoU) to bound the extent of objects. CFA pursues optimal feature assignment by constructing convex-hull sets and dynamically splitting positive or negative convex-hulls. By simultaneously considering overlapping convex-hulls and objects and penalizing convex-hulls shared by multiple objects, CFA alleviates spatial feature aliasing towards optimal feature adaptation. Experiments on DOTA and SKU110K-R datasets show that CFA significantly outperforms the baseline approach, achieving new state-of-the-art detection performance. Code is available at github.com/SDL-GuoZonghao/BeyondBoundingBox. Zonghao Guo, Chang Liu 0042, Xiaosong Zhang 0004, Jianbin Jiao, Xiangyang Ji, Qixiang Ye |
CVPR | 4 |
| 2021 | Anti-Aliasing Semantic Reconstruction for Few-Shot Semantic SegmentationabstractEncouraging progress in few-shot semantic segmentation has been made by leveraging features learned upon base classes with sufficient training data to represent novel classes with few-shot examples. However, this feature sharing mechanism inevitably causes semantic aliasing between novel classes when they have similar compositions of semantic concepts. In this paper, we reformulate few-shot segmentation as a semantic reconstruction problem, and convert base class features into a series of basis vectors which span a class-level semantic space for novel class reconstruction. By introducing contrastive loss, we maximize the orthogonality of basis vectors while minimizing semantic aliasing between classes. Within the reconstructed representation space, we further suppress interference from other classes by projecting query features to the support vector for precise semantic activation. Our proposed approach, referred to as anti-aliasing semantic reconstruction (ASR), provides a systematic yet interpretable solution for few-shot learning problems. Extensive experiments on PASCAL VOC and MS COCO datasets show that ASR achieves strong results compared with the prior works. Code will be released at github.com/Bibkiller/ASR. Binghao Liu, Yao Ding 0006, Jianbin Jiao, Xiangyang Ji, Qixiang Ye |
CVPR | 3 |
| 2021 | Self-Motivated Communication Agent for Real-World Vision-Dialog NavigationabstractVision-Dialog Navigation (VDN) requires an agent to ask questions and navigate following the human responses to find target objects. Conventional approaches are only allowed to ask questions at predefined locations, which are built upon expensive dialogue annotations, and inconvenience the real-word human-robot communication and cooperation. In this paper, we propose a Self-Motivated Communication Agent (SCoA) that learns whether and what to communicate with human adaptively to acquire instructive information for realizing dialogue annotation-free navigation and enhancing the transferability in real-world unseen environment. Specifically, we introduce a whether-to-ask (WeTA) policy, together with uncertainty of which action to choose, to indicate whether the agent should ask a question. Then, a what-to-ask (WaTA) policy is proposed, in which, along with the oracle’s answers, the agent learns to score question candidates so as to pick up the most informative one for navigation, and meanwhile mimic oracle’s answering. Thus, the agent can navigate in a self-Q&A manner even in real-world environment where the human assistance is often unavailable. Through joint optimization of communication and navigation in a unified imitation learning and reinforcement learning framework, SCoA asks a question if necessary and obtains a hint for guiding the agent to move towards the target with less communication cost. Experiments on seen and unseen environments demonstrate that SCoA shows not only superior performance over existing baselines without dialog annotations, but also competing results compared with rich dialog annotations based counterparts. Yi Zhu 0004, Yue Weng, Fengda Zhu, Xiaodan Liang, Qixiang Ye, Yutong Lu, Jianbin Jiao |
ICCV | 7 |
| 2021 | Conformer: Local Features Coupling Global Representations for Visual RecognitionabstractWithin Convolutional Neural Network (CNN), the convolution operations are good at extracting local features but experience difficulty to capture global representations. Within visual transformer, the cascaded self-attention modules can capture long-distance feature dependencies but unfortunately deteriorate local feature details. In this paper, we propose a hybrid network structure, termed Conformer, to take advantage of convolutional operations and self-attention mechanisms for enhanced representation learning. Conformer roots in the Feature Coupling Unit (FCU), which fuses local features and global representations under different resolutions in an interactive fashion. Conformer adopts a concurrent structure so that local features and global representations are retained to the maximum extent. Experiments show that Conformer, under the comparable parameter complexity, outperforms the visual transformer (DeiT-B) by 2.3% on ImageNet. On MSCOCO, it outperforms ResNet-101 by 3.7% and 3.6% mAPs for object detection and instance segmentation, respectively, demonstrating the great potential to be a general backbone network. Code is available at github.com/pengzhiliang/Conformer. Zhiliang Peng, Shanzhi Gu, Lingxi Xie, Yaowei Wang 0001, Jianbin Jiao, Qixiang Ye |
ICCV | 6 |
| 2021 | Interpretable Credit Risk Assessment Based on Heuristic Knowledge Extraction MethodabstractBuilding explainable model has become an important issue for credit risk assessment. Results can be presented as rule-based knowledge and are therefore considered interpretable because they indicate causation of classification. However, traditional knowledge extraction methods are not suited to finance because credit data involves discrete and continuous data, with missing values and imbalanced labels. In this study, a novel multi-label classification method is proposed, which summarizes high-dimensional structured data observations into a rule-based classifier in the form of a rule list. The solution is a novel hybrid evolutionary algorithm (hEA) which avoids preprocessing original data from complex credit datasets. The proposed method has significant advantages over established interpretable classification methods in terms of classification performance on complex credit data. The complete code of our proposed method are available at https://github.com/wenge963/The-RATP-method. Zhiwen Xiao, Jianbin Jiao |
ICTAI | 2 |
| 2021 | Long-tailed Distribution AdaptationabstractRecognizing images with long-tailed distributions remains a challenging problem while there lacks an interpretable mechanism to solve this problem. In this study, we formulate Long-tailed recognition as Domain Adaption (LDA), by modeling the long-tailed distribution as an unbalanced domain and the general distribution as a balanced domain. Within the balanced domain, we propose to slack the generalization error bound, which is defined upon the empirical risks of unbalanced and balanced domains and the divergence between them. We propose to jointly optimize empirical risks of the unbalanced and balanced domains and approximate their domain divergence by intra-class and inter-class distances, with the aim to adapt models trained on the long-tailed distribution to general distributions in an interpretable way. Experiments on benchmark datasets for image recognition, object detection, and instance segmentation validate that our LDA approach, beyond its interpretability, achieves state-of-the-art performance. Zhiliang Peng, Zonghao Guo, Xiaosong Zhang 0004, Jianbin Jiao, Qixiang Ye |
ACM Multimedia | 5 |
| 2021 | Discretization-aware architecture search
Yunjie Tian, Chang Liu 0042, Lingxi Xie, Jianbin Jiao, Qixiang Ye |
Pattern Recognit. | 4 |
| 2021 | Genetic Feature Fusion for Object Skeleton DetectionabstractObject skeleton detection requires the convolutional neural networks to recognize objects and their parts in the cluttered background, overcome the image definition degradation brought by the pooling layers, and predict the location of skeleton pixels in different scale granularity. Most existing object skeleton detection methods take great efforts into the designing of side-output networks for multiscale feature fusion. Despite the great progress achieved by them, there are still many problems that hinder the development of object skeleton detection, such as the manually designed network is labor-intensive and the network initialization depends on models pretrained on large-scale datasets. To alleviate these issues, we propose a genetic NAS method to automatically search on a newly designed architecture search space for adaptive multiscale feature fusion. Furthermore, we introduce a symmetric encoder-decoder search space based on reversing the VGG network, in which the decoder can reuse the ImageNet pretrained model of VGG. The searched networks improve the performance of the state-of-the-art methods on commonly used skeleton detection benchmarks, which proves the efficacy of our method. Yunjie Tian, Jianbin Jiao |
Secur. Commun. Networks | 4 |
| 2021 | Explainable Fraud Detection for Few Labeled Time Series DataabstractFraud detection technology is an important method to ensure financial security. It is necessary to develop explainable fraud detection methods to express significant causality for participants in the transaction. The main contribution of our work is to propose an explainable classification method in the framework of multiple instance learning (MIL), which incorporates the AP clustering method in the self-training LSTM model to obtain a clear explanation. Based on a real-world dataset and a simulated dataset, we conducted two comparative studies to evaluate the effectiveness of the proposed method. Experimental results show that our proposed method achieves the similar predictive performance as the state-of-art method, while our method can generate clear causal explanations for a few labeled time series data. The significance of the research work is that financial institutions can use this method to efficiently identify fraudulent behaviors and easily give reasons for rejecting transactions so as to reduce fraud losses and management costs. Zhiwen Xiao, Jianbin Jiao |
Secur. Commun. Networks | 2 |
| 2021 | Rethinking Triplet Loss for Domain AdaptationabstractThe gap in data distribution motivates domain adaptation research. In this area, image classification intrinsically requires the source and target features to be co-located if they are of the same class. However, many works only take a global view of the domain gap. That is, to make the data distributions globally overlap; and this does not necessarily lead to feature co-location at the class level. To resolve this problem, we study metric learning in the context of domain adaptation. Specifically, we introduce a similarity guided constraint (SGC). In the implementation, SGC takes the form of a triplet loss. The triplet loss is integrated into the network as an additional objective term. Here, an image triplet consists of two images of the same class and another image of a different class. Albeit simple, the working mechanism of our method is interesting and insightful. Importantly, images in the triplets are sampled from the source and target domains. From a micro perspective, by enforcing this constraint on every possible triplet, images from different domains but of the same class are mapped nearby, and those of different classes are far apart. From a macro perspective, our method ensures that cross-domain similarities are preserved, leading to intra-class compactness and inter-class separability. Extensive experiment on four datasets shows our method yields significant improvement over the baselines and has a competitive accuracy with the state-of-the-art results. Weijian Deng, Liang Zheng 0001, Yifan Sun 0003, Jianbin Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Harmonic Feature Activation for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation remains an open problem because limited support (training) images are insufficient to represent the diverse semantics within target categories. Conventional methods typically model a target category solely using information from the support image(s), resulting in incomplete semantic activation. In this paper, we propose a novel few-shot segmentation approach, termed harmonic feature activation (HFA), with the aim to implement dense support-to-query semantic transform by incorporating the features of both query and support images. HFA is formulated as a bilinear model, which takes charge of the pixel-wise dense correlation (bilinear feature activation) between query and support images in a systematic way. HFA incorporates a low-rank decomposition procedure, which speeds up bilinear feature activation with negligible performance cost. In addition, a semantic diffusion procedure is fused with HFA, which further improves the global harmony and local consistency of the feature activation. Extensive experiments on commonly used datasets (PASCAL VOC and MS COCO) show that HFA improves the state-of-the-arts with significant margins. Code is available at https://github.com/Bibikiller/HFA. Binghao Liu, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Image Process. | 2 |
| 2021 | Adaptive Linear Span Network for Object Skeleton DetectionabstractConventional networks for object skeleton detection are usually hand-crafted. Despite the effectiveness, hand-crafted network architectures lack the theoretical basis and require intensive prior knowledge to implement representation complementarity for objects/parts in different granularity. In this paper, we propose an adaptive linear span network (AdaLSN), driven by neural architecture search (NAS), to automatically configure and integrate scale-aware features for object skeleton detection. AdaLSN is formulated with the theory of linear span, which provides one of the earliest explanations for multi-scale deep feature fusion. AdaLSN is materialized by defining a mixed unit-pyramid search space, which goes beyond many existing search spaces using unit-level or pyramid-level features. Within the mixed space, we apply genetic architecture search to jointly optimize unit-level operations and pyramid-level connections for adaptive feature space expansion. AdaLSN substantiates its versatility by achieving significantly higher accuracy and latency trade-off compared with the state-of-the-arts. It also demonstrates general applicability to image-to-mask tasks such as edge detection and road extraction. Code is available at https://github.com/sunsmarterjie/SDL-Skeletongithub.com/sunsmarterjie/SDL-Skeleton. Chang Liu 0047, Yunjie Tian, Zhiwen Chen 0002, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Image Process. | 4 |
| 2021 | SRN: Side-Output Residual Network for Object Reflection Symmetry Detection and BeyondabstractThis article establishes a baseline for object reflection symmetry detection in natural images by releasing a new benchmark named Sym-PASCAL and proposing an end-to-end deep learning approach for reflection symmetry. Sym-PASCAL spans challenges of multiobjects, object diversity, part invisibility, and clustered backgrounds, which is far beyond those in existing data sets. The end-to-end deep learning approach, referred to as a side-output residual network (SRN), leverages the output residual units (RUs) to fit the errors between the symmetry ground truth and the side outputs of multiple stages of a trunk network. By cascading RUs from deep to shallow, SRN exploits the "flow" of errors along multiple stages to effectively matching object symmetry at different scales and suppress the clustered backgrounds. SRN is interpreted as a boosting-like algorithm, which assembles features using RUs during network forward and backward propagations. SRN is further upgraded to a multitask SRN (MT-SRN) for joint symmetry and edge detection, demonstrating its generality to image-to-mask learning tasks. Experimental results verify that the Sym-PASCAL benchmark is challenging related to real-world images, SRN achieves state-of-the-art performance, and MT-SRN has the capability to simultaneously predict edge and symmetry mask without loss of performance. Wei Ke 0003, Jie Chen 0001, Jianbin Jiao, Guoying Zhao 0001, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Vision-Dialog Navigation by Exploring Cross-Modal MemoryabstractVision-dialog navigation posed as a new holy-grail task in vision-language disciplinary targets at learning an agent endowed with the capability of constant conversation for help with natural language and navigating according to human responses. Besides the common challenges faced in visual language navigation, vision-dialog navigation also requires to handle well with the language intentions of a series of questions about the temporal context from dialogue history and co-reasoning both dialogs and visual scenes. In this paper, we propose the Cross-modal Memory Network (CMN) for remembering and understanding the rich information relevant to historical navigation actions. Our CMN consists of two memory modules, the language memory module (L-mem) and the visual memory module (V-mem). Specifically, L-mem learns latent relationships between the current language interaction and a dialog history by employing a multi-head attention mechanism. V-mem learns to associate the current visual views and the cross-modal memory about the previous navigation actions. The cross-modal memory is generated via a vision-to-language attention and a language-to-vision attention. Benefiting from the collaborative learning of the L-mem and the V-mem, our CMN is able to explore the memory about the decision making of historical navigation actions which is for the current step. Experiments on the CVDN dataset show that our CMN outperforms the previous state-of-the-art model by a significant margin on both seen and unseen environments. Yi Zhu 0004, Fengda Zhu, Zhaohuan Zhan, Bingqian Lin, Jianbin Jiao, Xiaojun Chang, Xiaodan Liang |
CVPR | 5 |
| 2020 | Learning Saliency Propagation for Semi-Supervised Instance SegmentationabstractInstance segmentation is a challenging task for both modeling and annotation. Due to the high annotation cost, modeling becomes more difficult because of the limited amount of supervision. We aim to improve the accuracy of the existing instance segmentation models by utilizing a large amount of detection supervision. We propose ShapeProp, which learns to activate the salient regions within the object detection and propagate the areas to the whole instance through an iterative learnable message passing module. ShapeProp can benefit from more bounding box supervision to locate the instances more accurately and utilize the feature activations from the larger number of instances to achieve more accurate segmentation. We extensively evaluate ShapeProp on three datasets (MS COCO, PASCAL VOC, and BDD100k) with different supervision setups based on both two-stage (Mask R-CNN) and single-stage (RetinaMask) models. The results show our method establishes new states of the art for semi-supervised instance segmentation. Yanzhao Zhou, Xin Wang 0066, Jianbin Jiao, Trevor Darrell, Fisher Yu 0001 |
CVPR | 3 |
| 2020 | Prototype Mixture Models for Few-Shot Semantic Segmentation
Boyu Yang 0002, Chang Liu 0042, Jianbin Jiao, Qixiang Ye |
ECCV (8) | 4 |
| 2020 | Image captioning via semantic element embedding
Xiaodan Zhang 0003, Shengfeng He, Xinhang Song, Rynson W. H. Lau, Jianbin Jiao, Qixiang Ye |
Neurocomputing | 5 |
| 2020 | Spatial Preserved Graph Convolution Networks for Person Re-identificationabstractPerson Re-identification is a very challenging task due to inter-class ambiguity caused by similar appearances, and large intra-class diversity caused by viewpoints, illuminations, and poses. To address these challenges, in this article, a graph convolution network based model for person re-identification is proposed to learn more discriminative feature embeddings, where a graph-structured relationship between person images and person parts are together integrated. Graph convolution networks extract common characteristics of the same person, while pyramid feature embedding exploits parts relations and learns stable representation with each person image. We achieve a very competitive performance respectively on three widely used datasets, indicating that the proposed approach significantly outperforms the baseline methods and achieves the state-of-the-art performance. Zhaoju Li, Zongwei Zhou, Zhenjun Han, Junliang Xing, Jianbin Jiao |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2019 | SIXray: A Large-Scale Security Inspection X-Ray Benchmark for Prohibited Item Discovery in Overlapping ImagesabstractIn this paper, we present a large-scale dataset and establish a baseline for prohibited item discovery in Security Inspection X-ray images. Our dataset, named SIXray, consists of 1,059,231 X-ray images, in which 6 classes of 8,929 prohibited items are manually annotated. It raises a brand new challenge of overlapping image data, meanwhile shares the same properties with existing datasets, including complex yet meaningless contexts and class imbalance. We propose an approach named class-balanced hierarchical refinement (CHR) to deal with these difficulties. CHR assumes that each input image is sampled from a mixture distribution, and that deep networks require an iterative process to infer image contents accurately. To accelerate, we insert reversed connections to different network backbones, delivering high-level visual cues to assist mid-level features. In addition, a class-balanced loss function is designed to maximally alleviate the noise introduced by easy negative samples. We evaluate CHR on SIXray with different ratios of positive/negative samples. Compared to the baselines, CHR enjoys a better ability of discriminating objects especially using mid-level features, which offers the possibility of using a weakly-supervised approach towards accurate object localization. In particular, the advantage of CHR is more significant in the scenarios with fewer positive training samples, which demonstrates its potential application in real-world security inspection. Caijing Miao, Lingxi Xie, Fang Wan 0001, Chi Su, Hongye Liu, Jianbin Jiao, Qixiang Ye |
CVPR | 6 |
| 2019 | C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) is a challenging task when provided with image category supervision but required to simultaneously learn object locations and object detectors. Many WSOD approaches adopt multiple instance learning (MIL) and have non-convex loss functions which are prone to get stuck into local minima (falsely localize object parts) while missing full object extent during training. In this paper, we introduce a continuation optimization method into MIL and thereby creating continuation multiple instance learning (C-MIL), with the intention of alleviating the non-convexity problem in a systematic way. We partition instances into spatially related and class related subsets, and approximate the original loss function with a series of smoothed loss functions defined within the subsets. Optimizing smoothed loss functions prevents the training procedure falling prematurely into local minima and facilitates the discovery of Stable Semantic Extremal Regions (SSERs) which indicate full object extent. On the PASCAL VOC 2007 and 2012 datasets, C-MIL improves the state-of-the-art of weakly supervised object detection and weakly supervised object localization with large margins. Fang Wan 0001, Chang Liu 0042, Wei Ke 0003, Xiangyang Ji, Jianbin Jiao, Qixiang Ye |
CVPR | 5 |
| 2019 | Learning Instance Activation Maps for Weakly Supervised Instance SegmentationabstractDiscriminative region responses residing inside an object instance can be extracted from networks trained with image-level label supervision. However, learning the full extent of pixel-level instance response in a weakly supervised manner remains unexplored. In this work, we tackle this challenging problem by using a novel instance extent filling approach. We first design a process to selectively collect pseudo supervision from noisy segment proposals obtained with previously published techniques. The pseudo supervision is used to learn a differentiable filling module that predicts a class-agnostic activation map for each instance given the image and an incomplete region response. We refer to the above maps as Instance Activation Maps (IAMs), which provide a fine-grained instance-level representation and allow instance masks to be extracted by lightweight CRF. Extensive experiments on the PASCAL VOC12 dataset show that our approach beats the state-of-the-art weakly supervised instance segmentation methods by a significant margin and increases the inference speed by an order of magnitude. Our method also generalizes well across domains and to unseen object categories. Without fine-tuning for the specific tasks, our model trained on VOC12 dataset (20 classes) obtains top performance for weakly supervised object localization on the CUB dataset (200 classes) and achieves competitive results on three widely used salient object detection benchmarks. Yi Zhu 0004, Yanzhao Zhou, Huijuan Xu 0001, Qixiang Ye, David S. Doermann, Jianbin Jiao |
CVPR | 6 |
| 2019 | Selective Sparse Sampling for Fine-Grained Image RecognitionabstractFine-grained recognition poses the unique challenge of capturing subtle inter-class differences under considerable intra-class variances (e.g., beaks for bird species). Conventional approaches crop local regions and learn detailed representation from those regions, but suffer from the fixed number of parts and missing of surrounding context. In this paper, we propose a simple yet effective framework, called Selective Sparse Sampling, to capture diverse and fine-grained details. The framework is implemented using Convolutional Neural Networks, referred to as Selective Sparse Sampling Networks (S3Ns). With image-level supervision, S3Ns collect peaks, i.e., local maximums, from class response maps to estimate informative, receptive fields and learn a set of sparse attention for capturing fine-detailed visual evidence as well as preserving context. The evidence is selectively sampled to extract discriminative and complementary features, which significantly enrich the learned representation and guide the network to discover more subtle cues. Extensive experiments and ablation studies show that the proposed method consistently outperforms the state-of-the-art methods on challenging benchmarks including CUB-200-2011, FGVC-Aircraft, and Stanford Cars. Yao Ding 0006, Yanzhao Zhou, Yi Zhu 0004, Qixiang Ye, Jianbin Jiao |
ICCV | 5 |
| 2019 | DANet: Divergent Activation for Weakly Supervised Object LocalizationabstractWeakly supervised object localization remains a challenge when learning object localization models from image category labels. Optimizing image classification tends to activate object parts and ignore the full object extent, while expanding object parts into full object extent could deteriorate the performance of image classification. In this paper, we propose a divergent activation (DA) approach, and target at learning complementary and discriminative visual patterns for image classification and weakly supervised object localization from the perspective of discrepancy. To this end, we design hierarchical divergent activation (HDA), which leverages the semantic discrepancy to spread feature activation, implicitly. We also propose discrepant divergent activation (DDA), which pursues object extent by learning mutually exclusive visual patterns, explicitly. Deep networks implemented with HDA and DDA, referred to as DANets, diverge and fuse discrepant yet discriminative features for image classification and object localization in an end-to-end manner. Experiments validate that DANets advance the performance of object localization while maintaining high performance of image classification on CUB-200 and ILSVRC datasets. Haolan Xue, Chang Liu 0042, Fang Wan 0001, Jianbin Jiao, Xiangyang Ji, Qixiang Ye |
ICCV | 4 |
| 2019 | Min-Entropy Latent Model for Weakly Supervised Object DetectionabstractWeakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces significant randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy serves as a model to learn object locations and a metric to measure the randomness of object localization during learning. It aims to principally reduce the variance of learned instances and alleviate the ambiguity of detectors. MELM is decomposed into three components including proposal clique partition, object clique discovery, and object localization. MELM is optimized with a recurrent learning algorithm, which leverages continuation optimization to solve the challenging non-convexity problem. Experiments demonstrate that MELM significantly improves the performance of weakly supervised object detection, weakly supervised object localization, and image classification, against the state-of-the-art approaches. Fang Wan 0001, Pengxu Wei, Zhenjun Han, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | High performance person re-identification via a boosting ranking ensemble
Zhaoju Li, Zhenjun Han, Junliang Xing, Qixiang Ye, Xuehui Yu, Jianbin Jiao |
Pattern Recognit. | 6 |
| 2018 | Image-Image Domain Adaptation With Preserved Self-Similarity and Domain-Dissimilarity for Person Re-IdentificationabstractPerson re-identification (re-ID) models trained on one domain often fail to generalize well to another. In our attempt, we present a "learning via translation" framework. In the baseline, we translate the labeled images from source to target domain in an unsupervised manner. We then train re-ID models with the translated images by supervised methods. Yet, being an essential part of this framework, unsupervised image-image translation suffers from the information loss of source-domain labels during translation. Our motivation is two-fold. First, for each image, the discriminative cues contained in its ID label should be maintained after translation. Second, given the fact that two domains have entirely different persons, a translated image should be dissimilar to any of the target IDs. To this end, we propose to preserve two types of unsupervised similarities, 1) self-similarity of an image before and after translation, and 2) domain-dissimilarity of a translated source image and a target image. Both constraints are implemented in the similarity preserving generative adversarial network (SPGAN) which consists of an Siamese network and a CycleGAN. Through domain adaptation experiment, we show that images generated by SPGAN are more suitable for domain adaptation and yield consistent and competitive re-ID accuracy on two large-scale datasets. Weijian Deng, Liang Zheng 0001, Qixiang Ye, Guoliang Kang, Yi Yang 0001, Jianbin Jiao |
CVPR | 6 |
| 2018 | Min-Entropy Latent Model for Weakly Supervised Object DetectionabstractWeakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy is used as a metric to measure the randomness of object localization during learning, as well as serving as a model to learn object locations. It aims to principally reduce the variance of positive instances and alleviate the ambiguity of detectors. MELM is deployed as two sub-models, which respectively discovers and localizes objects by minimizing the global and local entropy. MELM is unified with feature learning and optimized with a recurrent learning algorithm, which progressively transfers the weak supervision to object locations. Experiments demonstrate that MELM significantly improves the performance of weakly supervised detection, weakly supervised localization, and image classification, against the state-of-the-art approaches. Fang Wan 0001, Pengxu Wei, Jianbin Jiao, Zhenjun Han, Qixiang Ye |
CVPR | 3 |
| 2018 | Weakly Supervised Instance Segmentation Using Class Peak ResponseabstractWeakly supervised instance segmentation with image-level labels, instead of expensive pixel-level masks, remains unexplored. In this paper, we tackle this challenging problem by exploiting class peak responses to enable a classification network for instance mask extraction. With image labels supervision only, CNN classifiers in a fully convolutional manner can produce class response maps, which specify classification confidence at each image location. We observed that local maximums, i.e., peaks, in a class response map typically correspond to strong visual cues residing inside each instance. Motivated by this, we first design a process to stimulate peaks to emerge from a class response map. The emerged peaks are then back-propagated and effectively mapped to highly informative regions of each object instance, such as instance boundaries. We refer to the above maps generated from class peak responses as Peak Response Maps (PRMs). PRMs provide a fine-detailed instance-level representation, which allows instance masks to be extracted even with some off-the-shelf methods. To the best of our knowledge, we for the first time report results for the challenging image-level supervised instance segmentation task. Extensive experiments show that our method also boosts weakly supervised pointwise localization as well as semantic segmentation performance, and reports state-of-the-art results on popular benchmarks, including PASCAL VOC 2012 and MS COCO. Yanzhao Zhou, Yi Zhu 0004, Qixiang Ye, Qiang Qiu 0001, Jianbin Jiao |
CVPR | 5 |
| 2017 | SRN: Side-Output Residual Network for Object Symmetry Detection in the WildabstractIn this paper, we establish a baseline for object symmetry detection in complex backgrounds by presenting a new benchmark and an end-to-end deep learning approach, opening up a promising direction for symmetry detection in the wild. The new benchmark, named Sym-PASCAL, spans challenges including object diversity, multi-objects, part-invisibility, and various complex backgrounds that are far beyond those in existing datasets. The proposed symmetry detection approach, named Side-output Residual Network (SRN), leverages output Residual Units (RUs) to fit the errors between the object symmetry ground-truth and the outputs of RUs. By stacking RUs in a deep-to-shallow manner, SRN exploits the flow of errors among multiple scales to ease the problems of fitting complex outputs with limited layers, suppressing the complex backgrounds, and effectively matching object symmetry of different scales. Experimental results validate both the benchmark and its challenging aspects related to real-world images, and the state-of-the-art performance of our symmetry detection approach. The benchmark and the code for SRN are publicly available at https://github.com/KevinKecc/SRN. Wei Ke 0003, Jie Chen 0001, Jianbin Jiao, Guoying Zhao 0001, Qixiang Ye |
CVPR | 3 |
| 2017 | Oriented Response NetworksabstractDeep Convolution Neural Networks (DCNNs) are capable of learning unprecedentedly effective image representations. However, their ability in handling significant local and global image rotations remains limited. In this paper, we propose Active Rotating Filters (ARFs) that actively rotate during convolution and produce feature maps with location and orientation explicitly encoded. An ARF acts as a virtual filter bank containing the filter itself and its multiple unmaterialised rotated versions. During back-propagation, an ARF is collectively updated using errors from all its rotated versions. DCNNs using ARFs, referred to as Oriented Response Networks (ORNs), can produce within-class rotation-invariant deep features while maintaining inter-class discrimination for classification tasks. The oriented response produced by ORNs can also be used for image and object orientation estimation tasks. Over multiple state-of-the-art DCNN architectures, such as VGG, ResNet, and STN, we consistently observe that replacing regular filters with the proposed ARFs leads to significant reduction in the number of network parameters and improvement in classification performance. We report the best results on several commonly used benchmarks. Yanzhao Zhou, Qixiang Ye, Qiang Qiu 0001, Jianbin Jiao |
CVPR | 4 |
| 2017 | A scalable convolutional neural network for task-specified scenarios via knowledge distillationabstractIn this paper, we explore the redundancy in convolutional neural network, which scales with the complexity of vision tasks. Considering that many front-end visual systems are interested in only a limited range of visual targets, the removing of task-specified network redundancy can promote a wide range of potential applications. We propose a task-specified knowledge distillation algorithm to derive a simplified model with pre-set computation cost and minimized accuracy loss, which suits the resource constraint front-end systems well. Experiments on the MNIST and CIFAR10 datasets demonstrate the feasibility of the proposed approach as well as the existence of task-specified redundancy. Qixiang Ye, Zhenjun Han, Jianbin Jiao |
ICASSP | 5 |
| 2017 | Soft Proposal Networks for Weakly Supervised Object LocalizationabstractWeakly supervised object localization remains challenging, where only image labels instead of bounding boxes are available during training. Object proposal is an effective component in localization, but often computationally expensive and incapable of joint optimization with some of the remaining modules. In this paper, to the best of our knowledge, we for the first time integrate weakly supervised object proposal into convolutional neural networks (CNNs) in an end-to-end learning manner. We design a network component, Soft Proposal (SP), to be plugged into any standard convolutional architecture to introduce the nearly cost-free object proposal, orders of magnitude faster than state-of-the-art methods. In the SP-augmented CNNs, referred to as Soft Proposal Networks (SPNs), iteratively evolved object proposals are generated based on the deep feature maps then projected back, and further jointly optimized with network parameters, with image-level supervision only. Through the unified learning process, SPNs learn better object-centric filters, discover more discriminative visual evidence, and suppress background interference, significantly boosting both weakly supervised object localization and classification performance. We report the best results on popular benchmarks, including PASCAL VOC, MS COCO, and ImageNet. Yi Zhu 0004, Yanzhao Zhou, Qixiang Ye, Qiang Qiu 0001, Jianbin Jiao |
ICCV | 5 |
| 2017 | Keyword-driven image captioning via Context-dependent Bilateral LSTMabstractImage captioning has recently received much attention. Existing approaches, however, are limited to describing images with simple contextual information, which typically generate one sentence to describe each image with only a single contextual emphasis. In this paper, we address this limitation from a user perspective with a novel approach. Given some keywords as additional inputs, the proposed method would generate various descriptions according to the provided guidance. Hence, descriptions with different focuses can be generated for the same image. Our method is based on a new Context-dependent Bilateral Long Short-Term Memory (CDB-LSTM) model to predict a keyword-driven sentence by considering the word dependence. The word dependence is explored externally with a bilateral pipeline, and internally with a unified and joint training process. Experiments on the MS COCO dataset demonstrate that the proposed approach not only significantly outperforms the baseline method but also shows good adaptation and consistency with various keywords. Xiaodan Zhang 0003, Shengfeng He, Xinhang Song, Pengxu Wei, Shuqiang Jiang, Qixiang Ye, Jianbin Jiao, Rynson W. H. Lau |
ICME | 7 |
| 2017 | Beyond Group: Multiple Person Tracking via Minimal Topology-Energy-VariationabstractTracking multiple persons is a challenging task when persons move in groups and occlude each other. Existing group-based methods have extensively investigated how to make group division more accurately in a tracking-by-detection framework; however, few of them quantify the group dynamics from the perspective of targets' spatial topology or consider the group in a dynamic view. Inspired by the sociological properties of pedestrians, we propose a novel socio-topology model with a topology-energy function to factor the group dynamics of moving persons and groups. In this model, minimizing the topology-energy-variance in a two-level energy form is expected to produce smooth topology transitions, stable group tracking, and accurate target association. To search for the strong minimum in energy variation, we design the discrete group-tracklet jump moves embedded in the gradient descent method, which ensures that the moves reduce the energy variation of group and trajectory alternately in the varying topology dimension. Experimental results on both RGB and RGB-D data sets show the superiority of our proposed model for multiple person tracking in crowd scenes. Shan Gao 0003, Qixiang Ye, Junliang Xing, Arjan Kuijper, Zhenjun Han, Jianbin Jiao, Xiangyang Ji |
IEEE Trans. Image Process. | 6 |
| 2017 | Correlated Topic Vector for Scene ClassificationabstractScene images usually involve semantic correlations, particularly when considering large-scale image data sets. This paper proposes a novel generative image representation, correlated topic vector, to model such semantic correlations. Oriented from the correlated topic model, correlated topic vector intends to naturally utilize the correlations among topics, which are seldom considered in the conventional feature encoding, e.g., Fisher vector, but do exist in scene images. It is expected that the involvement of correlations can increase the discriminative capability of the learned generative model and consequently improve the recognition accuracy. Incorporated with the Fisher kernel method, correlated topic vector inherits the advantages of Fisher vector. The contributions to the topics of visual words have been further employed by incorporating the Fisher kernel framework to indicate the differences among scenes. Combined with the deep convolutional neural network (CNN) features and Gibbs sampling solution, correlated topic vector shows great potential when processing large-scale and complex scene image data sets. Experiments on two scene image data sets demonstrate that correlated topic vector improves significantly the deep CNN features, and outperforms existing Fisher kernel-based features. Pengxu Wei, Fang Wan 0001, Yi Zhu 0004, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Image Process. | 5 |
| 2016 | Multi-kernel metric learning for person re-identificationabstractIn this paper, we propose a new Multi-kernel Metric Learning (MKML) approach to enhance the performance of person re-identification using adaptive weighted Multi-kernel. The intuition behind our approach is that different features, i.e., low-level and middle-level features, have different nature and thus discriminating capability, utilizing different kernels could map these features into sub-spaces, which helps to improve the discrimination among features. The kernels are combined with an adaptive weighting strategy to get an efficient kernel space. The Fisher Discriminant Analysis (FDA) is used to learn the metric in the learned weighted kernels space that enhances the robustness of metric to discriminate among classes. Experiments on two challenging person reidentification datasets, i.e., VIPeR and CUHK01, demonstrated that our approach is effective. Muhammad Adnan Syed, Jianbin Jiao |
ICIP | 2 |
| 2016 | Collective motion pattern inference via Locally Consistent Latent Dirichlet Allocation
Jialing Zou, Qixiang Ye, Yanting Cui, Fang Wan 0001, Kun Fu 0001, Jianbin Jiao |
Neurocomputing | 6 |
| 2015 | Pedestrian detection via PCA filters based convolutional channel featuresabstractIn this paper, we propose a kind of image representation, named PCA filters based convolutional channel features (PCA-CCF) for pedestrian detection. The motivation is to use the convolutional network architecture with orthogonal PCA filters to enhance the state-of-the-art aggregate channel features (ACF). In PCA-CCF, the convolutional operation improves the feature robustness to pedestrian local deformation. The learned PCA filters reduce the correlations among features of each channel, and therefore, improve feature discrimination capability. With the proposed PCA-CCF features and cascaded AdaBoost classifiers, we develop a coarse-to-fine pedestrian detection approach. Experiments show that such approach achieves 3.04%, 17.87% and 6.28% performance gain on the INRIA, Caltech Reasonable and Caltech Overall pedestrian datasets, respectively. Wei Ke 0003, Pengxu Wei, Qixiang Ye, Jianbin Jiao |
ICASSP | 5 |
| 2015 | Orientation robust object detection in aerial images using deep convolutional neural networkabstractDetecting objects in aerial images is challenged by variance of object colors, aspect ratios, cluttered backgrounds, and in particular, undetermined orientations. In this paper, we propose to use Deep Convolutional Neural Network (DCNN) features from combined layers to perform orientation robust aerial object detection. We explore the inherent characteristics of DC-NN as well as relate the extracted features to the principle of disentangling feature learning. An image segmentation based approach is used to localize ROIs of various aspect ratios, and ROIs are further classified into positives or negatives using an SVM classifier trained on DCNN features. With experiments on two datasets collected from Google Earth, we demonstrate that the proposed aerial object detection approach is simple but effective. Haigang Zhu, Weiqun Dai, Kun Fu 0001, Qixiang Ye, Jianbin Jiao |
ICIP | 6 |
| 2015 | Rich Image Description Based on RegionsabstractAbstract Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In contrast to the previous image description methods that focus on describing the whole image, this paper presents a method of generating rich image descriptions from image regions. First, we detect regions with R-CNN (regions with convolutional neural network features) framework. We then utilize the RNN (recurrent neural networks) to generate sentences for image regions. Finally, we propose an optimization method to select one suitable region. The proposed model generates several sentence description of regions in an image, which has sufficient representative power of the whole image and contains more detailed information. Comparing to general image level description, generating more specific and accurate sentences on the different regions can satisfy more personal requirements for different people. Experimental evaluations validate the effectiveness of the proposed method. Xiaodan Zhang 0003, Xinhang Song, Xiong Lv, Shuqiang Jiang, Qixiang Ye, Jianbin Jiao |
ACM Multimedia | 6 |
| 2015 | Real-Time Multipedestrian Tracking in Traffic Scenes via an RGB-D-Based Layered Graph ModelabstractMultipedestrian tracking in traffic scenes is challenging due to cluttered backgrounds and serious occlusions. In this paper, we propose a layered graph model in image (RGB) and depth (D) domains for real-time robust multipedestrian tracking. The motivation is to investigate high-level constraints in RGB-D data association and to improve the optimization from the trajectory level to the layer level. To construct a layered graph, we define constraints in the depth domain so that pedestrian objects in the image domain are assigned to proper layers. We use pedestrian detection responses in the RGB domain as graph nodes, and we integrate 3-D motion, appearance, and depth features as graph edges. An online updating depth factor is defined to describe the depth relationships among the observations in and out of the layers, and the occlusion issue is processed with an analytical layer-level strategy. With a heuristic label switching algorithm, multiple pedestrian objects are optimally associated and tracked. Experiments and comparison on five public data sets show that our proposed approach significantly reduces pedestrian's ID switch and improves tracking accuracy in the cases of serious occlusions. Shan Gao 0003, Zhenjun Han, Ce Li 0005, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2014 | A cluster specific latent dirichlet allocation model for trajectory clustering in crowded videosabstractTrajectory analysis in crowded video scenes is challenging as trajectories obtained by existing tracking algorithms are often fragmented. In this paper, we propose a new approach to do trajectory inference and clustering on fragmented trajectories, by exploring a cluster specific Latent Dirichlet Allocation(CLDA) model. LDA models are widely used to learn middle level trajectory features and perform trajectory inference. However, they often require scene priors in the learning or inference process. Our cluster specific LDA model addresses this issue by using manifold based clustering as initialization and iterative statistical inference as optimization. The output middle level features of CLDA are input to a clustering algorithm to obtain trajectory clusters. Experiments on a public dataset show the effectiveness of our approach. Jialing Zou, Yanting Cui, Fang Wan 0001, Qixiang Ye, Jianbin Jiao |
ICIP | 5 |
| 2014 | Depth Structure Association for RGB-D Multi-target TrackingabstractMulti-target tracking in outdoor scenes plays an important role in many computer vision applications. Most previous work on visual information based multi-target tracking does not incorporate depth information and the absence of depth information often leads to mismatching or tracking failures. In this paper, we propose a Depth Structure Association (DSA) approach for RGB-D data based multi-target tracking. DSA encodes depth information in a chain structure, the structure is used by DSA together with appearance and motion information to address object occlusion issues in outdoor scenes. Additionally, the use of DSA has the advantages of regulating a much smaller solution space, greatly reducing the computational complexity. Experimental results on three datasets demonstrate that our DSA approach can significantly reduce object mismatch and tracking failure for long term occlusions. Shan Gao 0003, Zhenjun Han, David S. Doermann, Jianbin Jiao |
ICPR | 4 |
| 2014 | Locality-Constrained Sparse Reconstruction for Trajectory ClassificationabstractTrajectory classification has been extensively investigated in recent years, however, problems remain when processing incomplete trajectories of noises and local variations. In this paper, we propose a Locality-constrained Sparse Reconstruction (LSR) approach that explores both sparsity and local adaptability for robust trajectory classification. A trajectory dictionary with locality constrains is constructed with track lets partitioned from collected trajectories by control points of cubic B-spline curves. On the dictionary, the proposed LSR is used to calculate a discriminate code matrix. Then, a loss weighted decoding strategy is employed to perform multi-class trajectory classification. In addition, the approach can be used for anomalous trajectory detection with a thresholding strategy. Experiments on two datasets show that the results of the LSR approach improve the state of the art. Ce Li 0005, Zhenjun Han, Qixiang Ye, Shan Gao 0003, Lijin Pang, Jianbin Jiao |
ICPR | 6 |
| 2014 | A Belief Based Correlated Topic Model for Trajectory Clustering in Crowded Video ScenesabstractTrajectory clustering in crowded video scenes is very challenging. In this paper, we propose to use a belief based correlated topic model (BCTM) to learn discriminative middle level features for trajectory clustering. By constructing a scene prior based joint Gaussian distribution, the BCTM can uncover relations between trajectory clusters and the middle level features using a parameter estimation procedure. The method has distinct advantages over Correlated Topic Model (CTM) and Random Field Topic (RFT) model previously proposed. The inputs to the BCTM are either full trajectories or trajectory fragments obtained with an existing tracking algorithm. The output BCTM features are input to a hierarchical clustering algorithm to obtain trajectory clusters. Experiments on three benchmark datasets show that the proposed BCTM and trajectory clustering approach improves the state of the art. Jialing Zou, Qixiang Ye, Yanting Cui, David S. Doermann, Jianbin Jiao |
ICPR | 5 |
| 2013 | Robust Visual Object Tracking via Sparse Representation and Reconstruction
Zhenjun Han, Qixiang Ye, Jianbin Jiao |
CAIP (2) | 3 |
| 2013 | Minimum Entropy Models for Laser Line Extraction
Wei Ke 0003, Ce Li 0005, Jianbin Jiao |
CAIP (2) | 5 |
| 2013 | Visual abnormal behavior detection based on trajectory sparse reconstruction analysis
Ce Li 0005, Zhenjun Han, Qixiang Ye, Jianbin Jiao |
Neurocomputing | 4 |
| 2013 | Human Detection in Images via Piecewise Linear Support Vector MachinesabstractHuman detection in images is challenged by the view and posture variation problem. In this paper, we propose a piecewise linear support vector machine (PL-SVM) method to tackle this problem. The motivation is to exploit the piecewise discriminative function to construct a nonlinear classification boundary that can discriminate multiview and multiposture human bodies from the backgrounds in a high-dimensional feature space. A PL-SVM training is designed as an iterative procedure of feature space division and linear SVM training, aiming at the margin maximization of local linear SVMs. Each piecewise SVM model is responsible for a subspace, corresponding to a human cluster of a special view or posture. In the PL-SVM, a cascaded detector is proposed with block orientation features and a histogram of oriented gradient features. Extensive experiments show that compared with several recent SVM methods, our method reaches the state of the art in both detection accuracy and computational efficiency, and it performs best when dealing with low-resolution human regions in clutter backgrounds. Qixiang Ye, Zhenjun Han, Jianbin Jiao, Jianzhuang Liu |
IEEE Trans. Image Process. | 3 |
| 2012 | Pedestrian detection via part-based topology modelabstractIn this paper, we propose a part-based topology model and a pedestrian detection method, which obviously improve the detection accuracy. In Our method, pedestrian is divided into several parts. Firstly, histogram of oriented gradients (HOG) features and linear support vector machine (SVM) classifier are used to detect pedestrian parts. Secondly, a novel binary descriptor called log-polar pattern (LPP) is proposed to represent the spatial relation of a part pair. Then multiple LPPs are combined as a log-polar topology pattern (LTP) to model the global topology of a pedestrian. Finally, we put the LTP into One-Class SVM (OC-SVM) to determine whether the detected parts indicate a pedestrian or not. Experiments in INRIA dataset show that our method is robust to occlusion and multi-postures, which obviously reduces the miss rate. Wen Gao 0001, Qixiang Ye, Jianbin Jiao |
ICIP | 4 |
| 2012 | Evaluation of local feature descriptors and their combination for pedestrian representation
Jixiang Liang, Qixiang Ye, Jie Chen 0001, Jianbin Jiao |
ICPR | 4 |
| 2012 | Pedestrian detection in images via cascaded L1-norm minimization learning method
Jianbin Jiao, Baochang Zhang 0001, Qixiang Ye |
Pattern Recognit. | 2 |
| 2012 | Pedestrian Detection in Video Images via Error Correcting Output Code Classification of Manifold SubclassesabstractPedestrian detection in images and video frames is challenged by the view and posture problem. In this paper, we propose a new pedestrian detection approach by error correcting output code (ECOC) classification of manifold subclasses. The motivation is that pedestrians across views and postures form a manifold and that the ECOC method constructs a nonlinear classification boundary that can discriminate the manifold from negative samples. The pedestrian manifold is first constructed with a local linear embedding algorithm and then divided into subclasses with a -means clustering algorithm. The neighboring relationships of these subclasses are used to make the encoding rule for ECOCs, which we use to train multiple base classifiers with histogram of oriented gradient features and linear support vector machines. In the detection procedure, image windows are tested with all base classifiers, and their output codes are fed into an ECOC decoding procedure to decide whether it is a pedestrian or not. Experiments on three data sets show that the results of our approach improve the state of the art. Qixiang Ye, Jixiang Liang, Jianbin Jiao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2011 | Abnormal Behavior Detection via Sparse Reconstruction Analysis of TrajectoryabstractThis paper proposes a new method for abnormal behavior detection in surveillance videos via sparse reconstruction analysis. The motion trajectories of objects are firstly defined as fixed-length parametric vectors based on approximating cubic B-spline curves. Then the vectors are classified as behavior patterns and finally distinguished between normal and abnormal behaviors based on sparse reconstruction analysis, in which a classifier is constructed with sparse linear reconstruction coefficients by computing L1-norm minimization and sparse reconstruction residuals learning from labeled training samples. Experimental results on public dataset show the effectiveness of the proposed approach. Ce Li 0005, Zhenjun Han, Qixiang Ye, Jianbin Jiao |
ICIG | 4 |
| 2011 | Fast Pedestrian Detection with Laser and Image Data FusionabstractIn this paper, we proposed a pedestrian detection system based on laser and image data fusion. The high speed of laser data based location and precise of image based classification are fully explored. First, laser scanner point data is clustered into segments, each of which implies a pedestrian candidate. Then, the segments are projected to the image domain to form regions of interest (ROI) on the image, given camera calibration parameters. Finally two SVM classifiers on Histogram of Oriented Gradient (HOG) features are used to precisely locate pedestrians on the ROI. Experiments report over 30 times higher speed than the state-of-the-art method and a comparable detection rate. Jixiang Liang, Qixiang Ye, Zhenjun Han, Jianbin Jiao |
ICIG | 5 |
| 2011 | A fast object tracking approach based on sparse representationabstractThis paper proposes a new approach based on object sparse representation (OSR) for object tracking. The OSR method implemented by L1-norm minimization is robust to the partial occlusion and deterioration in object images. Firstly, we dynamically construct a set of samples in a predicted searching window in a new video frame, on which the sparse representation of the tracked object can be calculated by the OSR method. This procedure can automatically select the subset of the samples as a basis which most compactly expresses the object with small residuals and rejects all other possible but less compact representations. In terms of this sparse and compact representation, the instantaneous tracking result is achieved in the new video frame. Extensive comparative experiments demonstrate the effectiveness of the proposed approach especially in occlusion context. Zhenjun Han, Jianbin Jiao, Qixiang Ye |
ICIP | 2 |
| 2011 | Nonlinear L1-norm minimization learning for human detectionabstractView, appearance and pose variations make it difficult to detect human objects only by using linear classification methods. Inspired by the successful applications of L1-norm minimization learning (LML) for human detection, we propose a new nonlinear L1-norm minimization learning method (NL-LML). It integrates a nonlinear transformation with an LML optimization model for human detection. The NL-LML method first maps the samples into a space based on the kernel function, and then combines the reformulated samples in the transformed space with the LML model to learn a classifier. Histograms of orientated gradient (HOG) features are used as the feature descriptors, and the sliding window scheme is adopted to detect humans in images. Experiments on two human datasets validate the efficiency and effectiveness of the proposed method. Jianbin Jiao, Qixiang Ye |
ICIP | 2 |
| 2011 | Combined feature evaluation for adaptive visual object tracking
Zhenjun Han, Qixiang Ye, Jianbin Jiao |
Comput. Vis. Image Underst. | 3 |
| 2011 | Visual object tracking via sample-based Adaptive Sparse Representation (AdaSR)
Zhenjun Han, Jianbin Jiao, Baochang Zhang 0001, Qixiang Ye, Jianzhuang Liu |
Pattern Recognit. | 2 |
| 2010 | Cascaded L1-norm Minimization Learning (CLML) classifier for human detectionabstractThis paper proposes a new learning method, which integrates feature selection with classifier construction for human detection via solving three optimization models. Firstly, the method trains a series of weak-classifiers by the proposed L1-norm Minimization Learning (LML) and min-max penalty function models. Secondly, the proposed method selects the weak-classifiers by using the integer optimization model to construct a strong classifier. The L1-norm minimization and integer optimization models aim to find the minimal VC-dimension for weak and strong classifiers respectively. Finally, the method constructs a cascade of LML (CLML) classifier to reach higher detection rates and efficiency. Histograms of Oriented Gradients features of variable-size blocks (v-HOG) are employed as human representation to verify the proposed method. Experiments conducted on INRIA human test set show more superior detection rates and speed than state-of-the-art methods. Baochang Zhang 0001, Qixiang Ye, Jianbin Jiao |
CVPR | 4 |
| 2010 | Human detection in images via L1-norm Minimization LearningabstractIn recent years, sparse representation originating from signal compressed sensing theory has attracted increasing interest in computer vision research community. However, to our best knowledge, no previous work utilizes L1-norm minimization for human detection. In this paper we develop a novel human detection system based on L1-norm Minimization Learning (LML) method. The method is on the observation that a human object can be represented by a few features from a large feature set (sparse representation). And the sparse representation can be learned from the training samples by exploiting the L1-norm Minimization principle, which can also be called feature selection procedure. This procedure enables the feature representation more concise and more adaptive to object occlusion and deformation. After that a classifier is constructed by linearly weighting features and comparing the result with a calculated threshold. Experiments on two datasets validate the effectiveness and efficiency of the proposed method. Baochang Zhang 0001, Qixiang Ye, Jianbin Jiao |
ICASSP | 4 |
| 2010 | Fast pedestrian detection with multi-scale orientation features and two-stage classifiersabstractIn this paper, we propose an approach for fast pedestrian detection in images. Inspired by the histogram of oriented gradient (HOG) features, a set of multi-scale orientation (MSO) features are proposed as the feature representation. The features are extracted on square image blocks of various sizes (called units), containing coarse and fine features in which coarse ones are the unit orientations and fine ones are the pixel orientation histograms of the unit. A cascade of Adaboost is employed to train classifiers on the coarse features, aiming to high detection speed. A greedy searching algorithm is employed to select fine features, which are input into SVMs to train the fine classifiers, aiming to high detection accuracy. Experiments report that our approach obtains state-of-art results with 12.4 times faster than the SVM+HOG method. Qixiang Ye, Jianbin Jiao, Baochang Zhang 0001 |
ICIP | 2 |
| 2009 | A configurable method for multi-style license plate recognition
Jianbin Jiao, Qixiang Ye, Qingming Huang |
Pattern Recognit. | 1 |
| 2008 | Online feature evaluation for object tracking using Kalman FilterabstractAn online feature evaluation method for visual object tracking is put forward in this paper. Firstly, a combined feature set is built using color histogram (HC) bins and gradient orientation histogram (HOG) bins considering the color and contour representation of an object respectively. Then a novel method is proposed to evaluate the features’ weights in a tracking process using Kalman Filter, which is used to comprise the inter-frame predication and single-frame measurement of features’ discriminative power. In this way, we extend the traditional filter framework from modeling motion states to modeling feature evaluation. Experiments show this method can greatly improve the tracking stabilization when objects go across complex backgrounds. Zhenjun Han, Qixiang Ye, Jianbin Jiao |
ICPR | 3 |
| 2007 | Multi-posture Human Detection in Video Frames by Motion Contour Matching
Qixiang Ye, Jianbin Jiao |
ACCV (1) | 2 |
| 2007 | Text detection and restoration in natural scene images
Qixiang Ye, Jianbin Jiao |
J. Vis. Commun. Image Represent. | 2 |