EDBT 2026 Demo / reviewers in the wild / expert
Ming Tang 0001
dblp:73/4373-1
· DBLP profile ↗
119ranked-venue papers
10as first author
65since 2021 · last 2026
0000-0003-4976-3095ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 90 · 7 first-author · 48 since 2021Artificial intelligence and machine learning · 67 · 8 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly DetectionabstractAnomaly detection is a critical task across numerous domains and modalities, yet existing methods are often highly specialized, limiting their generalizability. These specialized models, tailored for specific anomaly types like textural defects or logical errors, typically exhibit limited performance when deployed outside their designated contexts. To overcome this limitation, we propose AnomalyMoE, a novel and universal anomaly detection framework based on a Mixture-of-Experts (MoE) architecture. Our key insight is to decompose the complex anomaly detection problem into three distinct semantic hierarchies: local structural anomalies, component-level semantic anomalies, and global logical anomalies. AnomalyMoE correspondingly employs three dedicated expert networks at the patch, component, and global levels, and is specialized in reconstructing features and identifying deviations at its designated semantic level. This hierarchical design allows a single model to concurrently understand and detect a wide spectrum of anomalies. Furthermore, we introduce an Expert Information Repulsion (EIR) module to promote expert diversity and an Expert Selection Balancing (ESB) module to ensure the comprehensive utilization of all experts. Experiments on 8 challenging datasets spanning industrial imaging, 3D point clouds, medical imaging, video surveillance, and logical anomaly detection demonstrate that AnomalyMoE establishes new state-of-the-art performance, significantly outperforming specialized methods in their respective domains. Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
AAAI | 6 |
| 2026 | Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and DetectionabstractDespite substantial progress in anomaly synthesis, existing diffusion-based and coarse inpainting pipelines commonly suffer from structural deficiencies such as micro-structural discontinuities, limited semantic controllability, and inefficient generation. To overcome these limitations, we introduce ARAS, a language-conditioned, auto-regressive anomaly synthesis approach that precisely injects local, text-specified defects into normal images via token-anchored latent editing. Leveraging a hard-gated auto-regressive operator and a training-free, context-preserving masked sampling kernel, ARAS significantly enhances defect realism, preserves fine-grained material textures, and provides continuous semantic control over synthesized anomalies. Integrated within our Quality-Aware Re-weighted Anomaly Detection (QARAD) framework, we propose a dynamic weighting strategy that emphasizes high-quality synthetic samples by computing an image-text similarity score with a dual-encoder model. Extensive experiments across three datasets, MVTec AD, VisA, and BTAD, demonstrate that our QARAD outperforms SOTA methods in both image- and pixel-level anomaly detection tasks, achieving improved accuracy, robustness, and a 5× synthesis speedup compared to diffusion-based alternatives. Bingke Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
AAAI | 4 |
| 2026 | Improving Generalization in LLM Structured Pruning via Function-Aware Neuron Grouping
Tao Yu 0013, Yongqi An, Kuan Zhu, Guibo Zhu, Ming Tang 0001, Jinqiao Wang |
AAAI | 5 |
| 2026 | GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) have demonstrated impressive progress in single-image grounding and general multi-image understanding. Recently, some methods begin to address multi-image grounding. However, they are constrained by single-target localization and limited types of practical tasks, due to the lack of unified modeling for generalized grounding tasks. Therefore, we propose GeM-VG, an MLLM capable of Generalized Multi-image Visual Grounding. To support this, we systematically categorize and organize existing multi-image grounding tasks according to cognitive demands and introduce the MG-Data-240K dataset, addressing the limitations of existing datasets regarding target quantity and image relation. To tackle the challenges of robustly handling diverse multi-image grounding tasks, we further propose a hybrid reinforcement finetuning strategy that integrates chain-of-thought (CoT) reasoning and direct answering, considering their complementary strengths. This strategy adopts an R1-like algorithm guided by a carefully designed rule-based reward, effectively enhancing the model’s overall perception and reasoning capabilities. Extensive experiments demonstrate the superior generalized grounding capabilities of our model. For multi-image grounding, it outperforms the previous leading MLLMs by 2.0% and 9.7% on MIG-Bench and MC-Bench, respectively. In single-image grounding, it achieves a 9.1% improvement over the base model on ODINW. Furthermore, our model retains strong capabilities in general multi-image understanding. Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang 0089, Yufei Zhan, Ming Tang 0001, Jinqiao Wang |
AAAI | 6 |
| 2026 | HB-Mamba: Hierarchical Bi-directional State Space Modeling for LiDAR Semantic Segmentation in Autonomous Drivingabstract3D semantic segmentation remains a pivotal challenge for autonomous driving due to the inherent sparsity of points. Existing CNN-based and Transformer-based methods struggle with either limited receptive fields or quadratic computational complexity. Although some Mamba-based 3D models are designed efficiently with linear complexity, they often overlook the long-term decay problem in Selective State-space Models when processing extremely long sequences in large-scale scenes. In this paper, we propose a Hierarchical Bi-directional Mamba (HB-Mamba) for point cloud semantic segmentation. By decoupling feature extraction into a Global Memory branch and a Local Detail branch, our architecture effectively captures long-range semantics and preserves fine-grained geometric information. Besides, we further introduce a Spatial-Channel Fusion Block to dynamically fuse these multi-scale representations. Experimental results on the nuScenes-Lidarseg benchmark demonstrate that HB-Mamba achieves state-of-the-art performance among Lidar-only methods, reaching 82.8% mIoU on the test set and 81.33% mIoU on the validation set, outperforming the leading transformer-based model PTv3 by 0.1% and 1.01%, respectively. Wei Li 0315, Haiyun Guo, Manli Tao, Honghui Dong, Ming Tang 0001, Jinqiao Wang |
ICMR | 5 |
| 2026 | Seg-LLaVA: Empowering pixel-level understanding with large vision language model
Fan Yang 0089, Yousong Zhu, Yufei Zhan, Hongyin Zhao, Xin Li 0034, Yaowei Wang 0001, Ming Tang 0001, Jinqiao Wang |
Pattern Recognit. | 7 |
| 2026 | FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable LocalizationabstractAnomaly detection methods typically require extensive normal samples from the target class for training, limiting their applicability in scenarios that require rapid adaptation, such as cold start. Zero-shot and few-shot anomaly detection do not require labeled samples from the target class in advance, making them a promising research direction. Existing zero-shot and few-shot approaches often leverage powerful multimodal models to detect and localize anomalies by comparing image-text similarity. However, their handcrafted generic descriptions fail to capture the diverse range of anomalies that may emerge in different objects, and simple patch-level image-text matching often struggles to localize anomalous regions of varying shapes and sizes. To address these issues, this paper proposes the FiLo++ method, which consists of two key components. The first component, Fused Fine-Grained Descriptions (FusDes), utilizes large language models to generate anomaly descriptions for each object category, combines both fixed and learnable prompt templates and applies a runtime prompt filtering method, producing more accurate and task-specific textual descriptions. The second component, Deformable Localization (DefLoc), integrates the vision foundation model Grounding DINO with position-enhanced text descriptions and a Multi-scale Deformable Cross-modal Interaction (MDCI) module, enabling accurate localization of anomalies with various shapes and sizes. In addition, we design a position-enhanced patch matching approach to improve few-shot anomaly detection performance. Experiments on multiple datasets demonstrate that FiLo++ achieves significant performance improvements compared with existing methods. Code will be available at https://github.com/CASIA-IVA-Lab/FiLo. Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Continual Instruction Tuning for Large Multimodal ModelsabstractInstruction tuning has become a widely adopted approach for aligning large multimodal models (LMMs) with human intent. It enables multi-task joint training through unified data formats. However, as new vision-language tasks constantly emerge, exhaustive joint training of all tasks becomes impractical. Continual learning offers a more flexible and resource-efficient alternative, enabling incremental training of LMMs on emerging tasks. This study investigates two fundamental questions when applying continual learning to instruction tuning of LMMs: 1) Do LMMs suffer from catastrophic forgetting during continual instruction tuning? 2) Can existing continual learning methods be effectively applied to continual instruction tuning of LMMs? A comprehensive study was conducted to answer these questions. First, we establish the first benchmark for continual instruction tuning of LMMs and reveal the phenomenon of catastrophic forgetting in this setup. Second, we integrate and adapt traditional continual learning approaches to this setting, demonstrating the effectiveness of these strategies to varying degrees in different scenarios. Third, we explore task-similarity dynamics between pairs of vision-language tasks and propose task-similarity-informed regularization and model expansion methods. Experimental results show that our approach can consistently boost the model's performance. Jinghan He, Haiyun Guo, Kuan Zhu, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Image Process. | 4 |
| 2025 | Cracking the Code of Hallucination in LVLMs with Vision-aware Head DivergenceabstractJinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang, Tat-Seng Chua, Jinqiao Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang 0001, Tat-Seng Chua, Jinqiao Wang |
ACL (1) | 7 |
| 2025 | UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly DetectionabstractVisual Anomaly Detection (VAD) aims to identify abnormal samples in images that deviate from normal patterns, covering multiple domains, including industrial, logical, and medical fields. Due to the domain gaps between these fields, existing VAD methods are typically tailored to each domain, with specialized detection techniques and model architectures that are difficult to generalize across different domains. Moreover, even within the same domain, current VAD approaches often require large amounts of normal samples to train class-specific models, resulting in poor generalizability and hindering unified evaluation across domains. To address this issue, we propose a generalized few-shot VAD method, UniVAD, capable of detecting anomalies across various domains, with a training-free unified model. UniVAD only needs few normal samples as references during testing to detect anomalies in previously unseen objects, without training on the specific domain. Specifically, UniVAD employs a Contextual Component Clustering (C3) module based on clustering and vision foundation models to segment components within the image accurately, and leverages Component-Aware Patch Matching (CAPM) and Graph-Enhanced Component Modeling (GECM) modules to detect anomalies at different semantic levels, which are aggregated to produce the final detection result. We conduct experiments on nine datasets spanning industrial, logical, and medical fields, and the results demonstrate that UniVAD achieves state-of-the-art performance in few-shot anomaly detection tasks across multiple domains, outperforming domain-specific anomaly detection models. Code is available at https://github.com/FantasticGNU/UniVAD. Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
CVPR | 5 |
| 2025 | PhysVLM: Enabling Visual Language Models to Understand Robotic Physical ReachabilityabstractUnderstanding the environment and a robot’s physical reachability is crucial for task execution. While state-of-the-art vision-language models (VLMs) excel in environmental perception, they often generate inaccurate or impractical responses in embodied visual reasoning tasks due to a lack of understanding of robotic physical reachability. To address this issue, we propose a unified representation of physical reachability across diverse robots, i.e., Space-Physical Reachability Map (S-P Map), and PhysVLM, a vision-language model that integrates this reachability information into visual reasoning. Specifically, the S-P Map abstracts a robot’s physical reachability into a generalized spatial representation, independent of specific robot configurations, allowing the model to focus on reachability features rather than robot-specific parameters. Subsequently, PhysVLM extends traditional VLM architectures by incorporating an additional feature encoder to process the S-P Map, enabling the model to reason about physical reachability without compromising its general vision-language capabilities. To train and evaluate PhysVLM, we constructed a large-scale multi-robot dataset, Phys100K, and a challenging benchmark, EQA-phys, which includes tasks for six different robots in both simulated and real-world environments. Experimental results demonstrate that PhysVLM outperforms existing models, achieving a 14% improvement over GPT-4o on EQA-phys and surpassing advanced embodied VLMs such as RoboMamba and SpatialVLM on the RoboVQA-val and OpenEQA benchmarks. Additionally, the S-P Map shows strong compatibility with various VLMs, and its integration into GPT-4o-mini yields a 7.1% performance improvement. Manli Tao, Chaoyang Zhao, Haiyun Guo, Honghui Dong, Ming Tang 0001, Jinqiao Wang |
CVPR | 6 |
| 2025 | Extracting Sparse Specialist Models from Generalist ModelsabstractRecently, several generalist models such as Contrastive Language Image Pre-training (CLIP) have demonstrated their capabilities of performing diverse downstream tasks through zero-shot or few-shot guidance. When these generalist models are used for the specific downstream task where only a fraction of features is relevant, they would suffer from a significant redundancy of parameters. While existing methods aim to achieve sparsity and specialization, they often require additional training and large datasets. In this paper, we propose a novel framework to extract a sparse specialist model from a generalist model using only few-shot samples, without any training. Our task-specific pruning framework defines task relevance metrics and employs weighted layer-wise pruning, preserving relevant features while removing redundancies. Experiments show that our method maintains nearly identical zero-shot accuracy compared to the original generalist models at 30% sparsity, with only minimal decline at 50%. Tao Yu 0013, Xu Zhao 0003, Yongqi An, Guibo Zhu, Ming Tang 0001, Jinqiao Wang |
ICASSP | 5 |
| 2025 | MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video ParsingabstractThe weakly-supervised audio-visual video parsing (AVVP) aims to predict all modality-specific events and locate their temporal boundaries. Despite significant progress, due to the limitations of the weakly-supervised and the deficiencies of the model architecture, existing methods are lacking in simultaneously improving both the segment-level prediction and the event-level prediction. In this work, we propose a audio-visual Mamba network with pseudo labeling aUGmentation (MUG) for emphasising the uniqueness of each segment and excluding the noise interference from the alternate modalities. Specifically, we annotate some of the pseudo-labels based on previous work. Using unimodal pseudo-labels, we perform cross-modal random combinations to generate new data, which can enhance the model's ability to parse various segment-level event combinations. For feature processing and interaction, we employ a audio-visual mamba network. The AV-Mamba enhances the ability to perceive different segments and excludes additional modal noise while sharing similar modal information. Our extensive experiments demonstrate that MUG improves state-of-the-art results on LLP dataset in all metrics (e.g,, gains of 2.1% and 1.2% in terms of visual Segment-level and audio Segment-level metrics). Our code is available at https://github.com/WangLY136/MUG. Langyu Wang, Bingke Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
ICCV | 5 |
| 2025 | Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-ReferringabstractLarge Vision Language Models have achieved fine-grained object perception, but the limitation of image resolution remains a significant obstacle to surpassing the performance of task-specific experts in complex and dense scenarios. Such limitation further restricts the model's potential to achieve nuanced visual and language referring in domains such as GUI Agents, counting, \textit{etc}. To address this issue, we introduce a unified high-resolution generalist model, Griffon v2, enabling flexible object referring with visual and textual prompts. To efficiently scale up image resolution, we design a simple and lightweight down-sampling projector to overcome the input tokens constraint in Large Language Models. This design inherently preserves the complete contexts and fine details and significantly improves multimodal perception ability, especially for small objects. Building upon this, we further equip the model with visual-language co-referring capabilities through a plug-and-play visual tokenizer. It enables user-friendly interaction with flexible target images, free-form texts, and even coordinates. Experiments demonstrate that Griffon v2 can localize objects of interest with visual and textual referring, achieve state-of-the-art performance on REC and phrase grounding, and outperform expert models in object detection, object counting, and REG. Data and codes are released at https://github.com/jefferyZhan/Griffon. Yufei Zhan, Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang 0089, Ming Tang 0001, Jinqiao Wang |
ICCV | 6 |
| 2025 | Systematic Outliers in Large Language ModelsabstractOutliers have been widely observed in Large Language Models (LLMs), significantly impacting model performance and posing challenges for model compression. Understanding the functionality and formation mechanisms of these outliers is critically important. Existing works, however, largely focus on reducing the impact of outliers from an algorithmic perspective, lacking an in-depth investigation into their causes and roles. In this work, we provide a detailed analysis of the formation process, underlying causes, and functions of outliers in LLMs. We define and categorize three types of outliers—activation outliers, weight outliers, and attention outliers—and analyze their distributions across different dimensions, uncovering inherent connections between their occurrences and their ultimate influence on the attention mechanism. Based on these observations, we hypothesize and explore the mechanisms by which these outliers arise and function, demonstrating through theoretical derivations and experiments that they emerge due to the self-attention mechanism's softmax operation. These outliers act as implicit context-aware scaling factors within the attention mechanism. As these outliers stem from systematic influences, we term them systematic outliers. Our study not only enhances the understanding of Transformer-based LLMs but also shows that structurally eliminating outliers can accelerate convergence and improve model compression. The code is avilable at \url{https://github.com/an-yongqi/systematic-outliers}. Yongqi An, Xu Zhao 0003, Tao Yu 0013, Ming Tang 0001, Jinqiao Wang |
ICLR | 4 |
| 2025 | FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes RecognitionabstractPedestrian attribute recognition (PAR) is a fundamental perception task in intelligent transportation and security. To tackle this fine-grained task, most existing methods focus on extracting regional features to enrich attribute information. However, a regional feature is typically used to predict a fixed set of pre-defined attributes in these methods, which limits the performance and practicality in two aspects: 1) Regional features may compromise fine-grained patterns unique to certain attributes in favor of capturing common characteristics shared across attributes. 2) Regional features cannot generalize to predict unseen attributes in the test time. In this paper, we propose the Fine-grained Optimization with semantiC gUided underStanding (FOCUS) approach for PAR, which adaptively extracts fine-grained attribute-level features for each attribute individually, regardless of whether the attributes are seen or not during training. Specifically, we propose the Multi-Granularity Mix Tokens (MGMT) to capture latent features at varying levels of visual granularity, thereby enriching the diversity of the extracted information. Next, we introduce the Attribute-guided Visual Feature Extraction (AVFE) module, which leverages textual attributes as queries to retrieve their corresponding visual attribute features from the Mix Tokens using a cross-attention mechanism. To ensure that textual attributes focus on the appropriate Mix Tokens, we further incorporate a Region-Aware Contrastive Learning (RACL) method, encouraging attributes within the same region to share consistent attention maps. Extensive experiments on PA100K, PETA, and RAPv1 datasets demonstrate the effectiveness and strong generalization ability of our method. Hongyan An, Kuan Zhu, Haiyun Guo, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
ICME | 6 |
| 2025 | FLARE: A Framework for Stellar Flare Forecasting Using Stellar Physical Properties and Historical RecordsabstractStellar flare events are critical observational samples for astronomical research; however, recorded flare events remain limited. Stellar flare forecasting can provide additional flare event samples to support research efforts. Despite this potential, no specialized models for stellar flare forecasting have been proposed to date. In this paper, we present extensive experimental evidence demonstrating that both stellar physical properties and historical flare records are valuable inputs for flare forecasting tasks. We then introduce FLARE (Forecasting Light-curve-based Astronomical Records via features Ensemble), the first-of-its-kind large model specifically designed for stellar flare forecasting. FLARE integrates stellar physical properties and historical flare records through a novel Soft Prompt Module and Residual Record Fusion Module. Experiments on the Kepler light curve dataset demonstrate that FLARE achieves superior performance compared to other methods across all evaluation metrics. Finally, we validate the forecast capability of our model through a comprehensive case study. Bingke Zhu, Minghui Jia, Yihan Tao, A-Li Luo, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
IJCAI | 8 |
| 2025 | LightPlanner: Unleashing the Reasoning Capabilities of Lightweight Large Language Models in Task PlanningabstractIn recent years, lightweight large language models (LLMs) have garnered significant attention in the robotics field due to their low computational resource requirements and suitability for edge deployment. However, in task planning—particularly for complex tasks that involve dynamic semantic logic reasoning—lightweight LLMs have underperformed. To address this limitation, we propose a novel task planner, LightPlanner, which enhances the performance of lightweight LLMs in complex task planning by fully leveraging their reasoning capabilities. Unlike conventional planners that use fixed skill templates, LightPlanner controls robot actions via parameterized function calls, dynamically generating parameter values. This approach allows for fine-grained skill control and improves task planning success rates in complex scenarios. Furthermore, we introduce hierarchical deep reasoning. Before generating each action decision step, LightPlanner thoroughly considers three levels: action execution (feedback verification), semantic parsing (goal consistency verification), and parameter generation (parameter validity verification). This ensures the correctness of subsequent action controls. Additionally, we incorporate a memory module to store historical actions, thereby reducing context length and enhancing planning efficiency for long-term tasks. We train the LightPlanner-1.5B model on our LightPlan-40k dataset, which comprises 40,000 action controls across tasks with 2 to 13 action steps. Experiments demonstrate that our model achieves the highest task success rate despite having the smallest number of parameters. In tasks involving spatial semantic reasoning, the success rate exceeds that of ReAct by 14.9%. Moreover, we demonstrate LightPlanner’s potential to operate on edge devices. Manli Tao, Chaoyang Zhao, Honghui Dong, Ming Tang 0001, Jinqiao Wang |
IROS | 5 |
| 2025 | Referring Expression Instance Retrieval and A Strong End-to-End BaselineabstractText-Image Retrieval (TIR) retrieves a target image from a gallery based on an image-level description, while Referring Expression Comprehension (REC) localizes a target object within a given image using an instance-level description. However, real-world applications often present more complex demands. Users typically query an instance-level description across a large gallery and expect to receive both relevant image and the corresponding instance location. In such scenarios, TIR struggles with fine-grained descriptions and object-level localization, while REC is limited in its ability to efficiently search large galleries and lacks an effective ranking mechanism. In this paper, we introduce a new task called Referring Expression Instance Retrieval (REIR), which supports both instance-level retrieval and localization based on fine-grained referring expressions. First, we propose a large-scale benchmark for REIR, named REIRCOCO, constructed by prompting advanced vision-language models to generate high quality referring expressions for instances in the MSCOCO and RefCOCO datasets. Second, we present a baseline method, Contrastive Language Instance Alignment with Relation Experts (CLARE), which employs a dual-stream architecture to address REIR in an end-to-end manner. Given a referring expression, the textual branch encodes it into a query embedding, enhanced by a Mix of Relation Experts (MORE) module designed to better capture inter-instance relationships. The visual branch detects candidate objects and extracts their instance-level visual features. The most similar candidate to the query is selected for bounding box prediction. CLARE is first trained on object detection and REC datasets to establish initial grounding capabilities, then optimized via Contrastive Language Instance Alignment (CLIA) for improved retrieval across images. Experimental results demonstrate that CLARE outperforms existing methods on the REIR benchmark and generalizes well to both TIR and REC tasks, showcasing its effectiveness and versatility. Xiangzhao Hao, Kuan Zhu, Haiyun Guo, Ming Tang 0001, Jinqiao Wang |
ACM Multimedia | 7 |
| 2025 | FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential SegmentationabstractRecent Large Vision Language Models (LVLMs) demonstrate promising capabilities in unifying visual understanding and generative modeling, enabling both accurate content understanding and flexible editing. However, current approaches treat \textbf{\textit{"what to see"}} and \textbf{\textit{"how to edit"}} separately: they either perform isolated object segmentation or utilize segmentation masks merely as conditional prompts for local edit generation tasks, often relying on multiple disjointed models. To bridge these gaps, we introduce FOCUS, a unified LVLM that integrates segmentation-aware perception and controllable object-centric generation within an end-to-end framework. FOCUS employs a dual-branch visual encoder to simultaneously capture global semantic context and fine-grained spatial details. In addition, we leverage a MoVQGAN-based visual tokenizer to produce discrete visual tokens that enhance generation quality. To enable accurate and controllable image editing, we propose a progressive multi-stage training pipeline, where segmentation masks are jointly optimized and used as spatial condition prompts to guide the diffusion decoder. This strategy aligns visual encoding, segmentation, and generation modules, effectively bridging segmentation-aware perception with fine-grained visual synthesis.
Extensive experiments across three core tasks, including multimodal understanding, referring segmentation accuracy, and controllable image generation, demonstrate that FOCUS achieves strong performance by jointly optimizing visual perception and generative capabilities. Fan Yang 0089, Yousong Zhu, Xin Li 0034, Yufei Zhan, Hongyin Zhao, Shurong Zheng, Yaowei Wang 0001, Ming Tang 0001, Jinqiao Wang |
NeurIPS | 8 |
| 2025 | PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical EnvironmentsabstractVisual reasoning in multimodal large language models (MLLMs) has primarily been studied in passive, static settings, limiting their effectiveness in real-world physical environments where an embodied agent must contend with incomplete information due to occlusion or a limited field of view. Humans, in contrast, leverage their embodiment to actively explore and interact with their environment—moving, examining, and manipulating objects—to gather information through a closed-loop process integrating perception, reasoning, and action. Inspired by this capability, we introduce the Active Visual Reasoning (AVR) task, extending visual reasoning to a paradigm of embodied interaction in partially observable environments. AVR necessitates embodied agents to: (1) actively acquire information via sequential physical actions, (2) integrate observations across multiple steps for coherent reasoning, and (3) dynamically adjust decisions based on evolving visual feedback. To rigorously evaluate AVR, we introduce CLEVR-AVR, a simulation benchmark featuring multi-round interactive environments designed to assess both reasoning correctness and information-gathering efficiency. We present AVR-152k, a large-scale dataset that offers rich Chain-of-Thought (CoT) annotations detailing iterative reasoning for uncertainty identification, action-conditioned information gain prediction, and information-maximizing action selection, crucial for training agents in a higher-order Markov Decision Process. Building on this, we develop PhysVLM-AVR, an embodied MLLM achieving state-of-the-art performance on CLEVR-AVR, embodied reasoning (OpenEQA, RoboVQA), and passive visual reasoning (GeoMath, Geometry30K). Our analysis also reveals that current embodied MLLMs, despite detecting information incompleteness, struggle to actively acquire and integrate new information through interaction, highlighting a fundamental gap in active reasoning capabilities. Xuantang Xiong, Manli Tao, Chaoyang Zhao, Honghui Dong, Ming Tang 0001, Jinqiao Wang |
NeurIPS | 7 |
| 2025 | AMITA: Attribute-Guided Masked Image-Text Alignment for Multi-Label Image RepresentationabstractMulti-label image classification, which involves recognizing multiple objects within a single image, is a fundamental task in computer vision. Recently, Visual-Language Models (VLMs) have made remarkable progress in this area. Many approaches combine textual and visual modalities to understand the entire image. In this paper, we find that there is a direct correlation between the accurate localization of objects and the accuracy of multi-label classification. However, previous research methods did not specifically address localization accuracy, resulting in sub-optimal accuracy. Therefore, we propose the AMITA, namely Attribute-guided Masked Image-Text Alignment for multi-label image representation. AMITA improves localization accuracy by segmenting object masks, thereby enhancing the accuracy of multi-label image classification. Additionally, AMITA introduces an AutoFocus method to handle the localization problem of small objects. AutoFocus conducts recognition by resizing and cropping the image respectively, and automatically selects the images useful for the classification target. Moreover, AMITA incorporates Attribute-guided Prompting to strengthen the semantic distinction among different categories. It uses large language models to obtain the attributes of different categories and carefully designs prompts to enhance the attribute differences among different categories. Finally, extensive experiments on three popular datasets, including MS-COCO, Pascal VOC 2007, and NUS-WIDE, demonstrate the superiority of AMITA. Jinyi Fang, Bingke Zhu, Jingling Yuan, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language ModelsabstractVision-language models (VLMs), such as CLIP, play a foundational role in various cross-modal applications. To fully leverage the potential of VLMs in adapting to downstream tasks, context optimization methods such as prompt tuning are essential. However, one key limitation is the lack of diversity in prompt templates, whether they are hand-crafted or learned through additional modules. This limitation restricts the capabilities of pretrained VLMs and can result in incorrect predictions in downstream tasks. To address this challenge, we propose context optimization with multi-knowledge representation (CoKnow), a framework that enhances prompt learning for VLMs with rich contextual knowledge. To facilitate CoKnow during inference, we train lightweight semantic knowledge mappers, which are capable of generating multi-knowledge representations for an input image without requiring additional priors. Experimentally, we conduct extensive experiments on 11 publicly available datasets, demonstrating that CoKnow outperforms a series of previous methods. Enming Zhang, Bingke Zhu, Yingying Chen 0003, Qinghai Miao, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Multim. | 5 |
| 2024 | Fluctuation-Based Adaptive Structured Pruning for Large Language ModelsabstractNetwork Pruning is a promising way to address the huge computing resource demands of the deployment and inference of Large Language Models (LLMs). Retraining-free is important for LLMs' pruning methods. However, almost all of the existing retraining-free pruning approaches for LLMs focus on unstructured pruning, which requires specific hardware support for acceleration. In this paper, we propose a novel retraining-free structured pruning framework for LLMs, named FLAP (FLuctuation-based Adaptive Structured Pruning). It is hardware-friendly by effectively reducing storage and enhancing inference speed. For effective structured pruning of LLMs, we highlight three critical elements that demand the utmost attention: formulating structured importance metrics, adaptively searching the global compressed model, and implementing compensation mechanisms to mitigate performance loss. First, FLAP determines whether the output feature map is easily recoverable when a column of weight is removed, based on the fluctuation pruning metric. Then it standardizes the importance scores to adaptively determine the global compressed model structure. At last, FLAP adds additional bias terms to recover the output feature maps using the baseline values. We thoroughly evaluate our approach on a variety of language benchmarks. Without any retraining, our method significantly outperforms the state-of-the-art methods, including LLM-Pruner and the extension of Wanda in structured pruning. The code is released at https://github.com/CASIA-IVA-Lab/FLAP. Yongqi An, Xu Zhao 0003, Tao Yu 0013, Ming Tang 0001, Jinqiao Wang |
AAAI | 4 |
| 2024 | AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) such as MiniGPT-4 and LLaVA have demonstrated the capability of understanding images and achieved remarkable performance in various visual tasks. Despite their strong abilities in recognizing common objects due to extensive training datasets, they lack specific domain knowledge and have a weaker understanding of localized details within objects, which hinders their effectiveness in the Industrial Anomaly Detection (IAD) task. On the other hand, most existing IAD methods only provide anomaly scores and necessitate the manual setting of thresholds to distinguish between normal and abnormal samples, which restricts their practical implementation. In this paper, we explore the utilization of LVLM to address the IAD problem and propose AnomalyGPT, a novel IAD approach based on LVLM. We generate training data by simulating anomalous images and producing corresponding textual descriptions for each image. We also employ an image decoder to provide fine-grained semantic and design a prompt learner to fine-tune the LVLM using prompt embeddings. Our AnomalyGPT eliminates the need for manual threshold adjustments, thus directly assesses the presence and locations of anomalies. Additionally, AnomalyGPT supports multi-turn dialogues and exhibits impressive few-shot in-context learning capabilities. With only one normal shot, AnomalyGPT achieves the state-of-the-art performance with an accuracy of 86.1%, an image-level AUC of 94.1%, and a pixel-level AUC of 95.3% on the MVTec-AD dataset. Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
AAAI | 5 |
| 2024 | Knowledge Distillation Dealing with Sample-Wise Long-Tail Problem
Tao Yu 0013, Xu Zhao 0003, Yongqi An, Ming Tang 0001, Jinqiao Wang |
ACCV (10) | 4 |
| 2024 | Self-Supervised Representation Learning from Arbitrary ScenariosabstractCurrent self-supervised methods can primarily be categorized into contrastive learning and masked image modeling. Extensive studies have demonstrated that combining these two approaches can achieve state-of-the-art performance. However, these methods essentially reinforce the global consistency of contrastive learning without taking into account the conflicts between these two approaches, which hinders their generalizability to arbitrary scenarios. In this paper, we theoretically prove that MAE serves as a patch-level contrastive learning, where each patch within an image is considered as a distinct category. This presents a significant conflict with global-level contrastive learning, which treats all patches in an image as an identical category. To address this conflict, this work abandons the non-generalizable global-level constraints and proposes explicit patch-level contrastive learning as a solution. Specifically, this work employs the encoder of MAE to generate dual-branch features, which then perform patch-level learning through a decoder. In contrast to global-level data aug-mentation in contrastive learning, our approach leverages patch-level feature augmentation to mitigate interference from global-level learning. Consequently, our approach can learn heterogeneous representations from a single image while avoiding the conflicts encountered by previous methods. Massive experiments affirm the potential of our method for learning from arbitrary scenarios. Zhaowen Li, Yousong Zhu, Zhiyang Chen 0002, Zongxin Gao, Rui Zhao 0001, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
CVPR | 7 |
| 2024 | Griffon: Spelling Out All Object Locations at Any Granularity with Large Language Models
Yufei Zhan, Yousong Zhu, Zhiyang Chen 0002, Fan Yang 0089, Ming Tang 0001, Jinqiao Wang |
ECCV (42) | 5 |
| 2024 | SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language ModelsabstractContinual learning (CL) is crucial for language models to dynamically adapt to the evolving real-world demands.To mitigate the catastrophic forgetting problem in CL, data replay has been proven a simple and effective strategy, and the subsequent data-replay-based distillation can further enhance the performance.However, existing methods fail to fully exploit the knowledge embedded in models from previous tasks, resulting in the need for a relatively large number of replay samples to achieve good results.In this work, we first explore and emphasize the importance of attention weights in knowledge retention, and then propose a SElective attEntion-guided Knowledge Retention method (SEEKR) for data-efficient replay-based continual learning of large language models (LLMs).Specifically, SEEKR performs attention distillation on the selected attention heads for finer-grained knowledge retention, where the proposed forgettabilitybased and task-sensitivity-based measures are used to identify the most valuable attention heads.Experimental results on two continual learning benchmarks for LLMs demonstrate the superiority of SEEKR over the existing methods on both performance and efficiency.Explicitly, SEEKR achieves comparable or even better performance with only 1/10 of the replayed data used by other methods, and reduces the proportion of replayed data to 1%.The code is available at https: //github.com/jinghan1he/SEEKR. Jinghan He, Haiyun Guo, Kuan Zhu, Ming Tang 0001, Jinqiao Wang |
EMNLP | 5 |
| 2024 | The Devil is in Details: Delving Into Lite FFN Design for Vision TransformersabstractTransformer has demonstrated exceptional performance on a variety of vision tasks. However, its high computational complexity can become problematic. In this paper, we conduct a systematic analysis of the complexity of each component in vision transformers, and identify an easily overlooked detail: that the Feed-Forward Network (FFN) is the primary computational bottleneck, even more so than the Multi-Head Self-Attention (MHSA) mechanism. Inspired by this, we further propose a lightweight FFN module, named SparseFFN, that can reduce dense computations in both channel and spatial dimension. Specifically, SparseFFN consists of two components: Channel-Sparse FFN (CS-FFN) and Spatial-Sparse FFN (SS-FFN), which can be seamlessly incorporated into various vision transformers and even pure MLP models with significantly fewer FLOPs. Extensive experiments demonstrate the effectiveness and efficiency of the proposed method. For example, our approach can reduce model complexity by 23%-39% for most of vision transformers and MLP models while keeping comparable accuracy. Zhiyang Chen 0002, Yousong Zhu, Zhaowen Li, Fan Yang 0089, Chaoyang Zhao, Jinqiao Wang, Ming Tang 0001 |
ICASSP | 7 |
| 2024 | BFRFormer: Transformer-Based Generator for Real-World Blind Face RestorationabstractBlind face restoration is a challenging task due to the unknown and complex degradation. Although face prior-based methods and reference-based methods have recently demonstrated high-quality results, the restored images tend to contain over-smoothed results and lose identity-preserved details when the degradation is severe. It is observed that this is attributed to short-range dependencies, the intrinsic limitation of convolutional neural networks. To model long-range dependencies, we propose a Transformer-based blind face restoration method, named BFRFormer, to reconstruct images with more identity-preserved details in an end-to-end manner. In BFRFormer, to remove blocking artifacts, the wavelet discriminator and aggregated attention module are developed, and spectral normalization and balanced consistency regulation are adaptively applied to address the training instability and over-fitting problem, respectively. Extensive experiments show that our method outperforms state-of-the-art methods on a synthetic dataset and four real-world datasets. The source code, Casia-Test dataset, and pre-trained models is released at https://github.com/s8Znk/BFRFormer. Guojing Ge, Qi Song 0003, Guibo Zhu, Yuting Zhang 0007, Jinglu Chen, Miao Xin, Ming Tang 0001, Jinqiao Wang |
ICASSP | 7 |
| 2024 | FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality LocalizationabstractZero-shot anomaly detection (ZSAD) methods detect anomalies without prior access to known normal or abnormal samples within target categories. Existing methods typically rely on pretrained multimodal models, computing similarities between manually crafted textual features representing ''normal'' or ''abnormal'' semantics and image patch features to detect anomalies. However, the generic descriptions of ''abnormal'' often fail to precisely match diverse types of anomalies across different object categories. Additionally, computing feature similarities for single patches struggles to pinpoint specific locations of anomalies with various sizes and scales. To address these issues, we propose a novel ZSAD method called FiLo, comprising two components: adaptively learned Fine-Grained Description (FG-Des) and position-enhanced High-Quality Localization (HQ-Loc). FG-Des introduces fine-grained anomaly descriptions for each category using Large Language Models (LLMs) and employs adaptively learned textual templates to enhance the accuracy and interpretability of anomaly detection. HQ-Loc, utilizing Grounding DINO for preliminary localization, position-enhanced text prompts, and Multi-scale Multi-shape Cross-modal Interaction (MMCI) module, facilitates more accurate localization of anomalies of different sizes and shapes. Experimental results on datasets like MVTec and VisA demonstrate that FiLo significantly improves the performance of ZSAD in both detection and localization, achieving state-of-the-art performance with an image-level AUC of 83.9% and a pixel-level AUC of 95.9% on the VisA dataset. Code is available at https://github.com/CASIA-IVA-Lab/FiLo. Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Hao Li 0115, Ming Tang 0001, Jinqiao Wang |
ACM Multimedia | 6 |
| 2024 | Efficient Masked Autoencoders With Self-ConsistencyabstractInspired by the masked language modeling (MLM) in natural language processing tasks, the masked image modeling (MIM) has been recognized as a strong self-supervised pre-training method in computer vision. However, the high random mask ratio of MIM results in two serious problems: 1) the inadequate data utilization of images within each iteration brings prolonged pre-training, and 2) the high inconsistency of predictions results in unreliable generations, i.e., the prediction of the identical patch may be inconsistent in different mask rounds, leading to divergent semantics in the ultimately generated outcomes. To tackle these problems, we propose the efficient masked autoencoders with self-consistency (EMAE) to improve the pre-training efficiency and increase the consistency of MIM. In particular, we present a parallel mask strategy that divides the image into K non-overlapping parts, each of which is generated by a random mask with the same mask ratio. Then the MIM task is conducted parallelly on all parts in an iteration and the model minimizes the loss between the predictions and the masked patches. Besides, we design the self-consistency learning to further maintain the consistency of predictions of overlapping masked patches among parts. Overall, our method is able to exploit the data more efficiently and obtains reliable representations. Experiments on ImageNet show that EMAE achieves the best performance on ViT-Large with only 13% of MAE pre-training time using NVIDIA A100 GPUs. After pre-training on diverse datasets, EMAE consistently obtains state-of-the-art transfer ability on a variety of downstream tasks, such as image classification, object detection, and semantic segmentation. Zhaowen Li, Yousong Zhu, Zhiyang Chen 0002, Wei Li 0314, Rui Zhao 0001, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Objformer: Boosting 3D object detection via instance-wise interaction
Manli Tao, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
Pattern Recognit. | 3 |
| 2024 | ImFusion: Boosting Two-Stage 3D Object Detection via Image CandidatesabstractMulti-modal fusion methods combine the advantages of both point clouds and RGB images to boost the performance of 3D object detection. Despite the significant progress, we find that existing two-stage multi-modal fusion methods suffer from the 3D proposal missing in the first stage and projected-style feature fusion mechanism. To solve these problems, we propose a two-stage multi-modal feature fusion network, which improves the recall rate of hard targets in the first stage of network with pseudo 3D proposals generated from image candidates. Then, considering the complementary information between similar image foreground features across multiple objects, we design a multi-modal cross-target fusion module to pay more attention to the foreground objects. It enables a 3D proposal can aggregate the semantic features of multiple image candidates belonging to the same category. Finally, these enhanced fused proposals are processed in the second stage to further boost the performance of 3D detector. Experimental results on SUN RGB-D and KITTI datasets show the effectiveness of our proposed method. Manli Tao, Chaoyang Zhao, Jinqiao Wang, Ming Tang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | EFCPose: End-to-End Multi-Person Pose Estimation With Fully Convolutional HeadsabstractMainstream methods of multi-person pose estimation are not end-to-end. Recently, some methods build an end-to-end framework based on the DETR framework, aiming to eliminate the need for hand-crafted modules like heuristic grouping and NMS post-processing. However, these DETR-based methods suffer from a heavy memory burden of processing the high-resolution backbone feature maps with transformers. In this paper, we propose an end-to-end multi-person pose estimation method with a fully convolutional network, termed EFCPose. Different from DETR-based methods, it directly predicts instance-aware poses in a pixel-wise manner with lightweight convolutional heads, avoiding the heavy memory burden. Overall, our method adopts the center-offset formulation and a one-to-one label assignment strategy to achieve the multi-person pose estimation in an end-to-end manner. The main contribution of our fully convolutional heads includes two aspects. On the one hand, we propose an unaligned center-offset representation to learn more reliable semantic centers to replace the inconsistent geometric centers, improving the performance of instance detection. On the other hand, we propose a novel regression strategy named limb-aware adaptive regression, which leverages separate adaptive points to convert challenging long-range offsets into simplified short-range offsets and incorporates limb constraints to elevate the regression quality of joint offsets. Compared with current DETR-based end-to-end methods, EFCPose avoids high computational complexity and achieves higher accuracy. Extensive experiments on COCO Keypoint and CrowdPose benchmarks show that EFCPose outperforms other state-of-the-art bottom-up and single-stage methods without flipping augmentation. Yingying Chen 0003, Zhiyang Chen 0002, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | AAformer: Auto-Aligned Transformer for Person Re-IdentificationabstractIn person re-identification (re-ID), extracting part-level features from person images has been verified to be crucial to offer fine-grained information. Most of the existing CNN-based methods only locate the human parts coarsely, or rely on pretrained human parsing models and fail in locating the identifiable nonhuman parts (e.g., knapsack). In this article, we introduce an alignment scheme in transformer architecture for the first time and propose the auto-aligned transformer (AAformer) to automatically locate both the human parts and nonhuman ones at patch level. We introduce the "Part tokens ([PART]s)," which are learnable vectors, to extract part features in the transformer. A [PART] only interacts with a local subset of patches in self-attention and learns to be the part representation. To adaptively group the image patches into different subsets, we design the auto-alignment. Auto-alignment employs a fast variant of optimal transport (OT) algorithm to online cluster the patch embeddings into several groups with the [PART]s as their prototypes. AAformer integrates the part alignment into the self-attention and the output [PART]s can be directly used as part features for retrieval. Extensive experiments validate the effectiveness of [PART]s and the superiority of AAformer over various state-of-the-art methods. Kuan Zhu, Haiyun Guo, Shiliang Zhang, Yaowei Wang 0001, Jing Liu 0001, Jinqiao Wang, Ming Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | ZBS: Zero-Shot Background Subtraction via Instance-Level Background Modeling and Foreground SelectionabstractBackground subtraction (BGS) aims to extract all moving objects in the video frames to obtain binary foreground segmentation masks. Deep learning has been widely used in this field. Compared with supervised-based BGS methods, unsupervised methods have better generalization. However, previous unsupervised deep learning BGS algorithms perform poorly in sophisticated scenarios such as shadows or night lights, and they cannot detect objects outside the pre-defined categories. In this work, we propose an unsuper-vised BGS algorithm based on zero-shot object detection called Zero-shot Background Subtraction (ZBS). The proposed method fully utilizes the advantages of zero-shot object detection to build the open-vocabulary instance-level background model. Based on it, the foreground can be effectively extracted by comparing the detection results of new frames with the background model. ZBS performs well for sophisticated scenarios, and it has rich and extensible categories. Furthermore, our method can easily generalize to other tasks, such as abandoned object detection in unseen environments. We experimentally show that ZBS surpasses state-of-the-art unsupervised BGS methods by 4.70% F-Measure on the CDnet 2014 dataset. The code is released at https://github.com/CASIA-IVA-Lab/ZBS. Yongqi An, Xu Zhao 0003, Tao Yu 0013, Haiyun Gu, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
CVPR | 6 |
| 2023 | Explicit Attention Modeling for Pedestrian Attribute RecognitionabstractRecent studies on pedestrian attribute recognition have achieved significant improvements by utilizing complex networks and attention mechanisms. However, most of these studies learn the attention map implicitly through the class activation map. In this paper, we propose an explicit attention modeling approach for pedestrian attribute recognition. We construct a mask branch to learn the attention maps with a lightweight feature pyramid network. The features inside the specific mask are then averaged to obtain the scores for attribute recognition. Additionally, we introduce spatial and semantic distillation to improve the consistency of attention masks and attribute scores. Our experiments demonstrate that the proposed explicit attention modeling can achieve state-of-the-art performance on PA100K, PETA, and PAR datasets with negligible parameters. Jinyi Fang, Bingke Zhu, Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001 |
ICME | 5 |
| 2023 | FreConv: Frequency Branch-and-Integration Convolutional NetworksabstractRecent researches indicate that utilizing the frequency information of input data can enhance the performance of networks. However, the existing popular convolutional structure is not designed specifically for utilizing the frequency information contained in datasets. In this paper, we propose a novel and effective module, named FreConv (frequency branch-and-integration convolution), to replace the vanilla convolution. FreConv adopts a dual-branch architecture to extract and integrate high- and low-frequency information. In the high-frequency branch, a derivative-filter-like architecture is designed to extract the high-frequency information while a light extractor is employed in the low-frequency branch because the low-frequency information is usually redundant. FreConv is able to exploit the frequency information of input data in a more reasonable way to enhance feature representation ability and reduce the memory and computational cost significantly. Without any bells and whistles, experimental results on various tasks demonstrate that FreConv-equipped networks consistently outperform state-of-the-art baselines. Zhaowen Li, Xu Zhao 0003, Peigeng Ding, Zongxing Gao, Ming Tang 0001, Jinqiao Wang |
ICME | 6 |
| 2023 | Bi-Level Implicit Semantic Data Augmentation for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) aims at finding the target vehicle identity from multi-camera surveillance videos, which plays an important role in the intelligent transportation system (ITS). It suffers from the subtle discrepancy among vehicles from the same vehicle model and large variation across different viewpoints of the same vehicle. To enhance the robustness of Re-ID models, many methods exploit additional detection or segmentation models to extract discriminative local features. Some others employ data-driven methods to enrich the diversity of the training data, such as the data augmentation and 3D-based data generation, so that the Re-ID model can obtain stronger robustness against intra-class variations. However, these methods either rely on extra annotations or greatly increase the computational cost. In this paper, we propose the Bi-level Implicit semantic Data Augmentation (BIDA) framework to solve this problem from two aspects. (1) We implicitly augment the images semantically in the feature space according to the identity-level and superclass-level intra-class variations, which can generate more diverse semantic augmentations beyond the intra-identity variations. (2) We introduce the similarity ranking constraints on the augmented training set by extending the sample-wise triplet loss to the distribution-wise one, which can effectively reduce meaningless semantic transformations and improve the discrimination of the feature. We conduct extensive experiments on VeRi-776, VehicleID and Cityflow benchmarks to reveal the effectiveness of our method. And we achieve new state-of-the-art performance on VeRi-776. Wei Li 0315, Haiyun Guo, Honghui Dong, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Human Parsing With Part-Aware Relation ModelingabstractIn this paper, a Part-aware Relation Modeling (PRM) is developed to handle the task of human parsing. For pixel-level recognition, it is essential to generate features with adaptive context for various sizes and shapes of human parts. To address the issue, we adaptively capture contexts based on the part-aware relation mechanism. PRM mainly consists of three modules, including a part class module, a part-relation aggregation module, and a part-relation dispersion module. The part class module selectively enhances spatial details of the high-level features to obtain enhanced original features, and then extracts the high-level representations of every human part from a categorical perspective. The part-relation aggregation module is developed to extract the representative global context by exploring associated semantics of human parts, adaptively augmenting the context for human parts. The part-relation dispersion module is designed to generate the discriminative and effective local context and neglect the distracting one by making the affinity of human parts disperse. It ensures that features of the same class will be close to each other and away from those of different classes. By fusing the outputs of the two part-relation modules and the first outputs of the part class module, our PRM produces adaptive contextual features for diverse sizes of human parts, boosting the parsing accuracy. Extensive experiments are conducted to validate the effectiveness of our network, and a new state-of-the-art segmentation performance is achieved on three challenging human parsing datasets,i.e., PASCAL-Person-Part, LIP, and CIHP. PRM is also extended to other tasks like animal parsing, and exhibits its generality. Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang, Xiangyu Zhu 0001, Zhen Lei 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Learning Semantics-Consistent Stripes With Self-Refinement for Person Re-IdentificationabstractAligning human parts automatically is one of the most challenging problems for person re-identification (re-ID). Recently, the stripe-based methods, which equally partition the person images into the fixed stripes for aligned representation learning, have achieved great success. However, the stripes with fixed height and position cannot well handle the misalignment problems caused by inaccurate detection and occlusion and may introduce much background noise. In this article, we aim at learning adaptive stripes with foreground refinement to achieve pixel-level part alignment by only using person identity labels for person re-ID and make two contributions. 1) A semantics-consistent stripe learning method (SCS). Given an image, SCS partitions it into adaptive horizontal stripes and each stripe is corresponding to a specific semantic part. Specifically, SCS iterates between two processes: i) clustering the rows to human parts or background to generate the pseudo-part labels of rows and ii) learning a row classifier to partition a person image, which is supervised by the latest pseudo-labels. This iterative scheme guarantees the accuracy of the learned image partition. 2) A self-refinement method (SCS+) to remove the background noise in stripes. We employ the above row classifier to generate the probabilities of pixels belonging to human parts (foreground) or background, which is called the class activation map (CAM). Only the most confident areas from the CAM are assigned with foreground/background labels to guide the human part refinement. Finally, by intersecting the semantics-consistent stripes with the foreground areas, SCS+ locates the human parts at pixel-level, obtaining a more robust part-aligned representation. Extensive experiments validate that SCS+ sets the new state-of-the-art performance on three widely used datasets including Market-1501, DukeMTMC-reID, and CUHK03-NP. Kuan Zhu, Haiyun Guo, Songyan Liu, Jinqiao Wang, Ming Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingabstractSelf-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the scene and instances, as well as the semantic difference of instances in the scene. To address the above problems, we propose a Unified Self-supervised Visual Pre-training (UniVIP), a novel self-supervised framework to learn versatile visual representations on either single-centric-object or non-iconic dataset. The framework takes into account the representation learning at three levels: 1) the similarity of scene-scene, 2) the correlation of scene-instance, 3) the discrimination of instance-instance. During the learning, we adopt the optimal transport algorithm to automatically measure the discrimination of instances. Massive experiments show that Uni-VIP pre-trained on non-iconic COCO achieves state-of-the-art transfer performance on a variety of downstream tasks, such as image classification, semi-supervised learning, object detection and segmentation. Furthermore, our method can also exploit single-centric-object dataset such as ImageNet and outperforms BYOL by 2.5% with the same pre-training epochs in linear probing, and surpass current self-supervised object detection methods on COCO dataset, demonstrating its universality and potential. Zhaowen Li, Yousong Zhu, Fan Yang 0089, Wei Li 0314, Chaoyang Zhao, Yingying Chen 0003, Zhiyang Chen 0002, Jiahao Xie 0002, Rui Zhao 0001, Ming Tang 0001, Jinqiao Wang |
CVPR | 11 |
| 2022 | C2AM Loss: Chasing a Better Decision Boundary for Long-Tail Object DetectionabstractLong-tail object detection suffers from poor performance on tail categories. We reveal that the real culprit lies in the extremely imbalanced distribution of the classifier's weight norm. For conventional softmax cross-entropy loss, such imbalanced weight norm distribution yields ill conditioned decision boundary for categories which have small weight norms. To get rid of this situation, we choose to maxi-mize the cosine similarity between the learned feature and the weight vector of target category rather than the inner-product of them. The decision boundary between any two categories is the angular bisector of their weight vectors. Whereas, the absolutely equal decision boundary is sub-optimal because it reduces the model's sensitivity to vari-ous categories. Intuitively, categories with rich data diver-sity should occupy a larger area in the classification space while categories with limited data diversity should occupy a slightly small space. Hence, we devise a Category-Aware Angular Margin Loss (C2AM Loss) to introduce an adaptive angular margin between any two categories. Specif-ically, the margin between two categories is proportional to the ratio of their classifiers' weight norms. As a result, the decision boundary is slightly pushed towards the cat-egory which has a smaller weight norm. We conduct comprehensive experiments on LVIS dataset. C2AM Loss brings 4.9~5.2 AP improvements on different detectors and back-bones compared with baseline. Tong Wang 0015, Yousong Zhu, Yingying Chen 0003, Chaoyang Zhao, Jinqiao Wang, Ming Tang 0001 |
CVPR | 7 |
| 2022 | Regularizing Vector Embedding in Bottom-Up Human Pose Estimation
Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
ECCV (6) | 4 |
| 2022 | PASS: Part-Aware Self-Supervised Pre-Training for Person Re-Identification
Kuan Zhu, Haiyun Guo, Tianyi Yan, Yousong Zhu, Jinqiao Wang, Ming Tang 0001 |
ECCV (14) | 6 |
| 2022 | Transfering Low-Frequency Features for Domain AdaptationabstractPrevious unsupervised domain adaptation methods did not handle the cross-domain problem from the perspective of frequency for computer vision. The images or feature maps of different domains can be decomposed into the low-frequency component and high-frequency component. This paper pro-poses the assumption that low-frequency information is more domain-invariant while the high-frequency information con-tains domain-related information. Hence, we introduce an approach, named low-frequency module (LFM), to extract domain-invariant feature representations. The LFM is constructed with the digital Gaussian low-pass filter. Our method is easy to implement and introduces no extra hyperparame-ter. We design two effective ways to utilize the LFM for domain adaptation, and our method is complementary to other existing methods and formulated as a plug-and-play unit that can be combined with these methods. Experimental results demonstrate that our LFM outperforms state-of-the-art meth-ods for various computer vision tasks, including image clas-sification and object detection. Zhaowen Li, Xu Zhao 0003, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
ICME | 4 |
| 2022 | Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual TasksabstractVisual tasks vary a lot in their output formats and concerned contents, therefore it is hard to process them with an identical structure. One main obstacle lies in the high-dimensional outputs in object-level visual tasks. In this paper, we propose an object-centric vision framework, Obj2Seq. Obj2Seq takes objects as basic units, and regards most object-level visual tasks as sequence generation problems of objects. Therefore, these visual tasks can be decoupled into two steps. First recognize objects of given categories, and then generate a sequence for each of these objects. The definition of the output sequences varies for different tasks, and the model is supervised by matching these sequences with ground-truth targets. Obj2Seq is able to flexibly determine input categories to satisfy customized requirements, and be easily extended to different visual tasks. When experimenting on MS COCO, Obj2Seq achieves 45.7% AP on object detection, 89.0% AP on multi-label classification and 65.0% AP on human pose estimation. These results demonstrate its potential to be generally applied to different visual tasks. Code has been made available at: https://github.com/CASIA-IVA-Lab/Obj2Seq. Zhiyang Chen 0002, Yousong Zhu, Zhaowen Li, Fan Yang 0089, Wei Li 0314, Chaoyang Zhao, Rui Zhao 0001, Jinqiao Wang, Ming Tang 0001 |
NeurIPS | 11 |
| 2022 | Global Patch Cross-Attention for Point Cloud Analysis
Manli Tao, Chaoyang Zhao, Jinqiao Wang, Ming Tang 0001 |
PRCV (3) | 4 |
| 2022 | Dynamic Orthogonal Projection Constrained Discriminative TrackingabstractDue to the end-to-end feature learning with convolutional neural networks (CNNs), modern discriminative trackers improve the state of the art significantly. To achieve a strong discrimination, the learned features are usually high-dimensional, resulting in a massive number of parameters contained in the discriminative model and the increase of risk of over-fitting in the online tracking. In this letter, we try to alleviate the risk of over-fitting by means of the adaptive dimensionality reduction (DR) through CNNs. Specifically, an orthogonality constrained ridge regression model is proposed to reduce the dimensionality of features, and a dynamic sub-network (DOPNet) is designed to learn to perform DR. After trained with an orthogonality loss and a regression one, DOPNet generates a set of orthogonal bases (i. e., weights in FC layers) dynamically to reduce the feature dimensionality for a discriminative model in the online tracking. Based on the novel discriminative model and DOPNet, an effective and efficient tracker, DOPTracker, is developed. DOPTracker achieves the state-of-the-art results on four benchmarks, OTB-2015, VOT-2018, NfS, and GOT-10 k while running at 30 FPS. Ming Tang 0001, Guibo Zhu, Jinqiao Wang, Hanqing Lu |
IEEE Signal Process. Lett. | 2 |
| 2022 | Grammar-Induced Wavelet Network for Human ParsingabstractMost existing methods of human parsing still face a challenge: how to extract the accurate foreground from similar or cluttered scenes effectively. In this paper, we propose a Grammar-induced Wavelet Network (GWNet), to deal with the challenge. GWNet mainly consists of two modules, including a blended grammar-induced module and a wavelet prediction module. We design the blended grammar-induced module to exploit the relationship of different human parts and the inherent hierarchical structure of a human body by means of grammar rules in both cascaded and paralleled manner. In this way, conspicuous parts, which are easily distinguished from the background, can amend the segmentation of inconspicuous ones, improving the foreground extraction. We also design a Part-aware Convolutional Recurrent Neural Network (PCRNN) to pass messages which are generated by grammar rules. To further improve the performance, we propose a wavelet prediction module to capture the basic structure and the edge details of a person by decomposing the low-frequency and high-frequency components of features. The low-frequency component can represent the smooth structures and the high-frequency components can describe the fine details. We conduct extensive experiments to evaluate GWNet on PASCAL-Person-Part, LIP, and PPSS datasets. GWNet obtains state-of-the-art performance on these human parsing datasets. Yingying Chen 0003, Ming Tang 0001, Zhen Lei 0001, Jinqiao Wang |
IEEE Trans. Image Process. | 3 |
| 2022 | Multi-Granularity Mutual Learning Network for Object Re-IdentificationabstractObject re-identification (re-ID), which is key and fundamental technology for intelligent transportation systems, is a challenging task including person re-ID and vehicle re-ID. It aims to retrieve a given target object from the gallery images captured by different cameras. In this task, it is necessary to extract fine-grained and discriminative features to deal with complex inter-class and intra-class variations caused by the changes of camera viewpoints and object poses. Existing methods focus on learning discriminative local features to improve the re-ID performance. Some state-of-the-art methods use key point detection model to locate local features, which also increases the additional computational cost as side effect. Another type of method focuses on how to learn features of different granularity from rigid stripes of different scales. However, there is little attention paid to how to effectively coalesce multi-granularity features without additional calculation cost. To tackle this issue, this paper proposes the Multi-granularity Mutual Learning Network (MMNet) and makes two contributions. 1) We introduce the multi-granularity jigsaw puzzle module into object re-ID to impel the network to learn local discriminative features from multiple visual granularities by breaking spatial correlation in original images. 2) We propose a parameter-free multi-scale feature reconstruction module to facilitate mutual learning of features at multiple grain levels, thereby both global features and local features have strong representation capabilities. Extensive experiments demonstrate the effectiveness of our proposed modules and the superiority of our method over various state-of-the-art methods on both person and vehicle re-ID benchmarks. Mingfei Tu, Kuan Zhu, Haiyun Guo, Qinghai Miao, Chaoyang Zhao, Guibo Zhu, Honglin Qiao, Gaopan Huang, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2021 | Adaptive Class Suppression Loss for Long-Tail Object DetectionabstractTo address the problem of long-tail distribution for the large vocabulary object detection task, existing methods usually divide the whole categories into several groups and treat each group with different strategies. These methods bring the following two problems. One is the training inconsistency between adjacent categories of similar sizes, and the other is that the learned model is lack of discrimination for tail categories which are semantically similar to some of the head categories. In this paper, we devise a novel Adaptive Class Suppression Loss (ACSL) to effectively tackle the above problems and improve the detection performance of tail categories. Specifically, we introduce a statistic-free perspective to analyze the long-tail distribution, breaking the limitation of manual grouping. According to this perspective, our ACSL adjusts the suppression gradients for each sample of each class adaptively, ensuring the training consistency and boosting the discrimination for rare categories. Extensive experiments on long-tail datasets LVIS and Open Images show that the our ACSL achieves 5.18% and 5.2% improvements with ResNet50-FPN, and sets a new state of the art. Code and models are available at https://github.com/CASIA-IVA-Lab/ACSL. Tong Wang 0015, Yousong Zhu, Chaoyang Zhao, Wei Zeng 0006, Jinqiao Wang, Ming Tang 0001 |
CVPR | 6 |
| 2021 | Improving Multiple Object Tracking With Single Object TrackingabstractDespite considerable similarities between multiple object tracking (MOT) and single object tracking (SOT) tasks, modern MOT methods have not benefited from the development of SOT ones to achieve satisfactory performance. The major reason for this situation is that it is inappropriate and inefficient to apply multiple SOT models directly to the MOT task, although advanced SOT methods are of the strong discriminative power and can run at fast speeds.In this paper, we propose a novel and end-to-end trainable MOT architecture that extends CenterNet by adding an SOT branch for tracking objects in parallel with the existing branch for object detection, allowing the MOT task to benefit from the strong discriminative power of SOT methods in an effective and efficient way. Unlike most existing SOT methods which learn to distinguish the target object from its local backgrounds, the added SOT branch trains a separate SOT model per target online to distinguish the target from its surrounding targets, assigning SOT models the novel discrimination. Moreover, similar to the detection branch, the SOT branch treats objects as points, making its online learning efficient even if multiple targets are processed simultaneously. Without tricks, the proposed tracker achieves MOTAs of 0.710 and 0.686, IDF1s of 0.719 and 0.714, on MOT17 and MOT20 benchmarks, respectively, while running at 16 FPS on MOT17. Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Guibo Zhu, Jinqiao Wang, Hanqing Lu |
CVPR | 2 |
| 2021 | High-Performance Discriminative Tracking with TransformersabstractEnd-to-end discriminative trackers improve the state of the art significantly, yet the improvement in robustness and efficiency is restricted by the conventional discriminative model, i.e., least-squares based regression. In this paper, we present DTT, a novel single-object discriminative tracker, based on an encoder-decoder Transformer architecture. By self- and encoder-decoder attention mechanisms, our approach is able to exploit the rich scene information in an end-to-end manner, effectively removing the need for hand-designed discriminative models. In online tracking, given a new test frame, dense prediction is performed at all spatial positions. Not only location, but also bounding box of the target object is obtained in a robust fashion, streamlining the discriminative tracking pipeline. DTT is conceptually simple and easy to implement. It yields state-of-the-art performance on four popular benchmarks including GOT-10k, LaSOT, NfS, and TrackingNet while running at over 50 FPS, confirming its effectiveness and efficiency. We hope DTT may provide a new perspective for single-object visual tracking. Ming Tang 0001, Linyu Zheng, Guibo Zhu, Jinqiao Wang, Xuetao Feng, Hanqing Lu |
ICCV | 2 |
| 2021 | Attention-Guided Knowledge Distillation for Efficient Single-Stage DetectorabstractKnowledge distillation has been successfully applied in image classification for model acceleration. There are also some works employing this technique to object detection, but they all treat different feature regions equally when performing feature mimic. In this paper, we propose an end-to-end attention-guided knowledge distillation method to train efficient single-stage detectors with much smaller backbones. More specifically, we introduce an attention mechanism to prioritize the transfer of important knowledge by focusing on a sparse set of hard samples, leading to a more thorough distillation process. In addition, the proposed distillation method also provides an easy way to train efficient detectors without tedious ImageNet pre-training procedure. Extensive experiments on PASCAL VOC and CityPersons datasets demonstrate the effectiveness of the proposed approach. We achieve 57.96% and 69.48% mAP on VOC07 with the backbone of 1/8 VGG16 and 1/4 VGG16, greatly outperforming their ImageNet pre-trained counterparts by 11.7% and 7.1% respectively. Tong Wang 0015, Yousong Zhu, Chaoyang Zhao, Xu Zhao 0003, Jinqiao Wang, Ming Tang 0001 |
ICME | 6 |
| 2021 | DPT: Deformable Patch-based Transformer for Visual RecognitionabstractTransformer has achieved great success in computer vision, while how to split patches in an image remains a problem. Existing methods usually use a fixed-size patch embedding which might destroy the semantics of objects. To address this problem, we propose a new Deformable Patch (DePatch) module which learns to adaptively split the images into patches with different positions and scales in a data-driven way rather than using predefined fixed patches. In this way, our method can well preserve the semantics in patches. The DePatch module can work as a plug-and-play module, which can easily be incorporated into different transformers to achieve an end-to-end training. We term this DePatch-embedded transformer as Deformable Patch-based Transformer (DPT) and conduct extensive evaluations of DPT on image classification and object detection. Results show DPT can achieve 81.8% top-1 accuracy on ImageNet classification, and 43.7% box AP with RetinaNet, 44.3% with Mask R-CNN on MSCOCO object detection. Code has been made available at: https://github.com/CASIA-IVA-Lab/DPT. Zhiyang Chen 0002, Yousong Zhu, Chaoyang Zhao, Guosheng Hu, Wei Zeng 0006, Jinqiao Wang, Ming Tang 0001 |
ACM Multimedia | 7 |
| 2021 | Multi-initialization Optimization Network for Accurate 3D Human Pose and Shape Estimationabstract3D human pose and shape recovery from a monocular RGB image is a challenging task. Existing learning based methods highly depend on weak supervision signals, e.g. 2D and 3D joint location, due to the lack of in-the-wild paired 3D supervision. However, considering the 2D-to-3D ambiguities existed in these weak supervision labels, the network is easy to get stuck in local optima when trained with such labels. In this paper, we reduce the ambituity by optimizing multiple initializations. Specifically, we propose a three-stage framework named Multi-Initialization Optimization Network (MION). In the first stage, we strategically select different coarse 3D reconstruction candidates which are compatible with the 2D keypoints of input sample. Each coarse reconstruction can be regarded as an initialization leads to one optimization branch. In the second stage, we design a mesh refinement transformer (MRT) to respectively refine each coarse reconstruction result via a self-attention mechanism. Finally, a Consistency Estimation Network (CEN) is proposed to find the best result from mutiple candidates by evaluating if the visual evidence in RGB image matches a given 3D reconstruction. Experiments demonstrate that our Multi-Initialization Optimization Network outperforms existing 3D mesh based methods on multiple public benchmarks. Zhiwei Liu 0004, Xiangyu Zhu 0001, Lu Yang 0006, Ming Tang 0001, Zhen Lei 0001, Guibo Zhu, Xuetao Feng, Yan Wang 0068, Jinqiao Wang |
ACM Multimedia | 5 |
| 2021 | MST: Masked Self-Supervised Transformer for Visual RepresentationabstractTransformer has been widely used for self-supervised pre-training in Natural Language Processing (NLP) and achieved great success. However, it has not been fully explored in visual self-supervised learning. Meanwhile, previous methods only consider the high-level feature and learning representation from a global perspective, which may fail to transfer to the downstream dense prediction tasks focusing on local features. In this paper, we present a novel Masked Self-supervised Transformer approach named MST, which can explicitly capture the local context of an image while preserving the global semantic information. Specifically, inspired by the Masked Language Modeling (MLM) in NLP, we propose a masked token strategy based on the multi-head self-attention map, which dynamically masks some tokens of local patches without damaging the crucial structure for self-supervised learning. More importantly, the masked tokens together with the remaining tokens are further recovered by a global image decoder, which preserves the spatial information of the image and is more friendly to the downstream dense prediction tasks. The experiments on multiple datasets demonstrate the effectiveness and generality of the proposed method. For instance, MST achieves Top-1 accuracy of 76.9% with DeiT-S only using 300-epoch pre-training by linear evaluation, which outperforms supervised methods with the same epoch by 0.4% and its comparable variant DINO by 1.0%. For dense prediction tasks, MST also achieves 42.7% mAP on MS COCO object detection and 74.04% mIoU on Cityscapes segmentation only with 100-epoch pre-training. Zhaowen Li, Zhiyang Chen 0002, Fan Yang 0089, Wei Li 0314, Yousong Zhu, Chaoyang Zhao, Rui Zhao 0001, Ming Tang 0001, Jinqiao Wang |
NeurIPS | 10 |
| 2021 | High-Performance Discriminative Tracking with Target-Aware Feature Embeddings
Ming Tang 0001, Linyu Zheng, Guibo Zhu, Jinqiao Wang, Hanqing Lu |
PRCV (1) | 2 |
| 2021 | Fast Kernelized Correlation Filter without Boundary EffectabstractIn recent years, correlation filter based trackers (CF trackers) have attracted much attention from the vision community because of their top performance in both localization accuracy and efficiency. The society of visual tracking, however, still needs to deal with the following difficulty on CF trackers: avoiding or eliminating the boundary effect completely, in the meantime, exploiting non-linear kernels and running efficiently. In this paper, we propose a fast kernelized correlation filter without boundary effect (nBEKCF) to solve this problem. To avoid the boundary effect thoroughly, a set of real and dense patches is sampled through the traditional sliding window and used as the training samples to train nBEKCF to fit a Gaussian response map. Non-linear kernels can be applied naturally in nBEKCF due to its different theoretical foundation from the existing CF trackers'. To achieve the fast training and detection, a set of cyclic bases is introduced to construct the filter. Two algorithms, ACSII and CCIM, are developed to significantly accelerate the calculation of kernel correlation matrices. ACSII and CCIM fully exploit the density of training samples and cyclic structure of bases, and totally run in space domain. The efficiency of CCIM exceeds that of the FFT counterpart remarkably in our task. Extensive experiments on six public datasets, OTB-2013, OTB-2015, NfS, VOT2018, GOT10k, and TrackingNet, show that compared to the CF trackers designed to relax the boundary effect, BACF and SRDCF, our nBEKCF achieves higher localization accuracy without tricks, in the meanwhile, runs at higher FPS. Ming Tang 0001, Linyu Zheng, Jinqiao Wang |
WACV | 1 |
| 2021 | Unsupervised cycle-consistent person pose transfer
Songyan Liu, Haiyun Guo, Kuan Zhu, Jinqiao Wang, Ming Tang 0001 |
Neurocomputing | 5 |
| 2021 | Siamese Regression Tracking With Reinforced Template UpdatingabstractSiamese networks are prevalent in visual tracking because of the efficient localization. The networks take both a search patch and a target template as inputs where the target template is usually from the initial frame. Meanwhile, Siamese trackers do not update network parameters online for real-time efficiency. The fixed target template and CNN parameters make Siamese trackers not effective to capture target appearance variations. In this paper, we propose a template updating method via reinforcement learning for Siamese regression trackers. We collect a series of templates and learn to maintain them based on an actor-critic framework. Among this framework, the actor network that is trained by deep reinforcement learning effectively updates the templates based on the tracking result on each frame. Besides the target template, we update the Siamese regression tracker online to adapt to target appearance variations. The experimental results on the standard benchmarks show the effectiveness of both template and network updating. The proposed tracker SiamRTU performs favorably against state-of-the-art approaches. Fei Zhao 0008, Ting Zhang 0006, Yibing Song, Ming Tang 0001, Xiaobo Wang 0002, Jinqiao Wang |
IEEE Trans. Image Process. | 4 |
| 2021 | Antidecay LSTM for Siamese Tracking With Adversarial LearningabstractVisual tracking is one of the fundamental tasks in computer vision with many challenges, and it is mainly due to the changes in the target's appearance in temporal and spatial domains. Recently, numerous trackers model the appearance of the targets in the spatial domain well by utilizing deep convolutional features. However, most of these CNN-based trackers only take the appearance variations between two consecutive frames in a video sequence into consideration. Besides, some trackers model the appearance of the targets in the long term by applying RNN, but the decay of the target's features degrades the tracking performance. In this article, we propose the antidecay long short-term memory (AD-LSTM) for the Siamese tracking. Especially, we extend the architecture of the standard LSTM in two aspects for the visual tracking task. First, we replace all of the fully connected layers with convolutional layers to extract the features with spatial structure. Second, we improve the architecture of the cell unit. In this way, the information of the target appearance can flow through the AD-LSTM without decay as long as possible in the temporal domain. Meanwhile, since there is no ground truth for the feature maps generated by the AD-LSTM, we propose an adversarial learning algorithm to optimize the AD-LSTM. With the help of adversarial learning, the Siamese network can generate the response maps more accurately, and the AD-LSTM can generate the feature maps of the target more robustly. The experimental results show that our tracker performs favorably against the state-of-the-art trackers on six challenging benchmarks: OTB-100, TC-128, VOT2016, VOT2017, GOT-10k, and TrackingNet. Fei Zhao 0008, Ting Zhang 0006, Yi Wu 0001, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Part-Aware Context Network for Human ParsingabstractRecent works have made significant progress in human parsing by exploiting rich contexts. However, human parsing still faces a challenge of how to generate adaptive contextual features for the various sizes and shapes of human parts. In this work, we propose a Part-aware Context Network (PCNet), a novel and effective algorithm to deal with the challenge. PCNet mainly consists of three modules, including a part class module, a relational aggregation module, and a relational dispersion module. The part class module extracts the high-level representations of every human part from a categorical perspective. We design a relational aggregation module to capture the representative global context by mining associated semantics of human parts, which adaptively augments the context for human parts. We propose a relational dispersion module to generate the discriminative and effective local context and neglect disturbing one by making the affinity of human parts dispersed. The relational dispersion module ensures that features in the same class will be close to each other and away from those of different classes. By fusing the outputs of the relational aggregation module, the relational dispersion module and the backbone network, our PCNet generates adaptive contextual features for various sizes of human parts, improving the parsing accuracy. We achieve a new state-of-the-art segmentation performance on three challenging human parsing datasets, i.e., PASCAL-Person-Part, LIP, and CIHP. Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001 |
CVPR | 5 |
| 2020 | Large Batch Optimization for Object Detection: Training COCO in 12 minutes
Tong Wang 0015, Yousong Zhu, Chaoyang Zhao, Wei Zeng 0006, Yaowei Wang 0001, Jinqiao Wang, Ming Tang 0001 |
ECCV (21) | 7 |
| 2020 | Adaptive Variance Based Label Distribution Learning for Facial Age Estimation
Xin Wen 0005, Biying Li, Haiyun Guo, Zhiwei Liu 0004, Guosheng Hu, Ming Tang 0001, Jinqiao Wang |
ECCV (23) | 6 |
| 2020 | Blended Grammar Network for Human Parsing
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001 |
ECCV (24) | 5 |
| 2020 | Learning Feature Embeddings for Discriminant Model Based Tracking
Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
ECCV (15) | 2 |
| 2020 | Identity-Guided Human Semantic Parsing for Person Re-identification
Kuan Zhu, Haiyun Guo, Zhiwei Liu 0004, Ming Tang 0001, Jinqiao Wang |
ECCV (3) | 4 |
| 2020 | High-Speed And Accurate Scale Estimation For Visual Tracking With Gaussian Process RegressionabstractRecent years have seen remarkable progress in the visual tracking domain. However, it remains a challenging task to estimate the scale of target efficiently and accurately. In this paper, we present a novel and high-performance scale estimation approach for tracking-by-detection framework. The proposed approach, named GPAS, formulates the scale estimation as a Gaussian process regression problem based on scale pyramid representation. In general, it enjoys the following there advantages. (i) Efficient. It only takes 2ms to estimate the scale of a target on a single CPU. (ii) Accurate. Without bells and whistles, its accuracy surpasses all previous hand-crafted features based scale estimation methods by large margins. (iii) Generic. It can be incorporated into any tracking-by-detection framework based trackers easily. Experiment results show that compared to the latest and classical scale estimation method, fDSST, our GPAS significantly improves the performance by 6.2% in mean distance precision, 8.9% in mean overlap precision, and 5.5% in mean AUC on 28 sequences of OTB2013 with significant scale variations. Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
ICME | 2 |
| 2020 | Task Decoupled Knowledge Distillation For Lightweight Face DetectorsabstractFace detection is a hot topic in computer vision. The face detection methods usually consist of two subtasks, i.e. the classification subtask and the regression subtask, which are trained with different samples. However, current face detection knowledge distillation methods usually couple the two subtasks, and use the same set of samples in the distillation task. In this paper, we propose a task decoupled knowledge distillation method, which decouples the detection distillation task into two subtasks and uses different samples in distilling the features of different subtasks. We firstly propose a feature decoupling method to decouple the classification features and the regression features, without introducing any extra calculations at inference time. Specifically, we generate the corresponding features by adding task-specific convolutions in the teacher network and adding adaption convolutions on the feature maps of the student network. Then we select different samples for different subtasks to imitate. Moreover, we also propose an effective probability distillation method to joint boost the accuracy of the student network. We apply our distillation method on a lightweight face detector, EagleEye. Experimental results show that the proposed method effectively improves the student detector's accuracy by 5.1%, 5.1%, and 2.8% AP in Easy, Medium, Hard subsets respectively. Xiaoqing Liang, Xu Zhao 0003, Chaoyang Zhao, Nanfei Jiang, Ming Tang 0001, Jinqiao Wang |
ACM Multimedia | 5 |
| 2020 | Siamese Attentive Graph TrackingabstractRecently, deep Siamese matching networks have attracted increasing attention for visual tracking. Despite the demonstrated successes, Siamese trackers do not take full advantage of the structural information of target objects. They tend to drift in the presence of non-rigid deformation or partly occlusion. In this paper, we propose to advance Siamese trackers with graph convolutional networks, which pay more attention to the structural layout of target objects, to learn features robust to large appearance changes over time. Specifically, we divide the target object into several sub-parts and design an attentive graph convolutional network to model the relationship between parts. We incrementally update the attention coefficients of the graph with the attention scheme at each frame in an end-to-end manner. To further improve localization accuracy, we propose a learnable cascade regression algorithm based on deep reinforcement learning to refine the predicted bounding boxes. Extensive experiments on seven challenging benchmark datasets, i.e., OTB-100, TC-128, VOT2018, VOT2019, TrackingNet, GOT-10k and LaSOT, demonstrate that the proposed tracking method performs favorably against state-of-the-art approaches. Fei Zhao 0008, Ting Zhang 0006, Chao Ma 0004, Ming Tang 0001, Jinqiao Wang, Xiaobo Wang 0002 |
ACM Multimedia | 4 |
| 2020 | A novel data augmentation scheme for pedestrian detection with attribute preserving GAN
Songyan Liu, Haiyun Guo, Jian-Guo Hu, Xu Zhao 0003, Chaoyang Zhao, Tong Wang 0015, Yousong Zhu, Jinqiao Wang, Ming Tang 0001 |
Neurocomputing | 9 |
| 2020 | Semantic-spatial fusion network for human parsing
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001 |
Neurocomputing | 5 |
| 2020 | Siamese Deformable Cross-Correlation Network for Real-Time Visual Tracking
Linyu Zheng, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang, Hanqing Lu |
Neurocomputing | 3 |
| 2020 | A Comparison of Correlation Filter-Based Trackers and Struck TrackersabstractIn recent years, two types of trackers, namely correlation filter-based tracker (CF tracker) and structured output tracker (struck), have exhibited the state-of-the-art performance. However, there seems to be a lack of analytic work on their relations in the computer vision community. In this paper, we investigate two state-of-the-art CF trackers, i.e., spatial regularization discriminative correlation filter (SRDCF) and correlation filter with limited boundaries (CFLB), and struck, and reveal their relations. Specifically, after extending the CFLB to its multiple channel versions, we prove the relation between SRDCF and CFLB on the condition that the spatial regularization factor of SRDCF is replaced by the masking matrix of CFLB. We also prove the asymptotical approximate relation between SRDCF and struck on the conditions that the spatial regularization factor of SRDCF is replaced by an indicator function of object bounding box, the weights of SRDCF in its loss item are replaced by those of struck, the linear kernel is employed by struck, and the search region tends to infinity. The extensive experiments on public benchmarks OTB50 and OTB100 are conducted to verify our theoretical results. Moreover, we explain how detailed differences among SRDCF, CFLB, and Struck would give rise to slightly different performances on visual sequences. Jinqiao Wang, Linyu Zheng, Ming Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Semantic Alignment: Finding Semantically Consistent Ground-Truth for Facial Landmark DetectionabstractRecently, deep learning based facial landmark detection has achieved great success. Despite this, we notice that the semantic ambiguity greatly degrades the detection performance. Specifically, the semantic ambiguity means that some landmarks (e.g. those evenly distributed along the face contour) do not have clear and accurate definition, causing inconsistent annotations by annotators. Accordingly, these inconsistent annotations, which are usually provided by public databases, commonly work as the ground-truth to supervise network training, leading to the degraded accuracy. To our knowledge, little research has investigated this problem. In this paper, we propose a novel probabilistic model which introduces a latent variable, i.e. the `real' ground-truth which is semantically consistent, to optimize. This framework couples two parts (1) training landmark detection CNN and (2) searching the `real' ground-truth. These two parts are alternatively optimized: the searched `real' ground-truth supervises the CNN training; and the trained CNN assists the searching of `real' ground-truth. In addition, to recover the unconfidently predicted landmarks due to occlusion and low quality, we propose a global heatmap correction unit (GHCU) to correct outliers by considering the global face shape as a constraint. Extensive experiments on both image-based (300W and AFLW) and video-based (300-VW) databases demonstrate that our method effectively improves the landmark detection accuracy and achieves the state of the art performance. Zhiwei Liu 0004, Xiangyu Zhu 0001, Guosheng Hu, Haiyun Guo, Ming Tang 0001, Zhen Lei 0001, Neil Robertson 0002, Jinqiao Wang |
CVPR | 5 |
| 2019 | Learning Discriminative and Complementary Patches for Face RecognitionabstractThe ensemble of convolutional neural networks (CNNs) has widely been used in many computer vision tasks including face recognition. Many existing ensembles of face recognition CNNs apply a two-stage pipeline to target performance improvement [10], [20], [22], [23], [29]: (1) it trains multiple CNNs separately with many face patches covering different facial areas; (2) the features derived from different models are aggregated off-line by different fusion methods. The well-known face recognition work, DeepID2 [20] trains 200 networks based on 200 arbitrarily chosen facial areas and chooses the best 25 ones to achieve impressive performance. However, it is very time-consuming to train so many networks. In addition, a brute-force like way of choosing facial patches is used without knowing which face patches are complementary and discriminative. It might be lack of generalization capability for cross-database applications. To solve that, we propose a novel end-to-end CNN ensemble architecture which automatically learns the complementary and discriminative patches for face recognition. Specifically, we propose a novel Patch Generation Engine (PGE) with Patch Search Spatial Transformer Network (PS-STN) and ROI shrunk loss to perform the patch selection process. ROI shrunk loss enlarges the distance of learned features in spatial space and feature space and learn complementary features. In order to get final aggregated feature, we use a supervised fusion module named Two Stage Discriminative Fusion Module (TSDFM) which effective to capture the global and local information and further guide the PGE to learn better patches. Extensive experiments conducted on LFW and YTF datasets show the effectiveness of our novel end-to-end ensemble method. Zhiwei Liu 0004, Ming Tang 0001, Guosheng Hu, Jinqiao Wang |
FG | 2 |
| 2019 | Fast-deepKCF Without Boundary EffectabstractIn recent years, correlation filter based trackers (CF trackers) have received much attention because of their top performance. Most CF trackers, however, suffer from low frame-per-second (fps) in pursuit of higher localization accuracy by relaxing the boundary effect or exploiting the high-dimensional deep features. In order to achieve real-time tracking speed while maintaining high localization accuracy, in this paper, we propose a novel CF tracker, fdKCF*, which casts aside the popular acceleration tool, i.e., fast Fourier transform, employed by all existing CF trackers, and exploits the inherent high-overlap among real (i.e., noncyclic) and dense samples to efficiently construct the kernel matrix. Our fdKCF* enjoys the following three advantages. (i) It is efficiently trained in kernel space and spatial domain without the boundary effect. (ii) Its fps is almost independent of the number of feature channels. Therefore, it is almost real-time, i.e., 24 fps on OTB-2015, even though the high-dimensional deep features are employed. (iii) Its localization accuracy is state-of-the-art. Extensive experiments on four public benchmarks, OTB-2013, OTB-2015, VOT2016, and VOT2017, show that the proposed fdKCF* achieves the state-of-the-art localization performance with remarkably faster speed than C-COT and ECO. Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu |
ICCV | 2 |
| 2019 | Bi-Directional Message Passing Based Scanet for Human Pose EstimationabstractArticulated human pose estimation is one of the fundamental computer vision problems. In this paper, a Bi-directional Message Passing(BDMP) module is proposed to fuse convolutional features of different scales in the up-sampling process of the hourglass model for human pose estimation. Moreover, a novel module which integrates Spatial and Channelwise Attention Network(SCANet) is proposed to refine the features obtained from the message passing stage. We design a Semantics-aware Channel-wise Attention(SACWA) module to reduce the feature redundancy and enrich the semantic information simultaneously. A Sharper Spatial Attention(SSA) module based on the Gumbel-Softmax sampling is proposed to exclude the interference from cluttered background and overcomes the gradient degradation induced by the softmax normalization. The proposed framework achieves leading position on MPII benchmark against the state-of-the-arts methods with much less parameters. Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001, Hanqing Lu |
ICME | 4 |
| 2019 | Adversarial image generation by combining content and styleabstractImages can be considered as the combination of two parts: the content and the style. The authors’ approach can leverage this property by extracting a certain unique style from the reference images and combining it to generate images with new contents. With a well‐defined style feature extraction module, they propose a novel framework to generate images with various styles and the same content. To train the style specific image generation model efficiently, a double‐cycle training strategy is proposed: they input two natural‐content pairs simultaneously, extract their style features, and exchange them twice to obtain the reconstruction of the input natural images. What is more, they apply the triplet margin loss to the style feature extracted from the images before and after style exchange and an adversarial discriminator to force the style‐exchanged images to be real. They perform experiments on licence‐plate image, Chinese characters, and shoes or handbags images generating, obtain photo‐realistic results and remarkably improve the corresponding supervised recognition task. Songyan Liu, Chaoyang Zhao, Yunze Gao, Jinqiao Wang, Ming Tang 0001 |
IET Image Process. | 5 |
| 2019 | Reading scene text with fully convolutional sequence modeling
Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001, Hanqing Lu |
Neurocomputing | 4 |
| 2019 | Elite Loss for scene text detection
Xu Zhao 0003, Chaoyang Zhao, Haiyun Guo, Yousong Zhu, Ming Tang 0001, Jinqiao Wang |
Neurocomputing | 5 |
| 2019 | Pixelwise Deep Sequence Learning for Moving Object DetectionabstractMoving object detection is an essential, well-studied but still open problem in computer vision and plays a fundamental role in many applications. Traditional approaches usually reconstruct background images with hand-crafted visual features, such as color, texture, and edge. Due to lack of prior knowledge or semantic information, it is difficult to deal with complicated and rapid changing scenes. To exploit the temporal structure of the pixel-level semantic information, in this paper, we propose an end-to-end deep sequence learning architecture for moving object detection. First, the video sequences are input into a deep convolutional encoder-decoder network for extracting pixel-wise semantic features. Then, to exploit the temporal context, we propose a novel attention long short-term memory (Attention ConvLSTM) to model pixelwise changes over time. A spatial transformer network and a conditional random field layer are finally appended to reduce the sensitivity to camera motion and smooth the foreground boundaries. A multi-task loss is proposed to jointly optimization for frame-based classification and temporal prediction in an end-to-end network. Experimental results on CDnet 2014 and LASIESTA show 12.15% and 16.71% improvement to the state of the art, respectively. Yingying Chen 0003, Jinqiao Wang, Bingke Zhu, Ming Tang 0001, Hanqing Lu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Adversarial Deep TrackingabstractA number of visual tracking methods achieve the state-of-the-art performance based on deep learning recently. However, most of these trackers utilize the deep neural network in regression task or classification task separately. In this paper, we propose an adversarial deep tracking framework. The framework is composed of a fully convolutional Siamese neural network (regression network) and a discriminative classification network. Then, we jointly optimize the regression network and the classification network by adversarial learning. In the uniform framework, the regression network and classification network can be trained end-to-end as a whole using large amounts of video training data sets. During the testing phase, the regression network generates a response map which reflects the location and the size of the target within each candidate search patch, and the classification network discriminates which response map is the best in terms of the corresponding template patch and candidate search patch. In addition, we propose an attention visualization algorithm for our tracker, and it reflects the area that attracts the attention of our tracker during tracking. The experimental results on three large-scale visual tracking benchmarks (OTB-100, TC-128, and VOT2016) demonstrate the effectiveness of the proposed tracking algorithm and show that our tracker performs comparably against the state-of-the-art trackers. Fei Zhao 0008, Jinqiao Wang, Yi Wu 0001, Ming Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Two-Level Attention Network With Multi-Grain Ranking Loss for Vehicle Re-IdentificationabstractVehicle re-identification (re-ID) aims to identify the same vehicle across multiple non-overlapping cameras, which is rather a challenging task. On the one hand, subtle changes in viewpoint and illumination condition can make the same vehicle look much different. On the other hand, different vehicles, even different vehicle models, may look quite similar. In this paper, we propose a novel Two-level Attention network supervised by a Multi-grain Ranking loss (TAMR) to learn an efficient feature embedding for the vehicle re-ID task. The two-level attention network consisting of hard part-level attention and soft pixel-level attention can adaptively extract discriminative features from the visual appearance of vehicles. The former one is designed to localize the salient vehicle parts, such as windscreen and car head. The latter one gives an additional attention refinement at pixel level to focus on the distinctive characteristics within each part. In addition, we present a multi-grain ranking loss to further enhance the discriminative ability of learned features. We creatively take the multi-grain relationship between vehicles into consideration. Thus, not only the discrimination between different vehicles but also the distinction between different vehicle models is constrained. Finally, the proposed network can learn a feature space, where both intra-class compactness and inter-class discrimination are well guaranteed. Extensive experiments demonstrate the effectiveness of our approach and we achieve state-of-the-art results on two challenging datasets, including VehicleID and Vehicle-1M. Haiyun Guo, Kuan Zhu, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Image Process. | 3 |
| 2018 | Progressive Cognitive Human ParsingabstractHuman parsing is an important task for human-centric understanding. Generally, two mainstreams are used to deal with this challenging and fundamental problem. The first one is employing extra human pose information to generate hierarchical parse graph to deal with human parsing task. Another one is training an end-to-end network with the semantic information in image level. In this paper, we develop an end-to-end progressive cognitive network to segment human parts. In order to establish a hierarchical relationship, a novel component-aware region convolution structure is proposed. With this structure, latter layers inherit prior component information from former layers and pay its attention to a finer component. In this way, we deal with human parsing as a progressive recognition task, that is, we first locate the whole human and then segment the hierarchical components gradually. The experiments indicate that our method has a better location capacity for the small objects and a better classification capacity for the large objects. Moreover, our framework can be embedded into any fully convolutional network to enhance the performance significantly. Bingke Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
AAAI | 3 |
| 2018 | High-Speed Tracking With Multi-Kernel Correlation FiltersabstractCorrelation filter (CF) based trackers are currently ranked top in terms of their performances. Nevertheless, only some of them, such as KCF [26] and MKCF [48], are able to exploit the powerful discriminability of non-linear kernels. Although MKCF achieves more powerful discriminability than KCF through introducing multi-kernel learning (MKL) into KCF, its improvement over KCF is quite limited and its computational burden increases significantly in comparison with KCF. In this paper, we will introduce the MKL into KCF in a different way than MKCF. We reformulate the MKL version of CF objective function with its upper bound, alleviating the negative mutual interference of different kernels significantly. Our novel MKCF tracker, MKCFup, outperforms KCF and MKCF with large margins and can still work at very high fps. Extensive experiments on public data sets show that our method is superior to state-of-the-art algorithms for target objects of small move at very high speed. Ming Tang 0001, Jinqiao Wang |
CVPR | 1 |
| 2018 | Dense Chained Attention Network for Scene Text RecognitionabstractReading text in the wild is a challenging task in computer vision. Scene text suffers from various background noise, including shadow, irrelevant symbols and background texture. In order to reduce the disturbance of background noise, we propose a dense chained attention network with stacked attention modules for scene text recognition. Each attention module learns the attention map that is adapted to corresponding features to enhance the foreground text and suppress the background noise. Besides, the attention branch is designed with the convolution-deconvolution structure which rapidly captures global information to guide the discriminative feature selection. We stack multiple attention modules to gradually refine the attention maps and capture both the low-level appearance feature and the high-level semantic information. Extensive experiments on the standard benchmarks, the Street View Text, IIIT5K, and ICDAR datasets validate the superiority of the proposed method. The dense chained attention network achieves state-of-the-art or highly competitive recognition performance. Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001, Hanqing Lu |
ICIP | 4 |
| 2018 | Tree Hierarchical CNNs for Object ParsingabstractObject parsing is a challenging topic in computer vision, which is to distinguish all parts of visual objects. Although lots of works have been proposed, it is difficult to segment complicated objects from complex scenes. Therefore, in this paper we propose a tree hierarchical CNNs for object parsing. Rather than segment all parts of objects at once, we segment object parts step by step in a tree hierarchy and then merge the results together with a full convolutional network. In the tree hierarchy, the segmentation errors of the previous layers of the network outputs could be passed down to following layers and result in accumulated errors. In order to reduce the accumulated errors, we adopt a new part-aware fusion strategy, which fuses global-level feature maps from fully convolutional networks as well as the part-level object feature maps from the output of previous layer. It also contributes to improve the integrity and robustness of object parsing. Finally, the experiments on published datasets show the superiority of the proposed approach, especially for neighboring objects in complex scene. Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001, Hanqing Lu |
ICIP | 5 |
| 2018 | Learning Robust Gaussian Process Regression for Visual TrackingabstractRecent developments of Correlation Filter based trackers (CF trackers) have attracted much attention because of their top performance. However, the boundary effect imposed by the basic periodic assumption in their fast optimization seriously degrades the performance of CF trackers. Although there existed many recent works to relax the boundary effect in CF trackers, the cost was that they can not utilize the kernel trick to improve the accuracy further. In this paper, we propose a novel Gaussian Process Regression based tracker (GPRT) which is a conceptually natural tracking approach. Compared to all the existing CF trackers, the boundary effect is eliminated thoroughly and the kernel trick can be employed in our GPRT. In addition, we present two efficient and effective update methods for our GPRT. Experiments are performed on two public datasets: OTB-2013 and OTB-2015. Without bells and whistles, on these two datasets, our GPRT obtains 84.1% and 79.2% in mean overlap precision, respectively, outperforming all the existing trackers with hand-crafted features. Linyu Zheng, Ming Tang 0001, Jinqiao Wang |
IJCAI | 2 |
| 2017 | Joint background reconstruction and foreground segmentation via a two-stage convolutional neural networkabstractForeground segmentation in video sequences is a classic topic in computer vision. Due to the lack of semantic and prior knowledge, it is difficult for existing methods to deal with sophisticated scenes well. Therefore, in this paper, we propose an end-to-end two-stage deep convolutional neural network (CNN) framework for foreground segmentation in video sequences. In the first stage, a convolutional encoder-decoder sub-network is employed to reconstruct the background images and encode rich prior knowledge of background scenes. In the second stage, the reconstructed background and current frame are input into a multi-channel fully-convolutional sub-network (MCFCN) for accurate foreground segmentation. In the two-stage CNN, the reconstruction loss and segmentation loss are jointly optimized. The background images and foreground objects are output simultaneously in an end-to-end way. Moreover, by incorporating the prior semantic knowledge of foreground and background in the pre-training process, our method could restrain the background noise and keep the integrity of foreground objects at the same time. Experiments on CDNet 2014 show that our method outperforms the state-of-the-art by 4.9%. Xu Zhao 0003, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
ICME | 3 |
| 2017 | DenseTracker: A multi-task dense network for visual trackingabstractHow to track an arbitrary object in video is one of the main challenges in computer vision, and it has been studied for decades. Based on hand-crafted features, traditional trackers show poor discriminability for complex changes of object appearance. Recently, some trackers based on convolutional neural network (CNN) have shown some promising results by exploiting the rich convolutional features. In this paper, we propose a novel DenseTracker based on a mutli-task dense convolutional network. To learn a more compact and discriminative representation, we adopt a dense block structure to ensemble features from different layers. Then a multitask loss is designed to accurately predict the object position and scale by joint learning of box regression and pair-wise similarity. Furtherly, the DenseTracker is trained end-to-end on large-scale datasets including ImageNet Video (VID) and ALOV300++. The DenseTracker runs in 25 fps on GPU and achieves the state-of-the-art performance on two public benchmarks of OTB50 and VOT2016. Fei Zhao 0008, Ming Tang 0001, Yi Wu 0001, Jinqiao Wang |
ICME | 2 |
| 2017 | Fast Deep Matting for Portrait Animation on Mobile PhoneabstractImage matting plays an important role in image and video editing. However, the formulation of image matting is inherently ill-posed. Traditional methods usually employ interaction to deal with the image matting problem with trimaps and strokes, and cannot run on the mobile phone in real-time. In this paper, we propose a real-time automatic deep matting approach for mobile devices. By leveraging the densely connected blocks and the dilated convolution, a light full convolutional network is designed to predict a coarse binary mask for portrait image. And a feathering block, which is edge-preserving and matting adaptive, is further developed to learn the guided filter and transform the binary mask into alpha matte. Finally, an automatic portrait animation system based on fast deep matting is built on mobile devices, which does not need any interaction and can realize real-time matting with 15 fps. The experiments show that the proposed approach achieves comparable results with the state-of-the-art matting solvers. Bingke Zhu, Yingying Chen 0003, Jinqiao Wang, Si Liu 0001, Bo Zhang 0069, Ming Tang 0001 |
ACM Multimedia | 6 |
| 2017 | Two-Stream Deep Correlation Network for Frontal Face RecoveryabstractPose and textural variations are two dominant factors to affect the performance of face recognition. It is widely believed that generating the corresponding frontal face from a face image of an arbitrary pose is an effective step toward improving the recognition performance. In the literature, however, the frontal face is generally recovered by only exploring textural characteristic. In this letter, we propose a two-stream deep correlation network, which incorporates both geometric and textural features for frontal face recovery. Given a face image under an arbitrary pose as input, geometric and textural characteristics are first extracted from two separate streams. The extracted characteristics are then fused through the proposed multiplicative patch correlation layer. These two steps are integrated into one network for end-to-end training and prediction, which is demonstrated effective compared with state-of-the-art methods on the benchmark datasets. Ting Zhang 0006, Qiulei Dong, Ming Tang 0001, Zhanyi Hu |
IEEE Signal Process. Lett. | 3 |
| 2017 | Hierarchical and Networked Vehicle Surveillance in ITS: A SurveyabstractTraffic surveillance has become an important topic in intelligent transportation systems (ITSs), which is aimed at monitoring and managing traffic flow. With the progress in computer vision, video-based surveillance systems have made great advances on traffic surveillance in ITSs. However, the performance of most existing surveillance systems is susceptible to challenging complex traffic scenes (e.g., object occlusion, pose variation, and cluttered background). Moreover, existing related research is mainly on a single video sensor node, which is incapable of addressing the surveillance of traffic road networks. Accordingly, we present a review of the literature on the video-based vehicle surveillance systems in ITSs. We analyze the existing challenges in video-based surveillance systems for the vehicle and present a general architecture for video surveillance systems, i.e., the hierarchical and networked vehicle surveillance, to survey the different existing and potential techniques. Then, different methods are reviewed and discussed with respect to each module. Applications and future developments are discussed to provide future needs of ITS services. Bin Tian 0003, Brendan Tran Morris, Ming Tang 0001, Yuqiang Liu, Yanjie Yao, Chao Gou, Dayong Shen, Shaohu Tang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2015 | Multi-kernel Correlation Filter for Visual TrackingabstractCorrelation filter based trackers are ranked top in terms of performances. Nevertheless, they only employ a single kernel at a time. In this paper, we will derive a multi-kernel correlation filter (MKCF) based tracker which fully takes advantage of the invariance-discriminative power spectrums of various features to further improve the performance. Moreover, it may easily introduce location and representation errors to search several discrete scales for the proper one of the object bounding box, because normally the discrete candidate scales are determined and the corresponding feature pyramid are generated ahead of searching. In this paper, we will propose a novel and efficient scale estimation method based on optimal bisection search and fast evaluation of features. Our scale estimation method is the first one that uses the truly minimal number of layers of feature pyramid and avoids constructing the pyramid before searching for proper scales. Ming Tang 0001 |
ICCV | 1 |
| 2015 | Hierarchical and Networked Vehicle Surveillance in ITS: A SurveyabstractTraffic surveillance has become an important topic in intelligent transportation systems (ITSs), which is aimed at monitoring and managing traffic flow. With the progress in computer vision, video-based surveillance systems have made great advances on traffic surveillance in ITSs. However, the performance of most existing surveillance systems is susceptible to challenging complex traffic scenes (e.g., object occlusion, pose variation, and cluttered background). Moreover, existing related research is mainly on a single video sensor node, which is incapable of addressing the surveillance of traffic road networks. Accordingly, we present a review of the literature on the video-based vehicle surveillance systems in ITSs. We analyze the existing challenges in video-based surveillance systems for the vehicle and present a general architecture for video surveillance systems, i.e., the hierarchical and networked vehicle surveillance, to survey the different existing and potential techniques. Then, different methods are reviewed and discussed with respect to each module. Applications and future developments are discussed to provide future needs of ITS services. Bin Tian 0003, Brendan Tran Morris, Ming Tang 0001, Yuqiang Liu, Yanjie Yao, Chao Gou, Dayong Shen, Shaohu Tang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2014 | Robust visual tracking via augmented kernel SVM
Yancheng Bai, Ming Tang 0001 |
Image Vis. Comput. | 2 |
| 2014 | Object Tracking via Robust Multitask Sparse RepresentationabstractSparse representation has been applied to the object tracking problem. Mining the self-similarities between particles via multitask learning can improve tracking performance. However, some particles may be different from others when they are sampled from a large region. Imposing all particles share the same structure may degrade the results. To overcome this problem, we propose a tracking algorithm based on robust multitask sparse representation (RMTT) in this letter. When we learn the particle representations, we decompose the sparse coefficient matrix into two parts in our algorithm. Joint sparse regularization is imposed on one coefficient matrix while element-wise sparse regularization is imposed on another matrix. The former regularization exploits self-similarities of particles while the later one considers the differences between them. Experiments on the benchmark data show the superior performance over other state-of-art algorithms. Yancheng Bai, Ming Tang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2012 | Robust tracking via weakly supervised ranking SVMabstractAppearance model is a key component of tracking algorithms. Most existing approaches utilize the object information contained in the current and previous frames to construct the object appearance model and locate the object with the model in frame t + 1. This method may work well if the object appearance just fluctuates in short time intervals. Nevertheless, suboptimal locations will be generated in frame t + 1 if the visual appearance changes substantially from the model. Then, continuous changes would accumulate errors and finally result in a tracking failure. To copy with this problem, in this paper we propose a novel algorithm - online Laplacian ranking support vector tracker (LRSVT) - to robustly locate the object. The LRSVT incorporates the labeled information of the object in the initial and the latest frames to resist the occlusion and adapt to the fluctuation of the visual appearance, and the weakly labeled information from frame t + 1 to adapt to substantial changes of the appearance. Extensive experiments on public benchmark sequences show the superior performance of LRSVT over some state-of-the-art tracking algorithms. Yancheng Bai, Ming Tang 0001 |
CVPR | 2 |
| 2012 | Robust Tracking With Discriminative Ranking ListsabstractIn this paper, we propose a novel tracking algorithm, i.e., the discriminative ranking list-based tracker (DRLTracker). The DRLTracker models the target object and its local background by using ranking lists of patches of different scales within object bounding boxes. The ranking list of each of such patches is its K nearest neighbors. Patches of the same scale with ranking lists of high purity values (meaning high probabilities to be on the target object) and some confusable background patches constitute the object model under that scale. A pair of object models of two different scales collaborate to determine which patches may belong to the target object in the next frame. The DRLTracker can effectively alleviate the distraction problem, and its superior ability over several representative and state-of-the-art trackers is demonstrated through extensive experiments. Ming Tang 0001, Xi Peng 0003 |
IEEE Trans. Image Process. | 1 |
| 2011 | Robust visual tracking via ranking SVMabstractIn this paper, we tackle the tracking problem in a quite other viewpoint, ranking. First, the ranking SVM is employed to learn a ranking function. Then, the ranking function ranks every instance sampled from the next frame, and the instance with the most preferred ranking score is assumed to be the object. Experiments of extensively quantitative and qualitative comparisons on public videos show the superior performance of our tracker over several state-of-the-art tracking algorithms. Yancheng Bai, Ming Tang 0001 |
ICIP | 2 |
| 2010 | Robust Tracking with Discriminative Ranking Lists
Ming Tang 0001, Xi Peng 0003 |
ACCV (1) | 1 |
| 2009 | Combining Discriminative and Descriptive Models for Tracking
Ming Tang 0001 |
ACCV (1) | 3 |
| 2009 | Random patch based video tracking via boosting the relative spacesabstractIn this paper, we propose a new visual tracking method based on the recently popular tracking-as-classification idea. We concentrate on exploring the intra-class variance of the foreground target to construct and update a classification based tracker. In our approach, foreground target is represented by a set of model patches. Different types of features are jointly used to represent those patches. Individual weak learners are trained based on each model patch's relative space. AdaBoost framework is applied to choose those weak classifiers to combine a strong classifier as the tracker for next frame. Moreover, with the new tracking result, the tracker is adjusted adaptively according to the change of scene to keep itself discriminative during the entire sequence. We demonstrate the effectiveness of our approach with comparison results on common video sequences. Ming Tang 0001 |
ICASSP | 3 |
| 2009 | Learning local features for object categorizationabstractIn this paper, for every local feature, we propose to learn its similar local features across all positive images, instead of using heuristic distance as similarity measure. Specifically, multiple instance learning (MIL) is employed to simultaneously determine the similar points of a local feature and learn its corresponding discriminative function which can be regarded as some kind of similarity measure. For each local feature, a weak learner is constructed based on such similarity measure. Then AdaBoost selects the most discriminative local features and combines them to form a strong classifier. Experimental results show encouraging performance of our method. Ming Tang 0001, Shi Chen 0008, Jinqiao Wang, Hanqing Lu, Songde Ma |
ICME | 2 |
| 2008 | A novel contextual descriptors for category recognitionabstractIn this paper, we propose a novel contextual descriptor which combines the contextual information and local appearance. Based on Gibbs distribution, a local descriptor is designed. By assembling the contextual information and local descriptors, a new partial contextual descriptor (PCD) is finally presented. Combining Pyramid Match Kernel (PMK) and SVM, we test our new descriptor and obtain higher average precision of classification than using local appearance descriptor. Ming Tang 0001, Jian Cheng 0001, Jinqiao Wang, Hanqing Lu, Songde Ma |
ICME | 2 |
| 2008 | Boosting relative spaces for categorizing objects with large intra-class variationabstractIn this paper, a novel method for object categorization is proposed. We first analyze the phenomenon of large intra-class variation and attribute it to the "subcategory" problem. To reveal the local and distinct properties of the different subcategories, relative spaces are constructed. Then the weighted FLDs (Fisher Linear Discriminant) as weak learners trained in relative spaces are integrated with the boosting framework to form the final classifier. Experiments on 8 categories from Caltech database show the effectiveness of our algorithm. Ming Tang 0001, Jinqiao Wang, Hanqing Lu, Songde Ma |
ACM Multimedia | 2 |
| 2006 | Accelerated Convergence Using Dynamic Mean Shift
Kai Zhang 0001, James T. Kwok, Ming Tang 0001 |
ECCV (2) | 3 |
| 2005 | Applying Neighborhood Consistency for Fast Clustering and Kernel Density EstimationabstractNearest neighborhood consistency is an important concept in statistical pattern recognition, which underlies the well-known k-nearest neighbor method. In this paper, we combine this idea with kernel density estimation based clustering, and derive the fast mean shift algorithm (FMS). FMS greatly reduces the complexity of feature space analysis, resulting satisfactory precision of classification. More importantly, we show that with FMS algorithm, we are in fact relying on a conceptually novel approach of density estimation, the fast kernel density estimation (FKDE) for clustering. The FKDE combines smooth and non-smooth estimators and thus inherits advantages from both. Asymptotic analysis reveals the approximation of the FKDE to standard kernel density estimator. Data clustering and image segmentation experiments demonstrate the efficiency of FMS. Kai Zhang 0001, Ming Tang 0001, James T. Kwok |
CVPR (2) | 2 |
| 2004 | A new robust circular Gabor based object matching by using weighted Hausdorff distance
Zhenfeng Zhu, Ming Tang 0001, Hanqing Lu |
Pattern Recognit. Lett. | 2 |
| 2002 | Corrections to "General Scheme of Region Competition Based on Scale Space"
Ming Tang 0001, Songde Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | General Scheme of Region Competition Based on Scale SpaceabstractWe propose a general scheme of region competition (GSRC) for image segmentation based on scale space. First, we present a novel classification algorithm to cluster the image feature data according to the generally defined peaks under a certain scale and a scale space-based classification scheme to classify the pixels by grouping the resultant feature data clusters into several classes with a standard classification algorithm. Next, to reduce the resultant segmentation error, we develop a nonparametric probability model from which the functional for GSRC is derived. We also design a general and formal approach to automatically determine the initial regions. Finally, we propose the kernel procedure of GSRC which segments an image by minimizing the functional. The strategy adopted by GSRC is first to label pixels whose corresponding regions can be determined in large likelihood, and then to fine-tune the final regions with the help of the nonparametric probability model, boundary smoothing, and region competition. Although the description of the scheme is nonparametric in this paper, GSRC can also work parametrically if all nonparametric procedures in this paper are substituted with the parametric counterparts. Ming Tang 0001, Songde Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Semantically Homogeneous Segmentation with Nonparametric Region CompetitionabstractPresents a nonparametric region competition algorithm which combines scale-space clustering and region competition to segment the image. It also proposes a formal and general procedure to automatically find the initial regions. Our algorithm can also segment an image into regions which are not homogeneous in the sense of statistics, but is homogeneous in the sense of semantics with respect to the segmentation context. Ming Tang 0001, Jing Xiao 0002, Songde Ma |
ICPR | 1 |
| 2000 | Two-Step Classification Based on Scale SpaceabstractA new two-step classification scheme based on nonparametric estimation of density function and scale-space filtering is presented. This scheme is able to combine traditional supervised classification techniques with clustering. After nonparametric estimation of the underlying density function, this scheme utilizes scale-space filtering and a novel classification algorithm to extract the intrinsic basic structure of the data. Then, depending on applications, one of the traditional clustering or classification techniques may be employed to obtain a final high level data structure. Ming Tang 0001, Jing Xiao 0002, Songde Ma |
ICPR | 1 |
| 2000 | Model-based adaptive enhancement of far infrared image sequences
Ming Tang 0001, Songde Ma, Jing Xiao 0002 |
Pattern Recognit. Lett. | 1 |