Jian Li 0062

dblp:33/5448-62 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0002-0242-6481ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LLM-Oriented Token-Adaptive Knowledge Distillation
abstract
Knowledge Distillation (KD) is a key technique for compressing Large-scale Language Models (LLMs), but prevailing logit-based methods employ static strategies misaligned with the student’s dynamic learning process. By treating all tokens indiscriminately with a fixed temperature, these methods result in suboptimal knowledge transfer. To address this, we propose LLM-oriented token-Adaptive Knowledge Distillation (AdaKD), a framework that adapts the distillation process to each token’s real-time learning state. AdaKD consists of two synergistic modules driven by a unified token difficulty metric. First, the Loss-driven Adaptive Token Focusing (LATF) module dynamically concentrates distillation on valuable tokens by monitoring the student’s learning stability. Second, Inverse Difficulty Temperature Scaling (IDTS) introduces a counterintuitive token-level temperature: low for difficult tokens to target error correction, and high for easy tokens to learn the teacher’s smooth output distribution for better generalization. As a plug-and-play framework, AdaKD consistently improves performance across diverse distillation methods, model architectures, and benchmarks.
Xurong Xie, Zhucun Xue, Jiafu Wu, Jian Li 0062, Yabiao Wang, Xiaobin Hu, Yong Liu 0007, Jiangning Zhang
AAAI4
2026 Disco-RAG: Discourse-Aware Retrieval-Augmented Generation
abstract
Dongqi Liu, Hang Ding, Qiming Feng, Xurong Xie, Zhucun Xue, Chengjie Wang, Jian Li, Jiangning Zhang, Yabiao Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hang Ding, Qiming Feng, Xurong Xie, Zhucun Xue, Chengjie Wang 0001, Jian Li 0062, Jiangning Zhang, Yabiao Wang
ACL (1)7
2026 LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Weiheng Lu, Jian Li 0062, An Yu, Ming-Ching Chang
ICPR (5)2
2026 Towards fine-grained vision-language alignment for few-shot anomaly detection
Yuanting Fan, Jun Liu 0116, Xiaochen Chen, Bin-Bin Gao, Jian Li 0062, Yong Liu 0032, Jinlong Peng, Chengjie Wang 0001
Pattern Recognit.5
2026 Active Dataset Distillation via Dual-Space Informative Matching
abstract
Dataset distillation improves neural network training efficiency by compressing large real datasets into compact synthetic datasets. Existing methods typically optimize matching objectives, such as aligning gradients, features, and trajectories between the synthetic and original datasets to ensure the distilled data retains essential properties for model training. However, many of these approaches rely on predefined distillation pools to streamline the process or treat all real data points equally, overlooking the dynamic nature of the synthetic dataset's training requirements during optimization. To address these limitations, we propose Active Dataset Distillation via Dual-Space Informative Matching (ACDD), an active learning-based algorithm that dynamically selects the most informative real data subset to align with the synthetic dataset's evolving needs. By adaptively refining the distillation pool, ACDD enhances training efficiency and generalization while ensuring the synthetic dataset effectively captures the original data's key characteristics. ACDD operates through two interconnected loops: the dual-space active loop (DAL) and the distillation loop. DAL plays a key role by dynamically selecting samples that balance diversity and uncertainty, adding them to the target distillation pool to meet the evolving informational needs of the current distillation loop. As a result, ACDD enables the synthetic dataset to achieve superior performance compared to SOTA methods across multiple benchmarks, including SVHN, CIFAR-10, CIFAR-100, TinyImageNet, and ImageNet subset. Moreover, ACDD reduces the required real dataset to just 20%-40% of the original, demonstrating its efficiency and effectiveness in data distillation.
Ding Qi, Jian Li 0062, Shuguang Dou, Junyao Gao 0002, Yabiao Wang, Bo Zhao 0015, Cairong Zhao
IEEE Trans. Image Process.2
2025 Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
abstract
Multimodal Large Language Models (MLLMs) have displayed remarkable performance in multi-modal tasks, particularly in visual comprehension. However, we reveal that MLLMs often generate incorrect answers even when they understand the visual content. To this end, we manually construct a benchmark with 12 categories and design evaluation metrics that assess the degree of error in MLLM responses even when the visual content is seemingly understood. Based on this benchmark, we test 15 leading MLLMs and analyze the distribution of attention maps and logits of some MLLMs. Our investigation identifies two primary issues: 1) most instruction tuning datasets predominantly feature questions that "directly" relate to the visual content, leading to a bias in MLLMs’ responses to other indirect questions, and 2) MLLMs’ attention to visual tokens is notably lower than to system and question tokens. We further observe that attention scores between questions and visual tokens as well as the model’s confidence in the answers are lower in response to misleading questions than to straightforward ones. To address the first challenge, we introduce a paired positive and negative data construction pipeline to diversify the dataset. For the second challenge, we propose to enhance the model’s focus on visual content during decoding by refining the text and visual prompt. For the text prompt, we propose a content guided refinement strategy that performs preliminary visual content analysis to generate structured information before answering the question. Additionally, we employ a visual attention refinement strategy that highlights question-relevant visual tokens to increase the model’s attention to visual content that aligns with the question. Extensive experiments demonstrate that these challenges can be significantly mitigated with our proposed dataset and techniques.
Yexin Liu, Zhengyang Liang, Yueze Wang, Xianfeng Wu, Muyang He, Jian Li 0062, Zheng Liu 0011, Harry Yang, Ser-Nam Lim, Bo Zhao 0015
CVPR7
2025 Towards Universal Dataset Distillation via Task-Driven Diffusion
abstract
Dataset distillation (DD) condenses key information from large-scale datasets into smaller synthetic datasets, reducing storage and computational costs for training networks. However, most recent research has primarily focused on image classification tasks, with limited exploration in detection and segmentation. Two key challenges remain: (i) Task Optimization Heterogeneity, where existing methods focus on class-level information but fail to address the diverse needs of detection and segmentation, and (ii) Inflexible Image Generation, where current generation methods rely on global updates for single-class targets and lack localized optimization for specific object regions. To address these challenges, we propose UniDD, a universal dataset distillation framework built on a task-driven diffusion model for diverse DD tasks, as shown in Fig. 1. Our approach operates in two stages: Universal Task Knowledge Mining, which captures task-relevant information through task-specific proxy model training, and Universal Task-Driven Diffusion, where these proxies guide the diffusion process to generate task-specific synthetic images. Extensive experiments across ImageNet-1K, Pascal VOC, and MS COCO demonstrate that UniDD consistently outperforms state-of-the-art methods. In particular, on ImageNet-1K with IPC-10, UniDD surpasses previous diffusion-based methods by 6.1%, while also reducing deployment costs.
Ding Qi, Jian Li 0062, Junyao Gao 0002, Shuguang Dou, Ying Tai, Jianlong Hu, Bo Zhao 0015, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao
CVPR2
2025 MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
abstract
In the field of industrial inspection, Multimodal Large Language Models (MLLMs) have a high potential to renew the paradigms in practical applications due to their robust language capabilities and generalization abilities. However, despite their impressive problem-solving skills in many domains, MLLMs' ability in industrial anomaly detection has not been systematically studied. To bridge this gap, we present MMAD, a full-spectrum MLLM benchmark in industrial Anomaly Detection. We defined seven key subtasks of MLLMs in industrial inspection and designed a novel pipeline to generate the MMAD dataset with 39,672 questions for 8,366 industrial images. With MMAD, we have conducted a comprehensive, quantitative evaluation of various state-of-the-art MLLMs. The commercial models performed the best, with the average accuracy of GPT-4o models reaching 74.9\%. However, this result falls far short of industrial requirements. Our analysis reveals that current MLLMs still have significant room for improvement in answering questions related to industrial anomalies and defects. We further explore two training-free performance enhancement strategies to help models improve in industrial scenarios, highlighting their promising potential for future research. The code and data are available at https://github.com/jam-cc/MMAD.
Xi Jiang 0009, Jian Li 0062, Hanqiu Deng, Yong Liu 0032, Bin-Bin Gao, Chengjie Wang 0001, Feng Zheng 0001
ICLR2
2025 BoxSeg: Quality-Aware and Peer-Assisted Learning for Box-supervised Instance Segmentation
Jinxiang Lai, Jiawei Zhan, Jian Li 0062, Bin-Bin Gao, Jun Liu 0116, Jie Zhang 0006, Song Guo 0001
ACM Multimedia4
2025 MAP: Parameter-Efficient Tuning for Referring Expression Comprehension via Multi-Modal Adaptive Positional Encoding
abstract
This paper studies the challenging task of Referring Expression Comprehension (REC), which aims at detecting the text-referred target object in an input image. To achieve this, most recent works attempt to adapt powerful pretrained models through integrating additional structures (e.g., low-rank adaptation (LoRA) or adapter modules) to enable efficient parameter tuning. However, all these methods process pretrained features in a position-agnostic manner. This will limit their effectiveness in REC tasks, where the positional information is essential to correctly localize the target object. To this end, we propose a novel parameter-efficient tuning approach, named Multi-Modal Adaptive Positional Encoding (MAP), which addresses the above problem from a new perspective of positional encoding. More specifically, MAP first generates initial positional embeddings for different visual encoder layers from a set of learnable vectors, and then adjusts them adaptively based on spatial-wise visual-linguistic correlations of input data. In this way, the positional information of different image tokens can be appropriately modeled and utilized by MAP, thus making it more applicable to REC tasks. Extensive experiments on five widely-used datasets demonstrate that MAP achieves comparable results to full fine-tuning methods with much fewer extra parameters and outperforms other parameter-efficient tuning approaches. Our source code is available at: https://github.com/Mr-Bigworth/MAP.
Ruilin Yao, Tianyu Zou, Bo Zhang 0069, Jian Li 0062, Shengwu Xiong 0001, Shili Xiong
ACM Multimedia5
2025 VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model
abstract
With the growing requirement for natural human-computer interaction, speech-based systems receive increasing attention as speech is one of the most common forms of daily communication. However, the existing speech models still experience high latency when generating the first audio token during streaming, which poses a significant bottleneck for deployment. To address this issue, we propose VITA-Audio, an end-to-end large speech model with fast audio-text token generation. Specifically, we introduce a lightweight Multiple Cross-modal Token Prediction (MCTP) module that efficiently generates multiple audio tokens within a single model forward pass, which not only accelerates the inference but also significantly reduces the latency for generating the first audio in streaming scenarios. In addition, a four-stage progressive training strategy is explored to achieve model acceleration with minimal loss of speech quality. To our knowledge, VITA-Audio is the first multi-modal large language model capable of generating audio output during the first forward pass, enabling real-time conversational capabilities with minimal latency. VITA-Audio is fully reproducible and is trained on open-source data only. Experimental results demonstrate that our model achieves an inference speedup of 3~5x at the 7B parameter scale, but also significantly outperforms open-source models of similar model size on multiple benchmarks for automatic speech recognition (ASR), text-to-speech (TTS), and spoken question answering (SQA) tasks.
Zuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao, Lijiang Li, Peixian Chen, Mengdan Zhang, Jian Li 0062, Jinlong Peng, Haoyu Cao 0001, Ke Li 0015, Rongrong Ji, Xing Sun 0001
NeurIPS9
2025 TAPCNet: Tactile-Assisted Point Cloud Completion Network via Iterative Fusion Strategy
abstract
ABSTRACT With the development of the 3D point cloud field in recent years, point cloud completion of 3D objects has increasingly attracted researchers' attention. Point cloud data can accurately express the shape information of 3D objects at different resolutions, but the original point clouds collected directly by various 3D scanning equipment are often incomplete and have uneven density. Tactile is one distinctive way to perceive the 3D shape of an object. Tactile point clouds can provide local shape information for unknown areas during completion, which is a valuable complement to the point cloud data acquired with visual devices. In order to effectively improve the effect of point cloud completion using tactile information, the authors propose an innovative tactile‐assisted point cloud completion network, TAPCNet. This network is the first neural network customised for the input of tactile point clouds and incomplete point clouds, which can fuse two types of point cloud information in the feature domain. Besides, a new dataset named 3DVT was rebuilt, to fit the proposed network model. Based on the tactile fusion strategy and related modules, multiple comparative experiments were conducted by controlling the quantity of tactile point clouds on the 3DVT dataset. The experimental data illustrates that TAPCNet can outperform the state‐of‐the‐art methods in the benchmark.
Yangrong Liu, Jian Li 0062, Huaiyu Wang, Haorao Shen, Qin Wang 0011
IET Comput. Vis.2
2024 LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
abstract
Visual Spatial Description (VSD) aims to generate texts that describe the spatial relationships between objects within images. Traditional visual spatial relationship classification (VSRC) methods typically output the spatial relationship between two objects in an image, often neglecting world knowledge and lacking general language capabilities. In this paper, we propose a Large Language-and-Vision Assistant for Visual Spatial Description, named LLaVA-VSD, which is designed for the classification, description, and open-ended description of visual spatial relationships. Specifically, the model first constructs a visual spatial instruction-following dataset using given figure-caption pairs for the three tasks. It then employs LoRA to fine-tune a Large Language and Vision Assistant for VSD, which has 13 billion parameters and supports high-resolution images. Finally, a large language model is used to refine the generated sentences, enhancing their diversity and accuracy. LLaVA-VSD demonstrates excellent multimodal conversational capabilities and can follow open-ended instructions to assist with inquiries about object relationships in images.
Yizhang Jin, Jian Li 0062, Jiangning Zhang, Jianlong Hu, Zhenye Gan, Xin Tan 0002, Yong Liu 0032, Yabiao Wang, Chengjie Wang 0001, Lizhuang Ma
ACM Multimedia2
2024 Fetch and Forge: Efficient Dataset Condensation for Object Detection
abstract
Dataset condensation (DC) is an emerging technique capable of creating compact synthetic datasets from large originals while maintaining considerable performance. It is crucial for accelerating network training and reducing data storage requirements. However, current research on DC mainly focuses on image classification, with less exploration of object detection. This is primarily due to two challenges: (i) the multitasking nature of object detection complicates the condensation process, and (ii) Object detection datasets are characterized by large-scale and high-resolution data, which are difficult for existing DC methods to handle. As a remedy, we propose DCOD, the first dataset condensation framework for object detection. It operates in two stages: Fetch and Forge, initially storing key localization and classification information into model parameters, and then reconstructing synthetic images via model inversion. For the complex of multiple objects in an image, we propose Foreground Background Decoupling to centrally update the foreground of multiple instances and Incremental PatchExpand to further enhance the diversity of foregrounds. Extensive experiments on various detection datasets demonstrate the superiority of DCOD. Even at an extremely low compression rate of 1\%, we achieve 46.4\% and 24.7\% $\text{AP}_{50}$ on the VOC and COCO, respectively, significantly reducing detector training duration.
Ding Qi, Jian Li 0062, Jinlong Peng, Bo Zhao 0015, Shuguang Dou, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao
NeurIPS2
2023 Learning from Noisy Labels with Decoupled Meta Label Purifier
abstract
Training deep neural networks (DNN) with noisy labels is challenging since DNN can easily memorize inaccurate labels, leading to poor generalization ability. Recently, the meta-learning based label correction strategy is widely adopted to tackle this problem via identifying and correcting potential noisy labels with the help of a small set of clean validation data. Although training with purified labels can effectively improve performance, solving the meta-learning problem inevitably involves a nested loop of bi-level optimization between model weights and hyper-parameters (i.e., label distribution). As compromise, previous methods resort to a coupled learning process with alternating update. In this paper, we empirically find such simultaneous optimization over both model weights and label distribution can not achieve an optimal routine, consequently limiting the representation ability of backbone and accuracy of corrected labels. From this observation, a novel multi-stage label purifier named DMLP is proposed. DMLP decouples the label correction process into label-free representation learning and a simple meta label purifier, In this way, DMLP can focus on extracting discriminative feature and label correction in two distinctive stages. DMLP is a plug-and-play label purifier, the purified labels can be directly reused in naive end-to-end network retraining or other robust learning methods, where state-of-the-art results are obtained on several synthetic and real-world noisy datasets, especially under high noise levels. Code is available at https://github.com/yuanpengtu/DMLP.
Yuanpeng Tu, Boshen Zhang, Yuxi Li 0009, Liang Liu 0007, Jian Li 0062, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao
CVPR5
2023 Learning with Noisy labels via Self-supervised Adversarial Noisy Masking
abstract
Collecting large-scale datasets is crucial for training deep models, annotating the data, however, inevitably yields noisy labels, which poses challenges to deep learning algorithms. Previous efforts tend to mitigate this problem via identifying and removing noisy samples or correcting their labels according to the statistical properties (e.g., loss values) among training samples. In this paper, we aim to tackle this problem from a new perspective, delving into the deep feature maps, we empirically find that models trained with clean and mislabeled samples manifest distinguishable activation feature distributions. From this observation, a novel robust training approach termed adversarial noisy masking is proposed. The idea is to regularize deep features with a label quality guided masking scheme, which adaptively modulates the input data and label simultaneously, preventing the model to overfit noisy samples. Further, an auxiliary task is designed to reconstruct input data, it naturally provides noise-free self-supervised signals to rein-force the generalization ability of models. The proposed method is simple yet effective, it is tested on synthetic and real-world noisy datasets, where significant improvements are obtained over previous methods. Code is available at https://github.com/yuanpengtu/SANM.
Yuanpeng Tu, Boshen Zhang, Yuxi Li 0009, Liang Liu 0007, Jian Li 0062, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao
CVPR5
2023 Rethinking Mobile Block for Efficient Attention-based Models
abstract
This paper focuses on developing modern, efficient, lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Inverted Residual Block (IRB) serves as the infrastructure for lightweight CNNs, but no counterpart has been recognized by attention-based studies. This work rethinks lightweight infrastructure from efficient IRB and effective components of Transformer from a unified perspective, extending CNN-based IRB to attention-based models and abstracting a one-residual Meta Mobile Block (MMB) for lightweight model design. Following simple but effective design criterion, we deduce a modern Inverted Residual Mobile Block (iRMB) and build a ResNetlike Efficient MOdel (EMO) with only iRMB for down-stream tasks. Extensive experiments on ImageNet-1K, COCO2017, and ADE20K benchmarks demonstrate the superiority of our EMO over state-of-the-art methods, e.g., EMO-1M/2M/5M achieve 71.5, 75.1, and 78.4 Top-1 that surpass equal-order CNN-/Attention-based models, while trading-off the parameter, efficiency, and accuracy well: running 2.8-4.0× ↑ faster than EdgeNeXt on iPhone14.
Jiangning Zhang, Xiangtai Li, Jian Li 0062, Liang Liu 0007, Zhucun Xue, Boshen Zhang, Zhengkai Jiang 0001, Tianxin Huang, Yabiao Wang, Chengjie Wang 0001
ICCV3
2023 PVG: Progressive Vision Graph for Vision Recognition
abstract
Convolution-based and Transformer-based vision backbone networks process images into the grid or sequence structures, respectively, which are inflexible for capturing irregular objects. Though Vision GNN (ViG) adopts graph-level features for complex images, it has some issues, such as inaccurate neighbor node selection, expensive node information aggregation calculation, and over-smoothing in the deep layers. To address the above problems, we propose a Progressive Vision Graph (PVG) architecture for vision recognition task. Compared with previous works, PVG contains three main components: 1) Progressively Separated Graph Construction (PSGC) to introduce second-order similarity by gradually increasing the channel of the global graph branch and decreasing the channel of local branch as the layer deepens; 2) Neighbor nodes information aggregation and update module by using Max pooling and mathematical Expectation (MaxE) to aggregate rich neighbor information; 3) Graph error Linear Unit (GraphLU) to enhance low-value information in a relaxed form to reduce the compression of image detail information for alleviating the over-smoothing. Extensive experiments on mainstream benchmarks demonstrate the superiority of PVG over state-of-the-art methods, e.g., our PVG-S obtains 83.0% Top-1 accuracy on ImageNet-1K that surpasses GNN-based ViG-S by +0.9↑ with the parameters reduced by 18.5%, while the largest PVG-B obtains 84.2% that has +0.5↑ improvement than ViG-B. Furthermore, our PVG-S obtains +1.3↑ box AP and +0.4↑ mask AP gains than ViG-S on COCO dataset.
Jiafu Wu, Jian Li 0062, Jiangning Zhang, Boshen Zhang, Mingmin Chi, Yabiao Wang, Chengjie Wang 0001
ACM Multimedia2
2023 Highly efficient twin-field quantum key distribution with neural networks
Qingqing Jiang, Hua-Jian Ding, Ming-Shuo Sun, Jiaxin Xu, Shipeng Xie, Jian Li 0062, Guigen Zeng, Xingyu Zhou 0006, Qin Wang 0011
Sci. China Inf. Sci.9
2022 SCSNet: An Efficient Paradigm for Learning Simultaneously Image Colorization and Super-resolution
abstract
In the practical application of restoring low-resolution gray-scale images, we generally need to run three separate processes of image colorization, super-resolution, and dows-sampling operation for the target device. However, this pipeline is redundant and inefficient for the independent processes, and some inner features could have been shared. Therefore, we present an efficient paradigm to perform Simultaneously Image Colorization and Super-resolution (SCS) and propose an end-to-end SCSNet to achieve this goal. The proposed method consists of two parts: colorization branch for learning color information that employs the proposed plug-and-play Pyramid Valve Cross Attention (PVCAttn) module to aggregate feature maps between source and reference images; and super-resolution branch for integrating color and texture information to predict target images, which uses the designed Continuous Pixel Mapping (CPM) module to predict high-resolution images at continuous magnification. Furthermore, our SCSNet supports both automatic and referential modes that is more flexible for practical application. Abundant experiments demonstrate the superiority of our method for generating authentic images over state-of-the-art methods, e.g., averagely decreasing FID by 1.8 and 5.1 compared with current best scores for automatic and referential modes, respectively, while owning fewer parameters (more than x2) and faster running speed (more than x3).
Jiangning Zhang, Chao Xu 0023, Jian Li 0062, Yabiao Wang, Ying Tai, Yong Liu 0007
AAAI3
2021 LSTC: Boosting Atomic Action Detection with Long-Short-Term Context
abstract
In this paper, we place the atomic action detection problem intoa Long-Short Term Context (LSTC) to analyze how the temporalreliance among video signals affect the action detection results. Todo this, we decompose the action recognition pipeline into short-term and long-term reliance, in terms of the hypothesis that the twokinds of context are conditionally independent given the objectiveaction instance. Within our design, a local aggregation branch isutilized to gather dense and informative short-term cues, while ahigh order long-term inference branch is designed to reason theobjective action class from high-order interaction between actor andother person or person pairs. Both branches independently predictthe context-specific actions and the results are merged in the end.We demonstrate that both temporal grains are beneficial to atomicaction recognition. On the mainstream benchmarks of atomic actiondetection, our design can bring significant performance gain fromthe existing state-of-the-art pipeline.
Yuxi Li 0009, Boshen Zhang, Jian Li 0062, Yabiao Wang, Weiyao Lin, Chengjie Wang 0001, Feiyue Huang
ACM Multimedia3
2021 ASFD: Automatic and Scalable Face Detector
abstract
Along with current multi-scale based detectors, Feature Aggregation and Enhancement (FAE) modules have shown superior performance gains for cutting-edge object detection. However, these hand-crafted FAE modules show inconsistent improvements on face detection, which is mainly due to the significant distribution difference between its training and applying corpus, i.e. COCO vs. WIDER Face. To tackle this problem, we essentially analyse the effect of data distribution, and consequently propose to search an effective FAE architecture, termed AutoFAE by a differentiable architecture search, which outperforms all existing FAE modules in face detection with a considerable margin. Upon the found AutoFAE and existing backbones, a supernet is further built and trained, which automatically obtains a family of detectors under the different complexity constraints. Extensive experiments conducted on popular benchmarks, i.e. WIDER Face and FDDB, demonstrate the state-of-the-art performance-efficiency trade-off for the proposed automatic and scalable face detector (ASFD) family. In particular, our strong ASFD-D6 outperforms the best competitor with AP 96.7/96.2/92.1 on WIDER Face test, and the lightweight ASFD-D0 costs about 3.1 ms, i.e. more than 320 FPS, on the V100 GPU with VGA-resolution images.
Jian Li 0062, Yabiao Wang, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Yili Xia
ACM Multimedia1
2021 Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model
abstract
Inspired by biological evolution, we explain the rationality of Vision Transformer by analogy with the proven practical Evolutionary Algorithm (EA) and derive that both of them have consistent mathematical representation. Analogous to the dynamic local population in EA, we improve the existing transformer structure and propose a more efficient EAT model, and design task-related heads to deal with different tasks more flexibly. Moreover, we introduce the spatial-filling curve into the current vision transformer to sequence image data into a uniform sequential format. Thus we can design a unified EAT framework to address multi-modal tasks, separating the network architecture from the data format adaptation. Our approach achieves state-of-the-art results on the ImageNet classification task compared with recent vision transformer works while having smaller parameters and greater throughput. We further conduct multi-modal tasks to demonstrate the superiority of the unified EAT, \eg, Text-Based Image Retrieval, and our approach improves the rank-1 by +3.7 points over the baseline on the CSS dataset.
Jiangning Zhang, Chao Xu 0023, Jian Li 0062, Wenzhou Chen, Yabiao Wang, Ying Tai, Shuo Chen 0003, Chengjie Wang 0001, Feiyue Huang, Yong Liu 0007
NeurIPS3
2020 Fast Learning of Temporal Action Proposal via Dense Boundary Generator
abstract
Generating temporal action proposals remains a very challenging problem, where the main issue lies in predicting precise temporal proposal boundaries and reliable action confidence in long and untrimmed real-world videos. In this paper, we propose an efficient and unified framework to generate temporal action proposals named Dense Boundary Generator (DBG), which draws inspiration from boundary-sensitive methods and implements boundary classification and action completeness regression for densely distributed proposals. In particular, the DBG consists of two modules: Temporal boundary classification (TBC) and Action-aware completeness regression (ACR). The TBC aims to provide two temporal boundary confidence maps by low-level two-stream features, while the ACR is designed to generate an action completeness score map by high-level action-aware features. Moreover, we introduce a dual stream BaseNet (DSB) to encode RGB and optical flow information, which helps to capture discriminative boundary and actionness features. Extensive experiments on popular benchmarks ActivityNet-1.3 and THUMOS14 demonstrate the superiority of DBG over the state-of-the-art proposal generator (e.g., MGG and BMN).
Chuming Lin, Jian Li 0062, Yabiao Wang, Ying Tai, Donghao Luo 0001, Zhipeng Cui, Chengjie Wang 0001, Feiyue Huang, Rongrong Ji
AAAI2
2020 Learning Hierarchical Graph for Occluded Pedestrian Detection
abstract
Although pedestrian detection has made significant progress with the help of deep convolution neural networks, it is still a challenging problem to detect occluded pedestrians since the occluded ones can not provide sufficient information for classification and regression. In this paper, we propose a novel Hierarchical Graph Pedestrian Detector (HGPD), which integrates semantic and spatial relation information to construct two graphs named intra-proposal graph and inter-proposal graph, without relying on extra cues w.r.t visible regions. In order to capture the occlusion patterns and enhance features from visible regions, the intra-proposal graph considers body parts as nodes and assigns corresponding edge weights based on semantic relations between body parts. On the other hand, the inter-proposal graph adopts spatial relations between neighbouring proposals to provide additional proposal-wise context information for each proposal, which alleviates the lack of information caused by occlusion. We conduct extensive experiments on standard benchmarks of CityPersons and Caltech to demonstrate the effectiveness of our method. On CityPersons, our approach outperforms the baseline method by a large margin of 5.24pp on the heavy occlusion set, and surpasses all previous methods; on Caltech, we establish a new state of the art of 3.78% MR. Code is available at https://github.com/ligang-cs/PedestrianDetection-HGPD.
Jian Li 0062, Shanshan Zhang 0001, Jian Yang 0003
ACM Multimedia2
2019 DSFD: Dual Shot Face Detector
abstract
Recently, Convolutional Neural Network (CNN) has achieved great success in face detection. However, it remains a challenging problem for the current face detection methods owing to high degree of variability in scale, pose, occlusion, expression, appearance and illumination. In this Paper, we propose a novel detection network named Dual Shot face Detector(DSFD). which inherits the architecture of SSD and introduces a Feature Enhance Module (FEM) for transferring the original feature maps to extend the single shot detector to dual shot detector. Specially, progressive anchor loss (PAL) computed by using two set of anchors is adopted to effectively facilitate the features. Additionally, we propose an improved anchor matching (IAM) method by integrating novel data augmentation techniques and anchor design strategy in our DSFD to provide better initialization for the regressor. Extensive experiments on popular benchmarks: WIDER FACE (easy: 0.966, medium: 0.957, hard: 0.904) and FDDB ( discontinuous: 0.991, continuous: 0.862 ) demonstrate the superiority of DSFD over the state-of-the-art face detection methods (e.g., PyramidBox and SRN). Code will be made available upon publication.
Jian Li 0062, Yabiao Wang, Changan Wang, Ying Tai, Jianjun Qian, Jian Yang 0003, Chengjie Wang 0001, Feiyue Huang
CVPR1
2017 Object detection via feature fusion based single network
abstract
This paper presents a novel network, coined single unified fully convolutional network (SingleNet), for object detection. The proposed method mainly combines two ideas: (1) Our approach aggregates hierarchical features and map them into a uniform space. So, it can further enhance the feature representation ability to degrade the recognition error; (2) To approximate the ground-truth box, we design a set of dense boxes over different aspect ratios and scales per feature map pixel to regress bounding box. It's easy to improve the location performance using dense boxes scheme. Additionally, SingleNet is convenient to train and can be integrated into detection system. Experimental results (mAP: 0.776) on VOC 2007 test demonstrate the advantages of the proposed method over state-of-the-art methods.
Jian Li 0062, Jianjun Qian, Jian Yang 0003
ICIP1