EDBT 2026 Demo / reviewers in the wild / expert
Jianmin Li 0001
dblp:71/5930-1
· DBLP profile ↗
56ranked-venue papers
2as first author
22since 2021 · last 2025
0000-0002-4937-2433ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 2 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 11 since 2021Databases, data management, data science and information retrieval · 5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Physical Adversarial Examples for Person Detectors in Thermal Images Based on 3D ModelingabstractThermal Infrared detection is widely used in autonomous driving, medical AI, etc., but its security has only attracted attention recently. We propose infrared adversarial clothing designed to evade thermal person detectors in real-world scenarios. The design of the adversarial clothing is based on 3D modeling, which makes it easier to simulate multiangle scenes near the real world compared to 2D modeling. We optimized the patch layout pattern of 3D clothing based on the adversarial example technique and made physical adversarial clothing using the aerogel. The idea is to paste a set of square aerogel patches, which display black squares in thermal images, in the inner side of clothing at specific locations with specific orientations. To enhance realism, we propose a method to build infrared 3D models with real infrared photos and develop texture maps for 3D models to simulate varied infrared characteristics over time and location. In physical attacks, we achieved an attack success rate of 80.11% indoors and 76.85% outdoors against YOLOv9. In contrast, randomly placed patches yielded much lower success rates (26.53% indoors and 23.03% outdoors). The adversarial clothing also showed good transferability to unknown detectors with an ensemble attack method, demonstrating the effectiveness of our approach. Xiaopei Zhu, Zhanhao Hu, Jianmin Li 0001, Jun Zhu 0001, Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | CEDNet: A cascade encoder-decoder network for dense prediction
Chufeng Tang, Jianmin Li 0001, Xiaolin Hu 0001 |
Pattern Recognit. | 4 |
| 2024 | SAFDNet: A Simple and Effective Network for Fully Sparse 3D Object DetectionabstractLiDAR-based 3D object detection plays an essential role in autonomous driving. Existing high-performing 3D object detectors usually build dense feature maps in the backbone network and prediction head. However, the computational costs introduced by the dense feature maps grow quadratically as the perception range increases, making these models hard to scale up to long-range detection. Some recent works have attempted to construct fully sparse detectors to solve this issue; nevertheless, the resulting models either rely on a complex multi-stage pipeline or exhibit inferior performance. In this work, we propose a fully sparse adaptive feature diffusion network (SAFDNet) for LiDAR-based 3D object detection. In SAFDNet, an adaptive feature diffusion strategy is designed to address the center feature missing problem. We conducted extensive experiments on Waymo Open, nuScenes, and Argoverse2 datasets. SAFDNet performed slightly better than the previous SOTA on the first two datasets but much better on the last dataset, which features long-range detection, verifying the efficacy of SAFDNet in scenarios where long-range detection is required. Notably, on Argoverse2, SAFDNet surpassed the previous best hybrid detector HEDNet by 2.6% mAP while being 2.1 × faster, and yielded 2.1% mAP gains over the previous best sparse detector FSDv2 while being 1.3 × faster. The code will be available at https://github.com/zhanggang001/HEDNet. Junnan Chen, Guohuan Gao, Jianmin Li 0001, Si Liu 0001, Xiaolin Hu 0001 |
CVPR | 4 |
| 2024 | Infrared Adversarial Car StickersabstractInfrared physical adversarial examples are of great significance for studying the security of infrared AI systems that are widely used in our lives such as autonomous driving. Previous infrared physical attacks mainly focused on 2D infrared pedestrian detection which may not fully manifest its destructiveness to AI systems. In this work, we propose a physical attack method against infrared detectors based on 3D modeling, which is applied to a real car. The goal is to design a set of infrared adversarial stickers to make cars invisible to infrared detectors at various viewing angles, distances, and scenes. We build a 3D infrared car model with real infrared characteristics and propose an infrared adversarial pattern generation method based on 3D mesh shadow. We propose a 3D control points-based mesh smoothing algorithm and use a set of smoothness loss functions to enhance the smoothness of adversarial meshes and facilitate the sticker implementation. Besides, We designed the aluminum stickers and conducted physical experiments on two real Mercedes-Benz A200L cars. Our adversarial stickers hid the cars from Faster RCNN, an object detector, at various viewing angles, distances, and scenes. The attack success rate (ASR) was 91.49% for real cars. In comparison, the ASRs of random stickers and no sticker were only 6.21% and 0.66%, respectively. In addition, the ASRs of the designed stickers against six unseen object detectors such as YOLOv3 and Deformable DETR were between 73.35%-95.80%, showing good transferability of the attack performance across detectors. Xiaopei Zhu, Yuqiu Liu, Zhanhao Hu, Jianmin Li 0001, Xiaolin Hu 0001 |
CVPR | 4 |
| 2024 | InstructPix2NeRF: Instructed 3D Portrait Editing from a Single ImageabstractWith the success of Neural Radiance Field (NeRF) in 3D-aware portrait editing, a variety of works have achieved promising results regarding both quality and 3D consistency. However, these methods heavily rely on per-prompt optimization when handling natural language as editing instructions. Due to the lack of labeled human face 3D datasets and effective architectures, the area of human-instructed 3D-aware editing for open-world portraits in an end-to-end manner remains under-explored. To solve this problem, we propose an end-to-end diffusion-based framework termed $\textbf{InstructPix2NeRF}$, which enables instructed 3D-aware portrait editing from a single open-world image with human instructions. At its core lies a conditional latent 3D diffusion process that lifts 2D editing to 3D space by learning the correlation between the paired images' difference and the instructions via triplet data. With the help of our proposed token position randomization strategy, we could even achieve multi-semantic editing through one single pass with the portrait identity well-preserved. Besides, we further propose an identity consistency module that directly modulates the extracted identity signals into our diffusion process, which increases the multi-view 3D identity consistency. Extensive experiments verify the effectiveness of our method and show its superiority against strong baselines quantitatively and qualitatively. Source code and pretrained models can be found on our project page: https://mybabyyh.github.io/InstructPix2NeRF. Shilong Liu 0004, Yikai Wang 0001, Kaiwen Zheng 0003, Jinghui Xu, Jianmin Li 0001, Jun Zhu 0001 |
ICLR | 7 |
| 2024 | Correcting Pseudo Labels in Semi Supervised Object Detection with SAMabstractPseudo label method is a simple and effective method in semi supervised object detection(SSOD). However, pseudo labels inevitably contain noise, which significantly affects the training efficiency and accuracy of the SSOD model. Conventional pseudo label filtering schemes only consider the category accuracy of pseudo labels, but cannot balance the position quality. The current methods for optimizing pseudo labels also have drawbacks such as relying on the performance of the trained model. To mitigate this problem, we propose a basic process for correcting pseudo labels in SSOD using segment anything model (SAM), which can be simply applied to various existing SSOD models. Moreover, we optimize a specific method of generating pseudo labels using the mask obtained from SAM. We apply our method to several representative models, experiment results show that our method can improve the quality of the obtained pseudo labels and model performance. Jianmin Li 0001, Wenbo Ding 0004, Jiachen Zhong, Jianyong Ai |
ICME | 2 |
| 2024 | Full-Distance Evasion of Pedestrian Detectors in the Physical WorldabstractMany studies have proposed attack methods to generate adversarial patterns for evading pedestrian detection, alarming the computer vision community about the need for more attention to the robustness of detectors. However, adversarial patterns optimized by these methods commonly have limited performance at medium to long distances in the physical world. To overcome this limitation, we identify two main challenges. First, in existing methods, there is commonly an appearance gap between simulated distant adversarial patterns and their physical world counterparts, leading to incorrect optimization. Second, there exists a conflict between adversarial losses at different distances, which causes difficulties in optimization. To overcome these challenges, we introduce a Full Distance Attack (FDA) method. Our physical world experiments demonstrate the effectiveness of our FDA patterns across various detection models like YOLOv5, Deformable-DETR, and Mask RCNN. Codes available at https://github.com/zhicheng2T0/Full-Distance-Attack.git Zhi Cheng, Zhanhao Hu, Yuqiu Liu, Jianmin Li 0001, Hang Su 0006, Xiaolin Hu 0001 |
NeurIPS | 4 |
| 2024 | Hiding from thermal imaging pedestrian detectors in the physical world
Xiaopei Zhu, Xiao Li 0028, Jianmin Li 0001, Zheyao Wang, Xiaolin Hu 0001 |
Neurocomputing | 3 |
| 2023 | PREIM3D: 3D Consistent Precise Image Attribute Editing from a Single ImageabstractWe study the 3D-aware image attribute editing problem in this paper, which has wide applications in practice. Recent methods solved the problem by training a shared encoder to map images into a 3D generator's latent space or by per-image latent code optimization and then edited images in the latent space. Despite their promising results near the input view, they still suffer from the 3D inconsistency of produced images at large camera poses and imprecise image attribute editing, like affecting unspecified attributes during editing. For more efficient image inversion, we train a shared encoder for all images. To alleviate 3D inconsistency at large camera poses, we propose two novel methods, an alternating training scheme and a multi-view identity loss, to maintain 3D consistency and subject identity. As for imprecise image editing, we attribute the problem to the gap between the latent space of real images and that of generated images. We compare the latent space and inversion manifold of GAN models and demonstrate that editing in the inversion manifold can achieve better results in both quantitative and qualitative evaluations. Extensive experiments show that our method produces more 3D consistent images and achieves more precise image editing than previous work. Source code and pretrained models can be found on our project page: https://mybabyyh.github.io/Preim3D/. Jianmin Li 0001, Haoji Zhang 0001, Shilong Liu 0004, Zihao Xiao 0002, Kaiwen Zheng 0003, Jun Zhu 0001 |
CVPR | 2 |
| 2023 | HEDNet: A Hierarchical Encoder-Decoder Network for 3D Object Detection in Point Cloudsabstract3D object detection in point clouds is important for autonomous driving systems. A primary challenge in 3D object detection stems from the sparse distribution of points within the 3D scene. Existing high-performance methods typically employ 3D sparse convolutional neural networks with small kernels to extract features. To reduce computational costs, these methods resort to submanifold sparse convolutions, which prevent the information exchange among spatially disconnected features. Some recent approaches have attempted to address this problem by introducing large-kernel convolutions or self-attention mechanisms, but they either achieve limited accuracy improvements or incur excessive computational costs. We propose HEDNet, a hierarchical encoder-decoder network for 3D object detection, which leverages encoder-decoder blocks to capture long-range dependencies among features in the spatial space, particularly for large and distant objects. We conducted extensive experiments on the Waymo Open and nuScenes datasets. HEDNet achieved superior detection accuracy on both datasets than previous state-of-the-art methods with competitive efficiency. The code is available at https://github.com/zhanggang001/HEDNet. Junnan Chen, Guohuan Gao, Jianmin Li 0001, Xiaolin Hu 0001 |
NeurIPS | 4 |
| 2023 | Hiding from infrared detectors in real world with adversarial clothes
Xiaopei Zhu, Zhanhao Hu, Jianmin Li 0001, Xiaolin Hu 0001, Zheyao Wang |
Appl. Intell. | 4 |
| 2022 | AutoLoss-GMS: Searching Generalized Margin-based Softmax Loss Function for Person Re-identificationabstractPerson re-identification is a hot topic in computer vision, and the loss function plays a vital role in improving the discrimination of the learned features. However, most existing models utilize the hand-crafted loss functions, which are usually sub-optimal and challenging to be designed. In this paper, we propose a novel method, AutoLoss-GMS, to search the better loss function in the space of generalized margin-based softmax loss function for person reidentification automatically. Specifically, the generalized margin-based softmax loss function is first decomposed into two computational graphs and a constant. Then a general searching framework built upon the evolutionary algorithm is proposed to search for the loss function efficiently. The computational graph is constructed with a forward method, which can construct much richer loss function forms than the backward method used in existing works. In addition to the basic in-graph mutation operations, the cross-graph mutation operation is designed to further improve the offspring's diversity. The loss-rejection protocol, equivalence-check strategy and the predictor-based promising-loss chooser are developed to improve the search efficiency. Finally, experimental results demonstrate that the searched loss functions can achieve state-of-the-art performance and be transferable across different models and datasets in person re-identification. Hongyang Gu, Jianmin Li 0001, Guangyuan Fu, Chifong Wong, Xinghao Chen 0001, Jun Zhu 0001 |
CVPR | 2 |
| 2022 | Infrared Invisible Clothing: Hiding from Infrared Detectors at Multiple Angles in Real WorldabstractThermal infrared imaging is widely used in body temperature measurement, security monitoring, and so on, but its safety research attracted attention only in recent years. We proposed the infrared adversarial clothing, which could fool infrared pedestrian detectors at different angles. We simulated the process from cloth to clothing in the digital world and then designed the adversarial “QR code” pattern. The core of our method is to design a basic pattern that can be expanded periodically, and make the pattern after random cropping and deformation still have an adversarial effect, then we can process the flat cloth with an adversarial pattern into any 3D clothes. The results showed that the optimized “QR code” pattern lowered the Average Precision (AP) of YOLOv3 by 87.7%, while the random “QR code” pattern and blank pattern lowered the AP of YOLOv3 by 57.9% and 30.1%, respectively, in the digital world. We then manufactured an adversarial shirt with a new material: aerogel. Physical-world experiments showed that the adversarial “QR code” pattern clothing lowered the AP of YOLOv3 by 64.6%, while the random “QR code” pattern clothing and fully heat-insulated clothing lowered the AP of YOLOv3 by 28.3% and 22.8%, respectively. We used the model ensemble technique to improve the attack transferability to unseen models. Xiaopei Zhu, Zhanhao Hu, Jianmin Li 0001, Xiaolin Hu 0001 |
CVPR | 4 |
| 2022 | The MSR-Video to Text dataset with clean annotations
Haoran Chen 0011, Jianmin Li 0001, Simone Frintrop, Xiaolin Hu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2022 | Improving Image Segmentation with Boundary Patch Refinement
Xiaolin Hu 0001, Chufeng Tang, Hang Chen 0004, Xiao Li 0028, Jianmin Li 0001, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 5 |
| 2022 | Loss function search for person re-identification
Hongyang Gu, Jianmin Li 0001, Guangyuan Fu, Min Yue, Jun Zhu 0001 |
Pattern Recognit. | 2 |
| 2021 | Fooling Thermal Infrared Pedestrian Detectors in Real World Using Small BulbsabstractThermal infrared detection systems play an important role in many areas such as night security, autonomous driving, and body temperature detection. They have the unique advantages of passive imaging, temperature sensitivity and penetration. But the security of these systems themselves has not been fully explored, which poses risks in applying these systems. We propose a physical attack method with small bulbs on a board against the state of-the-art pedestrian detectors. Our goal is to make infrared pedestrian detectors unable to detect real-world pedestrians. Towards this goal, we first showed that it is possible to use two kinds of patches to attack the infrared pedestrian detector based on YOLOv3. The average precision (AP) dropped by 64.12% in the digital world, while a blank board with the same size caused the AP to drop by 29.69% only. After that, we designed and manufactured a physical board and successfully attacked YOLOv3 in the real world. In recorded videos, the physical board caused AP of the target detector to drop by 34.48%, while a blank board with the same size caused the AP to drop by 14.91% only. With the ensemble attack techniques, the designed physical board had good transferability to unseen detectors. Xiaopei Zhu, Xiao Li 0028, Jianmin Li 0001, Zheyao Wang, Xiaolin Hu 0001 |
AAAI | 3 |
| 2021 | Look Closer To Segment Better: Boundary Patch Refinement for Instance SegmentationabstractTremendous efforts have been made on instance segmentation but the mask quality is still not satisfactory. The boundaries of predicted instance masks are usually imprecise due to the low spatial resolution of feature maps and the imbalance problem caused by the extremely low proportion of boundary pixels. To address these issues, we propose a conceptually simple yet effective post-processing refinement framework to improve the boundary quality based on the results of any instance segmentation model, termed BPR. Following the idea of looking closer to segment boundaries better, we extract and refine a series of small boundary patches along the predicted instance boundaries. The refinement is accomplished by a boundary patch refinement network at higher resolution. The proposed BPR framework yields significant improvements over the Mask R-CNN baseline on Cityscapes benchmark, especially on the boundary-aware metrics. Moreover, by applying the BPR framework to the "PolyTransform + SegFix" baseline, we reached 1stplace on the Cityscapes leaderboard. Code is available at https://github.com/tinyalpha/BPR. Chufeng Tang, Hang Chen 0004, Xiao Li 0028, Jianmin Li 0001, Zhaoxiang Zhang 0001, Xiaolin Hu 0001 |
CVPR | 4 |
| 2021 | RefineMask: Towards High-Quality Instance Segmentation With Fine-Grained FeaturesabstractThe two-stage methods for instance segmentation, e.g. Mask R-CNN, have achieved excellent performance recently. However, the segmented masks are still very coarse due to the downsampling operations in both the feature pyramid and the instance-wise pooling process, especially for large objects. In this work, we propose a new method called RefineMask for high-quality instance segmentation of objects and scenes, which incorporates fine-grained features during the instance-wise segmenting process in a multi-stage manner. Through fusing more detailed information stage by stage, RefineMask is able to refine high-quality masks consistently. RefineMask succeeds in segmenting hard cases such as bent parts of objects that are oversmoothed by most previous methods and outputs accurate boundaries. Without bells and whistles, RefineMask yields significant gains of 2.6, 3.4, 3.8 AP over Mask R-CNN on COCO, LVIS, and Cityscapes benchmarks respectively at a small amount of additional computational cost. Furthermore, our single-model result outperforms the winner of the LVIS Challenge 2020 by 1.3 points on the LVIS test-dev set and establishes a new state-of-the-art. Code will be available at https://github.com/zhanggang001/RefineMask. Xin Lu 0002, Jingru Tan, Jianmin Li 0001, Zhaoxiang Zhang 0001, Quanquan Li, Xiaolin Hu 0001 |
CVPR | 4 |
| 2021 | Attack on Practical Speaker Verification System Using Universal Adversarial PerturbationsabstractIn authentication scenarios, applications of practical speaker verification systems usually require a person to read a dynamic authentication text. Previous studies played an audio adversarial example as a digital signal to perform physical attacks, which would be easily rejected by audio replay detection modules. This work shows that by playing our crafted adversarial perturbation as a separate source when the adversary is speaking, the practical speaker verification system will misjudge the adversary as a target speaker. A two-step algorithm is proposed to optimize the universal adversarial perturbation to be text-independent and has little effect on the authentication text recognition. We also estimated room impulse response (RIR) in the algorithm which allowed the perturbation to be effective after being played over the air. In the physical experiment, we achieved targeted attacks with success rate of 100%, while the word error rate (WER) on speech recognition was only increased by 3.55%. And recorded audios could pass replay detection for the live person speaking. Shuning Zhao, Jianmin Li 0001, Xingliang Cheng, Thomas Fang Zheng, Xiaolin Hu 0001 |
ICASSP | 4 |
| 2021 | Auto-ReID+: Searching for a multi-branch ConvNet for person re-identification
Hongyang Gu, Guangyuan Fu, Jianmin Li 0001, Jun Zhu 0001 |
Neurocomputing | 3 |
| 2021 | Vocabulary-Wide Credit Assignment for Training Image Captioning ModelsabstractReinforcement learning (RL) algorithms have been shown to be efficient in training image captioning models. A critical step in RL algorithms is to assign credits to appropriate actions. There are mainly two classes of credit assignment methods in existing RL methods for image captioning, assigning a single credit for the whole sentence and assigning a credit to every word in the sentence. In this article, we propose a new credit assignment method which is orthogonal to the above two. It assigns every word in vocabulary an appropriate credit at each generation step. It is called vocabulary-wide credit assignment. Based on this we propose a Vocabulary-Critical Sequence Training (VCST). VCST can be incorporated into existing RL methods for training image captioning models to achieve better results. Extensive experiments with many popular models validated the effectiveness of VCST. Jianmin Li 0001, Xiaolin Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Delving Deeper into the Decoder for Video CaptioningabstractVideo captioning is an advanced multi-modal task which aims to describe a video clip using a natural language sentence. The encoder-decoder framework is the most popular paradigm for this task in recent years. However, there exist some problems in the decoder of a video captioning model. We make a thorough investigation into the decoder and adopt three techniques to improve the performance of the model. First of all, a combination of variational dropout and layer normalization is embedded into a recurrent unit to alleviate the problem of overfitting. Secondly, a new online method is proposed to evaluate the performance of a model on a validation set so as to select the best checkpoint for testing. Finally, a new training strategy called professional learning is proposed which uses the strengths of a captioning model and bypasses its weaknesses. It is demonstrated in the experiments on Microsoft Research Video Description Corpus (MSVD) and MSR-Video to Text (MSR-VTT) datasets that our model has achieved the best results evaluated by BLEU, CIDEr, METEOR and ROUGE-L metrics with significant gains of up to 18% on MSVD and 3.5% on MSR-VTT compared with the previous state-of-the-art models. Haoran Chen 0011, Jianmin Li 0001, Xiaolin Hu 0001 |
ECAI | 2 |
| 2019 | Joint Cluster Unary Loss for Efficient Cross-Modal HashingabstractRecently, cross-modal deep hashing has received broad attention for solving cross-modal retrieval problems efficiently. Most cross-modal hashing methods generate $O(n^2)$ data pairs and $O(n^3)$ data triplets for training, but the training procedure is less efficient because the complexity is high for large-scale dataset. In this paper, we propose a novel and efficient cross-modal hashing algorithm named Joint Cluster Cross-Modal Hashing (JCCH). First, We introduce the Cross-Modal Unary Loss (CMUL) with $O(n)$ complexity to bridge the traditional triplet loss and classification-based unary loss, and the JCCH algorithm is introduced with CMUL. Second, a more accurate bound of the triplet loss for structured multilabel data is introduced in CMUL. The resultant hashcodes form several clusters in which the hashcodes in the same cluster share similar semantic information, and the heterogeneity gap on different modalities is diminished by sharing the clusters. Experiments on large-scale datasets show that the proposed method is superior over or comparable with state-of-the-art cross-modal hashing methods, and training with the proposed method is more efficient than others. Jianmin Li 0001, Bo Zhang 0010 |
ICMR | 2 |
| 2019 | Semantic Cluster Unary Loss for Efficient Deep HashingabstractHashing method maps similar data to binary hashcodes with smaller hamming distance, which has received broad attention due to its low storage cost and fast retrieval speed. With the rapid development of deep learning, deep hashing methods have achieved promising results in efficient information retrieval. Most existing deep hashing methods adopt pairwise or triplet losses to deal with similarities underlying the data, but their training are difficult and less efficient because O(n2) data pairs and O(n3) triplets are involved. To address these issues, we propose a novel deep hashing algorithm with unary loss which can be trained very efficiently. First of all, we introduce a Unary Upper Bound of the traditional triplet loss, thus reducing the complexity to O(n) and bridging the classificationbased unary loss and the triplet loss. Second, we propose a novel Semantic Cluster Deep Hashing (SCDH) algorithm by introducing a modified Unary Upper Bound loss, named Semantic Cluster Unary Loss (SCUL). The resultant hashcodes form several compact clusters, which means hashcodes in the same cluster have similar semantic information. We also demonstrate that the proposed SCDH is easy to be extended to semi-supervised settings by incorporating the state-of-the-art semi-supervised learning algorithms. Experiments on large-scale datasets show that the proposed method is superior to state-of-the-art hashing algorithms. Jianmin Li 0001, Bo Zhang 0010 |
IEEE Trans. Image Process. | 2 |
| 2018 | Visual instance mining from the graph perspective
Wei Li 0152, Jianmin Li 0001, Changhu Wang, Lei Zhang 0001, Bo Zhang 0010 |
Multim. Syst. | 2 |
| 2018 | Scalable Discrete Supervised Multimedia Hash Learning With ClusteringabstractThe hashing method maps similar data of various types to binary hashcodes with smaller hamming distance, and it has received broad attention due to its low-storage cost and fast retrieval speed. However, the existing limitations make the present algorithms difficult to deal with for large-scale data sets: 1) discrete constraints are involved in the learning of the hash function and 2) pairwise or triplet similarity is adopted to generate efficient hashcodes, resulting in both time and space complexity greater than O(n2). To address these issues, we propose a novel discrete supervised hash learning framework that can be scalable to large-scale data sets of various types. First, the discrete learning procedure is decomposed into a binary classifier learning scheme and binary codes learning scheme, which makes the learning procedure more efficient. Second, by adopting the asymmetric low-rank matrix factorization, we propose the fast clustering-based batch coordinate descent method, such that the time and space complexity are reduced to O(n). The proposed framework also provides a flexible paradigm to incorporate with arbitrary hash function, including deep neural networks and kernel methods, as well as any types of data to hash, including images and videos. Experiments on large-scale data sets demonstrate that the proposed method is superior or comparable with the state-of-the-art hashing algorithms. Jianmin Li 0001, Mengqing Jiang, Peijiang Yuan, Bo Zhang 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Accelerating convolutional neural networks by group-wise 2D-filter pruningabstractNetwork pruning is an effective way to accelerate Convolutional Neural Networks (CNNs). In recent years, structured pruning methods are proposed in favor of unstructured methods as they have shown greater speedup in practical use. Existing structured methods does pruning along two main dimensions: 3D-filter wise, i.e., remove a 3D-fllter as a whole, and filter-shape wise, i.e., remove a same position from all 3D-filters. In this work, we propose a new group-wise 2D-fllter pruning approach that is orthogonal and complementary to the existing methods. The proposed approach removes a portion of 2D-fllters from each 3D-filter according to the pruning patterns learned from the data, and leads to compressed models that do not require sophisticated implementation of convolution operations. A fine-tuning process is followed to recover the accuracy. The knowledge distillation (KD) framework is explored in the fine-tuning process to improve the performance. We present our method for learning the pruning pattens as well as the fine-tuning strategy based on knowledge distillation. The proposed approach is validated on two representative CNN models - ZF and VGG16, pre-trained on ILSVRC12. Experimental results demonstrate the effectiveness of our approach. In VGG16, we get even higher accuracy after speeding-up the network by 4 times. Niange Yu, Xiaolin Hu 0001, Jianmin Li 0001 |
IJCNN | 4 |
| 2016 | Scalable Discrete Supervised Hash Learning with Asymmetric Matrix FactorizationabstractHashing methods map similar data to binary hashcodes with smaller hamming distance, and it has received a broad attention due to its low storage cost and fast retrieval speed. However, the existing limitations make the present algorithms difficult to deal with large-scale datasets: (1) discrete constraints are involved in the learning of the hash function, (2) pairwise or triplet similarity is adopted to generate efficient hashcodes, resulting both time and space complexity are greater than O(n2). To address these issues, we propose a novel discrete supervised hash learning framework which can be scalable to large-scale datasets. First, the learning procedure is decomposed into a binary classifier learning scheme and hashcodes learning scheme. Second, we adopt the Asymmetric Low-rank Matrix Factorization and propose the Fast Clustering-based Batch Coordinate Descent method, such that the time and space complexity is reduced to O(n). The proposed framework also provides a flexible paradigm to incorporate with arbitrary hash function, including deep neural networks. Experiments on large-scale datasets demonstrate that the proposed method is superior or comparable with state-of-the-art hashing algorithms. Jianmin Li 0001, Jinma Guo, Bo Zhang 0010 |
ICDM | 2 |
| 2016 | Hash Learning with Convolutional Neural Networks for Semantic Based Image Retrieval
Jinma Guo, Jianmin Li 0001 |
PAKDD (1) | 3 |
| 2016 | BitHash: An efficient bitwise Locality Sensitive Hashing method with applications
Wenhao Zhang 0003, Jianqiu Ji, Jun Zhu 0001, Jianmin Li 0001, Hua Xu 0003, Bo Zhang 0010 |
Knowl. Based Syst. | 4 |
| 2015 | Regularizing neural networks with adaptive local dropabstractNeural network (NN) models have shown good performance on many image recognition benchmarks. Given large image datasets, these models typically have millions or billions of parameters that can easily lead to over-fitting without regularization. Dropout and DropConnect show their effectiveness of regularizing large fully connected layers within neural networks. In Dropout, each neural activation within the network is randomly set to zero with a probability during training. In DropConnect, a generalization of Dropout, each connection weight within the network is randomly set to zero with a probability instead. Both of the probabilities in Dropout and DropConnect are universal predefined constants. We propose Adaptive Local Drop (ALDrop), a novel regularization method that sets each connection weight within the network with a learned probability adaptive to the input image dataset using a locality-based measure. Experiments on several image recognition benchmarks show that our model outperforms Dropout and DropConnect. Binbin Cao, Jianmin Li 0001, Bo Zhang 0010 |
IJCNN | 2 |
| 2015 | Angular-Similarity-Preserving Binary Signatures for Linear SubspacesabstractWe propose a similarity-preserving binary signature method for linear subspaces. In computer vision and pattern recognition, linear subspace is a very important representation for many kinds of data, such as face images, action and gesture videos, and so on. When there is a large amount of subspace data and the ambient dimension is high, the cost of computing the pairwise similarity between the subspaces would be high and it requires a large storage space for storing the subspaces. In this paper, we first define the angular similarity and angular distance between the subspaces. Then, based on this similarity definition, we develop a similarity-preserving binary signature method for linear subspaces, which transforms a linear subspace into a compact binary signature, and the Hamming distance between two signatures provides an unbiased estimate of the angular similarity between the two subspaces. We also provide a lower bound of the signature length sufficient to guarantee uniform distance-preservation between every pair of subspaces in a set. Experiments on face recognition, gesture recognition, and action recognition verify the effectiveness of the proposed method. Jianqiu Ji, Jianmin Li 0001, Qi Tian 0001, Shuicheng Yan, Bo Zhang 0010 |
IEEE Trans. Image Process. | 2 |
| 2014 | Similarity-Preserving Binary Signature for Linear SubspacesabstractLinear subspace is an important representation for many kinds of real-world data in computer vision and pattern recognition, e.g. faces, motion videos, speeches. In this paper, first we define pairwise angular similarity and angular distance for linear subspaces. The angular distance satisfies non-negativity, identity of indiscernibles, symmetry and triangle inequality, and thus it is a metric. Then we propose a method to compress linear subspaces into compact similarity-preserving binary signatures, between which the normalized Hamming distance is an unbiased estimator of the angular distance. We provide a lower bound on the length of the binary signatures which suffices to guarantee uniform distance-preservation within a set of subspaces. Experiments on face recognition demonstrate the effectiveness of the binary signature in terms of recognition accuracy, speed and storage requirement. The results show that, compared with the exact method, the approximation with the binary signatures achieves an order of magnitude speed-up, while requiring significantly smaller amount of storage space, yet it still accurately preserves the similarity, and achieves high recognition accuracy comparable to the exact method in face recognition. Jianqiu Ji, Jianmin Li 0001, Shuicheng Yan, Qi Tian 0001, Bo Zhang 0010 |
AAAI | 2 |
| 2014 | Batch-Orthogonal Locality-Sensitive Hashingfor Angular SimilarityabstractSign-random-projection locality-sensitive hashing (SRP-LSH) is a widely used hashing method, which provides an unbiased estimate of pairwise angular similarity, yet may suffer from its large estimation variance. We propose in this work batch-orthogonal locality-sensitive hashing (BOLSH), as a significant improvement of SRP-LSH. Instead of independent random projections, BOLSH makes use of batch-orthogonalized random projections, i.e, we divide random projection vectors into several batches and orthogonalize the vectors in each batch respectively. These batch-orthogonalized random projections partition the data space into regular regions, and thus provide a more accurate estimator. We prove theoretically that BOLSH still provides an unbiased estimate of pairwise angular similarity, with a smaller variance for any angle in (0, π), compared with SRP-LSH. Furthermore, we give a lower bound on the reduction of variance. The extensive experiments on real data well validate that with the same length of binary code, BOLSH may achieve significant mean squared error reduction in estimating pairwise angular similarity. Moreover, BOLSH shows the superiority in extensive approximate nearest neighbor (ANN) retrieval experiments. Jianqiu Ji, Shuicheng Yan, Jianmin Li 0001, Guangyu Gao, Qi Tian 0001, Bo Zhang 0010 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Min-Max Hash for Jaccard SimilarityabstractMin-wise hash is a widely-used hashing method for scalable similarity search in terms of Jaccard similarity, while in practice it is necessary to compute many such hash functions for certain precision, leading to expensive computational cost. In this paper, we introduce an effective method, i.e. the min-max hash method, which significantly reduces the hashing time by half, yet it has a provably slightly smaller variance in estimating pair wise Jaccard similarity. In addition, the estimator of min-max hash only contains pair wise equality checking, thus it is especially suitable for approximate nearest neighbor search. Since min-max hash is equally simple as min-wise hash, many extensions based on min-wise hash can be easily adapted to min-max hash, and we show how to combine it with b-bit minwise hash. Experiments show that with the same length of hash code, min-max hash reduces the hashing time to half as much as that of min-wise hash, while achieving smaller mean squared error (MSE) in estimating pair wise Jaccard similarity, and better best approximate ratio (BAR) in approximate nearest neighbor search. Jianqiu Ji, Jianmin Li 0001, Shuicheng Yan, Qi Tian 0001, Bo Zhang 0010 |
ICDM | 2 |
| 2013 | Restricted Boltzmann Machine with Adaptive Local Hidden Units
Binbin Cao, Jianmin Li 0001, Jun Wu 0022, Bo Zhang 0010 |
ICONIP (2) | 2 |
| 2013 | Traffic sign detection by ROI extraction and histogram features-based recognitionabstractWe present a traffic sign detection model consisting of two modules. The first module is for ROI (region of interest) extraction. By supervised learning, it transforms the color images to gray images such that the characteristic colors for the traffic signs are more distinguishable in the gray images. It follows shape template matching, where a set of templates for each target category of signs are designed. After that, a set of ROIs are generated. The second module is for recognition. It validates if an ROI belongs to a target category of traffic signs by supervised learning. Local shape and color features are extracted. The supervised learning methods used in the model are SVMs. The overall model is applied on the GTSDB benchmark and achieves 100%, 98.85% and 92.00% AUC (area under the precision-recall curve) for Prohibitory, Danger and Mandatory signs, respectively. The testing speed is 0.4-1.0 second per image on a mainstream PC, which demonstrates the great potential of the proposed model in real-time applications. Mingyi Yuan, Xiaolin Hu 0001, Jianmin Li 0001, Huaping Liu 0001 |
IJCNN | 4 |
| 2013 | Traffic sign detection based on convolutional neural networksabstractWe propose an approach for traffic sign detection based on Convolutional Neural Networks (CNN). We first transform the original image into the gray scale image by using support vector machines, then use convolutional neural networks with fixed and learnable layers for detection and recognition. The fixed layer can reduce the amount of interest areas to detect, and crop the boundaries very close to the borders of traffic signs. The learnable layers can increase the accuracy of detection significantly. Besides, we use bootstrap methods to improve the accuracy and avoid overfitting problem. In the German Traffic Sign Detection Benchmark, we obtained competitive results, with an area under the precision-recall curve(AUC) of 99.73% in the category “Danger”, and an AUC of 97.62% in the category “Mandatory”. Yihui Wu, Jianmin Li 0001, Huaping Liu 0001, Xiaolin Hu 0001 |
IJCNN | 3 |
| 2012 | Super-Bit Locality-Sensitive HashingabstractSign-random-projection locality-sensitive hashing (SRP-LSH) is a probabilistic dimension reduction method which provides an unbiased estimate of angular similarity, yet suffers from the large variance of its estimation. In this work, we propose the Super-Bit locality-sensitive hashing (SBLSH). It is easy to implement, which orthogonalizes the random projection vectors in batches, and it is theoretically guaranteed that SBLSH also provides an unbiased estimate of angular similarity, yet with a smaller variance when the angle to estimate is within $(0,\pi/2]$. The extensive experiments on real data well validate that given the same length of binary code, SBLSH may achieve significant mean squared error reduction in estimating pairwise angular similarity. Moreover, SBLSH shows the superiority over SRP-LSH in approximate nearest neighbor (ANN) retrieval experiments. Jianqiu Ji, Jianmin Li 0001, Shuicheng Yan, Bo Zhang 0010, Qi Tian 0001 |
NIPS | 2 |
| 2010 | Learning Vocabulary-Based Hashing with AdaBoost
Yingyu Liang, Jianmin Li 0001, Bo Zhang 0010 |
MMM | 2 |
| 2009 | Vocabulary-based hashing for image searchabstractThis paper proposes a hash function family based on feature vocabularies and investigates the application in building indexes for image search. Each hash function is associated with a set of feature points, i.e. a vocabulary, and maps an input point to the ID of the nearest one in the vocabulary. The function family can be employed to build a high-dimensional index for approximate nearest neighbor search. Then we concentrate on its application in image search. Guiding rules for the construction of the vocabularies are derived, which improve the effectiveness of the approach in this context by taking advantage of the data distribution. The rules are applied to design an algorithm for vocabulary construction in practice. Experiments show promising performance of the approach and the effectiveness of the guiding rules. Comparison with the popular Euclidean locality-sensitive hashing also shows the advantage of our approach in image search. Yingyu Liang, Jianmin Li 0001, Bo Zhang 0010 |
ACM Multimedia | 2 |
| 2009 | Query representation by structured concept threads with application to interactive video retrieval
Dong Wang 0022, Jianmin Li 0001, Bo Zhang 0010, Xirong Li 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2008 | Scene understanding with discriminative structured predictionabstractSpatial priors play crucial roles in many high-level vision tasks, e.g. scene understanding. Usually, learning spatial priors relies on training a structured output model. In this paper, two special cases of discriminative structured output model, i.e. conditional random fields (CRFs) and max-margin Markov networks (M3N), are demonstrated to perform image scene understanding. The two models are empirically compared in a fair manner, i.e. using the common feature representation and the same optimization algorithm. Particularly, we adopt online exponentiated gradient (EG) algorithm to solve the convex duals of both models. We describe the general procedure of EG algorithm and present a two-stage training procedure to overcome the degeneration of EG when exact inference is intractable. Experiments on a large scale image region annotation task are carried out. The results show that both models yield encouraging results but CRFs slightly outperforms M3N. Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010 |
CVPR | 2 |
| 2007 | Retrieving Web Images to Enrich Music RepresentationabstractAudiovisual media which integrates visual media with audio to enrich music representation, such as music video (MV) or music slideshow, is now more welcome than only audio. In this paper, we proposed a novel approach to automatically retrieve web images appropriately suitable to a given music song. In this approach, an imageability measurement is first proposed to select those meaningful words and phrases from lyrics, as queries for further image search. Then, considering the possible ambiguities of queries may cause web image search engines to return images with various semantic concepts, we also introduced a search result clustering (SRC)based strategy to select those images which are more likely to be relevant to the content of music, using a naive Bayesian inference. Preliminary evaluations of the proposed approach on around 100 popular English music songs have shown promising results. Rui Cai 0002, Lei Zhang 0001, Jianmin Li 0001 |
ICME | 5 |
| 2007 | The importance of query-concept-mapping for automatic video retrievalabstractA new video retrieval paradigm of query-by-concept emerges recently. However, it remains unclear how to exploit the detected concepts in retrieval given a multimedia query. In this paper, we point out that it is important to map the query to a few relevant concepts instead of search with all concepts. In addition, we show that solving this problem through both text and image inputs are effective for search, and it is possible to determine the number of related concepts by a language modeling approach. Experimental evidence is obtained on the automatic search task of TRECVID 2006 using a large lexicon of 311 learned semantic concept detectors. Dong Wang 0022, Xirong Li 0001, Jianmin Li 0001, Bo Zhang 0010 |
ACM Multimedia | 3 |
| 2007 | Gradual transition detection with conditional random fieldsabstractIn this paper, we view gradual transition detection as a sequence labeling problem and propose to use Conditional Random Fields (CRFs) for this purpose. CRFs is a state-of-the-art sequence labeling approach. It provides a unified way to integrate various useful clues to form a decision system. Moreover, it has principled way for parameter estimation and inference. Compared to rule-based approaches, gradual transition detection with CRFs requires fewer human interactions while designing the system. The experiments on TRECVID platform show that CRFs can achieve comparable performance to that of the state-of-the-art approaches. Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010 |
ACM Multimedia | 2 |
| 2007 | Exploiting spatial context constraints for automatic image region annotationabstractIn this paper we conduct a relatively complete study on how to exploit spatial context constraints for automated image region annotation. We present a straight forward method to regularize the segmented regions into 2D lattice layout, so that simple grid-structure graphical models can be employed to characterize the spatial dependencies. We show how to represent the spatial context constraints in various graphical models and also present the related learning and inference algorithms. Different from most of the existing work, we specifically investigate how to combine the classification performance of discriminative learning and the representation capability of graphical models. To reliably evaluate the proposed approaches, we create a moderate scale image set with region-level ground truth. The experimental results show that (i) spatial context constraints indeed help for accurate region annotation, (ii) the approaches combining the merits of discriminative learning and context constraints perform best, (iii) image retrieval can benefit from accurate region-level annotation. Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010 |
ACM Multimedia | 2 |
| 2007 | A Formal Study of Shot Boundary DetectionabstractThis paper conducts a formal study of the shot boundary detection problem. First, a general formal framework of shot boundary detection techniques is proposed. Three critical techniques, i.e., the representation of visual content, the construction of continuity signal and the classification of continuity values, are identified and formulated in the perspective of pattern recognition. Meanwhile, the major challenges to the framework are identified. Second, a comprehensive review of the existing approaches is conducted. The representative approaches are categorized and compared according to their roles in the formal framework. Based on the comparison of the existing approaches, optimal criteria for each module of the framework are discussed, which will provide practical guide for developing novel methods. Third, with all the above issues considered, we present a unified shot boundary detection system based on graph partition model. Extensive experiments are carried out on the platform of TRECVID. The experiments not only verify the optimal criteria discussed above, but also show that the proposed approach is among the best in the evaluation of TRECVID 2005. Finally, we conclude the paper and present some further discussions on what shot boundary detection can learn from other related fields Jinhui Yuan, Wujie Zheng, Jianmin Li 0001, Fuzong Lin, Bo Zhang 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2006 | Multiple-Instance Learning Via Random Walk
Dong Wang 0022, Jianmin Li 0001, Bo Zhang 0010 |
ECML | 2 |
| 2006 | Learning concepts from large scale imbalanced data sets using support cluster machinesabstractThis paper considers the problem of using Support Vector Machines (SVMs) to learn concepts from large scale imbalanced data sets. The objective of this paper is twofold. Firstly, we investigate the effects of large scale and imbalance on SVMs. We highlight the role of linear non-separability in this problem. Secondly, we develop a both practical and theoretical guaranteed meta-algorithm to handle the trouble of scale and imbalance. The approach is named Support Cluster Machines (SCMs). It incorporates the informative and the representative under-sampling mechanisms to speedup the training procedure. The SCMs differs from the previous similar ideas in two ways, (a) the theoretical foundation has been provided, and (b) the clustering is performed in the feature space rather than in the input space. The theoretical analysis not only provides justification, but also guides the technical choices of the proposed approach. Finally, experiments on both the synthetic and the TRECVID data are carried out. The results support the previous analysis and show that the SCMs are efficient and effective while dealing with large scale imbalanced data sets. Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010 |
ACM Multimedia | 2 |
| 2005 | A unified shot boundary detection framework based on graph partition modelabstractIn this paper, we propose a unified shot boundary detection framework by extending the previous work of graph partition model with temporal constraints. To detect both the abrupt transitions (CUTs) and gradual transitions (GTs, excluding fade out/in) in a unified way, we incorporate temporal multi-resolution analysis into the model. Furthermore, instead of ad-hoc thresholding scheme, we construct a novel kind of feature to characterize shot transitions and employ support vector machine (SVM) with active leaning strategy to classify boundaries and non-boundaries. Extensive experiments have been carried out on the platform of TRECVID benchmark. The experimental results show that the proposed framework outperforms some others and achieves satisfactory results. Jinhui Yuan, Jianmin Li 0001, Fuzong Lin, Bo Zhang 0010 |
ACM Multimedia | 2 |
| 2004 | Improvements to Bennett?s Nearest Point Algorithm for Support Vector Machines
Jianmin Li 0001, Jianwei Zhang 0001, Bo Zhang 0010 |
ISNN (1) | 1 |
| 2003 | Nonlinear Speech Model Based on Support Vector Machine and Wavelet TransformabstractTo improve the naturalness of reconstructed speech, nonlinear speech models are paid more and more attention in recent years. A nonlinear speech model for speech synthesis based on support vector machine (SVM) is presented firstly. After speech signal is embedded into phase space, nonlinear map in the model is obtained with support vector regression. It is shown in the experiments that for some pieces of speech, not only can speech be perfectly reconstructed by the system, but also jitter and shimmer in the original signal is preserved. However, the output of the system is quite different from the original one for other pieces. The reason is that the sub-bands with different frequency in the original signal can not be perfectly described by a SVM-based autoregressive model trained with one set of training parameters. Consequently, a multi-band model is then proposed. After the original speech is decomposed into several bands through wavelet packet decomposition, a nonlinear dynamical model based on SVM is constructed for each sub-band signal. It is shown in the experiments that the stability of such system is improved. Jianmin Li 0001, Bo Zhang 0010, Fuzong Lin |
ICTAI | 1 |
| 2002 | Generation of Chinese prosodic phrasing rules by an extension matrix algorithmabstractThis paper presents a new rule induction algorithm based on the extension matrix theory, and uses it to learn prosodic phrasing rules automatically for Chinese text-to-speech. Firstly, the basic idea of our algorithm is introduced. Secondly, we collected 937 sentences from news programs and built a corpus for modeling Chinese prosody, a group of feature variables are also proposed. Lastly, the data is divided into two parts: training set and test set, and the experimental results show that our method achieves higher comprehensibility, better accuracy and fewer rules than other algorithms. And the generated rules are quite similar to hand-crafted ones. Weijun Chen 0001, Fuzong Lin, Jianmin Li 0001, Bo Zhang 0010 |
ICASSP | 3 |
| 2001 | Training prosodic phrasing rules for Chinese TTS systemsabstractThis paper describes several experiments designed to train prosodic phrasing models for Chinese TTS systems and to investigate the underlying rules that control Chinese prosody. First, we collected 559 sentences from news programs and built a large corpus for modeling Chinese prosody. Second, we selected 20 features and used classification and regression trees (CART) and transformational rule-based learning (TRBL) techniques to generate phrasing rules automatically. Lastly, we propose a computer aided error-driven method of designing rule templates, and integrate it into the TRBL algorithm. The experimental results show that we achieve a high success rate of 94.5%, and we also get a set of well comprehensible rule templates which may give us insights into the relationship between Chinese syntax and p rosody. Weijun Chen 0001, Fuzong Lin, Jianmin Li 0001, Bo Zhang 0010 |
INTERSPEECH | 3 |