Biao Leng

dblp:42/2913 · DBLP profile ↗
← Back
60ranked-venue papers
15as first author
26since 2021 · last 2026
0000-0003-3588-5622ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 31 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Efficient multi-agent communication via entity-aware causal network
Yifan Bo, Jinghan Feng, Shuo Zhang 0003, Biao Leng
Neural Networks5
2026 Scattering center guided mono-static radar cross section prediction
Zehao Tang, Shuo Zhang 0003, Biao Leng
Neural Networks4
2026 DuoNet: Joint optimization of representation learning and prototype classifier for unbiased scene graph generation
Zhaodi Wang 0003, Biao Leng, Shuo Zhang 0003
Pattern Recognit.2
2026 Focus-then-fusion: Learning discriminative cross-modal prototypes for few-shot classification
Rongshan Chen, Shuo Zhang 0003, Biao Leng
Pattern Recognit.4
2026 CDC: Enhancing Scene Graph Generation for IoST-Driven Social Behavioral Modeling With Cooperative Dual Classifier
abstract
Scene graph generation (SGG) plays an important role in the intelligence of social things (IoST) framework by extracting structured semantic representations from social device data, thereby supporting advanced scene understanding and behavioral-cultural modeling. However, the intrinsic long-tail nature of real-world social device data, coupled with the semantic entanglement between head and tail categories (e.g., “on” versus “standing on”), presents significant challenges for fine-grained SGG. This often results in biased models and suboptimal generalization to rare but semantically informative relations. To address these issues, we propose a novel cooperative dual classifier (CDC) framework for fine-grained SGG in IoST-driven social systems. CDC introduces a cooperative learning mechanism that combines two classifiers. The frozen prototype classifier is designed with maximum interclass margins to alleviate class imbalance. In parallel, a learnable classifier dynamically adjusts decision boundaries to improve discriminative precision. To further enhance the integration between the two classifiers, we introduce a weight knowledge transfer (WKT) module and a collaborative constraint term, facilitating robust adaptation to tail categories. Extensive experiments on the Visual Genome and GQA datasets demonstrate that CDC outperforms state-of-the-art SGG methods, particularly in modeling fine-grained relations under long-tail distributions. These results highlight the capability of CDC to advance semantic understanding of complex behavioral and cultural patterns within computational social systems.
Zhaodi Wang 0003, Yangyan Zeng, Biao Leng, Xiaokang Zhou
IEEE Trans. Comput. Soc. Syst.3
2026 SPANext: Subpattern-Aware Two-Stage Graph Learning Framework for Next Location Prediction
Meiyue You, Shuo Zhang 0003, Biao Leng
IEEE Trans. Intell. Transp. Syst.3
2026 CrossHypergraph: Consistent High-Order Semantic Network for Few-Shot Image Classification
abstract
Few-shot classification is a challenging task that recognizes novel classes by learning from few training instances. Metric-based models are currently the most effective solutions for few-shot classification. In these models, patch feature distances between query instances and support classes are calculated to achieve classification. However, it is difficult for patch-based methods to mine semantic information of support and query instances, leading to inaccurate feature similarity measures. To address these problems, we propose to construct CrossHypergraph based on hypergraph modeling. Specifically, we first align the local prototype vertices of support and query instances to model consistent hypergraph structures. Then avertex-hyperedge-vertex-based interactive feature updating mechanism is designed to generate CrossHypergraph representation with consistent high-order semantic information for support and query instances. Based on the CrossHypergraph, we propose a consistent high-order semantic network, in which the high-order semantic-based weighted metric strategy is designed to achieve accurate classification. The proposed method is evaluated on general, fine-grained, and cross-domain few-shot benchmarks, including miniImageNet, tieredImageNet, CIFAR-FS, FC100, and miniImageNet$\rightarrow$CUB datasets. Experimental results show that our CrossHypergraph-based few-shot classifier generates consistent high-order semantic features, and achieves state-of-the-art performance on both 1-shot and 5-shot tasks.
Shuo Zhang 0003, Biao Leng
IEEE Trans. Multim.4
2025 Unlocking the Potential of Reverse Distillation for Anomaly Detection
abstract
Knowledge Distillation (KD) is a promising approach for unsupervised Anomaly Detection (AD). However, the student network's over-generalization often diminishes the crucial representation differences between teacher and student in anomalous regions, leading to detection failures. To address this problem, the widely accepted Reverse Distillation (RD) paradigm designs the asymmetry teacher and student network, using an encoder as teacher and a decoder as student. Yet, the design of RD does not ensure that the teacher encoder effectively distinguishes between normal and abnormal features or that the student decoder generates anomaly-free features. Additionally, the absence of skip connections results in a loss of fine details during feature reconstruction. To address these issues, we propose RD with Expert, which introduces a novel Expert-Teacher-Student network for simultaneous distillation of both the teacher encoder and student decoder. The added expert network enhances the student's ability to generate normal features and optimizes the teacher's differentiation between normal and abnormal features, reducing missed detections. Additionally, Guided Information Injection is designed to filter and transfer features from teacher to student, improving detail reconstruction and minimizing false positives. Experiments on several benchmarks prove that our method outperforms existing unsupervised AD methods under RD paradigm, fully unlocking RD’s potential.
Biao Leng, Shuo Zhang 0003
AAAI3
2025 Uniform Distribution Based Learnable Quantization via Self-Knowledge Distillation
abstract
Quantization is an effective method for compressing DNNs with enormous parameters and improving inference efficiency especially on resource-constrained devices like mobile phones. However, it is difficult to quantize models to extremely low-bit while maintaining considerable accuracy compared to their full-precision counterparts, because the information representation capacity of quantized models declines drastically as bit-widths reducing. To alleviate this problem, we propose an improved learnable quantizer modified from LCQ and adopt a self-knowledge distillation framework to assist quantization. Our method parameterizes both domain and range of full-precision data, promoting the upper bound of quantizer with more quantization parameters. Besides, we utilize knowledge distillation mechanism to adapt the distribution of full-precision data, turning it into more quantization friendly. We demonstrate the effectiveness of our method on CIFAR-10 and ImageNet dataset with various network architectures. Experimental results show that our method bridge the gap between quantized and full-precision models.
Biao Leng
ICASSP2
2025 Wave-Spectrogram Cross-Modal Aggregation for Audio Deepfake Detection
abstract
Realistic deepfake audio has posed significant security risk. To address this threat, current countermeasures attempt to extract forgery traces from either waveform signals or corresponding spectrograms. However, features extracted from single modality are susceptible to non-spoof disturbances and tend to overfit specific forgery method, thus fail to generalize on out-of-domain forgeries. In this paper, we propose a novel cross-modal deepfake audio detection framework, which leverages multi-scale representations to improve generalization in distinguish unseen synthesized utterances. As spectrogram can reveal hidden features in audio signals, our model learns discriminative and intrinsic feature representations by meticulously aligning features of multiple modalities. Specifically, we design a multiscale fusion strategy to aggregate deepfake artifacts of different scales, overcoming the challenge of aligning heterogeneous multi-modal features. Furthermore, we employ single center loss to condense the embeddings of bonafide audio, enhancing the ability of classifier in detecting out-of-domain deepfake audios. As a result, our approach outperforms the state-of-the-art studies on challenging ASVspoof2021 Deepfake dataset and In-The-Wild dataset. Extensive experiments further demonstrate the effectiveness and generalization of proposed detection framework.
Zehui Jin, Linlong Lang, Biao Leng
ICASSP3
2025 Adaptive Contribution Modulation For Multi-Modal Manipulation Media Detection and Grounding
abstract
In response to the security risks posed by the realistic propagation of manipulated media data, detecting and grounding multi-modal media manipulation has received attention as a challenging task. However, there is multi-modal contribution imbalance on current approach for cross-modal learning, which affects model performance optimisation. To this end, we propose an Adaptive Contribution Modulation (ACM) framework to solve the problem of multi-modal contribution imbalance. To balance the image and text embedding features before fusion, we propose adaptive weight decision to computes dynamic weights for fusion features, which enable more adaptive and robust decision-making. Meanwhile, we propose contribution modulation block, which dynamically governs the contributions of different modalities for optimization. Based on cross-modal contrastive learning, we balance image and text embeddings contribution through multi-modal contribution balanced learning, which makes better use of the semantic correlation of all modalities. We conduct experiments on the DGM4 dataset, which demonstrate the superior performance of our approach through compared to state-of-the-art methods.
Yixiang Li, Biao Leng
ICASSP2
2025 Spatial-Frequency Cross-Domain Mutual Guidance for Visible-Infrared Image Fusion
abstract
The primary objective of Visible-Infrared Image Fusion (VIF) is to combine the rich texture and color information from visible light images with the comprehensive thermal radiation data provided by infrared images. However, most current fusion algorithms focus solely on spatial domain feature transformations, which results in fused images lacking sufficient detail and failing to effectively preserve crucial details from the source images. In this paper, we propose Spatial-Frequency Mutual Guidance for VIF. The framework comprises two branches. Each branch reconstructs input features in the frequency and spatial domains. We introduce a novel cross-domain mutual guidance mechanism. It fully integrates information between the frequency and spatial domains to enhance the detail quality of fused images. Furthermore, a weight allocation network is utilized to adaptively assign importance to visible and infrared images based on scene characteristics. Experiments on three VIF datasets show that our method outperforms recent advanced algorithms.
Zhili Lin, Biao Leng
ICASSP2
2025 Multi-agent reinforcement learning with graph representation for green edge-cloud computation offloading
Yifan Bo, Jinghan Feng, Shou Zhang, Biao Leng
Comput. Commun.4
2025 ImVoxelENet: Image to Voxels Epipolar Transformer for Multi-View RGB-Based 3D Object Detection
abstract
The task of detecting three-dimensional objects using only RGB images presents a considerable challenge within the domain of computer vision. The core issue lies in accurately performing epipolar geometry matching between multiple views to obtain latent geometric priors. Existing methods establish correspondences along epipolar line features in voxel space through various layers of convolution. However, this step often occurs in the later stages of the network, which limits overall performance. To address this challenge, we introduce a novel framework, ImVoxelENet, that integrates a geometric epipolar constraint. We start from the back-projection of pixel-wise features and design an attention mechanism that captures the relationship between forward and backward features along the ray for multiple views. This approach enables the early establishment of geometric correspondences and structural connections between epipolar lines. Using ScanNetV2 as a benchmark, extensive comparative and ablation experiments demonstrate that our proposed network achieves a 1.1% improvement in mAP, highlighting its effectiveness in enhancing 3D object detection performance. Our code is available at https://github.com/xug-coder/ImVoxelENet.
Biao Leng, Zhang Xiong 0001
Comput. Vis. Media3
2025 Cross-domain data-driven reinforcement learning for IGSO satellite coverage optimization
Yifan Bo, Biao Leng
Neurocomputing3
2025 HGFormer: Topology-Aware Vision Transformer With HyperGraph Learning
abstract
The computer vision community has witnessed an extensive exploration of vision transformers in the past two years. Drawing inspiration from traditional schemes, numerous works focus on introducing vision-specific inductive biases. However, the implicit modeling of permutation invariance and fully-connected interaction with individual tokens disrupts the regional context and spatial topology, further hindering higher-order modeling. This deviates from the principle of perceptual organization that emphasizes the local groups and overall topology of visual elements. Thus, we introduce the concept of hypergraph for perceptual exploration. Specifically, we propose a topology-aware vision transformer called HyperGraph Transformer (HGFormer). Firstly, we present a Center Sampling K-Nearest Neighbors (CS-KNN) algorithm for semantic guidance during hypergraph construction. Secondly, we present a topology-aware HyperGraph Attention (HGA) mechanism that integrates hypergraph topology as perceptual indications to guide the aggregation of global and unbiased information during hypergraph messaging. Using HGFormer as visual backbone, we develop an effective and unitive representation, achieving distinct and detailed scene depictions. Empirical experiments show that the proposed HGFormer achieves competitive performance compared to the recent SoTA counterparts on various visual benchmarks. Extensive ablation and visualization studies provide comprehensive explanations of our ideas and contributions.
Shuo Zhang 0003, Biao Leng
IEEE Trans. Multim.3
2024 Dual-Modeling Decouple Distillation for Unsupervised Anomaly Detection
abstract
Knowledge distillation based on student-teacher network is one of the mainstream solution paradigms for the challenging unsupervised Anomaly Detection task, utilizing the difference in representation capabilities of the teacher and student networks to implement anomaly localization. However, over-generalization of the student network to the teacher network may lead to negligible differences in representation capabilities of anomaly, thus affecting the detection effectiveness. Existing methods address the possible over-generalization by using differentiated students and teachers from the structural perspective or explicitly expanding distilled information from the content perspective, which inevitably results in an increased likelihood of underfitting of the student network and poor anomaly detection capabilities in anomaly center or edge. In this paper, we propose Dual-Modeling Decouple Distillation (DMDD) for the unsupervised Anomaly Detection. In DMDD, a Decouple Student-Teacher Network is proposed to decouple the initial student features into normality and abnormality features. We further introduce Dual-Modeling Distillation based on normal-anomalous image pairs, fitting normality features of anomalous image and the teacher features of the corresponding normal image, widening the distance between abnormality features and the teacher features in anomalous regions. Synthesizing these two distillation ideas, we achieve anomaly detection which focuses on both edge and center of anomaly. Finally, a Multi-perception Segmentation Network is proposed to achieve focused anomaly map fusion based on multiple attention. Experimental results on MVTec AD show that DMDD surpasses SOTA localization performance of previous knowledge distillation-based methods, reaching 98.85% on pixel-level AUC and 96.13% on PRO.
Biao Leng, Shuo Zhang 0003
ACM Multimedia3
2023 Efficient Lightweight Network with Transformer-Based Distillation for Micro-crack Detection of Solar Cells
Xiangying Xie, QiXiang Chen, Biao Leng
ICONIP (3)4
2022 SHRAG: Semantic Hierarchical Graph for Floorplan Representation
abstract
Representation learning from a floorplan is a fundamental step for various floorplan-related applications, such as retrieval, reconstruction, and generation. Previous works often use single-layer graphs to model room categories and adjacencies as nodes and edges, which can only represent a floorplan in a coarse level without detailed information such as room shape and door/window locations. Thus, we propose SHRAG, a hierarchical semantic graph for floor-plan representation, with two hierarchical graph modules to semantically encode both coarse and detailed information of a floorplan. First, a Detail Graph Module(DGM) was designed to learn detailed contour and attachable elements embedding for each room. Then, a Global Graph Module(GGM) was applied to concatenate and encode both detailed embeddings from DGM and coarse information including room categories and room relations. Results show that the representation learned from SHRAG achieved SOTA on floorplan retrieval tasks. Further, SHRAG performs more satisfying than other floorplan similarity metrics in a comprehensive user study. We also demonstrate that SHRAG can be easily adapted to real-world complex indoor scenes even with furniture.
Jiongchao Jin, Zhou Xue, Biao Leng
3DV3
2022 Unifying Visual Perception by Dispersible Points Learning
Jianming Liang, Guanglu Song, Biao Leng, Yu Liu 0015
ECCV (9)3
2022 Self-slimmed Vision Transformer
Zhuofan Zong, Kunchang Li 0002, Guanglu Song, Yali Wang 0001, Yu Qiao 0001, Biao Leng, Yu Liu 0015
ECCV (11)6
2022 IMC-NET: Learning Implicit Field with Corner Attention Network for 3D Shape Reconstruction
abstract
Learning implicit field as shape representation brings a revolution in 3D shape single view reconstruction, while it suffers from over-smoothness of corners and details. Many involve 2D supervision to solve corners and details. However, they are limited by the number of views and the accuracy of the renderer. Thus, we propose IMC-NET, an IMplicit field with Corner attention NETwork, to reconstruct details and corners of 3D shapes without any 2D supervision. In IMC-NET, we designed a corner attention branch to predict detailed corners of 3D shape under hand-crafted 3D corner points supervision. Then, an aggregation module was presented to differentiably merge the shape corners and implicit fields using a corner weight prediction. In addition, to help further shape reconstruction work on corners and details, we released a new dataset based on Shapenet models, named Shape-CornerNet, including 43784 3D shapes in 13 categories with detailed corner points. Afterward, we quantitatively and qualitatively evaluated our model on Shape-CornerNet. Our IMC-NET achieved state-of-the-art performance on Shape-CornerNet.
Jiongchao Jin, Huanqiang Xu, Pengliang Ji, Biao Leng
ICIP4
2021 PCNET: Parallelly Conquer the Large Variance of Person Re-Identification
abstract
Person re-identification has a wide range of applications, and many state-of-the-art methods are proposed to solve the problem under specific scenarios. However, it is still a challenging issue because of the large variance in practical applications, such as pose variations, misalignment, and image noises.In this paper, Parallelly Conquer Net (PCNet) is proposed to deal with large variance in a parallel manner. PCNet consists of three module: Pose Adaptation Module (PAM), Global Alignment Module (GAM), and Pixel-Wised Attention Module (PWAM). Each module is designed to deal with a sub-variance independently. Furthermore, the generated features are aggregated by parallel branches to utilize complementary information among them. Extensive experiments on three benchmarks (Market-1501, DukeMTMC-reID, and CUHK03) demonstrate the effectiveness of the method. The results show that PCNet can significantly improve the performance of person re-identification.
Meiyue You, Biao Leng, Guanglu Song
ICIP3
2021 Scale Semantic Flow Preserving Across Image Pyramid
Zhili Lin, Guanglu Song, Biao Leng
ICONIP (5)3
2021 Hint Pyramid Learning Via Salient Semantic Mining
abstract
Feature pyramid network (FPN) has dominated the task of object detection for many years with its congenital capability of semantic information interaction between different feature levels. Satisfactory performance can be easily achieved via the heavy backbones or tricky designs of anchors. However, the heavy structure brings about a huge increase of parameters and calculation. To address this challenge, this paper proposes a novel Hint Pyramid Learning(HPL) strategy to optimize a lightweight FPN-based detector such as one based on ResNet18. In HPL, the hint feature representation is learnt via salient semantic mining based on the teacher-student learning mechanism and self learning mechanism. Besides minimizing of the discrepancy of features belonging to the same feature pyramid layer, HPL also focuses on the relationships between feature pyramid layers. Unlike the popular mimicking algorithms such as hint learning and knowledge distillation which achieve comparable performance on classification tasks but failed to generalize to the detection tasks with FPN, HPL can be easily plugged into different FPN-based pipelines such as one-stage RetinaNet or two-stage Faster R-CNN and improves the performance straightforwardly. Without any meticulous designing, when the backbone of students is ResNet18 and the backbone of teachers is ResNet50, HPL effectively improves the performance by 2.1 AP for RetinaNet and 1.0 AP for Faster R-CNN on COCO benchmark.
Qianggang Cao, Biao Leng
IJCNN4
2021 RCNet: Reverse Feature Pyramid and Cross-scale Shift Network for Object Detection
abstract
Feature pyramid networks (FPN) are widely exploited for multi-scale feature fusion in existing advanced object detection frameworks. Numerous previous works have developed various structures for bidirectional feature fusion, all of which are shown to improve the detection performance effectively. We observe that these complicated network structures require feature pyramids to be stacked in a fixed order, which introduces longer pipelines and reduces the inference speed. Moreover, semantics from non-adjacent levels are diluted in the feature pyramid since only features at adjacent pyramid levels are merged by the local fusion operation in a sequence manner. To address these issues, we propose a novel architecture named RCNet, which consists of Reverse Feature Pyramid (RevFP) and Cross-scale Shift Network (CSN). RevFP utilizes local bidirectional feature fusion to simplify the bidirectional pyramid inference pipeline. CSN directly propagates representations to both adjacent and non-adjacent levels to enable multi-scale features more correlative. Extensive experiments on the MS COCO dataset demonstrate RCNet can consistently bring significant improvements over both one-stage and two-stage detectors with subtle extra computational overhead. In particular, RetinaNet is boosted to 40.2 AP, which is 3.7 points higher than baseline, by replacing FPN with our proposed model. On COCO test-dev, RCNet can achieve very competitive performance with a single-model single-scale 50.5 AP.
Zhuofan Zong, Qianggang Cao, Biao Leng
ACM Multimedia3
2020 KPNet: Towards Minimal Face Detector
Guanglu Song, Yu Liu 0015, Yuhang Zang, Xiaogang Wang 0001, Biao Leng, Qingsheng Yuan
AAAI5
2020 Improving Multi-view Stereo with Contextual 2D-3D Skip Connection
Biao Leng
ICONIP (2)3
2020 Weighted triple-sequence loss for video-based person re-identification
Biao Leng, Guanglu Song
Neurocomputing2
2020 Learning Discriminative and Generative Shape Embeddings for Three-Dimensional Shape Retrieval
abstract
As an important solution for 3D shape retrieval, a multi-view shape descriptor has achieved impressive performance. One crucial part of view-based shape descriptors is to interpret 3D structures through various 2D observations. Most existing methods like MVCNN believe that a strong classification model trained with deep learning, can often provide an efficient shape embedding for 3D shape retrieval. However, these methods pay much attention to discriminative models and none of them necessarily incorporate the underlying 3D properties of the objects from 2D images. In this paper, we present a novel encoder-decoder recurrent feature aggregation network (ERFA-Net) to address this problem. Aiming at emphasizing the 3D properties of 3D shapes in the fusion of multiple view features, 3D properties prediction tasks are introduced into the 3D shape retrieval. Specifically, an image sequence of the shape is recurrently aggregated into a discriminative shape embedding based on LSTM network, and then this latent shape embedding is trained to predict the original voxel grids and estimate images of unseen viewpoints. This generation task gives an effective supervision which makes the network exploit 3D properties of shapes through various 2D images. Our method achieves the state-of-the-art performance for 3D shape retrieval, on two large-scale 3D shape datasets, ModelNet and ShapeNetCore55. Extensive experiments show that the proposed 3D representation performs robust discrimination against view occlusion, and strong generation ability for various 3D shape tasks.
Biao Leng, Xiaochen Zhou
IEEE Trans. Multim.2
2019 Angular Triplet-Center Loss for Multi-View 3D Shape Retrieval
abstract
How to obtain the desirable representation of a 3D shape, which is discriminative across categories and polymerized within classes, is a significant challenge in 3D shape retrieval. Most existing 3D shape retrieval methods focus on capturing strong discriminative shape representation with softmax loss for the classification task, while the shape feature learning with metric loss is neglected for 3D shape retrieval. In this paper, we address this problem based on the intuition that the cosine distance of shape embeddings should be close enough within the same class and far away across categories. Since most of 3D shape retrieval tasks use cosine distance of shape features for measuring shape similarity, we propose a novel metric loss named angular triplet-center loss, which directly optimizes the cosine distances between the features. It inherits the triplet-center loss property to achieve larger inter-class distance and smaller intra-class distance simultaneously. Unlike previous metric loss utilized in 3D shape retrieval methods, where Euclidean distance is adopted and the margin design is difficult, the proposed method is more convenient to train feature embeddings and more suitable for 3D shape retrieval. Moreover, the angle margin is adopted to replace the cosine margin in order to provide more explicit discriminative constraints on an embedding space. Extensive experimental results on two popular 3D object retrieval benchmarks, ModelNet40 and ShapeNetCore 55, demonstrate the effectiveness of our proposed loss, and our method has achieved state-ofthe-art results on various 3D shape datasets.
Zhaoqun Li, Biao Leng
AAAI3
2019 Enhancing 2D Representation via Adjacent Views for 3D Shape Retrieval
abstract
Multi-view shape descriptors obtained from various 2D images are commonly adopted in 3D shape retrieval. One major challenge is that significant shape information are discarded during 2D view rendering through projection. In this paper, we propose a convolutional neural network based method, CenterNet, to enhance each individual 2D view using its neighboring ones. By exploiting cross-view correlations, CenterNet learns how adjacent views can be maximally incorporated for an enhanced 2D representation to effectively describe shapes. We observe that a very small amount of, e.g., six, enhanced 2D views, are already sufficient for a panoramic shape description. Thus, by simply aggregating features from six enhanced 2D views, we arrive at a highly compact yet discriminative shape descriptor. The proposed shape descriptor significantly outperforms state-of-the-art 3D shape retrieval methods on the ModelNet and ShapeNetCore55 benchmarks, and also exhibits robustness against object occlusion.
Zhaoqun Li, Biao Leng, Jingfei Jiang
ICCV4
2019 Rethinking Loss Design for Large-scale 3D Shape Retrieval
abstract
Learning discriminative shape representations is a crucial issue for large-scale 3D shape retrieval. In this paper, we propose the Collaborative Inner Product Loss (CIP Loss) to obtain ideal shape embedding that discriminative among different categories and clustered within the same class. Utilizing simple inner product operation, CIP loss explicitly enforces the features of the same class to be clustered in a linear subspace, while inter-class subspaces are constrained to be at least orthogonal. Compared to previous metric loss functions, CIP loss could provide more clear geometric interpretation for the embedding than Euclidean margin, and is easy to implement without normalization operation referring to cosine margin. Moreover, our proposed loss term can combine with other commonly used loss functions and can be easily plugged into existing off-the-shelf architectures. Extensive experiments conducted on the two public 3D object retrieval datasets, ModelNet and ShapeNetCore 55, demonstrate the effectiveness of our proposal, and our method has achieved state-of-the-art results on both datasets.
Zhaoqun Li, Biao Leng
IJCAI3
2019 Learning Discriminative 3D Shape Representations by View Discerning Networks
abstract
In view-based 3D shape recognition, extracting discriminative visual representation of 3D shapes from projected images is considered the core problem. Projections with low discriminative ability can adversely influence the final 3D shape representation. Especially under the real situations with background clutter and object occlusion, the adverse effect is even more severe. To resolve this problem, we propose a novel deep neural network, View Discerning Network, which learns to judge the quality of views and adjust their contributions to the representation of shapes. In this network, a Score Generation Unit is devised to evaluate the quality of each projected image with score vectors. These score vectors are used to weight the image features and the weighted features perform much better than original features in 3D shape recognition task. In particular, we introduce two structures of Score Generation Unit, Channel-wise Score Unit and Part-wise Score Unit, to assess the quality of feature maps from different perspectives. Our network aggregates features and scores in an end-to-end framework, so that final shape descriptors are directly obtained from its output. Our experiments on ModelNet and ShapeNet Core55 show that View Discerning Network outperforms the state-of-the-arts in terms of the retrieval task, with excellent robustness against background clutter and object occlusion.
Biao Leng, Xiaocheng Zhou, Kai Xu 0004
IEEE Trans. Vis. Comput. Graph.1
2018 Region-Based Quality Estimation Network for Large-Scale Person Re-Identification
abstract
One of the major restrictions on the performance of video-based person re-id is partial noise caused by occlusion, blur and illumination. Since different spatial regions of a single frame have various quality, and the quality of the same region also varies across frames in a tracklet, a good way to address the problem is to effectively aggregate complementary information from all frames in a sequence, using better regions from other frames to compensate the influence of an image region with poor quality. To achieve this, we propose a novel Region-based Quality Estimation Network (RQEN), in which an ingenious training mechanism enables the effective learning to extract the complementary region-based information between different frames. Compared with other feature extraction methods, we achieved comparable results of 92.4%, 76.1% and 77.83% on the PRID 2011, iLIDS-VID and MARS, respectively. In addition, to alleviate the lack of clean large-scale person re-id datasets for the community, this paper also contributes a new high-quality dataset, named "Labeled Pedestrian in the Wild (LPW)" which contains 7,694 tracklets with over 590,000 images. Despite its relatively large scale, the annotations also possess high cleanliness. Moreover, it's more challenging in the following aspects: the age of characters varies from childhood to elderhood; the postures of people are diverse, including running and cycling in addition to the normal walking state.
Guanglu Song, Biao Leng, Yu Liu 0015, Congrui Hetang, Shaofan Cai
AAAI2
2018 Emphasizing 3D Properties in Recurrent Multi-View Aggregation for 3D Shape Retrieval
abstract
Multi-view based shape descriptors have achieved impressive performance for 3D shape retrieval. The core of view-based methods is to interpret 3D structures through 2D observations. However, most existing methods pay more attention to discriminative models and none of them necessarily incorporate the 3D properties of the objects. To resolve this problem, we propose an encoder-decoder recurrent feature aggregation network (ERFA-Net) to emphasize the 3D properties of 3D shapes in multi-view features aggregation. In our network, a view sequence of the shape is trained to encode a discriminative shape embedding and estimate unseen rendered views of any viewpoints. This generation task gives an effective supervision which makes the network exploit 3D properties of shapes through various 2D images. During feature aggregation, a discriminative feature representation across multiple views is effectively exploited based on LSTM network. The proposed 3D representation has following advantages against other state-of-the-art: 1) it performs robust discrimination under the existence of noise such as view missing and occlusion, because of the improvement brought by 3D properties. 2) it has strong generative capabilities, which is useful for various 3D shape tasks. We evaluate ERFA-Net on two popular 3D shape datasets, ModelNet and ShapeNetCore55, and ERFA-Net outperforms the state-of-the-art methods significantly. Extensive experiments show the effectiveness and robustness of the proposed 3D representation.
Biao Leng, Xiaochen Zhou
AAAI2
2018 Beyond Trade-Off: Accelerate FCN-Based Face Detector With Higher Accuracy
abstract
Fully convolutional neural network (FCN) has been dominating the game of face detection task for a few years with its congenital capability of sliding-window-searching with shared kernels, which boiled down all the redundant calculation, and most recent state-of-the-art methods such as Faster-RCNN, SSD, YOLO and FPN use FCN as their backbone. So here comes one question: Can we find a universal strategy to further accelerate FCN with higher accuracy, so could accelerate all the recent FCN-based methods? To analyze this, we decompose the face searching space into two orthogonal directions, 'scale' and 'spatial'. Only a few coordinates in the space expanded by the two base vectors indicate foreground. So if FCN could ignore most of the other points, the searching space and false alarm should be significantly boiled down. Based on this philosophy, a novel method named scale estimation and spatial attention proposal (S2AP) is proposed to pay attention to some specific scales in image pyramid and valid locations in each scales layer. Furthermore, we adopt a masked-convolution operation based on the attention result to accelerate FCN calculation. Experiments show that FCN-based method RPN can be accelerated by about 4Ã- with the help of S2AP and masked-FCN and at the same time it can also achieve the state-of-the-art on FDDB, AFW and MALF face detection benchmarks as well.
Guanglu Song, Yu Liu 0015, Biao Leng
CVPR6
2018 Fast Portrait Matting Using Spatial Detail-Preserving Network
Shaofan Cai, Biao Leng, Guanglu Song, Zheng Ge
ICONIP (6)2
2018 Dimensionality Reduction in Multiple Ordinal Regression
abstract
Supervised dimensionality reduction (DR) plays an important role in learning systems with high-dimensional data. It projects the data into a low-dimensional subspace and keeps the projected data distinguishable in different classes. In addition to preserving the discriminant information for binary or multiple classes, some real-world applications also require keeping the preference degrees of assigning the data to multiple aspects, e.g., to keep the different intensities for co-occurring facial expressions or the product ratings in different aspects. To address this issue, we propose a novel supervised DR method for DR in multiple ordinal regression (DRMOR), whose projected subspace preserves all the ordinal information in multiple aspects or labels. We formulate this problem as a joint optimization framework to simultaneously perform DR and ordinal regression. In contrast to most existing DR methods, which are conducted independently of the subsequent classification or ordinal regression, the proposed framework fully benefits from both of the procedures. We experimentally demonstrate that the proposed DRMOR method (DRMOR-M) well preserves the ordinal information from all the aspects or labels in the learned subspace. Moreover, DRMOR-M exhibits advantages compared with representative DR or ordinal regression algorithms on three standard data sets.
Jiabei Zeng, Yang Liu 0007, Biao Leng, Zhang Xiong 0001, Yiu-Ming Cheung
IEEE Trans. Neural Networks Learn. Syst.3
2017 Weakly-supervised Learning of Mid-level Features for Pedestrian Attribute Recognition and Localization
Kai Yu 0003, Biao Leng, Zhang Zhang 0001, Dangwei Li, Kaiqi Huang
BMVC3
2017 A Multi-level Weighted Representation for Person Re-identification
Xianglai Meng, Biao Leng, Guanglu Song
ICANN (2)2
2017 Computer-Aided Diagnosis in Chest Radiography with Deep Multi-Instance Learning
Kang Qu, Xiangfei Chai, Biao Leng, Zhang Xiong 0001
ICONIP (4)5
2017 Spatial Quality Aware Network for Video-Based Person Re-identification
Biao Leng, Guanglu Song
ICONIP (3)2
2017 Data augmentation for unbalanced face recognition training sets
Biao Leng, Kai Yu 0003, Jingyan Qin
Neurocomputing1
2017 3D Object retrieval based on viewpoint segmentation
Biao Leng, Changchun Du, Jiabei Zeng, Zhang Xiong 0001
Multim. Syst.1
2016 Cascade shallow CNN structure for face verification and identification
Biao Leng, Yu Liu 0015, Kai Yu 0003, Songting Xu, Ziqing Yuan, Jingyan Qin
Neurocomputing1
2016 3D object understanding with 3D Convolutional Neural Networks
Biao Leng, Yu Liu 0015, Kai Yu 0003, Zhang Xiong 0001
Inf. Sci.1
2016 Analysis of Taxi Drivers' Behaviors Within a Battle Between Two Taxi Apps
abstract
A battle between two Chinese taxi booking mobile apps, namely, Didi and Kuaidadi, had recently occurred in early 2014. These two apps, which are backed by Internet giants Tencent and Alipay, gave promotion fees to taxi drivers for each deal made and also allowed each taxi passenger to save some money, when a customer had taken a taxi through the app and paid the fare through the mobile payment method. As expected, the taxi service pattern had been greatly changed during this battle. To address the debates on social justice, equity, and improvements of taxi service, we collect 37-day trip data of over 9000 taxis in Beijing to study the influence of this pattern change. In the first 18 days, the battle had not occurred and in the remaining 19 days, the battle is white-hot. We quantitatively demonstrate how several important service indices (e.g., the traveling distances and idle time lengths) of taxi drivers had been changed. The spatial-temporal traveling patterns of taxis are then studied. Based on comprehensive analysis, the benefits and drawbacks brought by money promotion are finally discussed. The obtained results indicate that productively employing big data can help answer some important questions attracting the interest of the whole society.
Biao Leng, Li Li 0013, Zhang Xiong 0001
IEEE Trans. Intell. Transp. Syst.1
2015 A powerful 3D model classification mechanism based on fusing multi-graph
Biao Leng, Changchun Du, Zhang Xiong 0001
Neurocomputing1
2015 A 3D model recognition mechanism based on deep Boltzmann machines
Biao Leng, Zhang Xiong 0001
Neurocomputing1
2015 A novel wavelet-SVM short-time passenger flow prediction in Beijing subway system
Yuxing Sun, Biao Leng
Neurocomputing2
2015 3-D object retrieval using topic model
Jiabei Zeng, Biao Leng, Zhang Xiong 0001
Multim. Tools Appl.2
2015 3D object retrieval with stacked local convolutional autoencoder
Biao Leng, Zhang Xiong 0001
Signal Process.1
2015 3D Object Retrieval With Multitopic Model Combining Relevance Feedback and LDA Model
abstract
View-based 3D model retrieval uses a set of views to represent each object. Discovering the complex relationship between multiple views remains challenging in 3D object retrieval. Recent progress in the latent Dirichlet allocation (LDA) model leads us to propose its use for 3D object retrieval. This LDA approach explores the hidden relationships between extracted primordial features of these views. Since LDA is limited to a fixed number of topics, we further propose a multitopic model to improve retrieval performance. We take advantage of a relevance feedback mechanism to balance the contributions of multiple topic models with specified numbers of topics. We demonstrate our improved retrieval performance over the state-of-the-art approaches.
Biao Leng, Jiabei Zeng, Zhang Xiong 0001
IEEE Trans. Image Process.1
2014 3D Object Classification Using Deep Belief Networks
Biao Leng, Zhang Xiong 0001
MMM (2)1
2013 Probability tree based passenger flow prediction and its application to the Beijing subway system
Biao Leng, Jiabei Zeng, Zhang Xiong 0001, Weifeng Lv, Yueliang Wan
Frontiers Comput. Sci.1
2011 ModelSeek: an effective 3D model retrieval system
Biao Leng, Zhang Xiong 0001
Multim. Tools Appl.1
2010 A 3D shape retrieval framework for 3D smart cities
Biao Leng, Zhang Xiong 0001, Xiangwei Fu
Frontiers Comput. Sci. China1
2008 A powerful relevance feedback mechanism for content-based 3D model retrieval
Biao Leng
Multim. Tools Appl.1
2007 OFS: A Feature Selection Method for Shape-based 3D Model Retrieval
abstract
We focus on improving the effectiveness of shape-based similarity retrieval in 3D model repositories. Motivated by retrieval performance of several individual 3D model descriptors for projected images in shape-based approaches, we present an optimized feature selection (OFS) method to choose a perfect feature vector based on each query model. Experimental results show that the OFS method for shape-based 3D model retrieval has achieved significant improvements on retrieval effectiveness of 3D shape search with several measures on a standard 3D database, and it provides a retrieval performance 45.5% better than the average precision of several descriptors. Compared to the currently best method light field descriptor (LFD), OFS has better retrieval effectiveness. Furthermore, the feature vector components of our approach are only 6.77% of that in LFD.
Biao Leng
CAD/Graphics2