EDBT 2026 Demo / reviewers in the wild / expert
Bin Feng 0001
dblp:04/4053-1
· DBLP profile ↗
38ranked-venue papers
4as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 6Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gait Recognition via Collaborating Discriminative and Generative Diffusion ModelsabstractGait recognition offers a non-intrusive biometric solution by identifying individuals through their walking patterns. Although discriminative models have achieved notable success in this domain, the full potential of generative models remains largely unexplored. In this paper, we introduce CoD², a novel framework that combines the data distribution modeling capabilities of diffusion models with the semantic representation learning strengths of discriminative models to extract robust gait features. We propose a Multi-level Conditional Control strategy that integrates both high-level identity-aware semantic conditions and low-level visual details. Specifically, the high-level condition, extracted by the discriminative extractor, guides the generation of identity-consistent gait sequences, while low-level visual details, such as appearance and motion, are preserved to enhance consistency. Moreover, the generated sequences facilitate the discriminative extractor's learning, enabling it to capture more comprehensive high-level semantic features. Extensive experiments on four datasets (SUSTech1K, CCPG, GREW, and Gait3D) demonstrate that CoD² achieves state-of-the-art performance and can be seamlessly integrated with existing discriminative methods, yielding consistent improvements. Haijun Xiong, Bin Feng 0001, Bang Wang 0001, Xinggang Wang, Wenyu Liu 0001 |
AAAI | 2 |
| 2026 | MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement LearningabstractOptical Chemical Structure Recognition (OCSR) plays a pivotal role in modern chemical informatics, enabling the automated conversion of chemical structure images from scientific literature, patents, and educational materials into machine-readable molecular representations. This capability is essential for large-scale chemical data mining, drug discovery pipelines, and Large Language Model (LLM) applications in related domains. However, existing OCSR systems face significant challenges in accurately recognizing stereochemical information due to the subtle visual cues that distinguish stereoisomers, such as wedge and dash bonds, ring conformations, and spatial arrangements. To address these challenges, we propose MolSight, a comprehensive learning framework for OCSR that employs a three-stage training paradigm. In the first stage, we conduct pre-training on large-scale but noisy datasets to endow the model with fundamental perception capabilities for chemical structure images. In the second stage, we perform multi-granularity fine-tuning using datasets with richer supervisory signals, systematically exploring how auxiliary tasks—specifically chemical bond classification and atom localization—contribute to molecular formula recognition. Finally, we employ reinforcement learning for post-training optimization and introduce a novel stereochemical structure dataset. Remarkably, we find that even with MolSight's relatively compact parameter size, the Group Relative Policy Optimization (GRPO) algorithm can further enhance the model's performance on stereomolecular. Through extensive experiments across diverse datasets, our results demonstrate that MolSight achieves state-of-the-art performance in (stereo)chemical optical structure recognition. Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001 |
AAAI | 3 |
| 2026 | Fourier-based adaptive counterfactual intervention for object re-identification
Haijun Xiong, Bin Feng 0001, Bang Wang 0001, Xinggang Wang, Wenyu Liu 0001 |
Neural Networks | 2 |
| 2025 | Mask-Adapter: The Devil is in the Masks for Open-Vocabulary SegmentationabstractRecent open-vocabulary segmentation methods adopt mask generators to predict segmentation masks and leverage pretrained vision-language models, e.g., CLIP, to classify these masks via mask pooling. Although these approaches show promising results, it is counterintuitive that accurate masks often fail to yield accurate classification results through pooling CLIP image embeddings within the mask regions. In this paper, we reveal the performance limitations of mask pooling and introduce Mask-Adapter, a simple yet effective method to address these challenges in open-vocabulary segmentation. Compared to directly using proposal masks, our proposed Mask-Adapter extracts semantic activation maps from proposal masks, providing richer contextual information and ensuring alignment between masks and CLIP. Additionally, we propose a mask consistency loss that encourages proposal masks with similar IoUs to obtain similar CLIP embeddings to enhance models’ robustness to varying predicted masks. Mask-Adapter integrates seamlessly into open-vocabulary segmentation methods based on mask pooling in a plug-and-play manner, delivering more accurate classification results. Extensive experiments across several zero-shot benchmarks demonstrate significant performance gains for the proposed Mask-Adapter on several well-established methods. Notably, Mask-Adapter also extends effectively to SAM and achieves impressive results on several open-vocabulary segmentation datasets. Code and models are available at https://github.com/hustvl/MaskAdapter. Yongkang Li 0005, Tianheng Cheng, Bin Feng 0001, Wenyu Liu 0001, Xinggang Wang |
CVPR | 3 |
| 2025 | Collaborative Spatial and Channel Attention for Structural Vibration-based Gait RecognitionabstractStructural vibration-based gait recognition aims to identify pedestrians through their unique vibration patterns and has gained significant attention for its non-invasive nature. However, existing methods primarily rely on manually crafted features and convolutional neural networks (CNNs) to extract local details, without considering global gait features in vibration signals. To address this limitation, we propose CSCA, a novel framework designed to capture discriminative global gait features through collaborative spatial and channel attention. Specifically, CSCA consists of two key components: Mamba-inspired Linear Channel Attention (MLCA) and Strip Pooling-guided Spatial Attention (SPSA). MLCA integrates Mamba with linear attention to enhance discriminative attention in the channel dimension by capturing global gait features. Additionally, SPSA integrates axial strip pooling with multikernel depth-wise convolutions to extract multi-level spatial semantic features. Experimental results on the VIBEID dataset demonstrate that CSCA achieves state-of-the-art performance across diverse environmental conditions, with identity recognition accuracy exceeding 96%. Code is available at https://github.com/xiaomush/CSCA. Junhao Lu, Haijun Xiong, Ziyu Lin, Bin Feng 0001 |
IJCB | 5 |
| 2025 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition. Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001 |
IJCB | 28 |
| 2025 | CGTGait: Collaborative Graph and Transformer for Gait Emotion RecognitionabstractSkeleton-based gait emotion recognition has received significant attention due to its wide-ranging applications. However, existing methods primarily focus on extracting spatial and local temporal motion information, failing to capture long-range temporal representations. In this paper, we propose CGTGait, a novel framework that collaboratively integrates graph convolution and transformers to extract discriminative spatiotemporal features for gait emotion recognition. Specifically, CGTGait consists of multiple CGT blocks, where each block employs graph convolution to capture frame-level spatial topology and the transformer to model global temporal dependencies. Additionally, we introduce a Bidirectional Cross-Stream Fusion (BCSF) module to effectively aggregate posture and motion spatiotemporal features, facilitating the exchange of complementary information between the two streams. We evaluate our method on two widely used datasets, Emotion-Gait and ELMD, demonstrating that our CGTGait achieves state-of-the-art or at least competitive performance while reducing computational complexity by approximately 82.2% (only requiring 0.34G FLOPs) during testing. Code is available at https://github.com/githubzjj1/CGTGait. Haijun Xiong, Junhao Lu, Ziyu Lin, Bin Feng 0001 |
IJCB | 5 |
| 2025 | STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-To-4D Gaussian SplattingabstractText-To-4D generation is rapidly developing and widely applied in various scenarios. However, existing methods often fail to incorporate adequate spatio-temporal modeling and prompt alignment within a unified framework, resulting in temporal inconsistencies, geometric distortions, or low-quality 4D content that deviates from the provided texts. Therefore, we propose STP4D, a novel approach that aims to integrate comprehensive spatio-temporal-prompt consistency modeling for high-quality text-to-4D generation. Specifically, STP4D employs three carefully designed modules: Time-Varying Prompt Embedding, Geometric Information Enhancement, and Temporal Extension Deformation, which collaborate to accomplish this goal. Furthermore, STP4D is among the first methods to exploit the Diffusion model to generate 4D Gaussians, combining the fine-grained modeling capabilities and the real-time rendering process of 4DGS with the rapid inference speed of the Diffusion model. Extensive experiments demonstrate that STP4D excels in generating high-fidelity 4D content with exceptional efficiency (approximately 4.6s per asset), surpassing existing methods in both quality and speed. Yunze Deng, Haijun Xiong, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001 |
ICME | 3 |
| 2025 | Dynamic 2D Gaussians: Geometrically Accurate Radiance Fields for Dynamic ObjectsabstractReconstructing objects and extracting high-quality surfaces play a vital role in the real world. Current 4D representations show the ability to render high-quality novel views for dynamic objects, but cannot reconstruct high-quality meshes due to their implicit or geometrically inaccurate representations. In this paper, we propose a novel representation that can reconstruct accurate meshes from sparse image input, named Dynamic 2D Gaussians (D-2DGS). We adopt 2D Gaussians for basic geometry representation and use sparse-controlled points to capture the 2D Gaussian's deformation. By extracting the object mask from the rendered high-quality image and masking the rendered depth map, we remove floaters that are prone to occur during reconstruction and can extract high-quality dynamic mesh sequences of dynamic objects. Experiments demonstrate that our D-2DGS is outstanding in reconstructing detailed and smooth high-quality meshes from sparse inputs. The code is available at https://github.com/hustvl/Dynamic-2DGS. Shuai Zhang 0050, Guanjun Wu, Zhoufeng Xie, Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001 |
ACM Multimedia | 5 |
| 2025 | Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical OperationsabstractWhile large language models (LLMs) with Chain-of-Thought (CoT) reasoning excel in mathematics and coding, their potential for systematic reasoning in chemistry, a domain demanding rigorous structural analysis for real-world tasks like drug design and reaction engineering, remains untapped. Current benchmarks focus on simple knowledge retrieval, neglecting step-by-step reasoning required for complex tasks such as molecular optimization and reaction prediction. To address this, we introduce ChemCoTBench, a reasoning framework that bridges molecular structure understanding with arithmetic-inspired operations, including addition, deletion, and substitution, to formalize chemical problem-solving into transparent, step-by-step workflows. By treating molecular transformations as modular "chemical operations", the framework enables slow-thinking reasoning, mirroring the logic of mathematical proofs while grounding solutions in real-world chemical constraints. We evaluate models on two high-impact tasks: Molecular Property Optimization and Chemical Reaction Prediction. These tasks mirror real-world challenges while providing structured evaluability. We further provide ChemCoTDataset, a pioneering 22,000-instance chemical reasoning dataset with expert-annotated chains of thought to facilitate LLM fine-tuning. By providing annotated trainable datasets, a reasoning taxonomy, and baseline evaluations, our work bridges the gap between abstract reasoning methods and practical chemical discovery, establishing a foundation for advancing LLMs as tools for AI-driven scientific innovation. Hao Li 0073, He Cao, Bin Feng 0001, Daniel Shao, Robert Tang, Zhiyuan Yan 0002, Yonghong Tian 0001, Li Yuan 0007, Yu Li 0003 |
NeurIPS | 3 |
| 2025 | Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting
Zhengqi Zhao, Xiaohu Huang, Hao Zhou 0039, Errui Ding, Jingdong Wang 0001, Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001 |
Int. J. Comput. Vis. | 9 |
| 2025 | MambaGait: Gait recognition approach combining explicit representation and implicit state space model
Haijun Xiong, Bin Feng 0001, Bang Wang 0001, Xinggang Wang, Wenyu Liu 0001 |
Image Vis. Comput. | 2 |
| 2024 | Causality-Inspired Discriminative Feature Learning in Triple Domains for Gait Recognition
Haijun Xiong, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001 |
ECCV (54) | 2 |
| 2024 | Licaf: Lidar-Camera Asymmetric Fusion For Gait RecognitionabstractGait recognition is a biometric technology that identifies individuals by using walking patterns. Due to the significant achievements of multimodal fusion in gait recognition, we consider employing LiDAR-camera fusion to obtain robust gait representations. However, existing methods often overlook intrinsic characteristics of modalities, and lack fine-grained fusion and temporal modeling. In this paper, we introduce a novel modality-sensitive network LiCAF for LiDAR-camera fusion, which employs an asymmetric modeling strategy. Specifically, we propose Asymmetric Cross-modal Channel Attention (ACCA) and Interlaced Crossmodal Temporal Modeling (ICTM) for cross-modal valuable channel information selection and powerful temporal modeling. Our method achieves state-of-the-art performance ($93.9 \%$ in Rank-1 and $98.8 \%$ in Rank-5) on the SUSTech1K dataset, demonstrating its effectiveness. Yunze Deng, Haijun Xiong, Bin Feng 0001 |
ICIP | 3 |
| 2024 | Gaitgs: Temporal Feature Learning in Granularity And Span Dimension for Gait RecognitionabstractGait recognition, a growing field in biological recognition technology, utilizes distinct walking patterns for accurate individual identification. However, existing methods lack the incorporation of temporal information. To reach the full potential of gait recognition, we advocate for the consideration of temporal features at varying granularities and spans. This paper introduces a novel framework, GaitGS, which aggregates temporal features simultaneously in both granularity and span dimensions. Specifically, the Multi-Granularity Feature Extractor (MGFE) is designed to capture micro-motion and macro-motion information at fine and coarse levels respectively, while the Multi-Span Feature Extractor (MSFE) generates local and global temporal representations. Through extensive experiments on two datasets, our method demonstrates state-of-the-art performance, achieving Rank-1 accuracy of 98.2%, 96.5%, and 89.7% on CASIA-B under different conditions, and 97.6% on OU-MVLP. The source code will be available at https://github.com/Haijun-Xiong/GaitGS. Haijun Xiong, Yunze Deng, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001 |
ICIP | 3 |
| 2024 | SubgDiff: A Subgraph Diffusion Model to Improve Molecular Representation LearningabstractMolecular representation learning has shown great success in advancing AI-based drug discovery. A key insight of many recent works is that the 3D geometric structure of molecules provides essential information about their physicochemical properties. Recently, denoising diffusion probabilistic models have achieved impressive performance in molecular 3D conformation generation. However, most existing molecular diffusion models treat each atom as an independent entity, overlooking the dependency among atoms within the substructures. This paper introduces a novel approach that enhances molecular representation learning by incorporating substructural information in the diffusion model framework. We propose a novel diffusion model termed SubgDiff for involving the molecular subgraph information in diffusion. Specifically, SubgDiff adopts three vital techniques: i) subgraph prediction, ii) expectation state, and iii) k-step same subgraph diffusion, to enhance the perception of molecular substructure in the denoising network. Experiments on extensive downstream tasks, especially the molecular force predictions, demonstrate the superior performance of our approach. Jiying Zhang, Zijing Liu, Yu Wang 0027, Bin Feng 0001, Yu Li 0003 |
NeurIPS | 4 |
| 2023 | Graph Contrastive Learning for Skeleton-based Action Recognition
Xiaohu Huang, Hao Zhou 0039, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001, Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001 |
ICLR | 10 |
| 2023 | Condition-Adaptive Graph Convolution Learning for Skeleton-Based Gait RecognitionabstractGraph convolutional networks have been widely applied in skeleton-based gait recognition. A key challenge in this task is to distinguish the individual walking styles of different subjects across various views. Existing state-of-the-art methods employ uniform convolutions to extract features from diverse sequences and ignore the effects of viewpoint changes. To overcome these limitations, we propose a condition-adaptive graph (CAG) convolution network that can dynamically adapt to the specific attributes of each skeleton sequence and the corresponding view angle. In contrast to using fixed weights for all joints and sequences, we introduce a joint-specific filter learning (JSFL) module in the CAG method, which produces sequence-adaptive filters at the joint level. The adaptive filters capture fine-grained patterns that are unique to each joint, enabling the extraction of diverse spatial-temporal information about body parts. Additionally, we design a view-adaptive topology learning (VATL) module that generates adaptive graph topologies. These graph topologies are used to correlate the joints adaptively according to the specific view conditions. Thus, CAG can simultaneously adjust to various walking styles and viewpoints. Experiments on the two most widely used datasets (i.e., CASIA-B and OU-MVLP) show that CAG surpasses all previous skeleton-based methods. Moreover, the recognition performance can be enhanced by simply combining CAG with appearance-based methods, demonstrating the ability of CAG to provide useful complementary information. Xiaohu Huang, Xinggang Wang, Zhidianqiu Jin, Botao He, Bin Feng 0001, Wenyu Liu 0001 |
IEEE Trans. Image Process. | 6 |
| 2021 | Instances as QueriesabstractWe present QueryInst, a new perspective for instance segmentation. QueryInst is a multi-stage end-to-end system that treats instances of interest as learnable queries, enabling query based object detectors, e.g., Sparse RCNN, to have strong instance segmentation performance. The attributes of instances such as categories, bounding boxes, instance masks, and instance association embeddings are represented by queries in a unified manner. In QueryInst, a query is shared by both detection and segmentation via dynamic convolutions and driven by parallellysupervised multi-stage learning. We conduct extensive experiments on three challenging benchmarks, i.e., COCO, CityScapes, and YouTube-VIS to evaluate the effectiveness of QueryInst in object detection, instance segmentation, and video instance segmentation tasks. For the first time, we demonstrate that a simple end-to-end query based framework can achieve the state-of-the-art performance in various instance-level recognition tasks. Code is available at https://github.com/hustvl/QueryInst. Shusheng Yang, Xinggang Wang, Yu Li 0003, Ying Shan, Bin Feng 0001, Wenyu Liu 0001 |
ICCV | 7 |
| 2021 | Context-Sensitive Temporal Feature Learning for Gait RecognitionabstractAlthough gait recognition has drawn increasing research attention recently, it remains challenging to learn discriminative temporal representation since the silhouette differences are quite subtle in spatial domain. Inspired by the observation that humans can distinguish gaits of different subjects by adaptively focusing on temporal sequences with different time scales, we propose a context-sensitive temporal feature learning (CSTL) network in this paper, which aggregates temporal features in three scales to obtain motion representation according to the temporal contextual information. Specifically, CSTL introduces relation modeling among multi-scale features to evaluate feature importances, based on which network adaptively enhances more important scale and suppresses less important scale. Besides that, we propose a salient spatial feature learning (SSFL) module to tackle the misalignment problem caused by temporal operation, e.g., temporal convolution. SSFL recombines a frame of salient spatial features by extracting the most discriminative parts across the whole sequence. In this way, we achieve adaptive temporal learning and salient spatial mining simultaneously. Extensive experiments conducted on two datasets demonstrate the state-of-the-art performance. On CASIA-B dataset, we achieve rank-1 accuracies of 98.0%, 95.4% and 87.0% under normal walking, bag-carrying and coat-wearing conditions. On OU-MVLP dataset, we achieve rank-1 accuracy of 90.2%. The source code will be published at https://github.com/OliverHxh/CSTL. Xiaohu Huang, Duowang Zhu, Hao Wang 0207, Xinggang Wang, Botao He, Wenyu Liu 0001, Bin Feng 0001 |
ICCV | 8 |
| 2021 | Crossover Learning for Fast Online Video Instance SegmentationabstractModeling temporal visual context across frames is critical for video instance segmentation (VIS) and other video understanding tasks. In this paper, we propose a fast on-line VIS model termed CrossVIS. For temporal information modeling in VIS, we present a novel crossover learning scheme that uses the instance feature in the current frame to pixel-wisely localize the same instance in other frames. Different from previous schemes, crossover learning does not require any additional network parameters for feature enhancement. By integrating with the instance segmentation loss, crossover learning enables efficient cross-frame instance-to-pixel relation learning and brings cost-free improvement during inference. Besides, a global balanced instance embedding branch is proposed for better and more stable online instance association. We conduct extensive experiments on three challenging VIS benchmarks, i.e., YouTube-VIS-2019, OVIS, and YouTube-VIS-2021 to evaluate our methods. CrossVIS achieves state-of-the-art online VIS performance and shows a decent trade-off between latency and accuracy. Code is available at https://github.com/hustvl/CrossVIS. Shusheng Yang, Xinggang Wang, Yu Li 0003, Ying Shan, Bin Feng 0001, Wenyu Liu 0001 |
ICCV | 7 |
| 2020 | Maximum Entropy Regularization and Chinese Text Recognition
Changxu Cheng, Wuheng Xu, Xiang Bai, Bin Feng 0001, Wenyu Liu 0001 |
DAS | 4 |
| 2020 | Deep multi-metric learning for text-independent speaker verification
Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001 |
Neurocomputing | 3 |
| 2020 | Pose Anchor: A Single-Stage Hand Keypoint Detection NetworkabstractThis paper presents an effective single network for hand keypoint detection, instead of relying on the frequently-used two-stage pipeline consisting of localizing the hand and detecting the key points. Our method trains a fully convolutional neural network in an end-to-end manner, based on a novelly proposed pose anchor network, which can be deemed as an extension of the region proposal network (RPN) in Faster Region-based convolutional network. Moreover, we generate our pose anchor in a data-driven way, i.e., a K-means cluster algorithm based on object keypoint similarity (OKS), instead of manually design. In this way, we can obtain multiple representative pose anchors with various gestures, angles, and scales. By introducing the pose anchor, we are capable of utilizing the prior knowledge of the hand structure, mitigating the problem of occlusion to some extent. We demonstrate the feasibility and effectiveness of our method with extensive experiments on the challenging large-scale multiview 3D hand pose dataset (LSM-HPD) and New Zealand Sign Language Dataset (NZSL). Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Cascaded Boundary Network for High-Quality Temporal Action Proposal GenerationabstractCreating high-quality temporal action proposals is fundamental yet challenging for accurate action detection in untrimmed videos due to the complexity of the background and variation in actions' durations and magnitudes. In this paper, we propose a cascaded boundary network (CBN) to predict the action boundaries by considering the importance of precise boundary information to develop accurate action proposals. Specifically, the first stage of CBN locates the temporal boundaries by predicting the probability that each frame corresponds to an action, the start position, and the end position. A temporal convolutional network is used in this stage to capture short-term context information. Next, the predicted probabilities are forwarded to the second stage, in which a long short-term memory (LSTM) network is utilized for further refinement by exploiting the correlation between the predicted probabilities to capture long-term context information. Finally, we combine the results from both stages to produce a long- and short-term information fusion. The experiments on THUMOS14 and ActivityNet-1.3 show that CBN achieves state-of-the-art recall performance. The performance improvement is especially remarkable for a small average number (AN) of retrieved proposals; e.g., the average recall at AN=50 on THUMOS14 is improved from 37.46% to 43.06%. Further experiments are performed by introducing proposals generated by CBN into an existing action detection framework. CBN also achieves state-of-the-art average mAP@tIoU on the THUMOS14 detection benchmark. Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Patch Aggregator for Scene Text Script IdentificationabstractScript identification in the wild is of great importance in a multi-lingual robust-reading system. The scripts deriving from the same language family share a large set of characters, which makes script identification a fine-grained classification problem. Most existing methods make efforts to learn a single representation that combines the local features by making a weighted average or other clustering methods, which may reduce the discriminatory power of some important parts in each script for the interference of redundant features. In this paper, we present a novel module named Patch Aggregator (PA), which learns a more discriminative representation for script identification by taking into account the prediction scores of local patches. Specifically, we design a CNN-based method consisting of a standard CNN classifier and a PA module. Experiments demonstrate that the proposed PA module brings significant performance improvements over the baseline CNN model, achieving the state-of-the-art results on three benchmark datasets for script identification: SIW-13, CVSI 2015 and RRC-MLT 2017. Changxu Cheng, Qiuhui Huang, Xiang Bai, Bin Feng 0001, Wenyu Liu 0001 |
ICDAR | 4 |
| 2018 | Structured random forest for label distribution learning
Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001 |
Neurocomputing | 3 |
| 2018 | Deep attention network for joint hand gesture localization and recognition using static RGB-D images
Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001 |
Inf. Sci. | 4 |
| 2017 | Learning extremely shared middle-level image representation for scene classification
Peng Tang 0005, Xinggang Wang, Bin Feng 0001, Fabio Roli, Wenyu Liu 0001 |
Knowl. Inf. Syst. | 4 |
| 2017 | Depth-Projection-Map-Based Bag of Contour Fragments for Robust Hand Gesture RecognitionabstractThis paper presents a novel and robust descriptor, depth-projection-map-based bag of contour fragments, which is applied to extraction of hand shape and structure information from depth maps. Our method projects depth maps onto three orthogonal planes to generate the depth projection maps. Then, the bag of contour fragment descriptors are extracted from the three depth projection maps and concatenated as a final shape representation of the original depth data. A support vector machine with a linear kernel is used as a shape classifier. The proposed description method is evaluated on three public datasets, as well as a new and more challenging dataset for hand gesture recognition. Results demonstrate that the proposed method significantly outperforms the previous methods on all tested datasets for both static digit recognition and letter gesture recognition. For the challenging HUST-ASL dataset, in particular, the proposed method improves on the previous state-of-the-art methods from 40.1% to 64.6%. Bin Feng 0001, Fangzi He, Xinggang Wang, Yongjiang Wu, Hao Wang 0207, Sihua Yi, Wenyu Liu 0001 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2017 | Learning Multi-Instance Deep Discriminative Patterns for Image ClassificationabstractFinding an effective and efficient representation is very important for image classification. The most common approach is to extract a set of local descriptors, and then aggregate them into a high-dimensional, more semantic feature vector, like unsupervised bag-of-features and weakly supervised part-based models. The latter one is usually more discriminative than the former due to the use of information from image labels. In this paper, we propose a weakly supervised strategy that using multi-instance learning (MIL) to learn discriminative patterns for image representation. Specially, we extend traditional multi-instance methods to explicitly learn more than one patterns in positive class, and find the "most positive" instance for each pattern. Furthermore, as the positiveness of instance is treated as a continuous variable, we can use stochastic gradient decent to maximize the margin between different patterns meanwhile considering MIL constraints. To make the learned patterns more discriminative, local descriptors extracted by deep convolutional neural networks are chosen instead of hand-crafted descriptors. Some experimental results are reported on several widely used benchmarks (Action 40, Caltech 101, Scene 15, MIT-indoor, SUN 397), showing that our method can achieve very remarkable performance. Peng Tang 0005, Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Representing conditional preference by boosted regression trees for recommendation
Chenhong Sui, Dewei Deng, Bin Feng 0001, Wenyu Liu 0001, Caihua Wu |
Inf. Sci. | 5 |
| 2016 | Online similarity learning for visual tracking
Sihua Yi, Nan Jiang 0016, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001 |
Inf. Sci. | 3 |
| 2015 | Conditional preference in recommender systems
Wenyu Liu 0001, Caihua Wu, Bin Feng 0001 |
Expert Syst. Appl. | 3 |
| 2014 | Bag of contour fragments for robust shape classification
Xinggang Wang, Bin Feng 0001, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki |
Pattern Recognit. | 2 |
| 2006 | Fast adaptive inter mode decision method for P slices in H.264abstractH.264 defines 7 different coding modes for macroblocks (MBs) in P slices. In order to achieve a coding performance as high as possible, the H.264 encoder calculates rate distortion costs of all possible modes to determine the best mode of a MB. The computation complexity is so large that make it difficult to be used in practical applications especially in real time environment. In this paper, a fast adaptive inter modes decision method is proposed to reduce the complexity of H.264 encoder. The candidate inter modes used in rate distortion optimization can be limited in a small mode group (MG) by using the characteristics of the motion compensated residual image. The overlapped mode groups and dynamic adjusted thresholds adopted in the proposed method can make the best mode lies within the chosen MG with great possibility which leads to extensively computation reduction without any perceivable loss in quality. The experimental results show that the proposed method can save the encoding time up to 52% on average with -0.05dB performance degradation that is negligible. Keywords-H.264; inter modes; mode group; adaptive; overlapped Bin Feng 0001, Guangxi Zhu, Wenyu Liu 0001 |
CCNC | 1 |
| 2006 | Fast adaptive inter-prediction mode decision method for H.264 based on spatial correlationabstractH.264 defines 7 different coding modes for macroblocks (MBs) in P slices. In order to achieve a coding performance as high as possible, the H.264 encoder calculates rate distortion costs of all possible modes to determine the best mode of a MB. The computation complexity is so large that make it difficult to be used in practical applications especially in real time environment. In this paper, a fast adaptive intermodes decision method is proposed to reduce the complexity of H.264 encoder. Firstly the candidate inter modes used in rate distortion optimization can be limited in a small mode group (MG) by using the characteristics of the motion compensated residual image. Then the two most probable modes of the chosen MG are obtained on the basis of the modes of the up MB and the left MB. By calculating and comparing the rate distortion cost of the two modes, the optimum mode of the MB is determined. The overlapped mode groups and dynamic adjusted thresholds adopted in the proposed method can make the best mode lie within the chosen MG with great possibility which leads to extensive computation reduction with acceptable loss in quality. The experimental results show that the proposed method can save the encoding time up to 65% on average with -0.24dB performance degradation. Bin Feng 0001, Guangxi Zhu, Wenyu Liu 0001 |
ISCAS | 1 |
| 2005 | Fast Adaptive Inter Mode Decision Method in H.264 Based on Spatial CorrelationabstractThe computation complexity of H.264 is so large that make it difficult to be used in practical applications especially in real time environment. In this paper, a fast adaptive inter modes decision method is proposed to reduce the complexity of H.264 encoder. Firstly the candidate inter modes can be limited in a small mode group (MG) by using the characteristics of the motion compensated residual image. Then the most probable mode (MPM) of the MB is predicted on the basis of the modes of the neighboring macroblocks. The overlapped mode groups and dynamic adjusted thresholds adopted in the proposed method can make the best mode lies within the chosen MG with great possibility which leads to extensively computation reduction with acceptable loss in quality. The experimental results show that the proposed method can save the encoding time up to 64% on average with -0.45dB performance degradation. Bin Feng 0001, Guangxi Zhu, Wenyu Liu 0001 |
ISM | 1 |