VLDB 2026 Research / reviewers in the wild / expert
Jianfeng Ren
dblp:67/4326
· DBLP profile ↗
80ranked-venue papers
19as first author
58since 2021 · last 2026
0000-0003-4619-6590ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 13 first-author · 34 since 2021Artificial intelligence and machine learning · 33 · 6 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning?abstractMultimodal Large Language Models (MLLMs) largely lag human-level performance on abstract visual reasoning (AVR), which requires models to infer latent rules from visual question sets and generalize them to novel scenarios. Most AVR benchmarks are constrained to narrow and repetitive 2D patterns, involving relatively simple spatial relationships and assessing limited dimensions of reasoning ability. Drawing inspiration from real-world paper folding challenges, we propose Paper Folding Puzzles (PFP), a rigorously designed benchmark specifically developed to assess spatial reasoning capabilities. It comprises 150K visual question-answering samples across five diverse tasks, ranging from basic 2D geometric reasoning to 3D spatial understanding. The developed benchmark dataset can be employed to assess core spatial reasoning abilities essential to human cognition, encompassing fundamental symmetry reasoning and 3D spatial comprehension. Furthermore, we conduct a comprehensive evaluation of 18 leading MLLMs (both closed- and open-source variants) on the PFP benchmark to assess their spatial reasoning capabilities. Our findings show that most MLLMs achieve near-chance performance on FPF, exhibiting substantial performance gaps (>30%) relative to human baselines across all tasks. This highlights a critical research gap in improving spatial reasoning capabilities of MLLMs. Dibin Zhou, Yantao Xu, Zongming Huang, Zengwei Yan, Yongwei Miao, Jianfeng Ren, Fuchang Liu |
AAAI | 7 |
| 2026 | GDGraph: Geometry-Enhanced Dual-View Graph for Molecular Representation LearningabstractLearning effective molecular representations is crucial for accurate property prediction in AI-aided drug discovery. However, most existing molecular pre-training methods are still primarily based on 2D topological graphs, limiting their ability to exploit 3D geometric information. Moreover, methods that do incorporate 3D geometry often do not distinguish between the roles of atom-centered and bond-centered representations. To address these limitations, we propose GDGraph, a geometryenhanced dual-view framework for molecular representation learning. GDGraph models molecular geometry from two complementary structural perspectives: an atom view for capturing global spatial dependencies and a bond view for modeling local geometric patterns. To support this dual-view design, we introduce a multi-scale geometric feature encoding scheme and a view-specific geometry-aware learning strategy, enabling each view to focus on the geometric dependencies it is best suited to capture. Extensive experiments demonstrate that GDGraph achieves strong and stable performance on molecular property prediction benchmarks, and effectively predicts geometrysensitive quantum chemical properties on the QM9 dataset. Yu Liu 0152, Jonathan D. Hirst, Jianfeng Ren, Bencan Tang, Dave Towey |
COMPSAC | 3 |
| 2026 | M-MambaS: Multimodal Mamba for small lesion segmentation
Gui Wang, Jianfeng Ren, LinLin Shen, Wooi Ping Cheah, Rong Qu |
Pattern Recognit. | 2 |
| 2026 | S3DL: Sample-Aggregated Structured Supervised Dictionary Learning
Haiyan Yu 0003, Yucheng Peng, Jianfeng Ren, LinLin Shen, Xin Chen 0003, Ruibin Bai |
IEEE Signal Process. Lett. | 3 |
| 2026 | LP2DH: A Locality-Preserving Pixel-Difference Hashing Framework for Dynamic Texture RecognitionabstractSpatiotemporal Local Binary Pattern (STLBP) is a widely used dynamic texture descriptor, but it suffers from extremely high dimensionality. To tackle this, STLBP features are often extracted on three orthogonal planes, which sacrifice inter-plane correlation. In this work, we propose a Locality-Preserving Pixel-Difference Hashing (LP2DH) framework that jointly encodes pixel differences in the full spatiotemporal neighborhood. LP2DH transforms Pixel-Difference Vectors (PDVs) into compact binary codes with maximal discriminative power. Furthermore, we incorporate a locality-preserving embedding to maintain the PDVs' local structure before and after hashing. Then, a curvilinear search strategy is utilized to jointly optimize the hashing matrix and binary codes via gradient descent on the Stiefel manifold. After hashing, dictionary learning is applied to encode the binary vectors into codewords, and the resulting histogram is utilized as the final feature representation. The proposed LP2DH achieves state-of-the-art performance on three major dynamic texture recognition benchmarks: 99.80% against DT-GoogleNet's 98.93% on UCLA, 98.52% against HoGF3D's 97.63% on DynTex++, and 96.19% compared to STS's 95.00% on YUPENN. The source code is available at: https://github.com/drx770/LP2DH. Ruxin Ding, Jianfeng Ren, Heng Yu 0001, Jiawei Li 0001, Xudong Jiang 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Predictive Reasoning With Augmented Anomaly Contrastive Learning for Compositional Visual RelationsabstractWhile visual reasoning for simple analogies has received significant attention, compositional visual relations (CVR) remain relatively unexplored due to their greater complexity. To solve CVR tasks, we propose Predictive Reasoning with Augmented Anomaly Contrastive Learning (PR-A$^{2}$CL), i.e., to identify an outlier image given three other images that follow the same compositional rules. To address the challenge of modelling abundant compositional rules, an Augmented Anomaly Contrastive Learning is designed to distil discriminative and generalizable features by maximizing similarity among normal instances while minimizing similarity between normal and anomalous outliers. More importantly, a predict-and-verify paradigm is introduced for rule-based reasoning, in which a series of Predictive Anomaly Reasoning Blocks (PARBs) iteratively leverage features from three out of the four images to predict those of the remaining one. Throughout the subsequent verification stage, the PARBs progressively pinpoint the specific discrepancies attributable to the underlying rules. Experimental results on SVRT, CVR and MC$^{2}$R datasets show that PR-A$^{2}$CL significantly outperforms state-of-the-art reasoning models. Chengtai Li, Yuting He 0002, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Heng Yu 0001, Xudong Jiang 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | Ranking-Based Self-Supervised Representation Learning for Skeleton-Based Action RecognitionabstractRecently, researchers have achieved significant results in the skeleton-based action recognition. To better model the skeleton sequences, we drive the encoder to learn more discriminative representations in the self-supervised setting. We find that instead of clustering feature vectors to assign pseudo labels for samples as in DeepCluster, ranking them is a more reasonable, reliable, and efficient way to learn more effective feature representations. With this intuition, we propose a novel self-supervised learning framework,DeepRank. Specifically, we rank triplets of skeleton sequences with the ranking labels, obtained from the relative distances among them. Besides, to deeply mine complementary discriminative information that exists in different modalities of skeleton sequences, we further proposeMulti-ViewDeepRank(MV-DeepRank) to enable encoders to comprehensively learn complementary features from multiple modalities. Extensive experimental results on the NTU RGB+D, NTU RGB+D 120, PKU-MMD I, and PKU-MMD II datasets under various evaluation settings demonstrate the generality, transferability, and superiority of our proposed self-supervised learning frameworks. Notably, our frameworks surpass the previous methods that employ the same backbone networks as ours by at least 1.8% (ST-GCN) and 2.1% (STTFormer) under the finetuning setting. Additionally, DeepRank gains a significant advantage on computational complexities,$O(1)$, over the contrastive learning-based methods,$O(\rm{batch size})$, and the clustering-based methods,$O(\rm{number of clusters})$. Bizhu Wu, Junliang Chen 0002, Jinheng Xie, Qiufu Li, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
IEEE Trans. Multim. | 5 |
| 2025 | DARR: A Dual-Branch Arithmetic Regression Reasoning Framework for Solving Machine Number ReasoningabstractAbstract visual reasoning (AVR) is a critical ability of humans, and it has been widely studied, but arithmetic visual reasoning, a unique task in AVR to reason over number sense, is less studied in the literature. To facilitate this research, we construct a Machine Number Reasoning (MNR) dataset to assess the model's ability in arithmetic visual reasoning over number sense and spatial layouts. To solve the MNR tasks, we propose a Dual-branch Arithmetic Regression Reasoning (DARR) framework, which includes an Intra-Image Arithmetic Regression Reasoning (IIARR) module and a Cross-Image Arithmetic Regression Reasoning (CIARR) module. The IIARR includes a set of Intra-Image Regression Blocks to identify the correct number orders and the underlying arithmetic rules within individual images, and an Order Gate to determine the correct number order. The CIARR establishes the arithmetic relations across different images through a `3-to-1' regressor and a set of `2-to-1' regressors, with a Selection Gate to select the most suitable `2-to-1' regressor and a gated fusion to combine the two kinds of regressors. Experiments on the MNR dataset show that the DARR outperforms state-of-the-art models for arithmetic visual reasoning. Chengtai Li, Yee Yang Tan, Yuting He 0002, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Heng Yu 0001, Xudong Jiang 0001 |
AAAI | 4 |
| 2025 | ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded GapsabstractSolving jigsaw puzzles has been extensively studied. While most existing models focus on solving either small-scale puzzles or puzzles with no gap between fragments, solving large-scale puzzles with gaps presents distinctive challenges in both image understanding and combinatorial optimization. To tackle these challenges, we propose a framework of Evolutionary Reinforcement Learning with Multi-head Puzzle Perception (ERL-MPP) to derive a better set of swapping actions for solving the puzzles. Specifically, to tackle the challenges of perceiving the puzzle with gaps, a Multi-head Puzzle Perception Network (MPPN) with a shared encoder is designed, where multiple puzzlet heads comprehensively perceive the local assembly status, and a discriminator head provides a global assessment of the puzzle. To explore the large swapping action space efficiently, an Evolutionary Reinforcement Learning (EvoRL) agent is designed, where an actor recommends a set of suitable swapping actions from a large action space based on the perceived puzzle status, a critic updates the actor using the estimated rewards and the puzzle status, and an evaluator coupled with evolutionary strategies evolves the actions aligning with the historical assembly experience. The proposed ERL-MPP is comprehensively evaluated on the JPLEG-5 dataset with large gaps and the MIT dataset with large-scale puzzles. It significantly outperforms all state-of-the-art models on both datasets. Xingke Song, Chenglin Yao, Jianfeng Ren, Ruibin Bai, Xin Chen 0003, Xudong Jiang 0001 |
AAAI | 4 |
| 2025 | S³-Mamba: Small-Size-Sensitive Mamba for Lesion SegmentationabstractSmall lesions play a critical role in early disease diagnosis and intervention of severe infections. Popular models often face challenges in segmenting small lesions, as it occupies only a minor portion of an image, while down-sampling operations may inevitably lose focus on local features of small lesions. To tackle the challenges, we propose a Small-Size-Sensitive Mamba (S³-Mamba), which promotes the sensitivity to small lesions across three dimensions: channel, spatial, and training strategy. Specifically, an Enhanced Visual State Space block is designed to focus on small lesions through multiple residual connections to preserve local features, and selectively amplify important details while suppressing irrelevant ones through channel-wise attention. A Tensor-based Cross-feature Multi-scale Attention is designed to integrate input image features and intermediate-layer features with edge features and exploit the attentive support of features across multiple scales, thereby retaining spatial details of small lesions at various granularities. Finally, we introduce a novel regularized curriculum learning to automatically assess lesion size and sample difficulty, and gradually focus from easy samples to hard ones like small lesions. Extensive experiments on three medical image segmentation datasets show the superiority of our S³-Mamba, especially in segmenting small lesions. Gui Wang, Yuexiang Li, Wenting Chen, Meidan Ding, Wooi Ping Cheah, Rong Qu, Jianfeng Ren, LinLin Shen |
AAAI | 7 |
| 2025 | Equivariant Deformable Convolutions for Unrolling Networks in Cardiac Cine MR ImagingabstractDeep unrolling methods have achieved notable success in accelerated cardiac cine MRI reconstruction. However, their effectiveness remains limited by constrained receptive fields and rigid convolutional sampling, which hinder scalability in high-resolution reconstruction tasks. To address these challenges, we propose an Equivariant Deformable Convolutional Unrolling Network (EDCU-Net) that integrates a Spatiotemporal Deformable Module (STDM) and a Rotation Equivariant Module (REM). EDCU-Net effectively enlarges the receptive field while maintaining low computational cost and high parameter efficiency, enabling more effective suppression of large-scale aliasing artifacts. Specifically, STDM adaptively adjusts sampling locations to better capture complex anatomical structures and dynamic cardiac motion, while reducing interpolation overhead. Furthermore, REM embeds deformable convolutions within an equivariant framework, allowing shared parameters across orientations and further promoting parameter efficiency and generalization capability. Extensive experiments on cardiac cine MRI data demonstrate that EDCU-Net consistently outperforms state-of-the-art methods in both reconstruction accuracy and visual quality. Yuliang Zhu, Zhuo-Xu Cui, Qingyong Zhu, Zhaochi Wen, Jianfeng Ren, Dong Liang 0001 |
BIBM | 7 |
| 2025 | Enhancing Cryptocurrency Trading Strategies: A Deep Reinforcement Learning Approach Integrating Multi-Source LLM Sentiment AnalysisabstractRecent advancements in large language models (LLMs) have demonstrated their potential to significantly impact finance trading, particularly through sentiment analysis. The cryptocurrency market, known for its volatility and unpredictability, often renders price-based trading approaches inadequate. This necessitates the adoption of more sophisticated techniques such as market sentiment analysis, which can benefit from the insights provided by LLMs. This study introduces an innovative method that integrates sentiment analysis derived from five distinct LLMs with deep reinforcement learning to devise a cryptocurrency trading strategy. Recognizing that LLM outputs cannot be guaranteed to be infallibly accurate, which contributing to the LLM hallucinations, this paper details the implementation of a stringent outlier detection and removal process. By adopting a “Trust-The-Majority” strategy, the research aims to ensure that trading decisions are informed by reliable sentiment data. In addition, sentiment scores are traditionally timestamped to the publication of news or social media posts. To more accurately reflect the actual impact of such information on market sentiment, this study applies the Ebbinghaus Forgetting Curve to model the waning influence of information over time. This allows for a more nuanced understanding of how news affects market dynamics. The enhanced sentiment scores, in conjunction with traditional market data such as OHLCV (Open, High, Low, Close, Volume), are utilized by a deep reinforcement learning model to make trading decisions. Experimental results demonstrate that the proposed multi-LLM sentiment-driven framework improves trading performance in the fast-paced cryptocurrency market. The methodology outlined in this paper offers a solid foundation for incorporating real-time market sentiment analysis into financial applications. Nanjiang Du, Yida Zhao, Yicheng Zhu, Siyu Xie, Luyao Yang, Yiru Tong, Shengzhe Xu, Wangying Zhang, Zecheng Tang, Jianfeng Ren, Tianxiang Cui |
CIFEr | 12 |
| 2025 | Chemically-aware Attention-based Multi-modal Fusion Framework for Molecular Representation LearningabstractLearning effective molecular representations is crucial for accurate property prediction in artificial intelligence (AI)-aided drug discovery. Graph and fingerprint representations have been widely used to encode molecular topological structures and chemical substructures. To enhance the feature embedding of each modality and leverage their complementary strengths, we propose a novel Chemically-aware Attention-based Multi-modal Fusion Framework (CAMFF) for molecular representation learning, which integrates molecular graphs and extended-connectivity fingerprints by exploiting various attention mechanisms. Specifically, the proposed CAMFF consists of three modules: 1) a graph embedding module incorporating multi-head attention to capture local heterogeneous interactions and all-pair self-attention to capture long-range atomic dependencies from molecular graph representations; 2) a fingerprint embedding module using a pre-trained Mol2Vec model to generate dense chemical substructure representations; and 3) a chemically-aware feature interaction and fusion module incorporating self-attention to enable interactions between various chemical substructures and cross-attention to ensure effective multi-modal alignment and fusion. To evaluate the effectiveness of CAMFF, we compare it with 14 state-of-the-art methods across 9 molecular property prediction benchmarks. CAMFF demonstrates competitive predictive performance and improves interpretability through attention-based visualization, showing its potential for real-world drug discovery. Yu Liu 0152, Jonathan D. Hirst, Jianfeng Ren, Bencan Tang, Dave Towey |
COMPSAC | 3 |
| 2025 | MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple GranularitiesabstractRecent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM Bizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
CVPR | 5 |
| 2025 | De2r: Unifying DVFS and Early-Exit for Embedded AI Inference via Reinforcement LearningabstractExecuting neural networks on resource-constrained embedded devices faces challenges. Efforts have been made at the application and system levels to reduce the execution cost. Among them, the early-exit networks reduce computational cost through intermediate exits, while Dynamic Voltage and Frequency Scaling (DVFS) offers system energy reduction. Existing works strive to unify early-exit and DVFS for combined benefits on both timing and energy flexibility, yet limitations exist: 1) varying time constraints that make different exit points become more, or less, important in terms of inference accuracy, are not taken care of, and 2) the optimal decisions of unifying DVFS and early-exit as a multi-objective optimization problem are not achieved due to the large configuration space. To address these challenges, we propose Dr2r, a reinforcement learning-based framework that jointly optimizes early-exit points and DVFS settings for continuous inference. In particular, Dr2r includes a cross-training mechanism that fine-tunes the early-exit network to accommodate dynamic time constraints and system conditions. Experimental results demonstrate that Dr2r achieves up to 22.03% energy reduction and 3.23% accuracy gain compared to contemporary techniques. Yuting He 0002, Jingjin Li, Chengtai Li, Qingyu Yang 0004, Zheng Wang 0027, Heshan Du, Jianfeng Ren, Heng Yu 0001 |
DATE | 7 |
| 2025 | DBCR: Exploiting Both Intra-cluster and Extra-cluster Relations for Compositional ReasoningabstractMost existing models for abstract visual reasoning perform poorly in compositional visual reasoning (CVR), due to complex nature of compositional rules and difficulties in distinguishing tiny rule differences between outliers and normal images. To tackle the challenges, we propose a Dual-Branch Compositional Reasoning (DBCR) model, exploiting both intra-cluster relations among the cluster of normal images and extra-cluster relations between normal images and outliers. Specifically, we design one branch of Intra-Cluster Regression Reasoning Blocks (ICR2Bs) to encapsulate common relations among normal images through hierarchical regressing reasoning, and the other branch of Contrastive Attention Reasoning Blocks (CARBs) to exploit extra-cluster differences between normal images and outliers through self-attention. Simultaneously minimizing the regression errors in ICR2Bs and maximizing the extra-cluster differences in CARBs help identify the correct cluster of normal images. Experimental results on two CVR datasets show that the proposed DBCR consistently outperforms state-of-the-art models. The code is available at https://github.com/He1mont/DBCR. Chengtai Li, Guosheng Su, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Xudong Jiang 0001 |
ICASSP | 3 |
| 2025 | Jointly Optimizing Data Discretization and Naive Bayes Classifier via Multi-Objective OptimizationabstractData discretization plays a critical role in enhancing the performance of the naive Bayes classifier. Traditional data discretization methods often utilize a two-stage framework, where data discretization and classification are optimized separately, leading to sub-optimal performance. To tackle the issue, we propose a novel multi-objective optimization framework that incorporates the optimization of the naive Bayes classifier into the objective function of optimizing data discretization. To solve this problem, we employ an alternative optimization method to jointly optimize both data discretization and classification. Additionally, to further enhance the optimization process, we leverage a genetic algorithm to explore and exploit a larger solution space. Experimental results on 20 datasets demonstrate that our method outperforms state-of-the-art methods. Jiacheng Tu, Haiyan Yu 0003, Ruxin Ding, Shihe Wang, Jianfeng Ren, Xudong Jiang 0001 |
ICASSP | 5 |
| 2025 | FineMotion: A Dataset and Benchmark with Both Spatial and Temporal Annotation for Fine-Grained Motion Generation and Editing
Bizhu Wu, Jinheng Xie, Meidan Ding, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
ICCV | 5 |
| 2025 | GCA-SUNet: A Gated Context-Aware Swin-UNet for Exemplar-Free CountingabstractExemplar-Free Counting aims to count objects of interest without intensive annotations of objects or exemplars. To achieve this, we propose a Gated Context-Aware Swin-UNet (GCA-SUNet) to directly map an input image to the density map of countable objects. Specifically, a set of Swin transformers form an encoder to derive a robust feature representation, and a Gated Context-Aware Modulation block is designed to suppress irrelevant objects or background through a gate mechanism and exploit the attentive support of objects of interest through a self-similarity matrix. The gate strategy is also incorporated into the bottleneck network and the decoder of the Swin-UNet to highlight the features most relevant to objects of interest. By explicitly exploiting the attentive support among countable objects and eliminating irrelevant features through the gate mechanisms, the proposed GCA-SUNet focuses on and counts objects of interest without relying on predefined categories or exemplars. Experimental results on the real-world datasets such as FSC-147 and CARPK demonstrate that GCA-SUNet significantly and consistently outperforms state-of-the-art methods. The code is available at https://github.com/Amordia/GCA-SUNet. Yipeng Xu, Jialu Zhang 0003, Jianfeng Ren, Xudong Jiang 0001 |
ICME | 5 |
| 2025 | CEARI: Co-Evolutionary Agents for Reassembling and Inpainting Puzzles with Gaps and Missing PiecesabstractPuzzle solving has recently become a popular research topic. Existing solvers often overlook puzzles with missing pieces. The missing pieces, together with gaps between pieces, pose significant challenges, amplified by a large solution space. To tackle the challenges, we propose Co-Evolutionary Agents for Reassembling and Inpainting (CEARI), one agent to inpaint missing contents and the other to reassemble the puzzle, with a shared perception network to perceive the puzzle status. The reassembly agent utilizes an evolutionary algorithm to explore the large solution space, to discover a sequence of fragment-swapping actions to efficiently reassemble the puzzle, while the inpainting agent evolves from using a local outpainting network at the early stage to using a global inpainting network at the latter stage. Furthermore, a co-evolutionary training paradigm is designed to iteratively evolve the two agents in a coherent and collaborative manner, improving reassembly accuracy and inpainting quality simultaneously. Experimental results on three datasets show that CEARI largely outperforms state-of-the-art methods in terms of both reassembly accuracy and inpainting quality. Xingke Song, Jianxu Shangguan, Yiran Li 0003, Jialu Zhang 0003, Jianfeng Ren, Ruibin Bai, Xin Chen 0003, Xudong Jiang 0001 |
ACM Multimedia | 5 |
| 2025 | DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMsabstractAbstract Visual Reasoning (AVR) entails discerning latent patterns in visual data and inferring underlying rules. Existing solutions often lack scalability and adaptability, as deep architectures tend to overfit training data, and static neural networks fail to dynamically capture diverse rules. To tackle the challenges, we propose a Dynamic and Scalable Reasoning Framework (DSRF) that greatly enhances the reasoning ability by widening the network instead of deepening it, and dynamically adjusting the reasoning network to better fit novel samples instead of a static network. Specifically, we design a Multi-View Reasoning Pyramid (MVRP) to capture complex rules through layered reasoning to focus features at each view on distinct combinations of attributes, widening the reasoning network to cover more attribute combinations analogous to complex reasoning rules. Additionally, we propose a Dynamic Domain-Contrast Prediction (DDCP) block to handle varying task-specific relationships dynamically by introducing a Gram matrix to model feature distributions, and a gate matrix to capture subtle domain differences between context and target features. Extensive experiments on six AVR tasks demonstrate DSRF’s superior performance, achieving state-of-the-art results under various settings. Code is available here: https://github.com/UNNCRoxLi/DSRF. Chengtai Li, Yuting He 0002, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Xudong Jiang 0001 |
NeurIPS | 3 |
| 2025 | A cascaded retrieval-while-reasoning multi-document comprehension framework with incremental attention for medical question answering
Jianfeng Ren, Ruibin Bai, Zheng Lu 0002 |
Expert Syst. Appl. | 2 |
| 2025 | Radar gait recognition using Dual-branch Swin Transformer with Asymmetric Attention FusionabstractVideo-based gait recognition suffers from potential privacy issues and performance degradation due to dim environments, partial occlusions, or camera view changes. Radar has recently become increasingly popular and overcome various challenges presented by vision sensors. To capture tiny differences in radar gait signatures of different people, a dual-branch Swin Transformer is proposed, where one branch captures the time variations of the radar micro-Doppler signature and the other captures the repetitive frequency patterns in the spectrogram. Unlike natural images where objects can be translated, rotated, or scaled, the spatial coordinates of spectrograms and CVDs have unique physical meanings, and there is no affine transformation for radar targets in these synthetic images. The patch splitting mechanism in Vision Transformer makes it ideal to extract discriminant information from patches, and learn the attentive information across patches, as each patch carries some unique physical properties of radar targets. Swin Transformer consists of a set of cascaded Swin blocks to extract semantic features from shallow to deep representations, further improving the classification performance. Lastly, to highlight the branch with larger discriminant power, an Asymmetric Attention Fusion is proposed to optimally fuse the discriminant features from the two branches. To enrich the research on radar gait recognition, a large-scale NTU-RGR dataset is constructed, containing 45,768 radar frames of 98 subjects. The proposed method is evaluated on the NTU-RGR dataset and the MMRGait-1.0 database. It consistently and significantly outperforms all the compared methods on both datasets. The codes are available at: https://github.com/wentaoheunnc/NTU-RGR . • The proposed method could well extract complementary information from both spectrograms and CVDs. • The proposed Swin-T could extract discriminant features with physical meanings. • The proposed asymmetric attention fusion could effectively combine features with known importance. • A large-scale benchmark dataset, NTU-RGR dataset, is developed to advance the radar gait recognition. Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001 |
Pattern Recognit. | 2 |
| 2025 | Two-stage Rule-induction visual reasoning on RPMs with an application to video predictionabstractRaven's Progressive Matrices (RPMs) are frequently used in evaluating human's visual reasoning ability. Researchers have made considerable efforts in developing systems to automatically solve the RPM problem, often through a black-box end-to-end convolutional neural network for both visual recognition and logical reasoning tasks. Based on the intrinsic natures of RPM problem, we propose a Two-stage Rule-Induction Visual Reasoner (TRIVR), which consists of a perception module and a reasoning module, to tackle the challenges of real-world visual recognition and subsequent logical reasoning tasks, respectively. For the reasoning module, we further propose a “2+1” formulation that models human's thinking in solving RPMs and significantly reduces the model complexity. It derives a reasoning rule from each RPM sample, which is not feasible for existing methods. As a result, the proposed reasoning module is capable of yielding a set of reasoning rules modeling human in solving the RPM problems. To validate the proposed method on real-world applications, an RPM-like Video Prediction (RVP) dataset is constructed, where visual reasoning is conducted on RPMs constructed using real-world video frames. Experimental results on various RPM-like datasets demonstrate that the proposed TRIVR achieves a significant and consistent performance gain compared with state-of-the-art models. Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001 |
Pattern Recognit. | 2 |
| 2024 | Scale Optimization Using Evolutionary Reinforcement Learning for Object Detection on Drone ImageryabstractObject detection in aerial imagery presents a significant challenge due to large scale variations among objects. This paper proposes an evolutionary reinforcement learning agent, integrated within a coarse-to-fine object detection framework, to optimize the scale for more effective detection of objects in such images. Specifically, a set of patches potentially containing objects are first generated. A set of rewards measuring the localization accuracy, the accuracy of predicted labels, and the scale consistency among nearby patches are designed in the agent to guide the scale optimization. The proposed scale-consistency reward ensures similar scales for neighboring objects of the same category. Furthermore, a spatial-semantic attention mechanism is designed to exploit the spatial semantic relations between patches. The agent employs the proximal policy optimization strategy in conjunction with the evolutionary strategy, effectively utilizing both the current patch status and historical experience embedded in the agent. The proposed model is compared with state-of-the-art methods on two benchmark datasets for object detection on drone imagery. It significantly outperforms all the compared methods. Code is available at https://github.com/UNNC-CV/EvOD/. Jialu Zhang 0003, Jianfeng Ren, Qian Zhang 0018, Yitian Zhao, Ruibin Bai, Xiangjian He, Jiang Liu 0001 |
AAAI | 4 |
| 2024 | Three-Branch Molecular Representation Learning Framework for Predicting Molecular Properties in Drug DiscoveryabstractGraph Neural Networks (GNNs) have been widely used to model molecules with a graph representation. However, GNNs face inherent challenges in accurately modeling long-range atomic interactions and identifying complex molecular substructures. This research proposes a novel Three-branch Molecular Representation Learning Framework (TMRLF) for predicting molecular properties: it integrates one branch of a GNN that extracts local molecular structural information with two branches of fully connected networks that capture the chemical substructure based on two fingerprints. Specifically, to better capture the long-range interactions, the GNN is designed with an attention mechanism to enhance the atomic interactions. As the Morgan fingerprint effectively captures functional groups of molecules and another well-used molecular fingerprint in the field of drug discovery, the Extended Reduced Graph (ErG) Fingerprint specifically targets molecular features with pharmacological relevance. These two fingerprints are both utilized to complement the chemical information and long-range information processing at the level of key structural features that GNNs lack. The proposed TMRLF extracts a robust feature representation of molecules, crucial for accurately predicting molecular properties and identifying potential drug candidates. Our proposed TMRLF is compared against six state-of-the-art models on eight benchmark datasets. It demonstrates superior capability in predicting molecular properties. Its effectiveness is further highlighted through proof-of-concept validation in identifying potential inhibitors for the Son of Sevenless Homolog 1 (SOSI) protein in real-world drug discovery scenarios. Yu Liu 0152, Lihui Duo, Jonathan D. Hirst, Jianfeng Ren, Bencan Tang, Dave Towey |
COMPSAC | 4 |
| 2024 | Dual-Branch StarNet with Mutual Attention and U-Net Denoising for Simultaneously Recognizing Keywords and Speakers
Yuting He 0002, Chengtai Li, Heng Yu 0001, Jianfeng Ren, Zheng Wang 0027, Heshan Du, Yinshui Xia |
ICONIP (5) | 4 |
| 2024 | Regression Residual Reasoning with Pseudo-labeled Contrastive Learning for Uncovering Multiple Complex Compositional Relations
Chengtai Li, Yuting He 0002, Jianfeng Ren, Ruibin Bai, Yitian Zhao, Heng Yu 0001, Xudong Jiang 0001 |
IJCAI | 3 |
| 2024 | Cardinality and Bounding Constrained Portfolio Optimization Using Safe Reinforcement LearningabstractPortfolio optimization is a strategic approach aiming at achieving an optimal balance between risk and returns through the judicious allocation of limited capital across various assets. In recent years, there has been a growing interest in leveraging Deep Reinforcement Learning (DRL) to tackle the complexities of portfolio optimization. Despite its potential, a notable limitation of DRL algorithms is their inherent difficulty in integrating conflicted objectives with the reward functions throughout the learning process. Typically, DRL's reward function prioritizes the maximization of returns or other performance indicators, often overlooking the integration of risk aspects. Furthermore, the standard DRL framework struggles to incorporate practical constraints, such as cardinality and bounding, into the decision process. Without these constraints, the investment strategies developed might be unrealistic and unmanageable. To this end, in this paper, we propose an adaptive and safe DRL framework, which can dynamically optimize the portfolio weights while strictly respecting practical constraints. In our method, any infeasible action (i.e., one that violates the constraints) decided by the RL agent will be mapped to a feasible region using a safety layer. The extended Markowitz Mean-Variance (M-V) model is explicitly encoded in the safety layer to ensure the feasibility of the actions from the alternative views. In addition, we utilize Projection-based Interior-point Policy Optimization (IPO) to resolve multiple objectives and constraints in the examined problem. Extensive results on real-world datasets show that our method is effective in strictly respecting constraints under dynamic market environments, in contrast to prevailing data- driven trading strategies and conventional model-based static solutions. Yiran Li 0003, Nanjiang Du, Xingke Song, Tianxiang Cui, Ning Xue, Amin Farjudian, Jianfeng Ren, Wooi Ping Cheah |
IJCNN | 8 |
| 2024 | SRE-CNN: A Spatiotemporal Rotation-Equivariant CNN for Cardiac Cine MR Imaging
Yuliang Zhu, Zhuo-Xu Cui, Jianfeng Ren, Dong Liang 0001 |
MICCAI (7) | 4 |
| 2024 | Visual-linguistic Cross-domain Feature Learning with Group Attention and Gamma-correct Gated Fusion for Extracting Commonsense KnowledgeabstractAcquiring commonsense knowledge about entity-pairs from images is crucial across diverse applications. Distantly supervised learning has made significant advancements by automatically retrieving images containing entity pairs and summarizing commonsense knowledge from the bag of images. However, the retrieved images may not always cover all possible relations, and the informative features across the bag of images are often overlooked. To address these challenges, a Multi-modal Cross-domain Feature Learning framework is proposed to incorporate the general domain knowledge from a large vision-text foundation model, ViT-GPT2, to handle unseen relations and exploit complementary information from multiple sources. Then, a Group Attention module is designed to exploit the attentive information from other instances of the same bag to boost the informative features of individual instances. Finally, a Gamma-corrected Gated Fusion is designed to select a subset of informative instances for a comprehensive summarization of commonsense entity relations. Extensive experimental results demonstrate the superiority of the proposed method over state-of-the-art models for extracting commonsense knowledge. Jialu Zhang 0003, Chenglin Yao, Jianfeng Ren, Xudong Jiang 0001 |
ACM Multimedia | 4 |
| 2024 | Hierarchical Perceptual and Predictive Analogy-Inference Network for Abstract Visual ReasoningabstractAdvances in computer vision research enable human-like high-dimensional perceptual induction over analogical visual reasoning problems, such as Raven's Progressive Matrices (RPMs). In this paper, we propose a Hierarchical Perception and Predictive Analogy-Inference network (HP^2AI), consisting of three major components that tackle key challenges of RPM problems. Firstly, in view of the limited receptive fields of shallow networks in most existing RPM solvers, a perceptual encoder is proposed, consisting of a series of hierarchically coupled Patch Attention and Local Context (PALC) blocks, which could capture local attributes at early stages and capture the global panel layout at deep stages. Secondly, most methods seek for object-level similarities to map the context images directly to the answer image, while failing to extract the underlying analogies. The proposed reasoning module, Predictive Analogy-Inference (PredAI), consists of a set of Analogy-Inference Blocks (AIBs) to model and exploit the inherent analogical reasoning rules instead of object similarity. Lastly, the Squeeze-and-Excitation Channel-wise Attention (SECA) in the proposed PredAI discriminates essential attributes and analogies from irrelevant ones. Extensive experiments over four benchmark RPM datasets show that the proposed HP^2AI achieves significant performance gains over all the state-of-the-art methods consistently on all four datasets. Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001 |
ACM Multimedia | 2 |
| 2024 | Seat belt detection using gated Bi-LSTM with part-to-whole attention on diagonally sampled patchesabstractOne of the high-risk behaviors leading to severe traffic injuries is not wearing a seat belt. It is therefore very important to be able to automatically detect seat belts from surveillance images, encourage drivers to wear seat belts, and enhance passenger safety. In this paper, a novel deep neural network, Gated Bi-directional Long Short-Term Memory network with part-to-whole attention (GBL-PA), is proposed for seat belt detection from surveillance images. The innovation of our model lies in its unique diagonal sampling strategy, which meticulously captures the seat belt’s fine details, typically oriented from top right to bottom left across the torso of vehicle occupants. Our framework’s novelty is further encapsulated by the part-to-whole attention mechanism, which intelligently harmonizes the detailed local information from seat belt–specific patches with the broader contextual insights from the regional proposals. The pioneering design of a Gated Bi-directional LSTM network facilitates the dynamic integration of interactions across patches to deliver an optimized final prediction. The superiority of GBL-PA is established through rigorous comparison with the state-of-the-art methods on a new, large benchmark dataset comprising 14,936 images from traffic surveillance footage. Our framework demonstrates a notable improvement, achieving a mean Average Precision (mAP) of 72.3%, which surpasses the second best, YOLOX, by 0.9% mAP. This significant and consistent outperformance across various metrics underscores the transformative potential of GBL-PA in the realm of traffic safety enforcement. The source code of our framework is available at ANONYMISED. Zheng Lu 0002, Jianfeng Ren, Qian Zhang 0018 |
Expert Syst. Appl. | 3 |
| 2024 | Progressively-orthogonally-mapped EfficientNet for action recognition on time-range-Doppler signatureabstractAlthough 2D radar signal representations, such as spectrograms and range-Doppler maps have been widely used for target recognition, 3D time-range-Doppler (TRD) has been less studied, partially because of the difficulties in extracting features from the TRD representation, i.e., shallow 3D neural networks have limited discriminant power, but repeatedly applying 3D convolutions will lead to an oversized 3D network. A hybrid 3D–2D network architecture, Progressively-Orthogonally-Mapped EfficientNet (POMEN), is proposed to address these challenges. More specifically, the proposed POMEN utilizes 3D convolutions in the earlier stages to capture the information embedded in the sparse 3D TRD representation, and to avoid the oversized feature map caused by excessively applying 3D convolutions, we propose to progressively map the 3D features into three sets of 2D features corresponding to the range-time signature, range-Doppler map and time-Doppler signature (spectrogram), respectively. Subsequently, 2D EfficientNet blocks were designed to extract discriminant information from the three sets of 2D feature maps. This hybrid 3D–2D network design effectively extracts features from the 3D TRD representation, thereby avoiding oversized features from full-sized 3D networks and the information loss of 2D networks on 2D representations. Finally, a homogeneous gated fusion network was designed to fuse the three sets of 2D features. The proposed method was evaluated on the UGRS, MIMOGR, and mmWRWD datasets. The experimental results for all datasets demonstrate that the proposed POMEN significantly and consistently outperforms the state-of-the-art models in both 2D and 3D representations. Chenglin Yao, Jianfeng Ren, Ruibin Bai, Heshan Du, Jiang Liu 0001, Xudong Jiang 0001 |
Expert Syst. Appl. | 2 |
| 2024 | A Max-Relevance-Min-Divergence criterion for data discretization with applications on naive BayesabstractIn many classification models, data is discretized to better estimate its distribution. Existing discretization methods often target at maximizing the discriminant power of discretized data, while overlooking the fact that the primary target of data discretization in classification is to improve the generalization performance. As a result, the data tend to be over-split into many small bins since the data without discretization retain the maximal discriminant information. Thus, we propose a Max-Dependency-Min-Divergence (MDmD) criterion that maximizes both the discriminant information and generalization ability of the discretized data. More specifically, the Max-Dependency criterion maximizes the statistical dependency between the discretized data and the classification variable while the Min-Divergence criterion explicitly minimizes the JS-divergence between the training data and the validation data for a given discretization scheme. The proposed MDmD criterion is technically appealing, but it is difficult to reliably estimate the high-order joint distributions of attributes and the classification variable. We hence further propose a more practical solution, Max-Relevance-Min-Divergence (MRmD) discretization scheme, where each attribute is discretized separately, by simultaneously maximizing the discriminant information and the generalization ability of the discretized data. The proposed MRmD is compared with the state-of-the-art discretization algorithms under the naive Bayes classification framework on 45 benchmark datasets. It significantly outperforms all the compared methods on most of the datasets. Shihe Wang, Jianfeng Ren, Ruibin Bai, Yuan Yao 0007, Xudong Jiang 0001 |
Pattern Recognit. | 2 |
| 2024 | GenFace: A Large-Scale Fine-Grained Face Forgery Benchmark and Cross Appearance-Edge LearningabstractThe rapid advancement of photorealistic generators has reached a critical juncture where the discrepancy between authentic and manipulated images is increasingly indistinguishable. Thus, benchmarking and advancing techniques detecting digital manipulation become an urgent issue. Although there have been a number of publicly available face forgery datasets, the forgery faces are mostly generated using GAN-based synthesis technology, which does not involve the most recent technologies like diffusion. The diversity and quality of images generated by diffusion models have been significantly improved and thus a much more challenging face forgery dataset shall be used to evaluate SOTA forgery detection literature. In this paper, we propose a large-scale, diverse, and fine-grained high-fidelity dataset, namely GenFace, to facilitate the advancement of deepfake detection, which contains a large number of forgery faces generated by advanced generators such as the diffusion-based model and more detailed labels about the manipulation approaches and adopted generators. In addition to evaluating SOTA approaches on our benchmark, we design an innovative Cross Appearance-Edge Learning (CAEL) detector to capture multi-grained appearance and edge global representations, and detect discriminative and general forgery traces. Moreover, we devise an Appearance-Edge Cross-Attention (AECA) module to explore the various integrations across two domains. Extensive experiment results and visualizations show that our detection model outperforms the state of the arts on different settings like cross-generator, cross-forgery, and cross-dataset evaluations. Code and datasets will be available athttps://github.com/Jenine-321/GenFace. Zitong Yu, Tianyi Wang 0006, Xiaobin Huang, LinLin Shen, Zan Gao 0001, Jianfeng Ren |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Data augmentation by morphological mixup for solving Raven's progressive matrices
Jianfeng Ren, Ruibin Bai |
Vis. Comput. | 2 |
| 2023 | Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical ReasoningabstractRaven’s Progressive Matrices (RPMs) have been widely used to evaluate the visual reasoning ability of humans. To tackle the challenges of visual perception and logic reasoning on RPMs, we propose a Hierarchical ConViT with Attention-based Relational Reasoner (HCV-ARR). Traditional solution methods often apply relatively shallow convolution networks to visually perceive shape patterns in RPM images, which may not fully model the long-range dependencies of complex pattern combinations in RPMs. The proposed ConViT consists of a convolutional block to capture the low-level attributes of visual patterns, and a transformer block to capture the high-level image semantics such as pattern formations. Furthermore, the proposed hierarchical ConViT captures visual features from multiple receptive fields, where the shallow layers focus on the image fine details while the deeper layers focus on the image semantics. To better model the underlying reasoning rules embedded in RPM images, an Attention-based Relational Reasoner (ARR) is proposed to establish the underlying relations among images. The proposed ARR well exploits the hidden relations among question images through the developed element-wise attentive reasoner. Experimental results on three RPM datasets demonstrate that the proposed HCV-ARR achieves a significant performance gain compared with the state-of-the-art models. The source code is available at: https://github.com/wentaoheunnc/HCV-ARR. Jialu Zhang 0003, Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001 |
AAAI | 3 |
| 2023 | Siamese-Discriminant Deep Reinforcement Learning for Solving Jigsaw Puzzles with Large Eroded GapsabstractJigsaw puzzle solving has recently become an emerging research area. The developed techniques have been widely used in applications beyond puzzle solving. This paper focuses on solving Jigsaw Puzzles with Large Eroded Gaps (JPwLEG). We formulate the puzzle reassembly as a combinatorial optimization problem and propose a Siamese-Discriminant Deep Reinforcement Learning (SD2RL) to solve it. A Deep Q-network (DQN) is designed to visually understand the puzzles, which consists of two sets of Siamese Discriminant Networks, one set to perceive the pairwise relations between vertical neighbors and another set for horizontal neighbors. The proposed DQN considers not only the evidence from the incumbent fragment but also the support from its four neighbors. The DQN is trained using replay experience with carefully designed rewards to guide the search for a sequence of fragment swaps to reach the correct puzzle solution. Two JPwLEG datasets are constructed to evaluate the proposed method, and the experimental results show that the proposed SD2RL significantly outperforms state-of-the-art methods. Xingke Song, Jiahuan Jin, Chenglin Yao, Shihe Wang, Jianfeng Ren, Ruibin Bai |
AAAI | 5 |
| 2023 | Dual-Stream Siamese Vision Transformer With Mutual Attention For Radar Gait VerificationabstractThe inconspicuousness of human gait characteristic in radar signal makes it hard to differentiate different identities. In this work, a Dual-stream Siamese Vision Transformer with Mutual Attention is proposed to verify whether a pair of radar gait sequences originate from the same person or not. The proposed Siamese Vision Transformer extracts pairwise discriminant spectral information from spectrograms and cadence velocity diagrams (CVDs). The proposed Mutual Attention scheme extracts the discriminant information from each stream through a self-attention mechanism and discovers the complement information cross the two streams through a cross-attention mechanism. The proposed method is evaluated on a large benchmark radar gait verification dataset. It significantly outperforms state-of-the-art solutions. Jiarui Li 0001, Jianfeng Ren, Xudong Jiang 0001 |
ICASSP | 4 |
| 2023 | Confidence-Based Event-Centric Online Video Question Answering on a Newly Constructed ATBS DatasetabstractDeep neural networks facilitate video question answering (VideoQA), but the real-world applications on video streams such as CCTV and live cast place higher demands on the solver. To address the challenges of VideoQA on long videos of unknown length, we define a new set of problems called Online Open-ended Video Question Answering (O2VQA). It requires an online state-updating mechanism for the solver to decide if the collected information is sufficient to conclude an answer. We then propose a Confidence-based Event-centric Online Video Question Answering (CEO-VQA) model to solve this problem. Furthermore, a dataset called Answer Target in Background Stream (ATBS) is constructed to evaluate this newly developed online VideoQA application. Compared to the baseline VideoQA method that watches the whole video, the experimental results show that the proposed method achieves a significant performance gain. Weikai Kong, Shuhong Ye, Chenglin Yao, Jianfeng Ren |
ICASSP | 4 |
| 2023 | Face Recognition on Point Cloud with Cgan-Top for DenoisingabstractFace recognition using 3D point clouds is gaining growing interest, while raw point clouds often contain a significant amount of noise due to imperfect sensors. In this paper, an end-to-end 3D face recognition on a noisy point cloud is proposed, which synergistically integrates the denoising and recognition modules. Specifically, a Conditional Generative Adversarial Network on Three Orthogonal Planes (cGAN-TOP) is designed to effectively remove the noise in the point cloud, and recover the underlying features for subsequent recognition. A Linked Dynamic Graph Convolutional Neural Network (LDGCNN) is then adapted to recognize faces from the processed point cloud, which hierarchically links both the local point features and neighboring features of multiple scales. The proposed method is validated on the Bosphorus dataset. It significantly improves the recognition accuracy under all noise settings, with a maximum gain of 14.81%. Junyu Liu, Jianfeng Ren, Xudong Jiang 0001 |
ICASSP | 2 |
| 2023 | Solving Jigsaw Puzzle of Large Eroded Gaps Using Puzzlet Discriminant NetworkabstractSolving Jigsaw puzzles has recently become an emerging research topic. Traditionally, boundary similarities are utilized for puzzle reassembly. In this paper, we solve Jigsaw Puzzles of Large Eroded Gaps (JPLEG), where boundary similarities are weak and image semantics are the only feasible clues. Inspired by human strategy in solving a puzzle, we introduce the concept of puzzlet, where fragments are gradually combined to form puzzlets of different sizes until the completion of the puzzle. Two sets of Puzzlet Discriminant Networks are designed to visually perceive whether these puzzlets are correctly reassembled. The puzzle reassembly is then formulated as a combinatorial optimization problem, and solved using a genetic algorithm. The proposed method is evaluated on two large datasets, which shows that it significantly outperforms the state-of-the-art methods for puzzle solving. Xingke Song, Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001 |
ICASSP | 3 |
| 2023 | Video Question Answering Using Clip-Guided Visual-Text AttentionabstractCross-modal learning of video and text plays a key role in Video Question Answering (VideoQA). In this paper, we propose a visual-text attention mechanism to utilize the Contrastive Language-Image Pre-training (CLIP) trained on lots of general domain language-image pairs to guide the cross-modal learning for VideoQA. Specifically, we first extract video features using a TimeSformer and text features using a BERT from the target application domain, and utilize CLIP to extract a pair of visual-text features from the general-knowledge domain through the domain-specific learning. We then propose a Cross-domain Learning to extract the attention information between visual and linguistic features across the target domain and general domain. The set of CLIP-guided visual-text features are integrated to predict the answer. The proposed method is evaluated on MSVD-QA and MSRVTTQA datasets and outperforms state-of-the-art methods. Shuhong Ye, Weikai Kong, Chenglin Yao, Jianfeng Ren, Xudong Jiang 0001 |
ICIP | 4 |
| 2023 | Optimal Low-Rank QR Decomposition with an Application on RP-TSOD
Haiyan Yu 0003, Jianfeng Ren, Ruibin Bai, LinLin Shen |
ICONIP (14) | 2 |
| 2023 | A semi-supervised adaptive discriminative discretization method improving discrimination power of regularized naive BayesabstractRecently, many improved naive Bayes methods have been developed with enhanced discrimination capabilities. Among them, regularized naive Bayes (RNB) produces excellent performance by balancing the discrimination power and generalization capability. Data discretization is important in naive Bayes. By grouping similar values into one interval, the data distribution could be better estimated. However, existing methods including RNB often discretize the data into too few intervals, which may result in a significant information loss. To address this problem, we propose a semi-supervised adaptive discriminative discretization framework for naive Bayes, which could better estimate the data distribution by utilizing both labeled data and unlabeled data through pseudo-labeling techniques. The proposed method also significantly reduces the information loss during discretization by utilizing an adaptive discriminative discretization scheme, and hence greatly improves the discrimination power of classifiers. The proposed RNB+, i.e., regularized naive Bayes utilizing the proposed discretization framework, is systematically evaluated on a wide range of machine-learning datasets. It significantly and consistently outperforms state-of-the-art NB classifiers. Shihe Wang, Jianfeng Ren, Ruibin Bai |
Expert Syst. Appl. | 2 |
| 2023 | GAN-in-GAN for Monaural Speech EnhancementabstractSome generative adversarial networks (GANs) have been developed to remove background noise in real-world audio recordings. MetricGAN and its variants focus on generating a clean spectrogram from a noisy one, but the final audio quality can't be guaranteed. SEGAN and its variants directly generate an enhanced audio from a noisy one, but their over-long input representations make it less effective in identifying and removing audio noise. In this paper, a novel GAN-in-GAN framework is proposed, where the inner GAN conducts spectrogram-to-spectrogram recovery under the supervision of metric discriminators to effectively clean the audio noise, and the outer GAN conducts an audio-to-audio recovery under the supervision of multi-resolution discriminators to optimize the final audio quality. To tackle the challenges of utilizing multiple adversarial losses for training the proposed GAN-in-GAN simultaneously, a novel gradient balancing scheme is proposed to facilitate a coherent training. The proposed method is compared with state-of-the-art methods on the VoiceBank+DEMAND dataset for audio denoising. It outperforms all the compared methods. Yicun Duan, Jianfeng Ren, Heng Yu 0001, Xudong Jiang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Mask Attack Detection Using Vascular-Weighted Motion-Robust rPPG SignalsabstractDetecting 3D mask attacks to a face recognition system is challenging. Although genuine faces and 3D face masks show significantly different remote photoplethysmography (rPPG) signals, rPPG-based face anti-spoofing methods often suffer from performance degradation due to unstable face alignment in the video sequence and weak rPPG signals. To enhance the rPPG signal in a motion-robust way, a landmark-anchored face stitching method is proposed to align the faces robustly and precisely at the pixel-wise level by using both SIFT keypoints and facial landmarks. To better encode the rPPG signal, a weighted spatial-temporal representation is proposed, which emphasizes the face regions with rich blood vessels. In addition, characteristics of rPPG signals in different color spaces are jointly utilized. To improve the generalization capability, a lightweight EfficientNet with a Gated Recurrent Unit (GRU) is designed to extract both spatial and temporal features from the rPPG spatial-temporal representation for classification. The proposed method is compared with the state-of-the-art methods on five benchmark datasets under both intra-dataset and cross-dataset evaluations. The proposed method shows a significant and consistent improvement in performance over other state-of-the-art rPPG-based methods for face spoofing detection. Chenglin Yao, Jianfeng Ren, Ruibin Bai, Heshan Du, Jiang Liu 0001, Xudong Jiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Spatial Context-Aware Object-Attentional Network for Multi-Label Image ClassificationabstractMulti-label image classification is a fundamental but challenging task in computer vision. To tackle the problem, the label-related semantic information is often exploited, but the background context and spatial semantic information of related objects are not fully utilized. To address these issues, a multi-branch deep neural network is proposed in this paper. The first branch is designed to extract the discriminant information from regions of interest to detect target objects. In the second branch, a spatial context-aware approach is proposed to better capture the contextual information of an object in its surroundings by using an adaptive patch expansion mechanism. It helps the detection of small objects that are easily lost without the support of context information. The third one, the object-attentional branch, exploits the spatial semantic relations between the target object and its related objects, to better detect partially occluded, small or dim objects with the support of those easily detectable objects. To better encode such relations, an attention mechanism jointly considering the spatial and semantic relations between objects is developed. Two widely used benchmark datasets for multi-labeling classification, MS COCO and PASCAL VOC, are used to evaluate the proposed framework. The experimental results demonstrate that the proposed method outperforms the state-of-the-art methods for multi-label image classification. Jialu Zhang 0003, Jianfeng Ren, Qian Zhang 0018, Jiang Liu 0001, Xudong Jiang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Attention-Based Dual-Stream Vision Transformer for Radar Gait RecognitionabstractRadar gait recognition is robust to light variations and less infringement on privacy. Previous studies often utilize either spectrograms or cadence velocity diagrams. While the former shows the time-frequency patterns, the latter encodes the repetitive frequency patterns. In this work, a dual-stream net-work with attention-based fusion is proposed to fully aggregate the discriminant information from these two representations. Both streams are analyzed through the Vision Trans-former, which well captures the gait characteristics embedded in these representations. The proposed method is validated on a large benchmark dataset for radar gait recognition, showing that it significantly outperforms state-of-the-art solutions. Shiliang Chen, Jianfeng Ren, Xudong Jiang 0001 |
ICASSP | 3 |
| 2022 | Dynamic Texture Recognition Using PDV Hashing and Dictionary Learning on Multi-Scale Volume Local Binary PatternabstractSpatial-temporal local binary pattern (STLBP) has been widely used in dynamic texture recognition. STLBP often encounters the high-dimension problem as its dimension increases exponentially, so that STLBP could only utilize a small neighborhood. To tackle this problem, we propose a method for dynamic texture recognition using PDV hashing and dictionary learning on multi-scale volume local binary pattern (PHD-MVLBP). Instead of forming very high-dimensional LBP-histogram features, it first uses hash functions to map the pixel difference vectors (PDVs) to binary vectors, then forms a dictionary using the derived binary vector, and encodes them using the derived dictionary. In such a way, the PDVs are mapped to feature vectors of the size of the dictionary, instead of LBP histograms of very high dimension. Such an encoding scheme could extract the discriminant information from videos in a much larger neighborhood effectively. The experimental results on two widely-used dynamic textures datasets, DynTex++ and UCLA, show the superior performance of the proposed approach over the state-of-the-art methods. Ruxin Ding, Jianfeng Ren, Heng Yu 0001, Jiawei Li 0001 |
ICASSP | 2 |
| 2022 | Spatial-Context-Aware Deep Neural Network for Multi-Class Image ClassificationabstractMulti-label image classification is a fundamental but challenging task in computer vision. Over the past few decades, solutions exploring relationships between semantic labels have made great progress. However, the underlying spatial-contextual information of labels is under-exploited. To tackle this problem, a spatial-context-aware deep neural network is proposed to predict labels taking into account both semantic and spatial information. This proposed framework is evaluated on Microsoft COCO and PASCAL VOC, two widely used benchmark datasets for image multi-labelling. The results show that the proposed approach is superior to the state-of-the-art solutions on dealing with the multi-label image classification problem. Jialu Zhang 0003, Qian Zhang 0018, Jianfeng Ren, Yitian Zhao, Jiang Liu 0001 |
ICASSP | 3 |
| 2022 | Boosting the Discriminant Power of Naive BayesabstractNaive Bayes has been widely used in many applications because of its simplicity and ability in handling both numerical data and categorical data. However, lack of modeling of correlations between features limits its performance. In addition, noise and outliers in the real-world dataset also greatly degrade the classification performance. In this paper, we propose a feature augmentation method employing a stack auto-encoder to reduce the noise in the data and boost the discriminant power of naive Bayes. The proposed stack auto-encoder consists of two auto-encoders for different purposes. The first encoder shrinks the initial features to derive a compact feature representation in order to remove the noise and redundant information. The second encoder boosts the discriminant power of the features by expanding them into a higher-dimensional space so that different classes of samples could be better separated in the higher-dimensional space. By integrating the proposed feature augmentation method with the regularized naive Bayes, the discrimination power of the model is greatly enhanced. The proposed method is evaluated on a set of machine-learning benchmark datasets. The experimental results show that the proposed method significantly and consistently outperforms the state-of-the-art naive Bayes classifiers. Shihe Wang, Jianfeng Ren, Xiaoyu Lian, Ruibin Bai, Xudong Jiang 0001 |
ICPR | 2 |
| 2022 | Cross-document attention-based gated fusion network for automated medical licensing exam
Jianfeng Ren, Zheng Lu 0002, Menglin Cui, Ruibin Bai |
Expert Syst. Appl. | 2 |
| 2022 | Rain-component-aware capsule-GAN for single image de-raining
Jianfeng Ren, Zheng Lu 0002, Jialu Zhang 0003, Qian Zhang 0018 |
Pattern Recognit. | 2 |
| 2021 | rPPG-Based Spoofing Detection for Face Mask Attack using Efficientnet on Weighted Spatial-Temporal RepresentationabstractFace spoofing detection against paper attack and video-replay attack has been well studied, whereas detecting 3D face mask attack remains challenging. Remote photoplethysmography (rPPG) signal is a recently developed liveness clue for face-spoofing detection. The main challenge of existing rPPG-based methods is that the signal can be easily distorted by background noise or object motion. To address this problem, in this work, we propose an rPPG-based face-spoofing detection method using multiple regions of interests (ROIs) covering entire face, and emphasize the regions containing richer rPPG signals using larger weights. The rPPG signals of these regions form a weighted spatial-temporal map. In view of the discriminant power of EfficientNet over other deep convolutional neural networks, we propose a domain-specific EfficientNet as the classification method. Extensive experiments on two databases namely 3DMAD and HKBU-Mars V2 demonstrate the superior performance of the proposed method over state-of-the-art rPPG-based face-spoofing-detection algorithms. Chenglin Yao, Shihe Wang, Jialu Zhang 0003, Heshan Du, Jianfeng Ren, Ruibin Bai, Jiang Liu 0001 |
ICIP | 6 |
| 2021 | Calibration of optimized minimum inductor bandpass filter with controllable bandwidth and stopband rejection
Yu Wang 0155, Jianfeng Ren, Chien-In Henry Chen |
Integr. | 2 |
| 2021 | A three-step classification framework to handle complex data distribution for radar UAV detection
Jianfeng Ren, Xudong Jiang 0001 |
Pattern Recognit. | 1 |
| 2019 | Blood vessel segmentation from fundus image by a cascade classification framework
Xiaohong Wang 0003, Xudong Jiang 0001, Jianfeng Ren |
Pattern Recognit. | 3 |
| 2017 | Regularized 2-D complex-log spectral analysis and subspace reliability analysis of micro-Doppler signature for UAV detection
Jianfeng Ren, Xudong Jiang 0001 |
Pattern Recognit. | 1 |
| 2017 | LBP-Structure Optimization With Symmetry and Uniformity Regularizations for Scene ClassificationabstractLocal binary pattern (LBP) and its variants have been widely used in many visual recognition tasks. Most existing approaches utilize predefined LBP structures to extract LBP features. Recently, data-driven LBP structures have shown promising results. However, due to the limited number of training samples, data-driven structures may overfit the training samples, hence could not generalize well on the novel testing samples. To address this problem, we propose two structural regularization constraints for LBP-structure optimization: symmetry constraint and uniformity constraint. These two constraints are inspired by predefined LBP structures, which convey the human prior knowledge on designing LBP structures. The LBP-structure optimization is casted as a binary quadratic programming problem and solved efficiently via the branch-and-bound algorithm. The evaluation on two scene-classification datasets demonstrates the superior performance of the proposed approach compared with both predefined LBP structures and unconstrained data-driven LBP structures. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
IEEE Signal Process. Lett. | 1 |
| 2017 | Sound-Event Classification Using Robust Texture Features for Robot HearingabstractSound-event classification often utilizes time-frequency analysis, which produces an image-like spectrogram. Recent approaches such as spectrogram image features and subband power distribution image features extract the image local statistics such as mean and variance from the spectrogram. They have demonstrated good performance. However, we argue that such simple image statistics cannot well capture the complex texture details of the spectrogram. Thus, we propose to extract the local binary pattern (LBP) from the logarithm of the Gammatone-like spectrogram. However, the LBP feature is sensitive to noise. After analyzing the spectrograms of sound events and the audio noise, we find that the magnitude of pixel differences, which is discarded by the LBP feature, carries important information for sound-event classification. We thus propose a multichannel LBP feature via pixel difference quantization to improve the robustness to the audio noise. In view of the differences between spectrograms and natural images, and the reliability issues of LBP features, we propose two projection-based LBP features to better capture the texture information of the spectrogram. To validate the proposed multichannel projection-based LBP features for robot hearing, we have built a new sound-event classification database, the NTU-SEC database, in the context of social interaction between human and robot. It is publicly available to promote research on sound-event classification in a social context. The proposed approaches are compared with the state of the art on the RWCP database and the NTU-SEC database. They consistently demonstrate superior performance under various noise conditions. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001, Nadia Magnenat-Thalmann |
IEEE Trans. Multim. | 1 |
| 2015 | Quantized fuzzy LBP for face recognitionabstractFace recognition under large illumination variations is challenging. Local binary pattern (LBP) is robust to illumination variation, but sensitive to noise. Fuzzy LBP (FLBP) partially solves the noise-sensitivity problem by incorporating fuzzy logic in the representation of local binary patterns. The fuzzy membership function is determined by both sign and magnitude of the pixel difference. However, the magnitude is easily altered by noise, hence could be unreliable. Thus, we propose to determine the fuzzy membership function by its sign only. We name the proposed approach as Quantized Fuzzy LBP (QFLBP). On two challenging face recognition datasets, it is shown more robust to noise, and demonstrates a superior performance to FLBP and many other LBP variants. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
ICASSP | 1 |
| 2015 | Learning LBP structure by maximizing the conditional mutual information
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
Pattern Recognit. | 1 |
| 2015 | LBP Encoding Schemes Jointly Utilizing the Information of Current Bit and Other LBP BitsabstractLocal binary pattern (LBP) is sensitive to image noise. Noise-resistant LBP (NRLBP) improves the robustness to noise by incorporating the prior knowledge of images and information of other LBP bits into encoding process. However, it encodes the small pixel difference in such a way that its sign and magnitude are ignored. Although the small pixel difference may be easily distorted by noise, some of its information is still useful for LBP encoding. In this letter, we propose two enhanced NRLBPs that jointly utilize the sign and the magnitude of the current pixel difference, and also the information of other LBP bits. The proposed approaches are validated on two benchmark databases and demonstrate a superior performance compared with NRLBP and other LBP variants. The performance gain is significant when the noise level is high. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
IEEE Signal Process. Lett. | 1 |
| 2015 | A Chi-Squared-Transformed Subspace of LBP Histogram for Visual RecognitionabstractLocal binary pattern (LBP) and its variants have been widely used in many recognition tasks. Subspace approaches are often applied to the LBP feature in order to remove unreliable dimensions, or to derive a compact feature representation. It is well-known that subspace approaches utilizing up to the second-order statistics are optimal only when the underlying distribution is Gaussian. However, due to its nonnegative and simplex constraints, the LBP feature deviates significantly from Gaussian distribution. To alleviate this problem, we propose a chi-squared transformation (CST) to transfer the LBP feature to a feature that fits better to Gaussian distribution. The proposed CST leads to the formulation of a two-class classification problem. Due to its asymmetric nature, we apply asymmetric principal component analysis (APCA) to better remove the unreliable dimensions in the CST feature space. The proposed CST-APCA is evaluated extensively on spatial LBP for face recognition, protein cellular classification, and spatial-temporal LBP for dynamic texture recognition. All experiments show that the proposed feature transformation significantly enhances the recognition accuracy. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Optimizing LBP Structure For Visual Recognition Using Binary Quadratic ProgrammingabstractLocal binary pattern (LBP) and its variants have shown promising results in visual recognition applications. However, most existing approaches rely on a pre-defined structure to extract LBP features. We argue that the optimal LBP structure should be task-dependent and propose a new method to learn discriminative LBP structures. We formulate it as a point selection problem: Given a set of point candidates, the goal is to select an optimal subset to compose the LBP structure. In view of the problems of current feature selection algorithms, we propose a novel Maximal Joint Mutual Information criterion. Then, the point selection is converted into a binary quadratic programming problem and solved efficiently via the branch and bound algorithm. The proposed LBP structures demonstrate superior performance to the state-of-the-art approaches on classifying both spatial patterns in scene recognition and spatial-temporal patterns in dynamic texture recognition. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001, Gang Wang 0012 |
IEEE Signal Process. Lett. | 1 |
| 2013 | Dynamic texture recognition using enhanced LBP featuresabstractThis paper addresses the challenge of recognizing dynamic textures based on spatial-temporal descriptors. Dynamic textures are composed of both spatial and temporal features. The histogram of local binary pattern (LBP) has been used in dynamic texture recognition. However, its performance is limited by the reliability issues of the LBP histograms. In this paper, two learning-based approaches are proposed to remove the unreliable information in LBP features by utilizing Principal Histogram Analysis. Furthermore, a super histogram is proposed to improve the reliability of the LBP histograms. The temporal information is partially transferred to the super histogram. The proposed approaches are evaluated on two widely used benchmark databases: UCLA and Dyntex++ databases. Superior performance is demonstrated compared with the state of the arts. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
ICASSP | 1 |
| 2013 | Learning binarized pixel-difference pattern for scene recognitionabstractLocal binary pattern (LBP) and its variants have been used in scene recognition. However, most existing approaches rely on a pre-defined LBP structure to extract features. Those pre-defined structures can be generalized as the patterns constructed from the binarized pixel differences in a local neighborhood. Instead of using a handcraft structure, we propose to learn binarized pixel-difference patterns (BPP). We cast the problem as a feature selection problem and solve it by an incremental search via the criterion of minimum-redundancy-maximum-relevance. Then, BPP features are extracted based on the structures derived. On two challenging scene recognition databases, the proposed approach significantly outperforms the state of the arts. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
ICIP | 1 |
| 2013 | Relaxed local ternary pattern for face recognitionabstractLocal binary pattern (LBP) is sensitive to noise. Local ternary pattern (LTP) partially solves this problem by encoding the small pixel difference into a third state. The small pixel difference may be easily overwhelmed by noise. Thus, it is difficult to precisely determine its sign and magnitude. In this paper, we propose the concept of uncertain state to encode the small pixel difference. We do not care its sign and magnitude, and encode it as both 0 and 1 with equal probability. The proposed Relaxed LTP is tested on the CMU-PIE database, the extended Yale B database and the O2FN mobile face database. Superior performance is demonstrated compared with LBP and LTP. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
ICIP | 1 |
| 2013 | The complexity of two supply chain scheduling problems
Jianfeng Ren, Donglei Du, Dachuan Xu 0001 |
Inf. Process. Lett. | 1 |
| 2013 | A complete and fully automated face verification system on mobile devices
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
Pattern Recognit. | 1 |
| 2013 | Noise-Resistant Local Binary Pattern With an Embedded Error-Correction MechanismabstractLocal binary pattern (LBP) is sensitive to noise. Local ternary pattern (LTP) partially solves this problem. Both LBP and LTP, however, treat the corrupted image patterns as they are. In view of this, we propose a noise-resistant LBP (NRLBP) to preserve the image local structures in presence of noise. The small pixel difference is vulnerable to noise. Thus, we encode it as an uncertain state first, and then determine its value based on the other bits of the LBP code. It is widely accepted that most of the image local structures are represented by uniform codes and noise patterns most likely fall into the non-uniform codes. Therefore, we assign the value of an uncertain bit hence as to form possible uniform codes. Thus, we develop an error-correction mechanism to recover the distorted image patterns. In addition, we find that some image patterns such as lines are not captured in uniform codes. Those line patterns may appear less frequently than uniform codes, but they represent a set of important local primitives for pattern recognition. Thus, we propose an extended noise-resistant LBP (ENRLBP) to capture line patterns. The proposed NRLBP and ENRLBP are more resistant to noise compared with LBP, LTP, and many other variants. On various applications, the proposed NRLBP and ENRLBP demonstrate superior performance to LBP/LTP variants. Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 1 |
| 2009 | Approximation Algorithm for Minimizing the Weighted Number of Tardy Jobs on a Batch Machine
Jianfeng Ren, Yuzhong Zhang, Xianzhao Zhang, Guo Sun |
COCOA | 1 |
| 2009 | Scheduling with Rejection to Minimize the Makespan
Yuzhong Zhang, Jianfeng Ren, Chengfei Wang |
COCOA | 2 |
| 2009 | Real-time implementation of robust face detection on mobile platformsabstractAlthough many face detection algorithms have been introduced in the literature, only a handful of them can meet the real-time constraints of mobile devices. This paper presents the real-time implementation of our previously introduced face detection algorithm on a mobile device. The steps taken to achieve such a real-time implementation are discussed. Real-time comparison results with the widely used Viola-Jones face detection algorithm in terms of detection rate and processing speed are presented to demonstrate the robustness of our real-time solution. Md Tauhidur Rahman 0001, Jianfeng Ren, Nasser Kehtarnavaz |
ICASSP | 2 |
| 2009 | A hybrid face detection approach for real-time depolyment on mobile devicesabstractAlthough there are many face detection algorithms in the literature, only a handful of them meet the real-time constraints of a software-based solution without using any dedicated hardware engine. This paper presents a real-time and robust solution for mobile platforms which in general have limited computation and memory resources as compared to PC platforms. This solution involves combining our two previous real-time implementations for mobile platforms to address the shortcoming of each implementation. The first implementation provides an online or on-the-fly light source calibration for the second implementation which is found to be robust to various face poses or orientations. The real-time results obtained on an actual mobile platform indicate both the real-time and robustness capabilities of this hybrid face detection solution. Md Tauhidur Rahman 0001, Nasser Kehtarnavaz, Jianfeng Ren |
ICIP | 3 |
| 2009 | Fast eye localization based on pixel differencesabstractA novel fast eye localization algorithm based on pixel differences is presented, which is suitable for face recognition system on mobile device. It is based on the fact that eyeball is dark and round. A binary eye map is obtained by choosing those pixels darker than surrounding; then it is filtered by a rank order filter; connected regions in the eye map are then labeled by their geometric centers; best suitable eyeball pair is selected based on a set of geometric constraints. If no eyeball pair is detected, the algorithm is repeated iteratively until one pair is found. The algorithm is fast since it converts the gray level image to a binary eye map at the beginning. The algorithm is tested on our own face database, which consists of 4095 images of size 250×200. Detection rate is 93.04% when the tolerance is 0.7 times of eyeball width. Jianfeng Ren, Xudong Jiang 0001 |
ICIP | 1 |
| 2008 | Fast adaptive early termination for mode selection in H.264 scalable video codingabstractIn the scalable video coding extension of the H.264/AVC standard, an additional inter-layer prediction is combined with the traditional inter-frame prediction and intra prediction to obtain coding efficiency. This additional prediction demands a higher video encoding computational complexity. In this paper, our previously developed fast adaptive early termination algorithm is utilized to speed up the mode selection component of the H.264 scalable video coding at the base and enhancement layers. The experimental results over different content video sequences show a time reduction by 20% to 50% for spatial and quality scalable situations as compared with JSVM9.8 while achieving more or less the same video quality. Jianfeng Ren, Nasser Kehtarnavaz |
ICIP | 1 |
| 2004 | Applying Multi-class SVMs into Scene Image Classification
Jianfeng Ren, Yuntao Shen, Songhui Ma |
IEA/AIE | 1 |