EDBT 2026 Demo / reviewers in the wild / expert
Jianqin Yin
dblp:21/2631
· DBLP profile ↗
62ranked-venue papers
4as first author
54since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 2 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 23 since 2021Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Position Encoding Mechanism in Diffusion U-Net for Training-free High-resolution Image GenerationabstractDenoising higher-resolution latents using a pre-trained U-Net often results in repetitive and disordered image patterns. In this work, we are motivated to reveal the intrinsic cause of such pattern disruption in high-resolution image generation. Through theoretical analysis and empirical studies, we reveal that the pre-trained U-Net fails to provide sufficient positional information for tokens at high-resolution. Specifically, 1) zero-padding serves as a critical mechanism for position encoding but lacks robustness across varying resolutions; and 2) tokens located farther from the feature map boundaries have increasing difficulty acquiring positional awareness, leading to pattern disruptions. Inspired by these findings, we propose a novel training-free approach for high-resolution generation, introducing a Progressive Boundary Complement (PBC) method. It creates dynamic virtual image boundaries inside the feature map to supplement position information at high resolution, enabling high-quality and rich-content high-resolution image synthesis. Extensive experiments show that our method significantly improves high-resolution image synthesis in terms of visual quality and content richness, achieving state-of-the-art performance. Pu Cao, Yiyang Ma, Lu Yang 0006, Yonghao Dang, Jianqin Yin |
AAAI | 6 |
| 2026 | ViMA: A Decoupled FPGA Accelerator for Vision Mamba with Resource-Space Exploration
Jianqin Yin |
ISCAS | 5 |
| 2026 | Physical and virtual safety evaluation in autonomous driving systems using three-dimensional adversarial implementations
Yixun Zhang, Jianqin Yin, Yiyang Ma, Yingchun Niu, Xubo Zhang |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Oriented bounding box detection algorithm for dense scenarios of robotic arm operation
Jinshun Dong, Dapeng Wan, Jianqin Yin, Meiqi Guo, Shoujun Lin, Lida Liu |
Expert Syst. Appl. | 5 |
| 2026 | Image-point class incremental learning via masked rendering and Probabilistic prototype modeling
Jianqin Yin, Jialu Gu |
Expert Syst. Appl. | 2 |
| 2026 | ActivityCLIP: Enhancing group activity recognition by mining complementary information from text to supplement image modality
Jianqin Yin, Yonghao Dang |
Pattern Recognit. | 2 |
| 2026 | Object - PSF: A unified representation framework for end-to-end panoptic segmentation forecasting
Jiajun Fu, Fuxing Yang, Jianqin Yin |
Pattern Recognit. Lett. | 3 |
| 2026 | Bidirectional Fusion and Adaptation Network for Pedestrian Detection in Urban ScenesabstractPedestrian detection holds significant practical application value in fields such as urban public safety and intelligent traffic management. However, due to the influences of multi-scale variations, occlusion interferences, and complex environmental factors, pedestrian detection tasks in real-world scenarios still face numerous challenges. To address this, an efficient pedestrian detection architecture, the Bidirectional Fusion and Adaptation Network (BFA-Net), is proposed in this paper. In this method, a Bidirectional Multi-layer Fusion Network (BMFN) is designed to enhance the model's perception capability of multi-scale features. A Hybrid Stride Downsampling (HSD) module is adopted to improve the semantic representation of target features. An Adaptive Multi-scale Slimming Framework (AMSF) is introduced to balance detection accuracy and computational efficiency. In addition, a high-quality pedestrian detection dataset, UrbanCrowd-Ped, is constructed. Experimental results demonstrate that BFA-Net achieves 78.3% [email protected] and 50.5% [email protected]:0.95 on the UrbanCrowd-Ped dataset, and 78.6% and 53.2% respectively on the CityPersons dataset, outperforming existing mainstream methods. This fully validates the effectiveness and application potential of the proposed method. Dapeng Wan, Jinshun Dong, Jianqin Yin |
IEEE Signal Process. Lett. | 4 |
| 2026 | BiTAA: A Bi-Task Adversarial Attack on 3D Gaussian Fields for Cross-Task Perception
Yixun Zhang, Jianqin Yin |
IEEE Signal Process. Lett. | 3 |
| 2026 | Facial 3D Regional Structural Motion Representation Using Lightweight Point Cloud Networks for Micro-Expression RecognitionabstractHuman-computer interaction (HCI) relies on understanding and adapting to users' emotional states. Micro-expressions (MEs), a critical component of emotional perception, are characterized by their spontaneity, rapidity, subtlety, and difficulty to control. They often reveal an individual's true emotions. A comprehensive and detailed representation of motion is necessary to capture the nuances of facial dynamics effectively. Presently, motion representation methods are predominantly confined to 2D analysis within RGB images, overlooking the critical role of facial structure and its movements in conveying emotions. To overcome this limitation, we introduce an innovative facial motion representation that encompasses 3D facial structure, regionalized RGB and structural motion features. Furthermore, we segment the face into eight distinct regions, selecting only the most significant motion points to delineate the primary motion characteristics of each area. To model the interactions among crucial facial motion regions, we employ an advanced, lightweight point cloud and graph convolution network (Lite-Point-GCN). Comprehensive testing on the$\mathrm{CAS(ME)^{3}}$dataset, using leave-one-subject-out (LOSO), demonstrates that our method outperforms existing state-of-the-art methods. Jianqin Yin, Yonghao Dang, Huaping Liu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | OMEGAS: Object Mesh Extraction From Large Scenes Guided by Gaussian SegmentationabstractRecent advancements in 3D reconstruction technologies have paved the way for high-quality and real-time rendering of complex 3D scenes. Despite these achievements, a notable challenge persists: it is difficult to precisely reconstruct specific objects from large scenes. Current scene reconstruction techniques frequently result in the loss of object detail textures and are unable to reconstruct object portions that are occluded or unseen in views. To address this challenge, we delve into the meticulous 3D reconstruction of specific objects within large scenes and propose a framework termed OMEGAS: Object Mesh Extraction from Large Scenes Guided by GAussian Segmentation. Specifically, we propose a novel 3D target segmentation technique based on 2D Gaussian Splatting, which segments 3D consistent target masks in multi-view scene images and generates a preliminary target model. Moreover, to reconstruct the unseen portions of the target, we propose a novel target replenishment technique driven by large-scale generative diffusion priors. We demonstrate that our method can accurately reconstruct specific targets from large scenes, both quantitatively and qualitatively. Our experiments show that OMEGAS significantly outperforms existing reconstruction methods across various scenarios. Pu Cao, Jianqin Yin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event LocalizationabstractDense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. For such challenging task settings, most studies only employ audio-visual event semantic constraints on the final outputs, lacking progressive cross-modal semantic bridging in intermediate layers. This causes semantic gaps that hinder alignment between representations in audio and visual features, making it difficult to distinguish between event-related and irrelevant background content. Moreover, they rarely consider the correlations between events, which limits the model to infer co-occurring events among complex scenarios. In this paper, we incorporate multi-stage semantic guidance and multi-event relationship modeling, which respectively enable progressive semantic understanding of audio-visual events and adaptive extraction of event dependencies, thereby better focusing on event-related information. Specifically, our event-aware semantic guided network (ESG-Net) includes a early semantic interaction (ESI) module and a mixture of dependency experts (MoDE) module. ESI applys multi-stage semantic guidance to explicitly constrain the model in learning semantic information through multi-stage feature fusion and several classification loss functions, ensuring multi-stage understanding of event-related content. MoDE promotes the extraction of multi-event dependencies through multiple serial mixture of experts with adaptive weight allocation. Extensive experiments demonstrate that our method significantly surpasses the state-of-the-art methods, while greatly reducing parameters and computational load. Our code will be released on https://github.com/uchiha99999/ESG-Net. Huilai Li, Yonghao Dang, Jianqin Yin |
IEEE Trans. Multim. | 5 |
| 2026 | Multi-Granularity Query Network With Adaptive Category Feature Embedding for Behavior RecognitionabstractBehavior recognition is a highly challenging task, particularly in scenarios requiring unified recognition across both human and animal subjects. Most existing approaches primarily focus on single-species datasets or rely heavily on prior information such as species labels, positional annotations, or skeletal keypoints, which limits their applicability in real-world scenarios where species labels may be ambiguous or annotations are insufficient. To address these limitations, we propose a query-based Multi-Granularity Behavior Recognition Network that directly mines cross-species shared spatiotemporal behavior patterns from raw video inputs. Specifically, we design a Multi-Granularity Query module to effectively fuse fine-grained and coarse-grained features, thereby enhancing the model's capability in capturing spatiotemporal dynamics at different granularities. Additionally, we introduce a Category Query Decoder that leverages learnable category query vectors to achieve explicit behavior category modeling and mapping. Without relying on any extra annotations, the proposed method achieves unified recognition of multi-species and multi-category behaviors, setting a new state-of-the-art on the Animal Kingdom dataset and demonstrating strong generalization ability on the Charades dataset. Nuoer Long, Yonghao Dang, Chengpeng Xiong, Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Chan-Tong Lam, Jianqin Yin, Peter H. N. de With, Yue Sun 0001 |
IEEE Trans. Multim. | 9 |
| 2025 | MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event DetectionabstractSound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes. Previous methods have demonstrated impressive capabilities. However, they are deficient in learning features of complex scenes from heterogeneous dataset. In this paper, we introduce a novel dual-branch architecture named Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection (MTDA-HSED). The MTDA-HSED architecture employs the Mutual-Assistance Audio Adapter (M3A) to effectively tackle the multi-scenario problem and uses the Dual-Branch Mid-Fusion (DBMF) module to tackle the multi-granularity problem. Specifically, M3A is integrated into the BEATs block as an adapter to improve the BEATs’ performance by fine-tuning it on the multi-scenario dataset. The DBMF module connects BEATs and CNN branches, which facilitates the deep fusion of information from the BEATs and the CNN branches. Experimental results show that the proposed methods exceed the baseline of mpAUC by 5% on the DESED and MAESTRO Real datasets. Haobo Yue, Da Mu, Jin Tang 0007, Jianqin Yin |
ICASSP | 6 |
| 2025 | Q-Frame: Query-Aware Frame Selection and Multi-Resolution Adaptation for Video-LLMsabstractMultimodal Large Language Models (MLLMs) have demonstrated significant success in visual understanding tasks. However, challenges persist in adapting these models for video comprehension due to the large volume of data and temporal complexity. Existing Video-LLMs using uniform frame sampling often struggle to capture the query-related crucial spatiotemporal clues of videos effectively. In this paper, we introduce Q-Frame, a novel approach for adaptive frame selection and multi-resolution scaling tailored to the video's content and the specific query. Q-Frame employs a training-free, plug-and-play strategy generated by a text-image matching network like CLIP, utilizing the Gumbel-Max trick for efficient frame selection. Q-Frame allows Video-LLMs to process more frames without exceeding computational limits, thereby preserving critical temporal and spatial information. We demonstrate Q-Frame's effectiveness through extensive experiments on benchmark datasets, including MLVU, LongVideoBench, and Video-MME, illustrating its superiority over existing methods and its applicability across various video understanding tasks. Shaojie Zhang 0004, Jianqin Yin, Zhenbo Luo, Jian Luan 0001 |
ICCV | 3 |
| 2025 | Towards Physically Realizable Adversarial Attacks in Embodied Vision NavigationabstractThe significant advancements in embodied vision navigation have raised concerns about its susceptibility to adversarial attacks exploiting deep neural networks. Investigating the adversarial robustness of embodied vision navigation is crucial, especially given the threat of 3D physical attacks that could pose risks to human safety. However, existing attack methods for embodied vision navigation often lack physical feasibility due to challenges in transferring digital perturbations into the physical world. Moreover, current physical attacks for object detection struggle to achieve both multi-view effectiveness and visual naturalness in navigation scenarios. To address this, we propose a practical attack method for embodied navigation by attaching adversarial patches to objects, where both opacity and textures are learnable. Specifically, to ensure effectiveness across varying viewpoints, we employ a multi-view optimization strategy based on object-aware sampling, which optimizes the patch’s texture based on feedback from the vision-based perception model used in navigation. To make the patch inconspicuous to human observers, we introduce a two-stage opacity optimization mechanism, in which opacity is fine-tuned after texture optimization. Experimental results demonstrate that our adversarial patches decrease the navigation success rate by an average of 22.39%, outperforming previous methods in practicality, effectiveness, and naturalness. Code is available at: github.com/chen37058/Physical-Attacks-in-Embodied-Nav. Jiawei Tu, Yonghao Dang, Jianqin Yin |
IROS | 7 |
| 2025 | AlignCAPE: Support and Query Feature Aligning for Category-Agnostic Pose EstimationabstractRecent advancements in category-agnostic pose estimation have focused on developing a unified model capable of localizing keypoint coordinates across arbitrary categories, which enables robots to accurately interact with diverse objects by understanding their poses. While existing methods predominantly concentrate on local features surrounding the keypoints of the support image, they often overlook the importance of global features, leading to potential misalignment between the support and query image. To address the inherent conflicts between the two images, we propose AlignCAPE, a novel approach designed to mitigate such misalignment and enhance the model performance. Our method formulates a two-stage pipeline, generating initial proposals in the first stage, followed by another stage to refine iteratively. Specifically, we introduce two modules, Feature Alignment Module(FAM) and Keypoint Perception Module(KPM). FAM utilizes bidirectional cross-attention operation to align the support image feature and query image feature, thereby compensating for the limitations of previous methods. KPM employs self-attention mechanism to capture the interactions among keypoints, facilitating to localize keypoints in the query image. Experiments on MP-100 benchmark demonstrate that our method outperforms the widely-used baseline model in CAPE by 0.68% in [email protected] metric under 1-shot setting. Zhuoran Chen, Shaojie Zhang 0004, Jianqin Yin |
IROS | 6 |
| 2025 | 3DWSNet: A Novel 3D Wavelet Spiking Neural Network for Event-based Action RecognitionabstractIn robotics applications, event cameras provide low-latency and high-dynamic-range sensing by asynchronously detecting brightness changes, making them well-suited for capturing fast motions and subtle cues in dynamic environments. However, most existing Spiking Neural Network (SNN)-based methods enhance spatial information by stacking multiple frames of events, while neglecting the explicit modeling of high-and low-frequency components in the event stream. To address this limitation, we proposes a 3D Wavelet Spiking Neural Network (3DWSNet), which integrates a 3D wavelet transform with a cascaded Wavelet Spiking Convolution (WSC) module as its core. Specifically, the 3D wavelet transform decomposes input data into eight frequency sub-bands across spatial and temporal dimensions, enabling the model to preserve fine-grained high-frequency details while enriching low-frequency motion representations. The cascaded WSC architecture further improves the extraction of multi-scale spatio-temporal features by integrating information from feature maps at different resolutions. Extensive experiments show that our 3DWSNet significantly outperforms SOTA SNN performances on the CIFAR-10, CIFAR-100, DVS128 Gesture, and CIFAR10-DVS datasets. The source code will be publicly released soon. Junkang Fang, Yonghao Dang, Wending Zhao, Jianqin Yin |
IROS | 6 |
| 2025 | MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion RepresentationabstractHuman action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action recognition, the mask-based reconstruction paradigm learns the spatial structure and motion patterns of the skeleton by masking joints and reconstructing the target from unlabeled data. However, existing methods focus on a limited set of joints and low-order motion patterns, limiting the model’s ability to understand complex motion patterns. To address this issue, we introduce MaskSem, a novel semantic-guided masking method for learning 3D hybrid high-order motion representations. This novel framework leverages Grad-CAM based on relative motion to guide the masking of joints, which can be represented as the most semantically rich temporal orgions. The semantic-guided masking process can encourage the model to explore more discriminative features. Furthermore, we propose using hybrid high-order motion as the reconstruction target, enabling the model to learn multi-order motion patterns. Specifically, low-order motion velocity and high-order motion acceleration are used together as the reconstruction target. This approach offers a more comprehensive description of the dynamic motion process, enhancing the model’s understanding of motion patterns. Experiments on the NTU60, NTU120, and PKU-MMD datasets show that MaskSem, combined with a vanilla transformer, improves skeleton-based action recognition, making it more suitable for applications in human-robot interaction. The source code of our MaskSem is available at https://github.com/JayEason66/MaskSem. Shaojie Zhang 0004, Yonghao Dang, Jianqin Yin |
IROS | 4 |
| 2025 | CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
Jianqin Yin |
Int. J. Comput. Vis. | 2 |
| 2025 | A generically Contrastive Spatiotemporal Representation Enhancement for 3D skeleton action recognition
Shaojie Zhang 0004, Jianqin Yin, Yonghao Dang |
Pattern Recognit. | 2 |
| 2025 | Pseudo-EV: Enhancing 3D Visual Grounding With Pseudo Embodied Viewpointabstract3D Visual Grounding based on natural language is a fundamental task in Embodied AI. One of the fundamental challenges in localizing objects in 3D scenes through natural language descriptions arises from the variable perception of spatial relationships among objects when viewed from different perspectives. To address this issue, we introduce a model named Pseudo-EV, which decomposes the problem of 3D visual grounding into two stages:(1)predicting an embodied viewpoint and(2)determining the target object within that viewpoint, thereby eliminating viewpoint ambiguity. Given the scarcity of annotations for embodied viewpoint prediction, we employ a large language model (LLM) to generate pseudo-labels for existing datasets as intermediate training targets. However, directly predicting viewpoints in continuous Euclidean space proves inefficient, leading to weaker alignment with textual queries and scene semantics, as well as higher training overhead. To overcome these limitations, we introduce two streamlined strategies: an Embodied Viewpoint with Semantic Structure and a Decoupled Target Prediction Strategy. Extensive experiments demonstrate that predicting intermediate embodied viewpoints substantially boosts the performance of 3D visual grounding, achieving state-of-the-art results on both ScanRefer and Nr3D/Sr3D. Moreover, our framework significantly reduces computational cost compared to other viewpoint-aware approaches. Liang Geng, Jianqin Yin, Gang Chen 0029, Qingxuan Jia |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Lifting by Image - Leveraging Image Cues for Accurate 3D Human Pose EstimationabstractThe "lifting from 2D pose" method has been the dominant approach to 3D Human Pose Estimation (3DHPE) due to the powerful visual analysis ability of 2D pose estimators. Widely known, there exists a depth ambiguity problem when estimating solely from 2D pose, where one 2D pose can be mapped to multiple 3D poses. Intuitively, the rich semantic and texture information in images can contribute to a more accurate "lifting" procedure. Yet, existing research encounters two primary challenges. Firstly, the distribution of image data in 3D motion capture datasets is too narrow because of the laboratorial environment, which leads to poor generalization ability of methods trained with image information. Secondly, effective strategies for leveraging image information are lacking. In this paper, we give new insight into the cause of poor generalization problems and the effectiveness of image features. Based on that, we propose an advanced framework. Specifically, the framework consists of two stages. First, we enable the keypoints to query and select the beneficial features from all image patches. To reduce the keypoints attention to inconsequential background features, we design a novel Pose-guided Transformer Layer, which adaptively limits the updates to unimportant image patches. Then, through a designed Adaptive Feature Selection Module, we prune less significant image patches from the feature map. In the second stage, we allow the keypoints to further emphasize the retained critical image features. This progressive learning approach prevents further training on insignificant image features. Experimental results show that our model achieves state-of-the-art performance on both the Human3.6M dataset and the MPI-INF-3DHP dataset. Jianqin Yin |
AAAI | 2 |
| 2024 | Full-Frequency Dynamic Convolution: A Physical Frequency-Dependent Convolution for Sound Event Detection
Haobo Yue, Da Mu, Yonghao Dang, Jianqin Yin, Jin Tang 0007 |
ICPR (20) | 5 |
| 2024 | LLM-Driven "Coach-Athlete" Pretraining Framework for Complex Text-To-Motion GenerationabstractAdvanced text-to-motion models should one by one generate all sub-actions in the complex motion following the given motion description. Existing text-to-motion models have performed well on simple motion descriptions but they usually fail when the motion description involves multiple actions, which indicates the lack of ability to generate sub-actions step-by-step. To solve this problem, we propose a "Coach-Athlete" text-to-motion pretraining framework. In the framework, we employ the large language model (LLM) as the Coach to create incremental text and sub-action sequence pairs from complex text-motion samples. With these data, the Coach guides the Athlete, a text-to-motion model, to predict sub-action sequences in complex motion from short to long. Our approach provides a new way to inject the LLM’s motion decomposition and synthesis capabilities into a small-scale text-to-motion model. Experimental results on the current largest text-to-motion dataset HumanML3D and representative dataset KIT-ML show that our pretraining framework can facilitate performances of several state-of-the-art text-to-motion models, especially for complex motions. Jiajun Fu, Yuxing Long, Jianqin Yin |
IJCNN | 4 |
| 2024 | A Two-Stream Hybrid CNN-Transformer Network for Skeleton-Based Human Interaction Recognition
Ruoqi Yin, Jianqin Yin |
PRCV (7) | 2 |
| 2024 | Weakly supervised point cloud semantic segmentation based on scene consistency
Yingchun Niu, Jianqin Yin, Liang Geng |
Appl. Intell. | 2 |
| 2024 | Physics-constrained attack against convolution-based human motion prediction
Chengxu Duan, Yonghao Dang, Jianqin Yin |
Neurocomputing | 5 |
| 2024 | Weakly supervised point cloud semantic segmentation with the fusion of heterogeneous network features
Yingchun Niu, Jianqin Yin |
Image Vis. Comput. | 2 |
| 2024 | DHRNet: A Dual-path Hierarchical Relation Network for multi-person pose estimation
Yonghao Dang, Jianqin Yin, Pengxiang Ding |
Knowl. Based Syst. | 2 |
| 2024 | April-GCN: Adjacency Position-velocity Relationship Interaction Learning GCN for Human motion prediction
Baoxuan Gu, Jin Tang 0007, Rui Ding 0018, Jianqin Yin |
Knowl. Based Syst. | 5 |
| 2024 | MLP-AIR: An effective MLP-based module for actor interaction relation learning in group activity recognition
Jianqin Yin, Shaojie Zhang 0004, Moonjun Gong |
Knowl. Based Syst. | 2 |
| 2024 | Transfer the global knowledge for current gaze estimation
Jianqin Yin |
Multim. Tools Appl. | 2 |
| 2024 | Lgvc: language-guided visual context modeling for 3D visual grounding
Liang Geng, Jianqin Yin, Yingchun Niu |
Neural Comput. Appl. | 2 |
| 2024 | Any region can be perceived equally and effectively on rotation pretext task using full rotation and weighted-region mixture
Rui Liu 0033, Min Wang 0032, Jianqin Yin, Jun Liu 0007 |
Neural Networks | 5 |
| 2024 | Kinematics modeling network for video-based human pose estimation
Yonghao Dang, Jianqin Yin, Shaojie Zhang 0004, Jiping Liu |
Pattern Recognit. | 2 |
| 2024 | Instance-Incremental Scene Graph Generation From Real-World Point Clouds via Normalizing FlowsabstractThis work introduces a new task of instance-incremental scene graph generation: Given a scene of the point cloud, representing it as a graph and automatically increasing novel instances. A graph denoting the object layout of the scene is finally generated. It is an important task since it helps to guide the insertion of novel 3D objects into a real-world scene in vision-based applications like augmented reality. It is also challenging because the complexity of the real-world point cloud brings difficulties in learning object layout experiences from the observation data (non-empty rooms with labeled semantics). We model this task as a conditional generation problem and propose a 3D autoregressive framework based on normalizing flows (3D-ANF) to address it. First, we represent the point cloud as a graph by extracting the label semantics and contextual relationships. Next, a model based on normalizing flows is introduced to map the conditional generation of graphic elements into the Gaussian process. The mapping is invertible. Thus, the real-world experiences represented in the observation data can be modeled in the training phase, and novel instances can be autoregressively generated based on the Gaussian process in the testing phase. To evaluate the performance of our method sufficiently, we implement this new task on the indoor benchmark dataset 3DSSG-O27R16 and our newly proposed graphical dataset of outdoor scenes GPL3D. Experiments show that our method generates reliable novel graphs from the real-world point cloud and achieves state-of-the-art performance on the datasets. Jianqin Yin, Jinghang Xu, Pengxiang Ding |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | SiT-MLP: A Simple MLP With Point-Wise Topology Feature Learning for Skeleton-Based Action RecognitionabstractGraph convolution networks (GCNs) have achieved remarkable performance in skeleton-based action recognition. However, previous GCN-based methods rely on elaborate human priors excessively and construct complex feature aggregation mechanisms, which limits the generalizability and effectiveness of networks. To solve these problems, we propose a novel Spatial Topology Gating Unit (STGU), an MLP-based variant without extra priors, to capture the co-occurrence topology features that encode the spatial dependency across all joints. In STGU, to learn the point-wise topology features, a new gate-based feature interaction mechanism is introduced to activate the features point-to-point by the attention map generated from the input sample. Based on the STGU, we propose the first MLP-based model, SiT-MLP, for skeleton-based action recognition in this work. Compared with previous methods on three large-scale datasets, SiT-MLP achieves competitive performance. In addition, SiT-MLP reduces the parameters significantly with favorable results. The code will be available at https://github.com/BUPTSJZhang/SiT-MLP. Shaojie Zhang 0004, Jianqin Yin, Yonghao Dang, Jiajun Fu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Rich Action-Semantic Consistent Knowledge for Early Action PredictionabstractEarly action prediction (EAP) aims to recognize human actions from a part of action execution in ongoing videos, which is an important task for many practical applications. Most prior works treat partial or full videos as a whole, ignoring rich action knowledge hidden in videos, i.e., semantic consistencies among different partial videos. In contrast, we partition original partial or full videos to form a new series of partial videos and mine the Action-Semantic Consistent Knowledge (ASCK) among these new partial videos evolving in arbitrary progress levels. Moreover, a novel Rich Action-semantic Consistent Knowledge network (RACK) under the teacher-student framework is proposed for EAP. Firstly, we use a two-stream pre-trained model to extract features of videos. Secondly, we treat the RGB or flow features of the partial videos as nodes and their action semantic consistencies as edges. Next, we build a bi-directional semantic graph for the teacher network and a single-directional semantic graph for the student network to model rich ASCK among partial videos. The MSE and MMD losses are incorporated as our distillation loss to enrich the ASCK of partial videos from the teacher to the student network. Finally, we obtain the final prediction by summering the logits of different subnetworks and applying a softmax layer. Extensive experiments and ablative studies have been conducted, demonstrating the effectiveness of modeling rich ASCK for EAP. With the proposed RACK, we have achieved state-of-the-art performance on three benchmarks. The code is available at https://github.com/lily2lab/RACK.git. Jianqin Yin, Di Guo 0002, Huaping Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Deeply Supervised Skin Lesions Diagnosis With Stage and Branch AttentionabstractAccurate and unbiased examinations of skin lesions are critical for the early diagnosis and treatment of skin diseases. Visual features of skin lesions vary significantly because the images are collected from patients with different lesion colours and morphologies by using dissimilar imaging equipment. Recent studies have reported that ensembled convolutional neural networks (CNNs) are practical to classify the images for early diagnosis of skin disorders. However, the practical use of these ensembled CNNs is limited as these networks are heavyweight and inadequate for processing contextual information. Although lightweight networks (e.g., MobileNetV3 and EfficientNet) were developed to achieve parameter reduction for implementing deep neural networks on mobile devices, insufficient depth of feature representation restricts the performance. To address the existing limitations, we develop a new lite and effective neural network, namely HierAttn. The HierAttn applies a novel deep supervision strategy to learn the local and global features by using multi-stage and multi-branch attention mechanisms with only one training loss. The efficacy of HierAttn was evaluated by using the dermoscopy images dataset ISIC2019 and smartphone photos dataset PAD-UFES-20 (PAD2020). The experimental results show that HierAttn achieves the best accuracy and area under the curve (AUC) among the state-of-the-art lightweight networks. Rui Liu 0033, Min Wang 0032, Jianqin Yin, Jun Liu 0007 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Leveraging the Video-Level Semantic Consistency of Event for Audio-Visual Event LocalizationabstractAudio-visual event (AVE) localization has attracted much attention in recent years. Most existing methods are often limited to independently encoding and classifying each video segment separated from the full video (which can be regarded as the segment-level representations of events). However, they ignore the semantic consistency of the event within the same full video (which can be considered as the video-level representations of events). In contrast to existing methods, we propose a novel video-level semantic consistency guidance network for the AVE localization task. Specifically, we propose an event semantic consistency modeling (ESCM) module to explore video-level semantic information for semantic consistency modeling. It consists of two components: a cross-modal event representation extractor (CERE) and an intra-modal semantic consistency enhancer (ISCE). CERE is proposed to obtain the event semantic information at the video level. Furthermore, ISCE takes video-level event semantics as prior knowledge to guide the model to focus on the semantic continuity of an event within each modality. Moreover, we propose a new negative pair filter loss to encourage the network to filter out the irrelevant segment pairs and a new smooth loss to further increase the gap between different categories of events in the weakly-supervised setting. We perform extensive experiments on the public AVE dataset and outperform the state-of-the-art methods in both fully- and weakly-supervised settings, thus verifying the effectiveness of our method. Jianqin Yin, Yonghao Dang |
IEEE Trans. Multim. | 2 |
| 2024 | Learning Constrained Dynamic Correlations in Spatiotemporal Graphs for Motion PredictionabstractHuman motion prediction is challenging due to the complex spatiotemporal feature modeling. Among all methods, graph convolution networks (GCNs) are extensively utilized because of their superiority in explicit connection modeling. Within a GCN, the graph correlation adjacency matrix drives feature aggregation, and thus, is the key to extracting predictive motion features. State-of-the-art methods decompose the spatiotemporal correlation into spatial correlations for each frame and temporal correlations for each joint. Directly parameterizing these correlations introduces redundant parameters to represent common relations shared by all frames and all joints. Besides, the spatiotemporal graph adjacency matrix is the same for different motion samples, and thus, cannot reflect samplewise correspondence variances. To overcome these two bottlenecks, we propose dynamic spatiotemporal decompose GC (DSTD-GC), which only takes 28.6% parameters of the state-of-the-art GC. The key of DSTD-GC is constrained dynamic correlation modeling, which explicitly parameterizes the common static constraints as a spatial/temporal vanilla adjacency matrix shared by all frames/joints and dynamically extracts correspondence variances for each frame/joint with an adjustment modeling function. For each sample, the common constrained adjacency matrices are fixed to represent generic motion patterns, while the extracted variances complete the matrices with specific pattern adjustments. Meanwhile, we mathematically reformulate GCs on spatiotemporal graphs into a unified form and find that DSTD-GC relaxes certain constraints of other GC, which contributes to a better representation capability. Moreover, by combining DSTD-GC with prior knowledge like body connection and temporal context, we propose a powerful spatiotemporal GCN called DSTD-GCN. On the Human3.6M, Carnegie Mellon University (CMU) Mocap, and 3D Poses in the Wild (3DPW) datasets, DSTD-GCN outperforms state-of-the-art methods by 3.9%-8.7% in prediction accuracy with 55.0%-96.9% fewer parameters. Codes are available at https://github.com/Jaakk0F/DSTD-GCN. Jiajun Fu, Fuxing Yang, Yonghao Dang, Jianqin Yin |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Collaborative Multi-Dynamic Pattern Modeling for Human Motion PredictionabstractThe dynamic information of the joints, such as the movement amplitude, is critical for forecasting precise human joint trajectories. Existing methods adopt global modeling in which all joints are treated as a whole to extract features for movement coordination. Though global modeling can exploit hidden relationships between joints, it also inevitably introduces undesired trajectory dependencies, which weakens the dynamic information effects and simplifies the constraints and kinetics model of joints. Therefore, we propose a dynamic pattern-based collaborative modeling framework (DPnet) that contains a keyframe enhanced module (KEM) and multi-channel feature extractor blocks (MFE-block). The KEM tackles the discontinuity between the last frame of observation and the first predicted one by duplicating the decisive frame. The MFE-block utilizes a multi-channel graph structure to enrich the dynamic information effects and recessive constraints of joints. To distinguish the dynamic information of each joint, we calculate the movement amplitude of the joints and propose three dynamic patterns, including active, inactive, and static patterns. We also propose a dynamic pattern-guided feature extractor (DP-FE) to alleviate the trajectory dependencies between joints with different dynamic patterns. We evaluate our approach on three standard benchmark datasets, including H3.6M [8], CMU-Mocap [44], and 3DPW [45]. Our approach achieves impressive results in both short-term and long-term predictions, confirming its effectiveness and efficiency. Jin Tang 0007, Rui Ding 0018, Baoxuan Gu, Jianqin Yin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Neighborhood Spatial Aggregation MC Dropout for Efficient Uncertainty-Aware Semantic Segmentation in Point CloudsabstractUncertainty-aware semantic segmentation of the point clouds includes predictive uncertainty estimation and uncertainty-guided model optimization. One key challenge in the task is the efficiency of pointwise predictive distribution establishment. The widely used Monte Carlo (MC) dropout establishes the distribution by computing the standard deviation of samples using multiple stochastic forward propagations, which is time-consuming for tasks based on point clouds containing massive points. Hence, a framework embedded with neighborhood spatial aggregation (NSA)-MC dropout, a variant of MC dropout, is proposed to establish distributions in just one forward pass. Specifically, our method uses the one-time stochastic inference of a point with neighbors to approximate the point’s repeated stochastic inferences, outputting pointwise distribution via the prediction variance of neighbors. Based on this, uncertainties acquire from the predictive distribution. The aleatoric uncertainty is integrated into the loss function to suppress noise, preventing the model from overfitting. Besides, the predictive uncertainty quantifies the prediction confidence. Experiments show that our plug-and-play NSA-MC dropout significantly improves the segmentation performance of backbones, ranging from multilayer perceptron (MLP)-, convolution-, and attention-based to transformer-based networks, without introducing additional parameters. Besides, it is several times faster than MC dropout to quantify the results’ credibility, and the inference time does not establish a coupling relation with the sampling times. Jianqin Yin, Yingchun Niu, Jinghang Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Temporal consistency two-stream CNN for human motion prediction
Jin Tang 0007, Jianqin Yin |
Neurocomputing | 3 |
| 2022 | Towards More Realistic Human Motion Prediction With Attention to Motion CoordinationabstractJoint relation modeling is a curial component in human motion prediction. Most existing methods rely on skeletal-based graphs to build the joint relations, where local interactive relations between joint pairs are well learned. However, the motion coordination, a global joint relation reflecting the simultaneous cooperation of all joints, is usually weakened because it is learned from part to whole progressively and asynchronously. Thus, the final predicted motions usually appear unrealistic. To tackle this issue, we learn a medium, called coordination attractor (CA), from the spatiotemporal features of motion to characterize the global motion features, which is subsequently used to build new relative joint relations. Through the CA, all joints are related simultaneously, and thus the motion coordination of all joints can be better learned. Based on this, we further propose a novel joint relation modeling module, Comprehensive Joint Relation Extractor (CJRE), to combine this motion coordination with the local interactions between joint pairs in a unified manner. Additionally, we also present a Multi-timescale Dynamics Extractor (MTDE) to extract enriched dynamics from the raw position information for effective prediction. Extensive experiments show that the proposed framework outperforms state-of-the-art methods in both short- and long-term predictions on H3.6M, CMU-Mocap, and 3DPW. Pengxiang Ding, Jianqin Yin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Sign Language Recognition Based on R(2+1)D With Spatial-Temporal-Channel AttentionabstractPrevious work utilized three-dimensional (3-D) convolutional neural networks (CNNs) tomodel the spatial appearance and temporal evolution concurrently for sign language recognition (SLR) and exhibited impressive performance. However, there are still challenges for 3-D CNN-based methods. First, motion information plays a more significant role than spatial content in sign language. Therefore, it is still questionable whether to treat space and time equally and model them jointly by heavy 3-D convolutions in a unified approach. Second, because of the interference from the highly redundant information in sign videos, it is still nontrivial to effectively extract discriminative spatiotemporal features related to sign language. In this study, deep R(2+1)D was adopted for separate spatial and temporal modeling and demonstrated that decomposing 3-D convolution filters into independent spatial and temporal convolutions facilitates the optimization process in SLR. A lightweight spatial–temporal–channel attention module, including two submodules called channel–temporal attention and spatial–temporal attention, was proposed to make the network concentrate on the significant information along spatial, temporal, and channel dimensions by combining squeeze and excitation attention with self-attention. By embedding this module into R(2+1)D, superior or comparable results to the state-of-the-art methods on the CSL-500, Jester, and EgoGesture datasets were obtained, which demonstrated the effectiveness of the proposed method. Xiangzu Han, Fei Lu 0005, Jianqin Yin, Guohui Tian, Jun Liu 0007 |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2022 | Relation-Based Associative Joint Location for Human Pose Estimation in VideosabstractVideo-based human pose estimation (VHPE) is a vital yet challenging task. While deep learning algorithms have made tremendous progress for the VHPE, lots of these approaches to this task implicitly model the long-range interaction between joints by expanding the receptive field of the convolution or designing a graph manually. Unlike prior methods, we design a lightweight and plug-and-play joint relation extractor (JRE) to explicitly and automatically model the associative relationship between joints. The JRE takes the pseudo heatmaps of joints as input and calculates their similarity. In this way, the JRE can flexibly learn the correlation between any two joints, allowing it to learn the rich spatial configuration of human poses. Furthermore, the JRE can infer invisible joints according to the correlation between joints, which is beneficial for locating occluded joints. Then, combined with temporal semantic continuity modeling, we propose a Relation-based Pose Semantics Transfer Network (RPSTN) for video-based human pose estimation. Specifically, to capture the temporal dynamics of poses, the pose semantic information of the current frame is transferred to the next with a joint relation guided pose semantics propagator (JRPSP). The JRPSP can transfer the pose semantic features from the non-occluded frame to the occluded frame. The proposed RPSTN achieves state-of-the-art or competitive results on the video-based Penn Action, Sub-JHMDB, PoseTrack2018, and HiEve datasets. Moreover, the proposed JRE improves the performance of backbones on the image-based COCO2017 dataset. Code is available at https://github.com/YHDang/pose-estimation. Yonghao Dang, Jianqin Yin, Shaojie Zhang 0004 |
IEEE Trans. Image Process. | 2 |
| 2022 | $\hbox {PISEP}{^2}$: pseudo-image sequence evolution-based 3D pose prediction
Jianqin Yin, Huaping Liu 0001, Yilong Yin |
Vis. Comput. | 2 |
| 2021 | Neighborhood Spatial Aggregation based Efficient Uncertainty Estimation for Point Cloud Semantic SegmentationabstractUncertainty estimation for point cloud semantic segmentation is to quantify the confidence degree for the predicted label of points, which is essential for decision-making tasks. This paper proposes a neighborhood spatial aggregation based method, NSA-MC dropout, to achieve efficient uncertainty estimation for point cloud semantic segmentation. Unlike the traditional uncertainty estimation method MC dropout de-pending on repeated inferences, our NSA-MC dropout achieves uncertainty estimation through one-time inference. Specifically, a space-dependent method is designed to sample the model many times by performing stochastic forward pass through the model just once, and it approximates the repeated inferences based sampling process in MC dropout. Besides, a neighborhood spatial aggregation module, called NSA, aggregates neighborhood probabilistic outputs for each point and works with space-dependent sampling to establish output distribution. Finally, we propose an uncertainty-aware framework NSA-MC dropout to capture the uncertainty of prediction results efficiently. Experimental results show that our method obtains comparable performance with MC dropout. More significantly, our NSA-MC dropout has little influence on the efficiency of semantic inference. It is much faster than MC dropout, and the inference time does not establish a coupling relation with the sampling times. Our code is available at https://github.com/chaoqi7/Uncertainty_Estimation_PCSS Jianqin Yin, Huaping Liu 0001, Jun Liu 0007 |
ICRA | 2 |
| 2021 | Dynamic tracking for microrobot with active magnetic sensor arrayabstractAccurate position feedback in a wide range is critical for medical microrobotics and robot-assisted examinations, such as colonoscopy, bronchoscopy and capsule endoscopy examination. Among the many modalities of positioning feedback, magnetic tracking is a preferable method due to the unique advantages of free line of sight, free energy storage and untethered connection. However, the field strength of the magnetic source decreases with the third power of the distance, limiting the effectiveness of position feedback at long distances. In order to maintain a consistently high tracking accuracy in a broad area, this paper presents a new dynamic tracking solution by applying a movable sensor array. In this new solution, the tracking accuracy of the magnet is first determined and optimized within a short range. When the target microrobot carrying the magnet exceeds this optimized range, the sensor array is relocated by an external robotic arm to keep the target in the effective tracking range. Moreover, we also propose a multi-point locating algorithm to minimize the varying background noise. Experimental results show that the proposed method increases the range of magnetic tracking and achieves a satisfactory level of tracking accuracy, which demonstrates significant potentials to improve the position feedback of microrobots in medical applications. Min Wang 0032, Kwan Yi Leung, Rui Liu 0033, Shuang Song 0002, Yixuan Yuan, Jianqin Yin, Max Q.-H. Meng, Jun Liu 0007 |
ICRA | 6 |
| 2021 | MASK-GD segmentation based robotic grasp detection
Mingshuai Dong, Shimin Wei, Xiuli Yu, Jianqin Yin |
Comput. Commun. | 4 |
| 2021 | TrajectoryCNN: A New Spatio-Temporal Feature Learning Network for Human Motion PredictionabstractHuman motion prediction is an increasingly interesting topic in computer vision and robotics. In this paper, we propose a new end-to-end feedforward network, TrajectoryCNN, to predict future poses. Compared with the most existing methods, we introduce a new trajectory space and focus on modeling motion dynamics of the input sequence with coupled spatio-temporal features, dynamic local-global features, and global temporal co-occurrence features in the new space. Specifically, the coupled spatio-temporal features describe the spatial and temporal structural information hidden in a natural human motion sequence, which can be easily mined using CNN by simultaneously covering the spatial and temporal dimensions of the sequence with the convolutional filters. The dynamic local-global features encode different correlations among joint trajectories of human motion (i.e. strong correlations among joint trajectories of one part and weak correlations among joint trajectories of different parts), which can be captured by stacking multiple residual trajectory blocks and incorporating our skeletal representation. The global temporal co-occurrence features represent different importance of different input poses to mine the motion dynamics for predicting future poses, which can be obtained automatically by learning free parameters for each pose with our TrajectoryCNN. Finally, we predict future poses with the captured motion dynamic features in a non-recursive manner. Extensive experiments show that our method achieves state-of-the-art performance on five benchmarks (e.g. Human3.6M, CMU-Mocap, 3DPW, G3D, and FNTU), which demonstrates the effectiveness of our proposed method. The code is available at https://github.com/lily2lab/TrajectoryCNN.git. Jianqin Yin, Jin Li 0002, Pengxiang Ding, Jun Liu 0007, Huaping Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Energy-Based Periodicity Mining With Deep Features for Action Repetition Counting in Unconstrained VideosabstractAction repetition counting is to estimate the occurrence times of the repetitive motion in one action, which is a relatively new, significant, but challenging problem. To solve this problem, we propose a new method superior to the traditional ways in two aspects, without preprocessing and applicable for arbitrary periodicity actions. Without preprocessing, the proposed model makes our scheme convenient for real applications; processing the arbitrary periodicity action makes our model more suitable for the actual circumstance. In terms of methodology, firstly, we extract action features using ConvNets and then use Principal Component Analysis algorithm to generate the intuitive periodic information from the chaotic high-dimensional features; secondly, we propose an energy-based adaptive feature mode selection scheme to adaptively select proper deep feature mode according to the background of the video; thirdly,we construct the periodic waveform of the action based on the high-energy rules by filtering the irrelevant information. Finally, we detect the peaks to obtain the times of the action repetition. Our work features two-fold: 1) We give a significant insight that features extracted by ConvNets for action recognition can well model the self-similarity periodicity of the repetitive action. 2) A high-energy based periodicity mining rule using features from ConvNets is presented, which can process arbitrary actions without preprocessing. Experimental results show that our method achieves superior or comparable performance on the three benchmark datasets, i.e. YT_Segments, QUVA, and RARV. Jianqin Yin, Yanchun Wu, Chaoran Zhu, Zijin Yin, Huaping Liu 0001, Yonghao Dang, Zhiyi Liu, Jun Liu 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | One-Shot SADI-EPE: A Visual Framework of Event Progress EstimationabstractIn many practical engineering applications, the number of actions that have been finished should be known, particularly for an untrimmed video sequence that includes an event with a series of actions, it is important to know the number of actions that have been finished. In this paper, we termed this process as visual event progress estimation (EPE). However, the research related to this problem is few in the research community. To solve this problem, a visual human action analysis-based framework, namely one-shot simultaneously action detection and identification (SADI)-EPE, is presented in this paper. The visual EPE is modeled as an online one-shot learning-based problem; sliding window and attention-based bag of key poses formulate our framework. Unlike most of the action analysis methods relying on a number of training data of some predefined classes, our method can realize SADI for any event if one sample of the event is given, which makes it feasible for practical applications. At the same time, not only SADI but also the progress estimation of the event can be realized by our algorithm. In terms of methodology, the key pose is defined by an invariant pose descriptor from skeletal data and silhouette data. Moreover, in order to extract representative and discriminative poses from one training sample, we present a new bidirectional kNN-based attention weighted key pose selection method, which can filter the unrelated actions and model different importance of various key poses. In addition, an attention-based multi-modal fusion scheme, which addresses the difficulty of high-dimensional features and few training samples, is proposed to augment the performance of our algorithm. Finally, we propose an evaluation criterion for the estimation problem. Extensive results demonstrated the efficacy of our proposed framework. Jianqin Yin, Fuchun Sun 0001, Huaping Liu 0001, Bin Wang 0045, Jun Liu 0007, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Estimating Cement Compressive Strength from Microstructure Images Using Broad Learning SystemabstractThe microstructure images of cement are often used as the main data source for estimating compressive strength. They contain ample physical properties during the hydration process. Different gray values represent different substances in the grayscale image of cement. Deep learning algorithm based on microstructure images have been proposed to estimate cement compressive strength (CCS). However, there are a large number of parameters that need to be adjusted in deep structure. The high-efficiency system named broad learning system (BLS) is tried to use to estimate the cement compressive strength. The original cement microstructure images and the extracted features are used as input respectively, the connection weights can be obtained directly by calculating pseudo inverse matrix of feature matrix of microstructure image. If the structure is not sufficient to gain suitable result, BLS only calculates the pseudo inverse matrix of additional nodes to improve accuracy. The experiment shows that the broad learning structure (BLS) is an effective and efficient method on estimating cement compressive strength by contrasting with deep learning structure. Yonghao Dang, Lin Wang 0004, Jianqin Yin, Xuehui Zhu, Zhiquan Feng, Jifeng Guo 0002 |
SMC | 3 |
| 2017 | Coronal Mass Ejections detection using multiple features based ensemble learning
Jianqin Yin, Hai Yao, Jiaben Lin, Yilong Yin, Zhiquan Feng |
Neurocomputing | 1 |
| 2016 | An HCI paradigm fusing flexible object selection and AOM-based animation
Zhiquan Feng, Bo Yang 0001, Hong Liu 0013, Jianqin Yin, Yuan Zhang 0007, Xiuyang Zhao |
Inf. Sci. | 6 |
| 2015 | Motion-towards-each-other-based hand gesture initialization
Zhiquan Feng, Bo Yang 0001, Jianqin Yin, Xiuyang Zhao, Shichang Feng |
Pattern Recognit. | 5 |
| 2013 | Real-time oriented behavior-driven 3D freehand tracking for direct interaction
Zhiquan Feng, Bo Yang 0001, Yi Li 0026, Yanwei Zheng, Xiuyang Zhao, Jianqin Yin, Qingfang Meng |
Pattern Recognit. | 6 |
| 2004 | A new color-based face detection and location by using support vector machineabstractFace detection and location are widely used in practice. Although there are lots of related researches, the results are not satisfactory, esp. for the situation of side-face and multiscale in complex background. Among these techniques, the method based on skin-color is an important approach, so far, however, the effect of detection and location is not favorable due to the unsatisfactory segmentation. In order to improve the segmentation, we employ a new proposed classifier - support vector machine (SVM), based on the output of which we put forward a new scheme to realize the face detection and location. Experimental results indicate the algorithm can work well on wide range of color, light and size variations in still images. Jianqin Yin, Jinping Li, Yanbin Han, Aizeng Cao |
ICARCV | 1 |
| 2004 | A universal scheme of hidden information detection from original signals via wavelet neural networksabstractA universal mathematical scheme for hidden information detection from various original n-dimensional signals is presented. The key points include the establishment of functional expression and the employment of cascade neural networks. The former indicates the detection of hidden information is dependent upon the shape of whole signals, not the specific values and the sampling number of the signals, the latter refers to the architecture of neural networks consisting of two neural networks, the first extracts main features of signals, and the second detects hidden information from the extracted features. Since wavelet functions play important role in signal analysis and feature extraction, thus wavelet neural networks constitute the first neural networks. The general scheme of feature extraction from original signals in L/sup 2/(R/sup M/) by wavelet neural networks is presented. The application in signal processing of chemical chromatographic spectra of solution shows satisfactory results. Jinping Li, Yanbin Han, Hongbo Zhong, Jianqin Yin |
ICARCV | 4 |