Zhidong Deng

dblp:32/6262 · DBLP profile ↗
← Back
77ranked-venue papers
10as first author
32since 2021 · last 2025
0000-0001-9970-1023ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 9 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 1 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 1 since 2021Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Exploring Timeline Control for Facial Motion Generation
abstract
This paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arranged in temporal intervals, allowing precise control over the timing of each action. To model the timeline control capability, We first annotate the time intervals of facial actions in natural facial motion sequences at a frame-level granularity. This process is facilitated by Toeplitz Inverse Covariance-based Clustering to minimize human labor. Based on the annotations, we propose a diffusion-based generation model capable of generating facial motions that are natural and accurately aligned with input timelines. Our method supports text-guided motion generation by using ChatGPT to convert text into timelines. Experimental results show that our method can annotate facial action intervals with satisfactory accuracy, and produces natural facial motions accurately aligned with timelines.
Yifeng Ma 0001, Jinwei Qi, Chaonan Ji, Peng Zhang 0080, Bang Zhang, Zhidong Deng, Liefeng Bo
CVPR6
2025 LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
abstract
Recent advances in large vision-language models (LVLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches by ViT results in a fragmented perception, thereby hindering the visual understanding capabilities of LVLMs. In this paper, we propose an innovative enhancement to address this limitation by introducing a Scene Graph Expression (SGE) module in LVLMs. This module extracts and structurally expresses the complex semantic information within images, thereby improving the foundational perception and understanding abilities of LVLMs. Extensive experiments demonstrate that integrating our SGE module significantly enhances the LVLM’s performance in vision-language tasks, indicating its effectiveness in preserving intricate semantic details and facilitating better visual understanding.
Jianzhong Ju, Jian Luan 0001, Zhidong Deng
ICASSP4
2025 Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
abstract
Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus on grounding objects in static observations from 3D reconstruction, such as meshes and point clouds, but lack the ability to actively perceive and explore their environment. To address this limitation, we introduce \underline{\textbf{M}}ove \underline{\textbf{t}}o \underline{\textbf{U}}nderstand (\textbf{\model}), a unified framework that integrates active perception with \underline{\textbf{3D}} vision-language learning, enabling embodied agents to effectively explore and understand their environment. This is achieved by three key innovations: 1) Online query-based representation learning, enabling direct spatial memory construction from RGB-D frames, eliminating the need for explicit 3D reconstruction. 2) A unified objective for grounding and exploring, which represents unexplored locations as frontier queries and jointly optimizes object grounding and frontier selection. 3) End-to-end trajectory learning that combines \textbf{V}ision-\textbf{L}anguage-\textbf{E}xploration pre-training over a million diverse trajectories collected from both simulated and real-world RGB-D sequences. Extensive evaluations across various embodied navigation and question-answering benchmarks show that MTU3D outperforms state-of-the-art reinforcement learning and modular navigation approaches by 14\%, 23\%, 9\%, and 2\% in success rate on HM3D-OVON, GOAT-Bench, SG3D, and A-EQA, respectively. \model's versatility enables navigation using diverse input modalities, including categories, language descriptions, and reference images. These findings highlight the importance of bridging visual grounding and exploration for embodied intelligence.
Xilin Wang, Zhuofan Zhang, Xiaojian Ma 0001, Yixin Chen 0003, Baoxiong Jia, Wei Liang 0008, Zhidong Deng, Siyuan Huang 0001, Qing Li 0003
ICCV10
2025 PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection
abstract
Single point supervised oriented object detection has gained attention and made initial progress within the community. Diverse from those approaches relying on one-shot samples or powerful pretrained models (e.g. SAM), PointOBB has shown promise due to its prior-free feature. In this paper, we propose PointOBB-v2, a simpler, faster, and stronger method to generate pseudo rotated boxes from points without relying on any other prior. Specifically, we first generate a Class Probability Map (CPM) by training the network with non-uniform positive and negative sampling. We show that the CPM is able to learn the approximate object regions and their contours. Then, Principal Component Analysis (PCA) is applied to accurately estimate the orientation and the boundary of objects. By further incorporating a separation mechanism, we resolve the confusion caused by the overlapping on the CPM, enabling its operation in high-density scenarios. Extensive comparisons demonstrate that our method achieves a training speed 15.58$\times$ faster and an accuracy improvement of 11.60\%/25.15\%/21.19\% on the DOTA-v1.0/v1.5/v2.0 datasets compared to the previous state-of-the-art, PointOBB. This significantly advances the cutting edge of single point supervised oriented detection in the modular track. Code and models will be released.
Botao Ren, Xue Yang 0005, Yi Yu 0010, Zhidong Deng
ICLR5
2025 Feedback RoI Features Improve Aerial Object Detection
abstract
Research in visual perception has shown that the human visual system utilizes high-level feedback information to guide lower-level processing, enabling adaptation to signals of varying characteristics. Inspired by this, we propose the Feedback multi-Level feature Extractor (Flex) to dynamically adjust feature selection in object detection based on image-wise and instance-level feedback information. This is particularly beneficial for applications such as aerial object detection, UAV-based target recognition and autonomous vehicle navigation, where global image quality issues like sensor degradation, foggy, or rainy conditions can impact detection performance. Flex adapts to variations in image quality, refining the feature extraction process to improve robustness against these challenges. Experimental results demonstrate that Flex consistently enhances a range of state-of-the-art methods on challenging aerial object detection datasets, including DOTA-v1.0, DOTA-v1.5, and HRSC2016. Furthermore, additional experiments on MS COCO confirm the module's effectiveness in general object detection tasks. Our quantitative and qualitative analyses reveal that the improvements are strongly correlated with image quality, aligning with our original motivation to address global image quality issues in real-world scenarios.
Botao Ren, Botian Xu, Hanwei Gao, Qiankun Yu, Zhidong Deng
ICRA6
2025 Rethinking the Evaluation of Scene Graph Generation
Hanwei Gao, Zhidong Deng
PRCV (6)3
2025 TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles
abstract
Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which may be difficult to find and hence limits their usage. In this work, we propose TalkCLIP, a framework that can generate talking heads where the expressions are specified by natural language, hence allowing for specifying expressions more conveniently. To model the mapping from text to expressions, we first construct a text-video paired talking head dataset where each video has diverse text descriptions that depict both coarse-grained emotions and fine-grained facial movements. Leveraging the proposed dataset, we introduce a CLIP-based style encoder that projects natural language-based descriptions to the representations of expressions. TalkCLIP can even infer expressions for descriptions unseen during training. TalkCLIP can also use text to modulate expression intensity and edit expressions. Extensive experiments demonstrate that TalkCLIP achieves the advanced capability of generating photo-realistic talking heads with vivid facial expressions guided by text descriptions.
Yifeng Ma 0001, Suzhen Wang 0001, Yu Ding 0001, Tangjie Lv, Changjie Fan, Zhipeng Hu, Zhidong Deng, Xin Yu 0002
IEEE Trans. Multim.8
2024 Incorporating Rotation Invariance with Non-invariant Networks for Point Clouds
abstract
Rotation invariance is a fundamental requirement of point cloud processing when input point clouds are not aligned. Many non-invariant networks performing well on aligned point clouds do not perform equivalent to rotated ones. Thus non-invariant and invariant networks are developed separately and only benefit a little from each other, leading to repetitive and wasteful research efforts. In this paper, we aim to bridge this gap and incorporate rotation invariance with non-invariant networks for point clouds. To this end, we propose a novel rotation invariant learning method based on efficient invariant poses (EIPs). EIPs do not rely on novel features or operations. Instead, they only rotate input point clouds into invariant poses and apply non-invariant networks in feature processing. As the name implies, EIPs have negligible complexities (efficient) and solid theoretical foundations (invariant). Experimental results demonstrate that EIPs have competitive performances on several tasks. Without using new features or operations, EIPs yield the best results on ScanObjectNN (PB T50 RS) classification and ShapeNetPart segmentation task. Our code is available at: https://github.com/JaronTHU/EIP.
Jiajun Fei, Zhidong Deng
3DV2
2024 An Online Calibration Method for Robust Multi-Modality 3D Object Detection
abstract
Multi-modality sensor fusion for 3D perception is a significant part for autonomous driving perception, which enables a comprehensive integration of different sensors and obtains a holistic understanding of the surrounding environment to improve accuracy, stability and reliability. Nevertheless, vibrations, collisions, and acceleration/deceleration in motion may result in a minor disturbance to the position of sensors, leading to offsets in sensor calibration. To address the limitations of the original method demanding considerable manual effort and time for hand-annotating checkerboards, we introduce an online Camera-LiDAR calibration method, which can perform calculations during the operation of autonomous vehicles to eliminate sensor biases. Different from the previous CNN-based online calibration algorithm, our approach differs in utilizing multi-scale features to accomplish alignment with a foundational ResNet and FPN Backbone alongside an attention-based multi-scale feature fusion module. Moreover, in the design of the loss function, we incorporate depth loss and point cloud loss in addition to the original smooth L1 norm loss to facilitate network backpropagation. Ultimately, our proposed method achieves an average translation error of 0.88 cm and a rotation error of 0.073° on the KITTI odometry dataset, which can significantly limit the discrepancies between sensors and thus enhance detection accuracy.
Yige Yao, Jianming Hu, Zhidong Deng
DSAA4
2024 Unifying 3D Vision-Language Understanding via Promptable Queries
Zhuofan Zhang, Xiaojian Ma 0001, Xuesong Niu, Yixin Chen 0003, Baoxiong Jia, Zhidong Deng, Siyuan Huang 0001, Qing Li 0003
ECCV (44)7
2024 Open-Vocabulary Skeleton Action Recognition with Diffusion Graph Convolutional Network and Pre-Trained Vision-Language Models
abstract
This study explores unsupervised open-vocabulary skeleton action recognition, aiming at addressing inaccurate spatial matching and poor interpretability of existing GCN models. We present Skeleton-DGCFA, an approach to make feature alignment (FA) of skeleton with image modalities based on a large pre-trained vision and language (VL) model along with our new diffusion graph convolutional (DGC) skeleton encoder. The DGC comprises spatial and temporal convolutional modules, allowing for the diffusion of different graph semantic features. Skeleton-DGCFA harnesses recent large-scale VL models and extends their zero-shot capabilities to the skeleton modality by capitalizing on its natural pairing with images. The open-vocabulary zero-shot capabilities improve with the strength of the pre-trained VL model and our DGC skeleton encoder. We establish a new state-of-the-art in the zero-shot skeleton action recognition tasks, significantly surpassing the vanilla zero-shot method by 27.0% and 19.7% on NTU-60 and NTU-120, respectively.
Zhidong Deng
ICASSP2
2024 A Novel Contrastive Diffusion Graph Convolutional Network for Few-Shot Skeleton-Based Action Recognition
abstract
Existing skeleton spatial-temporal models tend to deteriorate the positional distinguishability of skeleton joints and lead to inaccurate spatial matching and poor interpretability. This paper proposes a novel contrastive diffusion graph convolutional network (CD-GCN) for few-shot action recognition based on attentional diffusion and contrastive loss. In CD-GCN, a diffusion graph convolution is employed to provide robust representation, enabling the diffusion of diverse graph semantic features. Additionally, a new meta-classifier is designed to make use of contrastive loss to pull together skeleton features of the same class and push apart samples from different classes in order to have more suitable similarity metric. Experimental results demonstrate that the proposed method achieves 3%-5% and 2%-3% relative improvements for 1-shot and 5-shot tasks, respectively, over previous state-of-the-art methods.
Zhidong Deng
ICASSP2
2024 Fine-Tuning Point Cloud Transformers with Dynamic Aggregation
abstract
Point clouds play an important role in 3D analysis, which has broad applications in robotics and autonomous driving. The pre-training fine-tuning paradigm has shown great potential in the point cloud domain. Full fine-tuning is generally effective but leads to a heavy storage and computational burden, which becomes inefficient and unacceptable as the size of pre-trained models scales. Although efficient fine-tuning approaches have significant progress in other domains, they generally perform worse for point clouds. To overcome this dilemma, we revisit the official Point-MAE implementation and find the critical role of aggregation in fine-tuning performances. Inspired by such discoveries, we propose a novel dynamic aggregation (DA) method to replace previous static aggregation like mean or max pooling for pre-trained point cloud Transformers. Besides standard metrics such as accuracy or mIoU, we evaluate the number of tunable parameters and additional FLOPs for a fair comparison of our method to different fine-tuning approaches. We construct several DA variants and validate them through extensive experiments. Experimental results demonstrate that DA has competitive performances against full fine-tuning and other efficient fine-tuning approaches. The code is publicly available at https://github.com/JaronTHU/DynamicAggregation.
Jiajun Fei, Zhidong Deng
ICRA2
2024 Incorporating Scene Graphs into Pre-trained Vision-Language Models for Multimodal Open-vocabulary Action Recognition
abstract
This paper presents Action-SGFA, a novel action feature alignment approach to learn unified joint embeddings across four action modalities incorporating scene graph (SG) comprehension. A new training paradigm for Action-SGFA is also devised to improve pre-trained VL models using datasets with SG annotation. When learning from image-SG pairs, it captures structure-associated action knowledge for visual and textual encoders. SG supervision generates fine-grained captions based on various graph augmentations highlighting different compositional aspects of action scenes. Furthermore, our research reveals that all combinations of paired data are unnecessary to train such unified embeddings, and only image-paired data is sufficient to bind all action modalities together. Our Action-SGFA can leverage existing large VL models, enhancing their zero-shot capabilities of new modalities due to their natural pairings with images. The open-vocabulary zero-shot performance improves with the strength of the pre-trained VL model and the SG comprehension. We establish a new state-of-the-art in several zero-shot action recognition tasks across modalities, significantly surpassing the vanilla skeleton zero-shot method by 27.0% and 19.7% on NTU-60 and NTU-120, respectively. Additionally, in the context of RGB videos, we surpass the state-of-the-art method on Kinetics-400 by 2.1%.
Zhidong Deng
ICRA2
2024 StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking Heads
abstract
Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail to produce diverse speaking styles in the final videos. To address this challenge, we propose a one-shot style-controllable talking face generation method that can obtain speaking styles from reference speaking videos and drive the one-shot portrait to speak with the reference speaking styles and another piece of audio. Our method aims to synthesize the style-controllable coefficients of a 3D Morphable Model (3DMM), including facial expressions and head movements, in a unified framework. Specifically, the proposed framework first leverages a style encoder to extract the desired speaking styles from the reference videos and transform them into style codes. Then, the framework uses a style-aware decoder to synthesize the coefficients of 3DMM from the audio input and style codes. During decoding, our framework adopts a two-branch architecture, which generates the stylized facial expression coefficients and stylized head movement coefficients, respectively. After obtaining the coefficients of 3DMM, an image renderer renders the expression coefficients into a specific person's talking-head video. Extensive experiments demonstrate that our method generates visually authentic talking head videos with diverse speaking styles from only one portrait image and an audio clip.
Suzhen Wang 0001, Yifeng Ma 0001, Yu Ding 0001, Zhipeng Hu, Changjie Fan, Tangjie Lv, Zhidong Deng, Xin Yu 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 G2-SCANN: Gaussian-kernel graph-based SLD clustering algorithm with natural neighbourhood
abstract
For most clustering methods, not only the number of clusters must be set in advance, but also various hyperparameters such as initial centroids, number of nearest neighbours, the minimum number of points, neighbourhood radius, and cutoff distance all require pre-specification. As one of the most promising unsupervised learning methods in machine intelligence, existing clustering methods cannot simultaneously handle datasets with arbitrary shapes, different densities, distinct sizes, and overlapping. Background outliers and high dimensionality make clustering problems more challenging. In this paper, we propose a novel universal clustering methodology, called G2-SCANN, which yields the best clustering performance for all 30 synthetic and real datasets without any hyperparameter tuning if the exact number of clusters is known. Firstly, the shortest path length (SPL) in complex network or graph-based geodesic distance is used to give a locally backbone-structured description of graph vertex similarity. Accordingly, SPL-weighted local degree (SLD) is defined as vertex attributes of a SPL-weighted graph expressed by G2-SPL adjacency matrix with ε-natural neighbourhood. Secondly, the process of calculating SLD for every data point in a bottom-up way directly leads to division from a complete graph constituted by all data points to a group of SLD trees. This brings the interpretability and the elimination of lone trees. Thirdly, contrastive learning of largest SLD values for finding root vertices of each divisive tree is conducted and top-down category message is then transmitted from the root vertices to all the leaf ones of a SLD tree. It eventually produces tree-like clusters. Totally, the proposed G2-SCANN method leverages both local neighbouring similarity of data points and global information about data distribution and makes it perform better than other methods.
Zhidong Deng
Pattern Recognit.1
2023 StyleTalk: One-Shot Talking Head Generation with Controllable Speaking Styles
abstract
Different people speak with diverse personalized speaking styles. Although existing one-shot talking head methods have made significant progress in lip sync, natural facial expressions, and stable head motions, they still cannot generate diverse speaking styles in the final talking head videos. To tackle this problem, we propose a one-shot style-controllable talking face generation framework. In a nutshell, we aim to attain a speaking style from an arbitrary reference speaking video and then drive the one-shot portrait to speak with the reference speaking style and another piece of audio. Specifically, we first develop a style encoder to extract dynamic facial motion patterns of a style reference video and then encode them into a style code. Afterward, we introduce a style-controllable decoder to synthesize stylized facial animations from the speech content and style code. In order to integrate the reference speaking style into generated videos, we design a style-aware adaptive transformer, which enables the encoded style code to adjust the weights of the feed-forward layers accordingly. Thanks to the style-aware adaptation mechanism, the reference speaking style can be better embedded into synthesized videos during decoding. Extensive experiments demonstrate that our method is capable of generating talking head videos with diverse speaking styles from only one portrait image and an audio clip while achieving authentic visual effects. Project Page: https://github.com/FuxiVirtualHuman/styletalk.
Yifeng Ma 0001, Suzhen Wang 0001, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding 0001, Zhidong Deng, Xin Yu 0002
AAAI7
2023 Learnable Flow Model Conditioned on Graph Representation Memory for Anomaly Detection
abstract
Anomaly detection could be applied in a wide range of fields from industrial scene to medical imaging analysis. Although invertible flow models are developed to accomplish unsupervised anomaly detection, they are usually hard to train and have limited capabilities of accurately modeling the distribution of normal samples. To address this problem, we propose a novel enhanced flow model conditioned on graph representation memory (FlowGRM) for visual surface defect detection. In our FlowGRM, a graph neural network is trained to query graph embedding features from memory bank and incorporate them into a learnable flow model as conditional information. Such memorized conditional information, together with the invertible flow model, is jointly optimized to give rise to a probability density score. Experimental results obtained on the industrial visual anomaly detection datasets including MVTec and MTD show that the proposed model has state-of-the-art performance, which could even reach up to 99.6% and 99.2% AUROC scores on the above two bench-marks, respectively. Meanwhile, it requires fewer parameters and saves over 80% training time compared to existing flow models.
Wenlei Liu, Zhidong Deng
ICASSP3
2023 3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment
abstract
3D vision-language grounding (3D-VL) is an emerging field that aims to connect the 3D physical world with natural language, which is crucial for achieving embodied intelligence. Current 3D-VL models rely heavily on sophisticated modules, auxiliary losses, and optimization tricks, which calls for a simple and unified model. In this paper, we propose 3D-VisTA, a pre-trained Transformer for 3D Vision and Text Alignment that can be easily adapted to various downstream tasks. 3D-VisTA simply utilizes self-attention layers for both single-modal modeling and multi-modal fusion without any sophisticated task-specific design. To further enhance its performance on 3D-VL tasks, we construct ScanScribe, the first large-scale 3D scene-text pairs dataset for 3D-VL pre-training. ScanScribe contains 2,995 RGB-D scans for 1,185 unique indoor scenes originating from ScanNet and 3R-Scan datasets, along with paired 278K scene descriptions generated from existing 3D-VL tasks, templates, and GPT-3. 3D-VisTA is pre-trained on ScanScribe via masked language/object modeling and scene-text matching. It achieves state-of-the-art results on various 3D-VL tasks, ranging from visual grounding and dense captioning to question answering and situated reasoning. Moreover, 3D-VisTA demonstrates superior data efficiency, obtaining strong performance even with limited annotations during downstream task fine-tuning.
Xiaojian Ma 0001, Yixin Chen 0003, Zhidong Deng, Siyuan Huang 0001, Qing Li 0003
ICCV4
2023 Cross-Modality Time-Variant Relation Learning for Generating Dynamic Scene Graphs
abstract
Dynamic scene graphs generated from video clips could help enhance the semantic visual understanding in a wide range of challenging tasks such as environmental perception, autonomous navigation, and task planning of self-driving vehicles and mobile robots. In the process of temporal and spatial modeling during dynamic scene graph generation, it is particularly intractable to learn time-variant relations in dynamic scene graphs among frames. In this paper, we propose a Time-variant Relation-aware TRansformer (TR2), which aims to model the temporal change of relations in dynamic scene graphs. Explicitly, we leverage the difference of text embeddings of prompted sentences about relation labels as the supervision signal for relations. In this way, cross-modality feature guidance is realized for the learning of time-variant relations. Implicitly, we design a relation feature fusion module with a transformer and an additional message token that describes the difference between adjacent frames. Extensive experiments on the Action Genome dataset prove that our TR2 can effectively model the time-variant relations. TR2 significantly outperforms previous state-of-the-art methods under two different settings by 2.1 % and 2.6% respectively.
Jinfa Huang, Can Zhang 0001, Zhidong Deng
ICRA4
2023 Accommodating Self-attentional Heterophily Topology into High- and Low-pass Graph Convolutional Network for Skeleton-based Action Recognition
abstract
In the scene of human skeleton-based action recognition, graph convolutional network (GCN) is widely applied to model human action dynamics and achieves extraordinary results, where graph topology is in charge of feature extraction and plays a crucial role in learning representation in GCNs. This paper presents a novel high- and low-pass graph convolutional network (HLP-GCN), aiming to effectively aggregate joint features based on learned homophily and heterophily topologies using self-attention modules for skeleton-based action recognition tasks. Specifically, such an HLP-GCN framework first learns homophily and heterophily topologies dynamically so as to capture different correlations between input samples and then combine similar information arising from homophily topologies with dissimilar ones from heterophily topologies to get the final embeddings. Furthermore, a temporal modeling module is added to this HLPGCN framework, and a new pooling strategy called global node (GNode) is cascaded subsequently, in order to have the extraction of more robust spatiotemporal features for skeleton action recognition. Finally, the experimental results demonstrate that our method outperforms previous state-of-the-art methods by 0.2%, 0.4%, and 0.3% on three large-scale action recognition datasets including NTU RGB+D X-Set, NTU RGB+D 120 X-Sub, and NW-UCLA, respectively.
Zhidong Deng
IJCNN2
2023 Improving Scene Graph Generation with Superpixel-Based Interaction Learning
abstract
Recent advances in Scene Graph Generation (SGG) typically model the relationships among entities utilizing box-level features from pre-defined detectors. We argue that an overlooked problem in SGG is the coarse-grained interactions between boxes, which inadequately capture contextual semantics for relationship modeling, practically limiting the development of the field. In this paper, we take the initiative to explore and propose a generic paradigm termed Superpixel-based Interaction Learning (SIL) to remedy coarse-grained interactions at the box level. It allows us to model fine-grained interactions at the superpixel level in SGG. Specifically, (i) we treat a scene as a set of points and cluster them into superpixels representing sub-regions of the scene. (ii) We explore intra-entity and cross-entity interactions among the superpixels to enrich fine-grained interactions between entities at an earlier stage. Extensive experiments on two challenging benchmarks (Visual Genome and Open Image V6) prove that our SIL enables fine-grained interaction at the superpixel level above previous box-level methods, and significantly outperforms previous state-of-the-art methods across all metrics. More encouragingly, the proposed method can be applied to boost the performance of existing box-level approaches in a plug-and-play fashion. In particular, SIL brings an average improvement of 2.0% mR (even up to 3.4%) of baselines for the PredCls task on Visual Genome, which facilitates its integration into any existing box-level method.
Can Zhang 0001, Jinfa Huang, Botao Ren, Zhidong Deng
ACM Multimedia5
2022 DuMLP-Pin: A Dual-MLP-Dot-Product Permutation-Invariant Network for Set Feature Extraction
abstract
Existing permutation-invariant methods can be divided into two categories according to the aggregation scope, i.e. global aggregation and local one. Although the global aggregation methods, e. g., PointNet and Deep Sets, get involved in simpler structures, their performance is poorer than the local aggregation ones like PointNet++ and Point Transformer. It remains an open problem whether there exists a global aggregation method with a simple structure, competitive performance, and even much fewer parameters. In this paper, we propose a novel global aggregation permutation-invariant network based on dual MLP dot-product, called DuMLP-Pin, which is capable of being employed to extract features for set inputs, including unordered or unstructured pixel, attribute, and point cloud data sets. We strictly prove that any permutation-invariant function implemented by DuMLP-Pin can be decomposed into two or more permutation-equivariant ones in a dot-product way as the cardinality of the given input set is greater than a threshold. We also show that the DuMLP-Pin can be viewed as Deep Sets with strong constraints under certain conditions. The performance of DuMLP-Pin is evaluated on several different tasks with diverse data sets. The experimental results demonstrate that our DuMLP-Pin achieves the best results on the two classification problems for pixel sets and attribute sets. On both the point cloud classification and the part segmentation, the accuracy of DuMLP-Pin is very close to the so-far best-performing local aggregation method with only a 1-2% difference, while the number of required parameters is significantly reduced by more than 85% in classification and 69% in segmentation, respectively. The code is publicly available on https://github.com/JaronTHU/DuMLP-Pin.
Jiajun Fei, Wenlei Liu, Zhidong Deng, Mingyang Li 0001, Huanjun Deng
AAAI4
2022 Few-shot Graph Classification with Contrastive Loss and Meta-classifier
abstract
Few-shot graph-level classification based on graph neural networks is critical in many tasks including drug and material discovery. We present a novel graph contrastive relation network (GCRNet) by introducing a practical yet straightforward graph meta-baseline with contrastive loss to gain robust representation and meta-classifier to realize more suitable similarity metric, which is more adaptive for graph few-shot problems. Experimental results demonstrate that the proposed method achieves 8%-12% in 5-shot, 5%-8% in 10 shot, and 1%-5% in 20-shot improvements, respectively, compared to the existing state-of-the-art methods.
Zhidong Deng
IJCNN2
2021 MP-Mono: Monocular 3D Detection Using Multiple Priors for Autonomous Driving
abstract
Monocular 3D object detection is an important and challenging task in autonomous driving. Due to the ill-posed nature of 3D detection, recent studies use prior knowledge on object categories to estimate 3D parameters. However, for each object category in real driving scenes, there exist a couple of sub-categories with different shapes (i.e. length, width, and height). For example, vehicle generally contains the sub-categories of car, van, and truck. Obviously, single prior knowledge cannot cover such diverse sub-categories. In this paper, we propose MP-Mono that exploits multiple priors to improve object detection. Specifically, a data-heuristic strategy is presented to generate multiple 3D proposals, in which we leverage the unsupervised algorithm to cluster potential sub-categories from realistic datasets, and a height-guided inference policy is used to determine the initial distances of proposals, reducing the difficulty of network learning. Additionally, we propose a local-ground embedding method that learns local depth information to enhance monocular 3D detection. The experimental results on the KITTI dataset demonstrate that our MP-Mono achieves competitive performances compared to other monocular methods, verifying the effectiveness of multi-prior integration and local-ground embedding.
Guorun Yang, Zhe Wang 0006, Jianping Shi, Zhidong Deng, Yu Qiao 0001
3DV6
2021 Loopster++: Termination Analysis for Multi-path Linear Loop
Weimin Ge, Yao Zhang 0019, Xiaohong Li 0001, Zhidong Deng
CollaborateCom (1)5
2021 Predicting Entity Relations across Different Security Databases by Using Graph Attention Network
abstract
Security databases such as Common Vulnerabilities and Exposures (CVE), Common Weakness Enumeration (CWE), and Common Attack Pattern Enumeration and Classification (CAPEC) maintain diverse high-quality security concepts, which are treated as security entities. Meanwhile, security entities are documented with many potential relation types that profit for security analysis and comprehension across these three popular databases. To support reasoning security entity relationships, translation-based knowledge graph representation learning treats each triple independently for the entity prediction. However, it neglects the important semantic information about the neighbor entities around the triples. To address it, we propose a text-enhanced graph attention network model (text-enhanced GAT). This model highlights the importance of the knowledge in the 2−hop neighbors surrounding a triple, under the observation of the diversity of each entity. Thus, we can capture more structural and textual information from the knowledge graph about the security databases. Extensive experiments are designed to evaluate the effectiveness of our proposed model on the prediction of security entity relationships. Moreover, the experimental results outperform the state-of-the-art by Mean Reciprocal Rank (MRR) 0.132 for detecting the missing relationships.
Yude Bai, Zhenchang Xing, Sen Chen 0001, Xiaohong Li 0001, Zhidong Deng
COMPSAC6
2021 Learning With Memory For Few-Shot Semantic Segmentation
abstract
Despite great progress made in the few-shot semantic segmentation task, the existing works still suffer from problems of incompleteness and inconsistency of segmentation. In this paper, a novel attention-aided LSTM optimization network called LONet is proposed, which optimizes predictions without forgetting useful inner cues. Particularly, we calculate an attention map to align and match possible locations with query features to deal with incomplete segmentation. Then, an LSTM-based module is designed to overcome the segmentation inconsistency by memorizing and updating useful cues iteratively. Extensive experiments are conducted on two popular few-shot segmentation datasets including PASCAL-5iand FSS-1000. The experimental results on the FSS-1000 dataset demonstrate that our LONet exceeds the state-of-the-art results by 2.1% and 2.3%, respectively.
Hongchao Lu, Zhidong Deng
ICIP3
2021 Feature Enhanced Projection Network for Zero-shot Semantic Segmentation
abstract
In environmental perception of autonomous driving, zero-shot semantic segmentation that can make prediction of new categories without using any labeled training samples is considered as a challenging task. One key step in this task is to transfer knowledge across categories via auxiliary semantic word embeddings. In this paper, we propose a feature enhanced projection network (FEPNet) that takes full advantage of transferred knowledge to enrich semantic representations. In FEPNet, two projection layers are added to a segmentation network so as to map features into seen (S) and unseen (U) category spaces, respectively. During training, U-space features are transferred to S-space using similarity relations to enhance the representation of seen categories. In the inference stage, the representation of unseen categories is also strengthened by incorporating features transferred from S-space. Moreover, a novel strategy is proposed to effectively alleviate prediction bias by performing segmentation independently in separate areas that contain seen and unseen categories. We conduct extensive experiments on three benchmark datasets. The experimental results show that our FEPNet achieves new state-of-the-art results compared to existing approaches.
Hongchao Lu, Longwei Fang, Matthieu Lin, Zhidong Deng
ICRA4
2021 Lane Intrusion Behaviors Dataset: Action Recognition in Real-world Highway Scenarios for Self-driving
abstract
It is necessary for the development of self-driving to fulfill the requirements for safety, stability, and intelligence, especially in high-speed conditions. Therefore, the detection of pedestrians that may occur in highway scenarios during driving and understanding the meaning of their behaviors in advance are significantly important for the self-driving vehicle to make correct decisions. However, no existing datasets are available for behavior recognition in self-driving scenarios. In order to advance the task of interactive cognition between the vehicle and pedestrians, in this paper, we present a new dataset, called THU-IntrudBehavior, that collects lane intrusion behaviors of pedestrians that can be applied in real world highway scenarios. The dataset contains diverse behaviors of single or multiple pedestrians/cyclists that are simulated in different urban roads under various weather conditions. We describe annotations of each video and report several experimental results of baseline methods on our self-collected dataset. Our THU-IntrudBehavior dataset provides new support for behavior recognition in high-speed conditions for self-driving.
Hongchao Lu, Zhidong Deng, Ruiwen Zhang
IJCNN2
2021 A Deep Graph Wavelet Convolutional Neural Network for Semi-supervised Node Classification
abstract
Graph convolutional neural network provides good solutions for node classification and other tasks with non-Euclidean data. There are several graph convolutional models that attempt to develop deep networks but do not cause serious over-smoothing at the same time. Considering that the wavelet transform generally has a stronger ability to extract useful information than the Fourier transform, we propose a new deep graph wavelet convolutional network (DeepGWC) for semi-supervised node classification tasks. Based on the optimized static filtering matrix parameters of vanilla graph wavelet neural networks and the combination of Fourier bases and wavelet ones, DeepGWC is constructed together with the reuse of residual connection and identity mappings in network architectures. Extensive experiments on three benchmark datasets including Cora, Citeseer, and Pubmed are conducted. The experimental results demonstrate that our DeepGWC outperforms existing graph deep models with the help of additional wavelet bases and achieves new state-of-the-art performances eventually.
Zhidong Deng
IJCNN2
2021 Phase Space Reconstruction Network for Lane Intrusion Action Recognition
abstract
In a complex road traffic scene, illegal lane intrusion of pedestrians or cyclists constitutes one of the main safety challenges in autonomous driving application. In this paper, we propose a novel object-level phase space reconstruction network (PSRNet) for motion time series classification, aiming to recognize lane intrusion actions that occur 150m ahead through a monocular camera fixed on moving vehicle. In the PSRNet, the movement of pedestrians and cyclists, specifically viewed as an observable object-level dynamic process, can be reconstructed as trajectories of state vectors in a latent phase space and further characterized by a learnable Lyapunov exponent-like classifier that indicates discrimination in terms of average exponential divergence of state trajectories. Additionally, in order to first transform video inputs into one-dimensional motion time series of each object, a lane width normalization based on visual object tracking-by-detection is presented. Extensive experiments are conducted on the THU-IntrudBehavior dataset collected from real urban roads. The results show that our PSRNet could reach the best accuracy of 98.0%, which remarkably exceeds existing action recognition approaches by more than 30%.
Ruiwen Zhang, Zhidong Deng, Hongsen Lin, Hongchao Lu
IJCNN2
2020 A Boundary-aware Distillation Network for Compressed Video Semantic Segmentation
abstract
In recent years optical flow is often estimated to reuse features so as to accelerate video semantic segmentation. With addition of optical flow network, however, extra cost may incur and accuracy may thus be degraded because of repeated warping operation. In this paper, we propose a boundary-aware distillation network (BDNet) that replaces optical flow network with block motion vectors encoded in compressed video, resulting in negligible computational complexity. In order to make salient features, an auxiliary boundary-aware stream is added to the main stream to jointly estimate silhouette and segmentation of objects. To further correct warped features, a well-trained teacher network is employed to transfer knowledge to the main stream. Both boundary-aware stream and the teacher network are neglected during inference stage, so that video segmentation network enables to get faster without increasing any computational burden. By splitting the task into three components, our BDNet shows almost 10% time saving as well as 1.6% accuracy improvement over baseline on the Cityscapes dataset.
Hongchao Lu, Zhidong Deng
ICPR2
2019 DrivingStereo: A Large-Scale Dataset for Stereo Matching in Autonomous Driving Scenarios
abstract
Great progress has been made on estimating disparity maps from stereo images. However, with the limited stereo data available in the existing datasets and unstable ranging precision of current stereo methods, industry-level stereo matching in autonomous driving remains challenging. In this paper, we construct a novel large-scale stereo dataset named DrivingStereo. It contains over 180k images covering a diverse set of driving scenarios, which is hundreds of times larger than the KITTI Stereo dataset. High-quality labels of disparity are produced by a model-guided filtering strategy from multi-frame LiDAR points. For better evaluations, we present two new metrics for stereo matching in the driving scenes, i.e. a distance-aware metric and a semantic-aware metric. Extensive experiments show that compared with the models trained on FlyingThings3D or Cityscapes, the models trained on our DrivingStereo achieve higher generalization accuracy in real-world driving scenes, while the proposed metrics better evaluate the stereo methods on all-range distances and across different classes. Our dataset and code are available at https://drivingstereo-dataset.github.io.
Guorun Yang, Chaoqin Huang, Zhidong Deng, Jianping Shi, Bolei Zhou
CVPR4
2019 Fast Object Detection in Compressed Video
abstract
Object detection in videos has drawn increasing attention since it is more practical in real scenarios. Most of the deep learning methods use CNNs to process each decoded frame in a video stream individually. However, the free of charge yet valuable motion information already embedded in the video compression format is usually overlooked. In this paper, we propose a fast object detection method by taking advantage of this with a novel Motion aided Memory Network (MMNet). The MMNet has two major advantages: 1) It significantly accelerates the procedure of feature extraction for compressed videos. It only need to run a complete recognition network for I-frames, i.e. a few reference frames in a video, and it produces the features for the following P frames (predictive frames) with a light weight memory network, which runs fast; 2) Unlike existing methods that establish an additional network to model motion of frames, we take full advantage of both motion vectors and residual errors that are freely available in video streams. To our best knowledge, the MMNet is the first work that investigates a deep convolutional detector on compressed videos. Our method is evaluated on the large-scale ImageNet VID dataset, and the results show that it is 3× times faster than single image detector R-FCN and 10× times faster than high-performance detector MANet at a minor accuracy loss.
Shiyao Wang 0001, Hongchao Lu, Zhidong Deng
ICCV3
2019 A Novel Dual Successive Projection-Based Model-Free Adaptive Control Method and Application to an Autonomous Car
abstract
In this paper, a novel model-free adaptive control (MFAC) algorithm based on a dual successive projection (DuSP)-MFAC method is proposed, and it is analyzed using the introduced DuSP method and the symmetrically similar structures of the controller and its parameter estimator of MFAC. Then, the proposed DuSP-MFAC scheme is successfully implemented in an autonomous car "Ruilong" for the lateral tracking control problem via converting the trajectory tracking problem into a stabilization problem by using the proposed preview-deviation-yaw angle. This MFAC-based lateral tracking control method was tested and demonstrated satisfactory performance on real roads in Fengtai, Beijing, China, and through successful participation in the Chinese Smart Car Future Challenge Competition held in 2015 and 2016.
Shida Liu, Zhongsheng Hou, Yantao Tian, Zhidong Deng, Zhenxuan Li
IEEE Trans. Neural Networks Learn. Syst.4
2018 SRC-Disp: Synthetic-Realistic Collaborative Disparity Learning for Stereo Matching
Guorun Yang, Zhidong Deng, Hongchao Lu, Zeping Li
ACCV (5)2
2018 Fully Motion-Aware Network for Video Object Detection
Shiyao Wang 0001, Yucong Zhou, Zhidong Deng
ECCV (13)4
2018 SegStereo: Exploiting Semantic Information for Disparity Estimation
Guorun Yang, Hengshuang Zhao, Jianping Shi, Zhidong Deng, Jiaya Jia
ECCV (7)4
2018 Masked Label Learning for Optical Flow Regression
abstract
Optical flow estimation is a challenging task in computer vision. Recent methods formulate such task as a supervised-learning problem. But it often suffers from limited realistic ground truth. In this paper, a compact network, embedded with cost volume, residual encoder and deconvolutional decoder, is presented to regress optical flow in an end- to-end manner. To overcome the lack of flow labels, we propose a novel data-driven strategy called masked label learning, where a large amount of masked labels are generated from the FlowNet 2.0 model and filtered by warping calibration for model training. We also present an extended-Huber loss to handle large displacements. With pretraining on massive masked flow data, followed by finetuning on a small number of sparse labels, our method achieves state-of-the-art accuracy on KITTI flow benchmark.
Guorun Yang, Zhidong Deng, Shiyao Wang 0001, Zeping Li
ICPR2
2018 Densely Connected CNN with Multi-scale Feature Attention for Text Classification
abstract
Text classification is a fundamental problem in natural language processing. As a popular deep learning model, convolutional neural network (CNN) has demonstrated great success in this task. However, most existing CNN models apply convolution filters of fixed window size, thereby unable to learn variable n-gram features flexibly. In this paper, we present a densely connected CNN with multi-scale feature attention for text classification. The dense connections build short-cut paths between upstream and downstream convolutional blocks, which enable the model to compose features of larger scale from those of smaller scale, and thus produce variable n-gram features. Furthermore, a multi-scale feature attention is developed to adaptively select multi-scale features for classification. Extensive experiments demonstrate that our model obtains competitive performance against state-of-the-art baselines on five benchmark datasets. Attention visualization further reveals the model's ability to select proper n-gram features for text classification.
Shiyao Wang 0001, Minlie Huang, Zhidong Deng
IJCAI3
2018 Semantic Image Segmentation Based on Attentions to Intra Scales and Inner Channels
abstract
Multi-scale features provide different context information of objects, which is significant for better performance in semantic segmentation tasks. But different scale features contribute equally to final predicitons. In this paper, we propose a new attention mechanism that not only learns weights between different scales but also allocates importance to subregions of inner channels. The network architecture is built on a state-of-the-art feedforward network to generate strong semantic feature maps, accompanying with a top-down pathway that incorporates larger scale feature maps through lateral connections. The proposed intra-scale attention module softly weights features between scales pixel by pixel. To enhance impact of each feature map in intermediate layers on performance, we further present an inner-channel attention module to pay attention to subregions within each channel. Moreover, an extra supervision is presented to achieve excellent performance. Importantly, the inner-channel attention module adaptively changes features as the layer goes deeper and could be inserted into any other layers. Extensive experiments are conducted on PASCAL VOC2012 to verify the network effectiveness. The experimental results show that both intra-scale and inner-channel attention modules could yield better performance.
Hongchao Lu, Zhidong Deng, Xiaolong Liu 0010
IJCNN2
2018 Detection and Recognition of Traffic Planar Objects Using Colorized Laser Scan and Perspective Distortion Rectification
abstract
Reliable detection and recognition of planar objects including traffic sign, street sign, and road surface in dynamic cluttered natural scenes are a big challenge for self-driving cars. In this paper, we propose a comprehensive method for planar object detection and recognition. First, the data association of LIDAR and camera is set up to acquire colorized laser scans, which simultaneously contain both color and geometrical information. Second, we combine three color spaces of RGB, HSV, and CIE L*a*b* with laser reflectivity as an aggregation-based feature vector. Third, the 3-D geometrical characteristics of planar objects that contain planarity, size, and aspect ratio are exploited to further reduce false alarm. Fourth, in order to increase robustness to any viewpoint variation, we present a new virtual camera-based rectification method to synthesize fronto-parallel views of refined object descriptors in 3-D space. Finally, experimental results achieved under a variety of challenging conditions show that integration of color space aggregation and laser reflectivity is superior to individuals. Specifically, the proposed perspective distortion rectification method remarkably eliminates false recognition error by 45.5%. Overall, the detection rate of our comprehensive method has up to 95.87% and the recognition rate even reaches 95.07% for traffic signs ranging within 100 m, with about 33.25 ms average running time per frame.
Zhidong Deng, Lipu Zhou
IEEE Trans. Intell. Transp. Syst.1
2017 End-to-End Disparity Estimation with Multi-granularity Fully Convolutional Network
Guorun Yang, Zhidong Deng
ICONIP (3)2
2017 Drivable Road Detection Based on Dilated FPN with Feature Aggregation
abstract
In the field of ADAS and self-driving car, lane and drivable road detection play an essential role in reliably accomplishing other tasks, such as objects detection. For monocular vision based semantic segmentation of lane and road, we propose a dilated feature pyramid network (FPN) with feature aggregation, called DFFA, where feature aggregation is employed to combine multi-level features enhanced with dilated convolution operations and FPN under the framework of ResNet. Experimental results validate effectiveness and efficiency of the proposed deep learning model for semantic segmentation of lane and drivable road. Our DFFA achieves the best performance both on Lane Estimation Evaluation and Behavior Evaluation tasks in KITTI-ROAD and take the second place on UU ROAD task.
Xiaolong Liu 0010, Zhidong Deng, Guorun Yang
ICTAI2
2017 Tightly-coupled convolutional neural network with spatial-temporal memory for text classification
abstract
Although several traditional models like bag of words (BOW), n-grams, and their variants of TFIDF exhibit high performance in the field of text classification, neural network methods such as LSTM, GRU and convolutional neural network (CNN) are recently attracting increasing attention. Considering that CNN has surprising capabilities of extracting hierarchical features, combination of LSTM/GRU with CNN seems to be quite reasonable for semantic representation and sequence analysis. On the other hand, it is also a promising subject to enable CNN to have memory embeddings and/or recurrent pathway. In this paper, we propose a novel tightly-coupled convolutional neural network with spatial-temporal memory (TCNN-SM). It comprises feature-representation and memory functional columns. Feature-representation functional column in our TCNN-SM actually performs hierarchical feature extraction as regular CNN does while memory functional column retains memories of different granularity and fulfills selective memory for historical information. In order to validate effectiveness and efficiency of the proposed TCNN-SM, we conduct extensive experiments on AG's News public dataset. The experimental results show that our new TCNN-SM achieves 7.99% test error, which has the best performance among other existing deep learning methods and is very close to state of the art results yielded using classical n-grams algorithm.
Shiyao Wang 0001, Zhidong Deng
IJCNN2
2016 Collaborative Learning Network for Face Attribute Prediction
Shiyao Wang 0001, Zhidong Deng, Zhenyang Wang
ACCV (3)2
2016 Stochastic Area Pooling for Generic Convolutional Neural Network
abstract
This paper proposes a novel SAPNet model that incorporates a stochastic area pooling (SAP) method with a generic stacked T-shaped CNN architecture. In our SAP method, pooling area is randomly transformed and max pooling operation is then conducted on such areas, which are no longer regular identical fixed upright squares. It can be viewed as feature-level augmentation, substantially reducing model parameters while keeping generalization ability of CNN almost unchanged. Furthermore, we present a generic CNN architecture that structurally resembles three stacked T-shaped cubes. In such architecture, the number of kernels in convolutional layer preceding any pooling layer is doubled and all learnable weight layers are combined with batch normalization and dropout with a small ratio. Finally, on CIFAR-10, CIFAR-100, MNIST, and SVHN datasets, the experimental results show that our SAPNet requires fewer parameters than regular CNN models and still achieves superior recognition performances for all the four benchmarks.
Zhidong Deng, Zhenyang Wang, Shiyao Wang 0001
ECAI1
2016 Accelerating Convolutional Neural Networks with Dominant Convolutional Kernel and Knowledge Pre-regression
Zhenyang Wang, Zhidong Deng, Shiyao Wang 0001
ECCV (8)2
2016 Deep self-organizing reservoir computing model for visual object recognition
abstract
Reservoir computing becomes increasingly a hot spot in recent years. In this paper, we propose a deep self-organizing reservoir computing model for visual object recognition. First, through combination of Kohonen's self-organizing map and SHESN network, we present a self-organizing SHESN (SO-SHESN). In the new model, we adopt the same mechanism of generating reservoir as SHESN, but McCulloch-Pitts type reservoir neuron is replaced with radial basis function neuron. Correspondingly, unsupervised competitive learning is exploited to train both input weights and reservoir weights of SO-SHESN. Second, we propose a deep SO-SHESN model through a stack of well-trained reservoir layers. In such a stacked structure, a novel trial-and-readout learning algorithm is used for pre-training of layer-wise reservoir, in which each layer is trained independently from each other. Finally, the experimental results obtained on MNIST benchmark dataset show that our SO-SHESN achieves the test recognition error rate of 5.66%, which improves classical ESN and SHESN by 6.44% and 1.74%, respectively. Furthermore, the test error rate of our deep SO-SHESN could reach up to 1.39%, which outperforms SO-SHESN with single reservoir layer by 4.27% and approximately approaches the state-of-the-art result of 1% among existing traditional machine learning approaches with non-CNN features.
Zhidong Deng, Chengzhi Mao
IJCNN1
2016 SAM: A rethinking of prominent convolutional neural network architectures for visual object recognition
abstract
Convolutional neural networks play an increasingly important role in computer vision tasks, especially in the field of visual object recognition. Many prominent models, such as Inception, Maxout, ResNet, and NIN, have been proposed to significantly improve recognition performance. Inspired from those models, we propose a novel module called self-adaptive module (SAM). SAM consists of four passes and one selector. Specifically, the four passes include two direct passes with different receptive fields and depths, one residual pass, and one Maxout pass. Actually, the residual pass is used to speed up convergence, while we take advantage of the Maxout pass to enhance approximate capabilities of SAM. The selector is further designed to help choose reasonable output. Basically, SAM is intended to simplify design of any new deep learning architecture, since it no longer requires consideration of how to select receptive fields and depths. Our SAM is tested on the visual object recognition datasets including CIFAR-10, CIFAR-100, MNIST, and SVHN. The experimental results demonstrate that the SAM-Net has superior recognition performances on the four benchmarks, which achieve test errors of 5.76%, 28.56%, 0.31%, and 1.98%, respectively.
Zhenyang Wang, Zhidong Deng, Shiyao Wang 0001
IJCNN2
2015 A Computational Model of Match Decision-Making Problem Using Spiking SHESN with Reward-Modulated Reinforcement Learning
Zhidong Deng, Guorun Yang
ICONIP (1)1
2014 A new fuzzy shape context approach based on multi-clue and state reservoir computing
abstract
This paper first builds a rule-based fuzzy representation of shape context and then present a multi-clue based fuzzy shape context approach (MFSC) using combination of geometric information and graph transduction. The MFSC takes complexity of object shape into account. In this approach, the distance between arbitrary two sampled points on any shape is redefined and graph transduction is used to correct and compensate training error. Furthermore, we propose a new fuzzy shape context approach based on both multi-clue and state reservoir computing. The experimental results show that the accuracy of detection achieved by our new approach on Kimia-216 and Kimia-99 datasets reaches up to 99.35% and 98.56%, respectively, which outperforms that of all the state-of-the-art shape context approaches.
Zhidong Deng, Kelaiti Xiao
IJCNN1
2014 Event-Related Complexity Analysis and its Application in the Detection of Facial attractiveness
abstract
In this study, an event-related complexity (ERC) analysis method is proposed and used to explore the neural correlates of facial attractiveness detection in the context of a cognitive experiment. The ERC method gives a quantitative index for measuring the diverse brain activation properties that represent the neural correlates of event-related responses. This analysis reveals distinct effects of facial attractiveness processing and also provides further information that could not have been achieved from event-related potential alone.
Zhidong Deng, Zimu Zhang
Int. J. Neural Syst.1
2013 Distributed self-learning scheduling approach for wireless sensor network
Jianjun Niu, Zhidong Deng
Ad Hoc Networks2
2012 Extrinsic calibration of a camera and a lidar based on decoupling the rotation from the translation
abstract
In this paper, we propose a novel robust algorithm for the extrinsic calibration of a camera and a lidar. This algorithm utilizes checkerboard as a calibration object. Since the interaction between the estimation errors of the plane parameters obtained from checkerboard images downgrades the quality of extrinsic calibration results, a new geometric constraint is presented to decouple the rotation from the translation so as to reduce the effect of such an interaction. Weights that represent uncertainty of the unit normal vector to the checkerboard plane are introduced to totally evaluate the quality of each pair of image and lidar scan. Furthermore, we analyze the configuration of checkerboard pose and give a formula that is used to assess the configuration. We compare the proposed algorithm with the previous ones. Simulation and experimental results show that our algorithm is able to achieve more accurate extrinsic parameters than the existing algorithms. Meanwhile, we also design experiments to validate the effectiveness and efficiency of the presented weight and the assessment formula.
Lipu Zhou, Zhidong Deng
Intelligent Vehicles Symposium2
2011 A stochastic policy search model for matching behavior
ZhenBo Cheng, Zhidong Deng
Sci. China Inf. Sci.3
2009 An approach for RNA secondary structure prediction based on Bayesian network
abstract
RNA secondary structure prediction is a fundamental problem in bioinformatics. This paper proposes a new approach to predict RNA secondary structure based on Bayesian network. Compared to the existing sophisticated prediction approaches such as Zuker's algorithm and the stochastic context-free grammar (SCFG) model, Bayesian network can naturally incorporate a priori knowledge from different models sources, and moreover, they have great expression capabilities. Our approach provides an effective method of combining free energy information of Zuker algorithm with statistical information from SCFG probability model. Basically, the proposed approach is suitable to all kinds of existing SCFG grammar models. Taking the BJK grammar model as an example, this paper gives a complete description of our prediction algorithm. When performing on RNA datasets with known structures, the experimental results show that the prediction accuracy is considerably improved. The sensitivity and the correlation coefficient are increased by 7.91% and 5.70%, respectively, compared to the SCFG approach alone.
Tianhua Wu, Zhidong Deng
CIBCB2
2009 Function of EEG Temporal Complexity Analysis in Neural Activities Measurement
Xiuquan Li, Zhidong Deng, Jianwei Zhang 0001
ISNN (1)2
2009 Computational prediction of novel non-coding RNAs in Arabidopsis thaliana
abstract
BACKGROUND: Non-coding RNA (ncRNA) genes do not encode proteins but produce functional RNA molecules that play crucial roles in many key biological processes. Recent genome-wide transcriptional profiling studies using tiling arrays in organisms such as human and Arabidopsis have revealed a great number of transcripts, a large portion of which have little or no capability to encode proteins. This unexpected finding suggests that the currently known repertoire of ncRNAs may only represent a small fraction of ncRNAs of the organisms. Thus, efficient and effective prediction of ncRNAs has become an important task in bioinformatics in recent years. Among the available computational methods, the comparative genomic approach seems to be the most powerful to detect ncRNAs. The recent completion of the sequencing of several major plant genomes has made the approach possible for plants. RESULTS: We have developed a pipeline to predict novel ncRNAs in the Arabidopsis (Arabidopsis thaliana) genome. It starts by comparing the expressed intergenic regions of Arabidopsis as provided in two whole-genome high-density oligo-probe arrays from the literature with the intergenic nucleotide sequences of all completely sequenced plant genomes including rice (Oryza sativa), poplar (Populus trichocarpa), grape (Vitis vinifera), and papaya (Carica papaya). By using multiple sequence alignment, a popular ncRNA prediction program (RNAz), wet-bench experimental validation, protein-coding potential analysis, and stringent screening against various ncRNA databases, the pipeline resulted in 16 families of novel ncRNAs (with a total of 21 ncRNAs). CONCLUSION: In this paper, we undertake a genome-wide search for novel ncRNAs in the genome of Arabidopsis by a comparative genomics approach. The identified novel ncRNAs are evolutionarily conserved between Arabidopsis and other recently sequenced plants, and may conduct interesting novel biological functions.
Yang Yang 0030, Binglian Zheng, Zhidong Deng, Bao-Liang Lu, Tao Jiang 0001
BMC Bioinform.5
2009 Trading strategy design in financial investment through a turning points prediction scheme
Xiuquan Li, Zhidong Deng
Expert Syst. Appl.2
2009 Corrigendum "Trading strategy design in financial investment through a turning points prediction scheme" [Experts Systems with Applications 36 (4) (2009) 7818-7826]
Xiuquan Li, Zhidong Deng
Expert Syst. Appl.2
2008 Stability Analysis of AQM Algorithm Based on Generalized Predictive Control
Xinghui Zhang, Kuan-sheng Zou, Zengqiang Chen 0001, Zhidong Deng
ICIC (1)4
2007 A Machine Learning Approach to Predict Turning Points for Chaotic Financial Time Series
abstract
In this paper, a novel approach to predict turning points for chaotic financial time series is proposed based on chaotic theory and machine learning. The nonlinear mapping between different data points in primitive time series is derived and proven. Our definition of turning points produces an event characterization function, which can transform the profile of time series to a measure. The RBF neural network is further used as a nonlinear modeler. We discuss the threshold selection and give a procedure for threshold estimation using out-of sample validation. The proposed approach is applied to the prediction problem of two real-world financial time series. The experimental results validate the effectiveness of our new approach.
Xiuquan Li, Zhidong Deng
ICTAI (2)2
2007 A BP-SCFG Based Approach for RNA Secondary Structure Prediction with Consecutive Bases Dependency and Their Relative Positions Information
Zhidong Deng
ISBRA2
2007 A fuzzy model of predicting RNA secondary structure
Zhidong Deng
Sci. China Ser. F Inf. Sci.2
2007 Collective Behavior of a Small-World Recurrent Neural System With Scale-Free Distribution
abstract
This paper proposes a scale-free highly clustered echo state network (SHESN). We designed the SHESN to include a naturally evolving state reservoir according to incremental growth rules that account for the following features: (1) short characteristic path length, (2) high clustering coefficient, (3) scale-free distribution, and (4) hierarchical and distributed architecture. This new state reservoir contains a large number of internal neurons that are sparsely interconnected in the form of domains. Each domain comprises one backbone neuron and a number of local neurons around this backbone. Such a natural and efficient recurrent neural system essentially interpolates between the completely regular Elman network and the completely random echo state network (ESN) proposed by Jaeger et al. We investigated the collective characteristics of the proposed complex network model. We also successfully applied it to challenging problems such as the Mackey-Glass (MG) dynamic system and the laser time-series prediction. Compared to the ESN, our experimental results show that the SHESN model has a significantly enhanced echo state property and better performance in approximating highly complex nonlinear dynamics. In a word, this large scale dynamic complex network reflects some natural characteristics of biological neural systems in many aspects such as power law, small-world property, and hierarchical architecture. It should have strong computing power, fast signal propagation speed, and coherent synchronization.
Zhidong Deng, Yi Zhang 0010
IEEE Trans. Neural Networks1
2006 Complex Systems Modeling Using Scale-Free Highly-Clustered Echo State Network
abstract
Inspired by the universal laws governing different kinds of complex networks, we propose a scale-free highly-clustered echo state network (SHESN). Different from echo state network (ESN), the state reservoir of the SHESN is generated by natural growth rules and eventually forms a complex network with small-world, scale-free properties, and hierarchically distributed structure. We implemented a large-scale SHESN with 3,000 internal neurons and applied it to modeling the pH-neutralization process. Simulation results showed the superior performance of SHESN. Furthermore, we analyzed the natural characteristics of the SHESN and discussed our growth rules and the new state reservoir from a brain functional network perspective.
Zhidong Deng, Yi Zhang 0010
IJCNN1
2006 A Realistic 3-D Reverse Modeling System Based on Real-World Sampling Dataset
abstract
This paper presents an image-based three-dimensioal (3-D) reverse modeling system. We take advantage of stereo vision-based method to acquire geometric information through sampling the surface of a physical object mounted on end effectors of a 4-DOF planar robot. Based on real-world sampling datasets, a realistic 3-D graphical model can be automatically constructed. We provide a new approach to system calibration, geometric information acquisition, and surface parameterization, and also implement a prototype system of rapidly establishing textured model that is available to virtual reality application.
Zhidong Deng, Jianjun Niu, Jingdan Zhang
IROS1
2006 A Fuzzy Dynamic Programming Approach to Predict RNA Secondary Structure
Zhidong Deng
WABI2
2003 Fuzzy neural control of systems with unknown dynamic using Q-learning strategies
abstract
In this paper an efficient Q-learning paradigm implemented on a fuzzy CMAC network is proposed. The fuzzy CMAC network topological architecture is described. The continuous states of the system are partitioned into a number of fuzzy boxes. With the proposed fuzzy CMAC the Q-values of agents in the fired fuzzy boxes are evaluated and the control actions with maximum Q-values can be derived. The proposed hybrid adaptive and learning type of Fuzzy Neural control system based on the Q-learning is applied to the control of a pH-neutralization process.
D. P. Kwok, Zhidong Deng, T. P. Leung, Zengqi Sun, J. C. K. Wong
FUZZ-IEEE2
2003 Correction and rectification of light fields
Lifeng Wang 0001, Zhouchen Lin, Zhidong Deng
Comput. Graph.5
2002 A New Algebraic Modelling Approach to Distributed Problem-Solving in MAS
Dianxun Shuai, Zhidong Deng
J. Comput. Sci. Technol.2
2001 A Dynamic Weight-Fuzzy Neural Network With Nonlinear Dynamic System Control
abstract
A class of fuzzy neural network with dynamic weights is proposed and its corresponding network topological architecture with suitable supervised learning algorithm is presented. Using the proposed network in control system design, a priori knowledge of the control system is not essential and this includes the order of the control system. The proposed network is applied to the control of a highly nonlinear pH-neutralization process. Simulation shows that the proposed dynamic learning control strategy has better dynamic quality, stronger robustness, adaptability and intelligence while comparing to the conventional control techniques, which demand the explicit and quantitative mathematical model of the system under control.
D. P. Kwok, T. P. Leung, Zhidong Deng, Zengqi Sun
FUZZ-IEEE4
2001 Linguistic Reward-orientated Takagi-Sugeno Fuzzy Reinforcement Learning
abstract
This paper presents a new learning method to attack two significant sub-problems in reinforcement learning at the same time: continuous space and linguistic rewards. Linguistic reward-oriented Takagi-Sugeno fuzzy reinforcement learning (LRTSFRL) is constructed by combining Q-learning with Takagi-Sugeno type fuzzy inference systems. The proposed paradigm is capable of solving complicated learning tasks of continuous domains, also can be used to design Takagi-Sugeno fuzzy logic controllers. Experiments on the double inverted pendulum system demonstrate the performance and applicability of the presented scheme. Finally, the conclusion remark is drawn.
Xionghuawei Yan, Zhidong Deng, Zengqi Sun
FUZZ-IEEE2
2001 Hyper-Distributed Hyper-Parallel Self-Organizing Dynamic Scheduling Based on Solitary Wave
Dianxun Shuai, Huiping Gu, Zhidong Deng
J. Comput. Sci. Technol.4
1996 A fuzzy neural network and its application to controls
Zengqi Sun, Zhidong Deng
Artif. Intell. Eng.2