EDBT 2026 Demo / reviewers in the wild / expert
Yaran Chen
dblp:189/4413
· DBLP profile ↗
34ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0001-9356-0610ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 4 first-author · 18 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EmoSENSE: Modeling Sentiment-Semantic Knowledge With Hierarchical Reinforcement Learning for Emotional Image GenerationabstractEmotional image generation aims to create images that effectively reflect target emotions. A fundamental challenge in this task is the affective gap, which refers to the discrepancy between visual content and emotional states perceived by users. Existing methods generally assume strong and explicit associations between target emotions and specific objects (e.g., “monster” and “fear”), which limits their generalization ability when encountering uncommon emotion-object pairs. This limitation stems from two main factors: 1) Most existing approaches primarily focus on semantic alignment without explicitly modeling how emotions influence visual attributes such as brightness and colorfulness; 2) diffusion-based image generation methods have limited capability in handling diverse sentiment-semantic pairs. To address these challenges, we propose EmoSENSE, a novel hierarchical fuzzy reinforcement learning framework for the emotional image generation task. EmoSENSE consists of a high-level module and a low-level module, working collaboratively in a hierarchical structure to inject sentiment-semantic knowledge into emotional images. The high-level module quantifies sentiment-semantic correlations within a unified emotional space, connecting emotions to visual attributes. The low-level module refines this connection by optimizing a fuzzy-logic-based mapping between emotions and visual attributes through reinforcement learning, enabling flexible adaptation to diverse emotion-object pairs. Extensive qualitative and quantitative experiments on public dataset demonstrate that EmoSENSE significantly enhances both the visual quality and emotional expression ability of the generated images, achieving a 12.21% higher EmoAccuracy-8 classes than the previous state-of-the-art methods.https://github.com/forever3600/EmoSENSE. Junyi Guo, Qiufeng Wang 0001, Yaran Chen, Fangyu Wu 0001, Eng Gee Lim |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Unsupervised Zero-Shot Reinforcement Learning via Dual-Value Forward-Backward RepresentationabstractOnline unsupervised reinforcement learning (URL) can discover diverse skills via reward-free pre-training and exhibits impressive downstream task adaptation abilities through further fine-tuning.
However, online URL methods face challenges in achieving zero-shot generalization, i.e., directly applying pre-trained policies to downstream tasks without additional planning or learning.
In this paper, we propose a novel Dual-Value Forward-Backward representation (DVFB) framework with a contrastive entropy intrinsic reward to achieve both zero-shot generalization and fine-tuning adaptation in online URL.
On the one hand, we demonstrate that poor exploration in forward-backward representations can lead to limited data diversity in online URL, impairing successor measures, and ultimately constraining generalization ability.
To address this issue, the DVFB framework learns successor measures through a skill value function while promoting data diversity through an exploration value function, thus enabling zero-shot generalization.
On the other hand, and somewhat surprisingly, by employing a straightforward dual-value fine-tuning scheme combined with a reward mapping technique, the pre-trained policy further enhances its performance through fine-tuning on downstream tasks, building on its zero-shot performance.
Through extensive multi-task generalization experiments, DVFB demonstrates both superior zero-shot generalization (outperforming on all 12 tasks) and fine-tuning adaptation (leading on 10 out of 12 tasks) abilities, surpassing state-of-the-art URL methods. Jingbo Sun 0001, Songjun Tu, Haoran Li 0010, Xin Liu 0039, Yaran Chen, Dongbin Zhao |
ICLR | 6 |
| 2025 | LeAffordNav: Enhancing Open-vocabulary Mobile Manipulation with LLM-guided Exploration and Affordance-aware NavigationabstractOpen-vocabulary mobile manipulation is a fundamental task for robotic assistants. However, Inefficient exploration and hand-off errors between different skills pose significant challenges to completing mobile manipulation tasks. In this paper, we propose a novel method named LeAffordNav, which is composed of LLM-guided exploration and Affordance-aware Navigation to address these challenges. LLM-guided exploration introduces LLMs to combine commonsense inference and frontier-based exploration, and achieves the balance between exploration and finding the target object. Considering the manipulability of the robot arm and the accessibility of the robot, we propose Affordance-aware Navigation which predicts the affordance of the mobile manipulation to reduce the hand-off errors between navigation and manipulation. Experiments on the HomeRobot benchmark show that LeAffordNav achieves new state-of-the-art performance, with a 20% higher success rate than the previous best. The code is available at https: //github.com/Cyuanwen/LeAffordNav. Yuanwen Chen, Haoran Li 0010, Yaran Chen, Dongbin Zhao |
ICME | 3 |
| 2025 | GAPartManip: A Large-Scale Part-Centric Dataset for Material-Agnostic Articulated Object ManipulationabstractEffectively manipulating articulated objects in household scenarios is a crucial step toward achieving general embodied artificial intelligence. Mainstream research in 3D vision has primarily focused on manipulation through depth perception and pose detection. However, in real-world environments, these methods often face challenges due to imperfect depth perception, such as with transparent lids and reflective handles. Moreover, they generally lack the diversity in partbased interactions required for flexible and adaptable manipulation. To address these challenges, we introduced a largescale part-centric dataset for articulated object manipulation that features both photo-realistic material randomizations and detailed annotations of part-oriented, scene-level actionable interaction poses. We evaluated the effectiveness of our dataset by integrating it with several state-of-the-art methods for depth estimation and interaction pose prediction. Additionally, we proposed a novel modular framework that delivers superior and robust performance for generalizable articulated object manipulation. Our extensive experiments demonstrate that our dataset significantly improves the performance of depth perception and actionable interaction pose prediction in both simulation and real-world scenarios. More information and demos can be found at: https://pku-epic.github.io/GAPartManip/. Wenbo Cui, Chengyang Zhao, Songlin Wei, Jiazhao Zhang, Yaran Chen, Haoran Li 0010, He Wang 0010 |
ICRA | 6 |
| 2025 | Sample-Efficient Unsupervised Policy Cloning from Ensemble Self-Supervised Labeled VideosabstractCurrent advanced policy learning methodologies have demonstrated the ability to develop expert-level strategies when provided enough information. However, their requirements, including task-specific rewards, action-labeled expert trajectories, and huge environmental interactions, can be expensive or even unavailable in many scenarios. In contrast, humans can efficiently acquire skills within a few trials and errors by imitating easily accessible internet videos, in the absence of any other supervision. In this paper, we try to let machines replicate this efficient watching-and-learning process through Unsupervised Policy from Ensemble Self-supervised labeled Videos (UPESV), a novel framework to efficiently learn policies from action-free videos without rewards and any other expert supervision. UPESV trains a video labeling model to infer the expert actions in expert videos through several organically combined self-supervised tasks. Each task performs its duties, and they together enable the model to make full use of both action-free videos and reward-free interactions for robust dynamics understanding and advanced action prediction. Simultaneously, UPESV clones a policy from the labeled expert videos, in turn collecting environmental interactions for self-supervised tasks. After a sample-efficient, unsupervised, and iterative training process, UPESV obtains an advanced policy based on a robust video labeling model. Extensive experiments in sixteen challenging procedurally generated environments demonstrate that the proposed UPESV achieves state-of-the-art interaction-limited policy learning performance (outperforming five current advanced baselines on 12/16 tasks) without exposure to any other supervision except for videos. Xin Liu 0039, Yaran Chen, Haoran Li 0010 |
ICRA | 2 |
| 2025 | Robotic Sim-to-Real Transfer for Long-Horizon Pick-and-Place Tasks in the Robotic Sim2Real CompetitionabstractThis paper presents a fully autonomous robotic system that performs sim-to-real transfer in complex longhorizon tasks involving navigation, recognition, grasping, and stacking in an environment with multiple obstacles. The key feature of the system is the ability to overcome typical sensing and actuation discrepancies during sim-to-real transfer and to achieve consistent performance without any algorithmic modifications. To accomplish this, a lightweight noise-resistant visual perception system and a nonlinearityrobust servo system are adopted. We conduct a series of tests in both simulated and realworld environments. The visual perception system achieves the speed of 11 ms per frame due to its lightweight nature, and the servo system achieves sub-centimeter accuracy with the proposed controller. Both exhibit high consistency during sim-to-real transfer. Benefiting from these, our robotic system took first place in the mineral searching task of the Robotic Sim2Real Challenge hosted at ICRA 2024. Hongyu Cao, Lixuan Zhao, Yaran Chen |
ICRA | 5 |
| 2025 | Advancing Object-Goal Navigation through LLM-enhanced Object Affinities TransferabstractObject-goal navigation requires mobile robots to efficiently locate targets with visual and spatial information, yet existing methods struggle with generalization in unseen environments. Heuristic approaches with naive metrics fail in complex layouts, while graph-based and learning-based methods suffer from environmental biases and limited generalization. Although Large Language Models (LLMs) as planners or agents offer a rich knowledge base, they are cost-inefficient and lack targeted historical experience. To address these challenges, we propose the LLM-enhanced Object Affinities Transfer (LOAT) framework, integrating LLM-derived semantics with learning-based approaches to leverage experiential object affinities for better generalization in unseen settings. LOAT employs a dual-module strategy: one module accesses LLMs’ vast knowledge, and the other applies learned object semantic relationships, dynamically fusing these sources based on context. Evaluations in AI2-THOR and Habitat simulators show significant improvements in navigation success and efficiency, and real-world deployment demonstrates the zero-shot ability of LOAT to enhance object-goal navigation systems. Mengying Lin, Shugao Liu, Dingxi Zhang, Yaran Chen, Zhaoran Wang 0001, Haoran Li 0010, Dongbin Zhao |
IROS | 4 |
| 2025 | Adaptive search for broad attention based vision transformers
Nannan Li 0003, Yaran Chen, Dongbin Zhao |
Neurocomputing | 2 |
| 2025 | Balancing State Exploration and Skill Diversity in Unsupervised Skill DiscoveryabstractUnsupervised skill discovery seeks to acquire different useful skills without extrinsic reward via unsupervised reinforcement learning (RL), with the discovered skills efficiently adapting to multiple downstream tasks in various ways. However, recent advanced skill discovery methods struggle to well balance state exploration and skill diversity, particularly when the potential skills are rich and hard to discern. In this article, we propose contrastive dynamic skill discovery (ComSD) which generates diverse and exploratory unsupervised skills through a novel intrinsic incentive, named contrastive dynamic reward. It contains a particle-based exploration reward to make agents access far-reaching states for exploratory skill acquisition, and a novel contrastive diversity reward to promote the discriminability between different skills. Moreover, a novel dynamic weighting mechanism between the above two rewards is proposed to balance state exploration and skill diversity, which further enhances the quality of the discovered skills. Extensive experiments and analysis demonstrate that ComSD can generate diverse behaviors at different exploratory levels for multijoint robots, enabling state-of-the-art adaptation performance on challenging downstream tasks. It can also discover distinguishable and far-reaching exploration skills in the challenging tree-like 2-D maze. Xin Liu 0039, Yaran Chen, Guixing Chen, Haoran Li 0010, Dongbin Zhao |
IEEE Trans. Cybern. | 2 |
| 2025 | Cross-Domain Random Pretraining With Prototypes for Reinforcement LearningabstractUnsupervised cross-domain reinforcement learning (RL) pretraining shows great potential for challenging continuous visual control but poses a big challenge. In this article, we propose cross-domain random pretraining with prototypes (CRPTpro), a novel, efficient, and effective self-supervised cross-domain RL pretraining framework. CRPTpro decouples data sampling from encoder pretraining, proposing decoupled random collection to easily and quickly generate a qualified cross-domain pretraining dataset. Moreover, a novel prototypical self-supervised algorithm is proposed to pretrain an effective visual encoder that is generic across different domains. Without finetuning, the cross-domain encoder can be implemented for challenging downstream tasks defined in different domains, either seen or unseen. Compared with recent advanced methods, CRPTpro achieves better performance on downstream policy learning without extra training on exploration agents for data collection, greatly reducing the burden of pretraining. We conduct extensive experiments across multiple challenging continuous visual-control domains, including balance control, robot locomotion, and manipulation. CRPTpro significantly outperforms the next best Proto-RL(C) on 11/12 cross-domain downstream tasks with only 54.5% wall-clock pretraining time, exhibiting state-of-the-art pretraining performance with greatly improved pretraining efficiency. Xin Liu 0039, Yaran Chen, Haoran Li 0010, Boyu Li 0003, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | Common Sense Language-Guided Exploration and Hierarchical Dense Perception for Instruction Following Embodied AgentsabstractEmbodied Instruction Following (EIF) involves the task of locating and manipulating objects according to language instructions. Existing methods face challenges in small object navigation due to ineffective exploration and imperfect perception, which ultimately affects their performance. This study focuses on small object navigation in the EIF domain. We propose Common Sense Language-guided exploration (CSL), a novel approach that leverages common-sense knowledge from seen scenes and information from language instructions to infer the location of objects. The proposed CSL significantly improves exploration efficiency. Additionally, we propose Hierarchical Dense Perception (HDP), which uses hierarchical features to perform semantic segmentation and depth estimation. The use of HDP significantly improves the agent’s perceptual capabilities. Experiments on the ALFRED benchmark demonstrate the effectiveness of CSL-HDP. The proposed CSL-HDP achieves an absolute improvement of 9.29% (18.45% relative) on unseen test scenes compared to the previous state-of-the-art, securing the top position on the leaderboard. Code will be available at https://github.com/Cyuanwen/CSL-HDP. Yuanwen Chen, Yaran Chen, Dongbin Zhao, Yunzhen Zhao, Pengfei Hu 0004 |
ICME | 3 |
| 2024 | ATV3D: 3D Object Detection from Attention-based Three-view RepresentationabstractIn the fields of autonomous driving and robot perception, the majority of methods are designed for onboard camera object detection, while there are fewer methods specifically tailored to environmental cameras. However, environmental cameras have the capability to capture a significant amount of road geometry and vehicle position information, which can enhance the safety of autonomous driving. Nevertheless, there is a difference in perspective between environmental cameras and onboard cameras, resulting in poorer performance of many methods designed for onboard camera 3D object detection when apply to environmental camera. In this paper, we propose a 3D Object Detection Algorithm from Attention-based Three-view Representation (ATV3D). The algorithm projects the 2D image features onto three orthogonal views (left view, front view, bird’s eye view) to achieve a representation of the 3D information. Compared to voxel-based 3D detection methods, our proposed approach retains the ability to capture 3D features while reducing computational complexity. During the process of three-view representation, we design a feature projection module based on attention. Unlike inverse perspective mapping that requires precise camera parameters, the attention can implicitly learn the mapping relationship from 2D images to the three-view planes. This enables the extraction and transformation of image features without the calibrated camera parameters, effectively addressing challenges associated with obtaining camera parameters for environmental cameras and their susceptibility to natural factors. The experimental results on the DAIR-V2X dataset demonstrate that our method achieves a 3D detection mean average precision (mAP) of 73.6%, surpassing the performance of previous calibration-free environmental camera methods. Furthermore, our method achieves the highest detection accuracy on the indoor multi-view robot dataset Neurons Perception, providing evidence of its outstanding detection performance. Yaran Chen, Haoran Li 0010, Yunzhen Zhao, Pengfei Hu 0004 |
IJCNN | 2 |
| 2024 | BViT: Broad Attention-Based Vision TransformerabstractRecent works have demonstrated that transformer can achieve promising performance in computer vision, by exploiting the relationship among image patches with self-attention. They only consider the attention in a single feature layer, but ignore the complementarity of attention in different layers. In this article, we propose broad attention to improve the performance by incorporating the attention relationship of different layers for vision transformer (ViT), which is called BViT. The broad attention is implemented by broad connection and parameter-free attention. Broad connection of each transformer layer promotes the transmission and integration of information for BViT. Without introducing additional trainable parameters, parameter-free attention jointly focuses on the already available attention information in different layers for extracting useful information and building their relationship. Experiments on image classification tasks demonstrate that BViT delivers superior accuracy of 75.0%/81.6% top-1 accuracy on ImageNet with 5M/22M parameters. Moreover, we transfer BViT to downstream object recognition benchmarks to achieve 98.9% and 89.9% on CIFAR10 and CIFAR100, respectively, that exceed ViT with fewer parameters. For the generalization test, the broad attention in Swin Transformer, T2T-ViT and LVT also brings an improvement of more than 1%. To sum up, broad attention is promising to promote the performance of attention-based models. Code and pretrained models are available at https://github.com/DRL/BViT. Nannan Li 0003, Yaran Chen, Weifan Li, Zixiang Ding, Dongbin Zhao, Shuai Nie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Dense Attention: A Densely Connected Attention Mechanism for Vision TransformerabstractRecently, Vision Transformer has demonstrated its impressive capability in image understanding. The multi-head self-attention mechanism is fundamental to its formidable performance. However, self-attention has the drawback of high computational effort, which makes the training of the model require powerful computational resources or more time. This paper designs a novel and efficient attention mechanism Dense Attention to overcome the above problem. Dense attention aims to focus on features from multiple views through a dense connection paradigm. Benefiting from the attention of comprehensive features, dense attention can i) remarkably strengthen the image representation of the model, and ii) partially replace the multi-head self-attention mechanism to allow model slimming. To verify the effectiveness of dense attention, we implement it in the prevalent Vision Transformer models, including non-pyramid architecture DeiT and pyramid architecture Swin Transformer. The experimental results on ImageNet classification show that dense attention indeed contributes to performance improvement,$+\mathbf{1.8}/\mathbf{1.3}\%$for DeiT-T/s and$+\mathbf{0.7}/\!\!+\mathbf{1.2}\%$for Swin-T/s, respectively. Dense attention also demonstrates its transferability on CIFAR10 and CIFAR100 recognition benchmarks with classification accuracy of 98.9% and 89.6% respectively. Furthermore, dense attention can weaken the performance sacrifice caused by the pruning in the number of heads. Code and pre-trained models will be available11https://github.com/koala719/Dense-ViT. Nannan Li 0003, Yaran Chen, Dongbin Zhao |
IJCNN | 2 |
| 2023 | Stacked BNAS: Rethinking Broad Convolutional Neural Network for Neural Architecture SearchabstractDifferent from other deep scalable architecture-based neural architecture search (NAS) approaches, broad NAS (BNAS) proposes a broad scalable architecture which consists of convolution and enhancement blocks, dubbed broad convolutional neural network (BCNN), as the search space for amazing efficiency improvement. BCNN reuses the topologies of cells in the convolution block so that BNAS can employ few cells for efficient search. Moreover, multiscale feature fusion and knowledge embedding are proposed to improve the performance of BCNN with shallow topology. However, BNAS suffers some drawbacks: 1) insufficient representation diversity for feature fusion and enhancement and 2) time consumption of knowledge embedding design by human experts. This article proposes Stacked BNAS, whose search space is a developed broad scalable architecture named Stacked BCNN, with better performance than BNAS. On the one hand, Stacked BCNN treats mini BCNN as a basic block to preserve comprehensive representation and deliver powerful feature extraction ability. For multiscale feature enhancement, each mini BCNN feeds the outputs of deep and broad cells to the enhancement cell. For multiscale feature fusion, each mini BCNN feeds the outputs of deep, broad and enhancement cells to the output node. On the other hand, knowledge embedding search (KES) is proposed to learn appropriate knowledge embeddings in a differentiable way. Moreover, the basic unit of KES is an over-parameterized knowledge embedding module that consists of all possible candidate knowledge embeddings. Experimental results show that: 1) Stacked BNAS obtains better performance than BNAS-v2 on both CIFAR-10 and ImageNet; 2) the proposed KES algorithm contributes to reducing the parameters of the learned architecture with satisfactory performance; and 3) Stacked BNAS delivers a state-of-the-art efficiency of 0.02 GPU days. Zixiang Ding, Yaran Chen, Nannan Li 0003, Dongbin Zhao, C. L. Philip Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Neurons Perception Dataset for RoboMaster AI ChallengeabstractFrom virtual game to physical robot, games have witnessed the development of artificial intelligence (AI) technology, especially the data-driven technology represented by deep learning. Compared with virtual games, a physical robot game such as RoboMaster AI challenge needs to build a complete closed-loop architecture composed of perception, planning, control, and decision-making to support autonomous confrontation. Perception, as the eye of the robot, its performance in the complex environment depends on a massive dataset. Although there are many open perception datasets, these datasets are difficult to meet the needs of RoboMaster AI challenge due to the high dynamics of the task, the distinctiveness of the objects, and limited computing resources. In this paper, we release a dataset named Neurons11Neurons is a team dedicated to promoting the development of robot with deep neural network. We will release the code and dataset at https://github.com/DRL-CASIA/NeuronsDataset. perception dataset for RoboMaster AI challenge, which covers 3 tasks including monocular depth estimation, lightweight object detection, and multi-view 3D object detection, and makes up the data blank in this field. In addition, we also evaluate State-Of-The-Art (SOTA) methods on each task, hoping to provide an impartial benchmark for the development of perception algorithm. Haoran Li 0010, Zicheng Duan, Yaran Chen, Dongbin Zhao |
IJCNN | 5 |
| 2022 | ModuleNet: Knowledge-Inherited Neural Architecture SearchabstractAlthough neural the architecture search (NAS) can bring improvement to deep models, it always neglects precious knowledge of existing models. The computation and time costing property in NAS also means that we should not start from scratch to search, but make every attempt to reuse the existing knowledge. In this article, we discuss what kind of knowledge in a model can and should be used for a new architecture design. Then, we propose a new NAS algorithm, namely, ModuleNet, which can fully inherit knowledge from the existing convolutional neural networks. To make full use of the existing models, we decompose existing models into different modules, which also keep their weights, consisting of a knowledge base. Then, we sample and search for a new architecture according to the knowledge base. Unlike previous search algorithms, and benefiting from inherited knowledge, our method is able to directly search for architectures in the macrospace by the NSGA-II algorithm without tuning parameters in these modules. Experiments show that our strategy can efficiently evaluate the performance of a new architecture even without tuning weights in convolutional layers. With the help of knowledge we inherited, our search results can always achieve better performance on various datasets (CIFAR10, CIFAR100, and ImageNet) over original architectures. Yaran Chen, Ruiyuan Gao 0001, Fenggang Liu, Dongbin Zhao |
IEEE Trans. Cybern. | 1 |
| 2022 | BiFNet: Bidirectional Fusion Network for Road SegmentationabstractMultisensor fusion-based road segmentation plays an important role in the intelligent driving system since it provides a drivable area. The existing mainstream fusion method is mainly to feature fusion in the image space domain which causes the perspective compression of the road and damages the performance of the distant road. Considering the bird's eye views (BEVs) of the LiDAR remains the space structure in the horizontal plane, this article proposes a bidirectional fusion network (BiFNet) to fuse the image and BEV of the point cloud. The network consists of two modules: 1) the dense space transformation (DST) module, which solves the mutual conversion between the camera image space and BEV space and 2) the context-based feature fusion module, which fuses the different sensors information based on the scenes from corresponding features. This method has achieved competitive results on the KITTI dataset. Haoran Li 0010, Yaran Chen, Dongbin Zhao |
IEEE Trans. Cybern. | 2 |
| 2022 | Boost 3-D Object Detection via Point Clouds Segmentation and Fused 3-D GIoU-L₁ LossabstractThe 3-D object detection is crucial for many real-world applications, attracting many researchers’ attention. Beyond 2-D object detection, 3-D object detection usually needs to extract appearance, depth, position, and orientation information from light detection and ranging (LiDAR) and camera sensors. However, due to more degrees of freedom and vertices, existing detection methods that directly transform from 2-D to 3-D still face several challenges, such as exploding increase of anchors’ number and inefficient or hard-to-optimize objective. To this end, we present a fast segmentation method for 3-D point clouds to reduce anchors, which can largely decrease the computing cost. Moreover, taking advantage of 3-D generalized Intersection of Union (GIoU) and$L_{1}$losses, we propose a fused loss to facilitate the optimization of 3-D object detection. A series of experiments show that the proposed method has alleviated the abovementioned issues effectively. Yaran Chen, Haoran Li 0010, Ruiyuan Gao 0001, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | BNAS: Efficient Neural Architecture Search Using Broad Scalable ArchitectureabstractEfficient neural architecture search (ENAS) achieves novel efficiency for learning architecture with high-performance via parameter sharing and reinforcement learning (RL). In the phase of architecture search, ENAS employs deep scalable architecture as search space whose training process consumes most of the search cost. Moreover, time-consuming model training is proportional to the depth of deep scalable architecture. Through experiments using ENAS on CIFAR-10, we find that layer reduction of scalable architecture is an effective way to accelerate the search process of ENAS but suffers from a prohibitive performance drop in the phase of architecture estimation. In this article, we propose a broad neural architecture search (BNAS) where we elaborately design broad scalable architecture dubbed broad convolutional neural network (BCNN) to solve the above issue. On the one hand, the proposed broad scalable architecture has fast training speed due to its shallow topology. Moreover, we also adopt RL and parameter sharing used in ENAS as the optimization strategy of BNAS. Hence, the proposed approach can achieve higher search efficiency. On the other hand, the broad scalable architecture extracts multi-scale features and enhancement representations, and feeds them into global average pooling (GAP) layer to yield more reasonable and comprehensive representations. Therefore, the performance of broad scalable architecture can be promised. In particular, we also develop two variants for BNAS that modify the topology of BCNN. In order to verify the effectiveness of BNAS, several experiments are performed and experimental results show that 1) BNAS delivers 0.19 days which is 2.37× less expensive than ENAS who ranks the best in RL-based NAS approaches; 2) compared with small-size (0.5 million parameters) and medium-size (1.1 million parameters) models, the architecture learned by BNAS obtains state-of-the-art performance (3.58% and 3.24% test error) on CIFAR-10; and 3) the learned architecture achieves 25.3% top-1 error on ImageNet just using 3.9 million parameters. Zixiang Ding, Yaran Chen, Nannan Li 0003, Dongbin Zhao, Zhiquan Sun, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | BNAS-v2: Memory-Efficient and Performance-Collapse-Prevented Broad Neural Architecture SearchabstractIn this article, we propose BNAS-v2 to further improve the efficiency of broad neural architecture search (BNAS), which employs a broad convolutional neural network (BCNN) as the search space. In BNAS, the single-path sampling-updating strategy of an overparameterized BCNN leads to terrible unfair training issue, which restricts the efficiency improvement. To mitigate the unfair training issue, we employ a continuous relaxation strategy to optimize all paths of the overparameterized BCNN simultaneously. However, continuous relaxation leads to a performance collapse issue that leads to the unsatisfactory performance of the learned BCNN. For that, we propose the confident learning rate (CLR) and introduce the combination of partial channel connections and edge normalization. Experimental results show that 1) BNAS-v2 delivers state-of-the-art search efficiency on both CIFAR-10 (0.05 GPU days, which is$4\times $faster than BNAS) and ImageNet (0.19 GPU days) with better or competitive performance; 2) the above two solutions are effectively alleviating the performance collapse issue; and 3) BNAS-v2 achieves powerful generalization ability on multiple transfer tasks, e.g., MNIST, FashionMNIST, NORB, and SVHN. The code is available athttps://github.com/zixiangding/BNASv2. Zixiang Ding, Yaran Chen, Nannan Li 0003, Dongbin Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | IA-CNN: A generalised interpretable convolutional neural network with attention mechanismabstractIn recent years, convolutional neural network (CNN) has been widely used in security, autonomous driving, and healthcare. Even though CNN has achieved a great performance, the results produced by CNN are difficult to explain and sometimes irresponsible. The black-box nature of CNN makes it lack trust. In this paper, we propose an attention based CNN structure, named IA -CNN, which highly improves the interpretability of the CNN models. Each feature map of the last conv-layer only has one response (one key point) of the target object, which is directly connected to the output. We also combine the attention mechanism to weakly supervise the last conv-layer. In this way, our model can clearly show that which features the model extracted are the keys to the output prediction. Meanwhile, our IA-CNN structure can be used in various classical models with higher performance in the fine-grained classification and comparative performance in the ordinary classification task. Note that our IA-CNN structure is an end-to-end model, the last conv-layer of which can extract key points from images automatically and is connected to the output prediction linearly. Zhisong Zhang, Yaran Chen, Haoran Li 0010 |
IJCNN | 2 |
| 2021 | MGRL: Graph neural network based inference in a Markov network with reinforcement learning for visual navigation
Yaran Chen, Dongbin Zhao, Dong Li 0016 |
Neurocomputing | 2 |
| 2020 | Shift-Invariant Convolutional Network SearchabstractThe development of Neural Architecture Search (NAS) makes Convolutional Neural Networks (CNN) more diverse and effective. But previous NAS approaches don't pay attention to the shift-invariant of CNN. Without the shift-invariant, convolutional network is not robust enough when input data is disturbed or damaged. Besides, taking accuracy as the only optimization goal of NAS cannot meet the increasingly diverse needs. In this paper, we propose Shift-Invariant Convolutional Network Search (SICNS). It uses one-shot NAS to search for shift-invariant convolutional network by incorporating the low-pass filter into the one-shot model. Furthermore, SICNS optimizes multiple indicators simultaneously through the multi-objective evolutionary algorithm. Through training one-shot model and evolving the architecture, we obtain convolutional networks which are robust and powerful on image classification task. Especially, our work can achieve 4.52% test error on CIFAR-10 with 0.7M parameters. And in case the input data are disturbed, the accuracy of searched network is 2.96% higher than network without low-pass filter. Nannan Li 0003, Yaran Chen, Zixiang Ding, Dongbin Zhao |
IJCNN | 2 |
| 2020 | RailNet: An Information Aggregation Network for Rail Track SegmentationabstractAs the basis of scenes understanding for the track inspection task, track segmentation is challenging due to the various illumination conditions, track crossing, and plant coverage. Since the rail has a strong shape prior, strict rail spacing and special distribution in the image, making full use of the spatial information of the rail features becomes an important factor to improve the accuracy of rail segmentation. In this paper, an information aggregation module is proposed to enhance the spatial relationship between pixels of the rail features. In other words, this module expands the receptive field. Furthermore, we build an information aggregation network based on this module, which is called as RailNet. Finally, the RailNet is evaluated in an open train track dataset. Experimental results show that RailNet can achieve the best performance so far in the dataset of trains. Haoran Li 0010, Dongbin Zhao, Yaran Chen |
IJCNN | 4 |
| 2020 | ContourRend: A Segmentation Method for Improving Contours by Rendering
Yaran Chen, Dongbin Zhao, Zhong-Hua Pang |
ISNN | 3 |
| 2019 | Lane Change Decision-making through Deep Reinforcement Learning with Rule-based ConstraintsabstractAutonomous driving decision-making is a great challenge due to the complexity and uncertainty of the traffic environment. Combined with the rule-based constraints, a Deep Q-Network (DQN) based method is applied for autonomous driving lane change decision-making task in this study. Through the combination of high-level lateral decision-making and low-level rule-based trajectory modification, a safe and efficient lane change behavior can be achieved. With the setting of our state representation and reward function, the trained agent is able to take appropriate actions in a real-world-like simulator. The generated policy is evaluated on the simulator for 10 times, and the results demonstrate that the proposed rule-based DQN method outperforms the rule-based approach and the DQN method. Dongbin Zhao, Yaran Chen |
IJCNN | 4 |
| 2019 | Graph-FCN for Image Semantic Segmentation
Yaran Chen, Dongbin Zhao |
ISNN (1) | 2 |
| 2019 | Deep Kalman Filter with Optical Flow for Multiple Object TrackingabstractDeep matching and Kalman filter-based multiple object tracking (DK-tracking) have been demonstrated to be promising. However, most of existing DK-tracking trackers assume that objects are slow-varying movement with a constant velocity. The assumption is hard to be satisfied in the real world, especially in the image space due to the sight distance. In this paper, we propose a novel multiple object tracking method combining deep feature matching, Kalman filter and flow information, which is called DK-flow-tracking, to improve tracking performance. In DK-flow-tracking, optical flow in consecutive frames is used to provide accurate object motion information for guiding Kalman filter to track objects. Experiments are performed on public datasets: MOT2016, MOT2017, and the proposed method achieves better performances compared to the DK-tracking with the assumption of a constant velocity movement. Yaran Chen, Dongbin Zhao, Haoran Li 0010 |
SMC | 1 |
| 2018 | A temporal-based deep learning method for multiple objects detection in autonomous drivingabstractThis paper proposes a novel vision-based object detection method in autonomous driving, which introduces the temporal information into the deep learning-based detection method for moving object detection. Vision-based object detection is a critical technology for autonomous driving. The objects in the real world such as driving cars, don't have great changes in their positions and velocities. So the position change of objects between two consecutive frames is not large. This is usually ignored by traditional works, which usually use object detection methods on still-images to detect moving objects. Considering the relationship among consecutive frames (temporal information), we present a robust and real-time tracking method following image detection to refine the object detection results. Based on the three key attributes (distances, sizes and positions), the tracking method aims to build the association between the detected objects on the current frame and those in previous frames. The proposed object detection with temporal information dramatically improves the performance of existing object detection algorithms based on stillimage. With the proposed method, we won the champion in the preceding vehicle detection task in 2017 intelligent vehicle future challenge(2017 IVFC)1. Yaran Chen, Dongbin Zhao, Haoran Li 0010, Dong Li 0016, Ping Guo 0002 |
IJCNN | 1 |
| 2018 | DeepSign: Deep Learning based Traffic Sign RecognitionabstractThis paper investigates the traffic sign recognition task with deep learning methods. The proposed algorithm which is called DeepSign includes three modules: a detection module (PosNet) for locating the traffic sign in a static image, a classification module (PatchNet) for classifying the detected image patch, and a temporal filter for correcting the recognition results. The PosNet is a binary object detection convolution neural network which regards all traffic signs as one class and the background as the other class. Different from the traditional works which recognize the traffic sign on the static image, the proposed temporal filter exploits the contextual information to recover the missed detection region and correct the false classification. The experiments validate the effectiveness of the proposed algorithm. It achieved the third place on the traffic sign recognition task in 2017 China intelligent vehicle future challenge (2017 CIVFC). Dong Li 0016, Dongbin Zhao, Yaran Chen |
IJCNN | 3 |
| 2018 | Multi-task learning for dangerous object detection in autonomous driving
Yaran Chen, Dongbin Zhao, Le Lv |
Inf. Sci. | 1 |
| 2017 | Multi-task Learning with Cartesian Product-Based Multi-objective Combination for Dangerous Object Detection
Yaran Chen, Dongbin Zhao |
ISNN (1) | 1 |
| 2016 | Convolutional fitted Q iteration for vision-based control problemsabstractIn this paper a deep reinforcement learning (DRL) method is proposed to solve the control problem which takes raw image pixels as input states. A convolutional neural network (CNN) is used to approximate Q functions, termed as Q-CNN. A pretrained network, which is the result of a classification challenge on a vast set of natural images, initializes the parameters of Q-CNN. Such initialization assigns Q-CNN with the features of image representation, so it is more concentrated on the control tasks. The weights are tuned under the scheme of fitted Q iteration (FQI), which is an offline reinforcement learning method with the stable convergence property. To demonstrate the performance, a modified Food-Poison problem is simulated. The agent determines its movements based on its forward view. In the end the algorithm successfully learns a satisfied policy which has better performance than the results of previous researches. Dongbin Zhao, Yuanheng Zhu, Le Lv, Yaran Chen |
IJCNN | 4 |