EDBT 2026 Demo / reviewers in the wild / expert
Tianhong Dai
dblp:177/5221
· DBLP profile ↗
25ranked-venue papers
4as first author
23since 2021 · last 2025
0000-0001-8904-1551ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | POT: Prototypical Optimal Transport for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) leverages Class Activation Maps (CAMs) to extract spatial information from image-level labels. However, CAMs primarily highlight the most discriminative foreground regions, leading to incomplete results. Prototype-based methods attempt to address this limitation by employing prototype CAMs instead of classifier CAMs. Nevertheless, existing prototype-based methods typically use a single prototype for each class, which is insufficient to capture all attributes of the foreground features due to the significant intra-class variations across different images. Consequently, these methods still struggle with incomplete CAM predictions. In this paper, we propose a novel framework called Prototypical Optimal Transport (POT) for WSSS. POT enhances CAM predictions by dividing features into multiple clusters and activating each cluster using its prototype. In this process, a similarity-aware optimal transport is employed to assign features to the most probable clusters. This similarity-aware strategy ensures the prioritization of significant cluster prototypes, thereby improving the accuracy of feature assignment. Additionally, we introduce an adaptive OT-based consistency loss to refine feature representations. This framework effectively overcomes the limitations of single-prototype methods, providing more complete and accurate CAM predictions. Extensive experimental results on standard WSSS benchmarks (PASCAL VOC and MS COCO) demonstrate that our method significantly improves the quality of CAMs and achieves state-of-the-art performances. The source code will be released https://github.com/jianwang91/POT. Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
CVPR | 2 |
| 2025 | Enhancing Prompt Generation with Adaptive Refinement for Camouflaged Object Detection
Guangyu Ren, Tianhong Dai, Tania Stathaki, Hengyan Liu |
ICCV | 3 |
| 2025 | Class Token as Proxy: Optimal Transport-Assisted Proxy Learning for Weakly Supervised Semantic Segmentation
Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
ICCV | 2 |
| 2025 | Unifying Reconstruction and Density Estimation via Invertible Contraction Mapping in One-Class ClassificationabstractDue to the difficulty in collecting all unexpected abnormal patterns, One-Class Classification (OCC) has become the most popular approach to anomaly detection (AD). Reconstruction-based AD method relies on the discrepancy between inputs and reconstructed results to identify unobserved anomalies. However, recent methods trained only on normal samples may generalize to certain abnormal inputs, leading to well-reconstructed anomalies and degraded performance. To address this, we constrain reconstructions to remain on the normal manifold using a novel AD framework based on contraction mapping. This mapping guarantees that any input converges to a fixed point through iterations of this mapping. Based on this property, training the contraction mapping using only normal data ensures that its fixed point lies within the normal manifold. As a result, abnormal inputs are iteratively transformed toward the normal manifold, increasing the reconstruction error. In addition, the inherent invertibility of contraction mapping enables flow-based density estimation, where a prior distribution learned from the previous reconstruction is used to estimate the input likelihood for anomaly detection, further improving the performance. Using both mechanisms, we propose a bidirectional structure with forward reconstruction and backward density estimation. Extensive experiments on tabular data, natural image, and industrial image data demonstrate the effectiveness of our method. The code is available at URD. Tianhong Dai, Huihui Bai 0001, Yao Zhao 0001, Jimin Xiao |
NeurIPS | 2 |
| 2025 | Adaptive Patch Contrast for Weakly Supervised Semantic Segmentation
Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Jimin Xiao, Fei Ma 0002, Renrong Ouyang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Generative Prompt Controlled Diffusion for weakly supervised semantic segmentationabstractWeakly supervised semantic segmentation (WSSS), aiming to train segmentation models solely using image-level labels, has received significant attention. Existing approaches mainly concentrate on creating high-quality pseudo labels by utilizing existing images and their corresponding image-level labels. However, a major challenge arises when the available dataset is limited, as the quality of pseudo labels degrades significantly. In this paper, we tackle this challenge from a different perspective by introducing a novel approach called Generative Prompt Controlled Diffusion (GPCD) for data augmentation . This approach enhances the current labeled datasets by augmenting them with a variety of images, achieved through controlled diffusion guided by Generative Pre-trained Transformer (GPT) prompts. In this process, the existing images and image-level labels provide the necessary control information , while GPT enriches the prompts to generate diverse backgrounds. Moreover, we make an original contribution by integrating data source information as tokens into the Vision Transformer (ViT) framework, which improves the ability of downstream WSSS models to recognize the origins of augmented images. Our proposed GPCD approach clearly surpasses existing state-of-the-art methods, with its advantages being more pronounced when the available data is scarce, thereby demonstrating the effectiveness of our method. Our source code will be released. Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
Neurocomputing | 2 |
| 2024 | PSDPM: Prototype-based Secondary Discriminative Pixels Mining for Weakly Supervised Semantic SegmentationabstractImage-level Weakly Supervised Semantic Segmentation (WSSS) has received increasing attention due to its low an-notation cost. Class Activation Mapping (CAM) generated through classifier weights in WSSS inevitably ignores cer-tain useful cues, while the CAM generated through class prototypes can alleviate that. However, because of the dif-ferent goals of image classification and semantic segmentation, the class prototypes still focus on activating primary discriminative pixels learned from classification loss, leading to incomplete CAM. In this paper, we propose a plug-and-play Prototype-based Secondary Discriminative Pixels Mining (PSDPM) framework for enabling class prototypes to activate more secondary discriminative pixels, thus gen-erating a more complete CAM. Specifically, we introduce a Foreground Pixel Estimation Module (FPEM) for esti-mating potential foreground pixels based on the correlations between primary and secondary discriminative pix-els and the semantic segmentation results of baseline meth-ods. Then, we enable WSSS model to learn discriminative features from secondary discriminative pixels through a consistency loss calculated between FPEM result and class-prototype CAM. Experimental results show that our PSDPM improves various baseline methods significantly and achieves new state-of-the-art performances on WSSS benchmarks. Codes are available at https://github.com/xinqiaozhao/PSDPM. Xinqiao Zhao, Ziqian Yang, Tianhong Dai, Bingfeng Zhang, Jimin Xiao |
CVPR | 3 |
| 2024 | Image Augmentation with Controlled Diffusion for Weakly-Supervised Semantic SegmentationabstractWeakly-supervised semantic segmentation (WSSS), which aims to train segmentation models solely using image-level labels, has achieved significant attention. Existing methods primarily focus on generating high-quality pseudo labels using available images and their image-level labels. However, the quality of pseudo labels degrades significantly when the size of available dataset is limited. Thus, in this paper, we tackle this problem from a different view by introducing a novel approach called Image Augmentation with Controlled Diffusion (IACD). This framework effectively augments existing labeled datasets by generating diverse images through controlled diffusion, where the available images and image-level labels are served as the controlling information. Moreover, we also propose a high-quality image selection strategy to mitigate the potential noise introduced by the randomness of diffusion models. In the experiments, our proposed IACD approach clearly surpasses existing state-of-the-art methods. This effect is more obvious when the amount of available data is small, demonstrating the effectiveness of our method. Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
ICASSP | 2 |
| 2024 | Multimodal Frequeny Spectrum Fusion Schema for RGB-T Image Semantic SegmentationabstractSemantic segmentation confronts challenges with traditional networks tailored exclusively for RGB inputs, which may suffer from quality degradation under adverse conditions like low-level illumination or inclement weather. Recent advancements have shown promising outcomes by integrating RGB images with corresponding thermal infrared (TIR) images. However, effectively fusing features from both modalities remains a significant challenge. In this paper, we introduce a novel approach termed Multimodal Frequency Spectrum Fusion Schema (MFSFS) for semantic segmentation of RGB-T images. MFSFS leverages the advantages of the frequency spectrum to effectively extract and utilize multimodal feature information. To mitigate redundant information’s adverse effects during multimodal fusion in the frequency domain, we propose a diversity-oriented contrastive learning approach. Simulation results demonstrate that MFSFS achieves competitive performance while maintaining a relatively smaller model size. Hengyan Liu, Wenzhang Zhang, Tianhong Dai, Longfei Yin, Guangyu Ren |
ICCCN | 3 |
| 2024 | Cross Domain Policy Transfer with Effect Cycle-ConsistencyabstractTraining a robotic policy from scratch using deep reinforcement learning methods can be prohibitively expensive due to sample inefficiency. To address this challenge, transferring policies trained in the source domain to the target domain becomes an attractive paradigm. Previous research has typically focused on domains with similar state and action spaces but differing in other aspects. In this paper, our primary focus lies in domains with different state and action spaces, which has broader practical implications, i.e. transfer the policy from robot A to robot B. Unlike prior methods that rely on paired data, we propose a novel approach for learning the mapping functions between state and action spaces across domains using unpaired data. We propose effect cycle-consistency, which aligns the effects of transitions across two domains through a symmetrical optimization structure for learning these mapping functions. Once the mapping functions are learned, we can seamlessly transfer the policy from the source domain to the target domain. Our approach has been tested on three locomotion tasks and two robotic manipulation tasks. The empirical results demonstrate that our method can reduce alignment errors significantly and achieve better performance compared to the state-of-the-art method. Project page: https://ricky-zhu.github.io/effect_cycle_consistency. Tianhong Dai, Oya Çeliktutan |
ICRA | 2 |
| 2024 | Active Label Refinement for Robust Training of Imbalanced Medical Image Classification Tasks in the Presence of High Label Noise
Bidur Khanal, Tianhong Dai, Binod Bhattarai, Cristian A. Linte |
MICCAI (11) | 2 |
| 2024 | Top-K Pooling with Patch Contrastive Learning for Weakly-Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) using only image-level labels has gained significant attention due to cost-effectiveness. Recently, Vision Transformer (ViT) based methods without class activation map (CAM) have shown greater capability in generating reliable pseudo labels than previous methods using CAM. However, the current ViT-based methods utilize max pooling to select the patch with the highest prediction score to map the patch-level classification to the image-level one, which may affect the quality of pseudo labels due to the inaccurate classification of the patches. In this paper, we introduce a novel ViT-based WSSS method named top-K pooling with patch contrastive learning (TKP-PCL), which employs a top-K pooling layer to alleviate the limitations of previous max pooling selection. A patch contrastive error (PCE) is also proposed to enhance the patch embeddings to further improve the final results. The experimental results show that our approach is very efficient and outperforms other state-of-the-art WSSS methods on the PASCAL VOC 2012 and MS COCO 2014 dataset. Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
SMC | 2 |
| 2024 | Wavelet-based network for high dynamic range imagingabstractHigh dynamic range (HDR) imaging from multiple low dynamic range (LDR) images has been suffering from ghosting artifacts caused by scene and objects motion. Existing methods, such as optical flow based and end-to-end deep learning based solutions, are error-prone either in detail restoration or ghosting artifacts removal. Comprehensive empirical evidence shows that ghosting artifacts caused by large foreground motion are mainly low-frequency signals and the details are mainly high-frequency signals. In this work, we propose a novel frequency-guided end-to-end deep neural network (FHDRNet) to conduct HDR fusion in the frequency domain, and Discrete Wavelet Transform (DWT) is used to decompose inputs into different frequency bands. The low-frequency signals are used to avoid specific ghosting artifacts, while the high-frequency signals are used for preserving details. Using a U-Net as the backbone, we propose two novel modules: merging module and frequency-guided upsampling module. The merging module applies the attention mechanism to the low-frequency components to deal with the ghost caused by large foreground motion. The frequency-guided upsampling module reconstructs details from multiple frequency-specific components with rich details. In addition, a new RAW dataset is created for training and evaluating multi-frame HDR imaging algorithms in the RAW domain. Extensive experiments are conducted on public datasets and our RAW dataset, showing that the proposed FHDRNet achieves state-of-the-art performance. Tianhong Dai, Wei Li 0002, Xilei Cao, Jianzhuang Liu, Xu Jia 0012, Ales Leonardis, Youliang Yan, Shanxin Yuan |
Comput. Vis. Image Underst. | 1 |
| 2024 | Class Activation Map Calibration for Weakly Supervised Semantic SegmentationabstractImage-level weakly supervised semantic segmentation (WSSS) has received substantial attention due to its cost-effective annotation process. In WSSS, Class Activation Maps (CAMs) generated via classifier weights tend to focus on the most discriminative region, while the CAMs derived from class prototypes are significantly enhanced to cover more complete regions. However, the prototype CAMs still exhibit limitations such as incomplete localization maps on target objects and the presence of background noise. In this paper, we propose a novel WSSS framework called Classifier-Prototype Mutual Calibration (CPMC) that leverages the characteristics of both classifier and prototype CAMs to address the above issues. Specifically, an iterative refinement strategy based on context feature dependency is applied to refine the original classifier CAMs, which helps to generate improved prototype CAMs. Subsequently, local prototypes are constructed based on the false negative regions and false positive regions extracted from the previous two CAMs, which contribute to completing missing parts of the target object and suppressing background noise respectively. Therefore, CPMC can alleviate the aforementioned issues. Extensive experimental results on standard WSSS benchmarks (PASCAL VOC and MS COCO) show that our method significantly improves the quality of CAMs and achieves state-of-the-art performance. Our source code will be released. Jian Wang 0122, Tianhong Dai, Xinqiao Zhao, Ángel F. García-Fernández, Eng Gee Lim, Jimin Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Learning to Solve Tasks with Exploring Prior BehavioursabstractDemonstrations are widely used in Deep Reinforcement Learning (DRL) for facilitating solving tasks with sparse rewards. However, the tasks in real-world scenarios can often have varied initial conditions from the demonstration, which would require additional prior behaviours. For example, consider we are given the demonstration for the task of picking up an object from an open drawer, but the drawer is closed in the training. Without acquiring the prior behaviours of opening the drawer, the robot is unlikely to solve the task. To address this, in this paper we propose an Intrinsic Reward Driven Example-based Control (IRDEC). Our method can endow agents with the ability to explore and acquire the required prior behaviours and then connect to the task-specific behaviours in the demonstration to solve sparse-reward tasks without requiring additional demonstration of the prior behaviours. The performance of our method outperforms other baselines on three navigation tasks and one robotic manipulation task with sparse rewards. Codes are available at https://github.com/Ricky-Zhu/IRDEC. Siyuan Li 0003, Tianhong Dai, Chongjie Zhang, Oya Çeliktutan |
IROS | 3 |
| 2023 | Deep Reinforcement Learning for Real-Time Assembly Planning in Robot-Based Prefabricated ConstructionabstractThe adoption of robotics is promising to improve the efficiency, quality, and safety of prefabricated construction. Besides technologies that improve the capability of a single robot, the automated assembly planning for robots at construction sites is vital for further improving the efficiency and promoting robots into practices. However, considering the highly dynamic and uncertain nature of a construction environment, and the varied scenarios in different construction sites, it is always challenging to make appropriate and up-to-date assembly plans. Therefore, this paper proposes a Deep Reinforcement Learning (DRL) based method for automated assembly planning in robot-based prefabricated construction. Specifically, a re-configurable simulator for assembly planning is developed based on a Building Information Model (BIM) and an open game engine, which could support the training and testing of various optimization methods. Furthermore, the assembly planning problem is modelled as a Markov Decision Process (MDP) and a set of DRL algorithms are developed and trained using the simulator. Finally, experimental case studies in four typical scenarios are conducted, and the performance of our proposed methods have been verified, which can also serve as benchmarks for future research works within the community of automated construction.Note to Practitioners—This paper is conducted based on the comprehensive analysis of real-life assembly planning processes in prefabricated construction, and the methods proposed could bring many benefits to practitioners. Firstly, the proposed simulator could be easily re-configured to simulate diverse scenarios, which can be used to evaluate and verify the operations’ optimization methods and new construction technologies. Secondly, the proposed DRL-based optimization methods can be directly adopted in various robot-based construction scenarios, and can also be tailored to support the assembly planning in traditional human-based or human-robot construction environments. Thirdly, the proposed DRL methods and their performance in the four typical scenarios can serve as benchmarks for proposing new advanced construction technologies and optimization methods in assembly planning. Aiyu Zhu, Tianhong Dai, Gangyan Xu, Pieter Pauwels, Bauke de Vries |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2022 | LevDoom: A Benchmark for Generalization on Level Difficulty in Reinforcement LearningabstractDespite the recent success of deep reinforcement learning (RL), the generalization ability of RL agents remains an open problem for real-world applicability. RL agents trained on pixels may completely be derailed from achieving their objectives in unseen situations with different levels of visual changes. However, numerous existing RL suites do not address this as a primary objective or lack consistent level design of increased complexity. In this paper, we introduce the LevDoom benchmark, a suite containing semi-realistic 3D simulation environments with coherent levels of difficulty in the renowned video game Doom, designed to benchmark generalization in vision-based RL. We demonstrate how our benchmark reveals weaknesses of some popular Deep RL algorithms, which fail to prevail in modified environments. We further establish how our difficulty level design presents increasing complexity to these algorithms. Tristan Tomilin, Tianhong Dai, Mykola Pechenizkiy |
CoG | 2 |
| 2022 | Adaptive Intra-Group Aggregation for Co-Saliency DetectionabstractCo-salient object detection (CoSOD) together with the rapid development of deep learning has led to substantial progress in recent years. However, the feature aggregation between group feature representation and individual feature representation is still a challenging issue. In this work, we propose a novel adaptive intra-group aggregation (AIGA) method, which provides a new perspective to investigate the interaction relationship between group and single-image features and aggregate these features in an adaptive way. A novel scale-aware loss is proposed to help the model capture the scale prior of different groups and discriminatively process groups during the training phase. Extensive experiments demonstrate that the proposed method can effectively improve the performance without increasing extra parameters and achieve better accuracy on three prevalent benchmarks. Guangyu Ren, Tianhong Dai, Tania Stathaki |
ICASSP | 2 |
| 2022 | Progressive multi-scale fusion network for RGB-D salient object detectionabstractSalient object detection (SOD) aims at locating the most significant object within a given image. In recent years, great progress has been made in applying SOD on many vision tasks. The depth map could provide additional spatial prior and boundary cues to boost the performance. Combining the depth information with image data obtained from standard visual cameras has been widely used in recent SOD works, however, introducing depth information in a suboptimal fusion strategy may have negative influence in the performance of SOD. In this paper, we discuss about the advantages of the so-called progressive multi-scale fusion method and propose a mask-guided feature aggregation module (MGFA). The proposed framework can effectively combine the two features of different modalities and, furthermore, alleviate the impact of erroneous depth features, which are inevitably caused by the variation of depth quality. We further introduce a mask-guided refinement module (MGRM) to complement the high-level semantic features and reduce the irrelevant features from multi-scale fusion, leading to an overall refinement of detection. Experiments on five challenging benchmarks demonstrate that the proposed method outperforms 11 state-of-the-art methods under different evaluation metrics. Guangyu Ren, Yanchun Xie, Tianhong Dai, Tania Stathaki |
Comput. Vis. Image Underst. | 3 |
| 2022 | Analysing deep reinforcement learning agents trained with domain randomisationabstractDeep reinforcement learning (DRL) has the potential to train robots to perform complex tasks in the real world without requiring accurate models of the robot or its environment. However, agents trained with these algorithms typically lack the explainability of more traditional control methods. In this work, we use a combination of out-of-distribution generalisation tests and post hoc interpretability methods in order to understand what strategies DRL-trained agents use to perform a reaching task. To do so, we train agents under different conditions, using comparison to better interpret both quantitative and qualitative results; this allows us to not only provide local explanations, but also broad categorisations of behaviour. A key aim of our work is to understand how agents trained with visual domain randomisation (DR)—a technique which allows agents to generalise from simulation-based-training to the real world—differ from agents trained without. Our results show that the primary outcome of DR is more robust, entangled representations, accompanied by greater spatial structure in convolutional filters. Furthermore, even with an improved saliency method introduced in this work, we show that qualitative studies may not always correspond with quantitative measures, necessitating the combination of inspection tools in order to provide sufficient insights into the behaviour of trained agents. We conclude with recommendations for applying interpretability methods to DRL agents. Tianhong Dai, Kai Arulkumaran, Tamara Gerbert, Samyakh Tukra, Feryal M. P. Behbahani, Anil A. Bharath |
Neurocomputing | 1 |
| 2022 | Diversity-augmented intrinsic motivation for deep reinforcement learning
Tianhong Dai, Yali Du 0001, Anil A. Bharath |
Neurocomputing | 1 |
| 2021 | Diversity-Based Trajectory and Goal Selection with Hindsight Experience ReplayabstractHindsight experience replay (HER) is a goal relabelling technique typically used with off-policy deep reinforcement learning algorithms to solve goal-oriented tasks; it is well suited to robotic manipulation tasks that deliver only sparse rewards. In HER, both trajectories and transitions are sampled uniformly for training. However, not all of the agent's experiences contribute equally to training, and so naive uniform sampling may lead to inefficient learning. In this paper, we propose diversity-based trajectory and goal selection with HER (DTGSH). Firstly, trajectories are sampled according to the diversity of the goal states as modelled by determinantal point processes (DPPs). Secondly, transitions with diverse goal states are selected from the trajectories by using k-DPPs. We evaluate DTGSH on five challenging robotic manipulation tasks in simulated robot environments, where we show that our method can learn more quickly and reach higher performance than other state-of-the-art approaches on all tasks. Tianhong Dai, Hengyan Liu, Kai Arulkumaran, Guangyu Ren, Anil A. Bharath |
PRICAI (3) | 1 |
| 2021 | Coupled Network for Robust Pedestrian Detection With Gated Multi-Layer Feature Extraction and Deformable Occlusion HandlingabstractPedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, detecting ismall-scaled pedestrians and occluded pedestrians remains a challenging problem. In this paper, we propose a pedestrian detection method with a couple-network to simultaneously address these two issues. One of the sub-networks, the gated multi-layer feature extraction sub-network, aims to adaptively generate discriminative features for pedestrian candidates in order to robustly detect pedestrians with large variations on scale. The second sub-network targets on handling the occlusion problem of pedestrian detection by using deformable regional region of interest (RoI)-pooling. We investigate two different gate units for the gated sub-network, namely, the channel-wise gate unit and the spatio-wise gate unit, which can enhance the representation ability of the regional convolutional features among the channel dimensions or across the spatial domain, repetitively. Ablation studies have validated the effectiveness of both the proposed gated multi-layer feature extraction sub-network and the deformable occlusion handling sub-network. With the coupled framework, our proposed pedestrian detector achieves promising results on both two pedestrian datasets, especially on detecting small or occluded pedestrians. On the CityPersons dataset, the proposed detector achieves the lowest missing rates (i.e. 40.78% and 34.60%) on detecting small and occluded pedestrians, surpassing the second best comparison method by 6.0% and 5.87%, respectively. Tianrui Liu 0001, Wenhan Luo, Lin Ma 0002, Junjie Huang 0001, Tania Stathaki, Tianhong Dai |
IEEE Trans. Image Process. | 6 |
| 2020 | Gated Multi-Layer Convolutional Feature Extraction Network for Robust Pedestrian DetectionabstractPedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, it remains a challenging problem how to robustly detect pedestrians of varied sizes and with occlusions. In this paper, we propose a gated multi-layer convolutional feature extraction method which can adaptively generate discriminative features for candidate pedestrian regions. The proposed gated feature extraction framework consists of squeeze units, gate units and concatenation layers which perform feature dimension squeezing, feature manipulation and features combination from multiple CNN layers, respectively. We proposed two different gate models that can manipulate the regional feature maps in a channel-wise selection manner and a spatial-wise selection manner, respectively. Experiments on the challenging CityPersons dataset demonstrate the effectiveness of the proposed method, especially on detecting small-size and occluded pedestrians. Tianrui Liu 0001, Junjie Huang 0001, Tianhong Dai, Guangyu Ren, Tania Stathaki |
ICASSP | 3 |
| 2019 | LIIR: Learning Individual Intrinsic Reward in Multi-Agent Reinforcement LearningabstractA great challenge in cooperative decentralized multi-agent reinforcement learning (MARL) is generating diversified behaviors for each individual agent when receiving only a team reward. Prior studies have paid much effort on reward shaping or designing a centralized critic that can discriminatively credit the agents. In this paper, we propose to merge the two directions and learn each agent an intrinsic reward function which diversely stimulates the agents at each time step. Specifically, the intrinsic reward for a specific agent will be involved in computing a distinct proxy critic for the agent to direct the updating of its individual policy. Meanwhile, the parameterized intrinsic reward function will be updated towards maximizing the expected accumulated team reward from the environment so that the objective is consistent with the original MARL problem. The proposed method is referred to as learning individual intrinsic reward (LIIR) in MARL. We compare LIIR with a number of state-of-the-art MARL methods on battle games in StarCraft II. The results demonstrate the effectiveness of LIIR, and we show LIIR can assign each individual agent an insightful intrinsic reward per time step. Yali Du 0001, Lei Han 0001, Ji Liu 0002, Tianhong Dai, Dacheng Tao |
NeurIPS | 5 |