EDBT 2026 Demo / reviewers in the wild / expert
Kaichen Zhou
dblp:275/7059
· DBLP profile ↗
23ranked-venue papers
4as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Systems, architecture and hardware · 7 · 7 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel ViewsabstractNeural Radiance Fields (NeRF) have shown remarkable capabilities for photorealistic novel view synthesis. One major deficiency of NeRF is that dense inputs are typically required, and the rendering quality will drop drastically given sparse inputs. In this paper, we highlight the effectiveness of rendered semantics from dense novel views, and show that rendered semantics can be treated as a more robust form of augmented data than rendered RGB. Our method enhances NeRF’s performance by incorporating guidance derived from the rendered semantics. The rendered semantic guidance encompasses two levels: the supervision level and the feature level. The supervision-level guidance incorporates a bi-directional verification module that decides the validity of each rendered semantic label, while the feature-level guidance integrates a learnable codebook that encodes semantic-aware information, which is queried by each point via the attention mechanism to obtain semanticrelevant predictions. The overall semantic guidance is embedded into a self-improved pipeline.We also introduce a more challenging sparse-input indoor benchmark, where the number of inputs is limited to as few as 6. Experiments demonstrate the effectiveness of our method and it exhibits superior performance compared to existing approaches. Yingji Zhong, Kaichen Zhou, Zhihao Li 0002, Lanqing Hong, Zhenguo Li, Dan Xu 0002 |
AAAI | 2 |
| 2026 | TARGO and TARGO-Net: Benchmarking Target-Driven Object Grasping Under OcclusionsabstractPredicting 6-DoF grasp poses from a single RGB-D frame has recently achieved impressive accuracy, yet performance collapses when the target object is heavily occluded by clutter. In this paper, we establish the first benchmark dataset for TARget-driven Grasping under Occlusions, named TARGO, and our model that remains robust under occlusion, TARGO-Net. Our main contributions are: 1) We first recognize the visual occlusion challenge in 6-DoF grasping using single RGB-D images, and found that even the current SOTA models suffer under high occlusion. 2) We propose TARGO dataset, which can be used to train and test 6-DoF grasp models under different visual occlusion severities, and evaluate model robustness in real-world scenarios. 3) We further devise TARGO-Net, a transformer-based grasping model involving a target completion module and target-scene cross-attention, that performs most robustly across all visual occlusion levels. 4) We discover that other than visual occlusion, number of occluders and target minimum dimension also contribute to grasp success. The dataset and codes are publicly available at https://targo-benchmark.github.io . Yan Xia 0003, Ziyuan Qin 0002, Guanqi Zhan, Kaichen Zhou, Hao Dong 0003, Daniel Cremers |
Int. J. Comput. Vis. | 5 |
| 2025 | TransDiff: Diffusion-Based Method for Manipulating Transparent Objects Using a Single RGB-D ImageabstractManipulating transparent objects presents significant challenges due to the complexities introduced by their reflection and refraction properties, which considerably hinder the accurate estimation of their 3D shapes. To address these challenges, we propose a single-view RGB-D-based depth completion framework, TransDiff, that leverages the Denoising Diffusion Probabilistic Models(DDPM) to achieve material-agnostic object grasping in desktop. Specifically, we leverage features extracted from RGB images, including semantic segmentation, edge maps, and normal maps, to condition the depth map generation process. Our method learns an iterative denoising process that transforms a random depth distribution into a depth map, guided by initially refined depth information, ensuring more accurate depth estimation in scenarios involving transparent objects. Additionally, we propose a novel training method to better align the noisy depth and RGB image features, which are used as conditions to refine depth estimation step by step. Finally, we utilized an improved inference process to accelerate the denoising procedure. Through comprehensive experimental validation, we demonstrate that our method significantly outperforms the baselines in both synthetic and real-world benchmarks with acceptable inference time. The demo of our method can be found on: https://wang-haoxiao.github.io/TransDiff/ Haoxiao Wang, Kaichen Zhou, Binrui Gu, Zhiyuan Feng, Peilin Sun, Yicheng Xiao, Hao Dong 0003 |
ICRA | 2 |
| 2025 | SR3D: Unleashing Single-view 3D Reconstruction for Transparent and Specular Object GraspingabstractRecent advancements in 3D robotic manipulation have improved grasping of everyday objects, but transparent and specular materials remain challenging due to depth sensing limitations. While several 3D reconstruction and depth completion approaches address these challenges, they suffer from setup complexity or limited observation information utilization. To address this, leveraging the power of single-view 3D object reconstruction approaches, we propose a training-free framework SR3D that enables robotic grasping of transparent and specular objects from a single-view observation. Specifically, given single-view RGB and depth images, SR3D first uses the external visual models to generate 3D reconstructed object mesh based on RGB image. Then, the key idea is to determine the 3D object’s pose and scale to accurately localize the reconstructed object back into its original depth corrupted 3D scene. Therefore, we propose view matching and keypoint matching mechanisms, which leverage both the 2D and 3D’s inherent semantic and geometric information in the observation to determine the object’s 3D state within the scene, thereby reconstructing an accurate 3D depth map for effective grasp detection. Experiments in both simulation and real-world show the reconstruction effectiveness of SR3D. More demonstrations can be found at: https://sites.google.com/view/sr3dtech/ Mingxu Zhang, Xiaoqi Li 0020, Kaichen Zhou, Hojin Bae, Yan Shen 0035, Chuyan Xiong, Hao Dong 0003 |
IROS | 4 |
| 2025 | Non-Line-of-Sight 3D Object Reconstruction via mmWave Surface Normal EstimationabstractThis paper presents the design, implementation, and evaluation of mmNorm, a new and highly-accurate method for non-line-of-sight 3D object reconstruction using millimeter wave (mmWave) signals. In contrast to past approaches for millimeter-wave-based imaging that perform backprojection for 3D object reconstruction, mmNorm reconstructs the surface by estimating the object's surface normals. To do this, it introduces a novel algorithm that directly estimates the surface normal vector field from mmWave reflections. By then inverting the normal field, it can reconstruct structural isosurfaces, then solve for the exact surface through a novel mmWave optimization framework. Laura Dodds, Tara Boroushaki, Kaichen Zhou, Fadel Adib |
MobiSys | 3 |
| 2025 | LoopSparseGS: Loop-Based Sparse-View Friendly Gaussian SplattingabstractDespite the photorealistic novel view synthesis (NVS) performance achieved by the original 3D Gaussian splatting (3DGS), its rendering quality significantly degrades with sparse input views. This performance drop is mainly caused by the limited number of initial points generated from the sparse input, lacking reliable geometric supervision during the training process, and inadequate regularization of the oversized Gaussian ellipsoids. To handle these issues, we propose the LoopSparseGS, a loop-based 3DGS framework for the sparse novel view synthesis task. In specific, we propose a loop-based Progressive Gaussian Initialization (PGI) strategy that could iteratively densify the initialized point cloud using the rendered pseudo images during the training process. Then, the sparse and reliable depth from the Structure from Motion, and the window-based dense monocular depth are leveraged to provide precise geometric supervision via the proposed Depth-alignment Regularization (DAR). Additionally, we introduce a novel Sparse-friendly Sampling (SFS) strategy to handle oversized Gaussian ellipsoids leading to large pixel errors. Comprehensive experiments on four datasets demonstrate that LoopSparseGS outperforms existing state-of-the-art methods for sparse-input novel view synthesis, across indoor, outdoor, and object-level scenes with various image resolutions. Code is available at: https://github.com/pcl3dv/LoopSparseGS. Zhenyu Bao, Guibiao Liao, Kaichen Zhou, Kanglin Liu, Qing Li 0029, Guoping Qiu |
IEEE Trans. Image Process. | 3 |
| 2024 | Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical RepresentationabstractCoarse-to-fine 3D instance segmentation methods show weak performances compared to recent Grouping-based, Kernel-based and Transformer-based methods. We argue that this is due to two limitations: 1) Instance size over-estimation by axis-aligned bounding box(AABB) 2) False negative error accumulation from inaccurate box to the re-finement phase. In this work, we introduce Spherical Mask, a novel coarse-to-fine approach based on spherical repre-sentation, overcoming those two limitations with several benefits. Specifically, our coarse detection estimates each in-stance with a 3D polygon using a center and radial distance predictions, which avoids excessive size estimation of AABB. To cut the error propagation in the existing coarse-to-fine approaches, we virtually migrate points based on the polygon, allowing all foreground points, including false negatives, to be refined. During inference, the proposal and point mi-gration modules run in parallel and are assembled to form binary masks of instances. We also introduce two margin-based losses for the point migration to enforce corrections for the false positives/negatives and cohesion of foreground points, significantly improving the performance. Experimen-tal results from three datasets, such as ScanNetV2, S3DIS, and STPLS3D, show that our proposed method outperforms existing works, demonstrating the effectiveness of the new in-stance representation with spherical coordinates. The code is available at: https://github.com/yunshin/SphericalMask Sang-Yun Shin, Kaichen Zhou, Madhu Vankadari, Andrew Markham, Agathoniki Trigoni |
CVPR | 2 |
| 2024 | SSL-Net: A Synergistic Spectral and Learning-Based Network for Efficient Bird Sound ClassificationabstractEfficient and accurate bird sound classification is of important for ecology, habitat protection and scientific research, as it plays a central role in monitoring the distribution and abundance of species. However, prevailing methods typically demand extensively labeled audio datasets and have highly customized frameworks, imposing substantial computational and annotation loads. In this study, we present an efficient and general framework called SSL-Net, which combines spectral and learned features to identify different bird sounds. Encouraging empirical results gleaned from a standard field-collected bird audio dataset validate the efficacy of our method in extracting features efficiently and achieving heightened performance in bird sound classification, even when working with limited sample sizes. Furthermore, we present three feature fusion strategies, aiding engineers and researchers in their selection through quantitative analysis. Yiyuan Yang, Kaichen Zhou, Agathoniki Trigoni, Andrew Markham |
ICASSP | 2 |
| 2024 | Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation ModelsabstractSelf-supervised depth estimation algorithms rely heavily on frame-warping relationships, exhibiting substantial performance degradation when applied in challenging circumstances, such as low-visibility and nighttime scenarios with varying illumination conditions. Addressing this challenge, we introduce an algorithm designed to achieve accurate selfsupervised stereo depth estimation focusing on nighttime conditions. Specifically, we use pretrained visual foundation models to extract generalised features across challenging scenes and present an efficient method for matching and integrating these features from stereo frames. Moreover, to prevent pixels violating photometric consistency assumption from negatively affecting the depth predictions, we propose a novel masking approach designed to filter out such pixels. Lastly, addressing weaknesses in the evaluation of current depth estimation algorithms, we present novel evaluation metrics. Our experiments, conducted on challenging datasets including Oxford RobotCar and MultiSpectral Stereo, demonstrate the robust improvements realized by our approach. Madhu Vankadari, Samuel Hodgson, Sang-Yun Shin, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni |
ICRA | 4 |
| 2024 | Learning Generalizable Manipulation Policy with Adapter-Based Parameter Fine-TuningabstractThis study investigates the use of adapters in reinforcement learning for robotic skill generalization across multiple robots and tasks. Traditional methods are typically reliant on robot-specific retraining and face challenges such as efficiency and adaptability, particularly when scaling to robots with varying kinematics. We propose an alternative approach where a disembodied (virtual) hand manipulator learns a task (i.e., an abstract skill) and then transfers it to various robots with different kinematic constraints without retraining the entire model (i.e., the concrete, physical implementation of the skill). Whilst adapters are commonly used in other domains with strong supervision available, we show how weaker feedback from robotic control can be used to optimize task execution by preserving the abstract skill dynamics whilst adapting to new robotic domains. We demonstrate the effectiveness of our method with experiments conducted in the SAPIEN ManiSkill environment, showing improvements in generalization and task success rates. All code, data, and additional videos are at this GitHub link: https://kl-research.github.io/genrob. Kai Lu 0003, Kim Tien Ly, William Hebberd, Kaichen Zhou, Ioannis Havoutis, Andrew Markham |
IROS | 4 |
| 2024 | SCANet: Correcting LEGO Assembly Errors with Self-Correct Assembly NetworkabstractAutonomous assembly in robotics and 3D vision presents significant challenges, particularly in ensuring assembly correctness. Presently, predominant methods such as MEPNet focus on assembling components based on manually provided images. However, these approaches often fall short in achieving satisfactory results for tasks requiring long-term planning. Concurrently, we observe that integrating a self-correction module can partially alleviate such issues. Motivated by this concern, we introduce the Single-Step Assembly Error Correction Task, which involves identifying and rectifying misassembled components. To support research in this area, we present the LEGO Error Correction Assembly Dataset (LEGO-ECA), comprising manual images for assembly steps and instances of assembly failures. Additionally, we propose the Self-Correct Assembly Network (SCANet), a novel method to address this task. SCANet treats assembled components as queries, determining their correctness in manual images and providing corrections when necessary. Finally, we utilize SCANet to correct the assembly results of MEPNet. Experimental results demonstrate that SCANet can identify and correct MEPNet's misassembled results, significantly improving the correctness of assembly. Our code and dataset are available at https://github.com/Yaser-wyx/SCANet. Kaichen Zhou, Jinhong Chen, Hao Dong 0003 |
IROS | 2 |
| 2024 | WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization via Radiance FieldabstractDespite the advancements in deep learning for camera relocalization tasks, obtaining ground truth pose labels required for the training process remains a costly endeavor. While current weakly supervised methods excel in lightweight label generation, their performance notably declines in scenarios with sparse views. In response to this challenge, we introduce WSCLoc, a system capable of being customized to various deep learning-based relocalization models to enhance their performance under weakly-supervised and sparse view conditions. This is realized with two stages. In the initial stage, WSCLoc employs a multilayer perceptron-based structure called WFT-NeRF to co-optimize image reconstruction quality and initial pose information. To ensure a stable learning process, we incorporate temporal information as input. Furthermore, instead of optimizing SE(3), we opt for sim(3) optimization to explicitly enforce a scale constraint. In the second stage, we co-optimize the pre-trained WFT-NeRF and WFT-Pose. This optimization is enhanced by Time-Encoding based Random View Synthesis and supervised by inter-frame geometric constraints that consider pose, depth, and RGB information. We validate our approaches on two publicly available datasets, one outdoor and one indoor. Our experimental results demonstrate that our weakly-supervised relocalization solutions achieve superior pose estimation accuracy in sparse-view scenarios, comparable to state-of-the-art camera relocalization methods. We will make our code publicly available. Kaichen Zhou, Andrew Markham, Agathoniki Trigoni |
IROS | 2 |
| 2024 | RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and ManipulationabstractA fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they still face challenges in two areas: (1) insufficient reasoning ability to tackle complex tasks, and (2) high computational costs for VLA model fine-tuning and inference. The recently proposed state space model (SSM) known as Mamba demonstrates promising capabilities in non-trivial sequence modeling with linear inference complexity. Inspired by this, we introduce RoboMamba, an end-to-end robotic VLA model that leverages Mamba to deliver both robotic reasoning and action capabilities, while maintaining efficient fine-tuning and inference. Specifically, we first integrate the vision encoder with Mamba, aligning visual tokens with language embedding through co-training, empowering our model with visual common sense and robotic-related reasoning. To further equip RoboMamba with SE(3) pose prediction abilities, we explore an efficient fine-tuning strategy with a simple policy head. We find that once RoboMamba possesses sufficient reasoning capability, it can acquire manipulation skills with minimal fine-tuning parameters (0.1\% of the model) and time. In experiments, RoboMamba demonstrates outstanding reasoning capabilities on general and robotic evaluation benchmarks. Meanwhile, our model showcases impressive pose prediction results in both simulation and real-world experiments, achieving inference speeds 3 times faster than existing VLA models. Jiaming Liu 0003, Zhenyu Wang 0002, Pengju An, Xiaoqi Li 0020, Kaichen Zhou, Senqiao Yang, Renrui Zhang, Yandong Guo, Shanghang Zhang |
NeurIPS | 6 |
| 2024 | Beyond Fusion: Modality Hallucination-based Multispectral Fusion for Pedestrian DetectionabstractPedestrian detection is a fundamental task for many downstream applications. Visible and thermal images, as the two most important data types, are usually used to detect pedestrians under various environmental conditions. Many state-of-the-art works have been proposed to use two-stream (i.e., two-branch) architectures to combine visible and thermal information to improve detection performance. However, conventional visible-thermal fusion-based methods have no ability to obtain useful information from the visible branch under poor visibility conditions. The visible branch could even sometimes bring noise into the combined features. In this paper, we present a novel thermal and visible fusion architecture for pedestrian detection. Instead of simply using two branches to separately extract thermal and visible features and then fusing them, we introduce a hallucination branch to learn the mapping from the thermal to the visible domain, forming a novel three-branch feature extraction module. We then adaptively fuse feature maps from all three branches (i.e., thermal, visible, and hallucination). With this new integrated hallucination branch, our network can still get relatively good visible feature maps under challenging low-visibility conditions, thus boosting the overall detection performance. Finally, we experimentally demonstrate the superiority of the proposed architecture over conventional fusion methods. Qian Xie 0001, Ta Ying Cheng, Jia-Xing Zhong, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni |
WACV | 4 |
| 2024 | OV-NeRF: Open-Vocabulary Neural Radiance Fields With Vision and Language Foundation Models for 3D Semantic UnderstandingabstractThe development of Neural Radiance Fields (NeRFs) has provided a potent representation for encapsulating the geometric and appearance characteristics of 3D scenes. Enhancing the capabilities of NeRFs in open-vocabulary 3D semantic perception tasks has been a recent focus. However, current methods that extract semantics directly from Contrastive Language-Image Pretraining (CLIP) for semantic field learning encounter difficulties due to noisy and view-inconsistent semantics provided by CLIP. To tackle these limitations, we propose OV-NeRF, which exploits the potential of pre-trained vision and language foundation models to enhance semantic field learning through proposed single-view and cross-view strategies. First, from the single-view perspective, we introduce Region Semantic Ranking (RSR) regularization by leveraging 2D mask proposals derived from Segment Anything (SAM) to rectify the noisy semantics of each training view, facilitating accurate semantic field learning. Second, from the cross-view perspective, we propose a Cross-view Self-enhancement (CSE) strategy to address the challenge raised by view-inconsistent semantics. Rather than invariably utilizing the 2D inconsistent semantics from CLIP, CSE leverages the 3D consistent semantics generated from the well-trained semantic field itself for semantic field training, aiming to reduce ambiguity and enhance overall semantic consistency across different views. Extensive experiments validate our OV-NeRF outperforms current state-of-the-art methods, achieving a significant improvement of 20.31% and 18.42% in mIoU metric on Replica and ScanNet, respectively. Furthermore, our approach exhibits consistent superior results across various CLIP configurations, further verifying its robustness. Codes are available at:https://github.com/pcl3dv/OV-NeRF. Guibiao Liao, Kaichen Zhou, Zhenyu Bao, Kanglin Liu, Qing Li 0029 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Sample, Crop, Track: Self-Supervised Mobile 3D Object Detection for Urban Driving LiDARabstractDeep learning has led to great progress in the detection of mobile (i.e. movement-capable) objects in urban driving scenes in recent years. Supervised approaches typically require the annotation of large training sets; there has thus been great interest in leveraging weakly, semi- or self- supervised methods to avoid this, with much success. Whilst weakly and semi-supervised methods require some annotation, self-supervised methods have used cues such as motion to relieve the need for annotation altogether. However, a complete absence of annotation typically degrades their performance, and ambiguities that arise during motion grouping can inhibit their ability to find accurate object boundaries. In this paper, we propose a new self-supervised mobile object detection approach called SCT. This uses both motion cues and expected object sizes to improve detection performance, and predicts a dense grid of 3$D$oriented bounding boxes to improve object discovery. We significantly outperform the state-of-the-art self-supervised mobile object detection method TCR on the KITTI tracking benchmark, and achieve performance that is within 30 % of the fully supervised PV-RCNN++ method for IoUs$\leq$0.5. Our source code will be made available online. Sang-Yun Shin, Stuart Golodetz, Madhu Vankadari, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni |
ICRA | 4 |
| 2023 | Multi-body SE(3) Equivariance for Unsupervised Rigid Segmentation and Motion EstimationabstractA truly generalizable approach to rigid segmentation and motion estimation is fundamental to 3D understanding of articulated objects and moving scenes. In view of the closely intertwined relationship between segmentation and motion estimates, we present an SE(3) equivariant architecture and a training strategy to tackle this task in an unsupervised manner. Our architecture is composed of two interconnected, lightweight heads. These heads predict segmentation masks using point-level invariant features and estimate motion from SE(3) equivariant features, all without the need for category information. Our training strategy is unified and can be implemented online, which jointly optimizes the predicted segmentation and motion by leveraging the interrelationships among scene flow, segmentation mask, and rigid transformations. We conduct experiments on four datasets to demonstrate the superiority of our method. The results show that our method excels in both model performance and computational efficiency, with only 0.25M parameters and 0.92G FLOPs. To the best of our knowledge, this is the first work designed for category-agnostic part-level SE(3) equivariance in dynamic point clouds. Jia-Xing Zhong, Ta Ying Cheng, Kai Lu 0003, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni |
NeurIPS | 5 |
| 2023 | DynPoint: Dynamic Neural Point For View SynthesisabstractThe introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealing with uncontrolled or lengthy scenarios, and require extensive training time specific to each new scenario.
To tackle these limitations, we propose DynPoint, an algorithm designed to facilitate the rapid synthesis of novel views for unconstrained monocular videos.
Rather than encoding the entirety of the scenario information into a latent representation, DynPoint concentrates on predicting the explicit 3D correspondence between neighboring frames to realize information aggregation.
Specifically, this correspondence prediction is achieved through the estimation of consistent depth and scene flow information across frames.
Subsequently, the acquired correspondence is utilized to aggregate information from multiple reference frames to a target frame, by constructing hierarchical neural point clouds.
The resulting framework enables swift and accurate view synthesis for desired views of target frames.
The experimental results obtained demonstrate the considerable acceleration of training time achieved - typically an order of magnitude - by our proposed method while yielding comparable outcomes compared to prior approaches. Furthermore, our method exhibits strong robustness in handling long-duration videos without learning a canonical representation of video content. Kaichen Zhou, Jia-Xing Zhong, Sang-Yun Shin, Kai Lu 0003, Yiyuan Yang, Andrew Markham, Agathoniki Trigoni |
NeurIPS | 1 |
| 2022 | No Pain, Big Gain: Classify Dynamic Point Cloud Sequences with Static Models by Fitting Feature-level Space-time SurfacesabstractScene flow is a powerful tool for capturing the motion field of 3D point clouds. However, it is difficult to directly apply flow-based models to dynamic point cloud classification since the unstructured points make it hard or even impossible to efficiently and effectively trace point-wise correspondences. To capture 3D motions without explicitly tracking correspondences, we propose a kinematics-inspired neural network (Kinet) by generalizing the kinematic concept of ST-surfaces to the feature space. By unrolling the normal solver of ST-surfaces in the feature space, Kinet implicitly encodes feature-level dynamics and gains advantages from the use of mature back-bones for static point cloud processing. With only minor changes in network structures and low computing overhead, it is painless to jointly train and deploy our framework with a given static model. Experiments on NvGesture, SHREC'17, MSRAction-3D, and NTU-RGBD demonstrate its efficacy in performance, efficiency in both the number of parameters and computational complexity, as well as its versatility to various static backbones. Noticeably, Kinet achieves the accuracy of 93.27% on MSRAction-3D with only 3.20M parameters and 10.35G FLOPS. The code is available at https://github.com/jx-zhong-for-academic-purpose/Kinet. Jia-Xing Zhong, Kaichen Zhou, Qingyong Hu, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
CVPR | 2 |
| 2022 | DevNet: Self-supervised Monocular Depth Learning via Density Volume Construction
Kaichen Zhou, Lanqing Hong, Changhao Chen, Hang Xu 0004, Chaoqiang Ye, Qingyong Hu, Zhenguo Li |
ECCV (39) | 1 |
| 2022 | Smart Train Operation Algorithms Based on Expert Knowledge and Reinforcement LearningabstractDuring decades, the automatic train operation (ATO) system has been gradually adopted in many subway systems for its low-cost and intelligence. This article proposes two smart train operation (STO) algorithms by integrating the expert knowledge with reinforcement learning algorithms. Compared with previous works, the proposed algorithms can realize the control of continuous action for the subway system and optimize multiple critical objectives without using an offline speed profile. First, through learning historical data of experienced subway drivers, we extract the expert knowledge rules and build inference methods to guarantee the riding comfort, the punctuality, and the safety of the subway system. Then we develop two algorithms for optimizing the energy efficiency of train operation. One is the STO algorithm based on deep deterministic policy gradient named (STOD) and the other is the STO algorithm based on normalized advantage function (STON). Finally, we verify the performance of proposed algorithms via some numerical simulations with the real field data from the Yizhuang Line of the Beijing Subway and illustrate that the developed STO algorithm are better than expert manual driving and existing ATO algorithms in terms of energy efficiency. Moreover, STOD and STON can adapt to different trip times and different resistance conditions. Kaichen Zhou, Shiji Song, Anke Xue, Keyou You, Hui Wu 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | VMLoc: Variational Fusion For Learning-Based Multimodal Camera LocalizationabstractRecent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missing input are less well studied. In particular, we note that previous approaches towards deep fusion do not perform significantly better than models employing a single modality. We conjecture that this is because of the naive approaches to feature space fusion through summation or concatenation which do not take into account the different strengths of each modality. To address this, we propose an end-to-end framework, termed VMLoc, to fuse different sensor inputs into a common latent space through a variational Product-of-Experts (PoE) followed by attention-based fusion. Unlike previous multimodal variational works directly adapting the objective function of vanilla variational auto-encoder, we show how camera localization can be accurately estimated through an unbiased objective function based on importance weighting. Our model is extensively evaluated on RGB-D datasets and the results prove the efficacy of our model. The source code is available at https://github.com/Zalex97/VMLoc. Kaichen Zhou, Changhao Chen, Bing Wang 0013, Muhamad Risqi Utama Saputra, Agathoniki Trigoni, Andrew Markham |
AAAI | 1 |
| 2020 | Suggestive Annotation of Brain Tumour Images with Gradient-Guided Sampling
Chengliang Dai, Shuo Wang 0011, Yuanhan Mo, Kaichen Zhou, Elsa D. Angelini, Yike Guo, Wenjia Bai |
MICCAI (4) | 4 |