VLDB 2026 Research / reviewers in the wild / expert
Chaofan Zhang
dblp:231/6995
· DBLP profile ↗
19ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The rainbow number of cycles in maximal outerplanar graphs
Liman Jiao, Chaofan Zhang, Yongxin Lan |
Discret. Appl. Math. | 3 |
| 2026 | Beyond Appearance: Dual-Graph Object Encoding With Learnable Graph StructureabstractObject encoding is essential for enabling robots to efficiently perform tasks such as recognition and autonomous exploration. Existing object encoding approaches typically rely on appearance-based representations, leading to poor performance in environments with multiple visually similar objects, which can compromise downstream task accuracy. Inspired by human perception of objects through both appearance and structure, we introduce DGOE, a novel object encoding method, which adopts a dual-graph embedding scheme. Rather than treating object representation as a single-level appearance encoding problem, DGOE explicitly models object discriminability through a dual-level structural formulation, decomposing it into intrinsic inner-object structure and cross-object relational context. DGOE leverages graph structures as the carrier to integrate object intrinsic appearance and cross-object spatial features. This unified framework leads to more distinctive and reliable object representations. Additionally, we use multi-head self-attention to learn cross-object graph structures for extracting cross-object spatial features. Experimental evaluations on multiple publicly available datasets demonstrate the superior performance of our method in object-level matching tasks, underscoring its effectiveness and robustness. Cuiyun Fang, Xirun Cheng, Yingwei Xia, Guoqing Deng, Chaofan Zhang |
IEEE Signal Process. Lett. | 8 |
| 2026 | DexTac: Learning Contact-Aware Visuotactile Policies via Hand-by-Hand TeachingabstractFor contact-intensive tasks, the ability to generate policies that produce comprehensive tactile-aware motions is essential. However, existing data collection and skill learning systems for dexterous manipulation often suffer from low-dimensional tactile information. To address this limitation, we propose DexTac, a visuo-tactile manipulation learning framework based on kinesthetic teaching. DexTac captures multi-dimensional tactile data—including contact force distributions and spatial contact regions—directly from human demonstrations. By integrating these rich tactile modalities into a policy network, the resulting contact-aware agent enables a dexterous hand to autonomously select and maintain optimal contact regions during complex interactions. We evaluate our framework on a challenging unimanual injection task. Experimental results demonstrate that DexTac achieves a 91.67% success rate. Notably, in high-precision scenarios involving small-scale syringes, our approach outperforms force-only baselines by 31.67%. These results underscore that learning multi-dimensional tactile priors from human demonstrations is critical for achieving robust, human-like dexterous manipulation in contact-rich environments. Chaofan Zhang, Boyue Zhang 0002, Zhinan Peng, Shaowei Cui, Shuo Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Remote sensing image change detection network with multi-scale feature information mining and fusion
Songdong Xue, Minming Zhang, Gangzhu Qiao, Chaofan Zhang |
Pattern Anal. Appl. | 4 |
| 2025 | GelStereo Tip: A Spherical Fingertip Visuotactile Sensor for Multi-Finger Screwing ManipulationabstractDexterous hands are the key element for robots to achieve human-like manipulation capabilities. An outstanding challenge is to provide fingertips of dexterous hands with precise tactile deformation sensing capabilities. In this paper, we present the GelStereo Tip, a spherical and easy-to-integrate GelStereo-type visuotactile sensor capable of sensing high-resolution 3D elastomer deformation. Previous calibration method does not take into account the impact of imaging errors caused by the sensor’s compact and high-curvature structural characteristics on the accuracy of tactile sensing. Therefore, we propose a novel self-calibration method based on the Refractive Stereo Ray Tracing model, named GTSC, and demonstrate the accuracy of less than 0.3 mm for deformation sensing. Furthermore, we also propose a Contact Retention Tactile Controller to address the issue of fingertips being unable to overcome obstructive torque during the multi-finger bottle cap screwing. After integrating GelStereo Tip into fingertips of Allegro Hand, the controller adjusts the joint positions of the given trajectory using proportional control based on the difference between the sensor’s actual deformation and the reference state for contact retention. We believe that the GelStereo Tip sensor combined with robotic dexterous hands has great application potential in the field of multi-finger fingertip manipulation. Note to Practitioners—The motivation of this paper is to design a fingertip visuotactile sensor with high-precision 3D tactile deformation sensing capabilities for multi-finger robotic hands and to validate its sensing performance. Additionally, it aims to address the issue of overcoming resistance in multi-finger screwing manipulations. Currently, most sensors do not consider the refraction effect or ignore the impact of planar imaging errors in refractive calibration. This paper proposes a visuotactile sensor along with a corresponding self-calibration method to ensure its sensing accuracy. Experiments show that our sensor possesses high-precision and robust 3D deformation sensing capabilities. On the other hand, multi-finger hands often struggle to complete screwing tasks along the given trajectory due to disturbances from torque resistance. This paper proposes a tactile controller that evaluates the contact state through aforementioned tactile sensing to improve subsequent trajectory and achieve continuous screwing. Comparative experiments highlight the necessity of this controller and the reliability of tactile sensing. We hope that the design of our sensor, the self-calibration method, and the tactile controller applied to multi-finger screwing can provide new insights for other practitioners. Boyue Zhang 0002, Shaowei Cui, Chaofan Zhang, Jingyi Hu, Rui Wang 0031, Shuo Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | SGHD: Semantic-Geometric Combined 3-D Histogram Descriptor for Global Localization in Autonomous VehiclesabstractLiDAR-based global localization plays a crucial role in autonomous driving and robotic systems. However, real-time changes in environmental conditions present challenges for global localization. Current methods predominantly rely on sparse LiDAR point clouds with single-feature information (semantic or geometric), leading to insufficient accuracy and robustness in 6-DOF pose estimation and similarity matching. In this work, we propose a novel semantic-geometric combined 3D histogram descriptor, called the SGHD descriptor, with robust place recognition and 6-DoF pose estimation capabilities. We first extract semantic and geometric features of objects through operations such as point cloud clustering and semantic segmentation. Next, we describe the distribution of neighboring objects using topological structures. To improve object discriminability, we construct triangular descriptors for each object based on semantic information and relative distance. Subsequently, we consider the complementary nature of semantic and geometric information and propose a 3D histogram descriptor. This 3D histogram descriptor encodes semantic, spatial angle and relative distance information, enhancing the invariance of the triangle descriptor. We perform similarity discrimination and 6-DOF pose estimation through coarse-to-fine matching of local and global semantic maps. Overall, the proposed descriptor achieves high-accuracy and high-robustness 6-DOF pose estimation and place recognition for global localization. We extensively compared the proposed SGHD descriptor with state-of-the-art methods on the public SemanticKITTI dataset. Quantitative results demonstrate that SGHD exhibits higher adaptability and significant improvements compared to its counterparts. The implementation of SGHD will be released athttps://github.com/wf-hahaha/SGHD Cuiyun Fang, Nanhao Liang, Chaofan Zhang, Guoqing Deng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | TacFlex: Multimode Tactile Imprints Simulation for Visuotactile Sensors With Coating PatternsabstractVisuotactile sensors have been shown to provide rich contact information for robots. However, how to build a high-fidelity visuotactile simulator that supports multi-mode tactile imprints and various sensor configurations (such as coating patterns) remains a challenging problem. In this paper, we present TacFlex, an efficient and flexible simulator for visuotactile sensors, which physically simulates the elastomer deformation using Finite Element Methods (FEM), and focuses on linking the deformed elastomer mesh to diverse tactile imprints, including tactile images with arbitrary coating patterns and tactile 3D point clouds. We further propose a ray tracing-based rectification method to deal with multi-medium refraction effects to make the simulated tactile images more realistic. Extensive qualitative and quantitative experiments are conducted to demonstrate the effectiveness of TacFlex on several visuotactile sensors. Furthermore, we explore the Sim2Real performance of different tactile imprints provided by TacFlex in tactile perception and manipulation tasks, such as cylindrical object pose estimation and peg-in-hole. The perception/policy models trained in simulation are successfully deployed in the real world. Finally, we present the outlook on the potential of TacFlex in visuotactile manipulation learning. The TacFlex simulator is open-sourced to the community. See supplementary video, code, and results athttps://sites.google.com/view/tacflex/. Chaofan Zhang, Shaowei Cui, Jingyi Hu, Tiandong Zhang, Rui Wang 0031, Shuo Wang 0001 |
IEEE Trans. Robotics | 1 |
| 2024 | Learning-Based Slip Detection for Dexterous Manipulation Using GelStereo SensingabstractEndowing the robot with tactile perception can effectively improve manipulation dexterity, along with various benefits of human-like touch. Using GelStereo (GS) tactile sensing, which gives high-resolution contact geometry information, including 2-D displacement field, and 3-D point cloud of the contact surface, we present a learning-based slip detection system in this study. The results reveal that the well-trained network achieves 95.79% accuracy on the never-seen testing dataset, which surpasses the current model-based and learning-based methods using visuotactile sensing. We also propose a general framework for slip feedback adaptive control for dexterous robot manipulation tasks. The experimental results show the effectiveness and efficiency of the proposed control framework using GS tactile feedback when deployed on real-world grasping and screwing manipulation tasks on various robot setups. Shaowei Cui, Shuo Wang 0001, Rui Wang 0031, Chaofan Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | GelStereo Palm: A Novel Curved Visuotactile Sensor for 3-D Geometry SensingabstractRecently, visuotactile sensors have shown promising potential in robotics due to their high-resolution sensing ability. Unfortunately, the majority of available visuotactile sensors are limited to flat shapes, which severely limits their application possibilities. In this article, we propose a novel curved visuotactile sensor, the GelStereo Palm, which senses the 3-D contact geometry on a curved surface using a binocular vision system. Meanwhile, to solve the light refraction problem in the binocular stereo vision system under a curved medium, a refractive stereo ray tracing model for GelStereo Palm is presented. Moreover, a 3-D tactile point cloud sensing pipeline is introduced to reconstruct the 3-D contact geometry in real-time. Finally, extensive experiments are conducted to verify the accuracy and robustness of the 3-D contact geometry sensing of our GelStereo Palm sensor. Jingyi Hu, Shaowei Cui, Shuo Wang 0001, Chaofan Zhang, Rui Wang 0031, Lipeng Chen |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Learning-based Six-axis Force/Torque Estimation Using GelStereo Fingertip Visuotactile SensingabstractVisuotactile sensors have recently attracted much attention in robot communities due to the benefit of high spatial resolution sensing. However, force/torque estimation by visuotactile sensors remains a challenging problem. In this paper, we propose a learning-based six-axis force/torque estimation network using GelStereo visuotactile sensor, which can provide two-dimensional (2D) and three-dimensional (3D) displacements of markers embedded in the sensor surface. The convolutional neural networks are employed to extract multi-modal tactile deformation features; and a novel contact positional encoding method is proposed to eliminate the influence of translation invariance in convolutional operators. The well-trained model achieves the best RMSE of 0.290 N in force and 0.0084 Nm in torque. Furthermore, the proposed force/torque estimation network is integrated with a force-feedback policy for adaptive grasping tasks. The experimental results demonstrate the effectiveness of the proposed method and its potential application in robotic grasping and manipulation tasks. Chaofan Zhang, Shaowei Cui, Yinghao Cai, Jingyi Hu, Rui Wang 0031, Shuo Wang 0001 |
IROS | 1 |
| 2022 | No-reference Omnidirectional Image Quality Assessment Based on Joint NetworkabstractIn panoramic multimedia applications, the perception quality of the omnidirectional content often comes from the observer's perception of the viewports and the overall impression after browsing. Starting from this hypothesis, this paper proposes a deep-learning based joint network to model the no-reference quality assessment of omnidirectional images. On the one hand, motivated by different scenarios that lead to different human understandings, a convolutional neural network (CNN) is devised to simultaneously encode the local quality features and the latent perception rules of different viewports, which are more likely to be noticed by the viewers. On the other hand, a recurrent neural network (RNN) is designed to capture the interdependence between viewports from their sequence representation, and then predict the impact of each viewport on the observer's overall perception. Experiments on two popular omnidirectional image quality databases demonstrate that the proposed method outperforms the state-of-the-art omnidirectional image quality metrics. Chaofan Zhang, Shiguang Liu |
ACM Multimedia | 1 |
| 2022 | Semi-global shape-aware attention network for image segmentation and retrieval
Jiagang Zhu, Chaofan Zhang, Zheng Rong, Yihong Wu 0002 |
Neurocomputing | 3 |
| 2022 | Leveraging local and global descriptors in parallel to search correspondences for visual localization
Chaofan Zhang, Bingxi Liu 0001, Yihong Wu 0002 |
Pattern Recognit. | 2 |
| 2022 | Dual-Channel Multi-Task CNN for No-Reference Screen Content Image Quality AssessmentabstractNowadays the problem of image quality assessment (IQA) for screen content images (SCIs) has become a research hotspot as they are ubiquitous in multimedia applications. Although the quality assessment of natural images (NIs) has been continuously developed in the past few decades, few NI-oriented IQA methods can be directly applied on SCIs due to different visual characteristics between them. In this paper, we present a no-reference quality prediction approach considering the content information of SCIs, which is based on dual-channel multi-task convolutional neural network. First, we segment a SCI into small patches and classify them as the textual patches and the pictorial patches. Then, we devise a novel dual-channel convolutional neural network (CNN) to predict the quality of textual patches and pictorial patches. Finally, we propose an effective adaptive weighting strategy for quality score aggregation. The proposed CNN is built on an end-to-end multi-task learning framework, which assists the SCI quality prediction task through the histogram of oriented gradient (HOG) feature prediction task to learn a better mapping between the input patch and its quality score. The adaptive weighting strategy further improves the representation ability of each SCI patch. Experimental results on two largest SCI-oriented databases demonstrate that the proposed method outperforms most of the state-of-the-art no-reference IQA methods and the full-reference IQA methods. Chaofan Zhang, Ziqing Huang, Shiguang Liu, Jian Xiao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Highly Efficient Line Segment Tracking with an IMU-KLT Prediction and a Convex Geometric Distance MinimizationabstractLine segment features become popular in SLAM community. Usually, line-based SLAM systems utilize local appearance descriptors for line segment tracking. However, traditional descriptor-based line segment tracking algorithms suffer from the problem that accuracy and speed cannot be possessed simultaneously, which affects the performance of line-based SLAM systems negatively. We propose a novel line segment tracking method with an IMU-KLT line segment prediction and a convex geometric distance minimization to boost line segment tracking performance in both accuracy and speed. Particularly, the proposed convex geometric distance minimization uses a ℓ1-norm model to minimize geometric constraints between predicted line segments and extracted line segments efficiently. Furthermore, the line segment tracking is embedded into a VIO system and we adapt it to obtain more reliable point tracking. Experimental results on public datasets show that the proposed line segment tracking method achieves much higher accuracy and much less time cost than state-of-the-art level, where not only the number of correct matches increases but also the inlier ratios are increased by at least 35.1% along with a 3 times faster speed. Besides, the VIO system combining the proposed line segment tracking is improved in terms of accuracy. Hao Wei 0008, Fulin Tang, Chaofan Zhang, Yihong Wu 0002 |
ICRA | 3 |
| 2021 | Learning to Decompose and Restore Low-light Images with Wavelet TransformabstractLow-light images often suffer from low visibility and various noise. Most existing low-light image enhancement methods often amplify noise when enhancing low-light images, due to the neglect of separating valuable image information and noise. In this paper, we propose a novel wavelet-based attention network, where wavelet transform is integrated into attention learning for joint low-light enhancement and noise suppression. Particularly, the proposed wavelet-based attention network includes a Decomposition-Net, an Enhancement-Net and a Restoration-Net. In Decomposition-Net, to benefit denoising, wavelet transform layers are designed for separating noise and global content information into different frequency features. Furthermore, an attention-based strategy is introduced to progressively select suitable frequency features for accurately restoring illumination and reflectance according to Retinex theory. In addition, Enhancement-Net is introduced for further removing degradations in reflectance and adjusting illumination, while Restoration-Net employs conditional adversarial learning to adversarially improve the visual quality of final restored results based on enhanced illumination and reflectance. Extensive experiments on several public datasets demonstrate that the proposed method achieves more pleasing results than state-of-the-art methods. Chaofan Zhang, Zheng Rong, Yihong Wu 0002 |
MMAsia | 2 |
| 2021 | High Power-Efficient and Performance-Density FPGA Accelerator for CNN-Based Object Detection
Chaofan Zhang, Fulin Tang, Yihong Wu 0002, Xuezhi Yang |
PRCV (1) | 2 |
| 2019 | Accurate and Robust RGB-D Dense Mapping with Inertial Fusion and Deformation-Graph OptimizationabstractRGB-D dense mapping has become more and more popular, however, when encountering rapid movement or shake, the robustness and accuracy of most RGB-D dense mapping methods are degraded and the generated maps are overlapped or distorted, due to the drift of pose estimation. In this paper, we present a novel RGB-D dense mapping method, which can obtain accurate, robust and global consistency map even in the above complex conditions. Firstly, the improved ORBSLAM method, which tightly-couples RGB-D information and inertial information to estimate the current pose of robot, is firstly introduced for accurate pose estimation rather than traditional frame-to-frame method in most RGB-D dense mapping methods. Besides, the TSDF (Truncated Signed Distance Function) method is used to effectively fuse depth frame into a global model, and to keep the global consistency of the generated map. Furthermore, since the drift error is inevitable, a deformation graph is constructed to minimize the consistent error in global model, to further improve the mapping performance. The performance of the proposed RGB-D dense mapping method was validated by extensive localization and mapping experiments on public datasets and real scene datasets, and it showed strongly accuracy and robustness over other state-of-the-art methods. What's more, the proposed method can achieve real-time performance implemented on GPU. Liming Bao, Chaofan Zhang, Yingwei Xia |
ICTAI | 3 |
| 2019 | Trajectory optimization for RLV in TAEM phase using adaptive Gauss pseudospectral method
Chaofan Zhang |
Sci. China Inf. Sci. | 2 |