VLDB 2026 Research / reviewers in the wild / expert
Yibin Tian
dblp:16/7591
· DBLP profile ↗
17ranked-venue papers
0as first author
16since 2021 · last 2026
0009-0008-4590-8318ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Denoising of Mammographic Microcalcifications Using Wavelet-Based Attention Deep Network for the Medical Internet of ThingsabstractBreast microcalcifications are key radiological indicators of malignancy, and their assessment in denser breast tissues is challenging. We propose a Wavelet-based Attention Deep Network (WADN) that isolates high-frequency components linked to microcalcification clusters using denoising wavelet transforms and integrates them into a deep attention architecture for interpretable, localization-assisted classification. The system is deployed in an edge-computing framework enabled by Jetson Orin Nano to support real-time analysis and remote radiologist access for the Medical Internet of Things (MIoT). For the private PINUM and public DDSM datasets, WADN outperforms standard convolutional neural networks (CNNs) and classical machine learning baselines, achieving accuracies of 0.95 and 0.93, respectively, with consistent gains in sensitivity, specificity, precision, F1-score, and area under the curve (AUC). On the IoT edge device, the method achieves accuracies of 0.91 for PINUM and 0.93 for DDSM, demonstrating the feasibility of ondevice screening. These results indicate that combining wavelet transforms with attention mechanisms and IoT edge deployment improves diagnostic accuracy while enabling practical real-time breast-cancer screening workflows. Khalil ur Rehman, Anaa Yasin, Jianqiang Li 0002, Ayesha Jabbar, Yibin Tian |
IEEE Internet Things J. | 5 |
| 2026 | A Feature Fusion Attention-Based Deep Learning Algorithm for Mammographic Architectural Distortion ClassificationabstractArchitectural Distortion (AD) is a common abnormality in digital mammograms, alongside masses and microcalcifications. Detecting AD in dense breast tissue is particularly challenging due to its heterogeneous asymmetries and subtle presentation. Factors such as location, size, shape, texture, and variability in patterns contribute to reduced sensitivity. To address these challenges, we propose a novel feature fusion-based Vision Transformer (ViT) attention network, combined with VGG-16, to improve accuracy and efficiency in AD detection. Our approach mitigates issues related to texture fixation, background boundaries, and deep neural network limitations, enhancing the robustness of AD classification in mammograms. Experimental results demonstrate that the proposed model achieves state-of-the-art performance, outperforming eight existing deep learning models. On the PINUM dataset, it attains 0.97 sensitivity, 0.92 F1-score, 0.93 precision, 0.94 specificity, and 0.96 accuracy. On the DDSM dataset, it records 0.93 sensitivity, 0.91 F1-score, 0.94 precision, 0.92 specificity, and 0.95 accuracy. These results highlight the potential of our method for computer-aided breast cancer diagnosis, particularly in low-resource settings where access to high-end imaging technology is limited. By enabling more accurate and timely AD detection, our approach could significantly improve breast cancer screening and early intervention worldwide. Khalil ur Rehman, Jianqiang Li 0002, Anaa Yasin, Shakila Basheer, Inam Ullah 0001, Kashif Jabbar, Yibin Tian |
IEEE J. Biomed. Health Informatics | 8 |
| 2026 | Fault-Tolerant Consensus Control for T-S Fuzzy Multiagent Systems Under Aperiodic Time-Constrained Sampling Communication and Stochastic Actuator FaultsabstractThe main objective of this study is to develop a fault-tolerant fuzzy time-constrained memory-sampled-data control (TCMSDC) mechanism for analyzing the consensus performance of nonlinear multiagent systems (MASs) affected by stochastic actuator faults and external disturbances. To achieve this, the nonlinear MASs are first converted into quasi-linear subsystems using the Takagi–Sugeno (T-S) fuzzy model. In contrast to existing memory-sampled-data consensus methods, the constructed TCMSDC signals vary over time within each sampling period, thereby enhancing consensus performance. Thereafter, a Markov variable process is applied to model multiple stochastic actuator faults in the considered MASs. Furthermore, an aperiodic-sampling-dependent asymmetric dual-looped functional (ASADLF) is constructed, which incorporates different sets of matrices designed for each sampling instance. Subsequently, a new nonorthogonal polynomial integral inequality (NOPII) is introduced to approximate the integral quadratic terms. By leveraging this ASADLF along with the proposed NOPII technique, less conservative consensus criteria are derived in the form of linear matrix inequalities (LMIs), and the TCMSDC gain parameters are obtained to guarantee mean-square asymptotic consensus with$H_{\infty }$performance for the considered MASs. Finally, the effectiveness and advantages of the proposed TCMSDC approach in enhancing fault tolerance and achieving reliable consensus under stochastic actuator faults are validated through numerical simulations on multi-ship steering autonomous surface vehicles (SSASVs) and mass–spring systems. A. Pratap 0001, Yibin Tian, Zhiguang Feng, Tingwen Huang, Yukang Cui 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | ITW-DehazeFormer: Imaging through Turbid Water Using Improved DehazeFormerabstractLight scattering and absorption degrade the quality of underwater images, and various image enhancement methods have been explored. However, the existing underwater image datasets lack corresponding high-quality references, and the degree of scattering and absorption is not strictly controlled. In this study, we constructed an image dataset with different degrees of light scattering and controlled water turbidity via a water tank. The Swin Transformer based dehazing network DehazeFormer has been improved, termed ITW-DehazeFormer, to enhance images acquired through turbid water. First, a histogram equalization pre-enhancement block is added. Second, the SKfusion block is replaced by a content-guided attention based fusion block to combine channel and spatial attention so that information interactions between different channels are guaranteed. Finally, a hybrid loss function combining space and frequency domain information is introduced. Experimental results show that ITW-DehazeFormer outperforms seven existing image enhancement methods in terms of several image quality metrics, including Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM) and multi-scale SSIM. Qicong Wang, Xiaopin Zhong, Dajiang Lu, Yibin Tian |
ICASSP | 4 |
| 2025 | LEFF-YOLO: A Lightweight Cherry Tomato Detection YOLOv8 Network with Enhanced Feature Fusion
Yibin Tian |
ICIC (21) | 2 |
| 2025 | Toroidal Adaptive Intensity and Spectrum Updating Image Reconstruction for Fourier Ptychographic MicroscopyabstractFourier Ptychographic Microscopy (FPM) is a popular computational imaging technique that reconstructs high-resolution and wide-field images from multiple low-resolution ones acquired with varying illumination angles, overcoming the optical diffraction limit. This study proposes an image reconstruction enhancement to FPM, named toroidal adaptive intensity and spectrum updating. It employs an image intensity update strategy that replaces low-resolution image intensities with varying fractional power based on illumination angles and image frequency content, which in turn adaptively determines how the reconstruction updates in the frequency domain. It consists of three key components: (a) determining two fractional powers for intensity updating, one of which is decided by illumination angles, and the other by wavelet transform frequency content; (b) obtaining a weighted sum of the two fractional powers; (c) an adaptive image intensity and spectrum updating scheme using the fractional power. Experimental validation was conducted on the USAF 1951 chart and 10 biological specimen slices using a Pentax microscope with a 15x15 LED array for illumination. It demonstrates reduced artifacts and improved local contrast compared to conventional methods. Dajiang Lu, Yibin Tian |
ICIP | 5 |
| 2025 | A Dental Periapical X-ray Images Segmentation Network Based on Pixel-Wise Contrastive Learning with Dual Attention MechanismsabstractPeriapical radiographs tend to have poor quality due to factors like acquisition techniques, equipment limitations, and patient differences. These factors lead to discrepancies in images, making precise segmentation challenging. However, accurate segmentation is crucial in dental practices. To address this issue, deep learning can be employed to improve segmentation accuracy and efficiency, thereby providing more reliable support for clinical diagnosis. In this study, we propose a deep learning network based on an encoder-decoder architecture, which integrates a dual attention mechanism and pixel-wise contrastive learning to address the tooth segmentation problem. We design the Dual Attention Contrast (DAC) module, which enhances feature maps through joint spatial and channel attention before utilizing optimized multi-scale features for both segmentation prediction and pixel-wise contrastive learning. This module, implemented with a multi-level deployment strategy, strengthens the network's ability to extract discriminative features from the multi-scale anatomical structures in periapical radiographs while constructing a globally structured feature space across datasets. The dual attention mechanism improves the recognition of key local features through spatial-channel collaborative calibration, while pixel-wise contrastive learning explicitly constrains the topological structure of the feature space, effectively mitigating the impact of inter-image variations on segmentation accuracy. Experimental results show that the proposed model outperforms current mainstream models across various evaluation metrics in periapical radiograph segmentation tasks. Yibin Tian, Zhiyuan Zhang 0004, Xueyang Zhang, Shan LianLei |
IJCNN | 3 |
| 2025 | HDiffTG: A Lightweight Hybrid Diffusion-Transformer-GCN Architecture for 3D Human Pose EstimationabstractWe propose HDiffTG, a novel 3D Human Pose Estimation (3DHPE) method that integrates Transformer, Graph Convolutional Network (GCN), and diffusion model into a unified framework. HDiffTG leverages the strengths of these techniques to significantly improve pose estimation accuracy and robustness while maintaining a lightweight design. The Transformer captures global spatiotemporal dependencies, the GCN models local skeletal structures, and the diffusion model provides step-by-step optimization for fine-tuning, achieving a complementary balance between global and local features. This integration enhances the model’s ability to handle pose estimation under occlusions and in complex scenarios. Furthermore, we introduce lightweight optimizations to the integrated model and refine the objective function design to reduce computational overhead without compromising performance. Evaluation results on the Human3.6M and MPI-INF-3DHP datasets demonstrate that HDiffTG achieves state-of-the-art (SOTA) performance on the MPI-INF-3DHP dataset while excelling in both accuracy and computational efficiency. Additionally, the model exhibits exceptional robustness in noisy and occluded environments. Source codes and models are available at https://github.com/CirceJie/HDiffTG Yajie Fu, Chaorui Huang, Junwei Li 0009, Hui Kong 0001, Yibin Tian, Huakang Li, Zhiyuan Zhang 0004 |
IJCNN | 5 |
| 2025 | TS-Diff: Two-Stage Diffusion Model for Low-Light RAW Image EnhancementabstractThis paper presents a novel Two-Stage Diffusion Model (TS-Diff) for enhancing extremely low-light RAW images. In the pre-training stage, TS-Diff synthesizes noisy images by constructing multiple virtual cameras based on a noise space. Camera Feature Integration (CFI) modules are then designed to enable the model to learn generalizable features across diverse virtual cameras. During the aligning stage, CFIs are averaged to create a target-specific CFIT, which is fine-tuned using a small amount of real RAW data to adapt to the noise characteristics of specific cameras. A structural reparameterization technique further simplifies CFITfor efficient deployment. To address color shifts during the diffusion process, a color corrector is introduced to ensure color consistency by dynamically adjusting global color distributions. Additionally, a novel dataset, QID, is constructed, featuring quantifiable illumination levels and a wide dynamic range, providing a comprehensive benchmark for training and evaluation under extreme low-light conditions. Experimental results demonstrate that TS-Diff achieves state-of-the-art performance on multiple datasets, including QID, SID, and ELD, excelling in denoising, generalization, and color consistency across various cameras and illumination levels. These findings highlight the robustness and versatility of TS-Diff, making it a practical solution for low-light imaging applications. Source codes and models are available at https://github.com/CircccleK/TS-Diff Zhiyuan Zhang 0004, Jiangnan Xia, Jianghan Cheng, Junwei Li 0009, Yibin Tian, Hui Kong 0001 |
IJCNN | 7 |
| 2025 | CSSA-Fusion: Channel Selective and Spatial Alignment Infrared-Visible Image FusionabstractInfrared-visible image fusion aims to integrate complementary information from two modalities to generate images with enriched semantic content. However, existing methods often neglect two critical aspects: the design of a local–global feature enhancement architecture and spatial alignment. To address these challenges, we propose Channel Selective and Spatial Alignment Fusion (CSSA-Fusion), a novel framework composed of two synergistic modules. The first is a selective channel and redundancy suppression module, which introduces a dual-branch selective channel attention mechanism to jointly capture local saliency and global channel importance for enhanced feature representation, and an informativeness–redundancy separation strategy to suppress redundant information while preserving discriminative features. The second is a directional feature processing module, consisting of a mechanism that decouples and recombines modality-specific and common representations to mitigate mutual interference, and a spatial alignment module that performs geometric alignment via horizontal and vertical coordinate decomposition to correct spatial discrepancies between modalities. Extensive experiments on benchmark datasets demonstrate that CSSA-Fusion consistently outperforms state-of-the-art deep learning methods on multiple quality metrics. The fused images exhibit superior visual quality with well-preserved textures and enhanced semantic details. Zhongrui Xiao, Zhiyuan Zhang 0004, Yibin Tian |
MMAsia | 6 |
| 2025 | MVTS: Multimodal Visual-Tactile Sensor Using a Single CameraabstractVision and tactile are the most widely used perception modes in robotic interactions. We propose a Multimodal Visual-Tactile Sensor (MVTS) using a single camera to synchronously capture visual and tactile information. Unlike the traditional design with a fixed opaque gel layer, the MVTS consists of a color camera, a transparent elastomer layer embedded with color markers, and multiple LEDs. It can acquire images for vision like a normal camera, even during object contact. It can also use the markers to acquire tactile information while interacting with objects. In order to simultaneously acquire and separate tactile and visual information in single-shot, we designed a multimodal information separation framework based on a multi-task learning deep neural network. It adopts a shared feature encoder and two separate and parallel decoders to restore visual images and extract tactile maps. The MVTS was evaluated on multiple tasks such as force estimation, contact surface reconstruction, texture reconstruction, and volume estimation of grasped simple-shaped objects, and the results show that it has good performance in multimodal information acquisition, which can help robots better interact with the environment. Dajiang Lu, Xiaopin Zhong, Yibin Tian, Zongze Wu 0001 |
SMC | 4 |
| 2025 | VDGPG: A Virtual Data-Guided Prompt Generation Framework for Incremental Learning with Application to Wafer Defect DetectionabstractAlthough convolutional neural networks have been widely used for wafer defect detection in semiconductor manufacturing, they typically rely on static offline datasets to train models. These models show strong reliability when detecting known defect types, but struggle with unknown ones, posing challenges in model adaptation and leading to high maintenance costs. Incremental Learning (IL) offers a solution that allows models to continuously adapt to new types of defects without accessing full historical data. This paper introduces a Virtual Data-Guided Prompt Generation (VDGPG) framework, a novel IL approach for wafer defect detection that integrates prompt-guided learning and dual-branch virtual data generation. Specifically, VDGPG assigns task-specific prompt vectors to individual attention heads, using a channel attention gating mechanism and similarity computation to select prompt vectors from a prompt pool. This enables the model to focus more effectively on the relevant features for each category of defects. The dual-branch virtual data generation module generates diverse virtual samples, with a special emphasis on contour edge generation, which guides the model to learn features of potential new categories proactively. Experiments using the public WM-811K dataset demonstrate that VDGPG achieves significant performance improvements in wafer defect detection over existing IL methods. Bingwen Liu, Yibin Tian, Shanglei Chai, Zhiyuan Zhang 0004 |
SMC | 2 |
| 2025 | Detection of Incomplete Root Canal Obturations in Dental X-ray Images via Spatial-Semantic Attention and Dynamic Feature CalibrationabstractTo address the challenges of low resolution, loss of small target features, and interference from complex anatomical structures in detecting incomplete root canal obturations in dental periapical radiographs, this article proposes an improved YOLOv8 model. First, we design a Convolution module with Space-to-Depth Transformation (SDT-Conv) that preserves feature map resolution through spatial depth-wise separable convolutions, effectively mitigating loss of small targets caused by downsampling operations. Second, we construct a Dynamic Iterative Token Aggregator (DITA) architecture that enhances global feature representation through hyper-token spatial aggregation and semantic correlation, while employing a spatial-semantic dual-stream attention mechanism to strengthen multiscale feature fusion capabilities, thereby providing richer feature information for the entire network. Finally, we embed an Efficient Multiscale Attention (EMA) dynamic calibration mechanism in the detection head, which optimizes feature responses through cross-channel weight adaptation, enabling the model to precisely localize small object boundaries. The experimental results demonstrate that the improved model achieves 81.5% mAP@50 on the validation set, representing a 12.8% improvement over YOLOv8n. It effectively overcomes the challenges posed by variations in obturation materials, dental structure occlusions, and low-contrast interference. Zhiqi Ren, Shanglei Chai, Zhiyuan Zhang 0004, Xueyang Zhang, Yibin Tian |
SMC | 5 |
| 2025 | YOLOv8-CTCD: An Improved YOLOv8 for Cherry Tomato Cluster Detection in Robotic HarvestingabstractCherry tomato harvesting is generally performed manually. Robotic harvesting is gaining increasing interest from both academia and industry. This paper proposes a cherry tomato cluster detection algorithm based on YOLOv8, named YOLOv8-CTCD. First, the YOLOv8 input channels are adjusted to enable 4-channel RGB-D images as input. Subsequently, a CARAFE-M module is designed to replace the upsampling method in YOLOv8n. It maintains a lightweight architecture while achieving a larger receptive field, allowing effective aggregation of contextual information. In addition, it assigns greater weight to more important features. Moreover, a C2f-MLCA module is introduced into YOLOv8, which integrates information from feature maps at different levels and enhances the network’s capability of feature extraction. It also integrates the SPPELAN module to strengthen its feature fusion capability. YOLOv8-CTCD has been evaluated using a private cherry tomato dataset obtained from a greenhouse farm. The experimental results show that it achieves an mAP@50 of 93.8% and an mAP@50:90 of 68%, which represents improvements of 2.1% and 3% over YOLOv8n, respectively. Shanglei Chai, Zhiyuan Zhang 0004, Yibin Tian |
SMC | 5 |
| 2024 | Teeth Segmentation from Bite-Wing X-Ray Images by Integrating Nested Dual UNet with Swin TransformersabstractIn medical practice, the precision of image segmentation is crucial for diagnosis and treatment evaluations. Specifically, in dentistry, accurate teeth segmentation from bite-wing images is important for automatic and objective evaluations of root canal treatments. This study introduces$\text{Swin}-\mathrm{U}^{2} \text{Net}$, a model merging the nested dual UNet with residual U-block and Swin Transformers. It combines the local feature extraction capability of the former and the global attention and context understanding of the latter. It has been evaluated for tooth root segmentation using 500 bite-wing dental x-ray images obtained from a root canal treatment clinic. It achieved the best segmentation outcome in terms of Intersection over Union (IOU) and the third best result in terms of Dice Similarity Coefficient (DSC) with the second least amount of network parameters among six UNet-like models, thus it is effective and efficient. Yibin Tian, Zhiyuan Zhang 0004, Xueyang Zhang, Bingran Du |
SMC | 2 |
| 2024 | Defect recognition in sonic infrared imaging by deep learning of spatiotemporal signals
Jinfang Xie, Xinlin Wu, Xiaoyan Han, Yibin Tian |
Eng. Appl. Artif. Intell. | 7 |
| 2020 | Global Context Aware Convolutions for 3D Point Cloud UnderstandingabstractRecent advances in deep learning for 3D point clouds have shown great promises in scene understanding tasks thanks to the introduction of convolution operators to consume 3D point clouds directly in a neural network. Point cloud data, however, could have arbitrary rotations, especially those acquired from 3D scanning. Recent works show that it is possible to design point cloud convolutions with rotation invariance property, but such methods generally do not perform as well as translation-invariant only convolution. We found that a key reason is that compared to point coordinates, rotation-invariant features consumed by point cloud convolution are not as distinctive. To address this problem, we propose a novel convolution operator that enhances feature distinction by integrating global context information from the input point cloud to the convolution. To this end, a globally weighted local reference frame is constructed in each point neighborhood in which the local point set is decomposed into bins. Anchor points are generated in each bin to represent global shape features. A convolution can then be performed to transform the points and anchor features into final rotation-invariant features. We conduct several experiments on point cloud classification, part segmentation, shape retrieval, and normals estimation to evaluate our convolution, which achieves state-of-the-art accuracy under challenging rotations. Zhiyuan Zhang 0004, Binh-Son Hua, Yibin Tian, Sai-Kit Yeung |
3DV | 4 |