VLDB 2026 Research / reviewers in the wild / expert
Hongqian Wang
dblp:153/0016
· DBLP profile ↗
11ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-1432-5012ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GFPP-MAE: gradient-guided frequency reconstruction and position predictions advance MAE for 3D CT image segmentation
Yuping Peng, Xing Xiao, Chengliang Wang 0002, Hongqian Wang |
Multim. Syst. | 5 |
| 2025 | SMBA-MIL: SAM-Enhanced Multi-branch Attention Multi-instance Learning for Whole Slide Image Classification
Biyun Zhou, Chengliang Wang 0002, Chao Liao, Hongqian Wang |
ICIC (5) | 6 |
| 2025 | mimicSAM: Guided Knowledge Transfer for SAM via Key Regions and Boundary InformationabstractSegment Anything Model (SAM) has received wide attention for its excellent segmentation capability and generalization performance, but its vast image encoder limits its deployment in real-time applications. To solve this problem, many researchers adopt knowledge distillation to lighten the image encoder of SAM, but the current work on knowledge distillation for SAM suffers from two problems: one is ignoring the learning of inter-sample relations, and another is ignoring the transfer of boundary information. Therefore, we propose a distillation method, mimcSAM, which is based on MobileSAM and introduces two modules in the feature map extraction phase: the Regional Attention Distillation Module (RADM) and the Boundary Gradient Distillation Module (BGDM). The RADM explicitly models relationships between samples by filtering key regions and constructing a queue of regional features, enabling the student model to align its relational understanding of key regions across samples. BGDM employs the Sobel operator to compare gradients on feature maps of teacher and student models, guiding the student to learn the teacher model’s sensitivity to boundaries. Experiments on the SA-1B, COCO, and LVIS datasets show that our method is 1.8% higher on mAP and 1.5% higher on mIoU than the baseline. Compared with the current relational distillation method, our method has the most significant improvement, which proves its superiority in improving the segmentation accuracy. Yixing Ma, Chengliang Wang 0002, Lingqiu Zeng, Hongqian Wang |
IJCNN | 6 |
| 2024 | US-SAM: An Automatic Prompt Sam For Ultrasound ImageabstractSegment anything model (SAM) has shown promising segmentation capabilities , however its performance significantly declines when applied to ultrasound images. Many works have emerged to address this issue, but still have two deficiencies: 1) SAM fine-tuning relies on manual prompts, which only allowing semiautomatics segmentation; 2) They exclusively utilize the frozen SAM encoder as the sole image encoder, which lacks pathological information. In this paper, we propose US-SAM which improve SAM with three modules: Pathological Extractor (PE), Fusion Module (FM), and Automatic Prompt Module (APM). PE extracts semantic information of lesions within the ultrasound image. Furthermore, FM fuses the features from PE and SAM image encoder. Finally, APM automatically generates prompts required by SAM using the fused features and uncertainty map. Through extensive experiments on two public ultrasound datasets BUSI and TN3k, our method outperforms other medical SAM methods by nearly 17% in Dice and IOU scores without any prompt from human. Yuteng Wang, Zhongshi He, Hongqian Wang |
ICME | 6 |
| 2024 | AIM-MIL: Adversarial Instance Mining for Robust Multi-instance Learning in Whole Slide Image Classification
Biyun Zhou, Chengliang Wang 0002, Hongqian Wang |
ICONIP (9) | 6 |
| 2024 | NFE-Net: Detection and Segmentation of Thyroid Nodules in Ultrasound Images Based on Nodule Feature EnhancedabstractDeep learning-based methods are commonly used for thyroid nodules detection and segmentation in ultrasound images, but the shape and size of the nodules vary greatly, there are also solid nodules that closely resemble the background and device-induced artifacts, making it difficult to accurately localize and segment the nodules. In this paper, NFE-Net is proposed to address the above difficulties, which designs receptive field enhancement path (RFEP) and texture-boundary guidance path (TBGP) in the neck of the Mask RCNN for nodule feature enhancement, so as to improve the localization and segmentation performance of network. RFEP enhances the learning ability of the network for nodule’s scale by introducing multi-scale receptive field enhancement module (RFEM); TBGP computes the channel correlation to extract the texture and boundary features of the nodule in information extraction module (IEM), further, uses true texture and boundary mask for deep supervised learning, finally fuses the upper and lower layers of the features by feature fusion block (FFB). Experimental results on the public thyroid datasets TN3K, DDTI show that our approach outperforms six state-of-the-art methods. Zhaoxin Long, Supeng Yin, Chengliang Wang 0002, Hongqian Wang |
IJCNN | 7 |
| 2024 | Predict EGFR Mutation Status on CT Images Using Texture and Contour Enhanced Masked AutoencodersabstractThe EGFR mutation status significantly influences targeted therapy for non-small cell lung cancer. In recent years, there has been significant progress in non-invasive EGFR mutation status prediction studies based on chest CT images. However, these studies commonly rely on extensive private data for training rather than small-scale publicly available dataset, thereby failing to overcome the dependency on high-cost large-scale annotated data. Additionally, these studies generally neglect the texture and contour features with strong discriminative power, leading to insufficient performance. This paper proposes a two-stage framework for EGFR mutation status prediction. Initially, we utilize self-supervised Masked Autoencoders (MAE) to pre-train the encoder on in-domain chest CT images, overcoming the problem of insufficient annotated data to reduce the dependency on high-cost annotated data. Subsequently, fine-tune the encoder to make it suitable for downstream EGFR prediction. Simultaneously, we propose texture and contour enhanced MAE (TCMAE), designing Multi-layer Features Aggregation Module (MFAM) to fully exploit multi-layer semantic features, introducing Texture and Contour Prediction Module (TCPM) to enhance the model’s capability in extracting texture and contour features through multitask learning, utilizing Modified Spectral Block (MSB) to adjust the weighting between high and low frequency features. Experiments demonstrate that despite using only a small-scale public dataset, NSCLC-Radiogenomics, the proposed method still achieves high accuracy. Yuping Peng, Zhongshi He, Chengliang Wang 0002, Hongqian Wang |
IJCNN | 7 |
| 2024 | Visual Navigation by Fusing Object Semantic FeatureabstractThe key of object goal visual navigation is to learn the spatial relationships between environmental objects and assess their semantic correlations with the target object. We propose an end-to-end visual navigation model based on deep reinforcement learning, called G2SNet, which consists of two feature maps and a specialized fusion feature network: GloVe Feature Map (GFM), Sbbox Feature Map (SFM), and GloVe fusion Network (GNet). GFM represents the position and the semantic information of the objects contained in the observation image, which addresses the issue of the interference in the target recognition caused by the complex background information in the observed image. SFM provides the object sizes in the field of view to assist the distance judgment. GNet relies entirely on network learning to compute the semantic correlations and spatial positional relationships among objects in the environment, enabling the agent to possess better generalization capabilities. This allows learning of spatial relationships between objects in GFM. Experiments on AI2-THOR demonstrate the effectiveness of our proposed three new structures, and the average SPL of the four known scenarios is increased by 16.6%. Chengliang Wang 0002, Zhongshi He, Hongqian Wang |
SMC | 6 |
| 2024 | Enhancing Autofocus Performance through Predictive Motion-Targeting and Self-Attention in a Deep Reinforcement Learning FrameworkabstractIn focusing tasks on moving targets, traditional methods that rely on maximizing contrast struggle to capture moving objects due to insufficient focusing speed. Deep learning-based methods have attempted to directly predict the optimal focal length for the target; however, due to low prediction accuracy, they often lead to out-of-focus situations when capturing moving objects. In recent years, some approaches have utilized reinforcement learning to automatically explore focal length adjustment patterns, thus achieving better results than traditional methods. However, these approaches have not considered the motion characteristics of the targets, leading to a need for further improvement in focusing performance. To overcome these limitations, we introduce a motion-based feature and deep reinforcement learning-driven autofocus algorithm named MF-DRLAF (Motion Features based Deep Reinforcement Learning Autofocus Model) for moving targets. This novel method tracks the object, predicts its motion state through feature extraction, and uses deep reinforcement learning to dynamically adjust the focus. We utilize a self-attention mechanism to adaptively learn various motion patterns and employ a feature pool structure to enhance processing efficiency. Experiments and real-world testing on a Google Pixel3 demonstrate that our approach significantly enhances autofocus performance on moving objects, highlighting its potential for broader imaging applications. This approach offers a promising direction for future development in autofocus technology. Xiaolin Wei, Ruilong Yang, Chengliang Wang 0002, Hongqian Wang |
SMC | 6 |
| 2024 | Improve Deep Learning Autofocus with Depth Information Supervision and Current Focal Distance CuesabstractTraditional autofocus methods search for the optimal focal distance (FD) by evaluating image quality from focal stacks, resulting in time-consuming focusing processes. Recently, deep learning has being adopted for single-shot autofocus methods, which can predict the optimal FD directly from a single input image. However, these methods often suffer from low prediction accuracy due to the lack of global features and structured global supervisory information, as they rely solely on the image's region of interest (ROI) as input and a single value for supervision. We propose a deep learning network named MPFS (Multi-Head Network with Per-Pixel Focal Distance Supervision), which takes a full-frame photograph as input and uses the optimal focal distance per pixel for supervision, this method effectively addresses the issues of missing global features and insufficient supervisory information by leveraging these enhancements. Additionally, the network integrates current camera focal distance information to mitigate the scale ambiguity caused by the lack of absolute scale information. To validate the effectiveness of the proposed method, we designed an experiment using a dataset annotated with optimal FD per pixel. Experimental results on this dataset indicate that our approach achieves a 0.22 decrease in the Mean Absolute Error (Mae) metric compared to the state-of-the-art models, with improvements of 0.02 and 0.004 in$\boldsymbol{d}_{\mathbf{1}}$and$\boldsymbol{d}_{\mathbf{2}}$metrics. Xiaolin Wei, Ruilong Yang, Chengliang Wang 0002, Hongqian Wang |
SMC | 6 |
| 2024 | SMDNet: A Pulmonary Nodule Classification Model Based on Positional Self-Supervision and Multi-Direction AttentionabstractAccurate classification of pulmonary nodules holds importance in the early diagnosis of lung cancer. Unlike 2D models, 3D models can simultaneously utilize multiple slices as input to capture features. However, 3D models face challenges in capturing nodule features in different directions and discerning feature differences in various positions of computed tomography (CT). We introduce a pulmonary nodule classification model, SMDNet. Firstly, a multi-direction attention is proposed to capture nodule features from sagittal, coronal, and axial axes. Secondly, distinct labels are assigned to the cubes at different cropping positions from CT for binary classification to capture local differences. Besides, gradient boosting decision tree (GBDT) is employed to combine shallow features with deep features to improve accuracy. Comparative experimental results on the largest publicly available dataset of pulmonary nodules, LIDC-IDRI, showed that SMDNet achieves a 4.81% improvement in accuracy under identical data processing. Chengliang Wang 0002, Hongqian Wang |
SMC | 6 |