VLDB 2026 Research / reviewers in the wild / expert
Bin Hu 0021
dblp:00/6381-21
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-2974-1166ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 43% 3D vision · 28% Image recognition and object detection · 14% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › generative adversarial network › GAN inversion
3D GAN inversion |
1.0 | 1 | 2026 | ReE3D: Boosting Novel View Synthesis for Monocular Images Using Residual Encoders · IEEE Trans. Multim. 2026 |
Machine learning › Generative modeling
generative adversarial network |
1.0 | 1 | 2026 | ReE3D: Boosting Novel View Synthesis for Monocular Images Using Residual Encoders · IEEE Trans. Multim. 2026 |
Computer vision › 3D vision
novel view synthesis |
1.0 | 1 | 2026 | ReE3D: Boosting Novel View Synthesis for Monocular Images Using Residual Encoders · IEEE Trans. Multim. 2026 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.7 | 1 | 2023 | RSNet: Relation Separation Network for Few-Shot Similar Class Recognition · IEEE Trans. Multim. 2023 |
Computer vision › Image recognition and object detection
visual recognition |
0.7 | 1 | 2023 | RSNet: Relation Separation Network for Few-Shot Similar Class Recognition · IEEE Trans. Multim. 2023 |
Image and video processing › super-resolution
image super-resolution |
0.6 | 1 | 2022 | Cross View Capture for Stereo Image Super-Resolution · IEEE Trans. Multim. 2022 |
Image and video processing › super-resolution › image super-resolution
stereo image super-resolution |
0.6 | 1 | 2022 | Cross View Capture for Stereo Image Super-Resolution · IEEE Trans. Multim. 2022 |
Computer vision › 3D vision
3d scene understanding |
0.3 | 1 | 2026 | ReE3D: Boosting Novel View Synthesis for Monocular Images Using Residual Encoders · IEEE Trans. Multim. 2026 |
Methods — techniques the papers use, named apart from their topics
residual encoder · 1.0latent code optimization · 1.0geometric loss · 1.0relation separation network · 0.7spatial perception module · 0.6cross-view attention · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Intelligent spatiotemporal computing via transformer-graph fusion for sustainable urban mobility in heterogeneous computer systems
Yibo Peng, Bin Hu 0021 |
Future Gener. Comput. Syst. | 4 |
| 2026 | DBL: Dual-Level balanced learning for long-Tailed classification
Zheng Wu 0004, Kehua Guo, Bin Hu 0021, Xiangyuan Zhu, Rui Ding 0017 |
Pattern Recognit. | 4 |
| 2026 | Exploring compression transferability for model pruning in intelligent transportation
Kehua Guo, Shengxiong Fan, Bin Hu 0021 |
Pattern Recognit. Lett. | 4 |
| 2026 | ReE3D: Boosting Novel View Synthesis for Monocular Images Using Residual EncodersabstractIn recent years, novel view synthesis from a monocular image has become a research hot-spot that attracts significant attention. Some recent work identifies latent vectors for high-quality view generation via iterative optimisation, which is a time-consuming process. In contrast, some others utilise an encoder learning a mapping function to approximately estimate optimal latent codes, which significantly reduces its processing time but sacrifices reconstruction quality. Consequently, how to balance synthesis quality and its generation efficiency still remains challenging. In this paper, we propose a residual-based encoder to incorporate with a 3D Generative Adversarial Networks (GAN), named ReE3D, for novel view synthesis. It applies an iterative prediction of latent codes to ensure much higher quality of novel view synthesis with an insignificant increase of processing time when compared to existing encoder-based 3D GAN inversion methods. Additionally, we enforce a novel geometric loss constraint on the encoder to predict view-invariant latent codes, thus effectively mitigating the trade-off between geometric and texture quality in 3D GAN inversion. Extensive experimental results demonstrate that our extended encoder-based method has achieved best trade-off performance in terms of novel view synthesis quality and its execution time. Our method has gained comparable synthesis quality with exponentially decreased processing time when compared to iterative optimisation methods, while improved synthesis performance of encoder-based methods significantly. Kehua Guo, Tianyu Chen 0004, Bin Hu 0021, Zheng Wu 0004, Shaojun Guo, Hui Fang 0003 |
IEEE Trans. Multim. | 4 |
| 2026 | Boosting Adversarial Training With Mitigating Hard Sample InterferenceabstractAdversarial training (AT) has shown impressive advantages in maintaining accuracy and enhancing robustness against adversarial examples. However, most existing AT techniques jointly optimize clean-example accuracy and adversarial-example robustness as dual objectives. This setting introduces an often overlooked issue; when optimizing hard samples near the decision boundary, the model may bolster robustness at the expense of accuracy or preserve accuracy to the detriment of robustness. To alleviate the accuracy-robustness sacrifices induced by hard samples, we propose mitigating hard sample interference (MHSI) from a sample-intervention perspective. MHSI aims to reduce the instability caused by hard samples during AT. Specifically, we introduce a weighted adaptive (WA) mechanism that strengthens the model's learning of clean samples, thereby reducing the negative impact of hard samples on accuracy. In addition, guided by an analysis of the gradient norm and the Hessian matrix, we design a dynamic calibration (DC) strategy that dynamically calibrates the probability outputs of hard samples to mitigate their damage to robustness. With these two modules, our approach significantly improves robustness without sacrificing accuracy. Extensive experiments on CIFAR-10, CIFAR-100, Tiny ImageNet, and SVHN demonstrate that MHSI effectively improves both accuracy and robustness and outperforms state-of-the-art methods under glass-box attacks. Notably, under an $l_{\infty }$ attack, MHSI yields up to a 6.22% robustness gain over the AT baseline. Our code is available at https://github.com/hubin111/MHSI. Bin Hu 0021, Kehua Guo, Shaojun Guo |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Enhancing robustness of backdoor attacks against backdoor defenses
Bin Hu 0021, Kehua Guo, Hui Fang 0003 |
Expert Syst. Appl. | 1 |
| 2025 | Backdoor Defense in Transportation Cyber-Physical Systems Using Frequency Domain Hybrid DistillationabstractIn the context of transportation cyber-physical systems (T-CPS), backdoor attacks leveraging traffic images have emerged as a significant security threat. As T-CPS increasingly relies on visual information, such as real-time images captured by traffic cameras, for tasks like traffic sign recognition and autonomous driving, the risk of image-based backdoor attacks has grown substantially. Although various detection-based defense techniques have shown some success in identifying backdoored models, they often fail to fully eliminate backdoor effects, leaving residual security risks. To address this challenge, we propose a Frequency-Domain Hybrid Distillation (FDHD) method for backdoor defense, which effectively weakens the association between backdoor triggers and target labels by combining distillation mechanisms in both the frequency and pixel domains. Furthermore, we design a loss function that integrates feature reconstruction with adaptive alignment, enhancing the student network’s ability to mimic the teacher network and thereby bolstering the backdoor defense capability. Extensive experiments conducted by FDHD on multiple benchmark datasets against the five latest attacks demonstrate that our proposed defense method effectively reduces backdoor threats while maintaining high accuracy in predicting clean samples. This approach will protect against image-based backdoor attacks in T-CPS and lay the foundation for enhancing future traffic safety. Bin Hu 0021, Kehua Guo, Zheng Wu 0004, Xianhong Wen, Xiaokang Zhou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Adaptive Alignment Contrastive Learning of Degradation Prediction for Blind Image Super-ResolutionabstractBlind super-resolution (BSR) is entering a new era focused on diverse and complex applications, where the tradeoff between generalization and performance prevents models from performing as they should. Model performance decreases when trained on multiple degraded images due to the inter-class and intra-class imbalances in degradation prediction, which consists of degradation sampling and estimation. The inter-class imbalance in degradation estimation causes inaccurate estimates, leading to severe artifacts in images. The intra-class imbalance in degradation sampling causes a long-tail problem, leading to model collapse and satisfactory results only in specific applications. To tackle these challenges, we propose adaptive alignment contrastive learning (AACL), which includes adaptive degradation sampling (ADS) and \(\sigma\) -alignment. ADS utilizes non-linear sampling by weighting the parameters of the degradation process for training uniformly degraded images, avoiding the long-tail problem. \(\sigma\) -alignment controls the SD among positive samples; we identify a subset with small degraded distance, which aids contrastive learning in extracting representations more effectively. We extend AACL to several CNN-based and Transformer-based methods by coming up with a 6 \(\times\) 6 fair architecture with degradation representation fusion block (DRFB) and degradation representation fusion group (DRFG). DRFB and DRFG are designed for degradation representation fusion and image reconstruction, respectively. We evaluate on six types of degradation, and the improvement experiments on synthesized images show that our method balances performance and generalization and is applicable to networks with different architectures. The comparison experiments show that our improved methods achieve promising results compared to SOTA methods. Code is available at: https://github.com/para999/AACL . Xianhong Wen, Bin Hu 0021, Xiangyuan Zhu, Tianyu Chen 0004, Kehua Guo |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Generalize Deep Neural Networks With Adaptive Regularization for ClassifyingabstractRegularization is a crucial technology to improve the generalization of deep neural networks. However, traditional regularization method approaches are all scenario-specific, because they are generally with ingeniously designed feature representations from input layer, hidden layer, and output layer, which increase the difficulty of model development and interpretation. To this end, a novel practical and flexible regularization method is presented to obtain higher generalization and interpretability. Specifically, the feature maps are decoupled by global suppression and partial suppression from various scales and locate the salient feature with strong low-resolution semantic information. Moreover, the guided discarding specification for feature decoupling by measuring the feature contributions to network decisions, leads to the logics with better interpretability. Subsequently, the max values of the feature map are suppressed by discarding the corresponding salient features. Comprehensive experiments demonstrate that the proposed adaptive regularization outperforms the state-of-the-art performance in image classification accuracy, generalization, and interpretability on several widely used datasets. And adaptive regularization helps the network to mine the connection between salient features, nonsalient features, and ground truth, encouraging the network to construct multiple layers of feature associations. Kehua Guo, Ze Tao, Bin Hu 0021, Xiaoyan Kui |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | Realistic medical image super-resolution with pyramidal feature multi-distillation networks for intelligent healthcare systems
Kehua Guo, Jianguang Ma, Feihong Zhu, Bin Hu 0021, Haoming Zhou |
Neural Comput. Appl. | 5 |
| 2023 | Medical Image Super-Resolution Based on Semantic Perception Transfer LearningabstractMedical images are an important basis for doctors to diagnose diseases, but some medical images have low resolution due to hardware technology and cost constraints. Super-resolution technology can reconstruct low-resolution medical images into high-resolution images and enhance the quality of low-resolution images, thus assisting doctors in diagnosing diseases. However, traditional super-resolution methods mainly learn the mapping relationships among modal pixels from low resolution to high resolution, lacking the learning of high-level semantic features, resulting in a lack of understanding and utilization of semantic information, such as reconstructed objects, object attributes, and spatial relationships between two objects. In this paper, we propose a medical image super-resolution method based on semantic perception transfer learning. First, we propose a novel semantic perception super-resolution method that empowers super-resolution models to perceive high-level semantics by transferring features of the image description generation network in natural language processing. Second, we construct a semantic feature extraction network and an image description generation network and comprehensively utilized image and text modal data to learn transferable, high-level semantic features. Third, we train an end-to-end, semantic perception super-resolution model by fusing dynamic perceptual convolution, a semantic extraction network, and distillation polarization self-attention. Experiments show that semantic perception transfer learning can effectively improve the quality of super-resolution reconstruction. Kehua Guo, Xiaokang Zhou, Bin Hu 0021, Feihong Zhu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | RSNet: Relation Separation Network for Few-Shot Similar Class RecognitionabstractAlthough deep learning methods have drastically improved the performance on visual recognition tasks in which large inter-class variances exist, similar-class recognition continues to pose significant challenges, mainly due to the close resemblance between similar classes. The challenge is further compounded in the case of few-shot learning because only a very small amount of training data is available; accordingly, a certain performance degradation has been observed when some few-shot methods are applied for classification tasks. To address the aforementioned issue, we propose a novel Relation Separation Network (RSNet) in this paper, aiming to boost few-shot learning by improving similar-class recognition performance. We assume that image features consist of common and private features, where the common features capture the basic attributes shared among similar classes and their private counterparts capture the unique attributes of each class. Our RSNet learns to decouple the common and private features of an image. As a result, the feature representation of an image is composed of two weakly associated but easily aligned components, and better classification performance is achieved by giving more attention to subtle features. Experimental results on the publicly available datasets miniImageNet, CUB, and CIFAR-FS show that the proposed model outperforms existing state-of-the-art methods. Specifically, compared to PT+MAP, RSNet improves the accuracy of classification on the CUB dataset by approximately 5% and that of similar-class classification by more than 10%. Kehua Guo, Changchun Shen, Bin Hu 0021, Min Hu 0007, Xiaoyan Kui |
IEEE Trans. Multim. | 3 |
| 2022 | RRL-GAT: Graph Attention Network-Driven Multilabel Image Robust Representation LearningabstractExploring the characterization laws of image data and improving the efficiency of image data characterization knowledge is essential to promote the development of the Internet of Things technology. Considering that images in the real world usually contain multiple objects, and the objects are closely dependent. For these reasons, it brings great challenges to the robust representation learning of multilabel images. In general, researchers model the relationship between objects based on a class activation map and use graph convolution to mine the dependencies between objects. However, graph structure data often contain noise, which means that the edges between nodes are sometimes not so reliable, and the relative importance of neighbors is also different. Based on this, our goal is to reduce noisy connections and false connections between objects, eliminate multilabel image representation bias, and learn robust representations. Therefore, we propose a robust representation learning method for multilabel images driven by graph attention network (RRL-GAT). Specifically, to reduce the accidental false connection of objects in the image, we propose the class attention graph convolution module (C-GAT) to mine the strong association structure between categories. Besides, for the dynamic correlation between objects in the image, we propose an adaptive graph attention convolution module (A-GAT) to capture the subtle dynamic dependencies in the image. The results on two authoritative data sets show that our method is significantly better than all current state-of-the-art methods. Besides, the visualization results show that RRL-GAT can capture the semantic relationship of a specific input image and has sufficient recognizability. Bin Hu 0021, Kehua Guo, Xiaokang Wang 0001, Jian Zhang 0048, Di Zhou 0009 |
IEEE Internet Things J. | 1 |
| 2022 | Lightweight Image Super-Resolution With Expectation-Maximization Attention MechanismabstractIn recent years, with the rapid development of deep learning, super-resolution methods based on convolutional neural networks (CNNs) have made great progress. However, the parameters and the required consumption of computing resources of these methods are also increasing to the point that such methods are difficult to implement on devices with low computing power. To address this issue, we propose a lightweight single image super-resolution network with an expectation-maximization attention mechanism (EMASRN) for better balancing performance and applicability. Specifically, a progressive multi-scale feature extraction block (PMSFE) is proposed to extract feature maps of different sizes. Furthermore, we propose an HR-size expectation-maximization attention block (HREMAB) that directly captures the long-range dependencies of HR-size feature maps. We also utilize a feedback network to feed the high-level features of each generation into the next generation’s shallow network. Compared with the existing lightweight single image super-resolution (SISR) methods, our EMASRN reduces the number of parameters by almost one-third. The experimental results demonstrate the superiority of our EMASRN over state-of-the-art lightweight SISR methods in terms of both quantitative metrics and visual quality. The source code can be downloaded athttps://github.com/xyzhu1/EMASRN. Xiangyuan Zhu, Kehua Guo, Bin Hu 0021, Min Hu 0007, Hui Fang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Cross View Capture for Stereo Image Super-ResolutionabstractStereo image super-resolution exploits additional features from cross view image pairs for high resolution (HR) image reconstruction. Recently, several new methods have been proposed to investigate cross view features along epipolar lines to enhance the visual perception of recovered HR images. Despite the impressive performance of these methods, global contextual features from cross view images are left unexplored. In this paper, we propose a cross view capture network (CVCnet) for stereo image super-resolution by using both global contextual and local features extracted from both views. Specifically, we design a cross view block to capture diverse feature embeddings from the views in stereo vision. In addition, a cascaded spatial perception module is proposed to redistribute each location in feature maps according to the weight it occupies to make the extraction of features more effective. Extensive experiments demonstrate that our proposed CVCnet outperforms the state-of-the-art image super-resolution methods to achieve the best performance for stereo image super-resolution tasks. The source code is available at https://github.com/xyzhu1/CVCnet. Xiangyuan Zhu, Kehua Guo, Hui Fang 0003, Bin Hu 0021 |
IEEE Trans. Multim. | 6 |
| 2021 | Toward Anomaly Behavior Detection as an Edge Network Service Using a Dual-Task Interactive Guided Neural NetworkabstractHow to use artificial intelligence technology to mine human abnormal behavior from considerable video data generated by the Internet-of-Things system has been intensively studied for a long time. Existing deep learning anomaly detection algorithms deployed in the cloud typically perform supervised learning based on constant kinds of abnormal behavior data. However, this supervised learning model with preset abnormal behavior categories ignores the diversity and unpredictability of abnormal occurrences in open scenarios. Thus, we propose an abnormal behavior detection algorithm as an edge network service by combining the advantages of cloud computing and the efficiency of edge networks. This method combines the double verification of global behavior detection and local fine-grained action cycle alignment to detect whether a behavior is abnormal. Moreover, to enable abnormal behavior detection models to predict test samples whose categories do not appear during the training stage, we propose an active label learning algorithm based on cycle clustering, which not only improves the efficiency of data transmission between the edge and the cloud but also makes model updates in the cloud more efficient. Extensive and quantitative experimental results show that our method can not only accurately detect abnormal human behavior at the edge of limited resources but also has strong robustness under the interference of test samples of unknown categories. Kehua Guo, Bin Hu 0021, Jianhua Ma 0002, Ze Tao, Jian Zhang 0048 |
IEEE Internet Things J. | 2 |