EDBT 2026 Demo / reviewers in the wild / expert
Linbo Qing
dblp:132/5829
· DBLP profile ↗
75ranked-venue papers
3as first author
58since 2021 · last 2026
0000-0003-3555-0005ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 1 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 2 first-author · 17 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A lightweight framework for robust object detection in adverse weather based on dual-teacher feature alignment
Hanjun Zheng, Shengjie Ye, Linbo Qing, Honggang Chen |
Neurocomputing | 4 |
| 2026 | RelPosGAR: Hierarchical relative position-aware interaction modeling for weakly supervised skeleton-based group activity recognition
Lindong Li, Linbo Qing, Liuyi Tao, Pingyu Wang, Honggang Chen, Owen Noel Newton Fernando, Weisi Lin |
Pattern Recognit. | 2 |
| 2026 | Beyond deceptive flatness: Dual-order solution for strengthening adversarial transferability
Pingyu Wang, Xingjian Zheng, Linbo Qing, Qi Liu 0005 |
Pattern Recognit. | 4 |
| 2026 | Progressive Reasoning-Based Group Activity RecognitionabstractGroup activity recognition (GAR) plays a crucial role in computer vision, enabling the exploration and comprehension of human behavior patterns. Existing methods mainly focus on dyad-level interactions within a group, but sociological studies have highlighted the importance of individual features, subgroup-level interactions, and overall group structure for understanding group activities. Therefore, we propose a new framework, the progressive group activity reasoning model (PGAR), which models these four aspects for GAR. Initially, we construct a person-person graph (PPG) using individual features to capture dyadic interactions. Subsequently, the PPG is fed into a novel ingredient graph model (Ingredient-GNN) for capturing subgroup-level interactions. Finally, we fuse the dyad-level and subgroup-level interactions with global information of group structure, obtained through an F-Formation modeling module, to form comprehensive representations for GAR. The F-Formation modeling module decouples the group structure into position, orientation, and skeleton graphs, and subsequently performs attribute recoupling at the individual level using the designed Tri-Coupling Transformer to form a global representation of the group structure. Extensive experiments on four public datasets demonstrate that our final model effectively integrates multi-level representations for group activity understanding, with our F-Formation modeling module outperforming comparable methods that rely solely on non-visual data. Lindong Li, Linbo Qing, Wang Tang, Pingyu Wang, Haosong Gou, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | AsyReC: A Multimodal Graph-Based Framework for Spatio-Temporal Asymmetric Dyadic Relationship ClassificationabstractDyadic social relationships, which refer to relationships between two individuals who know each other through repeated interactions (or not), are shaped by shared spatial and temporal experiences. Current computational methods for modeling these relationships face three major challenges: (1) the failure to model asymmetric relationships, e.g., one individual may perceive the other as afriendwhile the other perceives them as anacquaintance, (2) the disruption of continuous interactions by discrete frame sampling, which segments the temporal continuity of interaction in real-world scenarios, and (3) the limitation to consider periodic behavioral cues, such as rhythmic vocalizations or recurrent gestures, which are crucial for inferring the evolution of dyadic relationships. To address these challenges, we propose AsyReC, a multimodal graph-based framework for asymmetric dyadic relationship classification, with three core innovations: (i) a triplet graph neural network with node-edge dual attention that dynamically weights multimodal cues to capture interaction asymmetries (addressing challenge 1); (ii) a clip-level relationship learning architecture that preserves temporal continuity, enabling fine-grained modeling of real-world interaction dynamics (addressing challenge 2); and (iii) a periodic temporal encoder that projects time indices onto sine/cosine waveforms to model recurrent behavioral patterns (addressing challenge 3). Extensive experiments on two public datasets demonstrate state-of-the-art performance, while ablation studies validate the critical role of asymmetric interaction modelling and periodic temporal encoding in improving the robustness of dyadic relationship classification in real-world scenarios. Our code is publicly available at: https://github.com/tw-repository/AsyReC. Wang Tang, Fethiye Irmak Dogan, Linbo Qing, Hatice Gunes |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image DerainingabstractRecently, deep image deraining models based on paired datasets have made a series of remarkable progress. However, they cannot be well applied in real-world applications due to the difficulty of obtaining real paired datasets and the poor generalization performance. In this paper, we propose a novel Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image Deraining framework, CSUD, to tackle the aforementioned challenges. During training with unpaired data, CSUD is capable of generating high-quality pseudo clean and rainy image pairs which are used to enhance the performance of deraining network. Specifically, to preserve more image background details while transferring rain streaks from rainy images to the unpaired clean images, we propose a novel Channel Consistency Loss (CCLoss) by introducing the Channel Consistency Prior (CCP) of rain streaks into training process, thereby ensuring that the generated pseudo rainy images closely resemble the real ones. Furthermore, we propose a novel Self-Reconstruction (SR) strategy to alleviate the redundant information transfer problem of the generator, further improving the deraining performance and the generalization capability of our method. Extensive experiments on multiple synthetic and real-world datasets demonstrate that the deraining performance of CSUD surpasses other state-of-the-art unsupervised methods and CSUD exhibits superior generalization capability. Code is available at https://github.com/GuangluDong0728/CSUD. Guanglu Dong, Tianheng Zheng, Yuanzhouhan Cao, Linbo Qing, Chao Ren 0002 |
CVPR | 4 |
| 2025 | Graph-based interactive knowledge distillation for social relation continual learningabstractAs multimedia advances, there is a growing need for machines to adeptly understand diverse social relations . Traditional methods for recognizing these relations, which are limited to a fixed number of classes, are ill-equipped for continual learning as new social interactions emerge. To address this prob-lem, we propose a pioneering Graph-based Interactive Knowledge Distillation (GI-KD) method for social relation continual learning. GI-KD, embedded in a class incremental learning structure, creates a balanced system where previously learned social relations and new knowledge are positioned at either end of the scale. The old and new knowledge is learned dynamically by adjusting the tilt of the balance. To achieve this balance, we propose a novel Libra loss function, which evaluate the relative contribution of old and new information and thus guides the adaptive fine-tuning of the model. We evaluate the GI-KD on three public social relation recognition (SRR) datasets, under different data distribution strategies. Our method shows a remarkable average 3.6% increase in incremental accuracy over current CIL techniques, effectively reducing catastrophic forgetting. Furthermore, GI-KD improves mAP and Acc by 4.6%, 5.4%, and 4.5%, respectively, compared to current CIL techniques, highlighting its strength in both continual learning and SRR. Wang Tang, Linbo Qing, Pingyu Wang, Lindong Li, Yonghong Peng |
Neurocomputing | 2 |
| 2025 | Spatio-temporal interactive reasoning model for multi-group activity recognition
Jianglan Huang, Lindong Li, Linbo Qing, Wang Tang, Pingyu Wang, Li Guo 0018, Yonghong Peng |
Pattern Recognit. | 3 |
| 2025 | DRFormer: A Discriminable and Reliable Feature Transformer for Person Re-IdentificationabstractAs person image variations are likely to cause a part misalignment problem, most previous person Re-Identification (ReID) works may adopt local feature partition or additional landmark annotations to acquire aligned person features and boost ReID performance. However, such approaches either only achieve coarse-grained part alignments without considering detailed image variations within each part, or require extra annotated landmarks to train an available pose estimation model. In this work, we propose an effective Discriminable and Reliable Transformer (DRFormer) framework to learn part-aligned person representations with only person identity labels. Specifically, the DRFormer framework consists of Discriminable Feature Transformer (DFT) and Reliable Feature Transformer (RFT) modules, which generate discriminable and reliable high-order features, respectively. For reducing the dimension of high-order features, the DFT module utilizes a Self-Attentive Kronecker Product (SAKP) algorithm to promote the representational capabilities of compressed features via a self-attention strategy. For eliminating the background noise, the RFT module mines the foreground regions to adaptively aggregate foreground features via a Gumbel-Softmax strategy. Moreover, the proposed framework derives from an interpretable motivation and elegantly solves part misalignments without using feature partition or pose estimation. This paper theoretically and experimentally demonstrates the superiority of the proposed DRFormer framework, achieving state-of-the-art performance on various person ReID datasets. Pingyu Wang, Xingjian Zheng, Linbo Qing, Bonan Li, Zhicheng Zhao 0001, Honggang Chen |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | A Stable and Efficient Data-Free Model Attack With Label-Noise Data GenerationabstractThe objective of a data-free closed-box adversarial attack is to attack a victim model without using internal information, training datasets or semantically similar substitute datasets. Concerned about stricter attack scenarios, recent studies have tried employing generative networks to synthesize data for training substitute models. Nevertheless, these approaches concurrently encounter challenges associated with unstable training and diminished attack efficiency. In this paper, we propose a novel query-efficient data-free closed-box adversarial attack method. To mitigate unstable training, for the first time, we directly manipulate the intermediate-layer feature of a generator without relying on any substitute models. Specifically, a label noise-based generation module is created to enhance the intra-class patterns by incorporating partial historical information during the learning process. Additionally, we present a feature-disturbed diversity generation method to augment the inter-class distance. Meanwhile, we propose an adaptive intra-class attack strategy to heighten attack capability within a limited query budget. In this strategy, entropy-based distance is utilized to characterize the relative information from model outputs, while positive classes and negative samples are used to enhance low attack efficiency. The comprehensive experiments conducted on six datasets demonstrate the superior performance of our method compared to six state-of-the-art data-free closed-box competitors in both label-only and probability-only attack scenarios. Intriguingly, our method can realize the highest attack success rate on the online Microsoft Azure model under an extremely low query budget. Additionally, the proposed approach not only achieves more stable training but also significantly reduces the query count for a more balanced data generation. Furthermore, our method can maintain the best performance under the existing defense models and a limited query budget. Xingjian Zheng, Linbo Qing, Qi Liu 0005, Pingyu Wang, Yu Liu 0123, Jiyang Liao |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Hypergraph Mamba Reasoning-Based Social Relation RecognitionabstractRecognizing social relations from images is crucial for improving machine perception of social interactions. Current studies mainly focus on exploring single-type relation reasoning frameworks, such as the relation between father, mother and son in a family. However, real-world scenarios often involve complex hybrid relations, such as friendships and professional relations, which pose a challenge for current methods due to the difficulty of establishing robust logical connections between these relations. In fact, in this hybrid social relation recognition setting, the interactions extend beyond dyadic to multipartite structures. To effectively explore these multipartite interactions, we propose a novel Hypergraph Mamba (HGM) framework. Specifically, we construct two hypergraphs, i.e., Person-Person Hypergraphs (PPH) and Person-Object Hypergraphs (POH), to model these high-order multipartite interactions. The HGM module performs social relation reasoning within these hypergraph structures, which includes a Vertex Selection Algorithm to mitigate inference confusion by filtering out confounders, and a Vertex Interaction Operator to find optimal global vertex neighborhoods by capturing long-range vertex dependencies. In addition, a Multilevel Transformer is proposed to adaptively align the PPH and POH inferred knowledge and visual signals to facilitate information fusion. We validate the effectiveness of our proposed HGM model on several public datasets and perform extensive ablation studies to elucidate the reasons contributing to its superior performance. Experimental results indicate that our HGM model achieves superior accuracy in predicting social relations compared to the state-of-the-art methods. Codes and datasets are available at: https://github.com/tw-repository/HGM-SRR. Wang Tang, Linbo Qing, Pingyu Wang, Lindong Li, Ce Zhu |
IEEE Trans. Image Process. | 2 |
| 2025 | Enhancing pain intensity evaluation via an attention-driven channel-spatial fusion network
Linbo Qing, Lindong Li, Risheng Xu |
Vis. Comput. | 2 |
| 2025 | Facial expression recognition based on local-global information reasoning and spatial distribution of landmark features
Kunhong Xiong, Linbo Qing, Lindong Li, Li Guo 0018, Yonghong Peng |
Vis. Comput. | 2 |
| 2024 | An Attention Transformer-Based Method for the Modelling of Functional Connectivity and the Diagnosis of Autism Spectrum Disorder
Linbo Qing, Yanteng Zhang, Xiaohai He, Yonghong Peng |
ICPR (12) | 2 |
| 2024 | Semantic and geometric information propagation for oriented object detection in aerial images
Xiaohai He, Honggang Chen, Linbo Qing, Qizhi Teng |
Appl. Intell. | 4 |
| 2024 | Unveiling group activity recognition: Leveraging Local-Global Context-Aware Graph Reasoning for enhanced actor-scene interactions
Linbo Qing, Jianglan Huang, Li Guo 0018, Yonghong Peng |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | MSE-Net: A novel master-slave encoding network for remote sensing scene classification
Hongguang Yue, Linbo Qing, Zhengyong Wang, Li Guo 0018, Yonghong Peng |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | A Multi-Attention Feature Distillation Neural Network for Lightweight Single Image Super-ResolutionabstractIn recent years, remarkable performance improvements have been produced by deep convolutional neural networks (CNN) for single image super-resolution (SISR). Nevertheless, a high proportion of CNN-based SISR models are with quite a few network parameters and high computational complexity for deep or wide architectures. How to more fully utilize deep features to make a balance between model complexity and reconstruction performance is one of the main challenges in this field. To address this problem, on the basis of the well-known information multi-distillation model, a multi-attention feature distillation network termed as MAFDN is developed for lightweight and accurate SISR. Specifically, an effective multi-attention feature distillation block (MAFDB) is designed and used as the basic feature extraction unit in MAFDN. With the help of multi-attention layers including pixel attention, spatial attention, and channel attention, MAFDB uses multiple information distillation branches to learn more discriminative and representative features. Furthermore, MAFDB introduces the depthwise over-parameterized convolutional layer (DO-Conv)-based residual block (OPCRB) to enhance its ability without incurring any parameter and computation increase in the inference stage. The results on commonly used datasets demonstrate that our MAFDN outperforms existing representative lightweight SISR models when taking both reconstruction performance and model complexity into consideration. For example, for × 4 SR on Set5, MAFDN (597K/33.79G) obtains 0.21 dB/0.0037 and 0.10 dB/0.0015 PSNR/SSIM gains over the attention-based SR model AFAN (692K/50.90G) and the feature distillation-based SR model DDistill-SR (675K/32.83G), respectively. Yongfei Zhang, Xinying Lin, Linbo Qing, Xiaohai He, Yi Li 0069, Honggang Chen |
Int. J. Intell. Syst. | 5 |
| 2024 | Time-varying neurodynamic optimization approaches with fixed-time convergence for sparse signal reconstruction
Xingxing Ju, Xinsong Yang, Linbo Qing, Jinde Cao, Dianwei Wang |
Neurocomputing | 3 |
| 2024 | DVC-Net: a new dual-view context-aware network for emotion recognition in the wild
Linbo Qing, Hongqian Wen, Honggang Chen, Rulong Jin, Yongqiang Cheng 0001, Yonghong Peng |
Neural Comput. Appl. | 1 |
| 2024 | Finite-Time Stabilization of Uncertain Delayed T-S Fuzzy Systems via Intermittent ControlabstractThis article focuses on finite-time$\mathcal {L}_{2}$stabilization of T–S fuzzy systems with time delays and parameter uncertainties via intermittent control. To cope with the effects of parameters uncertainties, time delays, and intermittent divergence simultaneously, a new finite-time stability lemma for intermittently controlled systems is presented. Then, a weighted 2-norm Lyapunov–Krasovskii functional (LKF) is established, which has the advantage that it is convenient to derive less conservative linear matrix inequality sufficient conditions and to overcome the difficulty in analyzing$\mathcal {L}_{2}$performance under intermittent control frameworks. Another advantage of our result over existing results is that the growth increment of the LKF on the noncontrolled interval can be larger than the decreasing magnitude on the controlled interval. The merits of the theoretical results are examined by a numerical example and a coupled Chua's circuit. Rongqiang Tang, Xinsong Yang, Peng Shi 0001, Zhengrong Xiang, Linbo Qing |
IEEE Trans. Fuzzy Syst. | 5 |
| 2024 | Progressive Graph Reasoning-Based Social Relation RecognitionabstractIdentifying relationships between people from images is essential for studying social activities and interactions, and this has significant potential to further the understanding of human social behaviors. Existing image-based research mainly explores social relationships at the dyadic level, i.e., recognizing pairwise relationships based on visual features of persons, objects, and scenes and their logical constraints. Notably, social relational structures are hierarchically nested, i.e., individuals and dyads are nested within group structures, as indicated in the social relations model (SRM) of social psychology. However, existing computer vision-based studies fail to consider hierarchical nested structures, thus overlooking some of the most important interactions, which leads to poor relation reasoning. To improve the performance of reasoning neural networks, we propose a novel SRM framework for progressive graph reasoning (PGR) to explore social interactions. Specifically, we construct individual–dyad and dyad–group graphs to progressively explore the impact of individuals and groups on recognition of dyadic relationships. A transformer is utilized to fuse visual features and graph reasoning knowledge into a comprehensive representation of social relationships. We demonstrate the effectiveness of the proposed model based on PGR using several public datasets and perform extensive ablation studies to explore the reasons behind its superior performance. Experimental results demonstrate that our proposed model successfully predicts social relationships with higher accuracy than state-of-the-art methods. Codes and datasets are available at:https://github.com/tw-repository/PGRSRR. Wang Tang, Linbo Qing, Lindong Li, Ce Zhu |
IEEE Trans. Multim. | 2 |
| 2024 | DAG-YOLO: A Context-Feature Adaptive fusion Rotating Detection Network in Remote Sensing ImagesabstractObject detection in remote sensing image (RSI) research has seen significant advancements, particularly with the advent of deep learning. However, challenges such as orientation, scale, aspect ratio variations, dense object distribution, and category imbalances remain. To address these challenges, we present DAG-YOLO, a one-stage context-feature adaptive weighted fusion network that incorporates through three innovative parts. First, we integrate 1D Gaussian Angle-coding with YOLOv5 to convert the angle regression task into a classification task, establishing a more robust rotating object detection baseline, GLR-YOLO. Second, we introduce the Dual Branch Context Adaptive Modeling module, which enhances feature extraction capabilities by capturing global context information. Third, we design an adaptive detect head with the Adaptive Global Feature Aggregation and Reweighting (AGFAR) module. AGFAR addresses feature inconsistency among different output layers of the Feature Pyramid Network, retaining useful semantic information and elevating detection accuracy. Extensive experiments on public datasets DOTA-v1.0, DOTA-v1.5, and UCAS-AOD showcase mAP scores of 77.75%, 73.79%, and 90.27%, respectively. Our proposed method has the best performance among the current mainstream SOTA methods, which proves its effectiveness in RSI object detection. Zhenjiang Guo, Xiaohai He, Linbo Qing, Honggang Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Efficient Information Modulation Network for Image Super-ResolutionabstractRecent researches have shown that the success of Transformers comes from their macro-level framework and advanced components, not just their self-attention (SA) mechanism. Comparable results can be obtained by replacing SA with spatial pooling, shifting, MLP, fourier transform and constant matrix, all of which have spatial information encoding capability like SA. In light of these findings, this work focuses on combining efficient spatial information encoding technology with superior macro architectures in Transformers. We rethink spatial convolution to achieve more efficient encoding of spatial features and dynamic modulation value representations by convolutional modulation techniques. The large-kernel convolution and Hadamard product are utilizated in the proposed Multi-orders Long-range convolutional modulation (MOLRCM) layer to imitate the implementation of SA. Moreover, MOLRCM layer also achieve long-range correlations and self-adaptation behavior, similar to SA, with linear complexity. On the other hand, we also address the sub-optimality of vanilla feed-forward networks (FFN) by introducing spatial awareness and locality, improving feature diversity, and regulating information flow between layers in the proposed Spatial Awareness Dynamic Feature Flow Modulation (SADFFM) layer. Experiment results show that our proposed efficient information modulation network (EIMN) performs better both quantitatively and qualitatively. Codes and supplementary materials link: https://github.com/liux520/EIMN. Xiao Liu 0022, Xiangyu Liao, Xiuya Shi, Linbo Qing, Chao Ren 0002 |
ECAI | 4 |
| 2023 | Efficient Parallel Multi-Scale Detail and Semantic Encoding Network for Lightweight Semantic SegmentationabstractIn this work, we propose PMSDSEN, a parallel multi-scale encoder-decoder network architecture for semantic segmentation, inspired by the human visual perception system's ability to aggregate contextual information in various contexts and scales. Our approach introduces the efficient Parallel Multi-Scale Detail and Semantic Encoding (PMSDSE) unit to extract detailed local information and coarse large-range relationships in parallel, enabling the recognition of object boundaries and object-level areas. By stacking multiple PMSDSEs, our network learns fine-grained details and textures along with abstract category and semantic information, effectively utilizing a larger range of surrounding context information for robust segmentation. To further enhance the network's receptive field without increasing computational complexity, the Multi-Scale Semantic Extractor (MSSE) at the end of the encoder is utilized for multi-scale semantic context extraction and detailed information encoding. Additionally, the Dynamic Weighted Feature Fusion (DWFF) strategy is employed to integrate shallow layer detail information and deep layer semantic information during the decoder stage. Our method can obtain multi-scale context from local to global, achieving efficiently low-level feature extraction to high-level semantic interpretation at different scales and in different contexts. Without bells and whistles, PMSDSEN obtains a better trade-off between accuracy and complexity on popular benchmarks, including Cityscapes and Camvid. Specifically, PMSDSEN attains 73.2% mIoU with only 0.9M parameters on the Cityscapes test set. Codes and supplementary materials link: https://github.com/liux520/PMSDSEN. Xiao Liu 0022, Xiuya Shi, Lufei Chen, Linbo Qing, Chao Ren 0002 |
ACM Multimedia | 4 |
| 2023 | Image classification based on self-distillation
Linbo Qing, Xiaohai He, Honggang Chen, Qiang Liu 0021 |
Appl. Intell. | 2 |
| 2023 | Principal relation component reasoning-enhanced social relation recognition
Wang Tang, Linbo Qing, Lindong Li, Li Guo 0018, Yonghong Peng |
Appl. Intell. | 2 |
| 2023 | A collaborative perception method of human-urban environment based on machine learning and its application to the case area
Jianlin Huang, Linbo Qing, Longmei Han, Jiajia Liao, Li Guo 0018, Yonghong Peng |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Self-supervised cycle-consistent learning for scale-arbitrary real-world single image super-resolution
Honggang Chen, Xiaohai He, Yuanyuan Wu 0001, Linbo Qing, Ray E. Sheriff |
Expert Syst. Appl. | 5 |
| 2023 | Relationship existence recognition-based social group detection in urban public spaces
Lindong Li, Linbo Qing, Li Guo 0018, Yonghong Peng |
Neurocomputing | 2 |
| 2023 | RestorNet: An efficient network for multiple degradation image restoration
Honggang Chen, Haosong Gou, Zhengyong Wang, Xiaohai He, Linbo Qing, Ray E. Sheriff |
Knowl. Based Syst. | 7 |
| 2023 | Multi-relation graph convolutional network for Alzheimer's disease diagnosis using structural MRI
Xiaohai He, Linbo Qing, Xiang Chen 0008, Yan Liu 0078, Honggang Chen |
Knowl. Based Syst. | 3 |
| 2023 | BDNet: A BERT-based dual-path network for text-to-image cross-modal person re-identification
Qiang Liu 0021, Xiaohai He, Qizhi Teng, Linbo Qing, Honggang Chen |
Pattern Recognit. | 4 |
| 2023 | Dynamically Optimized Human Eyes-to-Face Generation via Attribute VocabularyabstractGenerating face from human eyes, named eyes-to-face generation, is an interesting research topic of face synthesis, which has great potential in the field of public security. One of the main challenges in eyes-to-face generation is the unbalanced information between inputs and outputs, where the outputs are complete facial images while the inputs only contain limited information in the region of eyes. The existing methods generate faces directly from eyes without considering the possibly available facial information (e.g. facial attributes), resulting in inaccurate predictions and high uncertainty in those features less correlated with eyes (e.g. hairstyle, moustache, facial contour). To address this challenge, we propose a two-stage solution (named EA2F-GAN) to dynamically optimize eyes-to-face generation via attribute vocabulary. In addition, a dataset named TEAF is constructed based on the public datasets CelebA and LFW, containing 138,934 triples of eye image, attribute vocabulary, and face image. Sufficient experimental results show that, by incorporating additional facial attributes, our proposed approach can synthesize realistic face with high consistency to the original one, significantly overwhelming state-of-the-art methods. Xiaodong Luo, Xiaohai He, Xiang Chen 0008, Linbo Qing, Honggang Chen |
IEEE Signal Process. Lett. | 4 |
| 2023 | Unveiling Social Relations: Leveraging Interpersonal Similarity Learning for Social Relation RecognitionabstractIdentifying social relationships from images is a challenging yet promising research area with great potential for improving human health and enhancing our understanding of social networks. However, present endeavors in this field tend to concentrate on leveraging visual features for the exploration of social relationships, while disregarding certain concealed information that lies beneath these features, such as interpersonal similarity. These methodologies may result in inadequate visual data encoding, thereby imposing limitations on the accuracy of social relationship recognition. In light of this, we propose a novel framework that utilizes interpersonal similarities within images to provide more information for identifying social relationships, thereby mitigating the issue of insufficient feature exploration. Furthermore, our proposed framework incorporates an innova-tive CF-Loss function that effectively incentivizes the identifica-tion of accurate social relationships while penalizing incorrect identifications, ultimately bolstering the model's capacity to dis-criminate between distinct social relationships. Our experimental findings demonstrate the superiority of our proposed framework over state-of-the-art methods on public datasets, confirming its effectiveness and accuracy in identifying social relationships. Wang Tang, Linbo Qing, Haosong Gou, Li Guo 0018, Yonghong Peng |
IEEE Signal Process. Lett. | 2 |
| 2023 | Non-Weighted L2-Gain Analysis for Synchronization of Switched Nonlinear Time-Delay Systems With Random Injection AttacksabstractThe current paper is devoted to studying global asymptotic$H_{\infty }$drive-response synchronization for a kind of switched nonlinear time-delay systems with output random injection attacks (IAs). An attack-decomposition method is proposed to derive an attack-free signal, by which an observer is designed to estimate the state of the driving system. Then, two mode-dependent event-triggering mechanisms (MDETMs) are respectively designed for observer-controller (O-C) and controller-actuator (C-A) channels to save the communication resources as much as possible. In order to analyze the effects of the switching on the$H_{\infty }$performance, a mode-dependent discretized Lyapunov-Krasovskii functional (LKF) is developed, which has the merit of monotone decreasing on any time-interval and switching instants. Sufficient criteria are given to ensure the$H_{\infty }$synchronization with non-weighted$\mathcal {L}_{2}$-gain, whether the attack is related to the output or not. Numerical simulations are provided to verify the non-weighted$\mathcal {L}_{2}$-gain performance with low conservatism. Xinsong Yang, Qihan Qi, Peng Shi 0001, Zhengrong Xiang, Linbo Qing |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | CBW-MSSANet: A CNN Framework With Compact Band Weighting and Multiscale Spatial Attention for Hyperspectral Image Change DetectionabstractChange detection (CD), aims to detect the changing area of the same scene at different times, which is an important application of remote sensing images. As the key data source of CD, hyperspectral image (HSI) is widely used in CD technology because of its rich spectral-spatial information. However, how to mine the multi-level spatial information of dual-temporal hyperspectral images (HSIs) and focus on the features of the pixels to be classified individually remains a problem in the spatial attention mechanism (SAM). To make full use of the spectral-spatial information of HSIs, in this paper we propose a CNN framework with compact band weighting and multi-scale spatial attention (CBW-MSSANet) for HSI pixel-level CD. The main contributions of this article are as follows: 1) a new method of pseudo-label training sample selection based on k-means (KM) centroid distance is designed; 2) apply the compact band weighting (CBW) module to HSI CD to take full advantage of the spectral information of HSIs; 3) a multi-scale spatial attention (MSSA) module is developed for pixel-level CD, which can mine multi-level spatial information and pay more attention to the features of the pixels to be classified, and combine the spatial information of adjacent pixels to make it more conducive to pixel-level CD. Experimental results on four real HSI datasets demonstrated that the performance of MSSA surpasses the classical single-scale SAM, and CBW-MSSANet is superior to some representative CD methods. Xianfeng Ou, Liangzhen Liu, Bing Tu, Linbo Qing, Guoyun Zhang, Zifei Liang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Dual adaptive alignment and partitioning network for visible and infrared cross-modality person re-identification
Qiang Liu 0021, Qizhi Teng, Honggang Chen, Bo Li 0074, Linbo Qing |
Appl. Intell. | 5 |
| 2022 | UAMNer: uncertainty-aware multimodal named entity recognition in social media posts
Luping Liu, Mozhi Zhang, Linbo Qing, Xiaohai He |
Appl. Intell. | 4 |
| 2022 | Medical visual question answering based on question-type reasoning and semantic space constraint
Xiaohai He, Luping Liu, Linbo Qing, Honggang Chen, Yan Liu 0078, Chao Ren 0002 |
Artif. Intell. Medicine | 4 |
| 2022 | Feature separation and double causal comparison loss for visible and infrared person re-identification
Qiang Liu 0021, Xiaohai He, Mozhi Zhang, Qizhi Teng, Bo Li 0074, Linbo Qing |
Knowl. Based Syst. | 6 |
| 2022 | Fact-based visual question answering via dual-process system
Luping Liu, Xiaohai He, Linbo Qing, Honggang Chen |
Knowl. Based Syst. | 4 |
| 2022 | CMAFGAN: A Cross-Modal Attention Fusion based Generative Adversarial Network for attribute word-to-face synthesis
Xiaodong Luo, Xiang Chen 0008, Xiaohai He, Linbo Qing, Xinyue Tan |
Knowl. Based Syst. | 4 |
| 2022 | Weakly-supervised contrastive learning-based implicit degradation modeling for blind image super-resolution
Yongfei Zhang, Ling Dong, Linbo Qing, Xiaohai He, Honggang Chen |
Knowl. Based Syst. | 4 |
| 2022 | Cross-modal multi-relationship aware reasoning for image-text matching
Xiaohai He, Linbo Qing, Luping Liu, Xiaodong Luo |
Multim. Tools Appl. | 3 |
| 2022 | DualG-GAN, a Dual-channel Generator based Generative Adversarial Network for text-to-face synthesis
Xiaodong Luo, Xiaohai He, Xiang Chen 0008, Linbo Qing |
Neural Networks | 4 |
| 2022 | An effective deep network using target vector update modules for image restoration
Sen Zhai, Chao Ren 0002, Zhengyong Wang, Xiaohai He, Linbo Qing |
Pattern Recognit. | 5 |
| 2022 | A video compression artifact reduction approach combined with quantization parameters estimation
Xin Shuai, Linbo Qing, Mozhi Zhang, Weiheng Sun, Xiaohai He |
J. Supercomput. | 2 |
| 2022 | A Feature-Enriched Deep Convolutional Neural Network for JPEG Image Compression Artifacts Reduction and its ApplicationsabstractThe amount of multimedia data, such as images and videos, has been increasing rapidly with the development of various imaging devices and the Internet, bringing more stress and challenges to information storage and transmission. The redundancy in images can be reduced to decrease data size via lossy compression, such as the most widely used standard Joint Photographic Experts Group (JPEG). However, the decompressed images generally suffer from various artifacts (e.g., blocking, banding, ringing, and blurring) due to the loss of information, especially at high compression ratios. This article presents a feature-enriched deep convolutional neural network for compression artifacts reduction (FeCarNet, for short). Taking the dense network as the backbone, FeCarNet enriches features to gain valuable information via introducing multi-scale dilated convolutions, along with the efficient 1 ×1 convolution for lowering both parameter complexity and computation cost. Meanwhile, to make full use of different levels of features in FeCarNet, a fusion block that consists of attention-based channel recalibration and dimension reduction is developed for local and global feature fusion. Furthermore, short and long residual connections both in the feature and pixel domains are combined to build a multi-level residual structure, thereby benefiting the network training and performance. In addition, aiming at reducing computation complexity further, pixel-shuffle-based image downsampling and upsampling layers are, respectively, arranged at the head and tail of the FeCarNet, which also enlarges the receptive field of the whole network. Experimental results show the superiority of FeCarNet over state-of-the-art compression artifacts reduction approaches in terms of both restoration capacity and model complexity. The applications of FeCarNet on several computer vision tasks, including image deblurring, edge detection, image segmentation, and object detection, demonstrate the effectiveness of FeCarNet further. Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | HybNet: a hybrid network structure for pain intensity estimation
Yibo Huang 0003, Linbo Qing, Shengyu Xu, Yonghong Peng |
Vis. Comput. | 2 |
| 2022 | HF-SRGR: a new hybrid feature-driven social relation graph reasoning model
Lindong Li, Linbo Qing, Jie Su 0011, Yongqiang Cheng 0001, Yonghong Peng |
Vis. Comput. | 2 |
| 2021 | Deep Deblocker Driven Adaptive Iteration Scheme for Compressed Image RecoveryabstractIt is challenging to propose a flexible and effective framework for various JPEG compressed image recovery (CIR) tasks. In this paper, we propose a novel deep deblocker-driven adaptive iteration scheme, which can quickly and flexibly address various CIR tasks. First, a novel fidelity (NF) is introduced into CIR, and then the CIR problem is divided into inversion and deblocking subproblems by our improved split Bregman iteration (ISBI) algorithm. Next, we design a set of compact yet effective deep deblockers. These deblockers are used as implicit priors and also used for NF in the CIR problem. The convergence of our method is proved as well. To the best of our knowledge, our method is the first work to use deblockers as implicit priors. Extensive experiments demonstrate the effectiveness of our CIR method. Chao Ren 0002, Xiaohai He, Linbo Qing, Yuanzhouhan Cao |
ICME | 3 |
| 2021 | An enhanced siamese angular softmax network with dual joint-attention for person re-identification
Jie Su 0011, Xiaohai He, Linbo Qing, Yongqiang Cheng 0001, Yonghong Peng |
Appl. Intell. | 3 |
| 2021 | Multi-scale features based interpersonal relation recognition using higher-order graph neural network
Linbo Qing, Lindong Li, Yongqiang Cheng 0001, Yonghong Peng |
Neurocomputing | 2 |
| 2021 | Deep recursive network for image denoising with global non-linear smoothness constraint prior
Chuncheng Wang, Chao Ren 0002, Xiaohai He, Linbo Qing |
Neurocomputing | 4 |
| 2021 | Bi-directional skip connection feature pyramid network and sub-pixel convolution for high-quality object detection
Shuqi Xiong, Honggang Chen, Linbo Qing, Xiaohai He |
Neurocomputing | 4 |
| 2021 | Remote sensing image recovery via enhanced residual learning and dual-luminance scheme
Chao Ren 0002, Xiaohai He, Linbo Qing, Yuanyuan Wu 0001, Yi-Fei Pu |
Knowl. Based Syst. | 3 |
| 2021 | Compressed image restoration via deep deblocker driven unified framework
Chao Ren 0002, Qizhi Teng, Xiaohai He, Linbo Qing, Truong Q. Nguyen |
Knowl. Based Syst. | 4 |
| 2020 | EyesGAN: Synthesize human face from human eyes
Xiaodong Luo, Xiaohai He, Linbo Qing, Xiang Chen 0008, Luping Liu |
Neurocomputing | 3 |
| 2019 | Large-Scale Street Space Quality Evaluation Based on Deep Learning Over Street View Image
Longmei Han, Shanshan Xiong, Linbo Qing, Haohao Ji, Yonghong Peng |
ICIG (2) | 4 |
| 2019 | Facial Expression Recognition Based on Group Domain Random Frame Extraction
Yibo Huang 0003, Linbo Qing, Xiaohai He |
ICIG (1) | 4 |
| 2019 | Machine learning-based H.264/AVC to HEVC transcoding via motion information reuse and coding mode similarity analysisabstractHigh‐efficiency video coding (HEVC), which is the latest video coding standard, is expected to have a dominant position in the market in the near future. However, most video resources are now encoded using the H.264/AVC standard. Consequently, there is a growing need for fast H.264/AVC to HEVC transcoders to facilitate the migration to the updated standard. This paper proposes a fast H.264/AVC to HEVC transcoding scheme, which constructs a three‐level classifier using an optimised tree‐augmented Naive Bayesian approach to predict the HEVC coding unit depth. A feature selection method is then proposed to improve prediction accuracy. A motion vector (MV) calculation method is also proposed to reduce the complexity of MV prediction in HEVC by reusing MVs from H.264/AVC. Experimental results show that, compared with other state‐of‐the‐art transcoding algorithms, the proposed algorithm considerably reduces coding complexity while causing only negligible rate‐distortion degradation. Xiaohai He, Linbo Qing, Shan Su, Shuhua Xiong |
IET Image Process. | 3 |
| 2019 | High-Order Statistical Modeling Based on a Decision Tree for Distributed Video CodingabstractAiming at low-complexity encoding, distributed video coding (DVC) based on the Wyner-Ziv theorem has attracted significant attention. However, there is still a compression performance gap between the state-of-the-art DVC and the conventional video coding. One of the most important factors is the efficient estimation of the source correlation statistics. The first-order Laplacian distribution has been widely used for source correlation modeling, but is not effective enough; high-order statistical modeling is necessary, but needs more context features for the estimation of a source symbol's conditional probability. How to analyze the strength of the correlation between the source symbol and the context features in order to utilize the features effectively is crucial for such modeling. In this paper, the estimation of the source statistical distribution is first treated as a classification problem. The symbols of the source can be classified into different classes when the relevant context features are given. Then decision tree learning is introduced to analyze the strength of the correlation between the source symbol and the context features. Specifically, by constructing the decision trees composed of the selected context features, the selected context features can be organized effectively to derive the rules to estimate the current symbol's conditional probability, upon which the high-order statistical modeling is designed. Experimental results show that the proposed model can achieve significant coding gain over existing DVC systems, especially for natural videos with high motion intensity. Linbo Qing, Wenjun Zeng 0001, Xiaohai He |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Improved Low-Bitrate HEVC Video Coding Using Deep Learning Based Super-Resolution and Adaptive Block PatchingabstractGood-quality video coding for low-bitrate applications is essential for narrow bandwidth transmission and limited capacity storage. In this paper, we propose an adaptive downsampling-based coding model to improve the low-bitrate compression efficiency of high-efficiency video coding (HEVC). At the encoder, the video sequence is adaptively divided into key frames (KFs) and nonkey frames (NKFs), which are encoded at the original resolution and at a reduced resolution, respectively. At the decoder, a super-resolution method based on deep learning and gradient transformation is used to upscale the NKFs. To improve the quality of NKFs without additional information during decoding, we use motion estimation to find the most similar blocks between the upscaled NKFs and the associated high-resolution KFs. Then, an adaptive patching-based method is used to warp the low-quality NKF blocks with the high-quality KF blocks. Experimental results indicate that for standard high-definition test video sequences, the maximum improvement in the peak signal-to-noise ratio can reach 3.54 dB, and the critical bitrate can reach 9.89 Mb/s at a low bitrate when compared to HEVC. These results demonstrate significant improvements compared to existing methods. Xiaohai He, Linbo Qing, Qizhi Teng, Songfan Yang |
IEEE Trans. Multim. | 3 |
| 2018 | CISRDCNN: Super-resolution of compressed images using deep convolutional neural networks
Honggang Chen, Xiaohai He, Chao Ren 0002, Linbo Qing, Qizhi Teng |
Neurocomputing | 4 |
| 2018 | Robust distributed video coding for wireless multimedia sensor networks
Linbo Qing, Xiaohai He, Xianfeng Ou |
Multim. Tools Appl. | 2 |
| 2018 | Adaptive Gradient Information and BFGS Based Inter Frame Rate Control for High Efficiency Video Coding
Yuyun Ye, Xiaohai He, Qizhi Teng, Linbo Qing, Dechun Xia |
Multim. Tools Appl. | 4 |
| 2018 | SGCRSR: Sequential gradient constrained regression for single image super-resolution
Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng, Chao Ren 0002 |
Signal Process. Image Commun. | 3 |
| 2018 | An Iterative Framework of Cascaded Deblocking and Superresolution for Compressed ImagesabstractSuperresolution (SR) of compressed images is chall-enging due to the combination of resolution loss and compression artifacts. To solve these intertwined problems, the conventional cascading framework splits the solution into independent deblocking and SR subprocesses, where some existing high-frequency (HF) components are often oversmoothed during deblocking and information exchange between cascaded deblocking and SR remains untouched. In this paper, we propose an iterative cascading framework after analyzing the correlation between the two subprocesses. Deblocking is provided with a shape-adaptive low-rank prior to well preserve edges and an extra prior to restore the lost HF components. The latter prior represents an important feedback link from SR to deblocking, which is a novel design in this framework. To provide an accurate and noise-robust feedback of the extra prior, an SR method via singular value decomposition projection is also developed. The extensive experimental results demonstrate the superior performance of the proposed method. Tao Li 0014, Xiaohai He, Linbo Qing, Qizhi Teng, Honggang Chen |
IEEE Trans. Multim. | 3 |
| 2017 | Tree-structured Bayesian compressive sensing via generalised inverse Gaussian distributionabstractCompressive sensing (CS) implements signal sampling and compression simultaneously, which significantly alleviates the pressure on the sampling end. However, the reconstruction algorithm is an underdetermined linear inverse problem. To solve this problem, it is crucial to involve prior knowledge regarding the reconstructed signal. In this study, the compressibility of wavelet coefficients is utilised as prior knowledge. Moreover, a generalised inverse Gaussian (GIG) distribution is integrated in the context of tree‐structured Bayesian CS (TSBCS), which also imposes the persistence property between the successive levels. Finally, variational Bayesian inference is used to infer the posterior probability distribution of the model parameters. Due to the overall algorithm is based on TSBCS, the proposal is referred to as TSBCS via a GIG distribution (TSBCS‐GIG). Experimental results show that the authors’ proposed TSBCS‐GIG algorithm outperforms other well‐known algorithms in both peak signal‐to‐noise ratio and visual quality. Maojiao Wang, Xiaohai He, Linbo Qing, Shuhua Xiong |
IET Signal Process. | 3 |
| 2017 | Single Image Super-Resolution via Adaptive Transform-Based Nonlocal Self-Similarity Modeling and Learning-Based Gradient RegularizationabstractSingle image super-resolution (SISR) is a challenging work, which aims to recover the missing information in an observed low-resolution (LR) image and generate the corresponding high-resolution (HR) version. As the SISR problem is severely ill-conditioned, effective prior knowledge of HR images is necessary to well pose the HR estimation. In this paper, an effective SISR method is proposed via the local structure-adaptive transform-based nonlocal self-similarity modeling and learning-based gradient regularization (LSNSGR). The LSNSGR exploits both the natural and learned priors of HR images, thus integrating the merits of conventional reconstruction-based and learning-based SISR algorithms. More specifically, on the one hand, we characterize nonlocal self-similarity prior (natural prior) in transform domain by using the designed local structure-adaptive transform; on the other hand, the gradient prior (learned prior) is learned via the jointly optimized regression model. The former prior is effective in suppressing visual artifacts, while the latter performs well in recovering sharp edges and fine structures. By incorporating the two complementary priors into the maximum a posteriori-based reconstruction framework, we optimize a hybrid L1- and L2-regularized minimization problem to achieve an estimation of the desired HR image. Extensive experimental results suggest that the proposed LSNSGR produces better HR estimations than many state-of-the-art works in terms of both perceptual and quantitative evaluations. Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng |
IEEE Trans. Multim. | 3 |
| 2016 | Depth-based distributed multi-view video coding with hierarchical Wyner-Ziv framesabstractDistributed multi-view video coding (DMVC) is a new emerging multi-view video coding (MVC) scheme, in which multi-view video are encoded separately and decoded dependently, so the burden of huge computation is shifted from the encoder to the decoder side. However, there is still a large gap between the DMVC and traditional MVC in terms of compression performance. In order to improve the coding performance of DMVC, wavelet domain DMVC framework based on hierarchical Wyner-Ziv frames is proposed. With the introduction of depth map to multi-view, a fusion algorithm based on error correction with the information from depth map and adjacent views is proposed. The experimental results show better quality of the SI for the WZ frames and significant improvement in rate distortion (RD) performance are achieved. Linbo Qing, Xiaohai He, Wenshi Xiong |
VCIP | 1 |
| 2015 | A fast inter-prediction algorithm for HEVC based on temporal and spatial correlation
Guo-Yun Zhong, Xiaohai He, Linbo Qing |
Multim. Tools Appl. | 3 |
| 2014 | Improving distributed video coding by exploiting context-adaptive modelingabstractThe statistical model of the bits to be encoded is crucial for the coding performance of distributed video coding (DVC). In this paper, a bit-level context-adaptive correlation model is proposed to exploit high-order statistical correlation for better channel coding performance, which consequently improves the video coding efficiency. In the proposed scheme, the wavelet domain DVC is considered and the coefficients are coded in a bit-plane fashion. The context for each bit to be coded is first formed. Then the probability distribution of each bit is estimated by using previously available data with the same context. For magnitude coding, the significant state of the following elements are included in the context, (1) the side information, (2) the local neighborhood, (3) the parent coefficients (if applicable). The condition of side information is considered as well. For sign coding, the context consists of the sign and the quality of the side information. The proposed model is implemented within a recently proposed DVC framework with decoderside multi-resolution motion refinement (MRMR). Experimental results show the effectiveness of the proposed scheme with significant coding gain over the original MRMR based DVC system, especially for videos with high motion intensity and for lower bit rates. Linbo Qing, Wenjun Zeng 0001 |
ICME | 1 |
| 2013 | Adaptive regularization-based space-time super-resolution reconstruction
Haiying Song, Linbo Qing, Yuanyuan Wu 0001, Xiaohai He |
Signal Process. Image Commun. | 2 |