VLDB 2026 Research / reviewers in the wild / expert
Guoli Song
dblp:143/0447
· DBLP profile ↗
28ranked-venue papers
7as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixture-of-experts network via frequency-causal reasoning for spinal CT segmentation
Guoli Song, Yuhan Ying, Xingang Zhao |
Expert Syst. Appl. | 1 |
| 2025 | A Variable Stiffness Supernumerary Robotic Limb with Pneumatic-Tendon Coupled Actuation *abstractSupernumerary robotic limbs (SRLs) can assist humans in achieving efficient and comfortable work in daily life or industrial assembly scenarios, requiring SRLs to switch between rigidity and flexibility to perform compliant movements while also providing stable support for humans to reduce fatigue from prolonged standing, existing SRLs struggle to achieve this transition. In this study, a variable stiffness supernumerary robotic limb (VSSRL) is implemented, capable of adjusting its position and stiffness through pneumatic-tendon coupled actuation. The position of the VSSRL is accurately modulated by tendons, while its stiffness is controlled by pneumatic-tendon coupled actuation, tendons significantly increase the overall stiffness of the VSSRL, and the fiber-reinforced actuators (FRAs) can dynamically adjust its stiffness in response to changes in dynamic loads. Furthermore, a kinematic model of the VSSRL and a stiffness model under the coupling of FRAs and tendons are developed. Then, the trajectory and stiffness of the VSSRL in task execution are assigned based on human motion, and a multi-objective control system for both position and stiffness of the VSSRL is designed based on reinforcement learning (RL) algorithm, achieving collaborative control of position and stiffness for the VSSRL. The accuracy of the control system is validated through experiments, which demonstrate that the load capacity of the VSSRL is significantly enhanced under the action of tendons and FRAs, and that the VSSRL is able to provide various modes of assistance for daily life activities. Mengcheng Zhao, Juanxia Zhou, Kaizhen Huang, Aihong Ji, Xuyan Hou, Guoli Song, Youfu Li 0001 |
IROS | 9 |
| 2025 | Probability-Based Force Control for Flexible Ureteroscopy RobotsabstractIn robotic flexible ureteroscopy, great care should be taken to prevent excessive collision force at the ureteroscope tip to avoid urinary tissue damage. However, the passive flexure of a flexible ureteroscope from reaction forces with the surrounding tissue introduces great uncertainties in force control. To address this, a probability-based force control method is proposed in this paper. A morphological wavelet-based statistic test is proposed to detect collision by identifying the change point of the probability distribution of force signal, which is measured by a fiber optical sensor at the ureteroscope tip. The force signal and its change points are applied as an admittance model inputs to generate position command. It is augmented to the robot controller for minimizing the collision force. Meanwhile, a probabilistic model approximates the real interactive system, and combines Bayesian inference for safely online learning the admittance parameters. Experimental results on a flexible ureteroscopy robot show that collision can be detected instantly, and the force is reduced remarkably during the advancement of the robotic ureteroscope. Note to Practitioners—This paper is motivated by the rapidly increasing applications for robotic flexible ureteroscopy. As ureteroscope enters into urethra, urinary bladder, ureter, and reaches renal pelvis, it sometimes penetrates soft tissue, especially in the case of no ureteral access sheath placement. It is important to control the ureteroscope tip force for surgery safety. Therefore, a force control method is proposed in this paper. Experiments conducted on a flexible ureteroscopy robot demonstrate that the proposed method is a protective measure to mitigate the risk of inner surface perforation. Yinan Deng, Tangwen Yang, Yuelin Zou, Jianchang Zhao, Bin Yao 0003, Guoli Song |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Development and Evaluation of an Automated Computational Approach for the Precise Placement of Pedicle Screws in Spinal Surgery Leveraging Three-Dimensional Point Cloud Registration MethodsabstractThis study focuses on enhancing the precision and efficacy of pedicle screw placement in spinal surgeries, particularly for patients with osteoporosis. It emphasizes the importance of accurate screw positioning to maximize pullout strength and biomechanical efficacy. The research highlights the relationship between CT values, Bone Mineral Density (BMD), and the Young’s modulus of bone tissue, suggesting that higher CT values, indicative of denser bone, lead to stronger mechanical properties. This understanding is crucial for assessing bone health, especially in osteoporosis and fracture risk analysis. The study introduces an innovative automated planning method for pedicle screw insertion in spinal vertebrae, beneficial for osteoporotic patients. This method uses PointNet++ combined with a Siamese network for semi-supervised segmentation of vertebrae in CT images, converting them into point clouds for individualized planning. The approach aims to improve accuracy and efficiency in pedicle screw placement, especially in complex or osteoporotic vertebrae. The method was validated using osteoporotic in vitro models, demonstrating its potential effectiveness for surgeons facing challenges in pedicle screw placement in osteoporotic patients. The complete workflow begins with semi-supervised vertebra segmentation using a Siamese PointNet++ architecture. The extracted vertebra point cloud is then aligned with a predefined pedicle screw model via an adaptively tuned Super4PCS registration algorithm to generate patient-specific trajectories. The paper also reviews traditional surgical path planning, which relied on surgeons’ experience and intuition, and the shift towards Computer-Aided Design (CAD) and Virtual Reality (VR) technologies for pre-operative planning. Despite these advancements, challenges remain in accurately simulating tissue properties and managing physiological variations. The study proposes a bifurcated approach to automated trajectory planning: segmenting individual vertebrae and planning pedicle screw implantation for each segmented vertebra. The method involves converting CT scans into 3D point clouds, using PointNet++ and the Siamese method for vertebra segmentation, and registering the pedicle screw point cloud with the vertebra model for precise surgical planning. The study demonstrates the method’s feasibility through clinical data from Shengjing Hospital, showing high applicability and consistency in surgical path planning, particularly in the lumbar and lower thoracic regions. The planning results were consistent with surgeons’ experience, indicating the algorithm’s adaptability and stability. Guoli Song, Andi Li, Yuhan Ying, Xingang Zhao |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | Human-Robot Interaction Control for Multi-Mode Exosuit with Reinforcement LearningabstractSoft exoskeleton robots have promising potential in walking assistance with comfortable wearing experience. In this study, an exosuit equipped with a twisted string actuator (TSA) is developed to provide powerful driving force and diverse operating modes for hemiplegic patients in daily life. It is challenging to establish the human-robot coupling dynamic model due to the soft structure of the exosuit and tight coupling, precise control and effective assistance are difficult to guaranteed in current exosuits. Considering the impedance characteristics of human-robot interaction, an adaptive impedance control method based on reinforcement learning (RL) is proposed, where human motion intention is utilized to optimize impedance parameters and adjust the robot's operating mode. A nonlinear disturbance observer is proposed to compensate for the effects of model estimation errors, joint friction, and external disturbances. Experimental verification demonstrates the effectiveness and superiority of the robotic system. Kaizhen Huang, Mengcheng Zhao, Aihong Ji, Guoli Song, Youfu Li 0001 |
IROS | 6 |
| 2023 | ACSeg: Adaptive Conceptualization for Unsupervised Semantic SegmentationabstractRecently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship among pixel-level representations typically contains rich class-aware information that semantically identical pixel embeddings in the representation space gather together to form sophisticated concepts. However, leveraging the learned models to ascertain semantically consistent pixel groups or regions in the image is non-trivial since over/ under-clustering overwhelms the conceptualization procedure under various semantic distributions of different images. In this work, we investigate the pixel-level semantic aggregation in self-supervised ViT pre-trained models as image Segmentation and propose the Adaptive Conceptualization approach for USS, termed ACSeg. Concretely, we explicitly encode concepts into learnable prototypes and design the Adaptive Concept Generator (ACG), which adaptively maps these prototypes to informative concepts for each image. Meanwhile, considering the scene complexity of different images, we propose the modularity loss to optimize ACG independent of the concept number based on estimating the intensity of pixel pairs belonging to the same concept. Finally, we turn the USS task into classifying the discovered concepts in an unsupervised manner. Extensive experiments with state-of-the-art results demonstrate the effectiveness of the proposed ACSeg. Kehan Li 0002, Zhennan Wang 0001, Zesen Cheng, Runyi Yu 0002, Yian Zhao, Guoli Song, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
CVPR | 6 |
| 2023 | Fuzzy Positive Learning for Semi-Supervised Semantic SegmentationabstractSemi-supervised learning (SSL) essentially pursues class boundary exploration with less dependence on human annotations. Although typical attempts focus on ameliorating the inevitable error-prone pseudo-labeling, we think differently and resort to exhausting informative semantics from multiple probably correct candidate labels. In this paper, we introduce Fuzzy Positive Learning (FPL) for accurate SSL semantic segmentation in a plug-and-play fashion, targeting adaptively encouraging fuzzy positive predictions and suppressing highly-probable negatives. Being conceptually simple yet practically effective, FPL can remarkably alleviate interference from wrong pseudo labels and progressively achieve clear pixel-level semantic discrimination. Concretely, our FPL approach consists of two main components, including fuzzy positive assignment (FPA) to provide an adaptive number of labels for each pixel and fuzzy positive regularization (FPR) to restrict the predictions of fuzzy positive categories to be larger than the rest under different perturbations. Theoretical analysis and extensive experiments on Cityscapes and VOC 2012 with consistent performance gain justify the superiority of our approach. Codes are provided in https://github.com/qpc1611094/FPL. Pengchong Qiao, Zhidan Wei, Yu Wang 0027, Zhennan Wang 0001, Guoli Song, Xiangyang Ji, Chang Liu 0030, Jie Chen 0001 |
CVPR | 5 |
| 2023 | Out-of-Distributed Semantic Pruning for Robust Semi-Supervised LearningabstractRecent advances in robust semi-supervised learning (SSL) typically filter out-of-distribution (OOD) information at the sample level. We argue that an overlooked problem of robust SSL is its corrupted information on semantic level, practically limiting the development of the field. In this paper, we take an initial step to explore and propose a unified framework termed OOD Semantic Pruning (OSP), which aims at pruning OOD semantics out from in-distribution (ID) features. Specifically, (i) we propose an aliasing OOD matching module to pair each ID sample with an OOD sample with semantic overlap. (ii) We design a soft orthogonality regularization, which first transforms each ID feature by suppressing its semantic component that is collinear with paired OOD sample. It then forces the predictions before and after soft orthogonality decomposition to be consistent. Being practically simple, our method shows a strong performance in OOD detection and ID classification on challenging benchmarks. In particular, OSP surpasses the previous state-of-the-art by 13.7% on accuracy for ID classification and 5.9% on AUROC for OOD detection on TinyImageNet dataset. The source codes are publicly available at https://github.com/rain305f/OSP. Yu Wang 0027, Pengchong Qiao, Chang Liu 0030, Guoli Song, Xiawu Zheng, Jie Chen 0001 |
CVPR | 4 |
| 2023 | Recurrent Fine-Grained Self-Attention Network for Video Crowd CountingabstractStriking a balance between exploring the spatio-temporal correlation and controlling model complexity is vital for video-based crowd counting methods. In this paper, we propose a Recurrent Fine-Grained Self-Attention Network (RFSNet) to achieve efficient and accurate counting in video scenes via the self-attention mechanism and a recurrent fine-tuning strategy. Specifically, we design a decoder which consists of patch-wise spatial self-attention and temporal self-attention. Compared with vanilla self-attention, it effectively leverages the dependencies in spatial and temporal domain respectively, while significantly reducing computational complexity. Moreover, the RFSNet recurrently feeds the features into the decoder to enhance the spatio-temporal representations. This strategy not only simplifies the model structure and reduces the number of parameters, but also improves the quality of estimated density maps. Our RFSNet achieves state-of-the-art performance on three video crowd counting benchmarks, and outperforms other methods by more than 20% on the challenging FDST dataset. Jifan Zhang, Zhe Wu 0006, Xinfeng Zhang 0001, Guoli Song, Yaowei Wang 0001, Jie Chen 0001 |
ICASSP | 4 |
| 2023 | Weakly-Supervised 3D Spatial Reasoning for Text-Based Visual Question AnsweringabstractText-based Visual Question Answering (TextVQA) aims to produce correct answers for given questions about the images with multiple scene texts. In most cases, the texts naturally attach to the surface of the objects. Therefore, spatial reasoning between texts and objects is crucial in TextVQA. However, existing approaches are constrained within 2D spatial information learned from the input images and rely on transformer-based architectures to reason implicitly during the fusion process. Under this setting, these 2D spatial reasoning approaches cannot distinguish the fine-grained spatial relations between visual objects and scene texts on the same image plane, thereby impairing the interpretability and performance of TextVQA models. In this paper, we introduce 3D geometric information into the spatial reasoning process to capture the contextual knowledge of key objects step-by-step. Specifically, (i) we propose a relation prediction module for accurately locating the region of interest of critical objects; (ii) we design a depth-aware attention calibration module for calibrating the OCR tokens' attention according to critical objects. Extensive experiments show that our method achieves state-of-the-art performance on TextVQA and ST-VQA datasets. More encouragingly, our model surpasses others by clear margins of 5.7% and 12.1% on questions that involve spatial reasoning in TextVQA and ST-VQA valid split. Besides, we also verify the generalizability of our model on the text-based image captioning task. Hao Li 0073, Jinfa Huang, Peng Jin 0001, Guoli Song, Qi Wu 0001, Jie Chen 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Robust and Hierarchical Spatial Relation Analysis for Traffic ForecastingabstractHow to model the complex spatial-temporal relation in traffic data is an important problem for precisely predicting the future status of a city traffic system. Existing traffic forecasting methods rarely consider the traffic state trend, and the robust spatial relation has not been well explored. To tackle these issues, we design a novel Robust And Hierarchical spatial Relation Analysis (RAHRA) method to calculate the local-period spatial relation, which applies temporal context information in both traffic state and trend similarities. This could capture abundant traffic patterns and learn stable and comprehensive spatial relations for accurate traffic forecasting. Furthermore, we introduce a Temporal Attention Module (TAM) to capture the temporal features and propose a Future Feature Inference Module (FFIM) to infer the future traffic information. Experiments on four real-world traffic datasets demonstrate that the proposed method outperforms the other state-of-the-art methods. Zhe Wu 0006, Xinfeng Zhang 0001, Guoli Song, Yaowei Wang 0001, Jie Chen 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Semi-Supervised CT Lesion Segmentation Using Uncertainty-Based Data Pairing and SwapMixabstractSemi-supervised learning (SSL) methods show their powerful performance to deal with the issue of data shortage in the field of medical image segmentation. However, existing SSL methods still suffer from the problem of unreliable predictions on unannotated data due to the lack of manual annotations for them. In this paper, we propose an unreliability-diluted consistency training (UDiCT) mechanism to dilute the unreliability in SSL by assembling reliable annotated data into unreliable unannotated data. Specifically, we first propose an uncertainty-based data pairing module to pair annotated data with unannotated data based on a complementary uncertainty pairing rule, which avoids two hard samples being paired off. Secondly, we develop SwapMix, a mixed sample data augmentation method, to integrate annotated data into unannotated data for training our model in a low-unreliability manner. Finally, UDiCT is trained by minimizing a supervised loss and an unreliability-diluted consistency loss, which makes our model robust to diverse backgrounds. Extensive experiments on three chest CT datasets show the effectiveness of our method for semi-supervised CT lesion segmentation. Pengchong Qiao, Guoli Song, Hu Han 0001, Yonghong Tian 0001, Yongsheng Liang 0001, Xi Li 0011, Shaohua Kevin Zhou, Jie Chen 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2022 | ViSTA: Vision and Scene Text Aggregation for Cross-Modal RetrievalabstractVisual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single Vision and Scene Text Aggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least 8.4% at Recall@ 1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework. Mengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu, Jie Chen 0001, Guoli Song, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
CVPR | 7 |
| 2022 | Locality Guidance for Improving Vision Transformers on Tiny Datasets
Kehan Li 0002, Runyi Yu 0002, Zhennan Wang 0001, Li Yuan 0007, Guoli Song, Jie Chen 0001 |
ECCV (24) | 5 |
| 2022 | Expectation-Maximization Contrastive Learning for Compact Video-and-Language RepresentationsabstractMost video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to the semantic similarities of text-video pairs. However, such learned shared latent spaces are not often optimal, and the modality gap between visual and textual representation can not be fully eliminated. In this paper, we propose Expectation-Maximization Contrastive Learning (EMCL) to learn compact video-and-language representations. Specifically, we use the Expectation-Maximization algorithm to find a compact set of bases for the latent space, where the features could be concisely represented as the linear combinations of these bases. Such feature decomposition of video-and-language representations reduces the rank of the latent space, resulting in increased representing power for the semantics. Extensive experiments on three benchmark text-video retrieval datasets prove that our EMCL can learn more discriminative video-and-language representations than previous methods, and significantly outperform previous state-of-the-art methods across all metrics. More encouragingly, the proposed method can be applied to boost the performance of existing approaches either as a jointly training layer or an out-of-the-box inference module with no extra training, making it easy to be incorporated into any existing methods. Peng Jin 0001, Jinfa Huang, Xian Wu 0001, Shen Ge, Guoli Song, David A. Clifton, Jie Chen 0001 |
NeurIPS | 6 |
| 2022 | Interpretation of convolutional neural networks reveals crucial sequence features involving in transcription during fiber developmentabstractBACKGROUND: Upland cotton provides the most natural fiber in the world. During fiber development, the quality and yield of fiber were influenced by gene transcription. Revealing sequence features related to transcription has a profound impact on cotton molecular breeding. We applied convolutional neural networks to predict gene expression status based on the sequences of gene transcription start regions. After that, a gradient-based interpretation and an N-adjusted kernel transformation were implemented to extract sequence features contributing to transcription. RESULTS: Our models had approximate 80% accuracies, and the area under the receiver operating characteristic curve reached over 0.85. Gradient-based interpretation revealed 5' untranslated region contributed to gene transcription. Furthermore, 6 DOF binding motifs and 4 transcription activator binding motifs were obtained by N-adjusted kernel-motif transformation from models in three developmental stages. Apart from 10 general motifs, 3 DOF5.1 genes were also detected. In silico analysis about these motifs' binding proteins implied their potential functions in fiber formation. Besides, we also found some novel motifs in plants as important sequence features for transcription. CONCLUSIONS: In conclusion, the N-adjusted kernel transformation method could interpret convolutional neural networks and reveal important sequence features related to transcription during fiber development. Potential functions of motifs interpreted from convolutional neural networks could be validated by further wet-lab experiments and applied in cotton molecular breeding. Hailiang Cheng, Javaria Ashraf, Youping Zhang, Qiaolian Wang, Limin Lv, Man He, Guoli Song, Dongyun Zuo |
BMC Bioinform. | 8 |
| 2021 | CDNet: Centripetal Direction Network for Nuclear Instance SegmentationabstractNuclear instance segmentation is a challenging task due to a large number of touching and overlapping nuclei in pathological images. Existing methods cannot effectively recognize the accurate boundary owing to neglecting the relationship between pixels (e.g., direction information). In this paper, we propose a novel Centripetal Direction Net-work (CDNet) for nuclear instance segmentation. Specifically, we define centripetal direction feature as a class of adjacent directions pointing to the nuclear center to rep-resent the spatial relationship between pixels within the nucleus. These direction features are then used to construct a direction difference map to represent the similarity within instances and the differences between instances. Finally, we propose a direction-guided refinement module, which acts as a plug-and-play module to effectively integrate auxiliary tasks and aggregate the features of different branches. Experiments on MoNuSeg and CPM17 datasets show that CDNet is significantly better than the other methods and achieves the state-of-the-art performance. The code is available at https://github.com/honglianghe/CDNet. Yao Ding 0006, Guoli Song, Lin Wang 0026, Qian Ren, Pengxu Wei, Jie Chen 0001 |
ICCV | 4 |
| 2021 | Learnable Oriented-Derivative Network for Polyp Segmentation
Mengjun Cheng, Zishang Kong, Guoli Song, Yonghong Tian 0001, Yongsheng Liang 0001, Jie Chen 0001 |
MICCAI (1) | 3 |
| 2021 | Hard-Boundary Attention Network for Nuclei Instance SegmentationabstractImage segmentation plays an important role in medical image analysis, and accurate segmentation of nuclei is especially crucial to clinical diagnosis. However, existing methods fail to segment dense nuclei due to the hard-boundary which has similar texture to nuclear inside. To this end, we propose a Hard-Boundary Attention Network (HBANet) for nuclei instance segmentation. Specifically, we propose a Background Weaken Module (BWM) to weaken the attention of our model to the nucleus background by integrating low-level features into high-level features. To improve the robustness of the model to the hard-boundary of nuclei, we further design a Gradient-based boundary adaptive Strategy (GS) which generates boundary-weakened data for model training in an adversarial manner. We conduct extensive experiments on MoNuSeg and CPM-17 datasets, and experimental results show that our HBANet outperforms the state-of-the-art methods. Yalu Cheng, Pengchong Qiao, Guoli Song, Jie Chen 0001 |
MMAsia | 4 |
| 2021 | Harmonized Multimodal Learning with Gaussian Process Latent Variable ModelsabstractMultimodal learning aims to discover the relationship between multiple modalities. It has become an important research topic due to extensive multimodal applications such as cross-modal retrieval. This paper attempts to address the modality heterogeneity problem based on Gaussian process latent variable models (GPLVMs) to represent multimodal data in a common space. Previous multimodal GPLVM extensions generally adopt individual learning schemes on latent representations and kernel hyperparameters, which ignore their intrinsic relationship. To exploit strong complementarity among different modalities and GPLVM components, we develop a novel learning scheme called Harmonization, where latent representations and kernel hyperparameters are jointly learned from each other. Beyond the correlation fitting or intra-modal structure preservation paradigms widely used in existing studies, the harmonization is derived in a model-driven manner to encourage the agreement between modality-specific GP kernels and the similarity of latent representations. We present a range of multimodal learning models by incorporating the harmonization mechanism into several representative GPLVM-based approaches. Experimental results on four benchmark datasets show that the proposed models outperform the strong baselines for cross-modal retrieval tasks, and that the harmonized multimodal learning method is superior in discovering semantically consistent latent representation. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Learning Feature Representation and Partial Correlation for Multimodal Multi-Label DataabstractUser-provided annotations in existing multimodal datasets sometimes are inappropriate for model learning and can hinder the task of cross-modal retrieval. To handle this issue, we propose a discriminative and noise-robust cross-modal retrieval method, called FLPCL, which consists of deep feature learning and partial correlation learning. Deep feature learning is implemented by utilizing label supervised information to guide the training of deep neural network for each modality, which aims to find modality-specific deep feature representations that preserve the similarity and discrimination information among multimodal data. Based on deep feature learning, partial correlation learning is proposed to infer direct association between different modalities by removing the effect of common underlying semantics from each modality. It is achieved by maximizing the canonical correlation of the feature representations of different modalities conditioned on the label modality. Different from existing works that build indirect association between modalities via incorporating semantic labels, our FLPCL method can learn more effective and robust multimodal latent representations by explicitly preserving both intra-modal and inter-modal relationship among multimodal data. Extensive experiments on three cross-modal datasets show that our method outperforms state-of-the-art methods on cross-modal retrieval tasks. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Augmented Adversarial Training for Cross-Modal RetrievalabstractCross-modal retrieval has received considerable attention in recent years. The core of cross-modal retrieval is to find a representation space to align data from different modalities according to their semantics. In this paper, we propose a cross-modal retrieval method that aligns data from different modalities by transferring one source modality to another target modality with augmented adversarial training. To preserve the semantic meaning in the modality transfer process, we employ the idea of conditional GANs and augment it. The key idea is to incorporate semantic information from the label space into the adversarial training process by sampling more semantic relevant and irrelevant source-target sample pairs. The augmented sample pairs improve the alignment from two aspects. First, relevant source-target sample pairs provide more training samples, leading to a better guidance of the alignment of fake targets and true paired targets. Second, relevant and irrelevant source-target sample pairs teach the discriminator to better distinguish true relevant pairs from fake relevant pairs, which guides the generator to better transfer from the source modality to the target modality. Extensive experiments compared with state-of-the-art methods show the promising power of our approach. Yiling Wu, Shuhui Wang, Guoli Song, Qingming Huang |
IEEE Trans. Multim. | 3 |
| 2020 | BCData: A Large-Scale Dataset and Benchmark for Cell Detection and Counting
Yao Ding 0006, Guoli Song, Lin Wang 0026, Ruizhe Geng, Yonghong Tian 0001, Yongsheng Liang 0001, Shaohua Kevin Zhou, Jie Chen 0001 |
MICCAI (5) | 3 |
| 2019 | Learning Fragment Self-Attention Embeddings for Image-Text MatchingabstractIn image-text matching task, the key to good matching quality is to capture the rich contextual dependencies between fragments of image and text. However, previous works either simply aggregate the similarity of all possible pairs of image regions and words, or take multi-step cross attention to attend to image regions and words with each other as context, which requires exhaustive similarity computation between all image region and word pairs. In this paper, we propose Self-Attention Embeddings (SAEM) to exploit fragment relations in images or texts by self-attention mechanism, and aggregate fragment information into visual and textual embeddings. Specifically, SAEM extracts salient image regions based on bottom-up attention, and takes WordPiece tokens as sentence fragments. The self-attention layers are built to model subtle and fine-grained fragment relation in image and text respectively, which consists of multi-head self-attention sub-layer and position-wise feed-forward network sub-layer. Consequently, the fragment self-attention mechanism can discover the fragment relations and identify the semantically salient regions in images or words in sentences, and capture their interaction more accurately. By simultaneously exploiting the fine-grained fragment relation in both visual and textual modalities, our method produces more semantically consistent embeddings for representing images and texts, and demonstrates promising image-text matching accuracy and high efficiency on Flickr30K and MSCOCO datasets. Yiling Wu, Shuhui Wang, Guoli Song, Qingming Huang |
ACM Multimedia | 3 |
| 2019 | Online Asymmetric Metric Learning With Multi-Layer Similarity Aggregation for Cross-Modal RetrievalabstractCross-modal retrieval has attracted intensive attention in recent years, where a substantial yet challenging problem is how to measure the similarity between heterogeneous data modalities. Despite using modality-specific representation learning techniques, most existing shallow or deep models treat different modalities equally and neglect the intrinsic modality heterogeneity and information imbalance among images and texts. In this paper, we propose an online similarity function learning framework to learn the metric that can well reflect the cross-modal semantic relation. Considering that multiple CNN feature layers naturally represent visual information from low-level visual patterns to high-level semantic abstraction, we propose a new asymmetric image-text similarity formulation which aggregates the layer-wise visual-textual similarities parameterized by different bilinear parameter matrices. To effectively learn the aggregated similarity function, we develop three different similarity combination strategies, i.e., average kernel, multiple kernel learning, and layer gating. The former two kernel-based strategies assign uniform weights on different layers to all data pairs; the latter works on the original feature representation and assigns instance-aware weights on different layers to different data pairs, and they are all learned by preserving the bi-directional relative similarity expressed by a large number of cross-modal training triplets. The experiments conducted on three public datasets well demonstrate the effectiveness of our methods. Yiling Wu, Shuhui Wang, Guoli Song, Qingming Huang |
IEEE Trans. Image Process. | 3 |
| 2017 | Multimodal Gaussian Process Latent Variable Models with HarmonizationabstractIn this work, we address multimodal learning problem with Gaussian process latent variable models (GPLVMs) and their application to cross-modal retrieval. Existing GPLVM based studies generally impose individual priors over the model parameters and ignore the intrinsic relations among these parameters. Considering the strong complementarity between modalities, we propose a novel joint prior over the parameters for multimodal GPLVMs to propagate multimodal information in both kernel hyperparameter spaces and latent space. The joint prior is formulated as a harmonization constraint on the model parameters, which enforces the agreement among the modality-specific GP kernels and the similarity in the latent space. We incorporate the harmonization mechanism into the learning process of multimodal GPLVMs. The proposed methods are evaluated on three widely used multimodal datasets for cross-modal retrieval. Experimental results show that the harmonization mechanism is beneficial to the GPLVM algorithms for learning non-linear correlation among heterogeneous modalities. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
ICCV | 1 |
| 2017 | Multimodal Similarity Gaussian Process Latent Variable ModelabstractData from real applications involve multiple modalities representing content with the same semantics from complementary aspects. However, relations among heterogeneous modalities are simply treated as observation-to-fit by existing work, and the parameterized modality specific mapping functions lack flexibility in directly adapting to the content divergence and semantic complicacy in multimodal data. In this paper, we build our work based on the Gaussian process latent variable model (GPLVM) to learn the non-parametric mapping functions and transform heterogeneous modalities into a shared latent space. We propose multimodal Similarity Gaussian Process latent variable model (m-SimGP), which learns the mapping functions between the intra-modal similarities and latent representation. We further propose multimodal distance-preserved similarity GPLVM (m-DSimGP) to preserve the intra-modal global similarity structure, and multimodal regularized similarity GPLVM (m-RSimGP) by encouraging similar/dissimilar points to be similar/dissimilar in the latent space. We propose m-DRSimGP, which combines the distance preservation in m-DSimGP and semantic preservation in m-RSimGP to learn the latent representation. The overall objective functions of the four models are solved by simple and scalable gradient decent techniques. They can be applied to various tasks to discover the nonlinear correlations and to obtain the comparable low-dimensional representation for heterogeneous modalities. On five widely used real-world data sets, our approaches outperform existing models on cross-modal content retrieval and multimodal classification. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Similarity Gaussian Process Latent Variable Model for Multi-modal Data AnalysisabstractData from real applications involve multiple modalities representing content with the same semantics and deliver rich information from complementary aspects. However, relations among heterogeneous modalities are simply treated as observation-to-fit by existing work, and the parameterized cross-modal mapping functions lack flexibility in directly adapting to the content divergence and semantic complicacy of multi-modal data. In this paper, we build our work based on Gaussian process latent variable model (GPLVM) to learn the non-linear non-parametric mapping functions and transform heterogeneous data into a shared latent space. We propose multi-modal Similarity Gaussian Process latent variable model (m-SimGP), which learns the nonlinear mapping functions between the intra-modal similarities and latent representation. We further propose multi-modal regularized similarity GPLVM (m-RSimGP) by encouraging similar/dissimilar points to be similar/dissimilar in the output space. The overall objective functions are solved by simple and scalable gradient decent techniques. The proposed models are robust to content divergence and high-dimensionality in multi-modal representation. They can be applied to various tasks to discover the non-linear correlations and obtain the comparable low-dimensional representation for heterogeneous modalities. On two widely used real-world datasets, we outperform previous approaches for cross-modal content retrieval and cross-modal classification. Guoli Song, Shuhui Wang, Qingming Huang, Qi Tian 0001 |
ICCV | 1 |