EDBT 2026 Demo / reviewers in the wild / expert
Kan Guo
dblp:139/0440
· DBLP profile ↗
24ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PlotGraph: Graph-First Screenplay Generation with Structural Consistency
Kan Guo, Haijun Lu, Jiaqian Ren, Daquan Feng |
ICPR (6) | 2 |
| 2026 | Multi-faceted contrastive learning with inter-frame difference for traffic video question answering
Kan Guo, Qi Zuo, Yongli Hu, Lanping Qian, Daxin Tian, Jiapu Wang, Guixian Qu, Tingzheng Jia, Junbin Gao |
Knowl. Based Syst. | 1 |
| 2026 | EmoFusion: Anchor-Guided Multimodal Emotion Recognition With Emotion-Specific Token Learning and Cross-Modal ComplementabstractExisting multimodal emotion recognition (MER) methods suffer from two critical flaws: (1) Single-modality encoders (e.g., wav2vec 2.0, BERT) lack emotion-specific adaptation. (2) Cross-modal fusion over-prioritizes alignment over emotion-discriminative enhancement. In this paper, we propose EmoFusion to address both limitations through injecting emotion-aware inductive biases into both unimodal feature extraction and multimodal fusion processes. First, we design modality-specific emotion-aware tokens ([EMO] for speech, emotion-guided [MASK] for text) optimized via prototypical contrastive learning, enabling unimodal networks to concentrate on affect-salient features. Second, we design an anchor-guided cross-modal attention mechanism that prioritizes interactions between emotion-aware tokens across modalities, effectively suppressing irrelevant signals while amplifying discriminative emotional cues. Extensive experiments on IEMOCAP and MSP-Podcast demonstrate that EmoFusion achieves 81.25% UA on IEMOCAP and 65.60% UA on MSP-Podcast, outperforms recent competitive methods. Jiaqian Ren, Xupu Cai, Kan Guo |
IEEE Signal Process. Lett. | 4 |
| 2026 | SSM-Det: State Space Model-Based Object Detector for Intelligent Transportation SystemabstractThe State Space Model (SSM) has been a growth of interest in computer vision due to its long-term dependency modeling with linear complexity. Despite massive endeavor, it has not been extensively explored in intelligent transportation system (ITS) yet. In this paper, we propose State Space Model-based object Detector (SSM-Det), that is meticulously curated with Direction-aware Visual State Space Encoder (D-VSSE). Specifically, it customizes multi-path pixel exchange and patch re-arrangement via four-direction scanning mechanism, promoting for information communication. To bridge the information bottleneck across high-low level, we further design Split-Fusion (SF) and Skip-Connection (SC) modules for contextual feature propagation before decoding: SF performs multi-channel semantic separation and re-weighting in global-local scope, while SC is responsible for cross-layer feature interaction in a cascaded manner. Empirical studies is conducted on both VisDrone2019-DET and SEU_PML benchmarks, and our proposed SSM-Det reports the state-of-the-art performance against all counterparts by a substantial margin, while maintaining the real-time inference speed. We hope this work contributes to the in-depth investigation of SSM-based detector for intelligent transportation applications. The code is available athttps://buaawjq.github.io/SSM-Det/. Chunmian Lin, Kan Guo, Jiangang Guo |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Spatial-Temporal Traffic Prediction Based on Multi-Scale Time Difference
Yongli Hu, Qi Zuo, Kan Guo, Zhongfan Sun, Tingzheng Jia |
ICIC (11) | 3 |
| 2025 | Large-Small Model Synergy with Multimodal Fine-Grained Heuristics for Knowledge-Based Visual Question AnsweringabstractMultimodal Large Language Models (MLLMs) possess extensive knowledge and strong reasoning capabilities, achieving remarkable performance in knowledge-based visual question answering, significantly surpassing traditional small-scale Vision-Language Models (VLMs). However, the distinct training paradigms of MLLMs and small-scale VLMs result in misaligned feature representation spaces and divergent answer prediction distributions. To bridge this gap, we propose a novel end-to-end large-small model synergy framework, where small VLMs and MLLMs collaborate via synergistic optimization of shared objectives while maintaining their co-evolving complementary specializations. Specifically, multimodal fine-grained heuristics are extracted from well-tuned small VLMs and subsequently projected into the textual space of MLLMs through dedicated visual and textual collaboration modules. This enables cross-modal guidance for both visual and textual inputs. Finally, a dual-objective synergy loss promotes alignment toward shared goals, while a visual discrepancy loss preserves specialization diversity. Extensive experiments demonstrate that our framework achieves state-of-the-art performance on both the OK-VQA and A-OKVQA benchmarks. Zhongfan Sun, Kan Guo, Yongli Hu, Daxin Tian, Qingqing Gao, Jiapu Wang, Junbin Gao |
ACM Multimedia | 2 |
| 2025 | ICL4RUL: In-Context Learning-Based Aircraft Engine Remaining Useful Life PredictionabstractAccurately predicting the remaining useful life (RUL) of an aircraft engine is critical for enhancing aircraft reliability and safety. To address the issues of recurrent neural network (RNN)’s inability to prioritize the significance of the most contributive temporal weights and the gradient vanishing problem arising from deep training, this study proposes a novel model that integrates a multi-head attention mechanism (MHA) into a residual network (ResNet) and introduces bi-directional long short-term memory (BiLSTM) with adaptive degradation temporal weighting (ADTW) module where two novel downsampling techniques are drsigned to predict the RUL of aircraft engines. This model, referred to as in-context learning-based aircraft engine RUL prediction (ICL4RUL), captures both temporal and spatial contextual features of an aircraft engine’s operating state. Specifically, it identifies intrinsic temporal evolution patterns in the time series data from each sensor across different operational phases, as well as the spatial correlations among sensors located in various subsystems. As a result, the model enhances the accuracy and stability of RUL prediction through its robust ability to extract contextual patterns. Through a comparative analysis of the NASA C-MAPSS dataset, the model demonstrates superior RUL prediction performance in terms of three evaluation metrics root mean square error (RMSE), Score and coefficient of determination (R2). Through ablation study and analysis of heatmaps, the advantages of ADTW and CrossResMHA are further demonstrated. Moreover, additional extensive experimental results on the real-world IEEE PHM2012 dataset are provided to further validate our model. The findings of this study can be effectively incorporated into the digital twin framework and Industrial Internet of Things (IIoT), optimizing decision-making processes related to aircraft engine maintenance. Guixian Qu, Shuiting Ding, Kan Guo |
IEEE Internet Things J. | 6 |
| 2025 | MDPM: Modulating domain-specific prompt memory for multi-domain traffic flow prediction with transformers
Zhuang Zhuang, Lingbo Liu, Kan Guo, Xingtong Yu, Heng Qi, Yanming Shen |
Knowl. Based Syst. | 3 |
| 2025 | DMP: Difference-Guided Motion Prediction for Vision-Centric Autonomous DrivingabstractVision-centric motion prediction concentrates on accurately determining the instance mask and its future trajectory from surround-view cameras, which manifests inherent merits such as holistic perspective and fully-differentiable spirit. Nonetheless, it is still impeded by sparse bird’s-eye view (BEV) representation and unfavorable temporal context across frames, resulting in a sub-optimal solution to decision-making and vehicle navigation. In this work, we propose a novelDifference-guideMotionPrediction for vision-centric autonomous driving, that is DMP, where it integrates BEV map refinement with spatial-temporal relation modeling in a hierarchical manner. Specifically, a bidirectional view projection strategy is introduced for the complementary BEV feature generation via depth-consistency correction. To promote spatiotemporal context aggregation, we design a difference-guided motion approach by offset approximation to align motion-aware cues between adjacent frames, and a dual-stream pyramid module is further developed for historical information fusion and future instance segmentation during specific durations. Extensive experiments on the large-scale nuScenes dataset demonstrate that it outperforms the baselines by a remarkable margin and delivers competitive motion prediction across diverse scenarios and range settings, suggesting its effectiveness and superiority. The details will be available athttps://github.com/pupu-chenyanyan/DMP-VAD. Chunmian Lin, Xuting Duan, Jianshan Zhou, Kan Guo, Dezong Zhao, Dongpu Cao, Daxin Tian |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Contrastive optimized graph convolution network for traffic forecasting
Kan Guo, Daxin Tian, Yongli Hu, Zhen (sean) Qian, Jianshan Zhou, Junbin Gao |
Neurocomputing | 1 |
| 2024 | CFMMC-Align: Coarse-Fine Multi-Modal Contrastive Alignment Network for Traffic Event Video Question AnsweringabstractTraffic video question answering (TrafficVQA) constitutes a specialized VideoQA task designed to enhance the basic comprehension and intricate reasoning capacities of videos, specifically focusing on traffic events. Recent VideoQA models employ pretrained visual and textual encoder models to bridge the feature space gap between visual and textual data. However, in addressing the unique challenges inherent to the TrafficVQA task, three pivotal issues must be addressed: (i) Dimension Gap: Between the pretrained image (appearance feature) and video (motion feature) models, there exists a conspicuous dimension difference in static and dynamic visual data; (ii) Scene Gap: The common real-world datasets and the traffic event datasets differ in visual scene content; (iii) Modality Gap: A pronounced feature distribution discrepancy emerges between traffic video and text data. To alleviate these challenges, we introduce the coarse-fine multimodal contrastive alignment network (CFMMC-Align). This model leverages sequence-level and token-level multimodal features, grounded in an unsupervised visual multimodal contrastive loss to mitigate dimension and scene gaps and a supervised visual-textual contrastive loss to alleviate modality discrepancies. Finally, the model is validated on the challenging public TrafficVQA dataset SUTD-TrafficQA and outperforms the state-of-the-art method by a substantial margin (50.2%compared to46.0%). The code is available at https://github.com/guokan987/CFMMC-Align. Kan Guo, Daxin Tian, Yongli Hu, Chunmian Lin, Jianshan Zhou, Xuting Duan, Junbin Gao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Person Foreground Segmentation by Learning Multi-Domain NetworksabstractSeparating the dominant person from the complex background is significant to the human-related research and photo-editing based applications. Existing segmentation algorithms are either too general to separate the person region accurately, or not capable of achieving real-time speed. In this paper, we introduce the multi-domain learning framework into a novel baseline model to construct the Multi-domain TriSeNet Networks for the real-time single person image segmentation. We first divide training data into different subdomains based on the characteristics of single person images, then apply a multi-branch Feature Fusion Module (FFM) to decouple the networks into the domain-independent and the domain-specific layers. To further enhance the accuracy, a self-supervised learning strategy is proposed to dig out domain relations during training. It helps transfer domain-specific knowledge by improving predictive consistency among different FFM branches. Moreover, we create a large-scale single person image segmentation dataset named MSSP20k, which consists of 22,100 pixel-level annotated images in the real world. The MSSP20k dataset is more complex and challenging than existing public ones in terms of scalability and variety. Experiments show that our Multi-domain TriSeNet outperforms state-of-the-art approaches on both public and the newly built datasets with real-time speed. Zhiyuan Liang, Kan Guo, Xiaogang Jin 0001, Jianbing Shen |
IEEE Trans. Image Process. | 2 |
| 2022 | Dynamic Graph Convolution Network for Traffic Forecasting Based on Latent Network of Laplace Matrix EstimationabstractTraffic forecasting is a challenging problem in the transportation research field as the complexity and non-stationary changing of the traffic data, thus the key to the issue is how to explore proper spatial and temporal characteristics. Based on this thought, many creative methods have been proposed, in which Graph Convolution Network (GCN) based methods have shown promising performance. However, these methods depend on the graph construction, which mainly uses the prior knowledge of the road network. Recently, some works realized the fact of the road network graph changing and tried to construct dynamic graphs for GCN, but they do not fully exploit the spatial and temporal properties of the traffic data in the graph construction. In this paper, we propose a novel dynamic graph convolution network for traffic forecasting, in which a latent network is introduced to extract spatial-temporal features for constructing the dynamic road network graph matrices adaptively. The proposed method is evaluated on several traffic datasets and the experimental results show that it outperforms the state of the art traffic forecasting methods. The website of the code ishttps://github.com/guokan987/DGCN.git. Kan Guo, Yongli Hu, Sean Qian, Junbin Gao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Dual Dynamic Spatial-Temporal Graph Convolution Network for Traffic PredictionabstractRecently, Graph Convolution Network (GCN) and Temporal Convolution Network (TCN) are introduced into traffic prediction and achieve state-of-the-art performance due to their good ability for modeling the spatial and temporal property of traffic data. In spite of having good performance, the current methods generally focus on the traffic measurement of road segments, i.e. the nodes of traffic flow graph, while the edges of the graph, which represent the correlation of traffic data of different road segments and form the affinity matrix for GCN, are usually constructed according to the structure of road network, but the spatial and temporal properties are not well exploited in their theories. In this paper, we propose a Dual Dynamic Spatial-Temporal Graph Convolution Network (DDSTGCN), which not only models the dynamic property of the nodes of the traffic flow graph but also captures the dynamic spatial-temporal feature of the edges by transforming the traffic flow graph into its dual hypergraph. The traffic prediction is enhanced by the collaborative convolutions on the traffic flow graph and its dual hypergraph. The proposed method is evaluated by extensive traffic prediction experiments on six real road datasets and the results show that it outperforms state-of-the-art related methods. Source codes are available athttps://github.com/j1o2h3n/DDSTGCN. Xiangheng Jiang, Yongli Hu, Fuqing Duan, Kan Guo, Boyue Wang, Junbin Gao |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Hierarchical Graph Convolution Network for Traffic ForecastingabstractTraffic forecasting is attracting considerable interest due to its widespread application in intelligent transportation systems. Given the complex and dynamic traffic data, many methods focus on how to establish a spatial-temporal model to express the non-stationary traffic patterns. Recently, the latest Graph Convolution Network (GCN) has been introduced to learn spatial features while the time neural networks are used to learn temporal features. These GCN based methods obtain state-of-the-art performance. However, the current GCN based methods ignore the natural hierarchical structure of traffic systems which is composed of the micro layers of road networks and the macro layers of region networks, in which the nodes are obtained through pooling method and could include some hot traffic regions such as downtown and CBD etc., while the current GCN is only applied on the micro graph of road networks. In this paper, we propose a novel Hierarchical Graph Convolution Networks (HGCN) for traffic forecasting by operating on both the micro and macro traffic graphs. The proposed method is evaluated on two complex city traffic speed datasets. Compared to the latest GCN based methods like Graph WaveNet, the proposed HGCN gets higher traffic forecasting precision with lower computational cost.The website of the code is https://github.com/guokan987/HGCN.git. Kan Guo, Yongli Hu, Sean Qian, Junbin Gao |
AAAI | 1 |
| 2021 | Optimized Graph Convolution Recurrent Neural Network for Traffic PredictionabstractTraffic prediction is a core problem in the intelligent transportation system and has broad applications in the transportation management and planning, and the main challenge of this field is how to efficiently explore the spatial and temporal information of traffic data. Recently, various deep learning methods, such as convolution neural network (CNN), have shown promising performance in traffic prediction. However, it samples traffic data in regular grids as the input of CNN, thus it destroys the spatial structure of the road network. In this paper, we introduce a graph network and propose an optimized graph convolution recurrent neural network for traffic prediction, in which the spatial information of the road network is represented as a graph. Additionally, distinguishing with most current methods using a simple and empirical spatial graph, the proposed method learns an optimized graph through a data-driven way in the training phase, which reveals the latent relationship among the road segments from the traffic data. Lastly, the proposed method is evaluated on three real-world case studies, and the experimental results show that the proposed method outperforms state-of-the-art traffic prediction methods. Kan Guo, Yongli Hu, Sean Qian, Hao Liu 0040, Ke Zhang 0016, Junbin Gao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Augmented Bi-path Network for Few-shot LearningabstractFew-shot Learning (FSL) which aims to learn from few labeled training data is becoming a popular research topic, due to the expensive labeling cost in many real-world applications. One kind of successful FSL method learns to compare the testing (query) image and training (support) image by simply concatenating the features of two images and feeding it into the neural network. However, with few labeled data in each class, the neural network has difficulty in learning or comparing the local features of two images. Such simple image-level comparison may cause serious mis-classification. To solve this problem, we propose Augmented Bi-path Network (ABNet) for learning to compare both global and local features on multi-scales. Specifically, the salient patches are extracted and embedded as the local features for every image. Then, the model learns to augment the features for better robustness. Finally, the model learns to compare global and local features separately, i.e., in two paths, before merging the similarities. Extensive experiments show that the proposed ABNet outperforms the state-of-the-art methods. Both quantitative and visual ablation studies are provided to verify that the proposed modules lead to more precise comparison results. Baoming Yan, Bo Zhao 0015, Kan Guo, Ming Zhang 0004, Yizhou Wang 0001 |
ICPR | 4 |
| 2020 | Semantic part segmentation of single-view point cloud
Haotian Peng, Liyuan Yin, Kan Guo, Qinping Zhao |
Sci. China Inf. Sci. | 4 |
| 2019 | Quadruplet Network With One-Shot Learning for Fast Visual Object TrackingabstractIn the same vein of discriminative one-shot learning, Siamese networks allow recognizing an object from a single exemplar with the same class label. However, they do not take advantage of the underlying structure of the data and the relationship among the multitude of samples as they only rely on the pairs of instances for training. In this paper, we propose a new quadruplet deep network to examine the potential connections among the training instances, aiming to achieve a more powerful representation. We design a shared network with four branches that receive a multi-tuple of instances as inputs and are connected by a novel loss function consisting of pair loss and triplet loss. According to the similarity metric, we select the most similar and the most dissimilar instances as the positive and negative inputs of triplet loss from each multi-tuple. We show that this scheme improves the training performance. Furthermore, we introduce a new weight layer to automatically select suitable combination weights, which will avoid the conflict between triplet and pair loss leading to worse performance. We evaluate our quadruplet framework by model-free tracking-by-detection of objects from a single initial exemplar in several visual object tracking benchmarks. Our extensive experimental analysis demonstrates that our tracker achieves superior performance with a real-time processing speed of 78 frames/s. Our source code is available. Xingping Dong, Jianbing Shen, Dongming Wu 0005, Kan Guo, Xiaogang Jin 0001, Fatih Porikli |
IEEE Trans. Image Process. | 4 |
| 2018 | 3D shape co-segmentation via sparse and low rank representations
Liyuan Yin, Kan Guo, Qinping Zhao |
Sci. China Inf. Sci. | 2 |
| 2018 | Image-guided 3D model labeling via multiview alignment
Kan Guo, Xiaowu Chen 0001, Qinping Zhao |
Graph. Model. | 1 |
| 2015 | Monocular Video Guided Garment Simulation
Xiaowu Chen 0001, Fei-Xiang Lu, Kan Guo, Qiang Fu 0004 |
J. Comput. Sci. Technol. | 5 |
| 2015 | 3D Mesh Labeling via Deep Convolutional Neural NetworksabstractThis article presents a novel approach for 3D mesh labeling by using deep Convolutional Neural Networks (CNNs). Many previous methods on 3D mesh labeling achieve impressive performances by using predefined geometric features. However, the generalization abilities of such low-level features, which are heuristically designed to process specific meshes, are often insufficient to handle all types of meshes. To address this problem, we propose to learn a robust mesh representation that can adapt to various 3D meshes by using CNNs. In our approach, CNNs are first trained in a supervised manner by using a large pool of classical geometric features. In the training process, these low-level features are nonlinearly combined and hierarchically compressed to generate a compact and effective representation for each triangle on the mesh. Based on the trained CNNs and the mesh representations, a label vector is initialized for each triangle to indicate its probabilities of belonging to various object parts. Eventually, a graph-based mesh-labeling algorithm is adopted to optimize the labels of triangles by considering the label consistencies. Experimental results on several public benchmarks show that the proposed approach is robust for various 3D meshes, and outperforms state-of-the-art approaches as well as classic learning algorithms in recognizing mesh labels. Kan Guo, Dongqing Zou, Xiaowu Chen 0001 |
ACM Trans. Graph. | 1 |
| 2013 | Garment Modeling from a Single ImageabstractAbstract Modeling of realistic garments is essential for online shopping and many other applications including virtual characters. Most of existing methods either require a multi‐camera capture setup or a restricted mannequin pose. We address the garment modeling problem according to a single input image. We design an all‐pose garment outline interpretation, and a shading‐based detail modeling algorithm. Our method first estimates the mannequin pose and body shape from the input image. It further interprets the garment outline with an oriented facet decided according to the mannequin pose to generate the initial 3D garment model. Shape details such as folds and wrinkles are modeled by shape‐from‐shading techniques, to improve the realism of the garment model. Our method achieves similar result quality as prior methods from just a single image, significantly improving the flexibility of garment modeling. Xiaowu Chen 0001, Qiang Fu 0004, Kan Guo |
Comput. Graph. Forum | 4 |