EDBT 2026 Demo / reviewers in the wild / expert
Gaoqi He
dblp:73/4983
· DBLP profile ↗
62ranked-venue papers
7as first author
51since 2021 · last 2026
0000-0001-8365-0970ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 1 first-author · 37 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KG-CPEN: Knowledge-Guided Compositional Prototype Evolution for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) serves as a pivotal bridge between Computer Vision and Natural Language Processing, aiming to parse unstructured imagery into structured semantic summaries. When dealing with long-tailed distributions, existing discriminative methods often suffer from severe data-dependency bias and ignore the compositional semantics of predicates, resulting in predictions inevitably collapsing into high-frequency head classes. To address this, we propose the Knowledge-Guided Compositional Prototype Evolution Network (KG-CPEN). To tackle the paucity of feature space caused by tail sample scarcity, we introduce a Knowledge-Guided Prototype Evolution mechanism, which utilizes external knowledge as a semantic engine to synthesize missing tail prototypes in the visual manifold, effectively replenishing feature representations in long-tailed and zero-shot scenarios. Furthermore, since traditional local greedy matching exacerbates head bias, we design an Optimal Transport Alignment module. This achieves unbiased matching by minimizing the distance between visual and prototype distributions on a global scale, thereby reducing decision layer difficulty and eliminating head class dominance. Extensive experiments on Visual Genome, GQA, and Open Images V6 consistently demonstrate that KG-CPEN establishes a new State-Of-The-Art (SOTA) for unbiased scene graph generation. Yujun Hu, Changbo Wang, Gaoqi He |
ICMR | 3 |
| 2026 | MEDP: Multimodal-Enhanced Dynamic Prototype learning for few-shot dynamic scene graph generation
Ziheng Huang, Weiliang Meng, Changbo Wang, Gaoqi He |
Knowl. Based Syst. | 5 |
| 2025 | Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationabstractRecent large-scale pre-trained diffusion models have demonstrated a powerful generative ability to produce high-quality videos from detailed text descriptions. However, exerting control over the motion of objects in videos generated by any video diffusion model remains a challenging problem. In this paper, we propose a novel zero-shot moving object trajectory control framework, Motion-Zero, to enable arbitrary single-object-trajectory control for the text-to-video diffusion model. To this end, an initial noise prior module is designed to provide a position-based prior to improve the stability of the appearance of the moving object and the accuracy of position. In addition, based on the attention map of the U-Net, spatial constraints are directly applied to the denoising process of diffusion models, which further ensures the positional consistency of moving objects during the inference. Furthermore, temporal consistency is guaranteed with a proposed shift temporal attention mechanism. Our method can be flexibly applied to various state-of-the-art video diffusion models without any training process. Extensive experiments demonstrate our proposed method can control the motion trajectories of arbitrary objects while preserving the original ability to generate high-quality videos. Changgu Chen, Junwei Shu, Gaoqi He, Changbo Wang, Yang Li 0041 |
AAAI | 3 |
| 2025 | Multi-granularity Feature Extraction Based on Long-Short Chains for Motion Retargeting
Weiliang Meng, Changbo Wang, Gaoqi He |
CGI (3) | 5 |
| 2025 | SandTouch: Empowering Virtual Sand Art in VR with AI Guidance and Emotional Relief
Junbin Ren, Zeyuan Fan, Chenhui Li 0001, Gaoqi He, Changbo Wang, Yang Gao 0025, Chen Li 0035 |
CHI | 5 |
| 2025 | Weakly Supervised Semantic Segmentation via Progressive Confidence Region ExpansionabstractWeakly supervised semantic segmentation (WSSS) has garnered considerable attention due to its effective reduction of annotation costs. Most approaches utilize Class Activation Maps (CAM) to produce pseudo-labels, thereby localizing target regions using only image-level annotations. However, the prevalent methods relying on vision transformers (ViT) encounter an "over-expansion" issue, i.e., CAM incorrectly expands high activation value from the target object to the background regions, as it is difficult to learn pixel-level local intrinsic inductive bias in ViT from weak supervisions. To solve this problem, we propose a Progressive Confidence Region Expansion (PCRE) framework for WSSS, it gradually learns a faithful mask over the target region and utilizes this mask to correct the confusion in CAM. PCRE has two key components: Confidence Region Mask Expansion (CRME) and Class-Prototype Enhancement (CPE). CRME progressively expands the mask in the small region with the highest confidence, eventually encompassing the entire target, thereby avoiding unintended coverage of background areas. CPE aims to enhance mask generation in CRME by leveraging the similarity between the learned, dataset-level class prototypes and patch features as supervision to optimize the mask output from CRME. Extensive experiments demonstrate that our method outperforms the existing single-stage and multi-stage approaches on the PASCAL VOC and MS COCO benchmark. Our code is available at https://github.com/xxf011/WSSS-PCRE. Xiangfeng Xu, Pinyi Zhang, Wenxuan Huang 0001, Yunhang Shen, Jingzhong Lin, Wei Li 0002, Gaoqi He, Jiao Xie, Shaohui Lin |
CVPR | 8 |
| 2025 | Dynamic Stereotype Theory Induced Micro-expression Recognition with Oriented DeformationabstractMicro-expression recognition (MER) aims to uncover genuine emotions and underlying psychological states. However, existing MER methods struggle with three main challenges. 1) Scarcity of micro-expression samples. 2) Difficulty in modeling nearly imperceptible facial movements. 3) Reliance on apex frame annotations. To address these issues, we propose a Self-supervised Oriented Deformation model for Apex-free Micro-expression Recognition (SODA4MER). Our approach enhances local deformation perception using muscle-group priors and amplifies subtle features through Dynamic Stereotype Theory (DST) based enhancement, while contrastive learning eliminates the need for manual apex annotations. Specifically, the Oriented deformation estimator of SODA4MER is first pretrained in a self-supervised manner. Secondly, a Gated Temporal Variance Gaussian model (GTVG) is introduced to adaptively integrate facial muscle-group priors, enhancing local deformation perception and mitigating noise from head movements. Then, contrastive learning is employed to achieve apex detection by identifying the frame with the most significant local deformation. Finally, guided by DST, we introduced a feature enhancement strategy that models the temporal dynamics of local deformation in the activation and decay phases, leading to richer deformation features. Our rigorous experiments confirm the competitive performance and practical applicability of SODA4MER. Bohao Zhang, Changbo Wang, Gaoqi He |
CVPR | 4 |
| 2025 | CoT Reasoning-Based Content Adaptation and Image Generation for Chinese Poetry
Jihao Chen, Songtao Chen, Gaoqi He |
ICIC (24) | 5 |
| 2025 | Shape-Preserving and Surface-Fitting Network for 3D Lane DetectionabstractCurrent transformer-based 3D lane detection methods typically use instance activation maps (IAM) and point-to-point loss to achieve small geometric deviations of lanes. However, these methods suffer from such lane visibility issues as wrong lane extensions and lane omissions because IAM places the lane vanishing points by mistake and loses the blurred lanes. And their performance is limited by lane continuity issues while using the point-to-point loss. In this paper, we propose a shape-preserving and surface-fitting (SPSF) network to improve the lane visibility and enhance the lane continuity. The proposed SPSF network consists of three key steps: 3D lane preliminary prediction, lane shape-preserving, and 3D lane surface-fitting. First, we design a novel transformer decoder with a mask-guided denoising block to predict preliminary 3D lanes after generating 2D lane masks. Next, after mapping the preliminary 3D lanes to 2D projected lanes, lane shapes are preserved using a two-stage mask-guided strategy to avoid the visibility issues. The two-stage mask-guided strategy includes mask-directed horizontal position adjustment and visibility correction. Finally, after fitting the surface of 3D Lanes, we improve the continuity of lanes through a surface-fitting loss. Various experiments show that our work achieves SOTA performance on two standard benchmarks. Jianhua Li 0009, Gaoqi He, Weiliang Meng |
ICME | 3 |
| 2025 | DiffLane: Diffusion Model-Based Lane Mask Generation for Accurate Video Lane DetectionabstractMask-based video lane detection methods currently have achieved promising performance. However, they generate irregular lane masks in complex scenes, resulting in inaccurate lane positioning. Diffusion models have achieved notable success in the field of image segmentation because of their ability to restore pixel-level details. In this paper, we propose a novel framework DiffLane, termed Diffusion Model-Based Lane Mask Generation for Accurate Video Lane Detection. The main idea of our work is to exploit the detail-restoring capability of diffusion models to generate high-quality lane masks. DiffLane includes the MultiFrame Fusion Enhancer (MFFE), the MultiScale De-noising Network (MSDN) and the Dynamic Lane Perception Unit (DLPU). In MFFE, the current frame is enhanced with visual information from the past two frames through global matching-based optical flow estimation. This enhanced frame serves as a condition for each denoising step. MSDN predicts noise through a multi-scale fusion strategy, enabling the diffusion model to remove noise and generate regular lane masks precisely. DLPU regresses the coefficient vectors from the generated lane masks with DSConv applied in two directions, completing the accurate video lane detection task. Extensive experiments on the VIL-100 and OpenLane-V datasets demonstrate that our method outperforms other state-of-the-art approaches. Weiliang Meng, Gaoqi He, Jianhua Li 0009 |
ICME | 4 |
| 2025 | DPSN: Dual Prior Knowledge Induced Tactile paving and Obstacle Joint Segmentation NetworkabstractAccurate semantic segmentation of both tactile paving and the obstacle is crucial for the safe mobility of visually impaired individuals. However, existing methods face two major challenges: (i) discontinuous segmentation fragments; (ii) Inaccurate obstacle recognition. To address challenge (i), we propose incorporating appearance priors of complete tactile pavings to prevent the model from directly learning irregular ground truth masks. To tackle challenge (ii), we propose introducing cross-modal semantic priors to complement the semantic information of obstacles. We implemented these strategies in proposed Dual Prior knowledge induced tactile paving and obstacle joint Segmentation Network (DPSN). Based on bilateral network architecture, DPSN merges obstacle category masks into tactile paving categories, constructing a complete tactile paving mask. Utilizing the complete mask, DPSN transfer appearance prior knowledge to detail features from boundary and structural perspectives. Concurrently, DPSN leverages the CLIP Text Encoder to guide visual feature decoding by attention mechanisms, transferring rich cross-modal semantic prior knowledge to the visual feature maps. Furthermore, we propose the TPO-Dataset, the first dataset for joint tactile paving and obstacle segmentation acquired from actual scenes. Experiments demonstrate that DPSN achieves state-of-the-art results on the TPO-Dataset, with relative gains of 27.16% in obstacle IoU and 30.53% in accuracy metrics compared to baseline methods. Notably, DPSN achieves real-time performance at 88.25 FPS on the maximum scale of 2048×512 resolution. Youqi Song, Zilong Jin, Changbo Wang, Gaoqi He |
IROS | 7 |
| 2025 | 3D Scene Graph Generation with Cross-Modal Alignment and Adversarial Learning
Yujun Hu, Changbo Wang, Weiliang Meng, Gaoqi He |
ICMR | 5 |
| 2025 | D3L: Curvature-Constrained Denoising Diffusion Model for 3D Lane DetectionabstractMonocular 3D lane detection is a challenging task for autonomous driving systems. Recent advances primarily focus on one-step methods for lane detection based on front-view features, which show promising results on straight lanes. However, curved lanes are difficult to handle with one-step prediction, which performs prediction in a single leap without gradual refinement. To address this issue, we propose a novel Denoising Diffusion Model for 3D Lane Detection framework (D3L). The main idea is to leverage the progressive generation capability of the diffusion model to generate accurate 3D curved lanes, and ensuring lane continuity through curvature constraints. The framework includes three creative components: coarse-to-fine denoiser (CFD), curvature-constrained loss (CCL) and multi-sampling aggregation strategy (MSAS). In CFD, both lane-level and point-level transformer blocks are integrated to accurately denoise 3D lanes, which effectively captures both global and local features. CCL is designed to reduce deviations in lane curvature, resulting in smoother lane continuity. This loss enhances both the accuracy and geometric consistency of lane detection, especially in complex curved scenes. MSAS is proposed to select the optimal lane point-by-point from multiple candidates, thus robustness of the lane prediction is significantly improved. Extensive experiments on two popular 3D lane detection benchmarks demonstrate that our D3 L outperforms the state-of-the-art methods. Weiliang Meng, Gaoqi He, Jianhua Li 0009 |
ACM Multimedia | 4 |
| 2025 | Regulatory Focus Theory Induced Micro-Expression Analysis with Structured Representation Learning
Bohao Zhang, Haoxin Xu, Jingzhong Lin, Changbo Wang, Gaoqi He |
ACM Multimedia | 5 |
| 2025 | TactPav: A Vision-Language Annotated Multi-modal Dataset for Tactile Paving Navigation
Youqi Song, Zilong Jin, Yunjie Xie, Changbo Wang, Gaoqi He |
PRCV (12) | 8 |
| 2025 | EmoDiffGes: Emotion-Aware Co-Speech Holistic Gesture Generation with Progressive Synergistic DiffusionabstractAbstract Co‐speech gesture generation, driven by emotional expression and synergistic bodily movements, is essential for applications such as virtual avatars and human‐robot interaction. Existing co‐speech gesture generation methods face two fundamental limitations: (1) producing inexpressive gestures due to ignoring the temporal evolution of emotion; (2) generating incoherent and unnatural motions as a result of either holistic body oversimplification or independent part modeling. To address the above limitations, we propose EmoDiffGes, a diffusion‐based framework grounded in embodied emotion theory, unifying dynamic emotion conditioning and part‐aware synergistic modeling. Specifically, a Dynamic Emotion‐Alignment Module (DEAM) is first applied to extract dynamic emotional cues and inject emotion guidance into the generation process. Then, a Progressive Synergistic Gesture Generator (PSGG) iteratively refines region‐specific latent codes while maintaining full‐body coordination, leveraging a Body Region Prior for part‐specific encoding and Progressive Inter‐Region Synergistic Flow for global motion coherence. Extensive experiments validate the effectiveness of our methods, showcasing the potential for generating expressive, coordinated, and emotionally grounded human gestures. Jingzhong Lin, Bohao Zhang, Changbo Wang, Gaoqi He |
Comput. Graph. Forum | 6 |
| 2025 | Precise Motion Inbetweening via Bidirectional Autoregressive Diffusion ModelsabstractABSTRACT Conditional motion diffusion models have demonstrated significant potential in generating natural and reasonable motions response to constraints such as keyframes, that can be used for motion inbetweening task. However, most methods struggle to match the keyframe constraints accurately, which resulting in unsmooth transitions between keyframes and generated motion. In this article, we propose Bidirectional Autoregressive Motion Diffusion Inbetweening (BAMDI) to generate seamless motion between start and target frames. The main idea is to transfer the motion diffusion model to autoregressive paradigm, which predicts subsequence of motion adjacent to both start and target keyframes to infill the missing frames through several iterations. This can help to improve the local consistency of generated motion. Additionally, bidirectional generation make sure the smoothness on both start frame target keyframes. Experiments show our method achieves state‐of‐the‐art performance compared with other diffusion‐based motion inbetweening methods. Jiawen Peng, Jingzhong Lin, Gaoqi He |
Comput. Animat. Virtual Worlds | 4 |
| 2025 | DCS-RISR: Dynamic channel splitting for efficient real-world image super-resolution
Junbo Qiao, Shaohui Lin, Yulun Zhang 0001, Wei Li 0002, Jie Hu 0021, Gaoqi He, Changbo Wang, Lizhuang Ma |
Neural Networks | 6 |
| 2025 | An Approach to Multi-AAV Ship Detection Based on Mobile Edge Computing ScenariosabstractAutonomous aerial vehicles (AAVs) are widely used for ship tracking and detection tasks. However, the real-time detection performance is limited by AAV battery capacity and computing power, resulting in a short operational duration. To address this challenge, this paper proposes a AAV ship detection system that focuses on two key aspects: algorithm improvement and computational resource allocation. Specifically, we introduce a lightweight ship detection method tailored for multi-AAV scenarios in a mobile edge computing environment. The proposed method first designs a multi-disentangled knowledge distillation approach based on an information decoupling framework and utilizes a newly designed teacher network to enhance the lightweight detection model. The teacher network disentangles two key types of entanglements: the relationship between the convolutional filters and target categories, and the relationship between the foreground and background regions in the feature maps. Additionally, a proximal policy optimization (PPO) reinforcement learning algorithm is designed to enable real-time decision-making for AAV motion, detection accuracy, and computational offloading. Finally, we validate the superiority of the proposed knowledge distillation method and demonstrate the robustness and effectiveness of the AAV path planning algorithm in various scenarios through a series of experiments. Compared to the improved student models YOLOv8-N and YOLOv10-N, our method improves [email protected] by 1.2% and 1.1% on the SeaShips7000 and FVessel validation sets. Furthermore, compared to the existing methods K-Means and DBSCAN, our approach achieves reward values approximately 2.0 times and 1.4 times higher, respectively. Tao Liu 0016, Zhengling Lei, Yuchi Huo, Xiaocai Zhang, Gaoqi He, Huafeng Wu |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2024 | Multi-Prototype Space Learning for Commonsense-Based Scene Graph GenerationabstractIn the domain of scene graph generation, modeling commonsense as a single-prototype representation has been typically employed to facilitate the recognition of infrequent predicates. However, a fundamental challenge lies in the large intra-class variations of the visual appearance of predicates, resulting in subclasses within a predicate class. Such a challenge typically leads to the problem of misclassifying diverse predicates due to the rough predicate space clustering. In this paper, inspired by cognitive science, we maintain multi-prototype representations for each predicate class, which can accurately find the multiple class centers of the predicate space. Technically, we propose a novel multi-prototype learning framework consisting of three main steps: prototype-predicate matching, prototype updating, and prototype space optimization. We first design a triple-level optimal transport to match each predicate feature within the same class to a specific prototype. In addition, the prototypes are updated using momentum updating to find the class centers according to the matching results. Finally, we enhance the inter-class separability of the prototype space through iterations of the inter-class separability loss and intra-class compactness loss. Extensive evaluations demonstrate that our approach significantly outperforms state-of-the-art methods on the Visual Genome dataset. Lianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu, Yang Li 0041, Changbo Wang, Gaoqi He |
AAAI | 8 |
| 2024 | Kumaraswamy Wavelet for Heterophilic Scene Graph GenerationabstractGraph neural networks (GNNs) has demonstrated its capabilities in the field of scene graph generation (SGG) by updating node representations from neighboring nodes. Actually it can be viewed as a form of low-pass filter in the spatial domain, which smooths node feature representation and retains commonalities among nodes. However, spatial GNNs does not work well in the case of heterophilic SGG in which fine-grained predicates are always connected to a large number of coarse-grained predicates. Blind smoothing undermines the discriminative information of the fine-grained predicates, resulting in failure to predict them accurately. To address the heterophily, our key idea is to design tailored filters by wavelet transform from the spectral domain. First, we prove rigorously that when the heterophily on the scene graph increases, the spectral energy gradually shifts towards the high-frequency part. Inspired by this observation, we subsequently propose the Kumaraswamy Wavelet Graph Neural Network (KWGNN). KWGNN leverages complementary multi-group Kumaraswamy wavelets to cover all frequency bands. Finally, KWGNN adaptively generates band-pass filters and then integrates the filtering results to better accommodate varying levels of smoothness on the graph. Comprehensive experiments on the Visual Genome and Open Images datasets show that our method achieves state-of-the-art performance. Lianggangxu Chen, Youqi Song, Shaohui Lin, Changbo Wang, Gaoqi He |
AAAI | 5 |
| 2024 | SPD-DDPM: Denoising Diffusion Probabilistic Models in the Symmetric Positive Definite SpaceabstractSymmetric positive definite(SPD) matrices have shown important value and applications in statistics and machine learning, such as FMRI analysis and traffic prediction. Previous works on SPD matrices mostly focus on discriminative models, where predictions are made directly on E(X|y), where y is a vector and X is an SPD matrix. However, these methods are challenging to handle for large-scale data. In this paper, inspired by denoising diffusion probabilistic model(DDPM), we propose a novel generative model, termed SPD-DDPM, by introducing Gaussian distribution in the SPD space to estimate E(X|y). Moreover, our model can estimate p(X) unconditionally and flexibly without giving y. On the one hand, the model conditionally learns p(X|y) and utilizes the mean of samples to obtain E(X|y) as a prediction. On the other hand, the model unconditionally learns the probability distribution of the data p(X) and generates samples that conform to this distribution. Furthermore, we propose a new SPD net which is much deeper than the previous networks and allows for the inclusion of conditional factors. Experiment results on toy data and real taxi data demonstrate that our models effectively fit the data distribution both unconditionally and conditionally. Yunchen Li, Gaoqi He, Yunhang Shen, Ke Li 0015, Xing Sun 0001, Shaohui Lin |
AAAI | 3 |
| 2024 | CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive Learningabstract3D Scene Graph Generation (3DSGG) aims to classify objects and their predicates within 3D point cloud scenes. However, current 3DSGG methods struggle with two main challenges. 1) The dependency on labor-intensive ground-truth annotations. 2) Closed-set classes training hampers the recognition of novel objects and predicates. Addressing these issues, our idea is to extract cross-modality features by CLIP from text and image data naturally related to 3D point clouds. Cross-modality features are used to train a robust 3D scene graph (3DSG)feature extractor. Specifically, we propose a novel Cross-Modality Contrastive Learning 3DSGG (CCL-3DSGG) method. Firstly, to align the text with 3DSG, the text is parsed into word level that are consistent with the 3DSG annotation. To enhance robustness during the alignment, adjectives are exchanged for different objects as negative samples. Then, to align the image with 3DSG, the camera view is treated as a positive sample and other views as negatives. Lastly, the recognition of novel object and predicate classes is achieved by calculating the cosine similarity between prompts and 3DSG features. Our rigorous experiments confirm the superior open-vocabulary capability and applicability of CCL-3DSGG in real-world contexts. Lianggangxu Chen, Jiale Lu, Shaohui Lin, Changbo Wang, Gaoqi He |
CVPR | 6 |
| 2024 | A General and Efficient Training for Transformer via Token ExpansionabstractThe remarkable performance of Vision Transformers (ViTs) typically requires an extremely large training cost. Existing methods have attempted to accelerate the training of ViTs, yet typically disregard method universality with accuracy dropping. Meanwhile, they break the training consistency of the original transformers, including the consistency of hyperparameters, architecture, and strategy, which prevents them from being widely applied to different Transformer networks. In this paper, we propose a novel token growth scheme Token Expansion (termed ToE) to achieve consistent training acceleration for ViTs. We introduce an “initialization-expansion-merging” pipeline to maintain the integrity of the intermediate feature distribution of original transformers, preventing the loss of crucial learnable information in the training process. ToE can not only be seamlessly integrated into the training and fine-tuning process of transformers (e.g., DeiT and LV-ViT), but also effective for efficient training frameworks (e.g., EfficientTrain), without twisting the original training hyperparameters, architecture, and introducing additional training strategies. Extensive experiments demonstrate that ToE achieves about 1.3× faster for the training of ViTs in a lossless manner, or even with performance gains over the full-token training baselines. Code is available at https://github.com/Osilly/TokenExpansion. Wenxuan Huang 0001, Yunhang Shen, Jiao Xie, Baochang Zhang 0001, Gaoqi He, Ke Li 0015, Xing Sun 0001, Shaohui Lin |
CVPR | 5 |
| 2024 | Knowledge Graph Information Bottleneck for Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) prediction is an important but challenging task in drug safety surveillance. With the accumulation of biological data, biomedical knowledge graphs (KGs) become available to model DDIs and related biological mechanisms. However, the presence of substantial noise in large-scale KGs hampers prediction performance and the identification of interpretable biological pathways. To fill the gaps, this paper proposes an information bottleneck-based (IB-based) framework that simultaneously denoises the KG and identifies key entities around drug pairs. Moreover, KG-based prediction methods rarely exploit the structural information of drug molecules. To this end, the proposed framework relates drug structures to IB objectives, together with a unique drug pair-centered readout to fuse molecular information into KG subgraph embeddings. Extensive experimental results and case studies demonstrate the effectiveness and interpretability of the framework. Gaoqi He, Kai Zhang 0001, Honglin Li 0003 |
IJCNN | 2 |
| 2024 | FIND: Fine-tuning Initial Noise Distribution with Policy Optimization for Diffusion ModelsabstractIn recent years, large-scale pre-trained diffusion models have demonstrated their outstanding capabilities in image and video generation tasks. However, existing models tend to produce visual objects commonly found in the training dataset, which diverges from user input prompts. The underlying reason behind the inaccurate generated results lies in the model's difficulty in sampling from specific intervals of the initial noise distribution corresponding to the prompt. Moreover, it is challenging to directly optimize the initial distribution, given that the diffusion process involves multiple denoising steps. In this paper, we introduce a Fine-tuning Initial Noise Distribution (FIND) framework with policy optimization, which unleashes the powerful potential of pre-trained diffusion networks by directly optimizing the initial distribution to align the generated contents with user-input prompts. To this end, we first reformulate the diffusion denoising procedure as a one-step Markov decision process and employ policy optimization to directly optimize the initial distribution. In addition, a dynamic reward calibration module is proposed to ensure training stability during optimization. Furthermore, we introduce a ratio clipping algorithm to utilize historical data for network training and prevent the optimized distribution from deviating too far from the original policy to restrain excessive optimization magnitudes. Extensive experiments demonstrate the effectiveness of our method in both text-to-image and text-to-video tasks, surpassing SOTA methods in achieving consistency between prompts and the generated content. Our method achieves 10 times faster than the SOTA approach. Changgu Chen, Libing Yang, Lianggangxu Chen, Gaoqi He, Changbo Wang, Yang Li 0041 |
ACM Multimedia | 5 |
| 2024 | Prototype-based contrastive substructure identification for molecular property predictionabstractSubstructure-based representation learning has emerged as a powerful approach to featurize complex attributed graphs, with promising results in molecular property prediction (MPP). However, existing MPP methods mainly rely on manually defined rules to extract substructures. It remains an open challenge to adaptively identify meaningful substructures from numerous molecular graphs to accommodate MPP tasks. To this end, this paper proposes Prototype-based cOntrastive Substructure IdentificaTion (POSIT), a self-supervised framework to autonomously discover substructural prototypes across graphs so as to guide end-to-end molecular fragmentation. During pre-training, POSIT emphasizes two key aspects of substructure identification: firstly, it imposes a soft connectivity constraint to encourage the generation of topologically meaningful substructures; secondly, it aligns resultant substructures with derived prototypes through a prototype-substructure contrastive clustering objective, ensuring attribute-based similarity within clusters. In the fine-tuning stage, a cross-scale attention mechanism is designed to integrate substructure-level information to enhance molecular representations. The effectiveness of the POSIT framework is demonstrated by experimental results from diverse real-world datasets, covering both classification and regression tasks. Moreover, visualization analysis validates the consistency of chemical priors with identified substructures. The source code is publicly available at https://github.com/VRPharmer/POSIT. Gaoqi He, Changbo Wang, Kai Zhang 0001, Honglin Li 0003 |
Briefings Bioinform. | 1 |
| 2024 | Improving rare relation inferring for scene graph generation using bipartite graph network
Jiale Lu, Lianggangxu Chen, Haoyue Guan, Shaohui Lin, Chunhua Gu, Changbo Wang, Gaoqi He |
Comput. Vis. Image Underst. | 7 |
| 2024 | SegCFT: Context-aware Fourier Transform for efficient semantic segmentation
Yinqi Zhang, Lingfu Jiang, Fuhai Chen, Jiao Xie, Baochang Zhang 0001, Gaoqi He, Shaohui Lin |
Neurocomputing | 6 |
| 2024 | KDPM: Knowledge-driven dynamic perception model for evacuation scene simulationabstractAbstract Evacuation scene simulation has become one important approach for public safety decision‐making. Although existing research has considered various factors, including social forces, panic emotions, and so forth, there is a lack of consideration of how complex environmental factors affect human psychology and behavior. The main idea of this paper is to model complex evacuation environmental factors from the perspective of knowledge and explore pedestrians' emergency response mechanisms to this knowledge. Thus, a knowledge‐driven dynamic perception model (KDPM) for evacuation scene simulation is proposed in this paper. This model combines three modules: knowledge dissemination, dynamic scene perception, and stress response. Both scenario knowledge and hazard source knowledge are extracted and expressed. The improved intelligent agent perception model is designed by adopting position determination. Moreover, a general adaptation syndrome (GAS) model is first presented by introducing a modified stress system model. Experimental results show that the proposed model is closer to the reality of real data sets. Kecheng Tang, Yuji Shen, Chen Li 0035, Gaoqi He |
Comput. Animat. Virtual Worlds | 5 |
| 2024 | A Transformative Topological Representation for Link Modeling, Prediction and Cross-Domain Network AnalysisabstractMany complex social, biological, or physical systems are characterized as networks, and recovering the missing links of a network could shed important lights on its structure and dynamics. A good topological representation is crucial to accurate link modeling and prediction, yet how to account for the kaleidoscopic changes in link formation patterns remains a challenge, especially for analysis in cross-domain studies. We propose a new link representation scheme by projecting the local environment of a link into a "dipole plane", where neighboring nodes of the link are positioned via their relative proximity to the two anchors of the link, like a dipole. By doing this, complex and discrete topology arising from link formation is turned to differentiable point-cloud distribution, opening up new possibilities for topological feature-engineering with desired expressiveness, interpretability and generalization. Our approach has comparable or even superior results against state-of-the-art GNNs, meanwhile with a model up to hundreds of times smaller and running much faster. Furthermore, it provides a universal platform to systematically profile, study, and compare link-patterns from miscellaneous real-world networks. This allows building a global link-pattern atlas, based on which we have uncovered interesting common patterns of link formation, i.e., the bridge-style, the radiation-style, and the community-style across a wide collection of networks with highly different nature. Kai Zhang 0001, Junchen Shen, Gaoqi He, Yu Sun 0076, Haibin Ling, Hongyuan Zha, Honglin Li 0003, Jie Zhang 0012 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Explicit Invariant Feature Induced Cross-Domain Crowd CountingabstractCross-domain crowd counting has shown progressively improved performance. However, most methods fail to explicitly consider the transferability of different features between source and target domains. In this paper, we propose an innovative explicit Invariant Feature induced Cross-domain Knowledge Transformation framework to address the inconsistent domain-invariant features of different domains. The main idea is to explicitly extract domain-invariant features from both source and target domains, which builds a bridge to transfer more rich knowledge between two domains. The framework consists of three parts, global feature decoupling (GFD), relation exploration and alignment (REA), and graph-guided knowledge enhancement (GKE). In the GFD module, domain-invariant features are efficiently decoupled from domain-specific ones in two domains, which allows the model to distinguish crowds features from backgrounds in the complex scenes. In the REA module both inter-domain relation graph (Inter-RG) and intra-domain relation graph (Intra-RG) are built. Specifically, Inter-RG aggregates multi-scale domain-invariant features between two domains and further aligns local-level invariant features. Intra-RG preserves taskrelated specific information to assist the domain alignment. Furthermore, GKE strategy models the confidence of pseudolabels to further enhance the adaptability of the target domain. Various experiments show our method achieves state-of-theart performance on the standard benchmarks. Code is available at https://github.com/caiyiqing/IF-CKT. Yiqing Cai, Lianggangxu Chen, Haoyue Guan, Shaohui Lin, Changhong Lu, Changbo Wang, Gaoqi He |
AAAI | 7 |
| 2023 | Learning Local Features of Motion Chain for Human Motion Prediction
Lianggangxu Chen, Chen Li 0035, Changbo Wang, Gaoqi He |
CGI (3) | 5 |
| 2023 | KDEM: A Knowledge-Driven Exploration Model for Indoor Crowd Evacuation Simulation
Yuji Shen, Bohao Zhang, Chen Li 0035, Changbo Wang, Gaoqi He |
CGI (3) | 5 |
| 2023 | Scene Graph Generation using Depth-based Multimodal NetworkabstractScene graph generation (SGG) provides an efficient way for scene understanding. However, it has been plagued by the inaccurate classification of relative spatial relationship and incorrect feature information aggregation from distant objects. In this paper, we innovatively introduce the depth information of objects into SGG and propose a multimodal edge-featured graph attention network (MEGA-Net). MEGA-Net primarily comprises three modules. First, the edge-aware message passing (EMP) module extracts multimodal features and fuses them as edge features in the graph network via a quadrilinear model. Multimodal features consist of depth features, visual features, spatial features, and linguistic features. The depth feature in EMP provides the relative spatial relationship among objects which prevents the tail spatial predicates from being recognized as the head predicates. Second, we propose a depth-based self-supervised graph attention (DSGAT) module to predict the correlation probability between object pairs. By encoding the depth ranking of different object pairs in 2D images, DSGAT learns more accurate directional attention to avoid unrelated neighbors. Third, we introduce a predicate aware loss (PA-Loss) to alleviate the feature redundancy problem caused by extra depth information. This is achieved by introducing semantic frequency information that reflects the priority between different types of relationships. Systematic experiments show that our method achieves state-of-the-art performance on two popular datasets, VG and VRD. Lianggangxu Chen, Jiale Lu, Changbo Wang, Gaoqi He |
ICME | 4 |
| 2023 | Beware of Overcorrection: Scene-induced Commonsense Graph for Scene Graph GenerationabstractA scene graph generation task is largely restricted under a class imbalance. Previous methods have alleviated the class imbalance problem by incorporating commonsense information into the classification, enabling the prediction model to rectify the incorrect head class into the correct tail class. However, the results of commonsense-based models are typically overcorrected, e.g., the visually correct head class is forcibly modified into the wrong tail class. We argue that there are two principal reasons for this phenomenon. First, existing models ignore the semantic gap between commonsense knowledge and real scenes. Second, current commonsense fusion strategies propagate the neighbors in the visual-linguistic contexts without long-range correlation. To alleviate overcorrection, we formulate the commonsense-based scene graph generation task as two sub-problems: scene-induced commonsense graph generation (SI-CGG) and commonsense-inspired scene graph generation (CI-SGG). In SI-CGG module, unlike conventional methods using fixed commonsense graph, we adaptively adjust the node embeddings in a commonsense graph according to their visual appearance and configure the new reasoning edge under a specific visual context. The CI-SGG module is proposed to propagate the information from scene-induced commonsense graph back to the scene graph. It updates the representations of each node in scene graph by the aggregation of neighbourhood information at different scales. Through maximum likelihood optimisation of the logarithmic Gaussian process, the scene graph automatically adapt to the different neighbors in the visual-linguistic contexts. Systematic experiments on the Visual Genome dataset show that our full method achieves state-of-the-art performance. Lianggangxu Chen, Jiale Lu, Youqi Song, Changbo Wang, Gaoqi He |
ACM Multimedia | 5 |
| 2023 | Prior Knowledge-driven Dynamic Scene Graph Generation with Causal InferenceabstractThe task of dynamic scene graph generation (DSGG) aims at constructing a set of frame-level scene graphs for the given video. It suffers from two kinds of spurious correlation problems. First, the spurious correlation between input object pair and predicate label is caused by the biased predicate sample distribution in dataset. Second, the spurious correlation between contextual information and predicate label arises from interference caused by background content in both the current frame and adjacent frames of the video sequence. To alleviate spurious correlations, our work is formulated into two sub-tasks: video-specific commonsense graph generation (VsCG) and causal inference (CI). VsCG module aims to alleviate the first correlation by integrating prior knowledge into prediction. Information of all the frames in current video is used to enhance the commonsense graph constructed from co-occurrence patterns of all training samples. Thus, the commonsense graph has been augmented with video-specific temporal dependencies. Then, a CI strategy with both intervention and counterfactual is used. The intervention component further eliminates the first correlation by forcing the model to consider all possible predicate categories fairly, while the counterfactual component resolves the second correlation by removing the bad effect from context. Comprehensive experiments on the Action Genome dataset show that the proposed method achieves state-of-the-art performance. Jiale Lu, Lianggangxu Chen, Youqi Song, Shaohui Lin, Changbo Wang, Gaoqi He |
ACM Multimedia | 6 |
| 2023 | Dynamic leader role modeling for self-organizing crowd evacuation simulationabstractAbstract The role of the leader in the self‐organizing crowd has a significant influence on the crowd evacuation when an emergency occurs. However, current crowd evacuation models either ignore the leadership characteristics of pedestrians or predefine an invariable leader identity. This article clearly distinguishes leaders from followers in a crowd and evaluates the impact of the leader's role on crowd evacuation under dynamic situations. First, the leadership model of the pedestrian is built based on the classic OCEAN personality parameters. Social dominance tendency is measured for more credibility. Then, the dynamic leader role model (DLRM) is proposed to describe the relationship between leaders and followers, considering the interaction among pedestrians during the evacuation. Finally, a novel crowd simulation algorithm based on the above DLRM is presented using the extended social force model. Various experiments in several typical scenarios verified that our proposed simulation model has better realism. Changbo Wang, Gaoqi He |
Comput. Animat. Virtual Worlds | 3 |
| 2023 | ORCANet: Differentiable multi-parameter learning for crowd simulationabstractAbstract Realistic crowd simulation has always been an important research field in computer graphics. While both agent‐based motion models and data‐driven behavior models have made some progress, they are still suffering from either huge effort of multi‐parameter tuning or limited realistic motion. In this article, we propose a novel and differentiable multi‐parameter learning method for crowd simulation, which is called ORCANet. The main idea is to learn from real data and inverse evaluating the multi‐parameter for subsequent simulation. ORCANet uses classic optimal reciprocal collision avoidance (ORCA) as a basic motion model which is integrated into the deep learning framework. Addressing the feature of linear programming and non‐differentiable operation, a Gaussian kernel is added to approximate the role of neighbor distance in collision avoidance, which turns the original discrete operation into a fully differentiable forward simulation. Furthermore, we leverage ORCANet to optimize the multi‐parameter combination in synthetic and real‐world datasets. ORCANet is proved to rapidly converge to correct parameter values and regenerate the input synthetic sequence. Moreover, experiments on real‐world datasets by the metric of pedestrian trajectories verified that a more realistic crowd simulation has been generated through ORCANet. Chen Li 0035, Changbo Wang, Gaoqi He |
Comput. Animat. Virtual Worlds | 4 |
| 2023 | Video-based spatio-temporal scene graph generation with efficient self-supervision tasks
Lianggangxu Chen, Yiqing Cai, Changhong Lu, Changbo Wang, Gaoqi He |
Multim. Tools Appl. | 5 |
| 2023 | Global Representation Guided Adaptive Fusion Network for Stable Video Crowd CountingabstractModern crowd counting methods in natural scenes, even when video datasets are available, are mostly based on images. Because of background interference or occlusion in the scene, these methods can easily lead to mutations and instability in density prediction. There has been minimal research on how to exploit the inherent consistency among adjacent frames to achieve high estimation accuracy of video sequences. In this study, we explore the long-term global temporal consistency in the video sequence and propose a novel Global Representation Guided Adaptive Fusion Network (GRGAF) for video crowd counting. The primary aim is to establish a long-term temporal representation among consecutive frames to guide the density estimation of local frames, which can alleviate the prediction instability caused by background noise and occlusions in crowd scenes. Moreover, in order to further enforce the temporal consistency, we apply the generative adversarial learning scheme and design a global-local joint loss, which can make the estimated density maps more temporally coherent. Extensive experiments on four challenging video-based crowd counting datasets (FDST, DroneCrowd, MALL and UCSD) demonstrate that our method makes effective use of spatio-temporal information of video and outperforms the other state-of-the-art approach. Yiqing Cai, Zhenwei Ma, Changhong Lu, Changbo Wang, Gaoqi He |
IEEE Trans. Multim. | 5 |
| 2022 | DH-GCN: Saliency-Aware Complex Scene Graph Generation Using Dual-Hierarchy Graph Convolutional NetworkabstractIn reality, complex scene plagues numerous scene graph generation models because realistic scene contains myriad of objects and complicated relationships. Most current methods suffer poor performance when encountering complex scenes. We find that there are two principal reasons for this phenomenon. First, the construction of graph loses sight of the hierarchy of objects. Second, there exists redundant information in feature optimization. To facilitate this issue, this paper proposes an innovative dual-hierarchy graph convolutional network (DH-GCN), which is a conceptually elegant and efficient top-down approach. In specific, DH-GCN leverages salient object detector to hierarchize objects and give gist nodes more accurate representation. Moreover, the dual-hierarchy message propagation is designed to refine the representation hierarchically and eliminate redundant information. Systematic experiments on Visual Genome dataset show the superiority of our method over strong baseline methods. Jiale Lu, Lianggangxu Chen, Yiqing Cai, Haoyue Guan, Changhong Lu, Changbo Wang, Gaoqi He |
ICME | 7 |
| 2022 | Skill-Oriented Hierarchical Structure for Deep Knowledge TracingabstractKnowledge tracing (KT) which aims to trace stu-dents' knowledge state is an effective technique in intelligent tutoring systems. Although most KT models have exploited the question side information, plentiful hierarchical information between skills hasn't been well extracted for making more accurate predictions. In this paper, a novel model called Skill-oriented Hierarchical structure for Deep Knowledge Tracing (SHDKT) is proposed to discover the relations between questions, which are implicit in the hierarchical skill structure. SHDKT comprises three modules. First, The skill concurrency graph (SCG) is constructed by incorporating students' response infor-mation into the question-skill bipartite graph, which contains both sequence and co-occurrence relations between skills. Second, a hierarchical skill representation module (HSRM) is proposed to exploit the hierarchical information of skills based on the SCG. Finally, a question representation module (QRM) is presented by learning explicit and implicit interactions of question side infor-mation. Hence we can predict the student response accurately through question representation. Extensive experiments on the KT datasets validate the effectiveness of our model. Zhenyuan Yang, Shimeng Xu, Changbo Wang, Gaoqi He |
ICTAI | 4 |
| 2022 | Multi-modal chemical information reconstruction from images and texts for exploring the near-drug spaceabstractIdentification of new chemical compounds with desired structural diversity and biological properties plays an essential role in drug discovery, yet the construction of such a potential space with elements of 'near-drug' properties is still a challenging task. In this work, we proposed a multimodal chemical information reconstruction system to automatically process, extract and align heterogeneous information from the text descriptions and structural images of chemical patents. Our key innovation lies in a heterogeneous data generator that produces cross-modality training data in the form of text descriptions and Markush structure images, from which a two-branch model with image- and text-processing units can then learn to both recognize heterogeneous chemical entities and simultaneously capture their correspondence. In particular, we have collected chemical structures from ChEMBL database and chemical patents from the European Patent Office and the US Patent and Trademark Office using keywords 'A61P, compound, structure' in the years from 2010 to 2020, and generated heterogeneous chemical information datasets with 210K structural images and 7818 annotated text snippets. Based on the reconstructed results and substituent replacement rules, structural libraries of a huge number of near-drug compounds can be generated automatically. In quantitative evaluations, our model can correctly reconstruct 97% of the molecular images into structured format and achieve an F1-score around 97-98% in the recognition of chemical entities, which demonstrated the effectiveness of our model in automatic information extraction from chemical patents, and hopefully transforming them to a user-friendly, structured molecular database enriching the near-drug space to realize the intelligent retrieval technology of chemical knowledge. Jie Wang 0146, Zihao Shen, Yichen Liao, Shiliang Li, Gaoqi He, Man Lan, Xuhong Qian, Kai Zhang 0001, Honglin Li 0003 |
Briefings Bioinform. | 6 |
| 2022 | VRPharmer: bringing virtual reality into pharmacophore-based virtual screening with interactive exploration and realistic visualizationabstractSUMMARY: Current pharmacophore-based virtual screening (VS) software has limited interactive capabilities and less intuitive screening processes. In this study, a novel tool named VRPharmer is proposed to perform the entire VS workflow in VR environments. VRPharmer enables users to interactively perceive computation processes and immersively observe molecular structures. Besides a typical screening mode (OPT mode), VRPharmer provides a unique interactive screening mode (SCORE mode) for freely exploring the optimal binding poses. Pharmacophore models are editable to study the impact of each feature and further refine the screening results. Moreover, molecular rendering algorithms are improved for precise representations. AVAILABILITY AND IMPLEMENTATION: VRPharmer is open-source software under the MIT license. The released version is available at https://github.com/VRPharmer/VRPharmer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianchao Zhou, Ziyan Feng, Zilong Jin, Chenfei Zhang, Shiliang Li, Gaoqi He, Honglin Li 0003 |
Bioinform. | 9 |
| 2022 | Exploring Contextual Relationships in 3D Cloud Points by Semantic Knowledge MiningabstractAbstract 3D scene graph generation (SGG) aims to predict the class of objects and predicates simultaneously in one 3D point cloud scene with instance segmentation. Since the underlying semantic of 3D point clouds is spatial information, recent ideas of the 3D SGG task usually face difficulties in understanding global contextual semantic relationships and neglect the intrinsic 3D visual structures. To build the global scope of semantic relationships, we first propose two types of Semantic Clue (SC) from entity level and path level, respectively. SC can be extracted from the training set and modeled as the co‐occurrence probability between entities. Then a novel Semantic Clue aware Graph Convolution Network (SC‐GCN) is designed to explicitly model each SC of which the message is passed in their specific neighbor pattern. For constructing the interactions between the 3D visual and semantic modalities, a visual‐language transformer (VLT) module is proposed to jointly learn the correlation between 3D visual features and class label embeddings. Systematic experiments on the 3D semantic scene graph (3DSSG) dataset show that our full method achieves state‐of‐the‐art performance. Lianggangxu Chen, Jiale Lu, Yiqing Cai, Changbo Wang, Gaoqi He |
Comput. Graph. Forum | 5 |
| 2022 | Authoring multi-style terrain with global-to-local control
Jian Zhang 0070, Chen Li 0035, Peichi Zhou, Changbo Wang, Gaoqi He, Hong Qin 0001 |
Graph. Model. | 5 |
| 2022 | BAW: learning from class imbalance and noisy labels with batch adaptation weighted loss
Siyuan Pan, Bin Sheng 0001, Gaoqi He, Huating Li, Guangtao Xue |
Multim. Tools Appl. | 3 |
| 2022 | F2-Bubbles: Faithful Bubble Set Construction and Flexible EditingabstractIn this paper, we propose F2-Bubbles, a set overlay visualization technique that addresses overlapping artifacts and supports interactive editing with intelligent suggestions. The core of our method is a new, efficient set overlay construction algorithm that approximates the optimal set overlay by considering set elements and their non-set neighbors. Thanks to the efficiency of the algorithm, interactive editing is achieved, and with intelligent suggestions, users can easily and flexibly edit visualizations through direct manipulations with local adaptations. A quantitative comparison with state-of-the-art set visualization techniques and case studies demonstrate the effectiveness of our method and suggests that F2-Bubbles is a helpful technique for set visualization. Yunhai Wang, Da Cheng, Jian Zhang 0070, Liang Zhou 0001, Gaoqi He, Oliver Deussen |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Social-Scene-Aware Generative Adversarial Networks for Pedestrian Trajectory Prediction
Binhao Huang, Zhenwei Ma, Lianggangxu Chen, Gaoqi He |
CGI | 4 |
| 2021 | Leveraging Intra-Domain Knowledge to Strengthen Cross-Domain Crowd CountingabstractUnsupervised cross-domain counting research using synthetic datasets becomes imminent when considering the laborious labeling for supervised methods. However, the existing methods only focus on learning domain shared knowledge to narrow the gap between the source domain and target domain (inter-domain gap). Nevertheless, these methods do not consider the enormous distribution gap among the target domain data itself (intra-domain gap). In this paper, we propose a two-step domain adaptation method with multi-level feature response branches, which further uses the intra-domain knowledge to strengthen the target domain’s adaptability. Specifically, we first use different feature response branches to learn inter-domain knowledge more robustly, reducing the prediction inconsistency of different scenarios. Subsequently, the trained model is used to generate pseudo-labels for the target domain. The entire model was retrained by using pseudo-labels. Various experiments on synthetic dataset GCC and three real public datasets validate our proposed method’s availability with higher accuracy. Yiqing Cai, Lianggangxu Chen, Zhenwei Ma, Changhong Lu, Changbo Wang, Gaoqi He |
ICME | 6 |
| 2020 | Novel Sketch-Based 3D Model Retrieval via Cross-domain Feature Clustering and Matching
Jian Zhang 0070, Chen Li 0035, Changbo Wang, Gaoqi He, Hong Qin 0001 |
ICANN (1) | 5 |
| 2020 | An advanced hybrid smoothed particle hydrodynamics-fluid implicit particle method on adaptive grid for condensation simulationabstractAbstract In this article, we propose a novel hybrid framework by combining smoothed particle hydrodynamics and adaptive narrow band fluid implicit particle method (NB‐FLIP) to faithfully model the multiphysical processes involving heat transfer and phase transition, and to precisely simulate the dynamics of condensed droplets moving along intricate objects. We first formulate a governing physical model built upon an improved phase transition model and an augmented on‐surface drop analysis method to achieve realistic condensation effects over intricate hydrophilic/hydrophobic interface. To achieve both high‐fidelity interactions and high‐resolution visual effects, we further develop an adaptive NB‐FLIP solver with octree‐dictated background grid in order to further enhance the performance of our framework. Experimental results have shown that our approach can be used to efficiently and realistically simulate the small‐scale interaction details between condensed drops and complex objects with arbitrary geometry. Jiajun Shi, Chen Li 0035, Changbo Wang, Hong Qin 0001, Gaoqi He |
Comput. Animat. Virtual Worlds | 5 |
| 2020 | Video flickering removal using temporal reconstruction optimization
Bin Sheng 0001, Ping Li 0016, Gaoqi He |
Multim. Tools Appl. | 5 |
| 2019 | Dynamic Region Division for Adaptive Learning Pedestrian CountingabstractAccurate pedestrian counting algorithm is critical to eliminate insecurity in the congested public scenes. However, counting pedestrians in crowded scenes often suffer from severe perspective distortion. In this paper, basing on the straightline double region pedestrian counting method, we propose a dynamic region division algorithm to keep the completeness of counting objects. Utilizing the object bounding boxes obtained by YoloV3 and expectation division line of the scene, the boundary for nearby region and distant one is generated under the premise of retaining whole head. Ulteriorly, appropriate learning models are applied to count pedestrians in each obtained region. In the distant region, a novel inception dilated convolutional neural network is proposed to solve the problem of choosing dilation rate. In the nearby region, YoloV3 is used for detecting the pedestrian in multi-scale. Accordingly, the total number of pedestrians in each frame is obtained by fusing the result in nearby and distant regions. A typical subway pedestrian video dataset is chosen to conduct experiment in this paper. The result demonstrate that proposed algorithm is superior to existing machine learning based methods in general performance. Gaoqi He, Zhenwei Ma, Binhao Huang, Bin Sheng 0001, Yubo Yuan 0001 |
ICME | 1 |
| 2019 | ADMM-Based Decentralized Electric Vehicle Charging with Trip Duration LimitsabstractWith the large-scale deployment of Electric Vehicles (EVs), the unbalanced distribution of charging needs and random charging behaviors cause charging stations (CSs) congestion. This degrades EV drivers' quality of experience by extending charging waiting time and increasing charging fee. Thus, EV owners are facing a critical issue on how to decrease the cost of charging, which consists of two parts: charging duration and charging fee. A great deal of existing work is confined to finding CSs to optimize the two parts individually. However, it still remains unexplored how to jointly minimize charging duration and charging fee under an overall time limit (i.e., deadline) of a scheduled trip. The problem is the focus of this paper. First, we formulate this problem as a 0-1 Integer Linear Programming problem and show its NP-Hardness. Then, we propose an efficient distributed algorithm based on the Alternating Direction Method of Multipliers (ADMM). The algorithm decomposes the original problem into sub-problems that can be solved locally and in parallel between charging stations and the global coordinator. Finally, we carry out extensive simulations based on real-life transport network data, and the results show that the proposed approach brings significant cost savings over existing ones. Gaoqi He, Zhifu Chai, Xingjian Lu, Fanxin Kong, Bin Sheng 0001 |
RTSS | 1 |
| 2019 | Physical-barrier detection based collective motion analysis
Gaoqi He, Dongxu Jiang, Yubo Yuan 0001, Xingjian Lu |
Frontiers Comput. Sci. | 1 |
| 2017 | A double-region learning algorithm for counting the number of pedestrians in subway surveillance videos
Gaoqi He, Dongxu Jiang, Xingjian Lu, Yubo Yuan 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2016 | JTangCMS: An efficient monitoring system for cloud platforms
Xingjian Lu, Jianwei Yin, Naixue Xiong, Shuiguang Deng, Gaoqi He, Huiqun Yu |
Inf. Sci. | 5 |
| 2016 | Shadow obstacle model for realistic corner-turning behavior in crowd simulationabstractThis paper describes a novel model known as the shadow obstacle model to generate a realistic corner-turning behavior in crowd simulation. The motivation for this model comes from the observation that people tend to choose a safer route rather than a shorter one when turning a corner. To calculate a safer route, an optimization method is proposed to generate the corner-turning rule that maximizes the viewing range for the agents. By combining psychological and physical forces together, a full crowd simulation framework is established to provide a more realistic crowd simulation. We demonstrate that our model produces a more realistic corner-turning behavior by comparison with real data obtained from the experiments. Finally, we perform parameter analysis to show the believability of our model through a series of experiments. Gaoqi He, Zhen Liu 0002, Xingjian Lu |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2013 | A review of behavior mechanisms and crowd evacuation animation in emergency exercisesabstractEmergency exercises are an efficient approach for preventing serious damage and harm, including loss of life and property and a wide range of adverse social effects, during various public emergencies. Among various factors affecting the value of emergency exercises, including their design, development, conduct, evaluation, and improvement planning, this paper emphasizes the focal role of evacuees and their behavior. We address two concerns: What are the intrinsic reasons behind human behavior? How do we model and exhibit human behavior? We review studies investigating the mechanisms of psychological behavior and crowd evacuation animation. A comprehensive analysis of logical patterns of behavior and crowd evacuation is presented first. The interactive effects of information (objective and subjective), psychology (panic, small groups, and conflicting roles), and six kinds of behavior contribute to a more effective understanding of an emergency scene and assist in making scientific decisions. Based on these studies, a wide range of perspectives on crowd formation and evacuation animation models is summarized. Collision avoidance is underlined as a special topic. Finally, this paper highlights some of the technical challenges and key questions to be addressed by future developments in this rapidly developing field. Gaoqi He, Zhi-Hua Chen, Chunhua Gu |
J. Zhejiang Univ. Sci. C | 1 |
| 2012 | Virtual Network Marathon with immersion, scientificalness, competitiveness, adaptability and learning
Mingliang Xu 0001, Lizhen Han, Yong Liu 0007, Pei Lv, Gaoqi He |
Comput. Graph. | 6 |