EDBT 2026 Demo / reviewers in the wild / expert
Fang Gao 0001
dblp:60/3980-1
· DBLP profile ↗
31ranked-venue papers
12as first author
30since 2021 · last 2026
0000-0003-1816-5420ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data-driven model-free graph multi-agent deep reinforcement learning for voltage-voltage-ampere reactive control in active distribution networksabstractThe high penetration of distributed generation (DG) and renewable energy sources (RES) exacerbates voltage-voltage-ampere reactive (volt–VAR control, VVC) due to fast uncertainty-driven voltage fluctuations, topology-coupled interactions across feeder regions, and the limited real-time practicality of model-dependent optimization and iterative power-flow solvers. This paper formulates VVC as a Markov game and proposes an integrated, data-driven, topology-aware multi-agent learning architecture. From an artificial-intelligence perspective, we develop a centralized training with decentralized execution (CTDE) multi-agent graph soft actor–critic framework (MAGSAC), where graph convolution is embedded into Actor–Critic networks to enhance topology-aware coordination and stabilize multi-agent learning under partial local observations. To avoid repeatedly invoking iterative power-flow solvers during interaction, an offline-trained power flow neural network (PNN) is embedded into the environment to provide fast voltage estimation. It is trained using Latin hypercube sampling and a Levenberg–Marquardt optimizer, reducing reliance on explicit network parameters. A polygon voltage fortress reward is further introduced for fine-grained penalty shaping near operational limits. From an engineering-application perspective, the proposed architecture enables coordinated reactive-power regulation of inverter-interfaced DGs and compensators under high DG variability, improving voltage compliance and operational economy. Case studies on modified IEEE 33-bus and IEEE 69-bus systems show that MAGSAC achieves zero voltage violation on the 33-bus typical-day test and reduces network losses by 38.13 % versus no control. On the 69-bus system, it attains the lowest voltage-violation ratio among compared strategies and reduces daily loss by 27.18 %. Additional tests under topology changes and measurement-data loss support robustness and generalization under the studied settings. Fang Gao 0001, Linfei Yin, Dejian Huang, Jiongkai Qin, Shilin Gao |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | SAM-MPA: A SAM-Based Motion Perception and Aggregation Framework for Referring Video SegmentationabstractReferring video object segmentation relies on natural language descriptions to identify and segment target objects in videos, and has achieved substantial progress in recent years. However, most prior studies process videos in a frame-by-frame manner, failing to fully exploit temporal information. Recently, the large-scale segmentation model Segment Anything Model (SAM) has attracted considerable attention due to its strong segmentation capability and impressive zero-shot generalization. Nevertheless, SAM still exhibits limitations when handling complex action-oriented descriptions. Motivated by these observations, we propose a novel SAM-based Motion-Perception Aggregation framework for referring video object segmentation, termed SAM-MPA, which consists of four modules. DINO-SAM leverages the powerful segmentation ability of SAM to perform initial video segmentation guided by textual prompts, generating object-level masks. The Kalman Filtering Motion Modeling module injects explicit object motion modeling into DINO-SAM, improving segmentation robustness under occlusion and fast-motion scenarios. The motion-aware aggregation module effectively captures object action cues at multiple temporal scales, thereby enhancing global video understanding. The text-token matching module further enforces semantic consistency between the segmentation results and the referring expressions. Extensive experiments on challenging RVOS benchmarks demonstrate that SAM-MPA provides a competitive and efficient SAM-based solution for motion-centric referring video object segmentation, while offering a favorable trade-off between performance and computational cost compared with conventional non-MLLM baselines. The code is available at https://github.com/GXU-LIPE/SAM-MPA. Fang Gao 0001, Ao Lu, Qingbao Huang, Jun Yu 0001 |
IEEE Internet Things J. | 1 |
| 2026 | Monocular Camera-Based Substation Safety Distance Monitoring and Early Warning MethodabstractTo ensure the stable operation of substations, staff members are frequently required to enter operational areas. However, due to the lack of proactive monitoring mechanisms, accidental intrusions into live zones often lead to electric shock incidents. To address this issue, this study proposes an intelligent method that integrates keypoint detection, ranging, and early warning into a unified framework. First, a Channel-aware Inception Depthwise Convolution (CIDW) feature extraction module is proposed to enhance the capability of small-object feature extraction. Second, a Context-Dynamic Attention (CDA) feature enhancement module is developed to improve the model’s ability to capture fine-grained details in critical regions. Third, a Context-Enhanced Mechanism (CEM) module is designed to improve the feature pyramid network, thereby strengthening the model’s perception of high-resolution features. To overcome the challenges that monocular cameras cannot directly obtain threedimensional information and that long-distance pixels exhibit nonlinear errors, an innovative Long-distance Error Correction (LEC) model is proposed. By combining an inverse coordinate transformation formula, the LEC model enables high-precision 3D ranging. Experimental results demonstrate that the proposed detector outperforms existing algorithms in terms of detection accuracy, speed, and lightweight design. The ranging accuracy reaches the millimeter level at short distances, with long-distance errors controlled within approximately 5 cm, making it an effective safety monitoring approach for the perception layer of intelligent substation IoT systems. Hanbo Zheng, Jinheng Li, Fang Gao 0001 |
IEEE Internet Things J. | 5 |
| 2026 | Bidirectional cross-modal image-guided point cloud completion with multi-scale progressive refinement
Fang Gao 0001, Yan Jin 0012, Shaodong Li, Hanbo Zheng, Jun Yu 0001 |
Pattern Recognit. | 1 |
| 2026 | Guided Self-Attention: Find the Generalized Necessarily Distinct Vectors for Grain Size GradingabstractWith the development of steel materials, metallographic analysis has become increasingly important. Unfortunately, grain size analysis is a manual process that requires experts to evaluate metallographic photographs, which is unreliable and time-consuming. To resolve this problem, we propose a novel classification method based on deep learning, namely, GSNets, a family of hybrid models, which can effectively introduce guided self-attention for classifying grain size. Concretely, we build our models from three insights: 1) Introducing our novel guided self-attention module can assist the model in finding the generalized necessarily distinct vectors capable of retaining intricate relational connections and rich local feature information; 2) By improving the pixelwise linear independence of the feature map, the highly condensed semantic representation will be captured by the model; 3) Our novel triple-stream merging module can significantly improve the generalization capability and efficiency of the model. Experiments show that our GSNet yields a classification accuracy of 90.1%, surpassing the state-of-the-art Swin Transformer V2 by 1.9% on the steel grain size dataset, which comprises 3599 images with 14 grain size levels. Furthermore, we intuitively believe our approach is applicable to broader applications such as object detection and semantic segmentation. Fang Gao 0001, XueTao Li, Jiabao Wang 0004, Shengheng Ma, Jun Yu 0001 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2025 | Towards Robust Autonomous Driving: Conditional Multimodal Large Language Models for Fine-Grained PerceptionabstractMultimodal large language models (MLLMs) have shown remarkable performance across various visual understanding tasks. However, most existing MLLMs still lack image detail perception, limiting their effectiveness in tasks that require detailed visual information. In this paper, we introduce Percept-DriveLM, a novel MLLM designed to tackle the fine-grained perception challenges in autonomous driving tasks. At the core of our model is the Visual Fusion Module, which integrates several innovative components: a dynamic resolution mechanism that combines both high and low resolution features, and an RoI conditional mechanism to incorporate object/region-level features identified by offline detectors, further refining the model's fine-grained perception abilities. Trained in a two-stage process, our model demonstrates exceptional performance, outperforming existing MLLMs with comparable parameter sizes and excelling in both autonomous driving perception and general vision-language tasks. The effectiveness of our approach is validated through extensive empirical studies. Code will be available at https://github.com/DebuggerSunfz/PerceptDriveLM. Fengzhao Sun, Jun Yu 0001, Jiaming Hou, Xilong Lu, Heng Song, Fang Gao 0001 |
ICRA | 7 |
| 2025 | A Dual-Arm Shared Control Framework Integrating Sub-goals and Predicted Trajectories for Asymmetric TasksabstractIn robotic operation, asymmetric tasks requiring dual-arm cooperation are the highly challenging research direction. Autonomous operation generally has a low success rate or poor generalization because of its excessive dependence on the accuracy of sub-goals from asymmetric tasks. Although teleoperation can significantly improve the performances above, during operation, the operators are prone to neglect crucial intermediate sub-goals that are conducive to fine-grained dual-arm cooperation. Therefore, we propose a dual-arm shared control framework which firstly introduces the Sub-goal Generation module to sufficiently concentrate on the intermediate states, thus improving the ability of fine-grained dual-arm cooperation and reducing the adjustment quantity during robotic asymmetric task operation. Also, we integrate the Trajectory Prediction module that computes the future trajectory based on the historical movement information to enhance the robot motion smoothness. Finally, through dynamic combination of Sub-goal Generation module, Trajectory Prediction module and operator movement in the shared control framework, we effectively decrease the sensitivity to the accuracy of sub-goals, thus significantly improving the success rate. In simulation, we conduct the comparative experiments with autonomous operation and teleoperation on four common asymmetric tasks to validate the advantages of our shared control framework. The effect of each element in our framework is verified by ablation study. Certainly, our shared control framework can also be applied in real-world scenario. Zhixiong Wang, Shaodong Li, Fang Gao 0001 |
IROS | 4 |
| 2025 | Factorizing value function with hierarchical residual Q-network in multi-agent reinforcement learning
Fang Gao 0001, Yunxiang Cai, Shaodong Li, Linfei Yin |
Neurocomputing | 1 |
| 2025 | TVTracker: Target-Adaptive Text-Guided Visual Fusion for Multimodal RGB-T TrackingabstractCurrent multi-modal sensor trackers mainly rely on visual cues for target tracking. However, in challenging scenarios, the visual information acquired by multi-modal sensors has limited descriptive ability when the target state undergoes significant changes, which may lead to unsatisfactory tracker performance. In this work, we propose TVTracker, a two-stage tracking framework. It leverages semantic information of the target state to enhance visual cues, enabling effective RGB-T tracking. The first stage is the generation of target text descriptions. Utilizing the Bootstrapping Language-Image Pre-training (BLIP) model, we generate target textual descriptions that match the images in the dataset. In the second stage, text-guided visual fusion is performed for target tracking. Target textual and visual features are extracted separately using the text and visual branches. Then the target textual features are integrated with the visual features to localize the target position and predict the target bounding box. In the text branch, we design the Target Text Adaptation Enhancement (TTAE) module to mitigate the interference of low-quality target textual features on visual features. In the visual branch, we develop the multi-modal visual information prompters, which include the Multi-modal Visual Shared Information Prompter (MVSIP) and the Multi-modal Visual Shared and Complementary Information Prompter (MVSCIP), to facilitate learning of multi-modal shared or complementary visual prompts. Experiments on the LasHeR, RGBT210, RGBT234, and VTUAV datasets demonstrate the effectiveness of TVTracker. Fang Gao 0001, Yan Jin 0012, Jingfeng Tang, Hanbo Zheng, Shengheng Ma, Jun Yu 0001 |
IEEE Internet Things J. | 1 |
| 2025 | Fine-grained hierarchical dynamics for image harmonization
Peng He 0004, Jun Yu 0001, Liuxue Ju, Fang Gao 0001 |
Neural Networks | 4 |
| 2025 | Visual and Textual Commonsense-Enhanced Layout Learning for Vision-and-Language NavigationabstractIn the Vision-and-Language Navigation (VLN) task, an agent must comprehend natural language instructions and execute precise navigation in complex environments. While significant progress has been made in the VLN field, the limited availability of navigation data hinders existing methods from fully learning the commonsense relationships between rooms and landmarks, which are crucial for environmental understanding and successful navigation. To address this issue, this work proposes a Visual and Textual Commonsense-Enhanced Layout Learning Model (ViTeC). We leverage the open-world knowledge embedded in large models by utilizing ChatGPT and BLIP-2 to provide commonsense information about environments. Specifically, BLIP-2 analyzes the room type corresponding to each panoramic image, while ChatGPT infers and provides knowledge about the most common landmarks within each room type. Moreover, to compensate for the agent’s lack of commonsense at the visual level, we employ Stable Diffusion to generate commonsense-based visual images, enhancing the agent’s visual perception. To ensure the agent effectively learns commonsense about the environment, we designed a Text Commonsense Layout Learning Module and a Visual Commonsense Layout Learning Module. These modules help the agent acquire environmental commonsense from both linguistic and visual perspectives, enabling it to utilize commonsense information effectively during navigation, thereby improving its environmental understanding and reasoning capabilities. Experimental results demonstrate that ViTeC achieves strong performance on REVERIE, R2R, and SOON datasets, exhibiting good generalization ability in complex environments. This validates the effectiveness of ViTeC in enhancing the agent’s environmental understanding and navigation capabilities. Fang Gao 0001, Jingfeng Tang, Jiabao Wang 0004, Shaodong Li, Shengheng Ma, Jun Yu 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Contrastive Learning With Multiple Prototypes for Unsupervised Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptive semantic segmentation aims to transfer knowledge from the annotated source domain to the unlabeled target domain. Recently, self-training methods have gained substantial attention, which leverage high-confidence predictions in the target domain as pseudo labels for supervision. However, limited exploration of intra-class variations across domains, including significant visual differences within each category, has led to misalignment between feature distribution across domains. In this article, we present a unified non-parametric distance-based online clustering method to efficiently maintain multiple centroid-based prototypes within each category subspace instead of one prototype for each category subspace, which enables prototypes to possess the capacity for richer feature representation. Then, considering the variance across different dimensions of a feature representation, we then extend the prototypes from centroid-based ones to distribution-based ones. Specifically, each subspace is modeled using a Gaussian mixture model which includes several anisotropic Gaussian distributions, aimed at prioritizing discriminative dimensions and obtaining a finer measurement of the pixel-to-prototype similarity. Meanwhile, a category-aware feature space is achieved through pixel-to-prototype contrastive learning to ensure the compactness of pixel features in the same subcategory and drive the separation between pixel features of different subcategories. What's more, multi-resolution features are utilized to promote diversity and robustness among intra-class prototypes. Experiments validate the competitiveness of our two prototype-based methods against existing state-of-the-art methods, with a mIoU of 76.8% on GTA$\rightarrow$Cityscapes, 68.4% on Synthia$\rightarrow$Cityscapes, 54.5% on Cityscapes$\rightarrow$DarkZurich and 56.4% on Cityscapes$\rightarrow$ACDC. Notably, our method is able to seamlessly integrate with existing UDA methods. Jun Yu 0001, Guochen Xie, Quansheng Liu, Zhen Kan, Lei Wang 0203, Qiang Ling 0001, Fang Gao 0001 |
IEEE Trans. Multim. | 9 |
| 2024 | A Method for Visual Spatial Description Based on Large Language Model Fine-tuning
Jiabao Wang 0004, Fang Gao 0001, Jingfeng Tang, Shaodong Li, Hanbo Zheng, Shengheng Ma, Feng Shuang 0002, Jun Yu 0001 |
ACM Multimedia | 2 |
| 2024 | Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global InteractionsabstractThe continual advancements in Generative Artificial Intelligence have created substantial hurdles for accurate deepfake detection, leading to limitations of currently popular detection methods across content-driven video-level deepfake detection scenarios. In this paper, we present the solutions to the Video-Level Deepfake Detection task. Our empirical findings demonstrate that modeling correlations of audio-visual modalities is important for video-level deepfake detection. Therefore, we introduce the model denoted Audio-Visual Local-Global Neural Network (i.e., AV-LGNN) in which the core design is the proposed AV-LGI Module (Audio-Visual Local-Global Interaction Module). The AV-LGI Module is composed of three stages: Local Intra-Region Interaction, Global Inter-Region Interaction, and Local-Global Interaction, which can better capture detailed information at local-level and efficiently learn the fine-grained correlations of inter-modalities in video deepfake detection under lower computational overheads. We further propose an adaptive modality selection strategy to facilitate model learning. Besides, a variety of data augmentation techniques are incorporated for audio-visual branches to enhance the robustness of the AV-LGNN. The experimental results verify the effectiveness of our model. Jia Zhang 0016, Mohan Jing, Keda Lu, Jun Yu 0001, Wen Su 0004, Fang Gao 0001, Jianqing Sun, Jiaen Liang |
ACM Multimedia | 8 |
| 2024 | Pin-CasNet: Detecting pin status in transmission lines based on cascade network
Fang Gao 0001, Rongwei Zhang, Jingfeng Tang, Shaomin Liu, Jun Yu 0001, Chang Wen Chen, Hanbo Zheng |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Language-Guided Dual-Modal Local Correspondence for Single Object TrackingabstractThis paper focuses on the advancement of single-object tracking technologies in computer vision, which have broad applications including robotic vision, video surveillance, and sports video analysis. Current methods relying solely on the target's initial visual information encounter performance bottlenecks and limited applications, due to the scarcity of target semantics in appearance features and the continuous change in the target's appearance. To address these issues, we propose a novel approach, combining visual-language dual-modal single-object tracking, that leverages natural language descriptions to enrich the semantic information of the moving target. We introduce a dual-modal single-object tracking algorithm based on local correspondence modeling. The algorithm decomposes visual features into multiple local visual semantic features and pairs them with local language features extracted from natural language descriptions. In addition, we also propose a new global relocalization method that utilizes visual language bimodal information to perceive target disappearance and misalignment and adaptively reposition the target in the entire image. This improves the tracker's ability to adapt to changes in target appearance over long periods of time, enabling long-term single target tracking based on bimodal semantic and motion information. Experimental results show that our model outperforms state-of-the-art methods, which demonstrates the effectiveness and efficiency of our approach. Jun Yu 0001, Zhongpeng Cai, Lei Wang 0203, Fang Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Cross-Domain Transformer with Adaptive Thresholding for Domain Adaptive Semantic Segmentation
Quansheng Liu, Lei Wang 0203, Jun Yu 0001, Fang Gao 0001 |
ICANN (8) | 4 |
| 2023 | Adaptive Fine-Grained Region Matching for Image Harmonization
Liuxue Ju, Chengdao Pu, Fang Gao 0001, Jun Yu 0001 |
ICIG (3) | 3 |
| 2023 | RatiO R-CNN: An Efficient and Accurate Detection Method for Oriented Object Detection
Chengdao Pu, Liuxue Ju, Fang Gao 0001, Jun Yu 0001 |
ICIG (3) | 3 |
| 2023 | Prototypical Contrastive Learning for Domain Adaptive Semantic SegmentationabstractThe goal of domain adaptive semantic segmentation is to train a model using labeled source domain data and produce accurate dense predictions on the unlabeled target domain. Previous methods adopt self-training, where reliable target domain predictions are used as pseudo labels for training. However, intra-class variations across domains, such as the varying visual appearance in each category, have not been fully explored, leading to misalignment in feature distribution between the source and target domains. In this paper, we propose to optimize the feature space with representative prototypes shared across domains. Specifically, we first adopt the non-parametric clustering to model multiple prototypes for each category feature space. Then, category-discriminative feature space is obtained via pixel-to-prototype contrastive learning. Through extensive experiments, our proposed method demonstrates competitive performance on GTA5→Cityscapes and Synthia→Cityscapes benchmark. It is noteworthy that our method is compatible with the existing UDA methods. Quansheng Liu, Chengdao Pu, Fang Gao 0001, Jun Yu 0001 |
IJCNN | 3 |
| 2023 | Efficient Micro-Expression Spotting Based on Main Directional Mean Optical Flow FeatureabstractHuman facial expressions can convey a great deal of information in daily life. Spotting macro-expression (MaE) and micro-expression (ME) intervals from long video sequences is a difficult challenge. In this paper, we propose an efficient framework for the expression spotting task. This framework consists of three main modules: Face Cropping and Alignment Module (FCAM), optical flow Feature Extraction Module (FEM), and expression Proposal Generation Module (PGM). The noise of optical flow features is reduced by face cropping and alignment, and the Main Directional Mean Optical Flow Feature of the regions of interest is extracted as the feature for expression spotting. Finally, the expression intervals are spotted by our designed expression proposal generation module. Our approach achieves very good results on the SAMM Long Videos and CAS(ME)^2. To demonstrate the transferability of our method, we tested it on the MEGC2023 unseen dataset and finally achieved the third place, proving the effectiveness of our method. Jun Yu 0001, Zhongpeng Cai, Shenshen Du, Xiaxin Shen, Lei Wang 0203, Fang Gao 0001 |
ACM Multimedia | 6 |
| 2023 | Improving 6D Object Pose Estimation Based on Semantic SegmentationabstractThe performance of 6D pose estimation, which is important for scene understanding, can be improved by more accurate object segmentation. RGB-D data including depth maps can provide more accurate position information than RGB data for semantic segmentation. In this work, we propose a novel two-stage RGBD-based pose estimation network, which can provide more precise semantic segmentation and effective point cloud features. Firstly, we use a lightweight semantic segmentation head to process the RBG-D data to get the pixel-level clustered mask, and then use a multi-scale and attention-based backbone to extract the point cloud features for pose estimation. We analyze the performance of our network on the YCB-Video dataset and the results show that our method is comparable to current state-of-the-art methods after optimization. Fang Gao 0001, Qiujun Li, Qingyi Sun |
SMC | 1 |
| 2023 | Dual-scale point cloud completion network based on high-frequency feature fusionabstractFor many vision tasks and intelligent robotics applications, it is common that the scanned 3D point cloud is not complete, so inferring from the residual defect shape to the intact shape becomes an essential task. Previous 3D completion neural network models generally use voxel-based or point-based methods to learn and process 3D data. For the voxel-based models, the computational cost and memory increase exponentially with the improvement of input resolution, and fine-grained features cannot be guaranteed in the completed point cloud due to limited computational resources . Point-based models suffer from the lack of precision in feature acquisition and crude reconstruction of complicated structures, making it extremely hard to accomplish elaborated semantic shapes. Combining advantages of voxel-based and point-based feature extraction through the high-frequency feature fusion module, this paper proposes a dual-scale point cloud completion network called DSNet, which performs global feature analysis at the voxel scale, and local feature analysis at the point cloud scale. The fused features are then integrated into the decoding and generation process, so as to complete the point cloud completion task from coarse to fine. Experimental results, at both quantitative and qualitative perspectives, in several prevailing datasets demonstrate that our approach surpasses state-of-the-art point cloud completion networks and has a good generalization performance . Code is available at https://github.com/engqing/DSNet. Fang Gao 0001, Pengbo Shi, Yan Jin 0012, Jun Yu 0001, Shaodong Li |
Image Vis. Comput. | 1 |
| 2023 | A viable framework for semi-supervised learning on realistic dataset
Guochen Xie, Jun Yu 0001, Qiang Ling 0001, Fang Gao 0001 |
Mach. Learn. | 5 |
| 2023 | Multi-Object Tracking: Decoupling Features to Solve the Contradictory Dilemma of Feature RequirementsabstractMulti-object tracking achieves the acquisition of target location information and identity information through two subtasks, detection and re-identification (ReID). The existing commonly used one-shot framework has speed advantages, but the two subtasks have different feature requirements, which leads to competitive learning in the training and thus weakens the feature quality. We propose a feature decoupling based multi-object tracking framework FDTrack for contradictory feature requirements. Through the mutual inhibition of the two subtasks, the features of the backbone network are decoupled. Then the decoupled features are self-constrained to enhance effective features. Considering the instability of the target state and the different confidence of the detections, a more reasonable association strategy is employed to maximize the matchings between detections, thus recovering low-confidence targets. FDTrack is extensively tested on the MOT17 and MOT20 benchmarks. The experimental results show that FDTrack surpasses the previous state-of-the-art (SOTA) methods and has good anti-interference and real-time performance. Moreover, our proposed modules have good portability and can be applied in other one-shot trackers to achieve performance improvement. Yan Jin 0012, Fang Gao 0001, Jun Yu 0001, Jiabao Wang 0004, Feng Shuang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Micro Expression Generation with Thin-plate Spline Motion Model and Face ParsingabstractMicro-expression generation aims at transfering the expression from the driving videos to the source images, which can be viewed as a motion transfer task. Recently, several works have been proposed to tackle this problem and achieve great performance. However, due to the intrinsic complexity of the face motion and different attributes of face regions, the task still remains challenging. In this paper, we propose an end-to-end unsupervised motion transfer network to tackle this challenge. As the motion of the face is non-rigid, we adopt an effective and flexible thin-plate spline motion estimation method to estimate the optical flow of the face motion. What's more, we find that several faces with eyeglasses show weird deformation in motion transfering. Thus, we introduce face parsing method to pay specific attention to the eyeglasses regions to ensure the reasonability of the deformation. We conduct several experiments on the provided datasets of the ACM MM 2022 micro-expression grand challenge (MEGC2022) and compare our method with several other typical methods. In comparison, our method shows the best performance. We (Team: USTC-IAT-United) also compare our method with other competitors' in MEGC2022, and the expert evaluation results show that our method performs best, which verifies the effectiveness of our method. Our code is available at https://github.com/HowToNameMe/micro-expression Jun Yu 0001, Guochen Xie, Zhongpeng Cai, Peng He 0004, Fang Gao 0001, Qiang Ling 0001 |
ACM Multimedia | 5 |
| 2022 | Efficient 6D object pose estimation based on attentive multi-scale contextual informationabstractAbstract 6D pose estimation has been pervasively applied to various robotic applications, such as service robots, collaborative robots, and unmanned warehouses. However, accurate 6D pose estimation is still a challenge problem due to the complexity of application scenarios caused by illumination changes, occlusion and even truncation between objects, and additional refinement is required for accurate 6D object pose estimation in prior work. Aiming at the efficiency and accuracy of 6D object pose estimation in these complex scenes, this paper presents a novel end‐to‐end network, which effectively utilises the contextual information within a neighbourhood region of each pixel to estimate the 6D object pose from RGB‐D images. Specifically, our network first applies the attention mechanism to extract effective pixel‐wise dense multimodal features, which are then expanded to multi‐scale dense features by integrating pixel‐wise features at different scales for pose estimation. The proposed method is evaluated extensively on the LineMOD and YCB‐Video datasets, and the experimental results show that the proposed method is superior to several state‐of‐the‐art baselines in terms of average point distance and average closest point distance. Fang Gao 0001, Qingyi Sun, Shaodong Li, Yong Li 0028, Jun Yu 0001, Feng Shuang 0002 |
IET Comput. Vis. | 1 |
| 2022 | Dual feature fusion network: A dual feature fusion network for point cloud completionabstractAbstract Point cloud data in the real world is often affected by occlusion and light reflection, leading to incompleteness of the data. Large‐region missing point clouds will cause great deviations in downstream tasks. A dual feature fusion network (DFF‐Net) is proposed to improve the accuracy of the completion of a large missing region of the point cloud. First, a dual feature encoder is designed to extract and fuse the global and local features of the input point cloud. Subsequently, a decoder is used to directly generate a point cloud of missing region that retains local details. In order to make the generated point cloud more detailed, a loss function with multiple terms is employed to emphasise the distribution density and visual quality of the generated point cloud. A large number of experiments show that the authors’ DFF‐Net is better than the previous state‐of‐the‐art methods in the aspect of point cloud completion. Fang Gao 0001, Pengbo Shi, Jiabao Wang 0004, Yaoxiong Wang, Jun Yu 0001, Yong Li 0028, Feng Shuang 0002 |
IET Comput. Vis. | 1 |
| 2021 | Radar Object Detection Using Data Merging, Enhancement and FusionabstractCompared to visible images, radar images are generally considered to be an active and robust solution, even in adverse driving situations, for object detection. However, the accuracy of radar object detection (ROD) is always poor. Owing to taking full advantage of data merging, enhancement and fusion, this paper proposes an effective ROD system with only radar images as the input. First, an aggregation module is designed to merge the data from all chirps in the same frame. Then, various gaussian noises with different parameters are employed to increase data diversity and reduce over-fitting based on the analysis of training data. Moreover, due to the process of inference with default parameters is not accurate enough, some hyperparameters are changed to increase the accuracy performance. Finally, a combination strategy is adopted to benefit from multi-model fusion. ROD2021 Challenge is supported by ACM ICMR 2021, and our team (ustc-nelslip) ranked 2nd in the test stage of this challenge. Diverse evaluations also verify the superiority of the proposed system. Jun Yu 0001, Xinlong Hao, Xinjian Gao, Yuyu Liu, Peng Chang 0002, Fang Gao 0001, Feng Shuang 0002 |
ICMR | 8 |
| 2021 | Quantum deep reinforcement learning for rotor side converter control of double-fed induction generator-based wind turbines
Linfei Yin, Lichun Chen, Dongduan Liu, Fang Gao 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2020 | Attention Based Beauty Product Retrieval Using Global and Local DescriptorsabstractBeauty product retrieval has drawn more and more attention for its wide application outlook and enormous economic benefits. However, this task is always challenging due to the variation of products, especially the disturbance of clustered background. In this paper, we first introduce attention mechanism into a global image descriptor, i.e., Maximum Activation of Convolutions (MAC), and propose Attention-based MAC (AMAC). With this enhancement, we can suppress the negative effect of background and highlight the foreground in an unsupervised manner. Then, AMAC and local descriptors are ensembled to complementarily increase the performance. Furthermore, we try to finetune multiple retrieval methods on the different datasets and adopt a query expansion strategy to obtain more improvements. Extensive experiments conducted on a dataset containing more the half million beauty products (Perfect-500K) demonstrate the effectiveness of the proposed method. Finally, our team (USTC-NELSLIP) wins the first place on the leaderboard of the 'AI Meets Beauty'Grand Challenge of ACM Multimedia 2020. The code is available at: https://github.com/gniknoil/Perfect500K-Beauty-Product-Retrieval-Challenge. Jun Yu 0001, Guochen Xie, Haonian Xie, Xinlong Hao, Fang Gao 0001, Feng Shuang 0002 |
ACM Multimedia | 6 |