VLDB 2026 Research / reviewers in the wild / expert
Feng Shuang 0002
dblp:121/6183-2 · also Shuang Feng 0002
· DBLP profile ↗
19ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-4733-4732ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Anchor-Based Multimodal Verification: A Dynamic Query Framework for Fake News Forensics in Short VideosabstractThe proliferation of maliciously altered short videos on social media platforms poses a significant threat to information security ecosystems, eroding public trust in digital media. Despite recent advancements in detecting fake video news, significant challenges remain in the forensic analysis of short videos, leading to issues of bias. First, as technology rapidly advances, fake videos are becoming increasingly semantically convincing, undermining the effectiveness of current classification methods. Second, the heterogeneous nature of video modalities (visual, textual, audio) creates critical challenges for models to learn discriminative feature representations. To address these challenges, we propose a dynamic query framework for fake news forensics in short videos, termed the Semantic Guided Adaptive Network (SGAN). Our approach is motivated by the need to utilize superficial alignment to identify suspicious manipulations through anchor-based verification and to leverage the adaptive capability of learnable queries to learn the heterogeneous boundary in each modality. Specifically, SGAN comprises a verification module and a flexible query learning module. The verification module employs text as the anchor to verify detailed context, mining fine-grained information while emphasizing key features, thereby providing candidate manipulations for downstream modules. The query learning module leverages learnable queries to map heterogeneous forensic features and integrates them through multi-level fusion for decision-making. Extensive experiments conducted on two widely used datasets demonstrate the effectiveness and generalization of the proposed method. Pijian Li, Qingbao Huang, Feng Shuang 0002, Yi Cai 0001, Haonan Cheng, Qing Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Visual primitives as words: Alignment and interaction for compositional zero-shot learning
Feng Shuang 0002, Jiahuan Li, Qingbao Huang, Wenye Zhao, Dongsheng Xu 0001, Haonan Cheng |
Pattern Recognit. | 1 |
| 2024 | Energy Consumption Modelling of Coaxial-Rotor in Vortex Ring State for Controllable High-speed DescendingabstractThe ability to fast climb and descend is crucial for Unmanned Aerial Vehicle (UAV) applications in the mountains. The slower descent speed will affect the UAV’s working efficiency in reaching the rescue area. However, during the fast descent of the rotorcraft, a chaotic flow field rampages as the rotorcraft falls into its wake flow. This is known as the vortex ring. Therefore, the safe descent velocity of consumer UAVs is usually limited to approximately 3m/s. This limitation reduces the potential of UAVs to execute tasks in mountainous and plateau regions. To broaden the task capability constrained by the maximum descending speed, it is necessary to jointly analyze the flow field and the energy consumption during descending. Existing research mainly focused on how to avoid entering the vortex ring instead of offering sufficient power to fly with it. In this paper, in order to achieve an efficient rotorcraft for rescuing in mountainous and plateaus, we break through the maximum-descending-speed of a coaxial rotors UAV. Hence, a power consumption managing pipeline is proposed to extend the power tolerance of the UAV. Specifically, a theoretic model for the coaxial rotors is proposed to analyze the induced velocity and energy consumption during vertical descending. Then, the theoretic model is verified to be consistent with the Computational Fluid Dynamics (CFD) and wind tunnel experiment results. Finally, we optimized the tolerance of the power and dynamic system according to the theoretic model. With this pipeline, our real-time flight achieved 8m/s controlled vertical-descent-speed (CVDS), which is a leading result in both quadrotors and coaxial UAVs. Taoze Ban, Jiannan Zhao, Feng Shuang 0002 |
ICRA | 5 |
| 2024 | A Method for Visual Spatial Description Based on Large Language Model Fine-tuning
Jiabao Wang 0004, Fang Gao 0001, Jingfeng Tang, Shaodong Li, Hanbo Zheng, Shengheng Ma, Feng Shuang 0002, Jun Yu 0001 |
ACM Multimedia | 7 |
| 2024 | Hate Speech Detection for the Power Domain
Qingbao Huang, Zehua Deng, Shizhen Chen, Feng Shuang 0002 |
NLPCC (4) | 5 |
| 2024 | 3D hand reconstruction via aggregating intra and inter graphs guided by prior knowledge for hand-object interaction scenario
Feng Shuang 0002, Shaodong Li |
J. Vis. Commun. Image Represent. | 1 |
| 2024 | Foodnet: multi-scale and label dependency learning-based multi-task network for food and ingredient recognition
Feng Shuang 0002, Zhouxian Lu, Yong Li 0028, Xia Gu, Shidi Wei |
Neural Comput. Appl. | 1 |
| 2024 | Multi-Granularity Feature Fusion for Image-Guided Story Ending GenerationabstractImage-guided Story Ending Generation aims at generating a reasonable and logical ending given a story context and an ending-related image. The existing models have achieved some success by fusing global image features with story context through an attention mechanism. However, they ignore the logical relationship between the story context and the image regions, and have not considered the high-level semantic features of the image such as visual sentiment. This may cause the generated ending inconsistent with the logic or sentiment of the given information. In this paper, we propose aMulti-Granularity featureFusion (MGF) model to solve this problem. Concretely, we first employ an image sentiment extractor to grasp the sentiment features of the image as part of the global image features. We then design a scene subgraph selector to capture the image features of the key region by picking the scene subgraph most relevant to the context. Finally, we fuse the textual and visual features from object level, region level, and global level, respectively. Our model is thereby capable of effectively capturing the key region features and visual sentiment of the image, so as to generate a more logical and sentimental ending. Experimental results show that our MGF model outperforms the state-of-the-art models on most metrics. Pijian Li, Qingbao Huang, Yi Cai 0001, Feng Shuang 0002, Qing Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | DINS: A Diverse Insulator Dataset for Object Detection and Instance SegmentationabstractIntelligent defect detection of insulators is faster, more accurate, standardized, and cheaper than manual detection with necessary massive inspection work. Insulator datasets are important for training detection models. Nevertheless, public datasets are scarce and lack variety, which hampers improving detection accuracy and achieving industrial-grade accuracy. We construct a comprehensive beyond the current insulator dataset—the diverse insulator dataset (DINS). DINS contains over 10 000 insulator images involving three insulator types (porcelain, glass, and composite) and defects. We annotate over 25 000 bounding boxes for object detection and 9000 masks, for instance, segmentation. DINS has much more scale and diversity than the current insulator datasets. Eventually, we discussed the effective augmentation methods for DINS and conducted some experiments demonstrating the usefulness of DINS with 98.3% mean average precision (mAP) of object detection and 97.2% mAP of instance segmentation. The datasets are available on GitHub. Benben Cui, Mingyuan Yang, Feng Shuang 0002 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Region-Focused Network for Dense CaptioningabstractDense captioning is a very critical but under-explored task, which aims to densely detect localized regions-of-interest (RoIs) and describe them with natural language in a given image. Although recent studies tried to fuse multi-scale features from different visual instances to generate more accurate descriptions, their methods still suffer from the lack of exploration of relation semantic information in images, leading to less informative descriptions. Furthermore, indiscriminately fusing all visual instance features will introduce redundant information, resulting in poor matching between descriptions and corresponding regions. In this work, we propose a Region-Focused Network (RFN) to address these issues. Specifically, to fully comprehend the images, we first extract the object-level features, and encode the interaction and position relations between objects to enhance the object representations. Then, to decrease the interference from redundant information about the target region, we extract the most relevant information to the region. Finally, a region-based Transformer is employed to compose and align the previous mined information and generate the corresponding descriptions. Extensive experiments on Visual Genome V1.0 and V1.2 datasets show that our RFN model outperforms the state-of-the-art methods, thus verifying its effectiveness. Our code is available at https://github.com/VILAN-Lab/DesCap . Qingbao Huang, Pijian Li, Youji Huang, Feng Shuang 0002, Yi Cai 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | ASCS-Reinforcement Learning: A Cascaded Framework for Accurate 3D Hand Pose Estimationabstract3D hand pose estimation can be achieved by cascading a feature extraction module and a feature exploitation module, where reinforcement learning (RL) is proved to be an effective way to perform feature exploitation. This paper points out the prospects of improving accuracy using better exploitation strategy, and proposes an Adaptive Step-Critic Shared RL (ASCS-RL) strategy for accurate feature exploitation in 3D hand pose estimation. Hand joint features are exploited in a multi-task manner, and divided into two groups according to the distributions of estimation error. An RL-based adaptive-step (AS-RL) strategy is then used to obtain the optimal step size for better exploitation. The exploitation process are finally performed using a critic-shared RL (CS-RL) strategy, where both groups share a universal critic mechanism. Ablation studies and extensive experiments are carried out to evaluate the performance of ASCS-RL on ICVL and NYU datasets. The results show the strategy achieves the state-of-the-art accuracy in monocular depth-based 3D hand pose estimation, especially the best on ICVL. Experiments also validates that ASCS-RL realizes better tradeoff between accuracy and running rapidity. Mingqi Chen, Feng Shuang 0002, Shaodong Li |
ICMR | 2 |
| 2023 | Cascading CNNs with S-DQN: A Parameter-Parsimonious Strategy for 3D Hand Pose Estimation
Mingqi Chen, Shaodong Li, Feng Shuang 0002 |
MMM (1) | 3 |
| 2023 | Lightweight 3D hand pose estimation by cascading CNNs with reinforcement learning
Mingqi Chen, Shaodong Li, Feng Shuang 0002 |
Pattern Recognit. Lett. | 3 |
| 2023 | Multi-Object Tracking: Decoupling Features to Solve the Contradictory Dilemma of Feature RequirementsabstractMulti-object tracking achieves the acquisition of target location information and identity information through two subtasks, detection and re-identification (ReID). The existing commonly used one-shot framework has speed advantages, but the two subtasks have different feature requirements, which leads to competitive learning in the training and thus weakens the feature quality. We propose a feature decoupling based multi-object tracking framework FDTrack for contradictory feature requirements. Through the mutual inhibition of the two subtasks, the features of the backbone network are decoupled. Then the decoupled features are self-constrained to enhance effective features. Considering the instability of the target state and the different confidence of the detections, a more reasonable association strategy is employed to maximize the matchings between detections, thus recovering low-confidence targets. FDTrack is extensively tested on the MOT17 and MOT20 benchmarks. The experimental results show that FDTrack surpasses the previous state-of-the-art (SOTA) methods and has good anti-interference and real-time performance. Moreover, our proposed modules have good portability and can be applied in other one-shot trackers to achieve performance improvement. Yan Jin 0012, Fang Gao 0001, Jun Yu 0001, Jiabao Wang 0004, Feng Shuang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Efficient 6D object pose estimation based on attentive multi-scale contextual informationabstractAbstract 6D pose estimation has been pervasively applied to various robotic applications, such as service robots, collaborative robots, and unmanned warehouses. However, accurate 6D pose estimation is still a challenge problem due to the complexity of application scenarios caused by illumination changes, occlusion and even truncation between objects, and additional refinement is required for accurate 6D object pose estimation in prior work. Aiming at the efficiency and accuracy of 6D object pose estimation in these complex scenes, this paper presents a novel end‐to‐end network, which effectively utilises the contextual information within a neighbourhood region of each pixel to estimate the 6D object pose from RGB‐D images. Specifically, our network first applies the attention mechanism to extract effective pixel‐wise dense multimodal features, which are then expanded to multi‐scale dense features by integrating pixel‐wise features at different scales for pose estimation. The proposed method is evaluated extensively on the LineMOD and YCB‐Video datasets, and the experimental results show that the proposed method is superior to several state‐of‐the‐art baselines in terms of average point distance and average closest point distance. Fang Gao 0001, Qingyi Sun, Shaodong Li, Yong Li 0028, Jun Yu 0001, Feng Shuang 0002 |
IET Comput. Vis. | 7 |
| 2022 | Dual feature fusion network: A dual feature fusion network for point cloud completionabstractAbstract Point cloud data in the real world is often affected by occlusion and light reflection, leading to incompleteness of the data. Large‐region missing point clouds will cause great deviations in downstream tasks. A dual feature fusion network (DFF‐Net) is proposed to improve the accuracy of the completion of a large missing region of the point cloud. First, a dual feature encoder is designed to extract and fuse the global and local features of the input point cloud. Subsequently, a decoder is used to directly generate a point cloud of missing region that retains local details. In order to make the generated point cloud more detailed, a loss function with multiple terms is employed to emphasise the distribution density and visual quality of the generated point cloud. A large number of experiments show that the authors’ DFF‐Net is better than the previous state‐of‐the‐art methods in the aspect of point cloud completion. Fang Gao 0001, Pengbo Shi, Jiabao Wang 0004, Yaoxiong Wang, Jun Yu 0001, Yong Li 0028, Feng Shuang 0002 |
IET Comput. Vis. | 8 |
| 2022 | DenseKPNET: Dense Kernel Point Convolutional Neural Networks for Point Cloud Semantic SegmentationabstractIn recent years, point clouds have been widely used in powerline inspection, smart cities, autonomous driving, and other fields. The deep learning-based point cloud processing methods have attracted more and more attention due to the developments of laser scanning technology and machine learning. However, the recent methods largely ignore global contextual relationships and do not make full use of the complementation between local feature and high-level geometric information. To the problem, we propose a novel deep neural network, namely, the Dense connection-based Kernel Point Network (DenseKPNet), which can greatly expand the receptive field of kernel point convolution to extract rich semantic context information and valuable geometric features from the local region effectively. Specifically, we first design a multiscale convolution kernel point module to extract initial geometric features from coarse to fine. Then, we design the dense connection module to efficiently learn more expressive local geometric features while capturing rich contextual information. In addition, we propose the kernel point convolution attention module (KPCAM), which can capture global interdependencies between points and strengthen the discriminativeness of effective features. We evaluate our method on public indoor and outdoor datasets. The qualitative and quantitative experimental results show the effectiveness of DenseKPNet. The mIoU of the proposed method on S3DIS and semantic3D datasets can reach 68.9% and 77.9%, respectively. Yong Li 0028, Xu Li 0025, Zhenxin Zhang, Feng Shuang 0002, Jincheng Jiang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Radar Object Detection Using Data Merging, Enhancement and FusionabstractCompared to visible images, radar images are generally considered to be an active and robust solution, even in adverse driving situations, for object detection. However, the accuracy of radar object detection (ROD) is always poor. Owing to taking full advantage of data merging, enhancement and fusion, this paper proposes an effective ROD system with only radar images as the input. First, an aggregation module is designed to merge the data from all chirps in the same frame. Then, various gaussian noises with different parameters are employed to increase data diversity and reduce over-fitting based on the analysis of training data. Moreover, due to the process of inference with default parameters is not accurate enough, some hyperparameters are changed to increase the accuracy performance. Finally, a combination strategy is adopted to benefit from multi-model fusion. ROD2021 Challenge is supported by ACM ICMR 2021, and our team (ustc-nelslip) ranked 2nd in the test stage of this challenge. Diverse evaluations also verify the superiority of the proposed system. Jun Yu 0001, Xinlong Hao, Xinjian Gao, Yuyu Liu, Peng Chang 0002, Fang Gao 0001, Feng Shuang 0002 |
ICMR | 9 |
| 2020 | Attention Based Beauty Product Retrieval Using Global and Local DescriptorsabstractBeauty product retrieval has drawn more and more attention for its wide application outlook and enormous economic benefits. However, this task is always challenging due to the variation of products, especially the disturbance of clustered background. In this paper, we first introduce attention mechanism into a global image descriptor, i.e., Maximum Activation of Convolutions (MAC), and propose Attention-based MAC (AMAC). With this enhancement, we can suppress the negative effect of background and highlight the foreground in an unsupervised manner. Then, AMAC and local descriptors are ensembled to complementarily increase the performance. Furthermore, we try to finetune multiple retrieval methods on the different datasets and adopt a query expansion strategy to obtain more improvements. Extensive experiments conducted on a dataset containing more the half million beauty products (Perfect-500K) demonstrate the effectiveness of the proposed method. Finally, our team (USTC-NELSLIP) wins the first place on the leaderboard of the 'AI Meets Beauty'Grand Challenge of ACM Multimedia 2020. The code is available at: https://github.com/gniknoil/Perfect500K-Beauty-Product-Retrieval-Challenge. Jun Yu 0001, Guochen Xie, Haonian Xie, Xinlong Hao, Fang Gao 0001, Feng Shuang 0002 |
ACM Multimedia | 7 |