EDBT 2026 Demo / reviewers in the wild / expert
Wenfeng Song
dblp:47/8351
· DBLP profile ↗
61ranked-venue papers
23as first author
48since 2021 · last 2026
0000-0002-5101-1071ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 18 first-author · 30 since 2021Artificial intelligence and machine learning · 24 · 9 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DECON: Reconstruction of Clothed-Geometric Multiple Humans from a Single Image via Geometry-Guided Decouplingabstract3D multi-human reconstruction from single images holds significant potential for advancing AR/VR applications. While remarkable progress has been made in single-human reconstruction, existing methods face challenges when reconstructing multiple humans. These challenges include: (1) severe inter-occlusion that disrupts individual body structures, and (2) the absence of physically plausible relative positioning among subjects. We present DECON, a novel DEcouple-and-reCONstruct framework that systematically addresses these limitations through two technical innovations: (1) a decouple-and-reconstruct framework with multi-view synthesis. It separates individuals and reconstructs detailed 3D bodies from a single image. (2) a Perspective-Aware Position Optimization (PAPO) approach. It ensures realistic positioning by fixing overlaps and gaps between subjects. Extensive experiments demonstrate our method's capability to reconstruct fully separated, anatomically complete 3D humans with clothed-geometric details and plausible interactions. Quantitative evaluations show a 54% reduction in Chamfer Distance and 35% in Point-to-Surface Distance compared to state-of-the-art methods. Yiming Jiang 0018, Wenfeng Song, Shuai Li 0001, Aimin Hao |
AAAI | 2 |
| 2026 | IntentMotion: Learning Intent-Aware Human Motion from Language in 3D ScenesabstractGenerating human motion in complex 3D scenes from text is a challenging task with broad applications. However, existing methods often overlook realistic physical contact, resulting in visually plausible but physically unrealistic motion, e.g., penetration. To alleviate this, we propose IntentMotion, a novel framework that generates human motion in 3D scenes from natural language instructions by explicitly modeling intent. We first introduce the Intention-Guided Contact Field (IGCF). This differentiable voxel-based contact region representation explicitly aligns parsed language roles with spatial contact regions through a hierarchical attention mechanism. IGCF is jointly trained with a diffusion-based motion generator, allowing contact predictions to adapt dynamically through gradient feedback. To improve the controllability and physics-aware motion, we further propose an Intention-Aware Diffusion Model (IADM), which decouples the high-level semantic planning from the low-level contact refinement in a coarse-to-fine process. The optimized contact cues are utilized to guide the synthesis of a coarse trajectory, followed by refining detailed pose sequences under IGCF supervision. Experiments on the HUMANISE and LINGO datasets demonstrate that our IntentMotion outperforms recent baselines in contact accuracy, semantic alignment, and generalization to unseen scenes. Wenfeng Song, Shi Zheng, Xingliang Jin, Aimin Hao, Fei Hou 0001, Xia Hou, Shuai Li 0001 |
AAAI | 1 |
| 2026 | From spectrum to prototype: Full-scale aggregated perception network for infrared small target detection
Duanyang Zhong, Meiwen Zhang, Yasheng Li, Wenfeng Song, Jisheng Xie, Senhao Yang |
Pattern Recognit. | 5 |
| 2026 | Channel-Wise Contribution Assessment for RGB-D Salient Object DetectionabstractIn RGB-D salient object detection (SOD), a common approach to improving accuracy is by using a dual-stream architecture to combine RGB and depth data. However, the effectiveness of depth information varies depending on the scene. In scenarios where depth maps provide a limited contribution, integrating them with RGB can be challenging and sometimes even detrimental to performance rather than enhancing it. Conventional RGB-D SOD methods often lack precision in assessing depth map quality, neglecting to account for the distinct contributions of various regions within the map, often resulting in a suboptimal fusion of regions where depth information is minimally beneficial or irrelevant. To address these issues, this article presents a novel channel-wise contribution assessment method that precisely evaluates the contributions of both the RGB and depth channels. By employing a controlled perturbation process to challenge the saliency detection model with specific, manageable disturbances, we are able to measure how much RGB and depth information each contributes to the final saliency map. Based on this analysis, we have developed a novel routing-style fusion of modality that dynamically adjusts the integration of the two modalities. This approach significantly lessens the negative impact of regions where depth data have a low, no, or even detrimental contribution, leading to a more effective and balanced fusion of RGB and depth information. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method consistently achieves competitive performance and improves the robustness of RGB-D salient object detection across diverse and challenging scenarios. Chenglizhao Chen, Mengke Song, Xinyu Liu 0029, Wenfeng Song |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2026 | HFHuman: High-Fidelity Human Reconstruction From Single Image With Multi-Modality FusionabstractAccurately reconstructing high-fidelity human models from single images is critical for virtual reality applications. Existing methods often rely on 3D features from the estimated parametric human model to provide geometric priors. This approach addresses challenges such as missing limbs or deformations, which often arise due to viewpoint limitations and self-occlusion. However, accurately predicting 3D features from monocular images remains a significant challenge. This limitation poses difficulties for the fusion of 2D and 3D information. In this paper, we introduce HFHuman, a novel approach for high-fidelity human reconstruction from a single image using multi-modality fusion. HFHuman effectively fuses multiple modalities, including geometric and depth, directly from images. Our method introduces three key innovations: (1) a depth and geometric parallel reconstruction framework that simultaneously handles whole-body geometry and detailed depth reconstruction, refining a parameterized 3D human model under progressive depth guidance; (2) a pixel-voxel feature fusion strategy that combines pixel-aligned features with voxel-aligned features using a multi-modality adaptor; and (3) a depth-refined technique that integrates RGB imagery with surface normals and depth mapping. By addressing the challenge of blending 2D and 3D modalities, HFHuman results in more accurate and realistic human reconstructions. Experimental results demonstrate that HFHuman outperforms state-of-the-art methods, setting a new standard for realistic 3D human body reconstruction. Yiming Jiang 0018, Wenfeng Song, Shuai Li 0001, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2026 | FCMD: Fine-Grained Text-Driven Cohesive Motion Generation With Diffusion ModelabstractGenerating continuous and expressive human motion from textual descriptions is a critical challenge in applications such as gaming and filmmaking. Existing methods often struggle to maintain global coherence, realistic frame continuity, and smooth transitions. To address these limitations, we propose FCMD, a novel diffusion-based model for generating cohesive motion sequences from fine-grained textual descriptions. FCMD introduces three key innovations: (1) Fine-grained Text Fusion, which integrates detailed textual cues with transitional narratives to enhance semantic consistency; (2) History Motion Guidance, ensuring motion accuracy and consistency across consecutive frames; and (3) Smooth Stitching Sampling, which leverages preceding and current motion information to achieve seamless transitions. Additionally, FCMD employs a large language model (LLM) to refine motion datasets by extracting fine-grained textual descriptions. Extensive experiments demonstrate that FCMD outperforms state-of-the-art methods in generating coherent, natural, and highly controllable motion sequences. Shuai Li 0001, Wenfeng Song, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | DynAvatar: Dynamic 3D Head Avatar Deformation With Expression Guided Gaussian SplattingabstractGenerating high-fidelity, expressive, and realistic 3D head avatars remains a fundamental challenge for immersive applications such as virtual reality, gaming, and telepresence. This task requires not only precise modeling of non-rigid facial deformations but also semantically controllable expression synthesis under diverse viewpoints and motion contexts. We present DynAvatar, a novel framework that integrates expression-guided deformation into the 3D Gaussian splatting pipeline to produce photorealistic and emotionally resonant head avatars. Our method introduces two key innovations: (1) an expression-guided Gaussian deformation module that tightly couples geometric displacement with high-level semantic cues, enabling fine-grained and anatomically meaningful facial animation; and (2) a spatial context embedding mechanism that encodes the canonical position of each Gaussian to preserve semantic coherence and spatial consistency during expression generation. Extensive experiments on both controlled and in-the-wild datasets demonstrate that DynAvatar significantly outperforms state-of-the-art methods in terms of visual realism, expression fidelity, and rendering quality. Wenfeng Song, Zhongyong Ye, Shuai Li 0001, Xia Hou, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | MaskScene: Hierarchical Conditional Masked Models for Real-Time 3D Indoor Scene SynthesisabstractIndoor scene synthesis is essential for creative industries. Driven by this demand, recent advances in scene synthesis using diffusion and autoregressive models have shown promising results. However, existing models struggle to achieve real-time performance, high visual fidelity, and flexible scene editing simultaneously. To tackle this challenge, we propose MaskScene, a novel hierarchical conditional masked model for real-time 3D indoor scene synthesis and editing. Specifically, MaskScene introduces a hierarchical scene representation that explicitly encodes scene relationships, semantics, and tokens. Based on this representation, we design a hierarchical conditional masked modeling architecture that enables parallel and iterative decoding, conditioned on both semantics and relationships. MaskScene leverages local object masking and hierarchical scene structures to infer occluded regions from partial observations, enabling rapid construction of 3D indoor environments that accurately reflect real-world scenes. Compared to state-of-the-art methods, MaskScene achieves 80× faster generation speed and improves scene quality by 10%, while also supporting zero-shot editing, such as scene completion and rearrangement, without extra fine-tuning. Qichuan Geng, Zhong Zhou, Wenfeng Song |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | CtrlAvatar: Controllable Avatars Generation via Disentangled Invertible NetworksabstractAs virtual experiences grow in popularity, the demand for realistic, personalized, and animatable human avatars increases. Traditional methods, relying on fixed templates, often produce costly avatars that lack expressiveness and realism. To overcome these challenges, we introduce Controllable Avatars generation via disentangled invertible networks (CtrlAvatar), a real-time framework for generating lifelike and customizable avatars. CtrlAvatar uses disentangled invertible networks to separate the deformation process into implicit body geometry and explicit texture components. This approach eliminates the need for repeated occupancy reconstruction, enabling detailed and coherent animations. The body geometry component ensures anatomical accuracy, while the texture component allows for complex, artifact-free clothing customization. This architecture ensures smooth integration between body movements and surface details. By optimizing transformations with position-varying offsets from the avatar’s initial Linear Blend Skinning vertices, CtrlAvatar achieves flexible, natural deformations that adapt to various scenarios. Extensive experiments show that CtrlAvatar outperforms other methods in quality, diversity, controllability, and cost-efficiency, marking a significant advancement in avatar generation. Wenfeng Song, Fei Hou 0001, Shuai Li 0001, Aimin Hao, Xia Hou |
AAAI | 1 |
| 2025 | ViMoGen: A Novel Motion Generator for Virtual Standard Patient
Xuehan Wang, Wenfeng Song, Shuai Li 0001, Xian'e Wang, Xia Hou |
ICXR | 2 |
| 2025 | Saliency-Free and Aesthetic-Aware Panoramic Video NavigationabstractMost of the existing panoramic video navigation approaches are saliency-driven, whereby off-the-shelf saliency detection tools are directly employed to aid the navigation approaches in localizing video content that should be incorporated into the navigation path. In view of the dilemma faced by our research community, we rethink if the "saliency clues" are really appropriate to serve the panoramic video navigation task. According to our in-depth investigation, we argue that using "saliency clues" cannot generate a satisfying navigation path, failing to well represent the given panoramic video, and the views in the navigation path are also low aesthetics. In this paper, we present a brand-new navigation paradigm. Although our model is still trained on eye-fixations, our methodology can additionally enable the trained model to perceive the "meaningful" degree of the given panoramic video content. Outwardly, the proposed new approach is saliency-free, but inwardly, it is developed from saliency but biasing more to be "meaningful-driven"; thus, it can generate a navigation path with more appropriate content coverage. Besides, this paper is the first attempt to devise an unsupervised learning scheme to ensure all localized meaningful views in the navigation path have high aesthetics. Thus, the navigation path generated by our approach can also bring users an enjoyable watching experience. As a new topic in its infancy, we have devised a series of quantitative evaluation schemes, including objective verifications and subjective user studies. All these innovative attempts would have great potential to inspire and promote this research field in the near future. Chenglizhao Chen, Guangxiao Ma, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | AttriDiffuser: Adversarially enhanced diffusion model for text-to-facial attribute image synthesis
Wenfeng Song, Zhongyong Ye, Xia Hou, Shuai Li 0001, Aimin Hao |
Pattern Recognit. | 1 |
| 2025 | UNI-IQA: A Unified Approach for Mutual Promotion of Natural and Screen Content Image Quality AssessmentabstractTo date, the image quality assessment (IQA) research field has mainly focused on natural images (NIs)-based IQA and screen content images (SCIs)-based IQA. Usually, these two research branches are quite independent due to the large differences between NIs and SCIs, where NIs, captured by cameras directly, contain pictorial information solely, yet, SCIs, synthesized or GPU-rendered, have pictures and textures. Moreover, the distortion types are also different, and subjective scores of different datasets assigned by participants are usually not well aligned. So, due to the above-mentioned “domain shifts” and “dataset misalignments”, our research community has widely believed that it could be very difficult to achieve joint mutual promotions between NIs- and SCIs-based IQA. In this paper, we argue that despite the “differences”, there still are some “common characteristics” — our human visual system perceives the “pictures” in both SCIs and NIs almost the same way. Thus, we can still achieve mutual performance promotion if we can appropriately use the “common characteristics” between SCIs and NIs. Our key idea is to devise a “content-aware” data switch, which, from the perspective of input’s contents (i.e., pictures or textures), aims at letting the model automatically enhance the commonness and compress the discrepancies between the two tasks. Notice that none of the existing fusion schemes can reach this goal since they are actually content-unaware, degenerating the “mutual interactions” into “mutual interferences”. This paper is the first attempt to achieve full end-to-end “mutual interactions” between NIs- and SCIs-based IQA. Using the proposed switch, we are also the first to achieve solid mutual promotions for the two tasks, reaching new SOTA results. Mengke Song, Chenglizhao Chen, Wenfeng Song, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | TalkingStyle: Personalized Speech-Driven 3D Facial Animation With Style PreservationabstractIt is a challenging task to create realistic 3D avatars that accurately replicate individuals' speech and unique talking styles for speech-driven facial animation. Existing techniques have made remarkable progress but still struggle to achieve lifelike mimicry. This article proposes "TalkingStyle", a novel method to generate personalized talking avatars while retaining the talking style of the person. Our approach uses a set of audio and animation samples from an individual to create new facial animations that closely resemble their specific talking style, synchronized with speech. We disentangle the style codes from the motion patterns, allowing our method to associate a distinct identifier with each person. To manage each aspect effectively, we employ three separate encoders for style, speech, and motion, ensuring the preservation of the original style while maintaining consistent motion in our stylized talking avatars. Additionally, we propose a new style-conditioned transformer decoder, offering greater flexibility and control over the facial avatar styles. We comprehensively evaluate TalkingStyle through qualitative and quantitative assessments, as well as user studies demonstrating its superior realism and lip synchronization accuracy compared to current state-of-the-art methods. Wenfeng Song, Xuan Wang 0024, Shi Zheng, Shuai Li 0001, Aimin Hao, Xia Hou |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Weakly Supervised Multimodal Affordance Grounding for Egocentric ImagesabstractTo enhance the interaction between intelligent systems and the environment, locating the affordance regions of objects is crucial. These regions correspond to specific areas that provide distinct functionalities. Humans often acquire the ability to identify these regions through action demonstrations and verbal instructions. In this paper, we present a novel multimodal framework that extracts affordance knowledge from exocentric images, which depict human-object interactions, as well as from accompanying textual descriptions that describe the performed actions. The extracted knowledge is then transferred to egocentric images. To achieve this goal, we propose the HOI-Transfer Module, which utilizes local perception to disentangle individual actions within exocentric images. This module effectively captures localized features and correlations between actions, leading to valuable affordance knowledge. Additionally, we introduce the Pixel-Text Fusion Module, which fuses affordance knowledge by identifying regions in egocentric images that bear resemblances to the textual features defining affordances. We employ a Weakly Supervised Multimodal Affordance (WSMA) learning approach, utilizing image-level labels for training. Through extensive experiments, we demonstrate the superiority of our proposed method in terms of evaluation metrics and visual results when compared to existing affordance grounding models. Furthermore, ablation experiments confirm the effectiveness of our approach. Code:https://github.com/xulingjing88/WSMA. Lingjing Xu, Yang Gao 0032, Wenfeng Song, Aimin Hao |
AAAI | 3 |
| 2024 | Frequency-Guided Network for Low-contrast Staining-free Dental Plaque SegmentationabstractTraditional dental plaque detection relies on medical staining reagents and professional intervention. Deep learning-based automatic staining-free dental plaque segmentation provides an alternative for patients to perform plaque detection at home without staining reagents. However, existing methods still struggle with low-contrast visual features between unstained plaque and healthy teeth. To address this, we propose a Frequency-Guided Network (FGN) for low-contrast staining-free dental plaque segmentation. We observe that dental plaque tends to concentrate specifically near the junction between the teeth and the gingiva. This junction demonstrates abrupt changes in pixel values, indicating high-frequency regions in the image. In other words, dental plaque tends to appear near the high-frequency regions of oral endoscope images. Exploiting this characteristic, we employ a frequency-guided decoupling module to separate the image into high-frequency and low-frequency regions automatically and expand the high-frequency region to encompass nearby potential dental plaque. Then we supervise two regions individually to specifically focus on the expended high-frequency region for localizing nearby dental plaque. Additionally, we propose a high-to-low frequency multiple tasks framework. In the first phase, the network segments the teeth region, and then we input the teeth mask into the second phase. In the second stage, the teeth mask allows us to have a higher frequency at the junction between the teeth and gums, thereby enhancing the effectiveness of frequency-guided decoupling. Furthermore, FGN integrates a frequency-driven refinement module to enhance the guidance quality of the teeth mask for the second phase. Extensive evaluations of the oral endoscope dataset demonstrate that our method outperforms existing high-performance segmentation methods. User studies also confirm that our approach achieves superior results to experienced dentists. https://frequency-guided-network.github.io/ Yiming Jiang 0018, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
BIBM | 2 |
| 2024 | mvEchoSeg: One-shot In-context Learning for Multi-view Echocardiography SegmentationabstractEchocardiography is the clinical standard for evaluation of cardiac morphology, and function, and providing hemodynamic parameters in patients with known or suspected heart disease. Due to the diversity and complexity of diagnostic tasks, comprehensive interpretation of echocardiography often requires multi-view imaging for integrated metric analysis. Deep learning methods have become the mainstream approach for echocardiography segmentation. However, achieving segmentation of multi-view echocardiography still requires a substantial amount of annotation. To remedy this, we propose a one-shot in-context learning network mvEchoSeg for multi-view echocardiography segmentation. This network requires only one annotated image for each view. Specifically, we propose a Task Prompt Identifier (TPI) module to identify the task type of the image and allocate the most precise task prompt for it, with minimal adaptation of CLIP and few-shot strategies. Additionally, we leverage a unified In-Context Model (ICM) capable of performing a diverse set of echocardiography segmentation tasks automatically. Using fine-tuning and low-rank adapters improved the performance of the pre-trained model, achieving significant results with minimal training cost. Furthermore, we collect a multi-view echocardiography dataset (MVECD) with 8 views to evaluate our method. The results show an improvement of more than 10% in the DICE score compared to the SOTA foundational medical image segmentation models. To our knowledge, this is the first exploration of a one-shot model for multi-view echocardiography segmentation. Our codes and models are available at https://github.com/stellating/mvEchoSeg. Ya Duan, Wenfeng Song, Nannan Li 0002, Aili Li, Shuai Li 0001 |
BIBM | 3 |
| 2024 | Radiographic Reports Generation via Retrieval Enhanced Cross-modal FusionabstractAccurate radiographic reports are crucial for effective clinical decision-making and patient safety, as they directly influence diagnosis and treatment plans. Existing models for radiographic report generation often struggle with integrating medical image and textual report features and addressing data imbalances. To overcome these limitations, we propose an enhanced cross-modal aligned retrieval-driven network (EARnet). Our model incorporates two key innovations: Enhanced Cross-modal Alignment (ECA) module and Case-based Retrieval Augmenter (CRA) module. ECA ensures the effective integration of medical visual and textual data by aligning these modalities. CRA helps mitigate data imbalance by improving the representation of medical information and ensuring a more balanced and comprehensive coverage of both normal and abnormal cases. This dual-module approach significantly improves the coherence and accuracy of the generated medical reports. Evaluation results demonstrate that our EARnet model substantially outperforms existing methods in terms of report quality and accuracy across multiple metrics. Our codes and models are available at https://github.com/lyf616/EARnet. Xia Hou, Wenfeng Song, Wenzhe You, Shuai Li 0001 |
BIBM | 3 |
| 2024 | Arbitrary Motion Style Transfer with Multi-Condition Motion Latent Diffusion ModelabstractComputer animation's quest to bridge content and style has historically been a challenging venture, with previous efforts often leaning toward one at the expense of the other. This paper tackles the inherent challenge of content-style duality, ensuring a harmonious fusion where the core narrative of the content is both preserved and elevated through stylistic enhancements. We propose a novel Multi-condition Motion Latent Diffusion Model (MCM-LDM) for Arbitrary Motion Style Transfer (AMST). Our MCM-LDM significantly emphasizes preserving trajectories, recognizing their fundamental role in defining the essence and fluidity of motion content. Our MCM-LDM's cornerstone lies in its ability first to disentangle and then intricately weave together motion's tripartite components: motion trajectory, motion content, and motion style. The critical insight of MCM-LDM is to embed multiple conditions with distinct priorities. The content channel serves as the primary flow, guiding the overall structure and movement, while the trajectory and style channels act as auxiliary components and synchronize with the primary one dynamically. This mechanism ensures that multi-conditions can seamlessly integrate into the main flow, enhancing the overall animation without overshadowing the core content. Empirical evaluations underscore the model's proficiency in achieving fluid and authentic motion style transfers, setting a new benchmark in the realm of computer animation. The source code and model are available at https://github.com/XingliangJin/MCM-LDM.git. Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou, Hong Qin 0001 |
CVPR | 1 |
| 2024 | HOIAnimator: Generating Text-Prompt Human-Object Animations Using Novel Perceptive Diffusion ModelsabstractTo date, the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehen-sive physics-centric model supporting real-world dynamics. To ameliorate, this paper advocates HOIAnimator, a novel and interactive diffusion model with perception ability and also ingeniously crafted to revolutionize the animation of complex interactions from linguistic narratives. The effectiveness of our model is anchored in two ground-breaking innovations: (1) Our Perceptive Diffusion Models (PDM) brings together two types of models: one focused on hu-man movements and the other on objects. This combination allows for animations where humans and objects move in concert with each other, making the overall motion more realistic. Additionally, we propose a Perceptive Message Passing (PMP) mechanism to enhance the communication bridging the two models, ensuring that the animations are smooth and unified; (2) We devise an Interaction Contact Field (ICF), a sophisticated model that implicitly captures the essence of HOls. Beyond mere predictive contact points, the ICF assesses the proximity of human and object to their respective environment, informed by a probabilistic distribution of interactions learned throughout the denoising phase. Our comprehensive evaluation showcases HOlani-mator's superior ability to produce dynamic, context-aware animations that surpass existing benchmarks in text-driven animation synthesis. Wenfeng Song, Shuai Li 0001, Yang Gao 0032, Aimin Hao, Xia Hau, Chenglizhao Chen, Hong Qin 0001 |
CVPR | 1 |
| 2024 | FusionCraft: Fusing Emotion and Identity in Cross-Modal 3D Facial Animation
Zhenyu Lv, Xuan Wang 0024, Wenfeng Song, Xia Hou |
ICIC (10) | 3 |
| 2024 | Infrared and Visible Image Fusion Method Based on Learnable Joint Sparse Low-Rank Decomposition
Wenfeng Song, Naiyun Huang, Xiaoqing Luo, Zhancheng Zhang, Tianyang Xu 0001, Xiaojun Wu 0001 |
ICPR (5) | 1 |
| 2024 | CoupNeRF: Property-aware Neural Radiance Fields for Multi-Material Coupled Scenario ReconstructionabstractAbstract Neural Radiance Fields (NeRFs) have achieved significant recognition for their proficiency in scene reconstruction and rendering by utilizing neural networks to depict intricate volumetric environments. Despite considerable research dedicated to reconstructing physical scenes, rare works succeed in challenging scenarios involving dynamic, multi‐material objects. To alleviate, we introduce CoupNeRF, an efficient neural network architecture that is aware of multiple material properties. This architecture combines physically grounded continuum mechanics with NeRF, facilitating the identification of motion systems across a wide range of physical coupling scenarios. We first reconstruct specific‐material of objects within 3D physical fields to learn material parameters. Then, we develop a method to model the neighbouring particles, enhancing the learning process specifically in regions where material transitions occur. The effectiveness of CoupNeRF is demonstrated through extensive experiments, showcasing its proficiency in accurately coupling and identifying the behavior of complex physical scenes that span multiple physics domains. Jin Li 0068, Yang Gao 0032, Wenfeng Song, Yacong Li, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Comput. Graph. Forum | 3 |
| 2024 | A review of automatic source code summarization
Xia Hou, Xiuming Qiao, Wenfeng Song |
Empir. Softw. Eng. | 4 |
| 2024 | Correction: Automatic Generation of 3D Scene Animation Based on Dynamic Knowledge Graphs and Contextual Encoding
Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Dynamic attention augmented graph network for video accident anticipation
Wenfeng Song, Shuai Li 0001, Tao Chang, Ke Xie 0005, Aimin Hao, Hong Qin 0001 |
Pattern Recognit. | 1 |
| 2024 | Rethinking Object Saliency Ranking: A Novel Whole-Flow Processing ParadigmabstractExisting salient object detection methods are capable of predicting binary maps that highlight visually salient regions. However, these methods are limited in their ability to differentiate the relative importance of multiple objects and the relationships among them, which can lead to errors and reduced accuracy in downstream tasks that depend on the relative importance of multiple objects. To conquer, this paper proposes a new paradigm for saliency ranking, which aims to completely focus on ranking salient objects by their "importance order". While previous works have shown promising performance, they still face ill-posed problems. First, the saliency ranking ground truth (GT) orders generation methods are unreasonable since determining the correct ranking order is not well-defined, resulting in false alarms. Second, training a ranking model remains challenging because most saliency ranking methods follow the multi-task paradigm, leading to conflicts and trade-offs among different tasks. Third, existing regression-based saliency ranking methods are complex for saliency ranking models due to their reliance on instance mask-based saliency ranking orders. These methods require a significant amount of data to perform accurately and can be challenging to implement effectively. To solve these problems, this paper conducts an in-depth analysis of the causes and proposes a whole-flow processing paradigm of saliency ranking task from the perspective of "GT data generation", "network structure design" and "training protocol". The proposed approach outperforms existing state-of-the-art methods on the widely-used SALICON set, as demonstrated by extensive experiments with fair and reasonable comparisons. The saliency ranking task is still in its infancy, and our proposed unified framework can serve as a fundamental strategy to guide future work. The code and data will be available at https://github.com/MengkeSong/Saliency-Ranking-Paradigm. Mengke Song, Dunquan Wu, Wenfeng Song, Chenglizhao Chen |
IEEE Trans. Image Process. | 4 |
| 2024 | Joints-Centered Spatial-Temporal Features Fused Skeleton Convolution Network for Action RecognitionabstractSkeleton-based action recognition is crucial for natural human-computer interaction, dynamic behavior analysis, and behavior surveillance. The key challenge is to effectively capture the intrinsic local-global clues of the activity. However, it remains challenging to efficiently leverage multidimensional information related to joints' local visual appearances, global spatial relationships, and coherent temporal cues. To address this challenge, we propose a joints-centered spatial-temporal feature-fused framework for action recognition, which exploits skeleton-based graph diffusion and convolution. Specifically, we employ Partial Differential Equation (PDE) based skeleton graph diffusion to automatically activate and diffuse the salient appearance features of joints. This approach simultaneously integrates the joints' appearance clues and their hierarchical relationships at both the super-pixel level and structure level. The diffused appearance-related features of the joints are further fused with skeleton-related spatial-temporal features, and the resulting fused features are fed into a skeleton convolution network for action recognition. Our method was extensively evaluated on two public datasets (NTU-RGBD and UWA3D), and the results demonstrate the improved accuracy and effectiveness of our approach. Our code will be public. Wenfeng Song, Tangli Chu, Shuai Li 0001, Nannan Li 0002, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | CenterFormer: A Novel Cluster Center Enhanced Transformer for Unconstrained Dental Plaque SegmentationabstractDental plaque segmentation is crucial for maintaining oral health. However, accurately segmenting dental plaque in unconstrained environments can be challenging due to its low contrast and high variability in appearance. While existing transformer-based networks rely on attention mechanisms for each pixel, they do not take into account the relationships between neighboring pixels. Consequently, feature extraction is limited, making it difficult to achieve accurate segmentation of low-contrast images. To address this issue, we propose a simple yet efficient cluster center transformer that improves dental plaque segmentation by clustering image pixels based on multiple levels of feature maps' intensity and texture information. By grouping similar pixels into regions, the proposed method enables the transformers to focus on the local contour and edge around the teeth regions, adapting to the low contrast and high variability of plaque appearance, leading to more accurate and efficient segmentation of dental plaque in dental images. Additionally, we designed Multiple Granularity Perceptions using a pyramid fusion mechanism to capture multiple scales of vision features, thereby enhancing the low-contrast vision features. The proposed method can benefit the dental diagnosis and treatment planning process by improving the accuracy and efficiency of dental plaque segmentation. Our proposed method achieved state-of-the-art results on the dental plaque dataset (Li et al., 2020), with intersection over union (IoU) of 60.91% and pixel accuracy (PA) of 76.81%, all of which were the highest among all methods, demonstrating its effectiveness in plaque segmentation in unconstrained environments. Wenfeng Song, Xuan Wang 0024, Shuai Li 0001, Aimin Hao |
IEEE Trans. Multim. | 1 |
| 2024 | Expressive 3D Facial Animation Generation Based on Local-to-Global Latent Diffusionabstract3D Facial animations, crucial to augmented and mixed reality digital media, have evolved from mere aesthetic elements to potent storytelling media. Despite considerable progress in facial animation of neutral emotions, existing methods still struggle to capture the authenticity of emotions. This paper introduces a novel approach to capture fine facial expressions and generate facial animations using audio synchronization. Our method consists of two key components: First, the Local-to-global Latent Diffusion Model (LG-LDM) tailored for authentic facial expressions, which can integrate audio, time step, facial expressions, and other conditions towards possible encoding of emotionally rich yet latent features in response to possibly noisy raw audio signals. The core of LG-LDM is our carefully designed Facial Denoiser Model (FDM) for aligning the local-to-global animation feature with audio. Second, we redesign an Emotion-centric Vector Quantized-Variational AutoEncoder framework (EVQ-VAE) to finely decode the subtle differences under different emotions and reconstruct the final 3D facial geometry. Our work significantly contributes to the key challenges of emotionally realistic 3D facial animation for audio synchronization and enhances the immersive experience and emotional depth in augmented and mixed reality applications. We provide a reproducibility kit including our code, dataset, and detailed instructions for running the experiments. This kit is available at https://github.com/wangxuanx/Face-Diffusion-Model. Wenfeng Song, Xuan Wang 0024, Yiming Jiang 0018, Shuai Li 0001, Aimin Hao, Xia Hou, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | RGB and LUT based Cross Attention Network for Image Enhancement
Tengfei Shi, Chenglizhao Chen, Yuanbo He, Wenfeng Song, Aimin Hao |
BMVC | 4 |
| 2023 | Sequential Texts Driven Cohesive Motions Synthesis with Natural TransitionsabstractThe intelligent synthesis/generation of daily-life motion sequences is fundamental and urgently needed for many VR/metaverse-related applications. However, existing approaches commonly focus on monotonic motion generation (e.g., walking, jumping, etc.) based on single instruction-like text, which is still not intelligent enough and can’t meet practical demands. To this end, we propose a cohesive human motion sequence synthesis framework based on free-form sequential texts while ensuring semantic connection and natural transitions between adjacent motions. At the technical level, we explore the local-to-global semantic features of previous and current texts to extract relevant information. This information is used to guide the framework in understanding the semantics of the current moment. Moreover, we propose learnable tokens to adaptively learn the influence range of the previous motions towards natural transitions. These tokens can be trained to encode the relevant information into well-designed transition loss. To demonstrate the efficacy of our method, we conduct extensive experiments and comprehensive evaluations on the public dataset as well as a new dataset produced by us. All the experiments confirm that our method outperforms the state-of-the-art methods in terms of semantic matching, realism, and transition fluency. Our project is public available. https://druthrie.github.io/sequential-texts-to-motion/ Shuai Li 0001, Sisi Zhuang, Wenfeng Song, Hejia Chen, Aimin Hao |
ICCV | 3 |
| 2023 | Tell Your Story: Text-Driven Face Video Synthesis with High Diversity via Adversarial LearningabstractFace synthesis is a rapidly growing area of research in computer vision. Text-driven face synthesis is particularly flexible, but challenges still exist in fusing the semantics of text and images, as well as generating diverse faces. To address these challenges, we propose a cross-modality adversarial learning framework to generate highly diverse face videos that correspond to given text descriptions. We encode text and images into a common latent space and align text and image features to control the synthesis of face attributes. We have designed a novel auto-encoder with a face identity discriminator that enlarges the margin between different individuals, increasing the variety of created faces while maintaining the semantic coherence of text and images. Our proposed method has been successfully tested on the recently released Multimodal VoxCeleb dataset. Our code is public available at https://github.com/sunmeng7/TYS.git. Xia Hou, Wenfeng Song |
ICIP | 3 |
| 2023 | Joint Probability Distribution Regression for Image CroppingabstractImage cropping aims at locating a candidate (rectangle region) with the highest aesthetic quality in professional photography. One solution of the previous methods is to generate a large number of candidates and then filter them, which leads to low efficiency. Another idea directly regresses the candidate coordinates to speed up but ignores the aesthetic subjectivity of the candidate’s evaluation, limiting the model’s performance. In this paper, we present an Aesthetic and Composition joint Probability Distribution regression Network (ACPD-Net) to explicitly investigate the process of generating the candidate with a joint probability distribution paradigm to improve the performance of cropping results in an efficient way. The joint probability distribution paradigm between location and size branch can identify the subjective aesthetic region and satisfy the objective composition rules in an end-to-end manner. Our method has been tested on the FCDB and FLMS datasets, which shows the superiority of ACPD-Net. The code is available at https://github.com/flyingbird93/ACPD-Net. Tengfei Shi, Chenglizhao Chen, Yuanbo He, Wenfeng Song, Aimin Hao |
ICIP | 4 |
| 2023 | Dual Temporal Transformers for Fine-Grained Dangerous Action RecognitionabstractRecognizing dangerous actions is a critical task in computer vision, especially for surveillance applications. While existing deep learning methods have been successful in confined environments, they struggle with the anomalous and salient variations of human postures in dangerous actions. Additionally, finer-grained dangerous actions require more discriminative cues, adding to the complexity of the task. To address these challenges, we propose a novel solution that models the intrinsic and invariant properties of dangerous actions at multiple temporal semantic levels. Concretely, we propose a Dual Temporal Transformers (DTT) to capture temporal interactions between distinct key points in the human body aggregation from shallow to deep layers, increasing the perception field from local to global, simultaneously. By doing so, our method avoids overfitting to unrelated or minor clues in videos and achieves a generalized representation of abnormal actions. We evaluate our approach on indoor and outdoor environments and found that DTT outperforms existing methods in terms of efficiency and accuracy. Our code and dataset are pubic available on https://github.com/AveryJohnsonJJ/DTT.git. Wenfeng Song, Xingliang Jin, Yang Gao 0032, Xia Hou |
ICIP | 1 |
| 2023 | Automatic Generation of 3D Scene Animation Based on Dynamic Knowledge Graphs and Contextual Encoding
Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Int. J. Comput. Vis. | 1 |
| 2023 | Graph Diffusion Convolutional Network for Skeleton Based Semantic Recognition of Two-Person ActionsabstractGraph Convolutional Networks (GCNs) have successfully boosted skeleton-based human action recognition. However, existing GCN-based methods mostly cast the problem as separated person's action recognition while ignoring the interaction between the action initiator and the action responder, especially for the fundamental two-person interactive action recognition. It is still challenging to effectively take into account the intrinsic local-global clues of the two-person activity. Additionally, message passing in GCN depends on adjacency matrix, but skeleton-based human action recognition methods tend to calculate the adjacency matrix with the fixed natural skeleton connectivity. It means that messages can only travel along a fixed path at different layers of the network or in different actions, which greatly reduces the flexibility of the network. To this end, we propose a novel graph diffusion convolutional network for skeleton based semantic recognition of two-person actions by embedding the graph diffusion into GCNs. At technical fronts, we dynamically construct the adjacency matrix based on practical action information, so that we can guide the message propagation in a more meaningful way. Simultaneously, we introduce the frame importance calculation module to conduct dynamic convolution, so that we can avoid the negative effect caused by the traditional convolution, wherein the shared weights may fail to capture key frames or be affected by noisy frames. Besides, we comprehensively leverage the multidimensional features related to joints' local visual appearances, global spatial relationship and temporal coherency, and for different features, different metrics are designed to measure the similarity underlying the corresponding real physical law of the motions. Moreover, extensive experiments and comprehensive evaluations on four public large-scale datasets (NTU-RGB+D 60, NTU-RGB+D 120, Kinetics-Skeleton 400, and SBU-Interaction) demonstrate that our method outperforms the state-of-the-art methods. Shuai Li 0001, Xinxue He, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | SC-GAN: Subspace Clustering based GAN for Automatic Expression Manipulation
Shuai Li 0001, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
Pattern Recognit. | 4 |
| 2023 | A Comprehensive Survey on Video Saliency Detection With Auditory Information: The Audio-Visual Consistency Perceptual is the Key!abstractVideo saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect. In contrast, our audio system is the most vital complementary part of our visual system. Also, audio-visual saliency detection (AVSD), one of the most representative research topics for mimicking human perceptual mechanisms, is currently in its infancy, and none of the existing survey papers have touched on it, especially from the perspective of saliency detection. Thus, the ultimate goal of this paper is to provide an extensive review to bridge the gap between audio-visual fusion and saliency detection. In addition, as another highlight of this review, we have provided a deep insight into key factors that could directly determine AVSD deep models’ performances. We claim that the audio-visual consistency degree (AVC) — a long-overlooked issue, can directly influence the effectiveness of using audio to benefit its visual counterpart when performing saliency detection. Moreover, to make the AVC issue more practical and valuable for future followers, we have newly equipped almost all existing publicly available AVSD datasets with additional frame-wise AVC labels. Based on these upgraded datasets, we have conducted extensive quantitative evaluations to ground our claim on the importance of AVC in the AVSD task. In a word, our ideas and new sets serve as a convenient platform with preliminaries and guidelines, all of which can potentially facilitate future works in further promoting state-of-the-art (SOTA) performance. Chenglizhao Chen, Mengke Song, Wenfeng Song, Li Guo 0016, Muwei Jian |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | An Intelligent Virtual Standard Patient for Medical Students Training Based on Oral Knowledge GraphabstractVirtual standard patient (VSP) is in high demand for medical students' diagnosis ability training in an efficient manner. Different from the traditional conversation system in medical dialogue generation, VSP needs a novel conversation paradigm to act as the patient instead of the doctor. However, existing conversation techniques still have limited ability in terms of generation of symptoms exhibited by patients with the personalized and knowledge-centered expressions. To alleviate these problems, we propose to construct a novel oral knowledge graph, which sufficiently provides medical clues of the certain disease. Accordingly, the VSP could accurately interact with the dentists for their underlying intention and express the symptoms characters in a natural style. To efficiently retrieve the related disease clues, the symptoms descriptions of the oral diseases are encoded into the oral knowledge graph, which could well organize the disease-centered symptom entities and speaking styles. Moreover, to transfer the common sense knowledge from existing large scale of medical knowledge graph to the specific oral knowledge graph, a coupled pre-trained Bert models is further designed to learn the related medical knowledge from coarse-level to fine-level hierarchically. Finally, a series of well-designed personalized templates are proposed to generate plausible and realistic answers in condition of the certain disease. We also conduct extensive user studies to demonstrate that the VSP satisfies the medical students' diagnosis practice requirement in terms of naturalness, realism, and topic relevance. Wenfeng Song, Xia Hou, Shuai Li 0001, Chenglizhao Chen, Danyang Gao, Xian'e Wang, Yuzhe Sun, Jianxia Hou, Aimin Hao |
IEEE Trans. Multim. | 1 |
| 2023 | FineStyle: Semantic-Aware Fine-Grained Motion Style Transfer with Dual Interactive-Flow FusionabstractWe present FineStyle, a novel framework for motion style transfer that generates expressive human animations with specific styles for virtual reality and vision fields. It incorporates semantic awareness, which improves motion representation and allows for precise and stylish animation generation. Existing methods for motion style transfer have all failed to consider the semantic meaning behind the motion, resulting in limited controls over the generated human animations. To improve, FineStyle introduces a new cross-modality fusion module called Dual Interactive-Flow Fusion (DIFF). As the first attempt, DIFF integrates motion style features and semantic flows, producing semantic-aware style codes for fine-grained motion style transfer. FineStyle uses an innovative two-stage semantic guidance approach that leverages semantic clues to enhance the discriminative power of both semantic and style features. At an early stage, a semantic-guided encoder introduces distinct semantic clues into the style flow. Then, at a fine stage, both flows are further fused interactively, selecting the matched and critical clues from both flows. Extensive experiments demonstrate that FineStyle outperforms state-of-the-art methods in visual quality and controllability. By considering the semantic meaning behind motion style patterns, FineStyle allows for more precise control over motion styles. Source code and model are available on https://github.com/XingliangJin/Fine-Style.git. Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | Person Re-Identification in Panoramic Views Based on Bayesian TransformersabstractThe panoramic view cameras offer more broad perspectives and continuous information for person re-identification (ReID). However, the panoramic-view videos suffer from objects distortion and bring more occlusion due to the fixed or moving capture points. This paper proposes a novel Bayesian Transformer Network (BTN) to adaptively capture the occlusion clues as Bayesian prior to guide the discriminative pedestrian-related feature extraction in the high-occlusion scenes. The Bayesian prior is built via a pre-trained CNN, which could recognize different occluded scenarios based on the severeness of noisy backgrounds. Moreover, to fully explore the occlusion prior, we propose to embed the semantic labels into a well-designed transformer network. By fostering the collaborative occlusion clues between the person and background, our method could achieve outstanding performance on both public benchmarks and panoramic view videos, which verifies the advantages of our BTN framework over existing methods. Wenfeng Song, Yang Gao 0032, Aimin Hao, Xia Hou |
ICIP | 1 |
| 2022 | Automatic image matting and fusing for portrait synthesis
Zhike Yi, Wenfeng Song, Shuai Li 0001, Aimin Hao |
Sci. China Inf. Sci. | 2 |
| 2022 | Improving RGB-D Salient Object Detection via Modality-Aware DecoderabstractMost existing RGB-D salient object detection (SOD) methods are primarily focusing on cross-modal and cross-level saliency fusion, which has been proved to be efficient and effective. However, these methods still have a critical limitation, i.e., their fusion patterns - typically the combination of selective characteristics and its variations, are too highly dependent on the network's non-linear adaptability. In such methods, the balances between RGB and D (Depth) are formulated individually considering the intermediate feature slices, but the relation at the modality level may not be learned properly. The optimal RGB-D combinations differ depending on the RGB-D scenarios, and the exact complementary status is frequently determined by multiple modality-level factors, such as D quality, the complexity of the RGB scene, and degree of harmony between them. Therefore, given the existing approaches, it may be difficult for them to achieve further performance breakthroughs, as their methodologies belong to some methods that are somewhat less modality sensitive. To conquer this problem, this paper presents the Modality-aware Decoder (MaD). The critical technical innovations include a series of feature embedding, modality reasoning, and feature back-projecting and collecting strategies, all of which upgrade the widely-used multi-scale and multi-level decoding process to be modality-aware. Our MaD achieves competitive performance over other state-of-the-art (SOTA) models without using any fancy tricks in the decoder's design. Codes and results will be publicly available at https://github.com/MengkeSong/MaD. Mengke Song, Wenfeng Song, Guowei Yang 0002, Chenglizhao Chen |
IEEE Trans. Image Process. | 2 |
| 2022 | Automatic Dental Plaque Segmentation Based on Local-to-Global Features Fused Self-Attention NetworkabstractThe accurate detection of dental plaque at an early stage will definitely prevent periodontal diseases and dental caries. However, it remains difficult for the current dental examination to accurately recognize dental plaque without using medical dyeing reagent due to the low contrast between dental plaque and healthy teeth. To combat this problem, this paper proposes a novel network enhanced by a self-attention module for intelligent dental plaque segmentation. The key motivation is to directly utilize oral endoscope images (bypassing the need for dyeing reagent) and get accurate pixel-level dental plaque segmentation results. The algorithm needs to conduct self-attention at the super-pixel level and fuse the super-pixels' local-to-global features. Our newly-designed network architecture will afford the simultaneous fusion of multiple-scale complementary information guided by the powerful deep learning paradigm. The critical fused information includes the statistical distribution of the plaques color, the heat kernel signature (HKS) based local-to-global structure relationship, and the circle-LBP based local texture pattern in the nearby regions centering around the plaque area. To further refine the fuzed multiple-scale features, we devise an attention module based on CNN, which could focalize the regions of interest in plaque more easily, especially for many challenging cases. Extensive experiments and comprehensive evaluations confirm that, for a small-scale training dataset, our method could outperform the state-of-the-art methods. Meanwhile, the user studies verify the claim that our method is more accurate than conventional dental practice conducted by experienced dentists. Shuai Li 0001, Zhennan Pang, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Correction to: Long-Short Temporal-Spatial Clues Excited Network for Robust Person Re-identification
Shuai Li 0001, Wenfeng Song, Zheng Fang 0008, Jiaying Shi, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Hierarchical Object Relationship Constrained Monocular Depth Estimation
Shuai Li 0001, Jiaying Shi, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
Pattern Recognit. | 3 |
| 2021 | A Global-Local Self-Adaptive Network for Drone-View Object DetectionabstractDirectly benefiting from the deep learning methods, object detection has witnessed a great performance boost in recent years. However, drone-view object detection remains challenging for two main reasons: (1) Objects of tiny-scale with more blurs w.r.t. ground-view objects offer less valuable information towards accurate and robust detection; (2) The unevenly distributed objects make the detection inefficient, especially for regions occupied by crowded objects. Confronting such challenges, we propose an end-to-end global-local self-adaptive network (GLSAN) in this paper. The key components in our GLSAN include a global-local detection network (GLDN), a simple yet efficient self-adaptive region selecting algorithm (SARSA), and a local super-resolution network (LSRN). We integrate a global-local fusion strategy into a progressive scale-varying network to perform more precise detection, where the local fine detector can adaptively refine the target's bounding boxes detected by the global coarse detector via cropping the original images for higher-resolution detection. The SARSA can dynamically crop the crowded regions in the input images, which is unsupervised and can be easily plugged into the networks. Additionally, we train the LSRN to enlarge the cropped images, providing more detailed information for finer-scale feature extraction, helping the detector distinguish foreground and background more easily. The SARSA and LSRN also contribute to data augmentation towards network training, which makes the detector more robust. Extensive experiments and comprehensive evaluations on the VisDrone2019-DET benchmark dataset and UAVDT dataset demonstrate the effectiveness and adaptivity of our method. Towards an industrial application, our network is also applied to a DroneBolts dataset with proven advantages. Our source codes have been available at https://github.com/dengsutao/glsan. Sutao Deng, Shuai Li 0001, Ke Xie 0005, Wenfeng Song, Xiao Liao, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Meta-RetinaNet for Few-shot Object Detection
Shaoqi Li, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
BMVC | 2 |
| 2020 | Meta Transfer Learning for Adaptive Vehicle Tracking in UAV Videos
Wenfeng Song, Shuai Li 0001, Shaoqi Li, Aimin Hao, Hong Qin 0001, Qinping Zhao |
MMM (1) | 1 |
| 2020 | Cross-View Contextual Relation Transferred Network for Unsupervised Vehicle Tracking in Drone VideosabstractRecently CNN-centric object tracking methods have been gaining tremendous success in ground-view videos, however, it remains hard to cope with vehicle tracking in unmanned aerial vehicle (UAV) videos. The key difficulties mainly stem from lacking large-scale well-labeled training datasets and view-invariant appearance model for fast-moving drone-view vehicles. We enhance the vehicle's cross-view feature by exploring relations between the pivotal context and the target to facilitate unsupervised vehicle tracking. The relation is modeled as the relevance of the target and its contextual regions in the tracking task. Specifically, we propose a contextual relation actor-critic (CRAC) framework integrates an actor-critic agent with a dual GAN learning mechanism, which aims to dynamically search the related contextual regions and transfer the relations from ground-view to drone-view videos while retaining the discriminative features. We demonstrate that CRAC could be applied to several state-of-the-art trackers by extensive experiments and ablation studies on four public benchmarks. All the experiments confirm that, our CRAC can improve the performance of state-of-the-art methods in terms of accuracy, robustness, and versatility. Wenfeng Song, Shuai Li 0001, Tao Chang, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
WACV | 1 |
| 2020 | Long-Short Temporal-Spatial Clues Excited Network for Robust Person Re-identification
Shuai Li 0001, Wenfeng Song, Zheng Fang 0008, Jiaying Shi, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Context-Interactive CNN for Person Re-IdentificationabstractDespite growing progresses in recent years, cross-scenario person re-identification remains challenging, mainly due to the pedestrians commonly surrounded by highly-complex environment contexts. In reality, the human perception mechanism could adaptively find proper contextualized spatial-temporal clues towards pedestrian recognition. However, conventional methods fall short in adaptively leveraging the long-term spatial-temporal information due to ever-increasing computational cost. Moreover, CNN-based deep learning methods are hard to conduct optimization due to the non-differentiable property of the built-in context search operation. To ameliorate, this paper proposes a novel Context-Interactive CNN (CI-CNN) to dynamically find both spatial and temporal contexts by embedding multi-task Reinforcement Learning (MTRL). The CI-CNN streamlines the multi-task reinforcement learning by using an actor-critic agent to capture the temporal-spatial context simultaneously, which comprises a context-policy network and a context-critic network. The former network learns policies to determine the optimal spatial context region and temporal sequence range. Based on the inferred temporal-spatial cues, the latter one focuses on the identification task and provides feedback for the policy network. Thus, CI-CNN can simultaneously zoom in/out the perception field in spatial and temporal domain for the context interaction with the environment. By fostering the collaborative interaction between the person and context, our method could achieve outstanding performance on various public benchmarks, which confirms the rationality of our hypothesis, and verifies the effectiveness of our CI-CNN framework. Wenfeng Song, Shuai Li 0001, Tao Chang, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Contextualized CNN for Scene-Aware Depth Estimation From Single RGB ImageabstractDirectly benefited from deep learning techniques, depth estimation from single image has gained great momentum in recent years. However, most of the existing approaches treat depth prediction as an isolated problem without taking into consideration high-level semantic context information, which results in inefficient utilization of training dataset and unavoidably requires a large number of captured depth data during the training phase. To ameliorate, this paper develops a novel scene-aware contextualized convolution neural network (CCNN), which characterizes the semantic context relationship at the class-level and refines depth at the pixel-level. Our newly-proposed CCNN is built upon the intrinsic exploitation of context-dependent depth association, including inner-object continuous depth and inter-object depth change priors nearby. Specifically, rather than conducting regression on depth in single CNN, we make the first attempt to integrate both class-level and pixel-level conditional random fields (CRFs) based probabilistic graphical model into the powerful CNN framework to simultaneously learn different-level features within the same CNN layer. With our CCNN, the former model will guide the latter one to learn the contextualized RGB-Depth mapping. Hence, CCNN has desirable properties in both class-level integrity and pixel-level discrimination, which makes it ideal to share such two-level convolutional features in parallel during the end-to-end training with the commonly-used back-propagation algorithm. We conduct extensive experiments and comprehensive evaluations on public benchmarks involving various indoor and outdoor scenes, and all the experiments confirm that, our method outperforms the state-of-the-art depth estimation methods, especially for the cases where only small-scale training data are readily available. Wenfeng Song, Shuai Li 0001, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Fine-Grained Thyroid Nodule Classification via Multi-Semantic Attention NetworkabstractThyroid nodule classification in ultrasound images has gained great momentum based on deep convolutional neural networks in recent years. Nevertheless, it is still challenging to intelligently classify the fine-grained thyroid nodules, which is significant for the subsequent clinical treatments. The difficulties mainly stem from four aspects: few fine-grained training dataset, highly-variable appearances of intra-class nodules, overall-similar characteristics of inter-class nodules, and the low resolution and contrast degree of the ultrasonic images as well as the influence of intrinsic speckle noises. In this paper, we propose a multi-semantic attention networks (MSAN) for fine-grained thyroid nodule classification in ultrasound images. Specifically, we employ a main network branch for coarse granularity feature extraction, which only focuses on the benign and malignant characteristics, and simultaneously employ multi-semantic network branches to extract discriminative features from the fine-grained pathological categories. Meanwhile, we introduce an self-attention scheme together with global average pooling (GAP) in our network, which facilitates to learn from the dynamically-selected nodule regions ranging from local to global. Extensive experiments demonstrate that, our MSAN gives rise to significant improvement of classification accuracy and outperforms the state-of-the-art methods. Shuai Li 0001, Wenfeng Song, Zhennan Pang, Aimin Hao, Hong Qin 0001 |
BIBM | 3 |
| 2019 | Few-Shot Learning for Monocular Depth Estimation Based on Local Object RelationshipabstractMonocular depth estimation has gained great momentum and achieved growing success recently. Nonetheless, due to the intrinsic difficulty associated with large-scale RGB-D data capture for training purpose and the inefficient utilization of existing training datasets, it is still challenging to accommodate flexibly-changing scenarios. To ameliorate, we propose a fewshot learning method for monocular depth estimation augmented by local object-object relationship. Our method is based on the insight that the depth changing between neighboring objects is relatively stable across diverse but similar scenarios. At the technical front, we first learn the object relationship based on the relative distance between single objects. Towards this goal, we design a CNN architecture to simultaneously encode the object spatial context into object-object relationship features and encode the original image into global context features. Hence we can complementally leverage few-shot dataset with only a few samples for depth estimation while preserving the global depth changing range and respecting the local object-object depth details. As a result, our novel approach could estimate depth from various indoor RGB images, which greatly alleviates the training dataset dependency in monocular depth estimation. Finally, we conduct extensive experiments and comprehensive evaluations on the widely-used public benchmarks, and all the experiments confirm that, our method outperforms the state-of-the-art depth estimation methods, especially for the cases where only smallscale training samples are available. Shuai Li 0001, Jiaying Shi, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
ICTAI | 3 |
| 2019 | Bidirectional Optimization Coupled Lightweight Networks for Efficient and Robust Multi-Person 2D Pose Estimation
Shuai Li 0001, Zheng Fang 0008, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
J. Comput. Sci. Technol. | 3 |
| 2019 | Multitask Cascade Convolution Neural Networks for Automatic Thyroid Nodule Detection and RecognitionabstractThyroid ultrasonography is a widely used clinical technique for nodule diagnosis in thyroid regions. However, it remains difficult to detect and recognize the nodules due to low contrast, high noise, and diverse appearance of nodules. In today's clinical practice, senior doctors could pinpoint nodules by analyzing global context features, local geometry structure, and intensity changes, which would require rich clinical experience accumulated from hundreds and thousands of nodule case studies. To alleviate doctors' tremendous labor in the diagnosis procedure, we advocate a machine learning approach to the detection and recognition tasks in this paper. In particular, we develop a multitask cascade convolution neural network (MC-CNN) framework to exploit the context information of thyroid nodules. It may be noted that our framework is built upon a large number of clinically confirmed thyroid ultrasound images with accurate and detailed ground truth labels. Other key advantages of our framework result from a multitask cascade architecture, two stages of carefully designed deep convolution networks in order to detect and recognize thyroid nodules in a pyramidal fashion, and capturing various intrinsic features in a global-to-local way. Within our framework, the potential regions of interest after initial detection are further fed to the spatial pyramid augmented CNNs to embed multiscale discriminative information for fine-grained thyroid recognition. Experimental results on 4309 clinical ultrasound images have indicated that our MC-CNN is accurate and effective for both thyroid nodules detection and recognition. For the correct diagnosis rate of malignant and benign thyroid nodules, its mean Average Precision (mAP) performance can achieve up to [Formula: see text] accuracy, which outperforms the common CNNs by [Formula: see text] on average. In addition, we conduct rigorous user studies to confirm that our MC-CNN outperforms experienced doctors, yet only consuming roughly [Formula: see text] ( 1/48) of doctors' examination time on average. Therefore, the accuracy and efficiency of our new method exhibit its great potential in clinical applications. Wenfeng Song, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
IEEE J. Biomed. Health Informatics | 1 |
| 2018 | Automatic Beautification for Group-Photo Facial Expressions Using Novel Bayesian GANs
Shuai Li 0001, Wenfeng Song, Hong Qin 0001, Aimin Hao |
ICANN (1) | 3 |
| 2018 | Learning from Weakly-Labeled Clinical Data for Automatic Thyroid Nodule Classification in Ultrasound ImagesabstractThis paper proposes a semi-supervised learning method based on weakly-labeled data to automatically classify ultrasound (US) thyroid nodules. Key to our new approach is the unification of multi-instance learning (MIL) with deep learning. Benefiting from that, our method can directly use off-the-shelf clinical data, which involves no labels to indicate nodule classes. To this end, we take the US images of a patient as a bag, and take the corresponding pathology report as the bag label. Specifically, we first propose a bag generating method, wherein the detected thyroid nodules are considered as instances corresponding to certain bag. After that, we design an effective EM algorithm to train a convolutional neural network (CNN) for nodule classification. We conduct extensive experiments and comprehensive evaluations on different datasets, and all the experiments confirm that, our method significantly outperforms state-of-the-art MIL algorithms, which exhibits great potential in clinical applications. Jianxiong Wang, Shuai Li 0001, Wenfeng Song, Hong Qin 0001, Aimin Hao |
ICIP | 3 |
| 2018 | Deep variance network: An iterative, improved CNN framework for unbalanced training datasets
Shuai Li 0001, Wenfeng Song, Hong Qin 0001, Aimin Hao |
Pattern Recognit. | 2 |