Fan Zhou 0001

dblp:63/3122-1 · DBLP profile ↗
← Back
65ranked-venue papers
1as first author
50since 2021 · last 2026
0000-0002-0400-9366ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 34 since 2021Artificial intelligence and machine learning · 12 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Software engineering, systems software and programming languages · 4Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 V-HOI: Velocity-Aware Human-Object Interaction Generation
Honghui Chen, Fan Zhou 0001, Ruomei Wang 0001, Baoquan Zhao
MMM (2)2
2026 NEGS-Avatar: Normal Embedded Gaussians for 2D avatar from monocular video
Zedan Zheng, Yudi Tan, Zhuo Su 0001, Fan Zhou 0001, Baoquan Zhao
Comput. Graph.4
2026 Action-aware anchor-based frame selection strategy for action recognition
Fan Zhou 0001, Ge Lin 0002, Zhuo Su 0001
Image Vis. Comput.3
2026 LVAD: A realistic data synthesis strategy and coarse-to-fine framework for low-light video anomaly detection
Linxuan Han, Fan Zhou 0001, Baoquan Zhao
Pattern Recognit.2
2026 PDDA: Prompt-Driven Domain Adaptation for Real-World Image Dehazing
abstract
Due to the complexity and diversity of practical environments, real-world image dehazing remains an unresolved problem, with one of the key challenges being how to bridge the distribution gap between synthetic and real domains. This paper proposes a Prompt-driven Domain Adaptation (PDDA) framework within the bi-level optimization perspective. Specifically, we introduce hyperparameter optimization-based bi-level modeling: the lower-level optimization emphasizes prior learning within the synthetic domain to stabilize dehazing performance, while the upper-level optimization focuses on enhancing cross-domain adaptability to ensure that the model can generalize across different domains. Given the scarcity of paired real haze images, we train learnable haze prompts by jointly optimizing the text-image similarity between positive/negative prompts and corresponding clear/haze images in the CLIP latent space to more effectively capture real-world haze characteristics. Based on the learned haze prompts, we construct an unsupervised cross-domain loss function that enhances the adaptability to complex real-world scenarios by integrating prompt learning with bi-level optimization strategy. Furthermore, we conduct a comprehensive exploration to uncover the inherent properties of PDDA, including architecture-irrelevant flexibility and domain-agnostic robustness. Extensive experiments across a wide range of benchmark datasets demonstrate that our method achieves both quantitative and qualitative improvements across diverse scenarios, showing robust performance not only in real-world daytime conditions but also exhibiting superior cross-domain adaptation capabilities in nighttime scenarios. Codes are available at https://github.com/YanZhang-zy/PDDA.git.
Yan Zhang 0002, Xin Li 0175, Fan Zhou 0001, Zhuo Su 0001
IEEE Trans. Image Process.4
2026 PSAM: Parameter-Free Spatiotemporal Attention Mechanism for Video Question Answering
abstract
Spatiotemporal attention learning has always been a challenging research task in video question answering (VideoQA). It needs to consider not only the modelling of local neighbourhood dependencies between the adjacent frames in a video but also the modelling of long-term dependencies between nonadjacent frames. Although the existing methods are usually good at modelling temporal dependencies in one aspect, they cannot simultaneously and effectively model the temporal dependencies between adjacent and nonadjacent frames. To address this issue, we first derive a novel statistic-driven difference-aware generation function, which can efficiently calculate the difference between a sequence feature value and the whole mean value to identify the significance of the feature. Subsequently, we design a novel parameter-free spatiotemporal attention mechanism (PSAM), which captures the most relevant cues scattered in the context of a spatiotemporal video by generating functions and utilizes a gating mechanism to adaptively integrate and filter relevant and irrelevant information. Finally, we use the PSAM and hierarchical modelling to construct a lightweight multiscale context fusion- and reasoning-based VideoQA model. Extensive experimental research results obtained on five benchmark datasets for the VideoQA task show that our VideoQA model has high Q&A performance and lightweight characteristics. Simultaneously, comprehensive ablation experimental results show that the PSAM can not only improve the performance of the model but also significantly reduce the number of model parameters. In addition, extensive experimental findings obtained on the benchmark dataset of joint tasks (video moment retrieval and video highlight detection) further demonstrate that the PSAM is a general and effective spatiotemporal attention mechanism.
Ruomei Wang 0001, Fan Zhou 0001, Yuanmao Luo
IEEE Trans. Multim.3
2026 Visual-Guided Long Temporal Context Learning Network for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised video anomaly detection (WVAD) aims to locate events or behaviors that deviate from normal patterns in untrimmed videos using video-level labels. Recent studies typically utilize supplementary modalities to assist anomaly detection. However, these methods suffer from two main issues: (1) The limitations of long-duration anomaly event temporal modeling. The model struggles to consistently maintain key information, resulting in the forgetting phenomenon, which affects the tracking of the event’s overall dynamic evolution and complicates anomaly event analysis and understanding. (2) The multi-modal fusion strategy is insufficient, particularly when there is temporal inconsistency between visual and audio information, causing the model to overlook key information, directly affecting the accurate detection and recognition of anomalous events. To address these issues, we propose a visual-guided long-term temporal context learning network (LTCLNet). The network consists of three key components: a cross-modal interaction module, a multi-modal fusion module, and a visual-guided parameter optimization strategy. First, to address the forgetting issue in long-duration anomaly detection, we designed a cross-modal interaction module. The key part of this module is the establishment of a cross-matrix mechanism. This mechanism achieves bidirectional temporal guidance across modalities. It allows the temporal modeling of each modality to dynamically integrate information from the other modality. This enables the model to continuously track the dynamic evolution of the event. The tracking is facilitated through shared temporal information between the visual and audio modalities. Secondly, to fully exploit the complementary characteristics between different modalities, we introduced a novel temporal reversal integration method in the multi-modal fusion module. This method reverses the feature sequences of each modality to enhance the model’s perception of temporal dynamic changes. By fusing the modality features before and after reversal, the shared temporal structure between modalities is strengthened, improving the model’s ability to capture anomalous information. Additionally, our proposed visual-guided parameter optimization strategy trains a parallel visual modality network as a semantic anchor, ensuring that the model stays aligned with a semantically stable and structurally clear visual flow during the learning process, thus ensuring stability and semantic coherence in the training. Extensive experiments on datasets such as XD-Violence demonstrate that our method significantly outperforms existing approaches, particularly achieving notable improvements in the accuracy and stability of long-term anomaly detection. Our code is publicly available at https://github.com/ibliever/LTCLNet .
Ruomei Wang 0001, Linxuan Han, Baoquan Zhao, Fan Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2025 MSPoint-Gait: Multi-Scale Point Cloud Analysis for 3D Gait Recognition via Cross-Modal Learning
abstract
Recent advances in LiDAR technology have enabled privacy-preserving gait recognition using 3D point cloud data. However, existing approaches struggle with the inherent challenges of point cloud processing and understanding such as spatial sparsity, irregular sampling, and complex temporal dynamics. In this paper, we present MSPoint-Gait, a novel framework that addresses these challenges through multi-scale analysis and cross-modal learning. At the core of our framework lies a Depth-Aware Attention Module (DAAM) that leverages rich 3D geometric information to generate attention-weighted depth representations, enabling fine-grained feature extraction from point cloud sequences. We further introduce a Multi-Scale Spatio-Temporal (MSST) network that hierarchically captures both local and global gait patterns through adaptive convolution kernels across multiple spatial and temporal scales. These components are unified through a novel cross-modal learning strategy that effectively bridges the semantic gap between raw point clouds and structured depth representations. The proposed frame-work achieves state-of-the-art performance on the challenging SUSTech1K dataset, with 91.9% Rank-1 and 98.0% Rank-5 accuracy, demonstrating significant improvements over existing methods across various walking conditions and viewpoints.
Xinzhu Li, Yikun Chen, Guanghui Yue 0001, Wei Zhou 0021, Ruomei Wang 0001, Xudong Mao, Juepeng Zheng, Fan Zhou 0001, Ziqi Qiu, Baoquan Zhao
ICME9
2025 Multi-granularity Frequency Difference-Aware Attention for Video Question Answering
abstract
Video Question Answering (VideoQA) demands complex reasoning about multi-granular information, requiring both fine-grained visual details and global event understanding from videos. While existing methods employ stacked cross-modal attention modules for multi-granular feature representation, they struggle to effectively separate different granularities due to the intertwined nature of visual information in videos. To address this challenge, we introduce a novel Multi-granularity Frequency Difference-Aware Attention (MFDA) mechanism that enhances VideoQA by modeling unified multi-granular relation-ships between multimodal features in the frequency domain. MFDA comprises three key components: a heterogeneous multi-granularity dynamic-aware module, a frequency distance-aware function, and a multi-granular complementary module. These components enable the VideoQA model to effectively parse and filter multi-granularity feature information, allocate attention weights, and provide precise visual semantic cues for answer prediction. Extensive experiments demonstrate that MFDA serves as a plug-and-play cross-modal attention mechanism that significantly improves existing VideoQA models’ performance, achieving state-of-the-art results across diverse question types while reducing computational complexity. Our code is available at https://github.com/haha94322/MFDA.
Fan Zhou 0001, Ruomei Wang 0001, Baoquan Zhao
ICME2
2025 MCSMoG: Multi-Conditional Diffusion for Stylized Motion Generation with Parametric Control
abstract
Stylized human motion synthesis remains a fundamental challenge in computer animation and graphics, with a wide spectrum of applications spanning gaming, film production, virtual reality, and beyond. While recent advances in text-driven motion generation have shown promise, existing approaches face critical limitations including the inability to maintain consistent trajectory control, the lack of fine-grained stylization intensity adjustment, and inadequate generalization across diverse motion styles. To address these challenges, We introduce MCSMoG, a novel framework for controllable stylized motion synthesis through multi-conditional guidance. First, a new Multi-Conditional Motion Latent Diffusion (MC-MLD) model is proposed to introduce additional trajectory guidance and achieve trajectory decoupling. Second, we develop a Style and Non-Style Feature Fusion Module that dynamically blends motion features through an adjustable parameter, providing control over stylization intensity. Third, we integrate MotionCLIP as our style encoder, enhancing the model’s generalization capability across diverse and unseen motion styles. Extensive experiments conducted on the combined HumanML3D and 100STYLE datasets demonstrate that our approach outperforms state-of-the-art methods, achieving a 4.6% reduction in FID scores and a 4.1% increase in motion diversity. User studies further confirm the superiority of our method in style fidelity, semantic consistency, and motion naturalness.
Xinzhu Li, Guanghui Yue 0001, Wei Zhou 0021, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao
ICME8
2025 Breaking the Synthetic Barrier: Towards Stable and Generalizable Real-World Image Dehazing
abstract
Existing learning-based dehazing methods perform well on synthetic data but struggle in real scenarios due to the domain gap, causing residual haze and detail loss. To address this, we propose a Multilevel Subspace Distribution Adapter (MSDA) to progressively reduce the feature distribution gap through hierarchical subspace modeling. We also introduce a Dual-Domain Synchronous Optimization (DDSO) strategy that jointly leverages synthetic supervision and adaptation to the real domain in a unified training scheme. Extensive experiments underscore the superiority of our approach and its excellence on no-reference image quality metrics.
Zhuo Su 0001, Jufeng Li, Yan Zhang 0002, Xin Li 0175, Fan Zhou 0001
ACM Multimedia7
2025 DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
abstract
Robust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing sufficient cues that can be exploited to handle viewpoint variations, and capture finer and meaningful details of gait. In this paper, we introduce a novel framework, termed DepthGait, that incorporates RGB-derived depth maps and silhouettes for enhanced gait recognition. Specifically, apart from the 2D silhouette representation of the human body, the proposed pipeline explicitly estimates depth maps from a given RGB image sequence and uses them as a new modality to capture discriminative features inherent in human locomotion. In addition, a novel multi-scale and cross-level fusion scheme has also been developed to bridge the modality gap between depth maps and silhouettes. Extensive experiments on standard benchmarks demonstrate that the proposed DepthGait achieves state-of-the-art performance compared to peer methods and attains an impressive mean rank-1 accuracy on the challenging datasets.
Xinzhu Li, Juepeng Zheng, Yikun Chen, Xudong Mao, Guanghui Yue 0001, Wei Zhou 0021, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao
ACM Multimedia9
2025 VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
abstract
The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and talks, have led to a dramatic increase in speech-rich video content. In contrast to other video types, such as surveillance footage, which typically contain abundant visual cues, speech-rich videos convey most of their meaningful information through the audio channel. This poses challenges for improving content consumption using existing visual-based video summarization, navigation, and exploration systems. In this paper, we present VisAug, a novel interactive system designed to enhance speech-rich video navigation and engagement by automatically generating informative and expressive visual augmentations based on the speech content of videos. Our findings suggest that this system has the potential to significantly enhance the consumption and engagement of information in an increasingly video-driven digital landscape.
Baoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang 0001, Fan Zhou 0001, Shujin Lin
ACM Multimedia5
2025 Part-aware distillation and aggregation network for human parsing
Yuntian Lai, Fan Zhou 0001, Zhuo Su 0001
Image Vis. Comput.3
2025 Single-View Clothed Human Reconstruction With Multi-View Consistency Representation
abstract
For single-view clothed human reconstruction, the fashionable PIFu-like framework depends on the pixel-aligned feature essentially, while this leads to depth ambiguity and inaccuracy of representation. Additionally, this task faces the inherent problem of the lack of invisible information. To solve these two problems, we propose depth-guided pixel-aligned feature and multi-view consistency prior to constrain representation learning of the single-view reconstruction task. The difference, between the depth values of points and estimated depth map, is used to filter pixel-aligned features. Thus, the image encoder can focus on capturing the feature of the visible part which is more effective in feature representation. The method introduces contrastive learning and masked autoencoder to achieve consistency of SMPL vertex features in each view which helps model to imagine invisible information. The experimental results show that the proposed method enables the feature to represent surface details more efficiently, thus achieves more reasonable and accurate representation learning. The qualitative and quantitative evaluations on the public and commercial datasets show that the proposed method can achieve better performance than previous implicit representation based methods.
Zhuo Su 0001, Yudi Tan, Zedan Zheng, Fan Zhou 0001, Baoquan Zhao
IEEE Trans. Vis. Comput. Graph.4
2025 Visual Boundary-Guided Pseudo-Labeling for Weakly Supervised 3D Point Cloud Segmentation in Indoor Environments
abstract
Accurate segmentation of 3D point clouds in indoor scenes remains a challenging task, often hindered by the labor-intensive nature of data annotation. While weakly supervised learning approaches have shown promise in leveraging partial annotations, they frequently struggle with imbalanced performance between foreground and background elements due to the complex structures and proximity of objects in indoor environments. To address this issue, we propose a novel foreground-aware label enhancement method utilizing visual boundary priors. Our approach projects 3D point clouds onto 2D planes and applies 2D image segmentation to generate pseudo-labels for foreground objects. These labels are subsequently back-projected into 3D space and used to train an initial segmentation model. We further refine this process by incorporating prior knowledge from projected images to filter the predicted labels, followed by model retraining. We introduce this technique as the Foreground Boundary Prior (FBP), a versatile, plug-and-play module designed to enhance various weakly supervised point cloud segmentation methods. We demonstrate the efficacy of our approach on the widely-used 2D-3D-Semantic dataset, employing both random-sample and bounding-box based weak labeling strategies. Our experimental results show significant improvements in segmentation performance across different architectural backbones, highlighting the method's effectiveness and portability.
Zhuo Su 0001, Yudi Tan, Boliang Guan, Fan Zhou 0001
IEEE Trans. Vis. Comput. Graph.5
2024 Improved Text-Driven Human Motion Generation via Out-of-Distribution Detection and Rectification
Yiyu Fu, Baoquan Zhao, Chenlei Lv, Guanghui Yue 0001, Ruomei Wang 0001, Fan Zhou 0001
CVM (1)6
2024 Clip-Medfake: Synthetic Data Augmentation With AI-Generated Content for Improved Medical Image Classification
abstract
Data augmentation is serving as a critical and fundamental technology to improve model generalization and performance in a wide spectrum of machine learning tasks. Despite the increasing interest in developing various pathways to artificially generate new data to reduce the overfitting issue during model training, enriching the diversity of training data in the field of medicine remains facing enormous challenges. By virtue of recent advancements in generative artificial intelligence, we present a novel data augmentation framework, CLIP-MedFake, to address the shortage of training data used in medical image classification. The proposed method first employs the Stable Diffusion model to generate new fake data based on a small amount of training data, and then adopts the paradigm of few-shot learning and uses the CLIP architecture as the backbone to pre-train the model with synthetic data and then fine-tune it with real medical images. Extensive experiment results on two publicly available datasets demonstrate the effectiveness of the proposed method in promoting medical image classification.
Honghui Chen, Baoquan Zhao, Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001
ICIP7
2024 Hierarchical Attention Feature Fusion and Refinement Network for Point Cloud Upsampling
abstract
This paper presents a novel hierarchical attention feature fusion and refinement network designed to address challenges in existing deep learning based point cloud upsampling methods. The network combines self-attention layers with a multi-level feature extraction architecture, effectively integrating local and global features, thereby enhancing the robustness and uniformity of the point cloud. Furthermore, a spatial refinement module is employed to predict the offset between the generated coarse dense point clouds and real point clouds, thereby enhancing consistency with the ground truth. Concurrently, a filter function is applied in the loss function to handle outliers of generated point clouds. Extensive experimental results across multiple datasets indicate that our method outperforms existing approaches.
Yaori Zhang, Shujin Lin, Fan Zhou 0001, Ruomei Wang 0001
ICME3
2024 Revitalizing Real Image Deraining via a Generic Paradigm towards Multiple Rainy Patterns
Xin Li 0175, Fan Zhou 0001, Yun Liang 0003, Zhuo Su 0001
IJCAI3
2024 Predicting Plain Text Imageability for Faithful Prompt-Conditional Image Generation
Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao
PRICAI (3)6
2024 Point Cloud Completion Method Assisted by Projected Image
abstract
Image-guided point cloud completion task aims to utilize image information to address the uncertainties in point cloud completion inference. Although acquiring 2D image data is relatively simpler than 3D data, it is still ineffective in scenarios with occlusions where image data cannot be reliably obtained as a reference. Therefore, we propose a point cloud completion model assisted by projected image data, which addresses the limitations of acquiring 2D images by constructing projected images of the point cloud. Extensive experiments demonstrate that our proposed method enhances the quality of point cloud completion and outperforms other advanced methods.
Shujin Lin, Zhaowen Li, Runxun Wu, Fan Zhou 0001
SMC4
2024 Video Q &A based on two-stage deep exploration of temporally-evolving features with enhanced cross-modal attention mechanism
Yuanmao Luo, Ruomei Wang 0001, Fan Zhou 0001
Neural Comput. Appl.4
2024 Advancing Real-World Image Dehazing: Perspective, Modules, and Training
abstract
Restoring high-quality images from degraded hazy observations is a fundamental and essential task in the field of computer vision. While deep models have achieved significant success with synthetic data, their effectiveness in real-world scenarios remains uncertain. To improve adaptability in real-world environments, we construct an entirely new computational framework by making efforts from three key aspects: imaging perspective, structural modules, and training strategies. To simulate the often-overlooked multiple degradation attributes found in real-world hazy images, we develop a new hazy imaging model that encapsulates multiple degraded factors, assisting in bridging the domain gap between synthetic and real-world image spaces. In contrast to existing approaches that primarily address the inverse imaging process, we design a new dehazing network following the "localization-and-removal" pipeline. The degradation localization module aims to assist in network capture discriminative haze-related feature information, and the degradation removal module focuses on eliminating dependencies between features by learning a weighting matrix of training samples, thereby avoiding spurious correlations of extracted features in existing deep methods. We also define a new Gaussian perceptual contrastive loss to further constrain the network to update in the direction of the natural dehazing. Regarding multiple full/no-reference image quality indicators and subjective visual effects on challenging RTTS, URHI, and Fattal real hazy datasets, the proposed method has superior performance and is better than the current state-of-the-art methods.
Long Ma 0002, Xiaozhe Meng, Fan Zhou 0001, Risheng Liu, Zhuo Su 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Bridging the Gap Between Haze Scenarios: A Unified Image Dehazing Model
abstract
In real-world scenarios, the haze presents diversity and complexity. However, current dehazing researches usually focus solely on specific categories or the removal of common white haze, frequently lacking the ability to adapt across various unknown haze types. In this study, our emphasis is on constructing a model that shows excellent adaptability across diverse haze conditions. Unlike approaches that solely rely on network structure design to enhance model adaptability, we comprehensively improve dehazing model adaptability from three key aspects: constructing the multitype haze dataset from designed haze degradation models, designing the network architecture, and formulating training strategies suitable for cross-scene generalization. Firstly, to meet the diverse haze training data requirements, we design a multitype haze degradation model to generate more realistic pairs of hazy images. Secondly, to ensure thorough haze removal and natural restoration of texture details in the recovered images, we construct a dual-branch ensemble network framework by leveraging pre-trained clear image prior features and the characteristics of 2D discrete wavelet priors. Finally, to further enhance the adaptability for removing various types of haze, we employ a sample reweighting decorrelation strategy during the network training phase to eliminate dependencies between haze and haze-free background features. Through extensive experiments, our approach shows remarkable performance across diverse haze scenarios. Our method not only outperforms state-of-the-art scene-specific dehazing methods in typical scenarios like daytime and nighttime, but it also excels in handling challenging scenarios such as dusty conditions, and color haze. See more resultshttps://github.com/fyxnl/Image-dehazing-CGID.
Zhuo Su 0001, Long Ma 0002, Xin Li 0175, Risheng Liu, Fan Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Subtask Prior-Driven Optimized Mechanism on Joint Video Moment Retrieval and Highlight Detection
abstract
Joint video moment retrieval and highlight detection is an emerging and challenging research task. It requires the generation of robust joint task features to satisfy the demands of video moment retrieval and video highlight detection. Moreover, it involves the interaction of multiple modalities. Presently, methods typically focus on the design of distinct enhancement modules and the addition of supplementary input data to improve the solution for joint video moment retrieval and highlight detection. However, they overlook subtask interference during joint training. Joint task learning leverages the correlations and complementarities between tasks, yet it also introduces task interference arising from the differences between tasks. In order to address task interference, we proposes a subtask prior-driven optimized mechanism. The mechanism consists of two stages. In the free stage, we train subtask model to get subtask prior features. In the constrained stage, the joint task model is constrained by the subtask. Besides, we propose a cross adaptive-gated mechanism. It addresses the issue of information loss in cross-modal fusion and filters out redundant information by conducting cross-modal interaction during feature compression and an adaptive gating process. Extensive experimental results exhibit the effectiveness of the subtask prior-driven optimized mechanism and the cross adaptive-gated transformer in joint video moment retrieval and highlight detection.
Ruomei Wang 0001, Fan Zhou 0001, Zhuo Su 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 DMAP: Decoupling-Driven Multi-Level Attribute Parsing for Interpretable Outfit Collocation
abstract
Outfit collocation requires considering the interrelationship and adaptability among the attributes of component items. However, with the numerous and diverse attributes of fashion items, accurately capturing attribute features and modeling the complex relationships between attributes become the key challenges. To address these challenges, we propose a novel scheme Decoupling-driven Multi-level Attribute Parsing for interpretable outfit collocation. First, we decouple a series of attribute features from the item's visual feature by fully supervised, which can improve the robustness of the model in processing both relevant and irrelevant attributes of items. Furthermore, employing a deep deconvolution neural network with attention mechanisms to reconstruct the decoupled attribute features into a visual image that is close to the original item image. It ensures all attribute features can be combined to contain complete item information. Next, graph attention networks are constructed to parse multi-level attribute compatibility relationships from three perspectives: intra-attribute, inter-attribute, and item integration relationships. Finally, we use multi-layer perceptrons to fuse the score distributions of the three and output the outfit compatibility score. Experiments conducted on the IQON3000 dataset demonstrate that our model outperforms existing state-of-the-art methods and exhibits good interpretability.
Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001, Ge Lin 0002
IEEE Trans. Multim.5
2023 High Fidelity Virtual Try-On via Dual Branch Bottleneck Transformer
Xiuxiang Li, Guifeng Zheng, Fan Zhou 0001, Zhuo Su 0001, Ge Lin 0002
ICIG (1)3
2023 MIM: Lightweight Multi-Modal Interaction Model for Joint Video Moment Retrieval and Highlight Detection
abstract
Joint video moment retrieval and highlight detection aims to find the relevant moments and highlight clips in a video with natural language. It is an emerging task though its individual problems have been studied for a while. The current methods utilize transformer to interact between modals, which leads to a huge cost of parameters and computation in spite of great performance. To address this problem, we present a cross-modal attention mechanism to capture related features from different modalities in a few-parameter way. Furthermore, a lightweight multi-modal interaction model (MIM) is proposed to solve video moment retrieval and highlight detection jointly. In the case of greatly reducing the number of parameters, we achieve competitive performance and faster convergence speed compared to previous method. Extensive experiments on four datasets demonstrate the effectiveness of our method.
Shujin Lin, Fan Zhou 0001, Ruomei Wang 0001
ICME4
2023 Image-Guided Point Cloud Completion with Multi-modal Fusion Transformers
abstract
The task of image-guided point cloud completion aims to leverage information from images to address uncertainty issues in the completion inference of point clouds. The key challenge in this setting lies in how to effectively combine features extracted from both modalities. Due to the large domain discrepancy between the image and point cloud, existing methods that use cross-modal attention to directly fuse features have increased attention on redundant information and noise from different modalities, resulting in poor feature fusion performance. Hence, by introducing multi-modal fusion transformers that use bottleneck tokens, we enabled point cloud feature to learn image feature through information bridges, leading to improved point cloud completion performance. Our method can not only benefit from RGB images, but also from sketches with less feature information but more emphasis on edge information. Extensive experiments demonstrate that our proposed method enhances the quality of point cloud completion and outperforms other state-of-the-art methods.
Zhaowen Li, Shujin Lin, Fan Zhou 0001
SMC3
2023 OaIF: Occlusion-Aware Implicit Function for Clothed Human Re-construction
abstract
Abstract Clothed human re‐construction from a monocular image is challenging due to occlusion, depth‐ambiguity and variations of body poses. Recently, shape representation based on an implicit function, compared to explicit representation such as mesh and voxel, is more capable with complex topology of clothed human. This is mainly achieved by using pixel‐aligned features, facilitating implicit function to capture local details. But such methods utilize an identical feature map for all sampled points to get local features, making their models occlusion‐agnostic in the encoding stage. The decoder, as implicit function, only maps features and does not take occlusion into account explicitly. Thus, these methods fail to generalize well in poses with severe self‐occlusion. To address this, we present OaIF to encode local features conditioned in visibility of SMPL vertices. OaIF projects SMPL vertices onto image plane to obtain image features masked by visibility. Vertices features integrated with geometry information of mesh are then feed into a GAT network to encode jointly. We query hybrid features and occlusion factors for points through cross attention and learn occupancy fields for clothed human. The experiments demonstrate that OaIF achieves more robust and accurate re‐construction than the state of the art on both public datasets and wild images.
Yudi Tan, Boliang Guan, Fan Zhou 0001, Zhuo Su 0001
Comput. Graph. Forum3
2023 Boundary-guided part reasoning network for human parsing
Zhuo Su 0001, Huiqiang Guan, Yuntian Lai, Fan Zhou 0001, Yun Liang 0003
Neurocomputing4
2023 Towards real-world haze removal with uncorrelated graph model
Xiaozhe Meng, Fan Zhou 0001, Yun Liang 0003, Zhuo Su 0001
J. Vis. Commun. Image Represent.3
2023 Attention-adaptive multi-scale feature aggregation dehazing network
Zhuo Su 0001, Fan Zhou 0001
J. Vis. Commun. Image Represent.4
2023 Transfer Learning Dehazing Network by Gaussian Process Mapping
Xiaozhe Meng, Zhuo Su 0001, Fan Zhou 0001
Signal Process. Image Commun.4
2023 Real-World Non-Homogeneous Haze Removal by Sliding Self-Attention Wavelet Network
abstract
In complex natural haze scenes, image haze removal still faces significant challenges in removing non-homogeneous and dense haze. The double complexity of haze distribution, on the one hand, is reflected in the interference of haze to the global image information, and on the other hand, it is reflected in the imbalance of image brightness and color caused by random haze distribution. In natural scenes with prominent edge and texture features, the above problems may cause severe degradation of image quality and performance of various tasks. Numerous studies on network learning show that the effect of haze removal is closely related to haze feature expression. Therefore, to improve the performance of dehazing, this paper proposes a sliding self-attention wavelet network. Specifically, we first design a sliding self-attention module to identify haze regions in images and capture rich haze-related feature information. Then, considering the uneven distribution of haze in images, discrete wavelet transform (DWT) and inverse transform (IDWT) are used for constructing a hierarchical encoder-decoder structure, which can fully use the multi-resolution characteristics of DWT, locally decompose feature maps of different scales, extract low and high-frequency information, and then gradually recover sharp edges and precise texture details from hazy images. Finally, to enable the proposed network to generate more realistic haze-free images on different complex haze scenes, we develop a DWT-based adversarial loss function to constrain the low and high-frequency components of generated images closer to the corresponding clear images. Experimental results on the relevant public benchmark datasets show that the proposed algorithm achieves favorable dehazing performance.
Xiaozhe Meng, Fan Zhou 0001, Weisi Lin, Zhuo Su 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 ERM: Energy-Based Refined-Attention Mechanism for Video Question Answering
abstract
Spatiotemporal attention learning remains a challenging video question answering (VideoQA) task as it requires a sufficient understanding of cross-modal spatiotemporal information. Existing methods usually leverage different cross-modal attention mechanisms to reveal potential associations between video and question. While these methods effectively remove irrelevant information from the spatiotemporal attention, they ignore the pseudo-related information within the cross-modal interaction attention. To address this problem, we proposed a novel energy-based refined-attention mechanism (ERM). ERM leverages the significant difference distribution as a discriminative criterion derived from question-guided cross-modal interaction information to determine question-related and question-irrelated cross-modal interaction information. The specific method is to measure the linear separability between the target neuron and other neurons in the neural network to confirm the importance of neurons. In addition, to solve the statistical bias caused by the differences between different modes in video tasks, the ERM proposed in this paper has learnable parameters. The correlation between different modes can be learned adaptively through learnable parameters. The advantages of the proposed ERM are that it is more flexible and modular while remaining lightweight. With the help of the ERM, we construct a lightweight VideoQA model that efficiently integrates the cross-modal feature representations in an energy-based manner. To evaluate the effectiveness of our method, we carried out extensive experiments on five publicly available datasets and compared them with state-of-the-art VideoQA methods. The experiment results demonstrate that our method brings a noticeable performance improvement compared to state-of-the-art VideoQA methods. ERM can be flexibly integrated into different VideoQA methods to improve their Q&A performance.
Ruomei Wang 0001, Fan Zhou 0001, Yuanmao Luo
IEEE Trans. Circuits Syst. Video Technol.3
2022 Multi-Target Domain Transfer Detector on Bad Weather Conditions
abstract
With the popularization of object detection in fields such as autonomous driving and surveillance, object detection networks are required to be used in more changeable scenarios. However, a multi-target domain adaptive object detection network is a challenging problem. Due to the lack of relevant datasets and the limitation of traditional domain adaptive methods, there is a lack of pertinent research in this field. Traditional single-target domain adaptive methods are often ineffective when used in multi-target domain adaptive problems. To allow the object detection network to be applied to multiple target domains, we propose the “domain transfer module” and the “multi-scale hybrid attention domain alignment module”. At the same time, we synthesized the “BlendedCityspce” dataset for training and testing of a multi-target domain adaptive target object network. The network provided in this paper has a good performance in the multi-target domain adaptation and reaches the state-of-the-art in the single-target domain adaptation. You can find our source code in https://github.com/Chonghuan-Liu/Multi-Target-Domain-Adaptative-Detector.
Chonghuan Liu, Zhuo Su 0001, Fan Zhou 0001
ICME4
2022 Unsupervised Domain Adaptation Image Dehazing with Contrastive Nearest-Farthest Subspace Distance
abstract
Haze can cause significant changes in the data domain of the image. Due to domain-shift problem, the model trained on the synthetic haze images has a weak or even invalid dehazing ef-fect on the real data. This paper proposes an unsupervised domain adaptation dehazing method based on the nearest-farthest subspace distance in response to this problem. For the middle convolved features, the matrix singular value de-composition is used to obtain the basis vectors of the support subspace. Narrowing the angle between the basis vectors of natural and synthetic haze domains could reduce the differ-ence between the data domains. In subspace, a newly defined distance with a new penalty term named subspace measure distance is employed to constrain the model. Furthermore, to avoid search blindness of the proposed method in the real data training phase, we propose the nearest-farthest subspace distance inspired by contrastive learning. In addition, we use a new training strategy. Finally, the experimental tests on the synthetic and real images prove the effectiveness of the pro-posed method.
Xiaozhe Meng, Zhuo Su 0001, Fan Zhou 0001
ICME4
2022 Cross-Rolling Attention Network for Fashion Landmark Detection
abstract
The variety and deformable characteristics of fashions are the challenges in improving the accuracy of fashion landmark detection. Considering the delicate characteristics, this paper proposes a new method for fashion landmark detection based on the cross-rolling attention module and the feature pyramid network. It imports the cross-rolling attention mechanism into the multi-scale structure to extract and fuse different features with rich context, which is beneficial to capture fine-grained fashion keypoints more accurately and improve detection performance. While constructing the network, we also focus on controlling the network scale to avoid a large increase in parameters and longtime training. Finally, we perform some experiments suggesting that the proposed method achieves superior performance on both DeepFashion and FLD datasets. It demonstrates the effectiveness of the cross-rolling attention module in the proposed method.
Fan Zhou 0001, Zhuo Su 0001
ICPR2
2022 NewsThumbnail: Automatic Generation of News Video Thumbnail
abstract
Reading news is an important way for people to obtain information. People can quickly sort out the context of events through a short news video. However, there are numerous news generated around the world every day. It’s challenging to locate the interesting video. Thumbnails are often used as video covers and play an important role in displaying video content and driving views. Video owners can choose from individual images or elaborate thumbnails to upload to the site. But manually selecting from a large number of frames is time-consuming, and customizing thumbnails requires a high degree of expertise. Therefore, this paper proposes an automatic generation method of news video thumbnail, which can screen out semantically similar contents according to user query and combine them into a thumbnail. In order to facilitate the screening of graphic materials, we also propose a video content structuring method based on multiple cues, which can accurately segment the video into theme units. At the same time, we designed a visual system to display thumbnails and designed a user survey to investigate the performance of this method in news retrieval and understanding. Compared with peer methods, the thumbnails generated by our method can help users better understand the video content and locate the videos they are interested in.
Shujin Lin, Fan Zhou 0001, Ruomei Wang 0001
SMC3
2022 Temporal-aware Mechanism with Bidirectional Complementarity for Video Q&A
abstract
Video question answering (Video Q&A) is a challenging task as it requires a sufficient understanding of the video and question information. Video is composed of frame sequence, which contains multi-scale temporal relationships and corresponding contextual information. A model competently tackle Video Q&A task that needs to be able to: 1) construct long-term and neighborhood dependencies in frame sequences to extract global and local contextual features that can reflect multi-scale temporal dependencies, and deduce the temporal-aware refined features, and 2) identify static and dynamic features from pertinent moments of a video, while filtering away question-irrelated dependencies of feature sequences, to yield the most precise and reasonable temporal-aware overall contextual features. In response to the above requirements, we propose a novel Video Q&A mechanism which consists of Bidirectional Complementary Attention(BCA) module and Adaptive Temporal-aware(ATA) module. Bidirectional complementary attention module stacks multi-head self-attention layer and convolutional layer in different orders to designed two kinds of attention units, which is able to make bidirectional multi-step reasoning based on complete global information and accurate local information to obtain temporal-aware refined features. Adaptive temporal-aware module is used to filter away question-irrelated dependencies in the feature sequence to yield the most precise and reasonable temporal-aware overall contextual features. Comprehensive comparative experiments are conducted on publicly available benchmark datasets. An extended ablation study is further conducted to show the usefulness of each module of the solution in acquiring its computational Q&A capabilities.
Yuanmao Luo, Ruomei Wang 0001, Fan Zhou 0001, Shujin Lin
SMC4
2022 Test-Driven Feature Extraction of Web Components
Yan-Cheng Chen, Xiangping Chen, Xiaohong Shi, Fan Zhou 0001
J. Comput. Sci. Technol.5
2022 Multi-receptive Field Aggregation Network for single image deraining
Songliang Liang, Xiaozhe Meng, Zhuo Su 0001, Fan Zhou 0001
J. Vis. Commun. Image Represent.4
2022 Popularity-Guided Cost Optimization for Live Streaming in Mobile Edge Computing
abstract
Live streaming service usually delivers the content in mobile edge computing (MEC) to reduce the network latency and save the backhaul capacity. Considering the limited resources, it is necessary that MEC servers collaborate with each other and form an overlay to realize more efficient delivery. The critical challenge is how to optimize the topology among the servers and allocate the link capacity so that the cost will be lower with delay constraints. Previous approaches rarely consider server collaborations for live streaming service, and the scheduling delay is usually ignored in MEC, leading to suboptimal performances. In this paper, we propose a popularity‐guided overlay model which takes the scheduling delay into consideration and utilizes MEC collaboration to achieve efficient live streaming service. The links and servers are shared among all channel streams and each stream is pushed from cloud servers to MEC servers via the trees. Considering the optimization problem is NP‐hard, we propose an effective optimization framework called cost optimization for live streaming (COLS) to predict the channel popularity by a LSTM model with multiscale input data. Finally, we compute topology graph by greedy scheme and allocate the capacity with convex programming. Experimental results show that the proposed approach achieves higher prediction accuracy, reducing the capacity cost by more than 40% with an acceptable delay compared with state‐of‐the‐art schemes.
Tao He 0012, Kunxin Zhu, Ruomei Wang 0001, Fan Zhou 0001
Wirel. Commun. Mob. Comput.5
2021 Fusing Temporally Distributed Multi-Modal Semantic Clues for Video Question Answering
abstract
Video Question Answering (VideoQA) is an intriguing topic, attracting increasing interest among the broad AI community. Yet videoQA is a difficult task. An algorithm competently tackle this task that needs to be able to: 1) extract rich semantics supplied in each modality of a video and incorporate them across modalities, and 2) identify and integrate such multimodal semantics from pertinent moments of a video, which may or may not be temporally adjacent or nearby, while filtering away irrelevant or even detractive portions of the video, to yield the most precise and sensible semantic context for executing the QA task. In response to the above requirements, a novel deep VideoQA solution is proposed in this paper, which comprises a multi-modal semantic clue extraction module, driven by a series of deep networks, each dedicated to digesting signals of a distinct modality type, to develop the first algorithmic QA capability, and a multi-modal temporal QA module empowered by a deep graph attention network to build the second algorithmic QA capability. Comprehensive experiments are conducted on publicly available benchmark data to validate advantages of the new solution in the end.
Ruomei Wang 0001, Songhua Xu, Fan Zhou 0001
ICME4
2021 Relation-Aware Attribute Network for Fine-Grained Clothing Recognition
Mingjian Yang, Zhuo Su 0001, Fan Zhou 0001
ICONIP (6)4
2021 News2Mapping: A news events correlation model for news videos
abstract
News video is an important way of news communication, and people can easily get news from all over the world through the Internet. However, it lacks in the organization of news video content about temporality, presentation and relevance and fails to express the correlation between news events. In this paper, we present an event correlation model of news videos to organize news content. News content is clustered through topics, and the relationship between events is illustrated through relationship mappings. To achieve this purpose, an XLNet-based language model is presented to extract news keywords and their relationships. The clustering algorithm is designed to obtain news event topic clustering and named entity clustering. At the same time, we build the relationship mappings in news events to visualize the correlation between news events better. The user study is also designed to investigate the performance of our method in news reading and understanding. Compared with peer methods, the news information organized by our method achieves a higher user satisfaction level.
Mingjie Zhou, Ruomei Wang 0001, Shujin Lin, Fan Zhou 0001, Shirou Ou
SMC4
2021 LGCPNet : Local-global combined point-based network for shape segmentation
Boliang Guan, Fan Zhou 0001, Shujin Lin, Ruomei Wang 0001
Comput. Graph.3
2021 MVSN: A Multi-view stack network for human parsing
Zhuo Su 0001, Minshi Chen, Enbo Huang, Ge Lin 0002, Fan Zhou 0001
Neurocomputing5
2020 Tao: A Trilateral Awareness Operation For Human Parsing
abstract
Human parsing, which segments a human image into semantic regions, is a fundamental task in human-centric analysis. Recently, numerous human parsing approaches based on convolutional neural networks (CNNs) have made significant progress. However, these models still suffer the high time-complexity and poor efficiency issues, since they always introduce too many and complex hypothetical priors. Therefore, in this paper, we attempt to solve above issues by mining the deep information hidden in data, such as the scale, spatial information and statistics of the image descriptors, etc. By this way, we propose an effective structure, named the trilateral awareness operation (TAO), which can boost the scale, spatial and fine-grained awareness of a CNN. In practice, it is a simple yet effective architecture without using bells and whistles (such as human pose and edge information), to address human parsing tasks. Comprehensive experiments and corresponding results on three public datasets have demonstrated that the proposed TAO is superior to the state-of-the-art methods.
Enbo Huang, Zhuo Su 0001, Fan Zhou 0001
ICME3
2020 Example-based web page recoloring method
Zhihao Zang, Xiangping Chen, Fan Zhou 0001
Frontiers Comput. Sci.4
2020 AutoWPR: An Automatic Web Page Recoloring Method
abstract
The color design is one of the important parts of GUI development. To gain an attractive color scheme, designers often seek inspiration from examples. However, transferring an example’s colors to a target web page is time-consuming and tedious. In this paper, we propose a method named AutoWPR to reuse the example web page’s colors for recoloring a web page. To preserve the semantic relations of web elements, we propose a clustering algorithm to group the related elements into a cluster. In order to make the recoloring result have similar color distributions to the example, we use the Random–Forest regression to learn human’s mappings and propose a top-down matching algorithm to generate a mapping between two web pages’ clusters. Then AutoWPR recolors the element with the matching element’s colors. We designed several experiments to evaluate the correctness of the clustering and matching algorithm. We also conducted some qualitative and quantitative experiments to evaluate the effectiveness of our results in helping recoloring. The results show that our method can generate a human-like recoloring result and help novice developers reuse the reference web page’s colors conveniently.
Xiangping Chen, Fan Zhou 0001
Int. J. Softw. Eng. Knowl. Eng.3
2020 Learning rebalanced human parsing model from imbalanced datasets
Enbo Huang, Zhuo Su 0001, Fan Zhou 0001, Ruomei Wang 0001
Image Vis. Comput.3
2020 Voxel-based quadrilateral mesh generation from point cloud
Boliang Guan, Shujin Lin, Ruomei Wang 0001, Fan Zhou 0001, Yongchuan Zheng
Multim. Tools Appl.4
2019 Learning mean progressive scattering using binomial truncated loss for image dehazing
abstract
In this study, the authors propose a novel progressive dehazing network to address the single image haze removal problem based on a new mean progressive scattering model. Different from methods that learn atmosphere light and transmission maps with different networks, these two variables are optimised in a unified network. Following the methodology of traditional prior‐based methods that estimate a coarse transmission map first, a progressive refinement branch in the decoder has been designed to restore the fine‐scale transmission map. To improve the prediction accuracy of the transmission map, a novel binomial truncated loss that assigns weights to error values according to the probabilities of error occurrences has been proposed. An ablation study is conducted to verify the effectiveness of the components in the proposed method. Experiments in the synthetic datasets and real images demonstrate that the proposed method outperforms other state‐of‐the‐art methods.
Bin Qiu, Xiwen Liang, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001
IET Image Process.5
2019 Conditional progressive network for clothing parsing
abstract
Clothing parsing is significant to many clothing applications. Recently, a lot of clothing parsing methods have been presented, which explore the innovation of the parsing pipeline or try to find more specific prior information. Although these methods perform well in some benchmarks, a few challenging problems have not been solved yet, such as the complicated mutual interference among labels. In this study, the authors propose a Conditional Progressive Network to parse clothing in different scales and prevent the mutual interference among labels. The authors’ solution consists of three sub‐networks, including Conditional Parsing Network (CPN), Pose Estimation Network (PEN) and Label Transform Network (LTN). Specifically, the CPN module generates the intermediate parsing result in the form of the multiple progressive stages, which combines with the previous outputs in each stage and the specific prior conditions. The PEN module provides a series of heat maps about the human pose information. The LTN module suppresses the redundant labels to avoid the mutual interference among labels. They demonstrate their solution in parsing the fashion clothing cases on the ATR and the Fashion dataset. In their experiments, their method obtains a better performance than the state‐of‐the‐art methods.
Zhuo Su 0001, Jiaming Guo, Gengwei Zhang, Xianghui Luo, Ruomei Wang 0001, Fan Zhou 0001
IET Image Process.6
2019 Automatically detecting the scopes of source code comments
abstract
Comments convey useful information about the system functionalities and many methods for software engineering tasks take comments as an important source for many software engineering tasks such as code semantic analysis, code reuse and so on. However, unlike structural doc comments, it is challenging to identify the relationship between the functional semantics of the code and its corresponding textual descriptions nested inside the code and apply it to automatic analyzing and mining approaches in software engineering tasks efficiently. In this paper, we propose a general method for the detection of source code comment scopes. Based on machine learning, our method utilized features of code snippets and comments to detect the scopes of source code comments automatically in Java programs. On the dataset of comment-statement pairs from 4 popular open source projects, our method achieved a high accuracy of 81.45% in detecting the scopes of comments. Furthermore, the results demonstrated the feasibility and effectiveness of our comment scope detection method on new projects. Moreover, our method was applied to two specific software engineering tasks in our studies: analyzing software repositories for outdated comment detection and mining software repositories for comment generation. As a general approach, our method provided a solution to comment-code mapping. It improved the performance of baseline methods in both tasks, which demonstrated that our method is conducive to automatic analyzing and mining approaches on software repositories.
Huanchao Chen, Yuan Huang 0002, Xiangping Chen, Fan Zhou 0001
J. Syst. Softw.5
2018 Automatically Detecting the Scopes of Source Code Comments
abstract
Comments which are an integral part of software development improve program comprehension and software maintainability. They convey useful information about the system functionalities and many text retrieval methods for software engineering tasks take comments as an important source for code semantic analysis. However, it is challenging to identify the relationship between the functional semantics of the code and its corresponding textual descriptions and apply it to automatic mining approaches in software engineering efficiently. In this paper, we use machine learning which utilizes features of code snippets and comments to detect the scopes of source code comments automatically in Java programs. Based on the dataset of comment-statement pairs from 4 popular open source projects, our method achieved a high accuracy of 81.15% in detecting the scopes of comments. Furthermore, the experimental results demonstrated the feasibility and effectiveness of our comment scope detection method.
Huanchao Chen, Xiangping Chen, Fan Zhou 0001
COMPSAC (1)4
2018 Automatic Detection of Outdated Comments During Code Changes
abstract
Comments are used as standard practice in software development to increase the readability of code and to express programmers' intentions in a more explicit manner. Nevertheless, keeping comments up-to-date is often neglected for programmers. In this paper, we proposed a machine learning based method for detecting the comments that should be changed during code changes. We utilized 64 features, taking the code before and after changes, comments and the relationship between the code and comments into account. Experimental results show that 74.6% of outdated comments can be detected using our method, and 77.2% of our detected outdated comments are real comments which require to be updated. In addition, the experimental results indicate that our model can help developers to discover outdated comments in historical versions of existing projects.
Huanchao Chen, Xiangping Chen, Fan Zhou 0001
COMPSAC (1)5
2017 A Data-Driven Approach for Sketch-Based 3D Shape Retrieval via Similar Drawing-Style Recommendation
abstract
Abstract Sketching is a simple and natural way of expression and communication for humans. For this reason, it gains increasing popularity in human computer interaction, with the emergence of multitouch tablets and styluses. In recent years, sketch‐based interactive methods are widely used in many retrieval systems. In particular, a variety of sketch‐based 3D model retrieval works have been presented. However, almost all of these works focus on directly matching sketches with the projection views of 3D models, and they suffer from the large differences between the sketch drawing and the views of 3D models, leading to unsatisfying retrieval results. Therefore, in this paper, during the matching procedure in the retrieval, we propose to match the sketch with each 3D model from historical users instead of projection views. Yet since the sketches between the current user and the historical users can have big difference, we also aim to handle users' personalized deviations and differences. To this end, we leverage recommendation algorithms to estimate the drawing style characteristic similarity between the current user and historical users. Experimental results on the Large Scale Sketch Track Benchmark(SHREC14LSSTB) demonstrate that our method outperforms several state‐of‐the‐art methods.
Fei Wang 0056, Shujin Lin, Hefeng Wu, Ruomei Wang 0001, Fan Zhou 0001
Comput. Graph. Forum6
2016 Retiling scheme: a novel approach of direct anisotropic quad-dominant remeshing
Ruomei Wang 0001, Fan Zhou 0001
Vis. Comput.2
2014 An adaptive neural fuzzy network clothing comfort evaluation model and application in digital home
Ruomei Wang 0001, Fan Zhou 0001, Daiguo Deng
Multim. Tools Appl.3
2013 Ontology Based Automatic Image Annotation Using Multi-class SVM
abstract
Image annotation is usually formed as a multiclass classification problem. Traditional methods learn the co-occurrence of keywords and images while they ignore the correlation between keywords, which turned out to be one of the reasons causing poor experiment results. In this paper, we propose an automatic image annotation approach by using multiclass SVM with ontology to achieve a higher accuracy. In our paper, we choose semantic dictionary Word Net in which hierarchy defined words are derived from the text ontology to calculate the correlations between keywords. Specifically, we use Bags of Visual Words model to present the image visual feature and apply a mixed kernel in multiclass SVM. Finally, we combine the probability outputs to get the final results. Compared to other state-of-the-art multiclass classification methods, our approach tested in typical Corel dataset maintain a high level of accuracy in classification.
Zhenzhen Wei, Fan Zhou 0001
ICIG3
2012 Image Resizing Based on Geometry Preservation with Seam Carving
abstract
When an image or a video is transformed to an aspect ratio deferent from its original size, information lost is inevitable no matter what method is used, thus, how to keep the most attractive contents and minimize the visual distortion during the resizing process is the key issue. To address this problem, this paper proposes an object geometry preservation method based on the seam carving method. We first define a framework that measures the importance of geometry feature in the source material, then a new energy function is presented with object geometry constraint, according to the new energy function, an optimized seam carving method is used to minimize distortion while resizing the source material. The experiment results show that our method is better to transform a variety of source images to a different display size than conventional resizing methods.
Fan Zhou 0001, Ruomei Wang 0001, Yun Liang 0003
TrustCom1