Baoquan Zhao

dblp:142/7378 · DBLP profile ↗
← Back
45ranked-venue papers
5as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-author · 30 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Training for Identity, Inference for Controllability: A Unified Approach to Tuning-Free Face Personalization
abstract
Tuning-free face personalization methods have developed along two distinct paradigms: text embedding approaches that map facial features into the text embedding space, and adapter-based methods that inject features through auxiliary cross-attention layers. While both paradigms have shown promise, existing methods struggle to simultaneously achieve high identity fidelity and flexible text controllability. We introduce UniID, a unified tuning-free framework that synergistically integrates both paradigms. Our key insight is that when merging these approaches, they should mutually reinforce only identity-relevant information while preserving the original diffusion prior for non-identity attributes. We realize this through a principled training-inference strategy: during training, we employ an identity-focused learning scheme that guides both branches to capture identity features exclusively; at inference, we introduce a normalized rescaling mechanism that recovers the text controllability of the base diffusion model while enabling complementary identity signals to enhance each other. This principled design enables UniID to achieve high-fidelity face personalization with flexible text controllability. Extensive experiments against six state-of-the-art methods demonstrate that UniID achieves superior performance in both identity preservation and text controllability.
Lianyu Pang, Qiping Wang 0004, Baoquan Zhao, Zhenguo Yang, Qing Li 0001, Xudong Mao
ICMR4
2026 V-HOI: Velocity-Aware Human-Object Interaction Generation
Honghui Chen, Fan Zhou 0001, Ruomei Wang 0001, Baoquan Zhao
MMM (2)4
2026 NEGS-Avatar: Normal Embedded Gaussians for 2D avatar from monocular video
Zedan Zheng, Yudi Tan, Zhuo Su 0001, Fan Zhou 0001, Baoquan Zhao
Comput. Graph.5
2026 Continual test-time adaptation for object detection with adaptive monitoring and randomized restoration
Shilei Cao 0005, Juepeng Zheng, Baoquan Zhao, Runmin Dong, Haohuan Fu
Expert Syst. Appl.4
2026 Cross-modal attention fusion of RGB and skeleton for multimodal-driven video anomaly detection
Boan Chen, Weide Liu, Jinmei Liu, Baoquan Zhao, Yang Liu 0246
Pattern Recognit.5
2026 LVAD: A realistic data synthesis strategy and coarse-to-fine framework for low-light video anomaly detection
Linxuan Han, Fan Zhou 0001, Baoquan Zhao
Pattern Recognit.3
2026 Part-Level Semantic Fusion for Sketch-Based 3D Voxel Reconstruction
abstract
Reconstructing 3D shapes from monocular freehand sketches is challenging due to their fragmented structures, varying line thickness, and discontinuity. These characteristics cause ambiguity, making it difficult for existing methods to extract sufficient geometric feature information to distinguish subtle shape variations and internal details of the object contours depicted in the sketches, resulting in poor overall reconstruction quality. To address these challenges, we introduce the Part-level Semantic Fusion (PSFusion) module, which combines local units of image features with global feature guidance to enhance the representation of complex geometric structures and subtle contour variations. This approach reduces ambiguity in complex edges and local details, improving shape preservation and edge contour accuracy. Additionally, we propose the coarse-to-fine 3D Decoder consisting of one 3D convolutional network as a coarse-grained regressor and one Shuffle-UNet-based fine-grained refiner, to capture feature dependencies across spatial and channel dimensions. Shuffle operations facilitate information exchange among sub-features, enhancing structural differentiation and cross-feature dependency modelling. This significantly improves the handling of subtle textures and complex intersection boundaries. Extensive experiments on three public benchmarks show that our method outperforms baseline approaches, as demonstrated by both quantitative and qualitative results.
Fei Wang 0056, Yanlong Pan, Junkun Jiang, Dazhi Jiang, Baoquan Zhao
IEEE Trans. Circuits Syst. Video Technol.5
2026 Expressive Human Volumetric Video Generation With Rich Text
abstract
Plain text has become the dominant interactive interface for text-driven human volumetric video generation. However, its limited customization options hinder users from expressing motion effects with accuracy. For example, plain text struggles to specify continuous variables such as motion amplitude, speed, and joint trajectories with precision, and it fails to convey stylized motion characteristics. Additionally, crafting detailed textual prompts for complex motion sequences is cumbersome, while excessively long prompts strain text encoders. To address these limitations, we propose a rich text-based framework that supports font styles, sizes, and trajectory sketching. By extracting motion-related attributes from rich text, our method enables fine-grained control over motion styles, precise speed regulation, and accurate joint trajectory manipulation. These capabilities are realized through gradient-guided noise editing and ControlNet-based motion optimization, which operate within the latent motion diffusion process. Specifically, we design a unified gradient-guided adaptation mechanism to ensure that the generated motion video adheres strictly to the specified constraints. Furthermore, we introduce realism-oriented optimization for stylistic and joint-level control, refining motion synthesis at a granular level to produce smoother, more natural movements. We present multiple comparative evaluations showcasing volumetric video generation from both rich text and plain text. Through quantitative analysis, we demonstrate that our method surpasses strong plain-text baselines, producing expressive, customizable human volumetric motion videos.
Guanghui Yue 0001, Wei Zhou 0021, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao
IEEE Trans. Circuits Syst. Video Technol.6
2026 Visual-Guided Long Temporal Context Learning Network for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised video anomaly detection (WVAD) aims to locate events or behaviors that deviate from normal patterns in untrimmed videos using video-level labels. Recent studies typically utilize supplementary modalities to assist anomaly detection. However, these methods suffer from two main issues: (1) The limitations of long-duration anomaly event temporal modeling. The model struggles to consistently maintain key information, resulting in the forgetting phenomenon, which affects the tracking of the event’s overall dynamic evolution and complicates anomaly event analysis and understanding. (2) The multi-modal fusion strategy is insufficient, particularly when there is temporal inconsistency between visual and audio information, causing the model to overlook key information, directly affecting the accurate detection and recognition of anomalous events. To address these issues, we propose a visual-guided long-term temporal context learning network (LTCLNet). The network consists of three key components: a cross-modal interaction module, a multi-modal fusion module, and a visual-guided parameter optimization strategy. First, to address the forgetting issue in long-duration anomaly detection, we designed a cross-modal interaction module. The key part of this module is the establishment of a cross-matrix mechanism. This mechanism achieves bidirectional temporal guidance across modalities. It allows the temporal modeling of each modality to dynamically integrate information from the other modality. This enables the model to continuously track the dynamic evolution of the event. The tracking is facilitated through shared temporal information between the visual and audio modalities. Secondly, to fully exploit the complementary characteristics between different modalities, we introduced a novel temporal reversal integration method in the multi-modal fusion module. This method reverses the feature sequences of each modality to enhance the model’s perception of temporal dynamic changes. By fusing the modality features before and after reversal, the shared temporal structure between modalities is strengthened, improving the model’s ability to capture anomalous information. Additionally, our proposed visual-guided parameter optimization strategy trains a parallel visual modality network as a semantic anchor, ensuring that the model stays aligned with a semantically stable and structurally clear visual flow during the learning process, thus ensuring stability and semantic coherence in the training. Extensive experiments on datasets such as XD-Violence demonstrate that our method significantly outperforms existing approaches, particularly achieving notable improvements in the accuracy and stability of long-term anomaly detection. Our code is publicly available at https://github.com/ibliever/LTCLNet .
Ruomei Wang 0001, Linxuan Han, Baoquan Zhao, Fan Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2026 SketchBodyNet++: Sketch-Based 3D Human Mesh Reconstruction via Hybrid Parametric Networks
abstract
Sketches are an efficient and effective tool for generating 3D human meshes with arbitrary body shapes and poses. However, current mesh reconstruction methods are mainly designed for natural images, which are hard to apply to sketches due to the abstract and sparse characteristics of the latter. Moreover, there is no dataset with sufficient sketch-mesh pairs for developing and evaluating relevant methods. To tackle these issues, we introduce a hybrid framework that fits parametric human models (e.g., skinned multi-person linear model) to sketches in a coarse-to-fine manner. Specifically, the proposed framework consists of three core components: (i) Given a sketch image as the input, a vision transformer-based Local Image Encoder (LIE) is introduced to model the local structures of the sketch and yields a coarse mesh estimation. (ii) A Global Point Encoder (GPE) taking the 2D coordinates of sketch contours as inputs, is also utilized to obtain the global representation of the sketch. (iii) As the local presentation can depict human poses more precisely while the global representation is more suitable for body shapes, we propose a graph-based refiner (GRefiner) to leverage the advantages of both representations and generate the final well-fitted mesh. Furthermore, we collect a large-scale dubbed Sketch3DS, containing approximately 10,000 paired sketches and human meshes with diverse poses and shapes. Extensive experiments on Sketch3DS demonstrate that the proposed approach outperforms existing methods, achieving accurate alignment between input sketches and constructed human meshes.
Fei Wang 0056, Baoquan Zhao
IEEE Trans. Vis. Comput. Graph.4
2025 CoRe: Context-Regularized Text Embedding Learning for Text-to-Image Personalization
abstract
Recent advances in text-to-image personalization have enabled high-quality and controllable image synthesis for user-provided concepts. However, existing methods still struggle to balance identity preservation with text alignment. Our approach is based on the fact that generating prompt-aligned images requires a precise semantic understanding of the prompt, which involves accurately processing the interactions between the new concept and its surrounding context tokens within the CLIP text encoder. To address this, we aim to embed the new concept properly into the input embedding space of the text encoder, allowing for seamless integration with existing tokens. We introduce Context Regularization (CoRe), which enhances the learning of the new concept's text embedding by regularizing its context tokens in the prompt. This is based on the insight that appropriate output vectors of the text encoder for the context tokens can only be achieved if the new concept's text embedding is correctly learned. CoRe can be applied to arbitrary prompts without requiring the generation of corresponding images, thus improving the generalization of the learned text embedding. Additionally, CoRe can serve as a test-time optimization technique to further enhance the generations for specific prompts. Comprehensive experiments demonstrate that our method outperforms several baseline methods in both identity preservation and text alignment.
Feize Wu, Yun Pang, Lianyu Pang, Jian Yin 0001, Baoquan Zhao, Qing Li 0001, Xudong Mao
AAAI6
2025 Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025
abstract
Human identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition.
Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001
IJCB31
2025 GameMLD: A Game-Sourced Motion-Language Dataset for Stylized Motion Generation
abstract
Text-guided character animation generation has emerged as a significant research area with broad applications in gaming, film, interactive media, and beyond. However, existing motion-language datasets face limitations in motion quality, stylistic diversity, and annotation depth, particularly for professional applications. In contrast to existing datasets based on motion capture or video reconstruction techniques, our dataset leverages professionally crafted game animations and employs a structured annotation framework that incorporates standardized game design terminology. The dataset contains 8,700 high-fidelity motion sequences paired with 26,100 multi-level textual descriptions, generated through our proposed annotation pipeline that combines domain expertise with large language models. Through comprehensive experiments and user studies, we demonstrate GameMLD’s advantages in motion quality, style expressiveness, and annotation quality. Additionally, we showcase its practical value by developing a text-driven character animation generation system that effectively supports game production pipelines. Our experiments with state-of-the-art motion synthesis models demonstrate significant improvements in both animation quality and style control. The GameMLD dataset and source code can be reached via this link.
Yiyu Fu, Ziming Cheng, Yihao Liao, Jiangfeiyang Wang, Ruomei Wang 0001, Guanghui Yue 0001, Chenlei Lv, Baoquan Zhao
ICME8
2025 MSPoint-Gait: Multi-Scale Point Cloud Analysis for 3D Gait Recognition via Cross-Modal Learning
abstract
Recent advances in LiDAR technology have enabled privacy-preserving gait recognition using 3D point cloud data. However, existing approaches struggle with the inherent challenges of point cloud processing and understanding such as spatial sparsity, irregular sampling, and complex temporal dynamics. In this paper, we present MSPoint-Gait, a novel framework that addresses these challenges through multi-scale analysis and cross-modal learning. At the core of our framework lies a Depth-Aware Attention Module (DAAM) that leverages rich 3D geometric information to generate attention-weighted depth representations, enabling fine-grained feature extraction from point cloud sequences. We further introduce a Multi-Scale Spatio-Temporal (MSST) network that hierarchically captures both local and global gait patterns through adaptive convolution kernels across multiple spatial and temporal scales. These components are unified through a novel cross-modal learning strategy that effectively bridges the semantic gap between raw point clouds and structured depth representations. The proposed frame-work achieves state-of-the-art performance on the challenging SUSTech1K dataset, with 91.9% Rank-1 and 98.0% Rank-5 accuracy, demonstrating significant improvements over existing methods across various walking conditions and viewpoints.
Xinzhu Li, Yikun Chen, Guanghui Yue 0001, Wei Zhou 0021, Ruomei Wang 0001, Xudong Mao, Juepeng Zheng, Fan Zhou 0001, Ziqi Qiu, Baoquan Zhao
ICME11
2025 Dynamic Feature-Focusing with Cross-Modal Semantic Alignment for Video Moment Retrieval and Highlight Detection
abstract
Video moment retrieval and highlight detection (MR&HD) is a challenging multimodal understanding task that requires precise temporal localization and saliency estimation. While existing approaches have achieved promising performance, they face two critical limitations, including static modeling strategies that fail to adapt to diverse video contents and durations, as well as significant semantic gaps between video and textual query representations that hinder effective cross-modal alignment. To address these challenges, this paper presents a novel adaptive semantic-guided framework for improved video MR&HD. First, we leverage video captions as semantic prior knowledge to enhance video representation and reduce the initial semantic gap. Second, we introduce an Adaptive Feature Focusing Module (AFF) that employs dynamic convolution to flexibly capture salient information across varying temporal scales. Third, we design a Multi-perspective Semantic Sensing Module (MSS) that combines attention mechanisms with text reconstruction to achieve robust cross-modal semantic alignment. Extensive experiments on four public benchmarks demonstrate that our approach consistently outperforms state-of-the-art methods. Ablation studies and visualizations further validate the effectiveness of each component and demonstrate our method’s capability in narrowing the semantic gap between heterogeneous features.
Xuehui Liang, Ruomei Wang 0001, Baoquan Zhao
ICME3
2025 Multi-granularity Frequency Difference-Aware Attention for Video Question Answering
abstract
Video Question Answering (VideoQA) demands complex reasoning about multi-granular information, requiring both fine-grained visual details and global event understanding from videos. While existing methods employ stacked cross-modal attention modules for multi-granular feature representation, they struggle to effectively separate different granularities due to the intertwined nature of visual information in videos. To address this challenge, we introduce a novel Multi-granularity Frequency Difference-Aware Attention (MFDA) mechanism that enhances VideoQA by modeling unified multi-granular relation-ships between multimodal features in the frequency domain. MFDA comprises three key components: a heterogeneous multi-granularity dynamic-aware module, a frequency distance-aware function, and a multi-granular complementary module. These components enable the VideoQA model to effectively parse and filter multi-granularity feature information, allocate attention weights, and provide precise visual semantic cues for answer prediction. Extensive experiments demonstrate that MFDA serves as a plug-and-play cross-modal attention mechanism that significantly improves existing VideoQA models’ performance, achieving state-of-the-art results across diverse question types while reducing computational complexity. Our code is available at https://github.com/haha94322/MFDA.
Fan Zhou 0001, Ruomei Wang 0001, Baoquan Zhao
ICME4
2025 MCSMoG: Multi-Conditional Diffusion for Stylized Motion Generation with Parametric Control
abstract
Stylized human motion synthesis remains a fundamental challenge in computer animation and graphics, with a wide spectrum of applications spanning gaming, film production, virtual reality, and beyond. While recent advances in text-driven motion generation have shown promise, existing approaches face critical limitations including the inability to maintain consistent trajectory control, the lack of fine-grained stylization intensity adjustment, and inadequate generalization across diverse motion styles. To address these challenges, We introduce MCSMoG, a novel framework for controllable stylized motion synthesis through multi-conditional guidance. First, a new Multi-Conditional Motion Latent Diffusion (MC-MLD) model is proposed to introduce additional trajectory guidance and achieve trajectory decoupling. Second, we develop a Style and Non-Style Feature Fusion Module that dynamically blends motion features through an adjustable parameter, providing control over stylization intensity. Third, we integrate MotionCLIP as our style encoder, enhancing the model’s generalization capability across diverse and unseen motion styles. Extensive experiments conducted on the combined HumanML3D and 100STYLE datasets demonstrate that our approach outperforms state-of-the-art methods, achieving a 4.6% reduction in FID scores and a 4.1% increase in motion diversity. User studies further confirm the superiority of our method in style fidelity, semantic consistency, and motion naturalness.
Xinzhu Li, Guanghui Yue 0001, Wei Zhou 0021, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao
ICME9
2025 DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
abstract
Robust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing sufficient cues that can be exploited to handle viewpoint variations, and capture finer and meaningful details of gait. In this paper, we introduce a novel framework, termed DepthGait, that incorporates RGB-derived depth maps and silhouettes for enhanced gait recognition. Specifically, apart from the 2D silhouette representation of the human body, the proposed pipeline explicitly estimates depth maps from a given RGB image sequence and uses them as a new modality to capture discriminative features inherent in human locomotion. In addition, a novel multi-scale and cross-level fusion scheme has also been developed to bridge the modality gap between depth maps and silhouettes. Extensive experiments on standard benchmarks demonstrate that the proposed DepthGait achieves state-of-the-art performance compared to peer methods and attains an impressive mean rank-1 accuracy on the challenging datasets.
Xinzhu Li, Juepeng Zheng, Yikun Chen, Xudong Mao, Guanghui Yue 0001, Wei Zhou 0021, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao
ACM Multimedia10
2025 VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
abstract
Cross-video question answering presents significant challenges beyond traditional single-video understanding, particularly in establishing meaningful connections across video streams and managing the complexity of multi-source information retrieval. We introduce VideoForest, a novel framework that addresses these challenges through person-anchored hierarchical reasoning, enabling effective cross-video understanding without requiring end-to-end training. VideoForest integrates three key innovations: 1) a human-anchored feature extraction mechanism that employs ReID and tracking algorithms to establish robust spatiotemporal relationships across multiple video sources; 2) a multi-granularity spanning tree structure that hierarchically organizes visual content around person-level trajectories; and 3) a multi-agent reasoning framework that efficiently traverses this hierarchical structure to answer complex queries. To evaluate our method, we develop CrossVideoQA, a comprehensive benchmark specifically designed for person-centric cross-video analysis. Experimental results demonstrate VideoForest's superior performance in cross-video reasoning tasks, achieving 71.93% accuracy in person recognition, 83.75% in behavior analysis, and 51.67% in summarization and reasoning.
Yiran Meng, Junhong Ye, Wei Zhou 0021, Guanghui Yue 0001, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao
ACM Multimedia7
2025 VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
abstract
The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and talks, have led to a dramatic increase in speech-rich video content. In contrast to other video types, such as surveillance footage, which typically contain abundant visual cues, speech-rich videos convey most of their meaningful information through the audio channel. This poses challenges for improving content consumption using existing visual-based video summarization, navigation, and exploration systems. In this paper, we present VisAug, a novel interactive system designed to enhance speech-rich video navigation and engagement by automatically generating informative and expressive visual augmentations based on the speech content of videos. Our findings suggest that this system has the potential to significantly enhance the consumption and engagement of information in an increasingly video-driven digital landscape.
Baoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang 0001, Fan Zhou 0001, Shujin Lin
ACM Multimedia1
2025 Cross-Modality Interactive Attention Network for AI-generated image quality assessment
Tianwei Zhou, Songbai Tan, Leida Li, Baoquan Zhao, Qiuping Jiang, Guanghui Yue 0001
Pattern Recognit.4
2025 Boundary-Guided Feature-Aligned Network for Colorectal Polyp Segmentation
abstract
Colorectal polyp segmentation in endoscopic images is very important for the prevention and treatment of colorectal cancer. Because of the high similarity between polyps and their surrounding tissues, most deep neural network (DNN) based methods often struggle with blurry boundaries and result in inaccurate segmentation. In this paper, we propose a Boundary-guided Feature-aligned Network (BFNet) for polyp segmentation by taking a boundary prediction task as an auxiliary. Firstly, BFNet aggregates multi-layer features extracted from the backbone to mine boundary cues. Secondly, a flexible feature aggregation (FFA) module is used at each layer to adaptively fuse cross-layer features for coarse polyp localization. In the FFA module, considering the spatial misalignment between features at different layers, the feature of the high layer is aligned to and fused with that of the current layer using the deformable convolution and flexible merge block. After that, a boundary-guided feature enhancement (BFE) module is applied to refine the localization at boundary areas. In the BFE module, the boundary information is extracted and highlighted in both channel and spatial dimensions using the attention mechanisms with the assistance of boundary cues. By applying deep supervision to the BFE modules, BFNet can produce accurate polyp segmentation. Experimental results show that our BFNet outperforms 14 state-of-the-art DNN-based polyp segmentation methods on both in-domain and out-of-domain tests.
Guanghui Yue 0001, Shangjie Wu, Cheng Zhao 0003, Tianwei Zhou, Baoquan Zhao
IEEE Trans. Circuits Syst. Video Technol.7
2025 Single-View Clothed Human Reconstruction With Multi-View Consistency Representation
abstract
For single-view clothed human reconstruction, the fashionable PIFu-like framework depends on the pixel-aligned feature essentially, while this leads to depth ambiguity and inaccuracy of representation. Additionally, this task faces the inherent problem of the lack of invisible information. To solve these two problems, we propose depth-guided pixel-aligned feature and multi-view consistency prior to constrain representation learning of the single-view reconstruction task. The difference, between the depth values of points and estimated depth map, is used to filter pixel-aligned features. Thus, the image encoder can focus on capturing the feature of the visible part which is more effective in feature representation. The method introduces contrastive learning and masked autoencoder to achieve consistency of SMPL vertex features in each view which helps model to imagine invisible information. The experimental results show that the proposed method enables the feature to represent surface details more efficiently, thus achieves more reasonable and accurate representation learning. The qualitative and quantitative evaluations on the public and commercial datasets show that the proposed method can achieve better performance than previous implicit representation based methods.
Zhuo Su 0001, Yudi Tan, Zedan Zheng, Fan Zhou 0001, Baoquan Zhao
IEEE Trans. Vis. Comput. Graph.5
2024 Improved Text-Driven Human Motion Generation via Out-of-Distribution Detection and Rectification
Yiyu Fu, Baoquan Zhao, Chenlei Lv, Guanghui Yue 0001, Ruomei Wang 0001, Fan Zhou 0001
CVM (1)2
2024 Clip-Medfake: Synthetic Data Augmentation With AI-Generated Content for Improved Medical Image Classification
abstract
Data augmentation is serving as a critical and fundamental technology to improve model generalization and performance in a wide spectrum of machine learning tasks. Despite the increasing interest in developing various pathways to artificially generate new data to reduce the overfitting issue during model training, enriching the diversity of training data in the field of medicine remains facing enormous challenges. By virtue of recent advancements in generative artificial intelligence, we present a novel data augmentation framework, CLIP-MedFake, to address the shortage of training data used in medical image classification. The proposed method first employs the Stable Diffusion model to generate new fake data based on a small amount of training data, and then adopts the paradigm of few-shot learning and uses the CLIP architecture as the backbone to pre-train the model with synthetic data and then fine-tune it with real medical images. Extensive experiment results on two publicly available datasets demonstrate the effectiveness of the proposed method in promoting medical image classification.
Honghui Chen, Baoquan Zhao, Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001
ICIP2
2024 Full-Reference Motion Quality Assessment Based on Efficient Monocular Parametric 3D Human Body Reconstruction
abstract
Human motion capture and analysis are pivotal to a wide spectrum of killer applications in various domains such as sports, performing arts, diagnostic tests in physical medicine, rehabilitation, and figure training. However, automatic reconstruction, assessment, and visualization of human motions from a monocular video are still suffering from grand challenges that are inadequately addressed by existing studies. In this paper, we present a novel full-reference human motion quality assessment and visualization system based on monocular parametric 3D human body reconstruction. Specifically, our method first reconstructs a 3D parametric model from each sampled frame of a monocular video and harvests a physically plausible motion sequence using the proposed optimization scheme; Secondly, a full-reference assessment metric is designed to evaluate the consistency between the reconstructed motion and the reference; Finally, a new interactive visualization system is developed to facilitate multi-grained motion quality evaluation and visual analysis. Extensive quantitative and qualitative experiments demonstrate the effectiveness and superiority of the proposed method.
Yiwei Yuan, Xiangyu Zeng 0005, Ling Xie, Yiyu Fu, Guanghui Yue 0001, Baoquan Zhao
ICME7
2024 Single Free-Hand Sketch Guided Free-Form Deformation For 3D Shape Generation
abstract
Sketch-guided point cloud reconstruction aims to provide an efficient and flexible pathway to generate plausible 3D shapes automatically from a free-hand sketch shaping the modeling intentions of end users. However, such a task is still in its infancy due to the complex, challenging, and highly variable patterns of sketches by nature. In this paper, we present a novel sketch-guided framework based on free-form deformation (FFD) for 3D point cloud generation. To capture sufficient meaningful features from a sketch, a dual-branch encoding architecture is devised to extract complementary semantic and geometric clues by formulating the input as a binary image and a 2D point cloud, respectively. The proposed encoder also learns useful features to guide content generation from a template point cloud before decoding the resultant global features into a set of control points for the use of FFD. We have also developed a large and diverse manually collected dataset, Sketch-3DPC, in which there are a total of 13,754 sketch and 3D point cloud pairs categorized into 11 classes. Both qualitative and quantitative experiment results demonstrate the superiority of the proposed methodology and dataset.
Fei Wang 0056, Jianqiang Sheng, Zhineng Zhang, Juepeng Zheng, Baoquan Zhao
ICME6
2024 AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
abstract
Recent advances in text-to-image models have enabled high-quality personalized image synthesis based on user-provided concepts with flexible textual control. In this work, we analyze the limitations of two primary techniques in text-to-image personalization: Textual Inversion and DreamBooth. When integrating the learned concept into new prompts, Textual Inversion tends to overfit the concept, while DreamBooth often overlooks it. We attribute these issues to the incorrect learning of the embedding alignment for the concept. To address this, we introduce AttnDreamBooth, a novel approach that separately learns the embedding alignment, the attention map, and the subject identity across different training stages. We also introduce a cross-attention map regularization term to enhance the learning of the attention map. Our method demonstrates significant improvements in identity preservation and text alignment compared to the baseline methods.
Lianyu Pang, Jian Yin 0001, Baoquan Zhao, Feize Wu, Fu Lee Wang, Qing Li 0001, Xudong Mao
NeurIPS3
2024 Predicting Plain Text Imageability for Faithful Prompt-Conditional Image Generation
Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao
PRICAI (3)7
2024 Dual-Constraint Coarse-to-Fine Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is an important yet challenging task, with great application values in industrial defect detection, medical care, etc. The challenges mainly come from the high intrinsic similarities between target objects and background. In this paper, inspired by the biological studies that object detection consists of two steps, i.e., search and identification, we propose a novel framework, named DCNet, for accurate COD. DCNet explores candidate objects and extra object-related edges through two constraints (object area and boundary) and detects camouflaged objects in a coarse-to-fine manner. Specifically, we first exploit an area-boundary decoder (ABD) to obtain initial region cues and boundary cues simultaneously by fusing multi-level features of the backbone. Then, an area search module (ASM) is embedded into each level of the backbone to adaptively search coarse regions of objects with the assistance of region cues from the ABD. After the ASM, an area refinement module (ARM) is utilized to identify fine regions of objects by fusing adjacent-level features with the guidance of boundary cues. Through the deep supervision strategy, DCNet can finally localize the camouflaged objects precisely. Extensive experiments on three benchmark COD datasets demonstrate that our DCNet is superior to 12 state-of-the-art COD methods. In addition, DCNet shows promising results on two COD-related tasks, i.e., industrial defect detection and polyp segmentation.
Guanghui Yue 0001, Houlu Xiao, Hai Xie, Tianwei Zhou, Wei Zhou 0021, Weiqing Yan, Baoquan Zhao, Tianfu Wang 0001, Qiuping Jiang
IEEE Trans. Circuits Syst. Video Technol.7
2024 Multitask Deep Neural Network With Knowledge-Guided Attention for Blind Image Quality Assessment
abstract
Blind image quality assessment (BIQA) targets predict the perceptual quality of an image without any reference information. However, known methods have considerable room for performance improvement due to limited efforts in distortion knowledge usage. This paper proposes a novel multitask learning based BIQA method termed KGANet, which takes image distortion classification as an auxiliary task and uses the knowledge learned from the auxiliary task to assist accurate quality prediction. Different from existing CNN-based methods, KGANet adopts a transformer as the backbone for feature extraction, which can learn more powerful and robust representations. Specifically, it comprises two essential components: a cross-layer information fusion (CIF) module and a knowledge-guided attention (KGA) module. Considering that both global and local distortions appear in an image, CIF fuses the features of the adjacent layers extracted by the backbone to obtain a multiscale feature representation. KGA incorporates the distortion probability estimated by the auxiliary task with the distortion embeddings, which are selected from subword unit embeddings based on a textual template, to form distortion knowledge. This knowledge further serves as guidance to enhance the features of each layer and strengthen the connection between the main and auxiliary task. We demonstrate the effectiveness of the proposed KGANet through extensive experiments on benchmark databases. Experimental results show that KGANet correlates well with subjective perceptual judgments and achieves superior performance over 12 state-of-the-art BIQA methods.
Tianwei Zhou, Songbai Tan, Baoquan Zhao, Guanghui Yue 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Intrinsic and Isotropic Resampling for 3D Point Clouds
abstract
With rapid development of 3D scanning technology, 3D point cloud based research and applications are becoming more popular. However, major difficulties are still exist which affect the performance of point cloud utilization. Such difficulties include lack of local adjacency information, non-uniform point density, and control of point numbers. In this paper, we propose a two-step intrinsic and isotropic (I&I) resampling framework to address the challenge of these three major difficulties. The efficient intrinsic control provides geodesic measurement for a point cloud to improve local region detection and avoids redundant geodesic calculation. Then the geometrically-optimized resampling uses a geometric update process to optimize a point cloud into an isotropic or adaptively-isotropic one. The point cloud density can be adjusted to global uniform (isotropic) or local uniform with geometric feature keeping (being adaptively isotropic). The point cloud number can be controlled based on application requirement or user-specification. Experiments show that our point cloud resampling framework achieves outstanding performance in different applications: point cloud simplification, mesh reconstruction and shape registration. We provide the implementation codes of our resampling method at https://github.com/vvvwo/II-resampling.
Chenlei Lv, Weisi Lin, Baoquan Zhao
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 KSS-ICP: Point Cloud Registration Based on Kendall Shape Space
abstract
Point cloud registration is a popular topic that has been widely used in 3D model reconstruction, location, and retrieval. In this paper, we propose a new registration method, KSS-ICP, to address the rigid registration task in Kendall shape space (KSS) with Iterative Closest Point (ICP). The KSS is a quotient space that removes influences of translations, scales, and rotations for shape feature-based analysis. Such influences can be concluded as the similarity transformations that do not change the shape feature. The point cloud representation in KSS is invariant to similarity transformations. We utilize such property to design the KSS-ICP for point cloud registration. To tackle the difficulty to achieve the KSS representation in general, the proposed KSS-ICP formulates a practical solution that does not require complex feature analysis, data training, and optimization. With a simple implementation, KSS-ICP achieves more accurate registration from point clouds. It is robust to similarity transformation, non-uniform density, noise, and defective parts. Experiments show that KSS-ICP has better performance than the state-of-the-art. Code (vvvwo/KSS-ICP) and executable files (vvvwo/KSS-ICP/tree/master/EXE) are made public.
Chenlei Lv, Weisi Lin, Baoquan Zhao
IEEE Trans. Image Process.3
2023 Interaction-Matrix Based Personalized Image Aesthetics Assessment
abstract
Personalized image aesthetics assessment (IAA) aims to estimate aesthetic experiences subject to the preferences of individual users, contrary to generic IAA that estimates aesthetic experiences subject to average preferences. Most existing personalized IAA methods treat personalized aesthetic experiences as deviations from a generic aesthetic experience, and therefore, personalized IAA models are designed to build upon the prior knowledge on generic IAA. However, we propose that acquiring knowledge on generic IAA is not necessary for building a personalized IAA model. Instead of modeling personalized IAA on the basis of generic IAA, this work proposes to directly estimate personalized aesthetic experiences from the interactions between image contents and user preferences (i.e., preference-content interaction), where interaction-matrices representing preference-content interactions are constructed without needs for prior generic IAA knowledge. To this end, we construct interaction-matrices from content features constructed from pre-trained image classification features and latent preference features. To realize a robust interaction-matrix based personalized IAA model, we discuss in detail on different strategies for constructing interaction-matrices and estimating personalized aesthetic scores from the interaction-matrices. Besides the personalized IAA scenario, we further propose strategies to adapt the proposed personalized IAA model to different scenarios of generic IAA. Extensive experiments show that: 1) our method significantly outperforms 5 previous relevant personalized IAA methods on FLICKR-AES dataset, especially the methods that require generic IAA knowledge as the basis; 2) in terms of generic IAA, the proposed approach also outperforms 13 generic IAA methods on AVA dataset.
Jingwen Hou, Weisi Lin, Guanghui Yue 0001, Weide Liu, Baoquan Zhao
IEEE Trans. Multim.5
2022 Voxel Structure-Based Mesh Reconstruction From a 3D Point Cloud
abstract
Mesh reconstruction from a 3D point cloud is an important topic in the fields of computer graphic, computer vision, and multimedia analysis. In this paper, we propose a voxel structure-based mesh reconstruction framework. It provides the intrinsic metric to improve the accuracy of local region detection. Based on the detected local regions, an initial reconstructed mesh can be obtained. With the mesh optimization in our framework, the initial reconstructed mesh is optimized into an isotropic one with the important geometric features such as external and internal edges. The experimental results indicate that our framework shows great advantages over peer ones in terms of mesh quality, geometric feature keeping, and processing speed. The source code of the proposed method is publicly available1.
Chenlei Lv, Weisi Lin, Baoquan Zhao
IEEE Trans. Multim.3
2021 Fine-Grained Patch Segmentation and Rasterization for 3-D Point Cloud Attribute Compression
abstract
Due to the high dimensionality of point cloud data and the irregularity and complexity of its geometric structure, effective attribute compression remains a very challenging task. Many recent efforts have focused on transforming point clouds into images and leveraging existing sophisticated image/video codecs to improve attribute coding efficiency. However, how to synthesize coherent and correlation-preserving attribute images is still inadequately addressed by existing studies, which are hindering the exertion of the merits of well-developed compression infrastructure. In this paper, we present a novel image synthesis method for effective point cloud attribute compression. Firstly, the proposed scheme segments a given point cloud into a collection of fine-grained patches by performing geometric structure analysis using heat kernel signature feature descriptor and complex points; Secondly, we transform the obtained patches from 3-D to 2-D using a low-dimensional embedding algorithm and then convert them into patch attribute images with the proposed patch rasterization and rectification method; And finally, we compactly assemble all the attribute images of patches together by formulating it as a bin nesting problem and harvest an attribute image of the whole point cloud for image/video-based compression. Experimental results demonstrate the effectiveness of the proposed method in point cloud attribute compression and its superiority over state-of-the-art codecs. The source code of this work is publicly available athttps://github.com/pccompession/UPCAC.
Baoquan Zhao, Weisi Lin, Chenlei Lv
IEEE Trans. Circuits Syst. Video Technol.1
2021 Approximate Intrinsic Voxel Structure for Point Cloud Simplification
abstract
A point cloud as an information-intensive 3D representation usually requires a large amount of transmission, storage and computing resources, which seriously hinder its usage in many emerging fields. In this paper, we propose a novel point cloud simplification method, Approximate Intrinsic Voxel Structure (AIVS), to meet the diverse demands in real-world application scenarios. The method includes point cloud pre-processing (denoising and down-sampling), AIVS-based realization for isotropic simplification and flexible simplification with intrinsic control of point distance. To demonstrate the effectiveness of the proposed AIVS-based method, we conducted extensive experiments by comparing it with several relevant point cloud simplification methods on three public datasets, including Stanford, SHREC, and RGB-D scene models. The experimental results indicate that AIVS has great advantages over peers in terms of moving least squares (MLS) surface approximation quality, curvature-sensitive sampling, sharp-feature keeping and processing speed. The source code of the proposed method is publicly available. (https://github.com/vvvwo/AIVS-project).
Chenlei Lv, Weisi Lin, Baoquan Zhao
IEEE Trans. Image Process.3
2021 Reconstructing 3D Model from Single-View Sketch with Deep Neural Network
abstract
In this paper, we introduce a novel 3D shape reconstruction method from a single‐view sketch image based on a deep neural network. The proposed pipeline is mainly composed of three modules. The first module is sketch component segmentation based on multimodal DNN fusion and is used to segment a given sketch into a series of basic units and build a transformation template by the knots between them. The second module is a nonlinear transformation network for multifarious sketch generation with the obtained transformation template. It creates the transformation representation of a sketch by extracting the shape features of an input sketch and transformation template samples. The third module is deep 3D shape reconstruction using multifarious sketches, which takes the obtained sketches as input to reconstruct 3D shapes with a generative model. It fuses and optimizes features of multiple views and thus is more likely to generate high‐quality 3D shapes. To evaluate the effectiveness of the proposed method, we conduct extensive experiments on a public 3D reconstruction dataset. The results demonstrate that our model can achieve better reconstruction performance than peer methods. Specifically, compared to the state‐of‐the‐art method, the proposed model achieves a performance gain in terms of the five evaluation metrics by an average of 25.5% on the man‐made model dataset and 23.4% on the character object dataset using synthetic sketches and by an average of 31.8% and 29.5% on the two datasets, respectively, using human drawing sketches.
Fei Wang 0056, Baoquan Zhao, Dazhi Jiang, Jianqiang Sheng
Wirel. Commun. Mob. Comput.3
2020 Content-Dependency Reduction With Multi-Task Learning In Blind Stitched Panoramic Image Quality Assessment
abstract
In this work, we investigate deep learning based solutions to blind quality assessment of stitched panoramic images (SPI). The main problem to tackle is that the ground truth data is usually insufficient. As a result, the learned model can easily overfit data with specific content. Because most distortions of SPIs lie within local regions, the problem cannot be alleviated by commonly-used patch-wise training, which assumes local quality equals global quality. We propose a multi-task learning strategy which encourages learned representation to be less dependent on image content. A siamese network with two weight-shared CNN branches is trained to simultaneously compare the quality of two images of the same scene and predict the quality score of each image. Since two images of the same scene are processed by the same CNN, the CNN tends to find their quality differences instead of content differences under the constraint of the quality ranking objective. Because two tasks share the same representations learned by the CNN, the regression task can be further benefited from the quality-sensitive representations. Extensive experiments demonstrate the effectiveness of the proposed model and its superiority over existing SPI quality assessment methods.
Jingwen Hou, Weisi Lin, Baoquan Zhao
ICIP3
2019 Range Image Based Point Cloud Colorization Using Conditional Generative Model
abstract
Nowadays, three-dimensional (3D) point cloud has been an emerging medium to represent real-world scenes and objects. However, there is a considerable proportion of point clouds whose color attribute information is not captured during the acquisition process due to the device or environment limitations. This poses a great challenge for efficient management and utilization of point clouds. To address this problem, we introduce an automatic colorization scheme based on a deep generative network for 3D point clouds. The proposed approach uses the range images of point could geometry and trains a conditional generative adversarial network to predict the color of those images. Later, the color of each pixel in the colorized image is projected back to its corresponding point in the 3D point cloud. The experimental results demonstrate the efficacy of the proposed colorization approach in facilitating users to recognize and handle 3D point cloud data better.
Jong-Uk Hou, Baoquan Zhao, Naushad Ansari, Weisi Lin
ICIP2
2019 Lecture2Note: Automatic Generation of Lecture Notes from Slide-Based Educational Videos
abstract
Given rapid development witnessed by open educational resources (OER) in the past few decades, a considerable number of online educational videos emerge on various MOOC platforms such as Coursera and YouTube. Nevertheless, most educational videos on the internet are lengthy and lack of elaborate annotations, which poses a challenge for learners to explore and locate content of interest efficiently. To address this, we present an automatic note-generating method to establish correspondences between visual entities in the slide-based lecture video and their descriptive speech texts by evaluating the semantic relationship. Firstly, the visual entities are extracted and recognised from the presentation slides. Then, each of visual entities is associated with its corresponding descriptive speech text. Finally, a placement optimisation scheme is put forward to pack the visual entities and speech texts into a note-like layout in a compact fashion, which can help learners to improve their learning efficiency. The experimental results show that the efficient performances about visual entity extraction and correspondence matching are efficient. The user study is also designed to investigate the performance of Lecture2Note in facilitating learning. Compared with peer methods, the auto-generated note created by our method achieves a higher user satisfaction level regarding a properly structured layout as well as efficient content navigation and exploration.
Chengpei Xu, Ruomei Wang 0001, Shujin Lin, Baoquan Zhao, Lijie Shao, Mengqiu Hu
ICME5
2019 A New Visual Interface for Searching and Navigating Slide-Based Lecture Videos
abstract
The rapid development of distance education technologies, e.g. MOOCs, provide learners unprecedented access to high-quality online lecture videos at scale, anytime and anywhere. Unfortunately, these valuable resources are often underutilized by online learners. One prevailing reason is the lack of support for and the resulting difficulty of exploring and locating content of interest among lengthy recordings of course lectures. To address this deficiency, we introduce a novel visual interface that supports efficient search and navigation of video content at fine granularities and with rich semantic clues. The interface is particularly designed for slide-based lecture videos (SBLV), which represent a significant portion of online lecture videos. The interface comprehensively derives versatile semantic clues for video content indexing and visual aid generation according to visual elements, text, and mathematical expressions included on lecture slides, speeches recorded, as well as mouse and cursor pointing actions captured during a lecture. Empowered by such semantically revealing indices and visual assistance, the interface is able to noticeably enhance online learners' capabilities in searching and browsing of content in need from SBLVs. The advantages of the new interface are demonstrated through benchmarked experimental results in comparison with peer methods.
Baoquan Zhao, Songhua Xu, Shujin Lin, Ruomei Wang 0001
ICME1
2017 A Novel System for Visual Navigation of Educational Videos Using Multimodal Cues
abstract
With recent developments and advances in distance learning and MOOCs, the amount of open educational videos on the Internet has grown dramatically in the past decade. However, most of these videos are lengthy and lack of high-quality indexing and annotations, which triggers an urgent demand for efficient and effective tools that facilitate video content navigation and exploration. In this paper, we propose a novel visual navigation system for exploring open educational videos. The system tightly integrates multimodal cues obtained from the visual, audio and textual channels of the video and presents them with a series of interactive visualization components. With the help of this system, users can explore the video content using multiple levels of details to identify content of interest with ease. Extensive experiments and comparisons against previous studies demonstrate the effectiveness of the proposed system.
Baoquan Zhao, Shujin Lin, Songhua Xu, Ruomei Wang 0001
ACM Multimedia1
2016 A new visual navigation system for exploring biomedical Open Educational Resource (OER) videos
abstract
OBJECTIVE: Biomedical videos as open educational resources (OERs) are increasingly proliferating on the Internet. Unfortunately, seeking personally valuable content from among the vast corpus of quality yet diverse OER videos is nontrivial due to limitations of today's keyword- and content-based video retrieval techniques. To address this need, this study introduces a novel visual navigation system that facilitates users' information seeking from biomedical OER videos in mass quantity by interactively offering visual and textual navigational clues that are both semantically revealing and user-friendly. MATERIALS AND METHODS: The authors collected and processed around 25 000 YouTube videos, which collectively last for a total length of about 4000 h, in the broad field of biomedical sciences for our experiment. For each video, its semantic clues are first extracted automatically through computationally analyzing audio and visual signals, as well as text either accompanying or embedded in the video. These extracted clues are subsequently stored in a metadata database and indexed by a high-performance text search engine. During the online retrieval stage, the system renders video search results as dynamic web pages using a JavaScript library that allows users to interactively and intuitively explore video content both efficiently and effectively.ResultsThe authors produced a prototype implementation of the proposed system, which is publicly accessible athttps://patentq.njit.edu/oer To examine the overall advantage of the proposed system for exploring biomedical OER videos, the authors further conducted a user study of a modest scale. The study results encouragingly demonstrate the functional effectiveness and user-friendliness of the new system for facilitating information seeking from and content exploration among massive biomedical OER videos. CONCLUSION: Using the proposed tool, users can efficiently and effectively find videos of interest, precisely locate video segments delivering personally valuable information, as well as intuitively and conveniently preview essential content of a single or a collection of videos.
Baoquan Zhao, Songhua Xu, Shujin Lin
J. Am. Medical Informatics Assoc.1
2014 3D Model Editing from Contour Drawings on Orthographic Projection Views
Yuhui Hu, Xuliang Guo, Baoquan Zhao, Shujin Lin
ICISP3