EDBT 2026 Demo / reviewers in the wild / expert
Meie Fang
dblp:259/6564 · also Mei-E Fang, Mei-e Fang
· DBLP profile ↗
47ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0003-4292-8889ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 21 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-End Knowledge Distillation for Unsupervised Domain Adaptation with Large Vision-language ModelsabstractKnowledge distillation based on large vision-language models (VLMs) has recently emerged as a significant solution to transfer knowledge from the source domain to the target domain in unsupervised domain adaptation (UDA) tasks. However, existing methods employ a two-stage training pipeline, which not only complicates the training procedure but also lacks interactions between the source and target domains, severely hindering real-time cross-domain knowledge transfer. To address these challenges, we propose End-to-End Knowledge Distillation for UDA with large VLMs (termed as EKDA). (1) EKDA employs a lightweight prompt learning mechanism to first embed the knowledge from the source domain into VLMs, and then simultaneously utilize the image encoder and text encoder of VLMs to perform knowledge distillation on the target domain, significantly reducing the domain gap. (2) EKDA designs a teacher-student alternating training strategy to implement real-time collaborative interactions across domains, enabling an end-to-end paradigm to provide accurate source domain-aware supervision for the target domain. We conduct extensive experiments on 4 widely recognized benchmark datasets including Office-31, Office-Home, VisDA-2017, and Mini-DomainNet. Experimental results demonstrate that EKDA achieves significant performance improvement over the state-of-the-art UDA approaches, while maintaining a much lower model complexity. Take Office-Home for example, EKDA has gained at least 2.7% performance improvement while reducing the learnable parameters by over 80% compared with the state-of-the-art UDA baselines. Yangtao Wang, Xingwei Deng, Yanzhao Xie, Weilong Peng, Siyuan Chen 0005, Xiaocui Li 0001, Maobin Tang, Meie Fang |
AAAI | 8 |
| 2026 | WaveSculpt: Text-to-3D generation with wavelet-guided score distillation
Weilong Peng, Jianhui Huo, Keke Tang, Yangtao Wang, Yan Wang 0022, Meie Fang |
Comput. Aided Geom. Des. | 6 |
| 2026 | Transferable and undefendable point cloud attacks via medial axis transform
Keke Tang, Yuze Gao, Weilong Peng, Meie Fang, Peican Zhu |
Comput. Aided Geom. Des. | 5 |
| 2026 | Intra-modal consistency for image-text retrieval through soft-label distillation
Yangtao Wang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, C. L. Philip Chen, Ping Li 0016, Wensheng Zhang 0002 |
Pattern Recognit. | 7 |
| 2026 | F3-SD: Focal feature fusion with self-distillation on large vision-language models for cross-modal retrieval
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
Pattern Recognit. | 7 |
| 2026 | Prompt-affinity multi-modal class centroids for unsupervised domain adaptionabstractIn recent years, the advancements in large vision-language models (VLMs) like CLIP have sparked a renewed interest in leveraging the prompt learning mechanism to preserve semantic consistency between source and target domains in unsupervised domain adaption (UDA). While these approaches show promising results, they encounter fundamental limitations when quantifying the similarity between source and target domain data , primarily stemming from the redundant and modality-missing class centroids . To address these limitations, we propose P rompt-affinity M ulti-modal C lass C entroids for UDA (termed as PMCC). Firstly, we fuse the text class centroids (directly generated from the text encoder of CLIP with manual prompts for each class) and image class centroids (generated from the image encoder of CLIP for each class based on source domain images) to yield the multi-modal class centroids. Secondly, we conduct the cross-attention operation between each source or target domain image and these multi-modal class centroids. In this way, these class centroids that contain rich semantic information of each class will serve as a bridge to effectively measure the semantic similarity between different domains. Finally, we design a logit bias head and employ a multi-modal prompt learning mechanism to accurately predict the true class of each image for both source and target domains. We conduct extensive experiments on 4 popular UDA datasets including Office-31, Office-Home, VisDA-2017, and DomainNet. The experimental results validate our PMCC achieves higher performance with lower model complexity than the state-of-the-art (SOTA) UDA methods. The code of this project is available at GitHub: https://github.com/246dxw/PMCC . Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
Pattern Recognit. | 6 |
| 2026 | Cross-domain distillation for unsupervised domain adaptation with large vision-language models
Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
Pattern Recognit. | 6 |
| 2026 | PTPD: Prototype-Guided Triplet Prompt Distillation with Vision-language models
Yanzhao Xie, Yangtao Wang, Rukai Wei, Dandan Shao, Maobin Tang, Meie Fang, Weilong Peng, Lisheng Fan, Wensheng Zhang 0002 |
Pattern Recognit. | 8 |
| 2026 | MKGPL: graph prompt learning with multi-view knowledge for few-shot recognition
Yanzhao Xie, Man Qiu, Yangtao Wang, Siyuan Chen 0005, Meie Fang, Maobin Tang, Wensheng Zhang 0002 |
Pattern Recognit. | 5 |
| 2026 | Boosting illuminant estimation in deep color constancy through brightness robustness enhancement
Mengda Xie, Chengzhi Zhong, Yiling He, Zhan Qin, Meie Fang |
Pattern Recognit. | 5 |
| 2026 | High Feature Distinguishability for Adaptive Image-text Matching with Dual-stream TransformersabstractRecently, most image-text matching (ITM) approaches have embraced a dual-stream transformer architecture to facilitate the learning and alignment of cross-modal semantic information. Despite the efficacy of this methodology in bridging the semantic disparity between images and texts, it exhibits two primary limitations. Firstly, it falls short in discriminating the nuanced similarities among features, which leads to misleading outcomes or even compromises the overall ITM process. Secondly, the conventional triplet training paradigm relies on a pre-determined, fixed margin coefficient, thereby impeding its capacity to accurately gauge the similarity relationships between positive and negative samples. In this article, we propose high feature D istinguishability for A daptive I mage-text M atching with dual-stream transformers (termed as DAIM). To address the first limitation, we design a feature discriminability module to bring similar features closer together but with a certain degree of distinction and push dissimilar features farther apart, resulting in high feature distinguishability for accurate ITM. To address the second limitation, we devise a margin optimization module to perceive the similarity distribution between positive and negative samples in real-time during training, thereby adaptively adjusting the margin coefficient to minimize the cross-modal semantic gap to the greatest extent possible. Based on this, we align the multi-level (i.e., representations from low-, middle-, and high-layer transformer encoders) semantic information of cross-modal data by adaptively optimizing the semantic distributions of positive and negative samples. We conduct extensive experiments on two commonly used benchmark datasets, including MSCOCO and Flickr30K. Experimental results verify that DAIM can achieve a higher performance (e.g., 4.7% RSUM gain on MSCOCO) than the state-of-the-art ITM methods. The open-sourced code of this project is available at: https://github.com/Hudjkfhdsjfhdjkg/DAIM.git . Yangtao Wang, Weibin Huang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2026 | BRep-GD: A Graph Diffusion Model for CAD Boundary Representation GenerationabstractIn modern computer-aided design (CAD), Boundary Representation (B-rep) is a widely used geometric modeling technique in industrial design and manufacturing. However, existing B-rep generation methods, which rely on tree-based hierarchies to represent and generate B-reps, fail to fully exploit the inherent graph structure of B-reps, resulting in suboptimal efficiency and model quality. To address this issue, we propose BRep-GD, a graph diffusion-based model specifically designed for B-rep generation. Unlike prior methods, BRep-GD treats B-reps as graphs, where nodes represent face elements, and edges represent boundary and vertex elements. By utilizing a continuous topological graph data structure, BRep-GD overcomes the challenges associated with directly applying graph diffusion models to B-rep generation. Specifically, BRep-GD introduces a graph diffusion method tailored to the features of CAD data, generating faces and edges sequentially. During edge generation, continuous topology decoupling is employed to avoid the need for global attention calculations, reducing computational complexity while ensuring geometric consistency and high-quality results. Experimental results demonstrate that BRep-GD outperforms existing state-of-the-art methods in both unconditional and class-conditional generation tasks, particularly in generating watertight solids and handling complex geometries. It significantly reduces isolated or inconsistent geometric components and improves generation efficiency. Fei-wei Qin, Chenqi Luo, Junhao Hou, Meie Fang, Ligang Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | PeelMesh: Efficient Interactive Segmentation via Geodesic-Driven Dynamic Topological Updates
Junjie Yin, Zixi Huang, Meie Fang, Ping Li 0016, Weiyin Ma |
CGI (1) | 4 |
| 2025 | From Pixels to Shapes: Generative AI for 2D Images and 3D Models
Jianhui Huo, Shijian Xu, Weilong Peng, Yangtao Wang, Yan Wang 0022, Meie Fang |
ICIC (19) | 7 |
| 2025 | Enhancing Cross-modal Semantic Consistency via Key Token Alignment for Image-text RetrievalabstractImage-text retrieval (ITR) plays a pivotal role in advancing intelligent transportation systems, facilitating efficient retrieval and utilization of multimedia data to enhance traffic management and safety significantly. However, existing ITR solutions have not effectively addressed the issues of image patch redundancy and text word redundancy, leading to erroneous image-text matching. In this paper, we propose SCTA that enhances cross-modal semantic consistency via key token alignment for ITR. Firstly, SCTA evaluates the importance of each image patch by calculating the self-attention scores within patches and cross-attention scores between patches and words. Secondly, SCTA implements aggregation operations on image and text separately, aiming to generate information-rich key image patch embeddings and text word token embeddings. Finally, SCTA completes fine-grained alignment by maximizing the similarity between patch-to-word and word-to-patch. Therefore, SCTA simultaneously addresses image patch redundancy and text word redundancy issues, enhancing semantic consistency by aligning the core semantic information between image-text pairs. Extensive experiments on multiple datasets including Flickr30K and MS-COCO verify the superior performance of SCTA compared with the SOTA fine-grained ITR methods. The code of this paper is released at GitHub: https://github.com/ICME2025ITR/SCTA. Huilong Lin, Yangtao Wang, Meie Fang, Yanzhao Xie, Xiaocui Li 0001, Weilong Peng, Siyuan Chen 0005, Maobin Tang, Ping Li 0016 |
ICME | 3 |
| 2025 | Drawing2CAD: Sequence-to-Sequence Learning for CAD Generation from Vector DrawingsabstractComputer-Aided Design (CAD) generative modeling is driving significant innovations across industrial applications. Recent works have shown remarkable progress in creating solid models from various inputs such as point clouds, meshes, and text descriptions. However, these methods fundamentally diverge from traditional industrial workflows that begin with 2D engineering drawings. The automatic generation of parametric CAD models from these 2D vector drawings remains underexplored despite being a critical step in engineering design. To address this gap, our key insight is to reframe CAD generation as a sequence-to-sequence learning problem where vector drawing primitives directly inform the generation of parametric CAD operations, preserving geometric precision and design intent throughout the transformation process. We propose Drawing2CAD, a framework with three key technical components: a network-friendly vector primitive representation that preserves precise geometric information, a dual-decoder transformer architecture that decouples command type and parameter generation while maintaining precise correspondence, and a soft target distribution loss function accommodating inherent flexibility in CAD parameters. To train and evaluate Drawing2CAD, we create CAD-VGDrawing, a dataset of paired engineering drawings and parametric CAD models, and conduct thorough experiments to demonstrate the effectiveness of our method. Code and dataset are available at https://github.com/lllssc/Drawing2CAD. Fei-wei Qin, Shichao Lu, Junhao Hou, Changmiao Wang, Meie Fang, Ligang Liu 0001 |
ACM Multimedia | 5 |
| 2025 | MeshPAD: Payload-aware mesh distortion for 3D steganography based on geometric deep learning
Weilong Peng, Keke Tang, Weixuan Tang 0002, Yong Su 0003, Meie Fang, Ping Li 0016 |
Expert Syst. Appl. | 5 |
| 2025 | Adaptive Multi-Lens Phase Modulation for Scale-Aware Privacy-Preserving Human Pose RecognitionabstractRecently, optical privacy protection has emerged as a promising approach for safeguarding visual privacy at the physical acquisition stage. However, existing methods often face a trade‐off between privacy strength and human pose recognition accuracy, particularly in long‐range and multi‐scale scenarios. To address this challenge, we propose a novel adaptive optical privacy‐preserving framework that integrates a learnable optical modulation system with a human pose recognition network. The core of our method lies in a sparse‐weighted multi‐lens model, where a lightweight multilayer perceptron (MLP) predicts a sparse set of coefficients to linearly combine predefined lens phase profiles based on facial region geometry. This enables dynamic control over the point spread function (PSF), adapting the degree of image degradation to subject scale in real time. Additionally, we introduce a privacy‐aware loss function that selectively reduces facial localization accuracy while preserving body pose information. Extensive experiments on MSCOCO and FLIC datasets demonstrate that the proposed method achieves a favorable balance between privacy protection and pose estimation, outperforming previous optical‐ and software‐based baselines. Weilong Peng, Quanwei Deng, Mingjie Li 0004, Yangtao Wang, Yan Wang 0022, Lisheng Fan, Meie Fang |
IET Softw. | 8 |
| 2025 | RetouchUAA: Unconstrained Adversarial Attack via Realistic Image RetouchingabstractDeep Neural Networks (DNNs) are susceptible to adversarial examples. Conventional attacks generate controlled noise-like perturbations that fail to reflect real-world scenarios and hard to interpretable. In contrast, recent unconstrained attacks mimic natural image transformations occurring in the real world for perceptible but inconspicuous attacks, yet compromise realism due to neglect of image post-processing and uncontrolled attack direction. In this paper, we propose RetouchUAA, an unconstrained attack that exploits a real-life perturbation: image retouching styles, highlighting its potential threat to DNNs. Compared to existing attacks, RetouchUAA offers several notable advantages. Firstly, RetouchUAA excels in generating interpretable and realistic perturbations through two key designs: the image retouching attack framework and the retouching style guidance module. The former custom-designed human-interpretability retouching framework for adversarial attack by linearizing images while modelling the local processing and retouching decision-making in human retouching behaviour, provides an explicit and reasonable pipeline for understanding the robustness of DNNs against retouching. The latter guides the adversarial image towards standard retouching styles, thereby ensuring its realism. Secondly, attributed to the design of the retouching decision regularization and the persistent attack strategy, RetouchUAA also exhibits outstanding attack capability and defense robustness, posing a heavy threat to DNNs. Experiments on ImageNet, Place365 and CUB200 reveal that RetouchUAA achieves nearly 100% white-box attack success against three DNNs, while achieving a better trade-off between image naturalness, transferability and defense robustness than baseline attacks. Mengda Xie, Yiling He, Zhan Qin, Meie Fang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | VGNet: Multimodal Feature Extraction and Fusion Network for 3D CAD Model RetrievalabstractThe reuse of 3D CAD models is crucial for industrial manufacturing because it shortens development cycles and reduces costs. Significant progress has been made in deep learning-based 3D model retrievals. There are many representations for 3D models, among which the multi-view representation has demonstrated a superior retrieval performance. However, directly applying these 3D model retrieval approaches to 3D CAD model retrievals may result in issues such as the loss of the engineering semantic and structural information. In this paper, we find that multiple views and B-rep can complement each other. Therefore, we propose the view graph neural network (VGNet), which effectively combines multiple views and B-rep to accomplish 3D CAD model retrieval. More specifically, based on the characteristics of the regular shape of 3D CAD models, and the richness of the attribute information in the B-rep attribute graph, we separately design two feature extraction networks for each modality. Moreover, to explore the latent relationships between the multiple views and B-rep attribute graphs, a multi-head attention enhancement module is designed. Furthermore, the multimodal fusion module is adopted to make the joint representation of the 3D CAD models more discriminative by using a correlation loss function. Experiments are carried out on a real manufacturing 3D CAD dataset and a public dataset to validate the effectiveness of the proposed approach. Fei-wei Qin, Gaoyang Zhan, Meie Fang, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Multim. | 3 |
| 2025 | Continuous Bijection Supervised Pyramid Diffeomorphic Deformation for Learning Tooth Meshes From CBCT ImagesabstractAccurate and high-quality tooth mesh generation from cone-beam computerized tomography (CBCT) is an essential computer-aided technology for digital dentistry. However, existing segmentation-based methods require complicated post-processing and significant manual correction to generate regular tooth meshes. In this paper, we propose a method of continuous bijection supervised pyramid diffeomorphic deformation (PDD) for learning tooth meshes, which could be used to directly generate high-quality tooth meshes from CBCT Images. Overall, we adopt a classic two-stage framework. In the first stage, we devise an enhanced detector to accurately locate and crop every tooth. In the second stage, a PDD network is designed to deform a sphere mesh from low resolution to high one according to pyramid flows based on diffeomorphic mesh deformations, so that the generated mesh approximates the ground truth infinitely and efficiently. To achieve that, a novel continuous bijection distance loss on the diffeomorphic sphere is also designed to supervise the deformation learning, which overcomes the shortcoming of loss based on nearest-neighbour mapping and improves the fitting precision. Experiments show that our method outperforms the state-of-the-art methods in terms of both different evaluation metrics and the geometry quality of reconstructed tooth surfaces. Zechu Zhang, Weilong Peng, Jinyu Wen, Keke Tang, Meie Fang, David Dagan Feng, Ping Li 0016 |
IEEE Trans. Multim. | 5 |
| 2025 | CADGCL: unsupervised retrieval of CAD models via boundary representationsabstractAbstract With the widespread application of CAD technology in the industrial manufacturing sector, the efficient retrieval of target models has become a critical research topic. Despite the outstanding performance of traditional supervised retrieval methods, their reliance on large amounts of labeled data significantly limits practical applications. Data labeling is not only time-consuming and costly but also difficult to ensure accuracy and consistency. To tackle this issue, this paper introduces CADGCL, an unsupervised method for retrieving CAD models based on boundary representations. The proposed method transforms CAD models represented by boundary representations into B-rep attributed graphs that integrate geometric information and topological structures. Graph contrastive learning facilitates unsupervised CAD model retrieval. To overcome the limitations of traditional GCL methods in data augmentation and negative sampling, two novel strategies are introduced: an edge perturbation strategy based on Edge Betweenness Centrality and a negative sampling strategy based on the Beta Mixture Model. These strategies effectively improve the performance of contrastive learning. Experimental results show that the proposed method outperforms existing approaches in mAP and F1 scores under unsupervised scenarios, validating its potential for applications in industrial manufacturing. Fei-wei Qin, Liangzhe Zhu, Zijian Xu 0009, Meie Fang, Ping Li 0016 |
Vis. Comput. | 4 |
| 2025 | Adversarial relighting attacks: physically interpretable manipulation of incident light for vision model vulnerability exploration
Chengzhi Zhong, Mengda Xie, Yiling He, Ping Li 0016, Meie Fang |
Vis. Comput. | 5 |
| 2025 | Topology-guided accelerated vector field streamline visualization
Junjie Yin, Yilun Yang, Meie Fang, Ping Li 0016 |
Vis. Comput. | 4 |
| 2024 | Image-text Retrieval with Main Semantics ConsistencyabstractImage-text retrieval (ITR) has been one of the primary tasks in cross-modal retrieval, serving as a crucial bridge between computer vision and natural language processing. Significant progress has been made to achieve global alignment and local alignment between images and texts by mapping images and texts into a common space to establish correspondences between these two modalities. However, the rich semantic content contained in each image may bring false matches, resulting in the matched text ignoring the main semantics but focusing on the secondary or other semantics of this image. To address this issue, this paper proposes a semantically optimized approach with a novel Main Semantics Consistency (MSC) loss function, which aims to rank the semantically most similar images (or texts) corresponding to the given query at the top position during the retrieval process. First, in each batch of image-text pairs, we separately compute (i) the image-image similarity, i.e., the similarity between every two images, (ii) the text-text similarity, i.e., the similarity between a group of texts (that belong to a certain image) and another group of texts (that belong to another image), and (iii) the image-text similarity, i.e., the similarity between each image and each text. Afterward, our proposed MSC effectively aligns the above image-image, image-text, and text-text similarity, since the main semantics of every two images will be highly close if their text descriptions remain highly semantically consistent. By this means, we can capture the main semantics of each image to be matched with its corresponding texts, prioritizing the semantically most related retrieval results. Extensive experiments on MSCOCO and FLICKR30K verify the superior performance of MSC compared with the SOTA image-text retrieval methods. The source code of this project is released at GitHub: https://github.com/xyi007/MSC. Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Jingjing Li 0001, Xiaocui Li 0001, Weilong Peng, Maobin Tang, Meie Fang |
CIKM | 9 |
| 2024 | IE-aware Consistency Losses for Detailed 3D Face Reconstruction from Multiple Images in the Wildabstract3D face reconstruction from multiple in-the-wild images in an unsupervised manner poses a significant challenge, primarily due to the pervasive presence of Intrinsic and Extrinsic inconsistencies in facial features. To tackle this, we introduce a novel set of IE-aware consistency losses designed to effectively mitigate these inconsistencies. Our Local Alignment Loss employs neighborhood search techniques to identify and optimize consistent pixel information, thereby reducing intrinsic inconsistencies. In parallel, our Region Subset Selection Loss filters out regions where significant discrepancies exist between the input and reconstructed images, effectively alleviating extrinsic inconsistencies. Extensive experimental results validate the effectiveness of our IE-aware consistency losses in reconstructing detailed 3D facial geometry from images captured in uncontrolled environments. Weilong Peng, Keke Tang, Kongyang Chen, Yangtao Wang, Ping Li 0016, Meie Fang |
ICME | 7 |
| 2024 | MIT: Multi-cue Injected Transformer for Two-Stage HOI Detection
Weilong Peng, Qingfeng Chen, Keke Tang, Meng Xing, Meie Fang |
PRCV (7) | 6 |
| 2024 | High-Quality Fusion and Visualization for MR-PET Brain Tumor Images via Multi-Dimensional FeaturesabstractThe fusion of magnetic resonance imaging and positron emission tomography can combine biological anatomical information and physiological metabolic information, which is of great significance for the clinical diagnosis and localization of lesions. In this paper, we propose a novel adaptive linear fusion method for multi-dimensional features of brain magnetic resonance and positron emission tomography images based on a convolutional neural network, termed as MdAFuse. First, in the feature extraction stage, three-dimensional feature extraction modules are constructed to extract coarse, fine, and multi-scale information features from the source image. Second, at the fusion stage, the affine mapping function of multi-dimensional features is established to maintain a constant geometric relationship between the features, which can effectively utilize structural information from a feature map to achieve a better reconstruction effect. Furthermore, our MdAFuse comprises a key feature visualization enhancement algorithm designed to observe the dynamic growth of brain lesions, which can facilitate the early diagnosis and treatment of brain tumors. Extensive experimental results demonstrate that our method is superior to existing fusion methods in terms of visual perception and nine kinds of objective image fusion metrics. Specifically, in the results of MR-PET fusion, the SSIM (Structural Similarity) and VIF (Visual Information Fidelity) metrics show improvements of 5.61% and 13.76%, respectively, compared to the current state-of-the-art algorithm. Our project is publicly available at: https://github.com/22385wjy/MdAFuse. Jinyu Wen, Amei Chen, Weilong Peng, Meie Fang, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Image Process. | 5 |
| 2024 | MsgFusion: Medical Semantic Guided Two-Branch Network for Multimodal Brain Image FusionabstractMultimodal image fusion plays an essential role in medical image analysis and application, where computed tomography (CT), magnetic resonance (MR), single-photon emission computed tomography (SPECT), and positron emission tomography (PET) are commonly-used modalities, especially for brain disease diagnoses. Most existing fusion methods do not consider the characteristics of medical images, and they adopt similar strategies and assessment standards to natural image fusion. While distinctive medical semantic information (MS-Info) is hidden in different modalities, the ultimate clinical assessment of the fusion results is ignored. Our MsgFusion first builds a relationship between the key MS-Info of the MR/CT/PET/SPECT images and image features to guide the CNN feature extractions using two branches and the design of the image fusion framework. For MR images, we combine the spatial domain feature and frequency domain feature (SF) to develop one branch. For PET/SPECT/CT images, we integrate the gray color space feature and adapt the HSV color space feature (GV) to develop another branch. A classification-based hierarchical fusion strategy is also proposed to reconstruct the fusion images to persist and enhance the salient MS-Info reflecting anatomical structure and functional metabolism. Fusion experiments are carried out on many pairs of MR-PET/SPECT and MR-CT images. According to seven classical objective quality assessments and one new subjective clinical quality assessment from 30 clinical doctors, the fusion results of the proposed MsgFusion are superior to those of the existing representative methods. Jinyu Wen, Fei-wei Qin, Jiao Du, Meie Fang, Xinhua Wei, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Multim. | 4 |
| 2023 | Camera-independent color constancy by scene semantics
Mengda Xie, Peng Sun 0009, Yubo Lang, Meie Fang |
Pattern Recognit. Lett. | 4 |
| 2023 | RepPVConv: attentively fusing reparameterized voxel features for efficient 3D point cloud perception
Keke Tang, Weilong Peng, Yanling Zhang, Meie Fang, Zheng Wang 0002, Peng Song 0001 |
Vis. Comput. | 5 |
| 2022 | Salient Object Detection Based on Multiscale Segmentation and Fuzzy Broad LearningabstractAbstract Saliency detection has been a hot topic in the field of computer vision. In this paper, we propose a novel approach that is based on multiscale segmentation and fuzzy broad learning. The core idea of our method is to segment the image into different scales, and then the extracted features are fed to the fuzzy broad learning system (FBLS) for training. More specifically, it first segments the image into superpixel blocks at different scales based on the simple linear iterative clustering algorithm. Then, it uses the local binary pattern algorithm to extract texture features and computes the average color information for each superpixel of these segmentation images. These extracted features are then fed to the FBLS to obtain multiscale saliency maps. After that, it fuses these saliency maps into an initial saliency map and uses the label propagation algorithm to further optimize it, obtaining the final saliency map. We have conducted experiments based on several benchmark datasets. The results show that our solution can outperform several existing algorithms. Particularly, our method is significantly faster than most of deep learning-based saliency detection algorithms, in terms of training and inferring time. Xiao Lin 0012, Zhi-Jie Wang 0009, Lizhuang Ma, Meie Fang |
Comput. J. | 5 |
| 2022 | Cross-modal retrieval based on deep regularized hashing constraintsabstractCross-modal retrieval has attracted great attention due to the increasing demand for tremendous amounts of multimodal data in recent years. These retrievals could either be text-to-image or image-to-text. To address the problem of inappropriate information included between images and texts, we propose two cross-modal recovery techniques established on a dual-branch neural network defined on a common subspace and the hashing learning method. First, a cross-modal recovery technique established on a multilabel information deep ranking model (MIDRM) is provided. In this method, we introduce a triplet-loss function into the dual-branch neural network model. This function takes advantage of the semantic information of the bimodal components, focusing on not only the similarities between similar images and text features but also the distances between dissimilar images and texts. Second, we establish a new cross-modal hashing technique said to be the deep regularized hashing constraint (DRHC). In this method, the regularized function is used to replace the binary constraint, and the discrete value is constrained to a certain numerical range so that the network can achieve end-to-end training. Overall, the time complexity is greatly improved, and the occupied storage space is also greatly reduced. Different experiments on our proposed MIDRM and DRHC models demonstrate their superior performance to those of the state-of-the-art methods on two widely used data sets. The experimental results show that our approach also increases the mean average precision of cross-modal recovery. Sakander Hayat, Muhammad Ahmad 0002, Jinyu Wen, Muhammad Umar Farooq 0002, Meie Fang, Wenchao Jiang |
Int. J. Intell. Syst. | 6 |
| 2022 | Re-transfer learning and multi-modal learning assisted early diagnosis of Alzheimer's disease
Meie Fang, Zhuxin Jin, Fei-wei Qin, Yong Peng 0001 |
Multim. Tools Appl. | 1 |
| 2021 | Single Image Deraining via detail-guided Efficient Channel Attention Network
Xiao Lin 0012, Qi Huang 0005, Xin Tan 0002, Meie Fang, Lizhuang Ma |
Comput. Graph. | 5 |
| 2019 | Automatic Landmark Placement for Large 3D Facial Image DatasetabstractFacial landmark placement is a key step in many biomedical and biometrics applications. This paper presents a computational method that efficiently performs automatic 3D facial landmark placement based on training images containing manually placed anthropological facial landmarks. After 3D face registration by an iterative closest point (ICP) technique, a visual analytics approach is taken to generate local geometric patterns for individual landmark points. These individualized local geometric patterns are derived interactively by a user's initial visual pattern detection. They are used to guide the refinement process for landmark points projected from a template face to achieve accurate landmark placement. Compared to traditional methods, this technique is simple, robust, and does not require a large number of training samples (e.g. in machine learning based methods) or complex 3D image analysis procedures. This technique and the associated software tool are being used in a 3D biometrics project that aims to identify links between human facial phenotypes and their genetic association. Jerry Wang, Shiaofen Fang, Meie Fang, Jeremy Wilson, Noah Herrick, Susan Walsh |
IEEE BigData | 3 |
| 2019 | MCCH: A novel convex hull prior based solution for saliency detection
Xiao Lin 0012, Zhi-Jie Wang 0009, Xin Tan 0002, Meie Fang, Naixue Xiong, Lizhuang Ma |
Inf. Sci. | 4 |
| 2017 | Re2l: An efficient output-sensitive algorithm for computing Boolean operations on circular-arc polygons and its applications
Zhi-Jie Wang 0009, Xiao Lin 0012, Meie Fang, Bin Yao 0002, Yong Peng 0001, Haibing Guan, Minyi Guo |
Comput. Aided Des. | 3 |
| 2016 | Efficient decolorization preserving dominant distinctions
Zhongping Ji, Meie Fang, Yigang Wang, Weiyin Ma |
Vis. Comput. | 2 |
| 2015 | High-quality topological structure extraction of volumetric data on C2-continuous framework
Weisi Gu, Meie Fang, Lizhuang Ma |
Comput. Aided Geom. Des. | 2 |
| 2014 | A generalized surface subdivision scheme of arbitrary order with a tension parameter
Meie Fang, Weiyin Ma, Guozhao Wang |
Comput. Aided Des. | 1 |
| 2014 | Volumetric data modeling and analysis based on seven-directional box spline
Meie Fang, Qunsheng Peng 0001 |
Sci. China Inf. Sci. | 1 |
| 2010 | A generalized curve subdivision scheme of arbitrary order with a tension parameter
Meie Fang, Weiyin Ma, Guozhao Wang |
Comput. Aided Geom. Des. | 1 |
| 2010 | N-way blending problem of circular quadrics
Meie Fang, Guozhao Wang, Weiyin Ma |
Sci. China Inf. Sci. | 1 |
| 2009 | Blending circular quadrics with parametric patchesabstractA method of blending circular quadrics with parametric patches is proposed in this paper. It needs n rational bicubic Bezier patches and two S-patches to blend n (n > 2) quadrics. The blend is G1continuous. Explicit formulae of control points of both Beacutezier patches and S-patches are derived from the corresponding G1-continuity conditions. In addition, the shape can be intuitively modified by adjusting the free parameters of the blending surfaces. Meie Fang, Guozhao Wang, Weiyin Ma |
CAD/Graphics | 1 |
| 2008 | omegaB-splines
Meie Fang, Guozhao Wang |
Sci. China Ser. F Inf. Sci. | 1 |
| 2007 | ω-BezierabstractA new kind of Bézier-like basis with a frequency parameter, called ω-Bezier basis, is presented. It unifies and extends the Bézier basis, C-Bézier basis and H-Bézier basis defined over polynomial space, trigonometric polynomial space and hyperbolic polynomial space respectively. The ω-Bezier basis is defined in the space spanned by {cosωt, sinωt, 1, t,..., tk-2}, where ω = α, αϵR, κ is an arbitrary nonnegative integer. The ω-Bezier basis persists all desirable properties of the existing Bézier-like bases. Furthermore, it also has some special properties advantageous for modeling free form curves and surfaces, for example shape adjustability relative to the frequency parameter. Meie Fang, Guozhao Wang |
CAD/Graphics | 1 |