EDBT 2026 Demo / reviewers in the wild / expert
Wei Ma 0008
dblp:32/32-8
· DBLP profile ↗
41ranked-venue papers
10as first author
25since 2021 · last 2026
0000-0001-9652-4260ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RealNet: Efficient and Unsupervised Detection of AI-Generated Images via Real-Only Representation LearningabstractDetecting AI-generated images remains a persistent challenge, as existing detectors often struggle to generalize to forgeries produced by previously unseen generative models. This generalization gap mainly stems from entanglement with semantic content and overfitting to model-specific artifacts. Moreover, many state-of-the-art methods rely on large pre-trained backbones or computationally intensive pipelines, which limit their applicability in real-world, resource-constrained environments. We propose RealNet, a lightweight and unsupervised framework that constructs a disentangled, forgery-aware representation space using only real images. RealNet first extracts semantic-agnostic representations through a dual adversarial denoising mechanism, producing compact features with low intra-class variance. These representations are then perturbed in feature space to generate pseudo-negative samples, which are combined with the original real features to train a lightweight discriminator, enabling robust detection without any dependence on synthetic images during training. Comprehensive evaluations across GAN, diffusion, and emerging VAR-based paradigms demonstrate that RealNet achieves superior cross-model generalization and robustness. RealNet surpasses previous state-of-the-art approaches by 4.51% in accuracy and 3.93% in average precision, while maintaining significantly lower computational cost. Furthermore, we introduce a medically relevant synthetic image dataset and show RealNet remains effective under severe distribution shifts, highlighting its potential for deployment in high-stakes real-world scenarios. Together, these advantages position RealNet as a practical, scalable and socially impactful solution for robust AI-generated image detection. Shuaibo Li, Laixin Zhang, Wei Ma 0008, Jianwei Guo 0003, Shibiao Xu, Zhijie Qiu, Hongbin Zha |
AAAI | 3 |
| 2026 | DualScope: Capturing Critical Spatial and Temporal Cues for Distracted Driving Activity RecognitionabstractAccurately recognizing distracted driving activities in real-world scenarios is essential for improving road and pedestrian safety. However, existing approaches are prone to attending to irrelevant scene context and are susceptible to interference from redundant frames, compromising their robustness in complex driving environments. To overcome these limitations, we propose DualScope, a novel framework that captures behaviorally critical information from both spatial and temporal perspectives. In the spatial domain, we introduce a Synergistic Behavior-Centric Distillation mechanism that leverages two key information sources: (1) position-aware knowledge derived from the SAM model, which enhances the perception of critical regions and their semantic interaction structures; and (2) fine-grained visual details obtained from cropped key regions, which improve the model's ability to capture detailed patterns within behavior-relevant areas. In the temporal domain, we present the Saliency-Aware Fine-to-Coarse Temporal Modeling module, comprising three components: a Fine-Grained Motion Encoder for capturing local inter-frame dependencies; a Dynamic Difference Extractor for generating salient motion dynamics; and a Saliency-Aware Temporal Pyramid Mamba for integrating these representations to enable multi-scale temporal modeling. This design effectively captures both short-term motions and long-term behavioral patterns. Furthermore, incorporating salient dynamics enhances the model's focus on significant behavioral variations. Extensive experiments on seven publicly available DDAR datasets demonstrate that DualScope consistently outperforms state-of-the-art methods, validating its effectiveness in capturing behavioral cues across spatial and temporal dimensions. Zhijie Qiu, Shuaibo Li, Laixin Zhang, Xuming Hu, Wei Ma 0008 |
AAAI | 5 |
| 2026 | Rotation and semantic co-aware Transformer for oriented object detection in remote sensing images
Shaojun Lv, Wei Ma 0008, Hongbin Zha |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Enhancing wheat pest detection: an edge-enhanced deformable attention network approach
Dongxue Liu, Yingchun Yuan, Qing En, Wei Ma 0008, Chunshan Wang, Zhenxue He, Fangfang Liang |
Vis. Comput. | 4 |
| 2025 | Training-free Fourier Phase Diffusion for Style TransferabstractDiffusion models have shown significant potential for image style transfer tasks. However, achieving effective stylization while preserving content in a training-free setting remains a challenging issue due to the tightly coupled representation space and inherent randomness of the models. In this paper, we propose a Fourier phase diffusion model that addresses this challenge. Given that the Fourier phase spectrum encodes an image's edge structures, we propose modulating the intermediate diffusion samples with the Fourier phase of a content image to conditionally guide the diffusion process. This ensures content retention while fully utilizing the diffusion model's style generation capabilities. To implement this, we introduce a content phase spectrum incorporation method that aligns with the characteristics of the diffusion process, preventing interference with generative stylization. To further enhance content preservation, we integrate homomorphic semantic features extracted from the content image at each diffusion stage. Extensive experimental results demonstrate that our method outperforms state-of-the-art models in both content preservation and stylization. Code is available at https://github.com/zhang2002forwin/Fourier-Phase-Diffusion-for-Style-Transfer. Wei Ma 0008, Hongbin Zha |
IJCAI | 2 |
| 2025 | Structured 3D gaussian splatting for novel view synthesis based on single RGB-LiDAR View
Zhiqun Zhao, Wei Ma 0008, Hongbin Zha |
Appl. Intell. | 3 |
| 2025 | Multi-Task Gradual Inference with a Single Encoder-Decoder Network for Automatic Portrait MattingabstractThis paper presents a multi-task gradual inference model, MTGINet, for automatic portrait matting. It handles the subtasks of automatic portrait matting, namely portrait-transition-background trimap segmentation and transition region matting, with a single encoder-decoder structure. First, we enrich the highest stage of features from the encoder with portrait shape context via a shape context aggregation (SCA) module for trimap segmentation. Then, we fuse the SCA-enhanced features with detailed clues from the encoder for transition-region-aware alpha matting. The gradual inference model naturally allows sufficient interaction between the subtasks via forward computation and backwards propagation during training, and therefore achieves high accuracy while maintaining low complexity. In addition, considering the discrepancies in feature requirements across subtasks, we adapt the features from the encoders before reusing them via a feature rectification module. In addition to the MTGINet model, we have constructed a new large-scale dataset, HPM-17K, for half-body portrait matting. It consists of 16,967 images with diverse backgrounds. Comparative experiments with existing deep models on the public P3M-10K dataset and our HPM-17K dataset demonstrate that the proposed model exhibits state-of-the-art performance. Wenbing Yang, Wei Ma 0008, Qing Mi, Hongbin Zha |
Comput. Vis. Media | 2 |
| 2025 | Learning and aggregating principal semantics for semantic edge detection in images
Lijun Dong, Wei Ma 0008, Hongbin Zha |
Expert Syst. Appl. | 2 |
| 2025 | Attention-based unsupervised prompt learning for SAM in leaf disease segmentation
Luda Tian, Yingchun Yuan, Qing En, Wei Ma 0008, Fangfang Liang |
Knowl. Based Syst. | 4 |
| 2025 | EdgeMaskFormer: Adapting Mask Transformer for Semantic Edge DetectionabstractSemantic Edge Segmentation (SED) is crucial for intelligent agents to understand and interact with their environments, as it enables them to locate and recognize semantic boundaries. The prevailing framework in the field of SED is multi-label learning, which identifies edges and their semantics by learning to assign multiple labels that indicate the categories of the objects forming the edges. However, this framework has demonstrated limited performance when dealing with complex scenarios. In this paper, we propose a mask classification framework specifically tailored for the SED task, termed EdgeMaskFormer. Within this framework, we develop a query-based edge semantic extractor to learn semantic embeddings for edge mask classification with assistance from regional semantic supervision. Additionally, we design a context-aware hierarchical edge extractor to serve as an edge mask head, which can capture multi-scale edges of different categories under the guidance from the semantic embeddings via dynamic convolution. Furthermore, we develop matching and supervision mechanisms specifically for edge mask classification in order to reduce edge noise and address the imbalance between edge and non-edge samples. Our extensive experiments on three public datasets demonstrate that the proposed approach achieves outstanding performance in semantic edge detection, particularly on those datasets with complex scenarios. Lijun Dong, Wei Ma 0008, Hongbin Zha |
IEEE Trans. Multim. | 2 |
| 2024 | UnionFormer: Unified-Learning Transformer with Multi-View Representation for Image Manipulation Detection and LocalizationabstractWe present UnionFormer, a novel framework that inte-grates tampering clues across three views by unified learning for image manipulation detection and localization. Specifically, we construct a BSFI-Net to extract tampering features from RGB and noise views, achieving enhanced responsive-ness to boundary artifacts while modulating spatial consis-tency at different scales. Additionally, to explore the incon-sistency between objects as a new view of clues, we combine object consistency modeling with tampering detection and localization into a three-task unified learning process, allowing them to promote and improve mutually. Therefore, we acquire a unified manipulation discriminative representation under multi-scale supervision that consolidates information from three views. This integration facilitates highly effective concurrent detection and localization of tampering. We perform extensive experiments on diverse datasets, and the results show that the proposed approach outperforms state-of-the-art methods in tampering detection and localization. Shuaibo Li, Wei Ma 0008, Jianwei Guo 0003, Shibiao Xu, Benchong Li, Xiaopeng Zhang 0001 |
CVPR | 2 |
| 2024 | Dual-modal non-local context guided multi-stage fusion for indoor RGB-D semantic segmentation
Wei Ma 0008, Fangfang Liang, Qing Mi |
Expert Syst. Appl. | 2 |
| 2023 | View-relation constrained global representation learning for multi-view-based 3D object recognition
Ruchang Xu, Qing Mi, Wei Ma 0008, Hongbin Zha |
Appl. Intell. | 3 |
| 2023 | A graph-based code representation method to improve code readability classification
Qing Mi, Han Weng, Qinghang Bao, Longjie Cui, Wei Ma 0008 |
Empir. Softw. Eng. | 6 |
| 2023 | DSC-MDE: Dual structural contexts for monocular depth estimation
Wubin Yan, Lijun Dong, Wei Ma 0008, Qing Mi, Hongbin Zha |
Knowl. Based Syst. | 3 |
| 2023 | Image Manipulation Localization Using Attentional Cross-Domain CNN FeaturesabstractAlong with the advancement of manipulation technologies, image modification is becoming increasingly convenient and imperceptible. To tackle the challenging image tampering detection problem, this article presents an attentional cross-domain deep architecture, which can be trained end-to-end. This architecture is composed of three convolutional neural network (CNN) streams to extract three types of features, including visual perception, resampling, and local inconsistency features, from spatial and frequency domains. The multitype and cross-domain features are then combined to formulate hybrid features to distinguish manipulated regions from nonmanipulated parts. Compared with other deep architectures, the proposed one spans a more complementary and discriminative feature space by integrating richer types of features from different domains in a unified end-to-end trainable framework and thus can better capture artifacts caused by different types of manipulations. In addition, we design and train a module called tampering discriminative attention network (TDA-Net) to highlight suspicious parts. These part-level representations are then integrated with the global ones to further enhance the discriminating capability of the hybrid features. To adequately train the proposed architecture, we synthesize a large dataset containing various types of manipulations based on DRESDEN and COCO. Experiments on four public datasets demonstrate that the proposed model can localize various manipulations and achieve the state-of-the-art performance. We also conduct ablation studies to verify the effectiveness of each stream and the TDA-Net module. Shuaibo Li, Shibiao Xu, Wei Ma 0008, Qiu Zong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | ReINView: Re-interpreting Views for Multi-view 3D Object RecognitionabstractMulti-view-based 3D object recognition is important in robot-environment interaction. However, recent methods simply extract features from each view via convolutional neural networks (CNNs) and then fuse these features together to make predictions. These methods ignore the inherent ambiguities of each view caused due to 3D-2D projection. To address this problem, we propose a novel deep framework for multi-view-based 3D object recognition. Instead of fusing the multi-view features directly, we design a re-interpretation module (ReINView) to eliminate the ambiguities at each view. To achieve this, ReINView re-interprets view features patch by patch by using their context from nearby views, considering that local patches are generally co-visible at nearby viewpoints. Since contour shapes are essential for 3D object recognition as well, ReINView further performs view-level re-interpretation, in which we use all the views as context sources since the target contours to be re-interpreted are globally observable. The re-interpreted multi-view features can better reflect the 3D global and local structures of the object. Experiments on both ModelNet40 and ModelNet10 show that the proposed model outperforms state-of-the-art methods in 3D object recognition. Ruchang Xu, Wei Ma 0008, Qing Mi, Hongbin Zha |
IROS | 2 |
| 2022 | Towards using visual, semantic and structural features to improve code readability classification
Qing Mi, Yiqun Hao, Liwei Ou, Wei Ma 0008 |
J. Syst. Softw. | 4 |
| 2022 | Multiview Feature Aggregation for Facade ParsingabstractFacade image parsing is essential to the semantic understanding and 3-D reconstruction of urban scenes. Considering the occlusion and appearance ambiguity in single-view images and the easy acquisition of multiple views, in this letter, we propose a multiview enhanced deep architecture for facade parsing. The highlight of this architecture is a cross-view feature aggregation module that can learn to choose and fuse useful convolutional neural network (CNN) features from nearby views to enhance the representation of a target view. Benefitting from the multiview enhanced representation, the proposed architecture can better deal with the ambiguity and occlusion issues. Moreover, our cross-view feature aggregation module can be straightforwardly integrated into existing single-image parsing frameworks. Extensive comparison experiments and ablation studies are conducted to demonstrate the good performance of the proposed method and the validity and transportability of the cross-view feature aggregation module. Wenguang Ma, Shibiao Xu, Wei Ma 0008, Hongbin Zha |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | All-Higher-Stages-In Adaptive Context Aggregation for Semantic Edge DetectionabstractConvolutional Neural Networks (CNNs) can reveal local variation details and multi-scale spatial context in images via low-to-high stages of feature expression; effective fusion of these raw features is key to Semantic Edge Detection (SED). The methods available in the field generally fuse features across stages in a position-aligned mode, which cannot satisfy the requirements of diverse semantic context in categorizing different pixels. In this paper, we propose a deep framework for SED, the core of which is a new multi-stage feature fusion structure, called All-HiS-In ACA (All-Higher-Stages-In Adaptive Context Aggregation). All-HiS-In ACA can adaptively select semantic context from all higher-stages for detailed features via a cross-stage self-attention paradigm, and thus can obtain fused features with high-resolution details for edge localization and rich semantics for edge categorization. In addition, we develop a non-parametric Inter-layer Complementary Enhancement (ICE) module to supplement clues at each stage with their counterparts in adjacent stages. The ICE-enhanced multi-stage features are then fed into the All-HiS-In ACA module. We also construct an Object-level Semantic Integration (OSI) module to further refine the fused features by enforcing the consistency of the features within the same object. Extensive experiments demonstrate the superior performance of the proposed method over state-of-the-art works. Qihan Bo, Wei Ma 0008, Yukun Lai, Hongbin Zha |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Progressive Feature Learning for Facade Parsing With OcclusionsabstractExisting deep models for facade parsing often fail in classifying pixels in heavily occluded regions of facade images due to the difficulty in feature representation of these pixels. In this paper, we solve facade parsing with occlusions by progressive feature learning. To this end, we locate the regions contaminated by occlusions via Bayesian uncertainty evaluation on categorizing each pixel in these regions. Then, guided by the uncertainty, we propose an occlusion-immune facade parsing architecture in which we progressively re-express the features of pixels in each contaminated region from easy to hard. Specifically, the outside pixels, which have reliable context from visible areas, are re-expressed at early stages; the inner pixels are processed at late stages when their surroundings have been decontaminated at the earlier stages. In addition, at each stage, instead of using regular square convolution kernels, we design a context enhancement module (CEM) with directional strip kernels, which can aggregate structural context to re-express facade pixels. Extensive experiments on popular facade datasets demonstrate that the proposed method achieves state-of-the-art performance. Wenguang Ma, Shibiao Xu, Wei Ma 0008, Xiaopeng Zhang 0001, Hongbin Zha |
IEEE Trans. Image Process. | 3 |
| 2021 | Siamese CNN-based rank learning for quality assessment of inpainted images
Xiangdong Meng, Wei Ma 0008, Chunhu Li, Qing Mi |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Pyramid ALKNet for Semantic Parsing of Building Facade ImageabstractThe semantic parsing of building facade images is a fundamental yet challenging task in urban scene understanding. Existing works sought to tackle this task by using facade grammars or convolutional neural networks (CNNs). The former can hardly generate parsing results coherent with real images while the latter often fails to capture relationships among facade elements. In this letter, we propose a pyramid atrous large kernel (ALK) network (ALKNet) for the semantic segmentation of facade images. The pyramid ALKNet captures long-range dependencies among building elements by using ALK modules in multiscale feature maps. It makes full use of the regular structures of facades to aggregate useful nonlocal context information and thereby is capable of dealing with challenging image regions caused by occlusions, ambiguities, and so on. Experiments on both rectified and unrectified facade data sets show that ALKNet has better performances than those of state-of-the-art methods. Wenguang Ma, Wei Ma 0008, Shibiao Xu, Hongbin Zha |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Context-aware network for RGB-D salient object detection
Fangfang Liang, Lijuan Duan, Wei Ma 0008, Yuanhua Qiao, Qixiang Ye |
Pattern Recognit. | 3 |
| 2021 | Line Flow Based Simultaneous Localization and MappingabstractIn this article, we propose a visual simultaneous localization and mapping (SLAM) method by predicting and updating line flows that represent sequential 2-D projections of 3-D line segments. While feature-based SLAM methods have achieved excellent results, they still face problems in challenging scenes containing occlusions, blurred images, and repetitive textures. To address these problems, we leverage a line flow to encode the coherence of line segment observations of the same 3-D line along the temporal dimension, which has been neglected in prior SLAM systems. Thanks to this line flow representation, line segments in a new frame can be predicted according to their corresponding 3-D lines and their predecessors along the temporal dimension. We create, update, merge, and discard line flows on-the-fly. We model the proposed line flow based SLAM (LF-SLAM) using a Bayesian network. Extensive experimental results demonstrate that the proposed LF-SLAM method achieves state-of-the-art results due to the utilization of line flows. Specifically, LF-SLAM obtains good localization and mapping results in challenging scenes with occlusions, blurred images, and repetitive textures. Qiuyuan Wang, Zike Yan, Junqiu Wang, Wei Ma 0008, Hongbin Zha |
IEEE Trans. Robotics | 5 |
| 2020 | FC-vSLAM: Integrating Feature Credibility in Visual SLAMabstractFeature-based visual SLAM (vSLAM) systems compute camera poses and scene maps by detecting and matching 2D features, mostly being points and line segments, from image sequences. These systems often suffer from unreliable detections. In this paper, we define feature credibility (FC) for both points and line segments, formulate it into vSLAMs and develop an FC-vSLAM system based on the widely used ORB-SLAM framework. Compared with existing credibility definitions, the proposed one, considering both temporal observation stability and perspective triangulation reliability, is more comprehensive. We formulate the credibility in our SLAM system to suppress the influences from unreliable features on the pose and map optimization. We also present a way to improve the line end observations by their multi-view correspondences, to improve the integrity of the 3D maps. Experiments on both the TUM and 7-Scenes datasets demonstrate that our feature credibility and the multi-view line optimization are effective; the developed FC-vSLAM system outperforms existing popular feature-based systems in both localization and mapping. Shuai Xie, Wei Ma 0008, Qiuyuan Wang, Ruchang Xu, Hongbin Zha |
3DV | 2 |
| 2020 | Learning across views for stereo image completionabstractStereo image completion (SIC) is to fill holes existing in a pair of stereo images. SIC is more complicated than single image repairing, which needs to complete the pair of images while keeping their stereoscopic consistency. In recent years, deep learning has been introduced into single image repairing but seldom used for SIC. The authors present a novel deep learning‐based approach for SIC. In their method, an X‐shaped fully convolutional network (called SICNet) is proposed and designed to complete stereo images, which is composed of two branches of convolutional neural network layers to encode the context of the left and right images separately, a fusion module for stereo‐interactive completion, and two branches of decoders to produce completed left and right images, respectively. In consideration of both inter‐view and intra‐view cues, they introduce auxiliary networks and define comprehensive losses to train SICNet to perform single‐view coherent and cross‐view consistent completion simultaneously. Extensive experiments are conducted to show the state‐of‐the‐art performances of the proposed approach and its key components. Wei Ma 0008, Mana Zheng, Wenguang Ma, Shibiao Xu, Xiaopeng Zhang 0001 |
IET Comput. Vis. | 1 |
| 2020 | CoCNN: RGB-D deep fusion for stereoscopic salient object detection
Fangfang Liang, Lijuan Duan, Wei Ma 0008, Yuanhua Qiao, Zhi Cai, Qixiang Ye |
Pattern Recognit. | 3 |
| 2018 | Walking into ancient paintings with virtual candlesabstractTaking a famous Chinese painting for a case study, the paper presents a virtual exhibition platform. Through the platform, users can walk into the scenes in the painting with virtual candles in hands, know the scenes which are endowed vitality by attaching actor performances, and see every detail of the artwork. The scenes change their light, shades and shadows in real time by the candles, just as real scenes. For real-time candle-moving and light-changing interaction, in implementation, we render the light effects at densely sampled user positions offline, and extract the light, shades and shadows as masks; during online processing, the system merges the artwork with masks chosen by the positions of candles. The system, novel in both design and techniques, has been partially used in the Palace Museum (Beijing). Wei Ma 0008, Qiuyuan Wang, Danqing Shi, Shuo Liu 0009, Congxin Cheng, Qingyuan Shi, Ying-Qing Xu |
VRST | 1 |
| 2018 | Stereoscopic saliency model using contrast and depth-guided-background prior
Fangfang Liang, Lijuan Duan, Wei Ma 0008, Yuanhua Qiao, Zhi Cai, Laiyun Qing |
Neurocomputing | 3 |
| 2018 | Interactive stereo image segmentation via adaptive prior selection
Wei Ma 0008, Shibiao Xu, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 1 |
| 2017 | Efficent Traffic-Sign Recognition with Scale-aware CNN
Shuo Liu 0009, Wei Ma 0008, Qiuyuan Wang, Zheng Liu 0002 |
BMVC | 3 |
| 2016 | Fast interactive stereo image segmentation
Wei Ma 0008, Luwei Yang, Lijuan Duan |
Multim. Tools Appl. | 1 |
| 2016 | Interactive Stereo Image Segmentation With RGB-D Hybrid ConstraintsabstractThis letter presents an approach to extracting a target object interactively from a given pair of stereo images. First, a user marks a few parts of the object and background in either of the two views with strokes. The marked pixels are used to generate the prior models of the foreground and background. Second, a graph is constructed with constraints formulated by the priors of foreground/background, similarities between intraview neighbor pixels and correspondences between interview pixels. Third, two segments of the foreground are extracted from the two views by optimization of the graph via graph cut. Traditional methods generally define the priors and neighbor similarities in RGB space. Differently, the proposed method integrates disparity distributions of foreground/background to enrich the priors and defines the similarity metric between neighbor pixels in RGB-D space. The proposed method that utilizes RGB-D hybrid constraints generates stereo segments with accuracies higher than those obtained by state-of-the-art methods. Wei Ma 0008, Luwei Yang, Shibiao Xu, Xiaopeng Zhang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2014 | A combined model for scan path in pedestrian searchingabstractTarget searching, i.e. fast locating target objects in images or videos, has attracted much attention in computer vision. A comprehensive understanding of factors influencing human visual searching is essential to design target searching algorithms for computer vision systems. In this paper, we propose a combined model to generate scan paths for computer vision to follow to search targets in images. The model explores and integrates three factors influencing human vision searching, top-down target information, spatial context and bottom-up visual saliency, respectively. The effectiveness of the combined model is evaluated by comparing the generated scan paths with human vision fixation sequences to locate targets in the same images. The evaluation strategy is also used to learn the optimal weighting coefficients of the factors through linear search. In the meanwhile, the performances of every single one of the factors and their arbitrary combinations are examined. Through plenty of experiments, we prove that the top-down target information is the most important factor influencing the accuracy of target searching. The effects from the bottom-up visual saliency are limited. Any combinations of the three factors have better performances than each single component factor. The scan paths obtained by the proposed model are optimal, since they are most similar to the human vision fixation sequences. Lijuan Duan, Zeming Zhao, Wei Ma 0008, Jili Gu, Zhen Yang 0004, Yuanhua Qiao |
IJCNN | 3 |
| 2014 | Data-Driven Synthetic Modeling of TreesabstractIn this paper, we develop a data-driven technique to model trees from a single laser scan. A multi-layer representation of the tree structure is proposed to guide the modeling process. In this process, a marching cylinder algorithm is first developed to construct visible branches from the laser scan data. Three levels of crown feature points are then extracted from the scan data to synthesize three layers of non-visible branches. Based on the hierarchical particle flow technique, the branch synthesis method has the advantage of producing visually convincing tree models that are consistent with scan data. User intervention is extremely limited. The robustness of this technique has been validated on both conifer and broadleaf trees. Xiaopeng Zhang 0001, Hongjun Li 0002, Mingrui Dai, Wei Ma 0008, Long Quan |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2010 | Interactive viewpoint-space navigation for visual-audio exhibition of paintingabstractIn this paper, we present a system for exhibiting a Chinese landscape painting about 900 years old. There are three parts in our system: (1) we allocate a voice dubbing or background music, which is treated as a point sound source, onto the 2D painting and obtain its position in the 2D space. All of the audio data are then located in a 3D hidden space, by projecting their 2D positions to the 3D space through a projection model. (2) A two-layer directed graph structure is proposed to well organize the audio data in a 4D space (with 1D temporal and 3D spatial). (3) The exhibition is defined as an active exploration in a viewpoint space, which faces both the image and the 3D world where the sound sources reside. The 3D space and the two-layer graph structure generate a natural and meaningful stereo audio field. Meanwhile, compared to videos with guided walk through, the active exploration makes the exhibition more attractive. Wei Ma 0008, Yang Liu 0006, Yizhou Wang 0001, Ying-Qing Xu, Hongbin Zha, Wen Gao 0001 |
ICME | 1 |
| 2009 | Modeling plants with sensor data
Wei Ma 0008, Bo Xiang, Hongbin Zha, Xiaopeng Zhang 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2008 | Decomposition of branching volume data by tip detectionabstractWe present an approach to decomposing branching volume data into sub-branches. First, a metric is proposed for evaluating local convexities in volumetric data, and it is a criterion for global selection of tip points. Second, a multi-path growing strategy is adopted to segment the volumes based on a DFS transformation starting from the tips. Experiments show that this approach is capable of generating desirable components and reasonable segmentation boundaries of a volume. Wei Ma 0008, Bo Xiang, Xiaopeng Zhang 0001, Hongbin Zha |
ICIP | 1 |
| 2008 | Convenient reconstruction of natural plants by imagesabstractConvenient reconstruction of natural plants is a difficult task because of their intrinsic complex geometry. In this paper, we propose a convenient image-based approach to modeling and rendering large-leaf plants by view synthesis based on an approximate geometric model. The model we adopt consists of a set of billboard clusters, obtained by automatically decomposing a volume based on the flat property of leaves and assigning billboards corresponding to all input image viewpoints for each cluster. The volume is recovered from a set of images. In the final rendering procedure, a view-dependent billboard in each cluster is generated by interpolating already constructed ones. The whole process can be finished in half an hour and requires no user intervention. Wei Ma 0008, Hongbin Zha |
ICPR | 1 |
| 2008 | Image-based plant modeling by knowing leaves from their apexesabstractIn the paper, we present a novel approach to modeling plants from images by detecting apex features. First, an effective algorithm is proposed to extract apex features in volumetric data recovered from the images. It provides position and pose information for assigning 3D generic leaves. Then, the 3D leaf shapes are determined by an optimization based on the volume. Finally, Branches are modeled by using a particle flow approach. The proposed method is simply with limited manual intervention and has the obvious benefit of knowing a leaf by its visible apex part. Wei Ma 0008, Hongbin Zha, Xiaopeng Zhang 0001, Bo Xiang |
ICPR | 1 |