VLDB 2026 Research / reviewers in the wild / expert
Fang Fang 0008
dblp:74/3719-8
· DBLP profile ↗
22ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-8969-8879ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geo-Object-Reader: a template filling method to jointly extract complex spatial information about geological objectsabstractExtracting spatial information from text subserves data-driven geospatial semantic research. Traditional methods consider words, phrases, or triples to extract spatial entities but often overlook specific spatiotemporal conditions, leading to fragmented representations and potentially inaccurate spatial perceptions. In this study, we present Geo-Object-Reader, a template-based method for the joint spatial information extraction (SIE) of spatial objects and spatiotemporal attributes. The joint extraction highlights an integrated representation of spatial object attributes, relations and their associated spatiotemporal conditions. This study develops three SIE templates tailored to the spatiotemporal characteristics of geospatial objects: spatial attribute template, non-spatial attribute template and 3D spatial relation template. These templates integrate specific spatiotemporal fields to ensure that the extracted attributes and relations are accurate and contextually relevant. A subsequent graph neural network approach captures the contextual information associated with these template fields to apprehend the complex interactions within geoscience texts. The final stage involves the use of a directed acyclic graph (DAG)-based filling strategy to enhance the efficiency of template filling. A dataset constructed based on Chinese geological reports was used to demonstrate that the proposed method provides holistic perspectives, bridging semantic gaps and forming more reliable knowledge chains compared to triples. Deping Chu, Bo Wan 0006, Fang Fang 0008, Shunping Zhou |
Int. J. Geogr. Inf. Sci. | 3 |
| 2025 | CLC²SIE: Cross-Level Consistency Constraint for Semisupervised Building Instance Extraction From High-Resolution Remote Sensing ImageryabstractSemi-supervised building instance extraction aims to learn from limited labeled data alongside an extensive collection of unlabeled data, offering a promising approach for extracting building instances from high-resolution (HR) remote sensing images (RSIs). However, complex RSIs frequently encounter challenges such as background interference and intricate noise, struggling in generating reliable pseudo labels. To alleviate this, we propose a novel cross-level consistency constraint semi-supervised building instance extraction method (CLC2SIE) to enhance pseudo label generation. Specifically, CLC2SIE contains two core modules: object-level dynamic consistency (OLDC) and pixel-level saliency consistency (PLSC). The OLDC module dynamically converts building features from background into valuable supplementary information, enhancing the model’s perception of building instances in complex scenes. Additionally, the PLSC module is designed to mitigate the boundary noise in pseudo labels by saliency guidance, which improves model’s awareness of building contours. By co-learning these modules in an end-to-end manner, CLC2SIE facilitates pseudo label generation and improves extraction performance. Experiments were conducted on the three public building datasets, i.e., WHU, CrowdAI and TCC, demonstrate that CLC2SIE achieves superior performance compared to state-of-the-art semi-supervised instance extraction methods at different labeling ratios. This study explores a novel semi-supervised learning (SSL) framework that exploits cross-level consistency to improve pseudo label generation, offering a methodological reference for various SSL applications in RSIs. Jun Pan 0001, Fang Fang 0008, Daoyuan Zheng, Shengwen Li |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Token-disentangling Mutual Transformer for multimodal emotion recognition
Guanghao Yin, Yuanyuan Liu 0004, Haoyu Zhang 0001, Fang Fang 0008, Chang Tang, Liangxiao Jiang |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | A multi-view ensemble machine learning approach for 3D modeling using geological and geophysical dataabstractGeophysical data are often integrated into geological data for 3D modeling of underground spaces. However, the existing single-view approach means it is difficult to adequately fuse the valid information between the two types of data, and the complexity of lithological decoding and classification is high. To address this issue, a multi-view ensemble machine learning (ML) framework is proposed. Initially, the original dataset of lithology prediction is constructed by aligning geological and geophysical data with different spatial scales. Next, the dataset is divided into three datasets of structural strength, density, and moisture content according to the lithology properties of the geophysical data. The proposed framework is then used to capture the lithologic characteristics under different views to achieve the prediction of lithologic labels. In this process, a self-attentive mechanism is used to adaptively fuse the valid information under each view. To validate the proposed framework, it is applied to a project in Jiaxing, Zhejiang Province, China. Compared with existing ML methods, the proposed multi-view ensemble ML framework improves modeling accuracy and constructs models with low uncertainty. The framework can be extended to other multi-source data fusion tasks across geoscience domains. Deping Chu, Jinming Fu, Bo Wan 0006, Lulan Li, Fang Fang 0008, Shengwen Li, Shengyong Pan, Shunping Zhou |
Int. J. Geogr. Inf. Sci. | 6 |
| 2023 | Spatial Extent-Aware Multimodal Fusion Method for Measuring Urban Socioeconomic StatusabstractThe automatic measurement of socioeconomic status (SES), such as household income, provides fundamental data for policymakers and business applications. Although several urban data sources were employed in previous studies, the spatial extent of the measured objects was ignored, which left room for further improvement in the accuracy of measuring SES. This study develops a multimodal semantic segmentation framework to fuse the spatial extent of ground objects, remote sensing, and near-sensing images for predicting SES levels. The framework first constructs ground feature layer (GFL) tiles by projecting ground-level features from sparse ground images. Then, the ground feature tiles aggregate regional ground-level features with the help of the spatial extents of ranges. Lastly, an improved deep semantic segmentation network is employed to fuse GFL and remote sensing images to predict SES levels. Experimental results on the London dataset show that the proposed method outperforms SOTA models and is robust. The framework provides a rewarding exploration in fusing image data and spatial vector data for image-based intelligent applications and can be applied to a series of socioeconomic applications. Fang Fang 0008, Shengwen Li, Daoyuan Zheng, Linyun Zeng, Bo Wan 0006 |
IGARSS | 1 |
| 2023 | A hierarchical constraint-based graph neural network for imputing urban area dataabstractUrban area data are strategically important for public safety, urban management, and planning. Previous research has attempted to estimate the values of unsampled regular areas, while minimal attention has been paid to the values of irregular areas. To address this problem, this study proposes a hierarchical geospatial graph neural network model based on the spatial hierarchical constraints of areas. The model first characterizes spatial relationships between irregular areas at different spatial scales. Then, it aggregates information from neighboring areas with graph neural networks, and finally, it imputes missing values in fine-grained areas under hierarchical relationship constraints. To investigate the performance of the proposed model, we constructed a new dataset consisting of the urban statistical values of irregular areas in New York City. Experiments on the dataset show that the proposed model outperforms state-of-the-art baselines and exhibits robustness. The model is adaptable to numerous geographic applications, including traffic management, public safety, and public resource allocation. Shengwen Li, Wanchen Yang, Suzhen Huang, Renyao Chen, Xuyang Cheng, Shunping Zhou, Junfang Gong, Haoyue Qian, Fang Fang 0008 |
Int. J. Geogr. Inf. Sci. | 9 |
| 2023 | A deep neural network model for coreference resolution in geological domain
Bo Wan 0006, Deping Chu, Jinming Fu, Fang Fang 0008, Shengwen Li |
Inf. Process. Manag. | 7 |
| 2023 | Semisupervised Building Instance Extraction From High-Resolution Remote Sensing ImageryabstractAutomatic building instance extraction from high-resolution (HR) remote sensing imagery (RSI) is crucial for urban planning and mapping. The dominant approaches are based on the full-supervised learning paradigm that requires a large number of labeled samples to train their models, which is very time-consuming and labor-intensive. To alleviate this problem, this study proposes a semi-supervised building instance extraction method that integrates teacher-student learning and pseudo-labeling to improve the building instance extraction from HR RSI. Specifically, the proposed method consists of three modules, the hybrid data augmentation (HDA) module, the pseudo label generation (PLG) module and the contour refinement (CR) module. The HDA module is designed to enrich the diversity of labeled samples to optimize the teacher model. The PLG module generates pseudo labels from unlabeled data, and to train student model with pseudo-labels. Finally, the CR module is designed to refine the contours of buildings. Experimental results on three challenging public datasets demonstrate that the proposed method achieves superior performance and exhibits great robustness at different proportions of labeled data and different building scenarios. This study provides a new approach for extracting building instances from HR RSI in scenarios with insufficient labeled samples, and a methodological reference for various applications of semi-supervised on RSIs. Fang Fang 0008, Shengwen Li, Qingyi Hao, Kaishun Wu, Bo Wan 0006 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Utilizing Bounding Box Annotations for Weakly Supervised Building Extraction From Remote-Sensing ImagesabstractImage-level weakly supervised semantic segmentation (WSSS) methods have greatly facilitated the extraction of buildings from remote sensing (RS) images. However, the lack of the locations and extents of individual buildings in image-level labels results in some limitations of the methods, especially in the cases of cluttered backgrounds, diverse building shapes and sizes. By utilizing bounding box annotations, a novel WSSS model is developed to improve building extraction from RS images in this paper. Specifically, during the training phase, a multiscale feature retrieval (MFR) module is designed to learn multiscale building features and suppress the background noise inside the bounding box. In the inference phase, multiscale class activation maps (CAM) are generated from multiscale features to achieve accurate building localization. Finally, a pseudo mask generation and correction (PGC) module refines the CAMs to generate and correct the building pseudo masks. Experiments are conducted to examine the proposed model in three datasets, namely, the WHU aerial building dataset, the CrowdAI building dataset, and a self-annotated building dataset. Experimental results demonstrate that the proposed method outperforms baselines, achieving 76.99%, 75.51% and 67.35% in terms of IoU scores on the three challenging datasets, respectively. This paper provides a methodological reference for the application of weakly supervised learning on RS images. Daoyuan Zheng, Shengwen Li, Fang Fang 0008, Bo Wan 0006, Yuanyuan Liu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | DSM-Assisted Unsupervised Domain Adaptive Network for Semantic Segmentation of Remote Sensing Imagery
Shunping Zhou, Shengwen Li, Daoyuan Zheng, Fang Fang 0008, Yuanyuan Liu 0004, Bo Wan 0006 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | ConGNN: Context-consistent cross-graph neural network for group emotion recognition in the wild
Yu Wang 0246, Shunping Zhou, Yuanyuan Liu 0004, Fang Fang 0008, Haoyue Qian |
Inf. Sci. | 5 |
| 2021 | Hierarchical Domain-Consistent Network For Cross-Domain Object DetectionabstractCross-domain object detection is a very challenging task due to multi-level domain shift in an unseen domain. To address the problem, this paper proposes a hierarchical domain-consistent network (HDCN) for cross-domain object detection, which effectively suppresses pixel-level, image-level, as well as instance-level domain shift via jointly aligning three-level features. Firstly, at the pixel-level feature alignment stage, a pixel-level subnet with foreground-aware attention learning and pixel-level adversarial learning is proposed to focus on local foreground transferable information. Then, at the image-level feature alignment stage, global domain-invariant features are learned from the whole image through image-level adversarial learning. Finally, at the instance-level alignment stage, a prototype graph convolution network is conducted to guarantee distribution alignment of instances by minimizing the distance of prototypes with the same category but from different domains. Moreover, to avoid the non-convergence problem during multi-level feature alignment, a domain-consistent loss is proposed to harmonize the adaptation training process. Comprehensive results on various cross-domain detection tasks demonstrate the broad applicability and effectiveness of the proposed approach. Yuanyuan Liu 0004, Fang Fang 0008, Zhanghua Fu, Zhanlong Chen |
ICIP | 3 |
| 2021 | Synthesizing location semantics from street view images to improve urban land-use classificationabstractLand-use maps are instrumental to inform urban planning and environmental research. Street view images (SVIs) have shown great potential for automated land-use classification for land-use mapping. However, previous studies overlooked SVI-derived location contextual information that may help improve land-use classification. This study proposes a novel land-use classification method that synthesizes location semantics from SVIs to account for contextual information from SVIs, land parcels and roads around the SVIs. The proposed method first generates land-use scene images (LUSIs) by using an SVI-derived straightforward algorithm. The LUSIs are then relocated to land parcels by using a displacement strategy and classified into land-use types by using a deep learning network. This study determines the land-use types of land parcels with classified LUSIs. Two case studies, consisting of LUSIs for five land-use types, show that introducing location semantics of SVIs can remarkably improve the classification accuracy of land-use types. Fang Fang 0008, Yafang Yu, Shengwen Li, Zejun Zuo, Yuanyuan Liu 0004, Bo Wan 0006, Zhongwen Luo |
Int. J. Geogr. Inf. Sci. | 1 |
| 2021 | Dynamic multi-channel metric network for joint pose-aware and identity-invariant facial expression recognition
Yuanyuan Liu 0004, Fang Fang 0008, Yongquan Chen, Rui Huang 0001, Run Wang 0002, Bo Wan 0006 |
Inf. Sci. | 3 |
| 2021 | Multiscale U-Shaped CNN Building Instance Extraction Framework With Edge Constraint for High-Spatial-Resolution Remote Sensing ImageryabstractBuilding extraction based on high-resolution remote sensing imagery has been widely used in automatic surveying and mapping. However, few methods have been developed for building instance extraction, i.e., extracting each building's footprint separately, which is required in a number of applications, such as the smallest unit of a cadastral database. In building instance extraction, there are two challenges: 1) buildings with various scales exist in the imagery and 2) precise building footprints are difficult to extract due to the blurry boundaries. In this article, to solve these problems, a multiscale U-shaped convolutional neural network building instance extraction framework with edge constraint (EMU-CNN) for high-spatial-resolution remote sensing imagery is proposed. The proposed framework consists of three components: 1) a multiscale fusion U-shaped network (MFUN); 2) a region proposal network (RPN); and 3) an edge-constrained multitask network (ECMN). First, in the proposed method, the MFUN includes three parallel branches to learn multiple building features with different scales. The RPN then detects the positions of the building instances, even for buildings that are connected with each other. Moreover, according to the instance positions, the ECMN is proposed to extract a precise mask and suppress overfitting. The experiments conducted on a self-annotated data set and two public data sets (the ISPRS Vaihingen semantic labeling contest data set and the WHU aerial image data set) show that the EMU-CNN method can achieve excellent performance and shows great robustness at different scales. Yuanyuan Liu 0004, Ailong Ma, Yanfei Zhong, Fang Fang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Aircraft detection in remote sensing image based on corner clustering and deep learning
Qiangwei Liu, Xiuqiao Xiang, Zhongwen Luo, Fang Fang 0008 |
Eng. Appl. Artif. Intell. | 5 |
| 2019 | An approach for computing routes without complicated decision points in landmark-based pedestrian navigationabstractDuring navigation, a pedestrian needs to recognize a landmark at a certain decision point. If a potential landmark located at a decision point is complicated to recognize, the complexity of the decision point is significantly increased. Thus, it is important to compute routes that avoid complicated decision points (CDPs) but still achieve optimal navigation performance. In this paper, we propose an approach for computing routes that avoid CDPs while optimizing the performance of landmark-based pedestrian navigation. The approach includes (1) a model for identifying CDPs based on the structures of pedestrian networks and landmark data in real scenes, and (2) a modified genetic algorithm for computing routes that avoid the identified CDPs and find the shortest route possible. To demonstrate the advantages and effectiveness of the proposed approach, we conducted an empirical study on the pedestrian network in a real-world scenario. The experimental results show that our approach can effectively avoid CDPs while still minimizing travel distance. Furthermore, our approach can provide the routes with the shortest travel distance if the distances of the routes without CDPs exceed a certain threshold. Run Wang 0002, Junhua Ding 0001, Xiaofang Pan, Shunping Zhou, Fang Fang 0008, Wenjie Zhen |
Int. J. Geogr. Inf. Sci. | 6 |
| 2019 | Visual Focus of Attention and Spontaneous Smile Recognition Based on Continuous Head Pose Estimation by Cascaded Multi-Task LearningabstractMulti-person Visual focus of attention (M-VFOA) and spontaneous smile (SS) recognition are important for persons’ behavior understanding and analysis in class. Recently, promising results have been reported using special hardware in constrained environment. However, M-VFOA and SS remain challenging problems in natural and crowd classroom environment, e.g. various poses, occlusion, expressions, illumination and poor image quality, etc. In this study, a robust and un-invasive M-VFOA and SS recognition system has been developed based on continuous head pose estimation in the natural classroom. A novel cascaded multi-task Hough forest (CM-HF) combined with weighted Hough voting and multi-task learning is proposed for continuous head pose estimation, tip of the nose location and SS recognition, which improves accuracies of recognition and reduces the training time. Then, M-VFOA can be recognized based on estimated head poses, environmental cues and prior states in the natural classroom. Meanwhile, SS is classified using CM-HF with local cascaded mouth-eyes areas normalized by the estimated head poses. The method is rigorously evaluated for continuous head pose estimation, multi-person VFOA recognition, and SS recognition on some public available datasets and real-class video sequences. Experimental results show that our method reduces training time greatly and outperforms the state-of-the-art methods for both performance and robustness with an average accuracy of 83.5% on head pose estimation, 67.8% on M-VFOA recognition and 97.1% on SS recognition in challenging environments. Yuanyuan Liu 0004, Xingmei Li, Fang Fang 0008, Fayong Zhang, Jingying Chen 0001, Zhizhong Zeng |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2018 | Urban Land-Use Classification From PhotographsabstractLand-use (LU) classification of urban areas is conventionally achieved via field survey or remote sensing technologies, which is labor-intensive and time-consuming. With the wide development of social networks such as microblog and ubiquitous network access, images are captured by residents and tourists. In this letter, we propose a method for an automatic urban LU classification using geotagged images from public venues. Our method identifies the LU type depicted in those images that are extrapolated to the local regions bounded by street blocks. Experiments were conducted with geotagged photographs and Open Street Map of an urban area in London, U.K. It was demonstrated that the proposed method achieved overall 76.5% accuracy across five LU types. More importantly, our method demonstrated a greater performance in dealing with a mixture of LU types. Fang Fang 0008, Xiaohui Yuan 0001, Yuanyuan Liu 0004, Zhongwen Luo |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Conditional convolution neural network enhanced random forest for facial expression recognition
Yuanyuan Liu 0004, Xiaohui Yuan 0001, Xi Gong, Zhong Xie, Fang Fang 0008, Zhongwen Luo |
Pattern Recognit. | 5 |
| 2017 | Urban function zoning using geotagged photos and openstreetmapabstractUrban function zoning is of great importance for urban structure optimization, urban resource allocation, and urban development planning. Since citizens usually act as a network of motion sensors of the city, their activities could reflect the environment around them. We considered taking advantage of VGI data to classify urban function zones. In this paper, we proposed a framework for automated urban function zoning which is based on VGI geo-tagged photos and OpenStreetMap (OSM) data. Through combining the high-level image features of geo-tagged photos with the road network data, we obtained the functional zoning map of the study area. The experiment result shows the effectiveness of the framework we proposed. Fang Fang 0008, Xiaohui Yuan 0001, Zhongwen Luo, Yuanyuan Liu 0004, Bo Wan 0006, Yishi Zhao |
IGARSS | 2 |
| 2017 | A Performance Evaluation Model for Taxi Cruising Path Recommendation System
Huimin Lv, Fang Fang 0008, Yishi Zhao, Yuanyuan Liu 0004, Zhongwen Luo |
PAKDD (2) | 2 |