Tian Feng 0001

dblp:65/2637-1 · DBLP profile ↗
← Back
34ranked-venue papers
6as first author
30since 2021 · last 2026
0000-0001-9691-3266ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SeCuRe: Toward sparse-view 3D curve reconstruction with Gaussian Splatting
Tian Feng 0001, Haojie Dong, Jinkang Ji, Junao Shen, Feiyi Fan
Comput. Graph.1
2026 CAStGS: Context-aware scene stylization with 3D Gaussian splatting
Junao Shen, Ruihong Ye, Tian Feng 0001
Comput. Graph.4
2025 In2NeCT: Inter-class and Intra-class Neural Collapse Tuning for Semantic Segmentation of Imbalanced Remote Sensing Images
abstract
Remote sensing images (RSIs) are frequently characterized by multi-scale inter-class objects and inconsistently distributed objects due to scene limitations, which would cause a significant data imbalance challenging the corresponding semantic segmentation. Recent methods have leveraged various deep learning techniques to capture high-quality representations for RSI semantic segmentation, but are hardly capable of addressing the afore-mentioned challenge given their limited explorations towards the mechanisms behind the representations. The recently discovered Neural Collapse (NC) phenomenon in computer vision models suggests the simplex equiangular tight frame (ETF) as the optimal representation structure, which has motivated us to observe that the optimal structure of last-layer representations is disrupted and inter-class representations for minor classes tend to become closer to each other beacuse of data imbalance. To address these issues, we propose Inter-class and Intra-class Neural Collapse Tuning (In2NeCT) to optimize the representations that satisfy the simplex ETF, which facilitates the discrimination of inter-class representations and the coherence of intra-class representations. Extensive experiments on three datasets demonstrate that our In2NeCT consistently leads to significant improvements in performance and outperforms the state-of-the-art methods.
Junao Shen, Qiyun Hu, Tian Feng 0001, Xinyu Wang 0036, Hui Cui 0002, Sensen Wu, Wei Zhang 0243
AAAI3
2025 DRoLaS: Diffusion-Based Coarse-to-Fine Conditional Synthesis of Hierarchical Road Layouts
Shenao Dong, Bo Li 0173, Junao Shen, Tian Feng 0001
ICMR6
2025 WETR: Wireframe parsing using deformable Transformers
Jingwen Cui, Jinkang Ji, Junao Shen, Tianai Shen, Bo Li 0173, Tian Feng 0001
Comput. Graph.7
2025 CaRoLS: Condition-adaptive multi-level road layout synthesis
Tian Feng 0001, Bo Li 0173, Junao Shen
Comput. Graph.1
2025 SuraGS: Toward efficient few-shot novel view synthesis via surface-aware Gaussian splatting
Junao Shen, Tian Feng 0001, Haojie Dong, Jinkang Ji, Xinyu Wang 0036, Tianjia Shao
Comput. Graph.2
2024 CGMGM: A Cross-Gaussian Mixture Generative Model for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment unseen objects in a query image using a few pixel-wise annotated support images, thus expanding the capabilities of semantic segmentation. The main challenge lies in extracting sufficient information from the limited support images to guide the segmentation process. Conventional methods typically address this problem by generating single or multiple prototypes from the support images and calculating their cosine similarity to the query image. However, these methods often fail to capture meaningful information for modeling the de facto joint distribution of pixel and category. Consequently, they result in incomplete segmentation of foreground objects and mis-segmentation of the complex background. To overcome this issue, we propose the Cross Gaussian Mixture Generative Model (CGMGM), a novel Gaussian Mixture Models~(GMMs)-based FSS method, which establishes the joint distribution of pixel and category in both the support and query images. Specifically, our method initially matches the feature representations of the query image with those of the support images to generate and refine an initial segmentation mask. It then employs GMMs to accurately model the joint distribution of foreground and background using the support masks and the initial segmentation mask. Subsequently, a parametric decoder utilizes the posterior probability of pixels in the query image, by applying the Bayesian theorem, to the joint distribution, to generate the final segmentation mask. Experimental results on PASCAL-5i and COCO-20i datasets demonstrate our CGMGM's effectiveness and superior performance compared to the state-of-the-art methods.
Junao Shen, Kun Kuang 0001, Xinyu Wang 0036, Tian Feng 0001, Wei Zhang 0243
AAAI5
2024 SESAME: Toward Medical Image Segmentation via Foundation Model-assisted Semi-supervised Learning
abstract
Medical image segmentation is essential for diagnosis but requires expensive and time-consuming labeled data. Semi-supervised learning (SSL) mitigates this issue by using unlabeled data to improve generalization. However, current SSL methods encounter issues with inadaptive perturbations and low-quality pseudo-labels. Vision foundation models, such as SAM, have shown promise in segmentation. We propose SESAME, an SSL method integrating SAM and U-Net to improve labeling accuracy through a foundation model-assisted pipeline. In particular, we introduce a reliability score to address low-quality pseudo-labels and employ strategies for utilizing both reliable and unreliable images. Reliable images are associated with refined pseudo-labels via a conflict resolving strategy, whereas unreliable ones undergo a mutual region swapping strategy. Extensive experiments demonstrate that our SESAME outperforms representative methods for medical image segmentation.
Qiyun Hu, Junao Shen, Jinkang Ji, Xinyu Wang 0036, Tian Feng 0001, Hui Cui 0002
BIBM5
2024 Cryptocurrency Topic Burst Prediction via Hybrid Twin-structured Multi-modal Learning
Keting Yin, Xiaen Sun, Tian Feng 0001
DASFAA (2)3
2024 ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models
abstract
Yuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu, Cheng Yang, Chufan Shi, Xinyu Zhu, Zihao Lin, Hanwen Wan, Yujiu Yang, Tetsuya Sakai, Tian Feng, Hayato Yamana. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Junjie Wang 0011, Cheng Yang 0007, Chufan Shi, Hanwen Wan, Yujiu Yang 0001, Tetsuya Sakai, Tian Feng 0001, Hayato Yamana
EMNLP12
2024 CROCFUN: Cross-Modal Conditional Fusion Network for Pansharpening
abstract
Pansharpening aims to reconstruct a high-fidelity multispectral (HR-MS) image by fusing a multispectral (MS) image and a panchromatic (PAN) image. However, conventional pansharpening methods often struggle to address the modal gap between PAN and MS images. In this paper, we propose a novel cross-modal conditional fusion network (CroCFuN) for pansharpening, which builds upon recent advancements in optimal transport (OT) theory. Specifically, we formulate the modal alignment in pansharpening as an OT problem and thus design a feature alignment module (FAM) to adjust the MS image features to align with the PAN image features. Meanwhile, we propose a feature-activation normalization fusion module (FANFM), which adopts the multi-stage conditional fusion strategy to generate high-quality fusion features. Experimental results demonstrate that our CroCFuN outperforms recent representative methods for pan-sharpening regarding both visual and quantitative qualities. Code will be available at https://github.com/Florina2333.
Mengting Ma, Chenlu Hu, Huanting Zhang, Tian Feng 0001, Wei Zhang 0243
ICASSP5
2024 Retinal Vessel Segmentation via Cross-attention Feature Fusion
abstract
Retinal vessel segmentation from fundus images is of significant importance for detecting and diagnosing common ocular diseases. Conventional deep learning-based methods for retinal vessel segmentation follow the U-Net framework with an encoder-decoder architecture and employ skip connections for the recovery of spatial information lost during downsampling. However, skip connections cannot consistently have positive contributions to segmentation performance, which is caused by the semantic incompatibility between encoder features and decoder features. Based on this observation, we propose CaFFNet, a Cross-attention Feature Fusion Network designed specifically for retinal vessel segmentation. Specifically, we improve skip connections by introducing a Cross-attention Feature Fusion (CaFF) module, which effectively mitigates the semantic gap between encoder and decoder feature maps by leveraging the cross-attention mechanism for feature fusion. Besides, we introduce a Dual-Branch Pooling Fusion (DBPF) module to address the loss of vessel spatial information during pooling and capture contextual details more effectively, so as to improve segmentation performance. Experimental results on three fundus image datasets demonstrate that our CaFFNet outperforms current representative methods for retinal vessel segmentation.
Tian Feng 0001, Junao Shen, Qiangguo Jin, Xinyu Wang 0036
ICME1
2024 WirePAuS: Auxiliary-free Single-shot Wireframe Parsing
abstract
Wireframe parsing aims to identify vectorized line segments as pairs of endpoints from an image. Conventional methods usually require field-specific knowledge for manual introduction of auxiliary processes or auxiliary learning tasks towards satisfactory performances. Such pipelines are, however, characterized by high complexity, insignificant efficiency, and limited space for further performance improvement. To address these issues, we propose WirePAuS, a novel Wireframe Parser with an Auxiliary-free Single-shot pipeline, which requires no auxiliary processes or auxiliary learning tasks. This is based on its capability to generate appropriate prior information from a prior-informed feature extractor, which incorporates frequency-domain and Hough-domain prior information on line segments in the backbone. Meanwhile, we devise a structurally compact pipeline that enables the parser to directly predict the focal midpoints as line objects and exploit their displacements for the corresponding endpoints. Our end-to-end trainable WirePAuS is capable to capture rich structural details during single-shot inference. Extensive experiments suggest that the proposed method reaches significantly improved performances for wireframe parsing and outperforms a series of state-of-the-art methods.
Jinkang Ji, Junao Shen, Xinyu Wang 0036, Tian Feng 0001, Sensen Wu
ICME4
2024 MuMoSNet: 3D MRI-based Brain Tumor Segmentation via Multi-modal and Multi-scale Feature Fusion
abstract
MRI images contain multi-modal information, introducing complexity to brain tumor segmentation. Recent studies have incorporated the Transformer model, given its exceptional capability to model long-range dependence, into convolutional neural networks (CNNs) to address limited receptive fields. However, such a hybrid strategy often neglects the inherent multimodal characteristics of MRI images and lacks the capacity to capture modality-specific features. In this paper, we propose a multi-modal and multi-scale feature fusion network (MuMoSNet) for brain tumor segmentation from 3D MRI images. Specifically, our MuMoSNet introduces a parallel ME-Transformer encoder alongside the CNN-based encoder in 3D U-Net to separately extract modality-specific features. Besides, we devise a multi-feature fusion (MuFF) module to learn affinity relationships between cross-modality shared features and modality-specific features, maximizing the exploration of multi-modal information. Extensive experiments on both BraTS21 and BraTS20 datasets suggest that our MuMoSNet outperforms current representative methods for brain tumor segmentation.
Hui Cui 0002, Junao Shen, Xinyu Wang 0036, Tian Feng 0001
ICME7
2024 A Multi-information Dual-Layer Cross-Attention Model for Esophageal Fistula Prognosis
Jianqiao Zhang 0003, Hao Xiong 0001, Qiangguo Jin, Tian Feng 0001, Jiquan Ma, Ping Xuan, Peng Cheng 0002, Zhiyu Ning, ChangYang Li, Hui Cui 0002
MICCAI (5)4
2024 DOCNet: Dual-Domain Optimized Class-Aware Network for Remote Sensing Image Segmentation
abstract
The spatial attention mechanism has been frequently employed for the semantic segmentation of remote sensing images, given its renowned capability to model long-range dependencies. As remote sensing images often exhibit intricate backgrounds, significant intraclass variability, and a foreground-background imbalance, spatial attention mechanism-based methods somehow tend to introduce an extensive amount of background context through intensive affinity operations, causing unsatisfactory segmentation outcomes. While several class-aware methods attempt to attenuate the interference of background context by generating class representations as representative features, they still encounter challenges related to independent correlation calculation and single-confidence scale class representations. We introduce a dual-domain optimized class-aware network designed to address these challenges. In the semantic domain, we use category confidence as a scaling criterion to derive class representations at multiple confidence scales, effectively modeling pixel-class relationships. In the spatial domain, we leverage pixel-class relationships and their consensus to enhance relevant correlations while suppressing erroneous ones. Experimental results on three datasets demonstrate that the proposed method surpasses previous state-of-the-art ones for remote sensing image segmentation. Code is available athttps://github.com/xwmaxwma/rssegmentation.
Rui Che, Xinyu Wang 0036, Mengting Ma, Sensen Wu, Tian Feng 0001, Wei Zhang 0243
IEEE Geosci. Remote. Sens. Lett.6
2024 Aggregative and Contrastive Dual-View Graph Attention Network for Hyperspectral Image Classification
abstract
Graph convolutional networks (GCNs) have recently gained prominence in hyperspectral images (HSIs) classification tasks given their superior performance on non-Euclidean data. However, GCN-based methods are heavily reliant on complete graph structural information, which can cause the aggregation and transmission of information across nodes from differing classes, thereby compromising the classification performance. Furthermore, the scarcity of labeled pixels in HSIs often limits the representational capability of such methods. To address these issues, we propose an aggregative and contrastive dual-view graph attention network (ACoD-GAT) for HSI classification. Specifically, we present a progressive aggregation module, including a pixel clustering submodule and a node aggregation submodule to exploit semantic information at various levels. Besides, we integrate multiscale manipulation with a diffusion matrix to construct the dual view to further extract semantic information from both local and global perspectives. Moreover, we design an unsupervised contrastive loss function and a supervised contrastive loss function to facilitate contrastive learning on the dual view, improving the representational capabilities of ACoD-GAT with very few labeled samples. The extensive experimental results on four benchmark datasets demonstrate the superiority of the proposed ACoD-GAT compared with other state-of-the-art methods.
Haoyu Jing, Sensen Wu, Laifu Zhang, Fanen Meng, Tian Feng 0001, Zhenhong Du
IEEE Trans. Geosci. Remote. Sens.5
2024 A Conditional Diffusion Model With Fast Sampling Strategy for Remote Sensing Image Super-Resolution
abstract
Conventional deep learning-based methods for single remote sensing image super-resolution (SRSISR) have made remarkable progress. However, the super-resolution (SR) outputs of these methods are yet to become sufficiently satisfactory in visual quality. Recent diffusion model-based generative deep learning models are capable to enhance the visual quality of output images, but this capability is limited due to their sampling efficiency. In this article, we propose FastDiffSR, an SRSISR method based on a conditional diffusion model. Specifically, we devise a novel sampling strategy to reduce the number of sampling steps required by the diffusion model while ensuring the sampling quality. Meanwhile, the residual image is adopted to reduce computational costs, demonstrating that integrating channel attention and spatial attention begets a further improvement in the visual quality of output images. Compared to the state-of-the-art (SOTA) convolutional neural network (CNN)-based, GAN-based, and Transformer-based SR methods, our FastDiffSR improves the learned perceptual image patch similarity (LPIPS) by 0.1–0.2 and achieves better visual results in some real-world scenes. Compared with existing diffusion-based SR methods, our FastDiffSR achieves significant improvements in pixel-level evaluation metric peak signal-noise ratio (PSNR) while having smaller model parameters and obtaining better SR results on Vaihingen data with faster inference time by 2.8–28 times, showing excellent generalization ability and time efficiency. Our code will be open source athttps://github.com/Meng-333/FastDiffSR.
Fanen Meng, Haoyu Jing, Laifu Zhang, Yingchao Ren, Sensen Wu, Tian Feng 0001, Renyi Liu, Zhenhong Du
IEEE Trans. Geosci. Remote. Sens.8
2024 Single Remote Sensing Image Super-Resolution via a Generative Adversarial Network With Stratified Dense Sampling and Chain Training
abstract
Super-resolution (SR) methods have significantly contributed to the improvement of the spatial resolution of remote sensing (RS) images. The development of deep learning empowers novel methods to learn informative feature representation from massive low-resolution (LR) and high-resolution (HR) image pairs. Conventional RS image SR methods, however, may fail in large-scale ($\times 8$and$\times 9$) SR tasks. Specifically, a larger scale factor corresponds to less information in LR images, which is a considerable challenge to SR. To address the issue, we propose a novel method for single RS image SR (SRSISR) based on stratified dense sampling to effectively extract image features. Specifically, the proposed SR dense-sampling residual attention network (SRDSRAN) combines dense sampling and residual learning to improve multilevel feature fusion and gradient propagation and employs local and global attentions to learn important features and long-range interdependence in the channel and spatial dimensions. Meanwhile, we also devise a discriminator model using local and global attentions and with the loss function integrating${L}_{1}$pixel loss,${L}_{1}$perceptual loss, and relativistic adversarial loss to obtain the perceptually realistic images. Besides, we introduce a chain training to promote performance and expedite the training process for large-scale SR. Experimental results on UC Merced image and other multispectral data demonstrated that our SRDSRAN outperformed the current state-of-the-art methods quantitatively and in visual quality and obtained a higher classification accuracy in scene classification, proving its potential for applications with other downstream tasks. The code of SRADSGAN will be available athttps://github.com/Meng-333/SRADSGAN.
Fanen Meng, Sensen Wu, Zhe Zhang 0040, Tian Feng 0001, Renyi Liu, Zhenhong Du
IEEE Trans. Geosci. Remote. Sens.5
2024 Causality-Guided Stepwise Intervention and Reweighting for Remote Sensing Image Semantic Segmentation
abstract
Semantic segmentation is one of the most significant tasks in remote sensing (RS) image interpretation, which focuses on learning global and local information to infer the semantic label of each pixel. Previous studies devise encoder-decoder structured deep learning (DL) models to extract global and local features from RS images with the help of pretraining knowledge to predict semantic labels. However, due to the common heterogeneity between the data for pretraining and the data to be semantically segmented, these models fail to learn general features appropriate to RS datasets. In this article, we propose a novel formulation of the above problem from a causal perspective, where the learned features from pretrained models result from causality and spurious correlations, and only the former carries general information that remains invariant regardless of the exact task and dataset. Based on the above formulation, we propose stepwise intervention and reweighting (SIR). It can reduce the confounding bias introduced by the pretraining knowledge and improve the model’s ability to learn general features, making semantic segmentation of RS images benefit more from pretraining. Besides, we conduct a detailed theoretical analysis of our methods and conduct extensive experiments on two widely used public RS datasets. Experimental results demonstrate that applying SIR to encoder-decoder semantic segmentation models achieves performance improvements, proving the effectiveness and application values of the proposed method.
Baohong Li, Laifu Zhang, Kun Kuang 0001, Sensen Wu, Tian Feng 0001, Zhenhong Du
IEEE Trans. Geosci. Remote. Sens.6
2023 Log-Can: Local-Global Class-Aware Network For Semantic Segmentation of Remote Sensing Images
abstract
Remote sensing images are known of having complex backgrounds, high intra-class variance and large variation of scales, which bring challenge to semantic segmentation. We present LoG-CAN, a multi-scale semantic segmentation network with a global class-aware (GCA) module and local class-aware (LCA) modules to remote sensing images. Specifically, the GCA module captures the global representations of class-wise context modeling to circumvent back-ground interference; the LCA modules generate local class representations as intermediate aware elements, indirectly associating pixels with global class representations to reduce variance within a class; and a multi-scale architecture with GCA and LCA modules yields effective segmentation of objects at different scales via cascaded refinement and fusion of features. Through the evaluation on the ISPRS Vaihingen dataset and the ISPRS Potsdam dataset, experimental results indicate that LoG-CAN outperforms the state-of-the-art methods for general semantic segmentation, while significantly reducing network parameters and computation. Code is available at https://github.com/xwmaxwma/rssegmentation.
Mengting Ma, Chenlu Hu, Zhiyuan Song, Tian Feng 0001, Wei Zhang 0243
ICASSP6
2023 SACANet: scene-aware class attention network for semantic segmentation of remote sensing images
abstract
Spatial attention mechanism has been widely used in semantic segmentation of remote sensing images given its capability to model long-range dependencies. Many methods adopting spatial attention mechanism aggregate contextual information using direct relationships between pixels within an image, while ignoring the scene awareness of pixels (i.e., being aware of the global context of the scene where the pixels are located and perceiving their relative positions). Given the observation that scene awareness benefits context modeling with spatial correlations of ground objects, we design a scene-aware attention module based on a refined spatial attention mechanism embedding scene awareness. Besides, we present a local-global class attention mechanism to address the problem that general attention mechanism introduces excessive background noises while hardly considering the large intra-class variance in remote sensing images. In this paper, we integrate both scene-aware and class attentions to propose a scene-aware class attention network (SACANet) for semantic segmentation of remote sensing images. Experimental results on three datasets show that SACANet outperforms other state-of-the-art methods and validate its effectiveness. Code is available at https://github.com/xwmaxwma/rssegmentation.
Rui Che, Tingfeng Hong, Mengting Ma, Tian Feng 0001, Wei Zhang 0243
ICME6
2023 STNet: Spatial and Temporal feature fusion network for change detection in remote sensing images
abstract
As an important task in remote sensing image analysis, remote sensing change detection (RSCD) aims to identify changes of interest in a region from spatially co-registered multi-temporal remote sensing images, so as to monitor the local development. Existing RSCD methods usually formulate RSCD as a binary classification task, representing changes of interest by merely feature concatenation or feature subtraction and recovering the spatial details via densely connected change representations, whose performances need further improvement. In this paper, we propose STNet, a RSCD network based on spatial and temporal feature fusions. Specifically, we design a temporal feature fusion (TFF) module to combine bitemporal features using a cross-temporal gating mechanism for emphasizing changes of interest; a spatial feature fusion module is deployed to capture fine-grained information using a cross-scale attention mechanism for recovering the spatial details of change representations. Experimental results on three benchmark datasets for RSCD demonstrate that the proposed method achieves the state-of-the-art performance. Code is available at https://github.com/xwmaxwma/rschange.
Tingfeng Hong, Mengting Ma, Tian Feng 0001, Wei Zhang 0243
ICME6
2023 DBDAN: Dual-Branch Dynamic Attention Network for Semantic Segmentation of Remote Sensing Images
Rui Che, Tingfeng Hong, Xinyu Wang 0036, Tian Feng 0001, Wei Zhang 0243
PRCV (4)5
2023 MAPMaN: Multi-Stage U-Shaped Adaptive Pattern Matching Network for Semantic Segmentation of Remote Sensing Images
abstract
Abstract Remote sensing images (RSIs) often possess obvious background noises, exhibit a multi‐scale phenomenon, and are characterized by complex scenes with ground objects in diversely spatial distribution pattern, bringing challenges to the corresponding semantic segmentation. CNN‐based methods can hardly address the diverse spatial distributions of ground objects, especially their compositional relationships, while Vision Transformers (ViTs) introduce background noises and have a quadratic time complexity due to dense global matrix multiplications. In this paper, we introduce Adaptive Pattern Matching (APM), a lightweight method for long‐range adaptive weight aggregation. Our APM obtains a set of pixels belonging to the same spatial distribution pattern of each pixel, and calculates the adaptive weights according to their compositional relationships. In addition, we design a tiny U‐shaped network using the APM as a module to address the large variance of scales of ground objects in RSIs. This network is embedded after each stage in a backbone network to establish a Multi‐stage U‐shaped Adaptive Pattern Matching Network (MAPMaN), for nested multi‐scale modeling of ground objects towards semantic segmentation of RSIs. Experiments on three datasets demonstrate that our MAPMaN can outperform the state‐of‐the‐art methods in common metrics. The code can be available at https://github.com/INiid/MAPMaN .
Tingfeng Hong, Xinyu Wang 0036, Rui Che, Chenlu Hu, Tian Feng 0001, Wei Zhang 0243
Comput. Graph. Forum6
2022 EDVAM: a 3D eye-tracking dataset for visual attention modeling in a virtual museum
abstract
Predicting visual attention facilitates an adaptive virtual museum environment and provides a context-aware and interactive user experience. Explorations toward development of a visual attention mechanism using eye-tracking data have so far been limited to 2D cases, and researchers are yet to approach this topic in a 3D virtual environment and from a spatiotemporal perspective. We present the first 3D Eye-tracking Dataset for Visual Attention modeling in a virtual Museum, known as the EDVAM. In addition, a deep learning model is devised and tested with the EDVAM to predict a user’s subsequent visual attention from previous eye movements. This work provides a reference for visual attention modeling and context-aware interaction in the context of virtual museums.
Yunzhan Zhou, Tian Feng 0001, Shihui Shuai, Lingyun Sun, Henry Been-Lirn Duh
Frontiers Inf. Technol. Electron. Eng.2
2022 Mood-Driven Colorization of Virtual Indoor Scenes
abstract
One of the challenging tasks in virtual scene design for Virtual Reality (VR) is causing it to invoke a particular mood in viewers. The subjective nature of moods brings uncertainty to the purpose. We propose a novel approach to automatic adjustment of the colors of textures for objects in a virtual indoor scene, enabling it to match a target mood. A dataset of 25,000 images, including building/home interiors, was used to train a classifier with the features extracted via deep learning. It contributes to an optimization process that colorizes virtual scenes automatically according to the target mood. Our approach was tested on four different indoor scenes, and we conducted a user study demonstrating its efficacy through statistical analysis with the focus on the impact of the scenes experienced with a VR headset.
Michael Solah, Haikun Huang, Jiachuan Sheng, Tian Feng 0001, Marc Pomplun, Lap-Fai Yu
IEEE Trans. Vis. Comput. Graph.4
2021 Predicting Esophageal Fistula Risks Using a Multimodal Self-attention Network
Yulu Guan, Hui Cui 0002, Yiyue Xu, Qiangguo Jin, Tian Feng 0001, Huawei Tu, Ping Xuan, Wanlong Li, Henry Been-Lirn Duh
MICCAI (5)5
2021 A review of computer graphics approaches to urban modeling from a machine learning perspective
abstract
Urban modeling facilitates the generation of virtual environments for various scenarios about cities. It requires expertise and consideration, and therefore consumes massive time and computation resources. Nevertheless, related tasks sometimes result in dissatisfaction or even failure. These challenges have received significant attention from researchers in the area of computer graphics. Meanwhile, the burgeoning development of artificial intelligence motivates people to exploit machine learning, and hence improves the conventional solutions. In this paper, we present a review of approaches to urban modeling in computer graphics using machine learning in the literature published between 2010 and 2019. This serves as an overview of the current state of research on urban modeling from a machine learning perspective.
Tian Feng 0001, Feiyi Fan, Tomasz Bednarz
Frontiers Inf. Technol. Electron. Eng.1
2018 Urban Zoning Using Higher-Order Markov Random Fields on Multi-View Imagery Data
Tian Feng 0001, Quang-Trung Truong, Duc Thanh Nguyen, Jing Yu Koh, Lap-Fai Yu, Alexander Binder, Sai-Kit Yeung
ECCV (8)1
2017 A Novel Riemannian Metric Based on Riemannian Structure and Scaling Information for Fixed Low-Rank Matrix Completion
abstract
Riemannian optimization has been widely used to deal with the fixed low-rank matrix completion problem, and Riemannian metric is a crucial factor of obtaining the search direction in Riemannian optimization. This paper proposes a new Riemannian metric via simultaneously considering the Riemannian geometry structure and the scaling information, which is smoothly varying and invariant along the equivalence class. The proposed metric can make a tradeoff between the Riemannian geometry structure and the scaling information effectively. Essentially, it can be viewed as a generalization of some existing metrics. Based on the proposed Riemanian metric, we also design a Riemannian nonlinear conjugate gradient algorithm, which can efficiently solve the fixed low-rank matrix completion problem. By experimenting on the fixed low-rank matrix completion, collaborative filtering, and image and video recovery, it illustrates that the proposed method is superior to the state-of-the-art methods on the convergence efficiency and the numerical performance.
Shasha Mao, Licheng Jiao, Tian Feng 0001, Sai-Kit Yeung
IEEE Trans. Cybern.4
2016 Crowd-driven mid-scale layout design
abstract
We propose a novel approach for designing mid-scale layouts by optimizing with respect to human crowd properties. Given an input layout domain such as the boundary of a shopping mall, our approach synthesizes the paths and sites by optimizing three metrics that measure crowd flow properties: mobility, accessibility, and coziness. While these metrics are straightforward to evaluate by a full agent-based crowd simulation, optimizing a layout usually requires hundreds of evaluations, which would require a long time to compute even using the latest crowd simulation techniques. To overcome this challenge, we propose a novel data-driven approach where nonlinear regressors are trained to capture the relationship between the agent-based metrics, and the geometrical and topological features of a layout. We demonstrate that by using the trained regressors, our approach can synthesize crowd-aware layouts and improve existing layouts with better crowd flow properties.
Tian Feng 0001, Lap-Fai Yu, Sai-Kit Yeung, KangKang Yin, Kun Zhou 0001
ACM Trans. Graph.1
2014 Modelling Mutual Information between Voiceprint and Optimal Number of Mel-Frequency Cepstral Coefficients in Voice Discrimination
abstract
In this paper, we study the relationship between the voiceprint and the optimal number of Mel-frequency Cepstral Coefficients (MFCCs). The voiceprint is modelled as sub-MFCCs matrix with the first d number of MFCCs. We model the relationship through information theory and formulate it as the mutual information maximization problem subject to the probabilities constraint. The solution of this optimization problem provides the optimal number of MFCCs, D among these d, which yields the highest classification accuracy of the voice discrimination, together with a confidence level. This study is dictated by the need to understand the use of MFCCs, which have proliferated since its invention to discriminate voice. We evaluate our model by comparing the leave-one-out cross validation (LOOCV) results of usual multi-class classifier, the Supervised Learning Gaussian Mixture Model (SLGMM), with a set of spoken words and A capella solo vocal performances. The experimental results show that our model is a more comprehensive feature selection criteria for the MFCCs than the de-facto technique, LOOCV.
Kin Wah Edward Lin, Tian Feng 0001, Natalie Agus, Clifford So, Simon Lui
ICMLA2