Min Tan 0005

dblp:08/3957-5 · DBLP profile ↗
← Back
34ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0002-1842-4050ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2026 SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images
abstract
Feed-forward 3D reconstruction from sparse, low-resolution (LR) images is a crucial capability for real-world applications, such as autonomous driving and embodied AI. However, existing methods often fail to recover fine texture details. This limitation stems from the inherent lack of high-frequency information in LR inputs. To address this, we propose SRSplat, a feed-forward framework that reconstructs high-resolution 3D scenes from only a few LR views. Our main insight is to compensate for the deficiency of texture information by jointly leveraging external high-quality reference images and internal texture cues. We first construct a scene-specific reference gallery, generated for each scene using Multimodal Large Language Models (MLLMs) and diffusion models. To integrate this external information, we introduce the Reference-Guided Feature Enhancement (RGFE) module, which aligns and fuses features from the LR input images and their reference twin image. Subsequently, we train a decoder to predict the Gaussian primitives using the multi-view fused feature obtained from RGFE. To further refine predicted Gaussian primitives, we introduce Texture-Aware Density Control (TADC), which adaptively adjusts Gaussian density based on the internal texture richness of the LR inputs. Extensive experiments demonstrate that our SRSplat outperforms existing methods on various datasets, including RealEstate10K, ACID, and DTU, and exhibits strong cross-dataset and cross-resolution generalization capabilities.
Changyue Shi, Chuxiao Yang, Jiajun Ding, Zhou Yu 0001, Min Tan 0005
AAAI9
2026 Emotional conflict adaptation for multimodal sentiment analysis
Tingting Han 0003, Lingyun Yu 0004, Min Tan 0005, Zhou Yu 0001, Hongxun Yao
Pattern Recognit.3
2026 HEART: Emotionally Grounded Video Captioning via Hierarchical Emotion-Aligned Representation
abstract
Emotional Video Captioning (EVC) seeks to generate video descriptions that are both factually accurate and emotionally expressive. However, existing approaches often lack structured semantic grounding and fine-grained temporal modeling, leading to incomplete or emotionally inconsistent captions. To address these issues, we proposeHEART(HierarchicalEmotion-AlignedRepresentation withTemporal structure), a unified framework that jointly models hierarchical visual semantics and multi-scale temporal context. Specifically, HEART introduces a Hierarchical Semantic Extraction Module that decomposes visual content into entity-, action-, and event-level representations, providing a rich foundation for multi-level emotional alignment. A Temporal Pyramid Module captures short- and long-range temporal dependencies through multi-scale convolution, enabling temporally coherent captioning. Together, these components enable HEART to generate captions that are both emotionally grounded and temporally complete. To support this framework, we construct EmoStruct, a new benchmark dataset with fine-grained emotional annotations at the subject and predicate levels. Experiments on EmoStruct and public datasets demonstrate that HEART significantly outperforms prior methods in both semantic and emotional dimensions.
Tingting Han 0003, Yuxuan Gong, Sicheng Zhao, Min Tan 0005, Zhou Yu 0001, Hongxun Yao
IEEE Trans. Affect. Comput.4
2026 Fuzzy Language Gaussian Splatting
abstract
Recent advancements in open-vocabulary 3D querying have achieved remarkable progress. However, existing approaches such as LERF and LangSplat still rely heavily on specific query text for accurate 3D target identification. They fail to accurately comprehend fuzzy query text (which describes a target's properties, functions, or traits rather than naming it directly), causing localization mistakes and poor 3D segmentation results. In this work, we introduce a novel task—open-vocabulary 3D fuzzy query—which aims to locate and generate precise 3D target masks based on fuzzy query text, a capability that is crucial for building a 3D intelligent system. To address this challenge, we proposeFuzzy Language Gaussian Splatting (FL-GS), a framework consisting of three key stages: We first leverage a multimodal large language model (MLLM) to identify and localize potential targets in representative views based on fuzzy query text, thereby generating initial masks using the segment anything model (SAM). Subsequently, we compute the pairwise similarity between all target masks and construct an undirected, unweighted graph, from which the maximum clique is identified, thus obtaining the most probable correct masks, referred to as refined masks. Finally, we propose single-dimensional mask encoding to efficiently achieve precise 3D target masks through supervision with refined masks. Furthermore, we manually annotated two key datasets and established a benchmark for this new task. Experimental results clearly demonstrate that FL-GS outperforms existing methods in open-vocabulary 3D fuzzy querying.
Jiajun Ding, Yaowei Liu, Hongxi Zhu, Min Tan 0005, Zhou Yu 0001
IEEE Trans. Fuzzy Syst.6
2025 DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency Regularization
abstract
The rapid evolution of the online fashion industry has intensified the demand for interactive fashion retrieval systems capable of precise and flexible searches based on user-specified attribute modifications. However, prevailing fashion retrieval methods often overlook the distinctive distributional properties of fashion images and struggle to preserve semantic consistency during attribute manipulation. To address these limitations, we propose DiSCo, a novel disentangled attribute manipulation retrieval framework via semantic reconstruction and consistency regularization. Our approach comprises three key components: (1) An attribute-aware manipulation network that constructs target fashion embeddings through cross-modal attribute modification deltas, leveraging dedicated fashion attribute encoders; (2) A cross-modal semantic reconstruction network that synthesizes target images directly from modified attribute descriptions, supervised by adversarial and attribute classification losses to ensure interpretable edits; (3) An adaptive fusion mechanism that dynamically integrates attribute-modified embeddings with reconstructed image features. Extensive evaluations on two benchmark datasets (DeepFashion and Shopping100K) demonstrate that DiSCo achieves superior retrieval accuracy over state-of-the-arts while maintaining high-fidelity editing. Quantitative and qualitative analyses further confirm that DiSCo generates more realistic fashion representations, underscoring its effectiveness in attribute-aware retrieval tasks.
Min Tan 0005, Guanhao Liu, Huijing Zhan, Yuyu Yin, Zhou Yu 0001, Jiajun Ding, Yinfu Feng
ACM Multimedia1
2025 FutureGS: Structured Gaussian Fields for Future-Aware Dynamic Scene Modeling
Mingyang Ding, Tingting Han 0003, Jiajun Ding, Min Tan 0005, Zhenzhong Kuang
ACM Multimedia7
2025 FedGA: Federated Learning via Gradient Adaptive Aggregation
abstract
In modern lives, the rapid proliferation of Internet of Things (IoT) devices has made them indispensable tools for data collection and analysis across various domains. However, growing concerns over data ownership and privacy have hindered effective data sharing among IoT devices, leading to the persistent challenge of data silos. Federated Learning (FL) has emerged as a promising solution to this problem by enabling collaborative model training without direct data exchange. Despite its potential, FL faces two critical limitations: severe catastrophic forgetting for historical knowledge and inefficient average aggregation. To address these challenges, this paper proposes FedGA, an innovative FL framework that leverages cosine similarity-based weighted aggregation to enhance model convergence speed. Furthermore, FedGA incorporates a mechanism to memorize historical models, thereby significantly alleviating catastrophic forgetting. Extensive experiments on three public datasets validate the effectiveness of FedGA, demonstrating its superior performance in both accuracy and training efficiency compared to state-of-the-art methods. The results highlight FedGA’s capability to overcome the key shortcomings of existing FL approaches, making it a robust solution for practical IoT applications.
Changfeng Hu, Min Tan 0005, Tingting Han 0003, Zhenzhong Kuang
SMC2
2025 ViewCloud: A lightweight multi-view point cloud representation for efficient 3D recognition and cross-domain retrieval
Zhihe Wu, Yaomin Wang, Zhenzhong Kuang, Jiajun Ding, Min Tan 0005, Xuefei Yin, Yanming Zhu 0001
Comput. Aided Des.5
2025 Modality-aware contrast and fusion for multi-modal summarization
Lixin Dai, Tingting Han 0003, Zhou Yu 0001, Jun Yu 0002, Min Tan 0005
Neurocomputing5
2025 MMGS: Multi-Model Synergistic Gaussian Splatting for Sparse View Synthesis
Changyue Shi, Chuxiao Yang, Jiajun Ding, Min Tan 0005
Image Vis. Comput.6
2025 SR4D: Dynamic scene super resolution from monocular videos
Chuxiao Yang, Changyue Shi, Suguo Zhu, Jiajun Ding, Yeru Wang, Min Tan 0005
Knowl. Based Syst.7
2025 DE-NAF: decoupled neural attenuation fields for sparse-view CBCT reconstruction
Tianning Zhao, Guoping Ding, Zhenyang Liu, Hangping Wei, Min Tan 0005, Jiajun Ding
Pattern Anal. Appl.6
2025 Benchmarking and Enhancing Geospatial Visual Reasoning Over Street Maps
abstract
Recent advances in large multimodal models (LMMs) have enabled substantial progress in various visual question answering (VQA) benchmarks, including the challenging text-centric ones that require a simultaneous understanding of both the visual and textual contents in the images. Despite the prominence of existing text-centric VQA benchmarks, they either have limited textual information or have a limited number of questions requiring complex reasoning skills beyond the basic OCR. To this end, we present SMVQA—a novel text-centric VQA benchmark based on street map images. SMVQA contains more than 10K real-world street map images from the open geospatial database OpenStreetMap. Each image in SMVQA is also associated with detailed geospatial annotations, enabling it to automatically generate up to 57.5K distinctive QA pairs of five representative question types. In addition to the standard test split, SMVQA introduces an extra test split to verify the generalization abilities over out-of-domain images and novel reasoning skills. The evaluation of the state-of-the-art open-source and commercial LMMs reflects the great challenge posed by SMVQA. The latest LMMs, such as GPT-4o, only achieve accuracies of 49.9%, showing plenty of room for improvement. To further improve the latest LMMs’ performance on SMVQA, we introduce a LMM-based agentic framework LHR, which consists of the localizing, highlighting, and reasoning stages. Specifically, LHR first prompts the LMM to localize region-of-interest (RoI) to the question and then highlight the RoI and perform chain-of-thought reasoning for answer prediction. By integrating LHR with GPT-4o, we observe a significant improvement over the vanilla counterpart, showing the effectiveness of our framework.
Wenwen Pan 0003, Haiting Zhou, Zhenwei Shao, Shuai Shao 0012, Suguo Zhu, Min Tan 0005, Jun Yu 0002, Zhou Yu 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 ScatDiff: Physical Diffusion Model for Electromagnetic Computational Imaging
abstract
Electromagnetic computational imaging offers a promising solution to electromagnetic inverse scattering problems. Whereas, it is challenged by its ill-posed nature and non-linearity. Traditional iterative methods are often slow and prone to local minima, while recent deep generative models overlook the physical principles that govern the transformation from scattering fields to images of constitute parameters, limiting their interpretability, generalization, and robustness. To address these issues, we propose ScatDiff, a novel Scatter-to-image Diffusion model that integrates electromagnetic data with fundamental physical principles. ScatDiff uses a time-aware, backpropagation-enhanced diffusion to generate noisy images embedded with electromagnetic priors, along with a denoising module that uses cross-attention to adaptively integrate scattering fields. Additionally, a physics-driven reconstruction module incorporates an induced current model into the loss function to enhance interpretability. Experiments on three MNIST variants, Gesture dataset, the “Austria” profile, and the “FoamDielExt” profile show that ScatDiff outperforms traditional iterative methods and deep learning models in both imaging quality and efficiency, with strong generalization and robustness under high noise conditions. Code and datasets are available on https://github.com/Scatdif.
Min Tan 0005, Kuiwen Xu, Zhou Yu 0001, Jun Yu 0002
IEEE Trans. Geosci. Remote. Sens.1
2024 Multi-Domain Deep Learning from a Multi-View Perspective for Cross-Border E-commerce Search
abstract
Building click-through rate (CTR) and conversion rate (CVR) prediction models for cross-border e-commerce search requires modeling the correlations among multi-domains. Existing multi-domain methods would suffer severely from poor scalability and low efficiency when number of domains increases. To this end, we propose a Domain-Aware Multi-view mOdel (DAMO), which is domain-number-invariant, to effectively leverage cross-domain relations from a multi-view perspective. Specifically, instead of working in the original feature space defined by different domains, DAMO maps everything to a new low-rank multi-view space. To achieve this, DAMO firstly extracts multi-domain features in an explicit feature-interactive manner. These features are parsed to a multi-view extractor to obtain view-invariant and view-specific features. Then a multi-view predictor inputs these two sets of features and outputs view-based predictions. To enforce view-awareness in the predictor, we further propose a lightweight view-attention estimator to dynamically learn the optimal view-specific weights w.r.t. a view-guided loss. Extensive experiments on public and industrial datasets show that compared with state-of-the-art models, our DAMO achieves better performance with lower storage and computational costs. In addition, deploying DAMO to a large-scale cross-border e-commence platform leads to 1.21%, 1.76%, and 1.66% improvements over the existing CGC-based model in the online AB-testing experiment in terms of CTR, CVR, and Gross Merchandises Value, respectively.
Yinfu Feng, Yunan Ye, Min Tan 0005, Rong Xiao 0005, Haihong Tang, Jiajun Ding, Jun Yu 0002
AAAI5
2024 DAS-COD: Depth-Aware Camouflaged Object Detection via Swin Transformer
abstract
The successful integration of depth information into salient object detection tasks has catalyzed research interest in depth-enhanced camouflaged object detection (COD) tasks. However, the challenges associated with acquiring depth information pose significant hurdles to this task, especially given the lack of RGB-D datasets tailored specifically for COD. Consequently, employing depth estimation techniques to generate pseudo-depth information emerges as a viable solution in the realm of depth-enhanced COD tasks. In this study, we propose an architecture, DAS-COD (Depth-Aware Swin Transformer COD), that integrates Swin Transformer model with depth estimation techniques for the purpose of camouflaged object detection. In particular, we use a dual-stream Swin Transformer backbone to extract feature maps from different modalities. These maps are then enhanced with a multi-modal feature enhancement module. Additionally, to address the inherent discrepancies between pseudo-depth maps and actual depth information, we incorporate an edge-aware module to significantly improve the accuracy of boundary delineation in the predicted outcomes. We tested our proposed method on three different COD datasets. Our results show that the model achieves state-of-the-art performance across these camouflaged object detection datasets.
Chenye Lu, Min Tan 0005, Xiaoyang Mao, Zilin Xia
SMC2
2024 Multimodal Federated Learning Via Local-Global Fusion
abstract
The number of Internet of Things (IoT) devices across diverse domains in modern life has made them significant sources for collecting and analyzing multi-modal data. However, concerns about ownership and data privacy associated with IoT devices make data sharing among multiple devices impractical. Recently, multimodal federated learning has emerged as an innovative solution where each device client can collectively train a satisfactory local model without exchanging local data. Nevertheless, most existing multimodal federated learning approaches prioritize training a powerful global server model while neglecting the performance of local client models. In this context, this paper introduces FedAF, a multimodal Federated learning approach via feature Fusion with adversarial representation learning, aimed at enhancing local representations and thereby improving local client models. Specifically, using the trained global model, FedAF integrates the global representation of each local client's data into its local feature obtained from the corresponding client model. Furthermore, domain adversarial learning is employed to align global and local representations by minimizing the discrepancy between local and global encoders, compelling the global encoder to adapt to local tasks. Comprehensive experiments on two unimodal classifications and one multimodal retrieval dataset demonstrate that FedAF achieves state-of-the-art performance compared to other federated learning methods and significantly improves local client models while maintaining the satisfactory performance of the global server model.
Zilin Xia, Min Tan 0005, Lingqiang Chu, Tingting Han 0003
SMC2
2024 GTADT: Gated tone-sensitive acne grading via augmented domain transfer
Min Tan 0005, Ruirui Wang, Ankur Purwar, Tao Jin 0004, Jun Yu 0002, Alex Chichung Kot
Multim. Tools Appl.1
2024 Semantic-aware hyper-space deformable neural radiance fields for facial avatar reconstruction
Kaixin Jin, Xiaoling Gu, Zhenzhong Kuang, Zizhao Wu, Min Tan 0005, Jun Yu 0002
Pattern Recognit. Lett.6
2024 FedSea: Federated Learning via Selective Feature Alignment for Non-IID Multimodal Data
abstract
The growing demands for privacy protection challenge the joint training of one model by leveraging multiple datasets. Federated learning (FL) provides a new way to overcome this challenge and has attracted many research interests, which enables multiple parties to collaboratively train a machine learning model without exchanging their local data. Despite some success, the non-independent and identically distributed (non-IID) data distributions in different parties remain challenging and easily damage the performance of FL methods, specifically for the heterogeneous multimodal data. Existing FL studies on non-IID data settings are often dedicated to the label space, neglecting the non-IID issues in feature space, thus limiting their performance when the parties with non-IID multimodal data. This paper proposes a newFederated learning method viaSelective featureAlignment (FedSea) to align representations across multiple parties in the feature space. FedSea uses a domain adversarial learning framework consisting of an affine-transform-based generator and a gradient-reversal-based client discriminator to perform IID transformation and reduce data source distinguishability, respectively. An attention-based mask module and a feature IID confidence quantification method are introduced to effectively address the diverse feature non-IID levels across multimodal data. Comprehensive experiments are conducted on three widely-used public datasets and one large-scale industrial dataset, showing FedSea has: 1) better performance than state-of-the-art FL methods on both multimodal and single-modal datasets; 2) superior feature alignment ability on non-IID datasets, and 3) good model interpretability.
Min Tan 0005, Yinfu Feng, Lingqiang Chu, Jingcheng Shi, Rong Xiao 0005, Haihong Tang, Jun Yu 0002
IEEE Trans. Multim.1
2023 EGRA-NeRF: Edge-Guided Ray Allocation for Neural Radiance Fields
Zhenbiao Gai, Zhenyang Liu, Min Tan 0005, Jiajun Ding, Jun Yu 0002, Mingzhao Tong, Junqing Yuan
Image Vis. Comput.3
2023 Electromagnetic Imaging Boosted Visual Object Recognition Under Difficult Visual Conditions
abstract
Object imaging and recognition under difficult visual conditions is extremely challenging due to the captured low-quality images, and traditional optical-based recognition methods always fail in this task. In this paper, we propose to utilize the visual-microwave image pairs captured by both visual cameras and microwave sensors for imaging and recognition. To address the heavy noises in the low-quality optical images, we retrieve the physically quantitative images from associated scattered field data, and enhance visual features by both optical and retrieval images. We develop a cross-modal Enhanced Attentive Visual-Microwave Fusion (EAVMF) object recognition model to jointly learn the cross-modal generator and multimodal recognizer. In addition, an attention module for the visual subnetwork is utilized to highlight the regions of interest. Two multimodal datasets with synthetic visual-microwave image pairs are built to simulate the difficult visual condition. The numerical results on these datasets demonstrate that: 1) both the multimodal fusion, cross-modal enhancement, and visual attention module can enhance the performance; and 2) compared with existing methods, the proposed EAVMF not only performs better in terms of accuracy but also has good scalability and one-shot learning ability.
Min Tan 0005, Tao Jin 0004, Danhui Ye, Kuiwen Xu, Xiaoling Gu, Jun Yu 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 Hierarchical Deep Click Feature Prediction for Fine-Grained Image Recognition
abstract
The click feature of an image, defined as the user click frequency vector of the image on a predefined word vocabulary, is known to effectively reduce the semantic gap for fine-grained image recognition. Unfortunately, user click frequency data are usually absent in practice. It remains challenging to predict the click feature from the visual feature, because the user click frequency vector of an image is always noisy and sparse. In this paper, we devise a Hierarchical Deep Word Embedding (HDWE) model by integrating sparse constraints and an improved RELU operator to address click feature prediction from visual features. HDWE is a coarse-to-fine click feature predictor that is learned with the help of an auxiliary image dataset containing click information. It can therefore discover the hierarchy of word semantics. We evaluate HDWE on three dog and one bird image datasets, in which Clickture-Dog and Clickture-Bird are utilized as auxiliary datasets to provide click data, respectively. Our empirical studies show that HDWE has 1) higher recognition accuracy, 2) a larger compression ratio, and 3) good one-shot learning ability and scalability to unseen categories.
Jun Yu 0002, Min Tan 0005, Hongyuan Zhang 0001, Yong Rui, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Fine-grained Image Classification via Multi-scale Selective Hierarchical Biquadratic Pooling
abstract
How to extract distinctive features greatly challenges the fine-grained image classification tasks. In previous models, bilinear pooling has been frequently adopted to address this problem. However, most bilinear pooling models neglect either intra or inter layer feature interaction. This insufficient interaction brings in the loss of discriminative information. In this article, we devise a novel fine-grained image classification approach named M ulti-scale S elective H ierarchical bi Q uadratic P ooling (MSHQP). The proposed biquadratic pooling simultaneously models intra and inter layer feature interactions and enhances part response by integrating multi-layer features. The subsequent coarse-to-fine multi-scale interaction structure captures the complementary information within features. Finally, the active interaction selection module adaptively learns the optimal interaction subset for a specific dataset. Consequently, we obtain a robust image representation with coarse-to-fine semantics. We conduct experiments on five benchmark datasets. The experimental results demonstrate that MSHQP achieves competitive or even match the state-of-the-art methods in terms of both accuracy and computational efficiency, with 89.0%, 94.9%, 93.4%, 90.4%, and 91.5% top-1 classification accuracy on CUB-200-2011, Stanford-Cars, FGVC-Aircraft, Stanford-Dog, and VegFru, respectively.
Min Tan 0005, Fu Yuan, Jun Yu 0002, Guijun Wang, Xiaoling Gu
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Fashion-Sketcher: A Model for Producing Fashion Sketches of Multiple Categories
Junkai Fang, Xiaoling Gu, Min Tan 0005
PRCV (2)3
2020 Fashion analysis and understanding with artificial intelligence
Xiaoling Gu, Fei Gao 0006, Min Tan 0005
Inf. Process. Manag.3
2020 Fine-grained image classification with factorized deep user click feature
Min Tan 0005, Zhiyou Peng, Jun Yu 0002, Fang Tang
Inf. Process. Manag.1
2019 Image Recognition by Predicted User Click Feature With Multidomain Multitask Transfer Deep Network
abstract
The click feature of an image, defined as a user click count vector based on click data, has been demonstrated to be effective for reducing the semantic gap for image recognition. Unfortunately, most of the traditional image recognition datasets do not contain click data. To address this problem, researchers have begun to develop a click prediction model using assistant datasets containing click information and have adapted this predictor to a common click-free dataset for different tasks. This method can be customized to our problem, but it has two main limitations: 1) the predicted click feature often performs badly in the recognition task since the prediction model is constructed independently of the subsequent recognition problem and 2) transferring the predictor from one dataset to another is challenging due to the large cross-domain diversity. In this paper, we devise a multitask and multidomain deep network with varied modals (MTMDD-VM) to formulate image recognition and click prediction tasks in a unified framework. Datasets with and without click information are integrated in the training. Furthermore, a nonlinear word embedding with a position-sensitive loss function is designed to discover the visual click correlation. We evaluate the proposed method on three public dog breed image datasets, and we utilize the Clickture-Dog dataset as the auxiliary dataset that provides click data. The experimental results show that: 1) the nonlinear word embedding and position-sensitive loss function largely enhance the predicted click feature in the recognition task, realizing a 32% improvement in accuracy; 2) the multitask learning framework improves accuracies in both image recognition and click prediction; and 3) the unified training using the combined dataset with and without click data further improves the performance. Compared with the state-of-the-art methods, the proposed approach not only performs much better in accuracy but also achieves good scalability and one-shot learning ability.
Min Tan 0005, Jun Yu 0002, Hongyuan Zhang 0001, Yong Rui, Dacheng Tao
IEEE Trans. Image Process.1
2018 Deep Point Convolutional Approach for 3D Model Retrieval
abstract
With the increasing popularity of 3D models, retrieving deformable 3D objects is becoming a crucial task. The state-of-the-art methods use complex deep neural networks to address this problem, which require lots of computational resources. In this paper, we develop a more effective solution by using point convolution. Our algorithm takes local point descriptors as the input and produces a global vector for shape retrieval. To save the efforts of designing complex deep convolutional neural network (CNN), we first use intrinsic point descriptors to describe the shape deformations. Then, a simple but effective point CNN network is developed to integrate the local shape information by performing subspace compression and fusion, which depends on an end-to-end learning process to link the local and global information for discriminative shape representation. The experimental results on popular benchmarks have verified that our algorithm is able to outperform the state-of-the-art methods.
Zhenzhong Kuang, Jun Yu 0002, Jianping Fan 0001, Min Tan 0005
ICME4
2018 Click data guided query modeling with click propagation and sparse coding
Min Tan 0005, Jun Yu 0002, Qingming Huang, Weichen Wu
Multim. Tools Appl.1
2018 User-Click-Data-Based Fine-Grained Image Recognition via Weakly Supervised Metric Learning
abstract
We present a novel fine-grained image recognition framework using user click data, which can bridge the semantic gap in distinguishing categories that are similar in visual. As query set in click data is usually large-scale and redundant, we first propose a click-feature-based query-merging approach to merge queries with similar semantics and construct a compact click feature. Afterward, we utilize this compact click feature and convolutional neural network (CNN)-based deep visual feature to jointly represent an image. Finally, with the combined feature, we employ the metriclearning-based template-matching scheme for efficient recognition. Considering the heavy noise in the training data, we introduce a reliability variable to characterize the image reliability, and propose a weakly-supervised metric and template leaning with smooth assumption and click prior (WMTLSC) method to jointly learn the distance metric, object templates, and image reliability. Extensive experiments are conducted on a public Clickture-Dog dataset and our newly established Clickture-Bird dataset. It is shown that the click-data-based query merging helps generating a highly compact (the dimension is reduced to 0.9%) and dense click feature for images, which greatly improves the computational efficiency. Also, introducing this click feature into CNN feature further boosts the recognition accuracy. The proposed framework performs much better than previous state-of-the-arts in fine-grained recognition tasks.
Min Tan 0005, Jun Yu 0002, Zhou Yu 0001, Fei Gao 0006, Yong Rui, Dacheng Tao
ACM Trans. Multim. Comput. Commun. Appl.1
2017 Fine-grained image recognition via weakly supervised click data guided bilinear CNN model
abstract
Bilinear convolutional neural networks (BCNN) model, the state-of-the-art in fine-grained image recognition, fails in distinguishing the categories with subtle visual differences. We design a novel BCNN model guided by user click data (C-BCNN) to improve the performance via capturing both the visual and semantical content in images. Specially, to deal with the heavy noise in large-scale click data, we propose a weakly supervised learning approach to learn the C-BCNN, namely W-C-BCNN. It can automatically weight the training images based on their reliability. Extensive experiments are conducted on the public Clickture-Dog dataset. It shows that: (1) integrating CNN with click feature largely improves the performance; (2) both the click data and visual consistency can help to model image reliability. Moreover, the method can be easily customized to medical image recognition. Our model performs much better than conventional BCNN models on both the Clickture-Dog and medical image dataset.
Guangjian Zheng, Min Tan 0005, Jun Yu 0002, Jianping Fan 0001
ICME2
2017 DeepSim: Deep similarity for image quality assessment
Fei Gao 0006, Panpeng Li, Min Tan 0005, Jun Yu 0002, Yani Zhu
Neurocomputing4
2016 Photo aesthetic quality assessment via label distribution learning
abstract
Automatic prediction of photo aesthetic quality is useful for many practical purposes. Current computational approaches typically solved this problem by assigning a categorical label (good or bad) to a photo. However, due to the subjectivity and complexity of humans aesthetic judgments, only a categorical label is insufficient to represent humans perceived aesthetic quality of a photo. This paper focuses on an interesting problem: is it possible to predict the crowed opinions about the aesthetic quality of a photo? The crowed opinion here is expressed by the distribution of scores given by a number of subjects. For each given photo, a deep convolutional neural network (DCNN) is utilized to calculate its feature representation. Afterwards, the crowed opinion prediction problem is formulated as one of label distribution learning (LDL). Experiments show that the proposed method is highly effective and outperforms state-of-the-art algorithms.
Fei Gao 0006, Di Huang 0007, Min Tan 0005, Jun Yu 0002
SMC4