Nan Gao 0001

dblp:87/4209-1 · DBLP profile ↗
← Back
32ranked-venue papers
13as first author
31since 2021 · last 2026
0000-0002-9694-2689ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 "Creating the World with Order!': Designing Tangible Toolkit to Support Creative Expression and Wellbeing for Individuals with ASD
abstract
Autistic people often experience heightened emotional and sensory demands that affect wellbeing. Creative drawing supports expression and emotion regulation, yet an open-ended process can introduce uncertainty and overstimulation, particularly for autistic individuals who prefer structure. We present the Structured Creativity Toolkit (SCT), a set of tangible artifacts including custom laser-cut stencils, composition templates, and guided color palettes, designed to scaffold creative expression and self-regulation. We first conducted a qualitative study to identify design opportunities and user needs, then evaluated SCT through a mixed-methods approach combining a controlled study with post-study interviews to examine engagement, user experience, and self-reported wellbeing. Our findings indicate that SCT increased drawing engagement and perceived drawing quality while supporting self-reported psychological and emotional benefits. This paper contributes: (1) a creativity-support toolkit for structured drawing, and (2) evidence that order-affirming scaffolds can translate preferences for structure into design resources that support autistic individuals.
Yibo Meng, Lyumanshan Ye, Bingyi Liu, Linghao Li, Nan Gao 0001
Creativity & Cognition7
2026 Division of Labor and Collaboration Between Parents in Family Education
abstract
Homework tutoring work is a demanding and often conflict-prone practice in family life, and parents often lack targeted support for managing its cognitive and emotional burdens. Through interviews with 18 parents of children in grades 1–3, we examine how homework-related labor is divided and coordinated between parents, and where AI might meaningfully intervene. We found three key insights: (1) Homework labor encompasses distinct dimensions: physical, cognitive, and emotional, with the latter two often remaining invisible. (2) We identified father-mother-child triadic dynamics in labor division, with children’s feedback as the primary factor shaping parental labor adjustments. (3) Building on prior HCI research, we propose an AI design that prioritizes relationship maintenance over task automation or broad labor mitigation. By employing labor as a lens that integrates care work, we explore the complexities of labor within family contexts, contributing to feminist and care-oriented HCI and to the development of context-sensitive coparenting practices.
Congrong Zhang, Jingying Deng, Xiaofan Hu, Jie Cai 0003, Nan Gao 0001, Chun Yu, Haining Zhang
CHI6
2026 An entropy-driven method for llm dataset evaluation and optimization
Meiping Wang, Rongduo Han, Liming Kang, Nan Gao 0001, Shihao Song, Yuelong Zhu, Chenghao He, Jing Xiang, Haining Zhang
Expert Syst. Appl.5
2026 Relational Mediators: LLM Chatbots as Boundary Objects in Psychotherapy CSCW033
abstract
As large language models (LLMs) are embedded into mental health technologies, they are often framed either as tools assisting therapists or autonomous therapeutic systems. Such perspectives overlook their potential to mediate relational complexities in therapy, particularly for systemically marginalized clients. Drawing on in-depth interviews with 12 therapists and 12 marginalized clients in China, including LGBTQ+ individuals or those from other marginalized backgrounds, we identify enduring relational challenges: difficulties building trust amid institutional barriers, the burden clients carry in educating therapists about marginalized identities, and challenges sustaining authentic self-disclosure across therapy and daily life. We argue that addressing these challenges requires AI systems capable of actively mediating underlying knowledge gaps, power asymmetries, and contextual disconnects. To this end, we propose the Dynamic Boundary Mediation Framework , which reconceptualizes LLM-enhanced systems as adaptive boundary objects that shift mediating roles across therapeutic stages. The framework delineates three forms of mediation: Epistemic (reducing knowledge asymmetries), Relational (rebalancing power dynamics), and Contextual (bridging therapy-life discontinuities). This framework offers a pathway toward designing relationally accountable AI systems that center the lived realities of marginalized users and more effectively support therapeutic relationships.
Jiatao Quan, Tian Qi Zhu, Baoying Wang, Wanda Pratt, Nan Gao 0001
Proc. ACM Hum. Comput. Interact.7
2026 InfoBFR: Real-World Blind Face Restoration via Information Bottleneck
abstract
Previous blind face restoration (BFR) methods have primarily leveraged facial priors from pretrained GAN or diffusion models. These neural BFR models suffer from diverse neural degradations, such as prior bias, topological distortion, textural distortion, and artifact residues, which limit their real-world generalization in complex real-world scenarios. In this paper, we propose an effective framework,InfoBFR, to address neural degradation from an information-theoretic perspective, which achieves BFR boosting in diverse wild and heterogeneous scenes. Specifically, on the basis of the results from pretrained BFR models, InfoBFR considers information compression by using a manifold information bottleneck (MIB) and manifold information compensation (MIC) with efficient diffusion LoRA to conduct information optimization. InfoBFR effectively synthesizes high-fidelity faces with texture and structure boosting. Comprehensive experimental results demonstrate the high boosting performance of InfoBFR (nearly 82%) for state-of-the-art GAN-based and diffusion-based BFR methods, as it can complete BFR tasks in approximately 70 ms and has 4M trainable parameters. It is promising that InfoBFR is the first unified postprocessing restorer universally employed by diverse BFR models to overcome the limitations of neural degradation.
Nan Gao 0001, Jia Li 0044, Huaibo Huang, Ran He 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 DiffGen: Optimizing I/O Trace Generation with Differentiated Modeling Techniques
Jian Liu 0053, Zhiyang Feng, Ziguang Fu, Guodao Sun, Yilong Zhang 0001, Nan Gao 0001, Ronghua Liang, Peng Chen 0008
ICA3PP (5)6
2025 Automated Construction of High-quality Evaluation Datasets Based on LLMs
Liming Kang, Rongduo Han, Meiping Wang, Nan Gao 0001, Haining Zhang
ICIC (23)5
2025 Dual Teacher with Dempster-Shafer Guidance for Decision Making in Semi-Supervised Small Object Detection
abstract
Small-scale object detection remains a major challenge in semi-supervised object detection (SSOD), particularly in medical image analysis. Conventional teacher models often struggle to accurately capture the features of low-contrast small lesions, leading to noisy pseudo-labels in both localization and classification, which introduces severe uncertainty and degrades detection performance. To address this issue, we propose Dual Teacher, a novel multimodal semi-supervised detection framework designed to enhance pseudo-label reliability and improve small-scale lesion detection. Specifically, we introduce two complementary teacher models: Hybrid-Scale Teacher, which exploits downsampled views to strengthen multi-scale feature learning, and Entropy-Based Multi-Modal Teacher, which leverages entropy maps to refine the quality of small-scale pseudo-labels. To effectively fuse predictions from both teachers and resolve conflicts, we propose a Dempster-Shafer-based Dual-Teacher pseudo-label fusion strategy that explicitly models uncertainty and optimizes classification confidence. Additionally, we introduce a class-adaptive threshold mechanism that dynamically adjusts pseudo-label selection based on dual-teacher predictions, further boosting the recall of small-scale lesions. Extensive experiments on the Dental Disease Dataset, ChestX-Det and M3FD demonstrate that our method consistently surpasses state-of-the-art SSOD approaches. Code is available at: https://github.com/z316910/Dual-Teacher.git.
Nan Gao 0001, Junchao Zhu, Yilong Zhang 0001, Ronghua Liang, Guodao Sun, Peng Chen 0008
ACM Multimedia1
2025 CaVIT: An integrated method for image style transfer using parallel CNN and vision transformer
Zaifang Zhang, Shunlu Lu, Nan Gao 0001, YuXiao Yang
Appl. Intell.4
2025 Towards machine learning fairness in classifying multicategory causes of deaths in colorectal or lung cancer patients
abstract
Classification of patient multicategory survival outcomes is important for personalized cancer treatments. Machine learning (ML) algorithms have increasingly been used to inform healthcare decisions, but these models are vulnerable to biases in data collection and algorithm creation. ML models have previously been shown to exhibit racial bias, but their fairness towards patients from different age and sex groups have yet to be studied. Therefore, we compared the multimetric performances of five ML models (random forests, multinomial logistic regression, linear support vector classifier, linear discriminant analysis, and multilayer perceptron) when classifying colorectal cancer patients (n = 589) of various age, sex, and racial groups using The Cancer Genome Atlas data. All five models exhibited biases for these sociodemographic groups. We then repeated the same process on lung adenocarcinoma (n = 515) to validate our findings. Surprisingly, most models tended to perform more poorly overall for the largest sociodemographic groups. Methods to optimize model performance, including testing the model on merged age, sex, or racial groups, and creating a model trained on and used for an individual or merged sociodemographic group, show potential to reduce disparities in model performance for different groups. This is supported by our regression analysis showing associations between model choice and methodology used with reduced performance disparities across demographic subgroups. Notably, these methods may be used to improve ML fairness while avoiding penalizing the model for exhibiting bias and thus sacrificing overall performance.
Catherine H. Feng, Mary L. Disis, Nan Gao 0001, Lanjing Zhang
Briefings Bioinform.4
2025 Coulson-type integral formulas for the general (skew) Estrada index of a vertex
Lu Qiao, Shenggui Zhang, Nan Gao 0001
Discret. Appl. Math.4
2025 Multi-granularity semantic relational mapping for image caption
Nan Gao 0001, Renyuan Yao, Peng Chen 0008, Ronghua Liang, Guodao Sun, Jijun Tang
Expert Syst. Appl.1
2025 TWDT: Training-free word-level controllable diffusion model for text generation
Nan Gao 0001, Yangjie Lu, Peng Chen 0008, Guodao Sun, Ronghua Liang, Yilong Zhang 0001
Knowl. Based Syst.1
2025 Generate anomalies from normal: a partial pseudo-anomaly augmented approach for video anomaly detection
Yuanjie Dang, Jiangyun Chen, Peng Chen 0008, Nan Gao 0001, Ruohong Huan
Vis. Comput.4
2024 ACPNet: Enhancing Small-Scale Dieases Detection in Panoramic X-rays
abstract
Deep learning-based disease detection can automatically identify dental diseases in panoramic X-rays and improve the accuracy and efficiency of doctors’ diagnoses. However, due to the complex data distribution of panoramic oral X-rays, significant scale differences among lesions, and the presence of many small-scale diseases, automated disease detection in panoramic oral X-rays faces considerable challenges. To alleviate the aforementioned issues, we propose ACPNet, which introduces a novel two-stage approach for detecting small-scale dental diseases in panoramic X-rays using the Contextual Attention Alignment Network (CAAN) and the Point-to-Patch Module (PTPM). To the best of our knowledge, we are the first to explore the detection of small-scale dental diseases in panoramic X-rays under limited sample conditions. Specifically, CAAN integrates deformable convolution with the global attention mechanism of transformer attention, enabling the model to more accurately extract small target foreground features in the complex background of panoramic X-rays. PTPM employs key point detection and cascade dynamic patches to adjust the bounding boxes of lesions, ensuring that small-scale diseases have sufficient high-quality proposals, thereby enhancing detector performance. Additionally, we collected a dataset containing 1157 instances of dental diseases to validate the effectiveness of our algorithm. Extensive experiments demonstrate that ACPNet achieves state-of-the-art performance, highlighting its superiority over baseline and other detection methods.
Nan Gao 0001, Junchao Zhu, Peng Chen 0008, Jijun Tang, Ronghua Liang
BIBM1
2024 Graph-Guided Multi-view Text Classification: Advanced Solutions for Fast Inference
Nan Gao 0001, Peng Chen 0008
ICANN (5)1
2024 Task-Agnostic Self-Distillation for Few-Shot Action Recognition
Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Nan Gao 0001, Ruohong Huan, Xiaofei He 0001
IJCAI5
2024 Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality Assessment
abstract
Current state-of-the-art video quality assessment (VQA) models typically integrate various perceptual features to comprehensively represent video quality degradation. These models either directly concatenate features or fuse different perceptual scores while ignoring the domain gaps between cross-aware features, thus failing to adequately learn the correlations and interactions between different perceptual features. To this end, we analyze the independent effects and information gaps of quality-and semantic-aware features on video quality. Based on an analysis of the spatial and temporal differences between two aware features, we propose a semantic-Aware and quality-Aware Interaction Network (A2INet) for blind VQA. For spatial gaps, we introduce a cross-aware guided interaction module to enhance the interaction between semantic-and quality-aware features in a local-to-global manner. Considering temporal discrepancies, we design a cross-aware temporal modeling module to further perceive temporal content variation and quality saliency information, and perceptual features are regressed into quality score by a temporal network and a temporal pooling. Extensive experiments on six benchmark VQA datasets show that our model achieves state-of-the-art performance, and ablation studies further validate the effectiveness of each module. We also present a simple video sampling strategy to balance the effectiveness and efficiency of the model. The code for the proposed method will be released at https://github.com/JianjunXiang/A2INet.
Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan, Nan Gao 0001
ACM Multimedia6
2024 A Unified Editing Method for Co-Speech Gesture Generation via Diffusion Inversion
Zeyu Zhao 0005, Nan Gao 0001, Guixuan Zhang, Jie Liu 0028, Shuwu Zhang
MMAsia2
2024 Focus on Subtle Actions: Semantic and Saliency Knowledge Co-Propagation Method for Weakly-Supervised Temporal Action Localization
Yuanjie Dang, Haoyu Shou, Peng Chen 0008, Nan Gao 0001, Ruohong Huan, Yilong Zhang 0001
PRCV (10)4
2024 TLCSFI: A Pose-Guided Person Re-Identification Method with Two-Level Channel-Spatial Feature Integration
abstract
Person re-identification methods currently encounter challenges in feature learning, primarily due to difficulties in expressing the correlation between local features and integrating global and local features effectively. To address these issues, a pose-guided person re-identification method with Two-Level Channel–Spatial Feature Integration (TLCSFI) is proposed. In TLCSFI, a two-level integration mechanism is implemented. At the first level, TLCSFI integrates the spatial information from local features to generate fine-grained spatial features. At the second level, the fine-grained spatial feature and the coarse-grained channel feature are integrated together to complete channel–spatial feature integration. In the method, a Pose-based Spatial Feature Integration (PSFI) module is introduced to generate the pose union feature, which calculates intra-body affinity to guide the integration of spatial information among local pose feature maps. Then, a Channel and Spatial Union Feature Integration (CSUFI) module is proposed to efficiently integrate the channel information of the global feature and the spatial information of the pose union feature. Two individual networks are designed to extract channel and spatial information, respectively, in CSUFI, which are then weighted and integrated. Experiments are conducted on three publicly available datasets to evaluate TLCSFI, and the experimental results demonstrate its competitive performance.
Ruohong Huan, Nan Gao 0001, Peng Chen 0008, Ronghua Liang
Int. J. Pattern Recognit. Artif. Intell.3
2024 Learning Reliable Dense Pseudo-Labels for Point-Level Weakly-Supervised Action Localization
abstract
Abstract Point-level weakly-supervised temporal action localization aims to accurately recognize and localize action segments in untrimmed videos, using only point-level annotations during training. Current methods primarily focus on mining sparse pseudo-labels and generating dense pseudo-labels. However, due to the sparsity of point-level labels and the impact of scene information on action representations, the reliability of dense pseudo-label methods still remains an issue. In this paper, we propose a point-level weakly-supervised temporal action localization method based on local representation enhancement and global temporal optimization. This method comprises two modules that enhance the representation capacity of action features and improve the reliability of class activation sequence classification, thereby enhancing the reliability of dense pseudo-labels and strengthening the model’s capability for completeness learning. Specifically, we first generate representative features of actions using pseudo-label feature and calculate weights based on the feature similarity between representative features of actions and segments features to adjust class activation sequence. Additionally, we maintain the fixed-length queues for annotated segments and design a action contrastive learning framework between videos. The experimental results demonstrate that our modules indeed enhance the model’s capability for comprehensive learning, particularly achieving state-of-the-art results at high IoU thresholds.
Yuanjie Dang, Guozhu Zheng, Peng Chen 0008, Nan Gao 0001, Ruohong Huan, Ronghua Liang
Neural Process. Lett.4
2024 A multi-stage recognizer for nested named entity with weakly labeled data
Nan Gao 0001, Bowei Yang, Peng Chen 0008, Li Ping Qian 0001
J. Supercomput.1
2024 Multi-Level Objective Alignment Transformer for Fine-Grained Oral Panoramic X-Ray Report Generation
abstract
Automatically generated oral panoramic X-ray report is highly beneficial for improving the efficiency of dental diagnosis. However, recent solutions adopt holistic methods, resulting in a cursory description of the oral condition. This may lead to reports lacking details, such as specific sites or lesion contours. Therefore, we propose a Multi-Level objective Alignment Transformer(MLAT) network, which integrates all tooth and disease objects into a positional alignment graph to extract fine-grained object-level features. Specifically, we introduce a novel Object-Level Collaborative Encoder (OLCE) module, which uses a positional alignment graph to construct object relationships. OLCE enhances object-level feature extraction by eliminating interference information between pathologically unrelated objects. In addition, we build a high-quality panoramic X-ray image-report dataset consisting of 562 sets of images and reports labeled by 13 experienced dental specialists. Experiments on the collected dataset show that the proposed MLAT significantly outperforms the state-of-the-art baselines by more than 5% in 4 different metrics, including BLEUs, Meteor, Rouge, and BERTScore.
Nan Gao 0001, Renyuan Yao, Ronghua Liang, Peng Chen 0008, Tianshuang Liu, Yuanjie Dang
IEEE Trans. Multim.1
2024 Discriminative Action Snippet Propagation Network for Weakly Supervised Temporal Action Localization
abstract
Weakly supervised temporal action localization (WTAL) aims to classify and localize actions in untrimmed videos with only video-level labels. Recent studies have attempted to obtain more accurate temporal boundaries by exploiting latent action instances in ambiguous snippets or propagating representative action features. However, empirically handcrafted ambiguous snippet extraction and the imprecise alignment of representative snippet propagation lead to challenges in modeling the completeness of actions for these methods. In this article, we propose a Discriminative Action Snippet Propagation Network (DASP-Net) to accurately discover ambiguous snippets in videos and propagate discriminative instance-level features throughout the video for improving action completeness. Specifically, we introduce a novel discriminative feature propagation module for capturing the global contextual attention and propagating the action concept across the whole video by perceiving the discriminative action snippets with instance information from the same video. Simultaneously, we incorporate denoised pseudo-labels as supervision, where we correct the controversial prediction based on the feature space distribution during training, thereby alleviating false detection caused by noise background features. Furthermore, we design an ambiguous feature mining module, which maximizes the feature affinity information of action and background in ambiguous snippets to generate more accurate latent action and background snippets and learns more precise action instance boundaries through contrastive learning of action and background snippets. Extensive experiments show that DASP-Net achieves state-of-the-art results on THUMOS14 and ActivityNet1.2 datasets.
Yuanjie Dang, Chunxia Huang, Peng Chen 0008, Nan Gao 0001, Ronghua Liang, Ruohong Huan
ACM Trans. Multim. Comput. Commun. Appl.5
2023 BTCN: Bridging the Gap Between Pre-trained and Downstream Models for Endoscopic Caries Detection
abstract
Although deep learning has been widely applied in the field of dental caries detection, there are still certain challenges that need to be addressed. The limitations of sharing the same backbone between the pre-trained model and the downstream model hinder the feature alignment capability of self-supervised learning (SSL) during the fine-tuning stage, leading to incomplete transfer from the pre-trained model to the downstream model. To address this challenge, we introduce an SSL pre-trained model called Bi-branches Transformer CNN Network (BTCN). BTCN adopts a parallel structure combining the CNN and Transformer branches. This parallel structure allows the pre-trained model to capture additional global representations, which helps alleviate feature differences during fine-tuning and better adapt to downstream detection models. Additionally, to further enhance the fusion quality of the bi-branches encoder, we introduced the Multi-layer Supervision Strategy (MSS) to increase the supervision on features at different layers. To validate the effectiveness of our approach, we collected a dedicated dataset for caries detection, comprising 1039 endoscopic images of dental caries. Through extensive experimental research, our results demonstrate the effectiveness of the proposed BTCN and MSS, showing significant improvements compared to the current state-of-the-art methods.
Nan Gao 0001, Peng Chen 0008, Yukai Li, Jijun Tang, Ronghua Liang, Tianshuang Liu
BIBM1
2023 NTAM: A New Transition-Based Attention Model for Nested Named Entity Recognition
Nan Gao 0001, Bowei Yang, Peng Chen 0008
NLPCC (2)1
2023 Boosting Short Text Classification by Solving the OOV Problem
abstract
In the field of natural language processing, text classification has received a lot of attention. Compared with long texts, short texts have fewer words and lack contextual semantic information. Existing approaches enrich short text information by linking the external knowledge graph, but they ignore the out-of-vocabulary (OOV) problem during entity linking, especially when dealing with domain-oriented data, which has some rare words or domain-specific nouns. In this paper, to alleviate the OOV problem caused by linking the external knowledge graph(KG), we propose a domain knowledge graph and entity complementation strategy to improve the performance of short text classification. Specifically, the external knowledge graph is used to enrich the information of short texts. The self-build domain knowledge graph is used to solve the problem of entities failing to link to the external knowledge graph. Finally, we conduct experiments on various datasets: 1. a labeled Chinese electronic domain dataset; 2. an open-source dataset to test the performance of our algorithm in different data distribution scenarios. The results demonstrate our dual knowledge graph model outperforms the state-of-the-art short text classification methods, especially when the OOV problem is severe.
Nan Gao 0001, Peng Chen 0008, Jijun Tang
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Disentangled and controllable sketch creation based on disentangling the structure and color enhancement
abstract
Abstract Existing sketch‐based image processing methods include sketch recognition, sketch synthesis and sketch‐based image retrieval. For sketch creation, a meaningful task is proposed namely disentangled and controllable sketch creation (DCSC) based on disentangling the structure and color enhancement. Specifically, as the first subtask, sketch structure enhancement (SSE) is used to enhance a non‐professional sketch (NPS) and obtain a professional sketch (PS), which is a process denoted as NPS2PS. A data set named SketchMan is first provided, consisting of NPSs and PSs with various postures in different scenes. SSE is trained as a conditional image‐to‐image translation problem, and there are three models: direct sketch‐to‐sketch (SS), grayscale guided SS and contour guided SS. Multiple IOU metrics are proposed based on Corner Point Map (CPM), Straight Line Map (SLM) and Segmented Area Map (SAM). As the second subtask, sketch color enhancement (SCE) is trained as a two‐stage framework containing a topology enhancement network (TE‐Net) that maps a sketch to the corresponding grayscale domain and a color injection network (CI‐Net) that injects the global color feature to the AdaIN residual blocks to perform adaptive sketch colorization. The TE‐Net and CI‐Net disentangle the topological and color features to perform more controllable and diverse SCE results. Experimental results demonstrate that our proposed methods are effective to address the challenging and meaningful DCSC task compared with other state‐of‐the‐art methods.
Nan Gao 0001, Hui Ren 0002, Jia Li 0044, Zhibin Su
IET Image Process.1
2022 Cutting the Unnecessary Long Tail: Cost-Effective Big Data Clustering in the Cloud
abstract
Clustering big data often requires tremendous computational resources where cloud computing is undoubtedly one of the promising solutions. However, the computation cost in the cloud can be unexpectedly high if it cannot be managed properly. The long tail phenomenon has been observed widely in the big data clustering area, which indicates that the majority of time is often consumed in the middle to late stages in the clustering process. In this research, we try to cut the unnecessary long tail in the clustering process to achieve a sufficiently satisfactory accuracy at the lowest possible computation cost. A novel approach is proposed to achieve cost-effective big data clustering in the cloud. By training the regression model with the sampling data, we can make widely used k-means and EM (Expectation-Maximization) algorithms stop automatically at an early point when the desired accuracy is obtained. Experiments are conducted on four popular data sets and the results demonstrate that both k-means and EM algorithms can achieve high cost-effectiveness in the cloud with our proposed approach. For example, in the case studies with the much more efficient k-means algorithm, we find that achieving a 99 percent accuracy needs only 47.71-71.14 percent of the computation cost required for achieving a 100 percent accuracy while the less efficient EM algorithm needs 16.69-32.04 percent of the computation cost. To put that into perspective, in the United States land use classification example, our approach can save up to $94,687.49 for the government in each use.
Dongwei Li, Shuliang Wang 0001, Nan Gao 0001, Qiang He 0001, Yun Yang 0001
IEEE Trans. Cloud Comput.3
2022 Generative Adversarial Networks for Spatio-temporal Data: A Survey
abstract
Generative Adversarial Networks (GANs) have shown remarkable success in producing realistic-looking images in the computer vision area. Recently, GAN-based techniques are shown to be promising for spatio-temporal-based applications such as trajectory prediction, events generation, and time-series data imputation. While several reviews for GANs in computer vision have been presented, no one has considered addressing the practical applications and challenges relevant to spatio-temporal data. In this article, we have conducted a comprehensive review of the recent developments of GANs for spatio-temporal data. We summarise the application of popular GAN architectures for spatio-temporal data and the common practices for evaluating the performance of spatio-temporal applications with GANs. Finally, we point out future research directions to benefit researchers in this area.
Nan Gao 0001, Hao Xue 0001, Wei Shao 0006, Sichen Zhao, Kyle Kai Qin, Arian Prabowo, Mohammad Saiedur Rahaman, Flora D. Salim
ACM Trans. Intell. Syst. Technol.1
2020 SketchMan: Learning to Create Professional Sketches
abstract
Human free-hand sketches have been studied in various fields including sketch recognition, synthesis and sketch-based image retrieval. We propose a new challenging task sketch enhancement (SE) defined in an ill-posed space, i.e. enhancing a non-professional sketch (NPS) to a professional sketch (PS), which is a creative generation task different from sketch abstraction, sketch completion and sketch variation. For the first time we release a database of NPS with PS for anime characters. We cast sketch enhancement as an image-to-image translation problem by exploiting the relationship to corresponding intensive or sparse pixel domains for sketch domain. Specifically, we explore three different routines based on conditional generative adversarial network (cGAN), i.e. Sketch-Sketch (SS), Sketch-Colorization-Sketch (SCS) and Sketch-Abstraction-Sketch (SAS). SS is a one-stage model that directly maps NPS to PS, while SCS and SAS are two-stage models where auxiliary inputs, grayscale parsing and shape parsing, are involved. Multiple metrics are used to evaluate the performance of the models in both the sketch domain and other low-level feature domains. With quantitative and qualitative analysis of the experiments, we have established solid baselines, which, we hope, could encourage more research conducted on this task. Our dataset is publicly available via https://github.com/LCXCUC/SketchMan2020.
Jia Li 0044, Nan Gao 0001, Wei Zhang 0031, Tao Mei 0001, Hui Ren 0002
ACM Multimedia2