EDBT 2026 Demo / reviewers in the wild / expert
Ping Li 0016
dblp:62/5860-16
· DBLP profile ↗
209ranked-venue papers
8as first author
140since 2021 · last 2026
0000-0002-1503-0240ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 144 · 4 first-author · 97 since 2021Artificial intelligence and machine learning · 51 · 1 first-author · 45 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Real-World High-Precision Image Matting and SegmentationabstractHigh-precision scene parsing tasks, including image matting and dichotomous segmentation, aim to accurately predict masks with extremely fine details (such as hair). Most existing methods focus on salient, single foreground objects. While interactive methods allow for target adjustment, their class-agnostic design restricts generalization across different categories. Furthermore, the scarcity of high-quality annotation has led to a reliance on inharmonious synthetic data, resulting in poor generalization to real-world scenarios. To this end, we propose a Foreground Consistent Learning model, dubbed as FCLM, to address the aforementioned issues. Specifically, we first introduce a Depth-Aware Distillation strategy where we transfer the depth-related knowledge for better foreground representation. Considering the data dilemma, we term the processing of synthetic data as domain adaptation problem where we propose a domain-invariant learning strategy to focus on foreground learning. To support interactive prediction, we contribute an Object-Oriented Decoder that can receive both visual and language prompts to predict the referring target. Experimental results show that our method quantitatively and qualitatively outperforms state-of-the-art methods. Haipeng Zhou, Zhaohu Xing, Hongqiu Wang, Jun Ma 0008, Ping Li 0016, Lei Zhu 0003 |
AAAI | 5 |
| 2026 | SynTaskNet: A synergistic multi-task network for joint segmentation and classification of small anatomical structures in ultrasound imaging
Abdulrhman H. Al-Jebrni, Saba Ghazanfar Ali, Bin Sheng 0001, Huating Li, Xiao Lin 0012, Ping Li 0016, Younhyun Jung, Jinman Kim, Lixin Jiang |
Comput. Vis. Image Underst. | 6 |
| 2026 | Slimmable neural architecture design based on cross architecture and token distillation
Guhao Qiu, Ping Li 0016, Bin Sheng 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | FOSP: Feature Orientated Sand Painting GenerationabstractABSTRACT Sand painting is a visually distinctive art form characterized by granular textures and diverse expressive techniques, where sand grains are manipulated through various hand movements such as waving, seeping, sweeping, and stroking. However, traditional stylized methods often fail to capture the fine textures and diverse techniques unique to sand painting. In this paper, we propose a Feature‐Oriented Sand Painting (FOSP) inspired by real sand painting techniques, aiming to produce sand paintings with authentic sandy textures and diverse techniques. Our FOSP comprises three main modules: overall sand waving generation, extraction and drawing of sand reduction regions, and simulation of sand grain accumulation with multiple techniques. The first module generates fine sand‐grain textures, whereas the latter two focus on contour rendering to enrich detail expression. Experiments validate that our FOSP can generate high‐quality sand paintings rapidly, outperforming existing methods in sand painting generation tasks. Meng Yang 0011, Mengting Zhu, Jianglang Kang, Weiliang Meng, Ping Li 0016 |
Comput. Animat. Virtual Worlds | 5 |
| 2026 | FDRM-Net: A Mamba structure for single image deraining with frequency guidance
Xiao Lin 0012, Lizhuang Ma, Ping Li 0016 |
J. Vis. Commun. Image Represent. | 5 |
| 2026 | A two-stage active cleaning strategy for long-tail label noise
Xiao Lin 0012, Zeyu Rong, Yan Li 0063, Qizhe Yang, Ping Li 0016 |
Neural Networks | 5 |
| 2026 | Autorep: Automatic network search with structured reparameterized based linear operation expansion and gradient proxy guided reduction
Guhao Qiu, Ruoxin Chen, Ping Li 0016, Bin Sheng 0001 |
Neural Networks | 5 |
| 2026 | Dataset Distillation via a Noise-Unconstrained Generative ModelabstractDataset distillation (DD) aims to synthesize a more compact dataset than the original one and models trained on it are expected to have the same generalization capabilities as on the original dataset. Previous work via a generative model (GM) faces several limitations. First, GM struggles to generate representative samples due to a lack of constraints. Second, it overlooks the relationships between generated samples, limiting its effectiveness. In this paper, a new noise-unconstrained GM-based DD framework is proposed. In the distillation stage, an adaptive matching coefficient is introduced to align generated images with representative class elements and the MiniMax loss function is extended to reduce the optimization difficulty. In the deployment stage, features among each generative image are ensembled by gradient-matching based DD. Theoretical analysis based on McDiarmid's inequality demonstrates that the proposed components can reduce the generalization error of the original baseline method. We also provide insights into the potential of generated images as an effective proxy dataset for DD. For example, on the ImageWoof dataset with 50 distilled images per class using a 6-layer ConvNet for evaluation, generated images outperform 25%, 50%, and 75% original images by 8.4%, 6.3%, and 8.3% in distillation performance. Our method effectively handles both low- and high-resolution datasets, with experiments on 11 benchmarks demonstrating its efficacy. Fei Ye 0004, Ping Li 0016, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Intra-modal consistency for image-text retrieval through soft-label distillation
Yangtao Wang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, C. L. Philip Chen, Ping Li 0016, Wensheng Zhang 0002 |
Pattern Recognit. | 9 |
| 2026 | Multiple weather degraded image restoration based on multi-component decomposition
Xiao Lin 0012, Duojiu Xu, Qizhe Yang, Yan Li 0063, Ping Li 0016 |
Pattern Recognit. | 7 |
| 2026 | Dynamic patch-level contrastive learning for image dehazing
Xiao Lin 0012, Dongchen Zhang, Yan Li 0063, Qizhe Yang, Ping Li 0016 |
Pattern Recognit. | 5 |
| 2026 | LODNeuS: A Flexible Lightweight Neural Implicit Surface Representation With Unconstrained Viewpoint RenderingabstractNeRF-like methods learn implicit 3D neural representations from 2D multiview images, enabling the synthesis of compelling novel views. However, to capture high-fidelity geometry, prior methods often rely on large-scale networks. This dependency hampers the potential applications of neural implicit representations, such as MR visualization. To address this, we introduce LODNeuS, an implicit surface representation based on feature voxel grids. LODNeuS captures multiple LODs of implicit geometry by maintaining voxel grids paired with a set of corresponding lightweight decoders. This allows for high-quality rendering with the ability to dynamically switch between detail levels. Another challenge is that existing methods, both volumetric and surface-based, tend to train and render their representations within a confined space, without explicitly restricting the sampling points properly. This lack of constraints can result in ambiguity, artifacts, and inefficient use of computational resources. We study this effect during free viewpoint rendering using conventional methods and develop an adaptive sampling scheme that emphasizes a valid geometric space for sampling point allocation. Our experimental results show that LODNeuS can match the visual quality of existing methods while offering flexible and lightweight inference. The benefits of adaptive sampling are also demonstrated in the free viewpoint rendering subsection. Our work extends the capabilities of neural implicit representations beyond previously defined limitations, broadening the scope of potential applications. Ping Li 0016, Lei Zhu 0003, Bin Sheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Addressing Client Drift in Federated Learning via Class-Prototype Similarity Distillation and Adaptive MaskabstractFederated learning (FL) enables multiple clients to learn collaboratively in a distributed way, allowing for privacy protection. However, the real-world nonindependent and identically distributed (non-IID) data will lead to client drift, which degrades the performance of FL. Interestingly, we find that the logit difference between the local and global models increases as the model is continuously updated, which is the primary factor behind performance degradation. This is mainly due to catastrophic forgetting caused by non-IID data between clients. To alleviate this problem, we propose a new algorithm, named FedCSD, a class-prototype similarity distillation in a federated framework to align the logits of local and global models. FedCSD does not simply transfer global knowledge to local clients, as an insufficiently trained global model cannot provide reliable knowledge, i.e., class similarity information, and its wrong soft labels will mislead the optimization of local models. Concretely, FedCSD leverages the similarity between local logits and the global prototype to refine the global logits, thereby enhancing its class similarity information. Furthermore, FedCSD adopts an adaptive mask to filter out the terrible soft labels of the global models, thereby preventing them from misleading local optimization. Extensive experiments demonstrate the superiority of our method over the state-of-the-art FL approaches in various non-IID settings. Code is publicly available at https://github.com/IAMJackYan/FedCSD. Yunlu Yan, Chun-Mei Feng 0001, Mang Ye, Wangmeng Zuo, Ping Li 0016, Rick Siow Mong Goh, Lei Zhu 0003, C. L. Philip Chen |
IEEE Trans. Cybern. | 5 |
| 2026 | Learn From Examples: In-Context Learning for Camouflaged Object DetectionabstractRecently, new paradigms of camouflaged object detection (COD), such as referring COD (Ref-COD) and collaborative COD (Co-COD), have been proposed to enhance task performance. However, there remains a lack of in-depth exploration of how to utilize reference information more effectively. In this paper, we introduce in-context learning camouflaged object detection (ICL-COD) as a novel paradigm of COD, which leverages camouflaged image samples and their corresponding annotations as visual examples to guide the model in better perceiving camouflage and recognizing camouflaged objects. We propose the ICL-Camo network, with the design of a context mining module (CMM) to mine fine-grained contextual information contained in the visual examples, and a context guiding module (CGM) that utilizes the contextual information mined from the examples as guidance to shift the attention of the target image features on potential camouflaged regions, thus enhancing its perception of camouflaged objects. Extensive experiments conducted on the COD benchmarks and other relevant tasks demonstrate the effectiveness of our proposed ICL-COD paradigm and ICL-Camo network. Code and results are available at: https://github.com/h0t-zer0/ICL-Camo. Chunyuan Chen, Weiyun Liang, Ji Du, Jing Xu 0008, Ping Li 0016, Grace Guiling Wang |
IEEE Trans. Image Process. | 5 |
| 2026 | RA-COD: Retrieval-Augmented Camouflaged Object DetectionabstractCamouflaged Object Detection (COD) is pivotal for segmenting objects that seamlessly blend into their surroundings. While prior endeavors demonstrate impressive performance through training on predefined labels, they heavily rely on labor-intensive data annotation and struggle to adapt to open-world scenarios. In this light, we propose RA-COD, a training-free paradigm that enables COD by retrieving the most similar samples from the prototype repository. The efficacy of RA-COD hinges on 1) capturing the nuanced resemblance between objects and their environments and 2) excelling in dense prediction tasks. To achieve (1), the crux lies in ensuring diversity and discriminability within the prototype repository. In this context, we propose GenPro, an automated pipeline for crafting Generative Prototypes. GenPro integrates a range of foundation models, including the Diffusion Model, Vision-Language Model, Segment Anything Model (SAM), and DINOv2, in a complementary manner that synergistically generates diverse and distinguishable prototype samples. To achieve (2), we propose C2F to retrieve camouflaged objects in a Coarse-to-Fine regime. We commence with pixel-level retrieval in the feature space, which generates a coarse mask that effectively captures class discrimination and object localization. Further refinement is achieved by extracting bounding boxes from this coarse mask to prompt SAM in generating mask proposals for region-level retrieval. Evaluations on four benchmarks showcase that RA-COD achieves state-of-the-art performance compared to existing training-free methods. Ji Du, Jiesheng Wu, Desheng Kong, Fangwei Hao, Jing Xu 0008, Ping Li 0016 |
IEEE Trans. Image Process. | 6 |
| 2026 | Sketch2Avatar: Geometry-Guided 3D Full-Body Human Generation in 360° From Hand-Drawn SketchesabstractGenerating full-body humans in 360$^{\circ }$∘ has broad applications in digital entertainment, online education and art design. Existing works primarily rely on coarse conditions such as body pose to guide the generation, lacking detailed control over the synthesized results. Regarding this limitation, sketches offer a promising alternative as an expressive condition that enables more explicit and precise control. However, current sketch-based generation methods focus on faces or common objects, how to transfer sketches into 360 $^{\circ }$∘ full-body humans remains unexplored. To bridge this gap, we propose Sketch2Avatar, the first generative model to achieve 3D full-body human generation from hand-drawn sketches. Our model is capable of synthesizing sketch-aligned and 360$^{\circ }$∘-consistent full-body human images by leveraging the geometry information extracted from sketches to guide the 3D representation generation and neural rendering. Specifically, we propose sketchguided 3D representation generation to model the 3D human and maintain the alignment between input sketches and generated humans. Our transformer-based generator incorporates spatial feature guidance and latent modulation derived from sketches to produce high-quality 3D representations. Additionally, our designed bodyaware neural rendering utilizes 3D human body priors from sketches, simplifying the learning of articulated body poses and complex body shapes. To train and evaluate our model, we construct a large-scale dataset comprising approximately 19 K 2D full-body human images and their corresponding sketches in a hand-drawn style. Experimental results demonstrate that our Sketch2Avatar can transfer hand-drawn sketches into photo-realistic 360$^{\circ }$∘ full-body human images with precise sketch-human alignment. Ablation studies further validate the effectiveness of our design choices. Qiang Li 0024, Jie Zhang 0090, Anthony Kong, Ping Li 0016 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | Large model-assisted video summarization via global entity unification and robust importance scoring
Donglei Chen, Shaoyu Huang, Xuemiao Xu, Yongwei Nie, Ping Li 0016, C. L. Philip Chen |
Vis. Comput. | 5 |
| 2025 | Accurate KV Cache Quantization with Outlier Tokens TracingabstractThe impressive capabilities of Large Language Models (LLMs) come at the cost of substantial computational resources during deployment. While KV Cache can significantly reduce recomputation during inference, it also introduces additional memory overhead. KV Cache quantization presents a promising solution, striking a good balance between memory usage and accuracy. Previous research has shown that the Keys are distributed by channel, while the Values are distributed by token. Consequently, the common practice is to apply channel-wise quantization to the Keys and token-wise quantization to the Values. However, our further investigation reveals that a small subset of unusual tokens exhibit unique characteristics that deviate from this pattern, which can substantially impact quantization accuracy. To address this, we develop a simple yet effective method to identify these tokens accurately during the decoding process and exclude them from quantization as outlier tokens, significantly improving overall accuracy. Extensive experiments show that our method achieves significant accuracy improvements under 2-bit quantization and can deliver a 6.4 times reduction in memory usage and a 2.3 times increase in throughput. Yi Su 0006, Yuechi Zhou, Quantong Qiu, Juntao Li 0005, Qingrong Xia, Ping Li 0016, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ACL (1) | 6 |
| 2025 | Line Drawing Abstraction Based on Line Importance Evaluation
Shilong Deng, Xueting Liu 0001, Chengze Li, Ping Li 0016, Zhenkun Wen, Huisi Wu |
CGI (1) | 4 |
| 2025 | Embedding Space Decomposition Meets Invertible Networks: A New Paradigm for Unpaired Low-Light Enhancement
Linbo Wang 0001, Zhuo Yan, Zhengyi Liu, Xianyong Fang, Ping Li 0016 |
CGI (3) | 5 |
| 2025 | PeelMesh: Efficient Interactive Segmentation via Geodesic-Driven Dynamic Topological Updates
Junjie Yin, Zixi Huang, Meie Fang, Ping Li 0016, Weiyin Ma |
CGI (1) | 5 |
| 2025 | MIGEdit: Multimodal Interactive Garment Editing
Sicheng Zheng, Bangchao Wang, Jinxing Liang, Li Li 0094, Tao Peng 0006, Junping Liu, Ping Li 0016, Xinrong Hu |
CGI (2) | 8 |
| 2025 | Training-Free Language-Guided Video Summarization via Multi-Grained Saliency Scoring
Yongwei Nie, Fei Ma 0006, Keke Tang, F. Richard Yu, Hongmin Cai, Ping Li 0016 |
CVM (3) | 7 |
| 2025 | 3DFaceController: Region-Controllable Face Synthesis via Decomposed and Recomposed Neural Radiance Fields
Kangneng Zhou, Yaxing Wang, Shuang Song 0005, Jie Zhang 0090, Ping Li 0016 |
CVM (2) | 5 |
| 2025 | Shift the Lens: Environment-Aware Unsupervised Camouflaged Object DetectionabstractCamouflaged Object Detection (COD) seeks to distinguish objects from their highly similar backgrounds. Existing work has essentially focused on isolating camouflaged objects from the environment, demonstrating ever-improving performance but at the cost of extensive annotations and complex optimizations. In this paper, we diverge from this paradigm and shift the lens to isolating the salient environment from the camouflaged object. We introduce EASE, an Environment-Aware unSupErvised COD framework that identifies the environment by referencing an environment prototype library and detects camouflaged objects by inverting the retrieved environmental features. Specifically, our approach (DiffPro) uses large multimodal models, diffusion models, and vision-foundation models to construct the environment prototype library. To retrieve environments from the library and refrain from confusing foreground and background, we incorporate three retrieval schemes: Kernel Density Estimation-based Adaptive Threshold (KDE-AT), Global-to-Local pixel-level retrieval (G2L), and Self-Retrieval (SR). Our experiments demonstrate significant improvements over current unsupervised methods, with EASE achieving an average gain of over 10% on the COD10K dataset. When integrated with SAM, EASE surpasses prompt-based segmentation approaches and performs competitively with state-of-the-art fully-supervised methods. Code is available at https://github.com/xiaohainku/EASE. Ji Du, Fangwei Hao, Mingyang Yu 0001, Desheng Kong, Jiesheng Wu, Jing Xu 0008, Ping Li 0016 |
CVPR | 8 |
| 2025 | NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration
Haotian Dong, Xin Wang 0118, Di Lin 0002, Yipeng Wu, Kairui Yang, Ping Li 0016, Qing Guo 0005 |
ICCV | 8 |
| 2025 | Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object DetectionabstractAt the core of Camouflaged Object Detection (COD) lies segmenting objects from their highly similar surroundings. Previous efforts navigate this challenge primarily through image-level modeling or annotation-based optimization. Despite advancing considerably, this commonplace practice hardly taps valuable dataset-level contextual information or relies on laborious annotations. In this paper, we propose RISE, a RetrIeval SElf-augmented paradigm that exploits the entire training dataset to generate pseudo-labels for single images, which could be used to train COD models. RISE begins by constructing prototype libraries for environments and camouflaged objects using training images (without ground truth), followed by K-Nearest Neighbor (KNN) retrieval to generate pseudo-masks for each image based on these libraries. It is important to recognize that using only training images without annotations exerts a pronounced challenge in crafting high-quality prototype libraries. In this light, we introduce a Clustering-then-Retrieval (CR) strategy, where coarse masks are first generated through clustering, facilitating subsequent histogram-based image filtering and cross-category retrieval to produce high-confidence prototypes. In the KNN retrieval stage, to alleviate the effect of artifacts in feature maps, we propose Multi-View KNN Retrieval (MVKR), which integrates retrieval results from diverse views to produce more robust and precise pseudo-masks. Extensive experiments demonstrate that RISE outperforms state-of-the-art unsupervised and prompt-based methods. Code is available at https://github.com/xiaohainku/RISE. Ji Du, Xin Wang 0118, Fangwei Hao, Mingyang Yu 0001, Chunyuan Chen, Jiesheng Wu, Jing Xu 0008, Ping Li 0016 |
ICCV | 9 |
| 2025 | Beware of Calibration Data for Pruning Large Language ModelsabstractAs large language models (LLMs) are widely applied across various fields, model
compression has become increasingly crucial for reducing costs and improving
inference efficiency. Post-training pruning is a promising method that does not
require resource-intensive iterative training and only needs a small amount of
calibration data to assess the importance of parameters. Recent research has enhanced post-training pruning from different aspects but few of them systematically
explore the effects of calibration data, and it is unclear if there exist better calibration data construction strategies. We fill this blank and surprisingly observe that
calibration data is also crucial to post-training pruning, especially for high sparsity. Through controlled experiments on important influence factors of calibration
data, including the pruning settings, the amount of data, and its similarity with
pre-training data, we observe that a small size of data is adequate, and more similar data to its pre-training stage can yield better performance. As pre-training data
is usually inaccessible for advanced LLMs, we further provide a self-generating
calibration data synthesis strategy to construct feasible calibration data. Experimental results on recent strong open-source LLMs (e.g., DCLM, and LLaMA-3)
show that the proposed strategy can enhance the performance of strong pruning
methods (e.g., Wanda, DSnoT, OWL) by a large margin (up to 2.68%). Yixin Ji, Yang Xiang 0003, Juntao Li 0005, Qingrong Xia, Ping Li 0016, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ICLR | 5 |
| 2025 | Enhancing Cross-modal Semantic Consistency via Key Token Alignment for Image-text RetrievalabstractImage-text retrieval (ITR) plays a pivotal role in advancing intelligent transportation systems, facilitating efficient retrieval and utilization of multimedia data to enhance traffic management and safety significantly. However, existing ITR solutions have not effectively addressed the issues of image patch redundancy and text word redundancy, leading to erroneous image-text matching. In this paper, we propose SCTA that enhances cross-modal semantic consistency via key token alignment for ITR. Firstly, SCTA evaluates the importance of each image patch by calculating the self-attention scores within patches and cross-attention scores between patches and words. Secondly, SCTA implements aggregation operations on image and text separately, aiming to generate information-rich key image patch embeddings and text word token embeddings. Finally, SCTA completes fine-grained alignment by maximizing the similarity between patch-to-word and word-to-patch. Therefore, SCTA simultaneously addresses image patch redundancy and text word redundancy issues, enhancing semantic consistency by aligning the core semantic information between image-text pairs. Extensive experiments on multiple datasets including Flickr30K and MS-COCO verify the superior performance of SCTA compared with the SOTA fine-grained ITR methods. The code of this paper is released at GitHub: https://github.com/ICME2025ITR/SCTA. Huilong Lin, Yangtao Wang, Meie Fang, Yanzhao Xie, Xiaocui Li 0001, Weilong Peng, Siyuan Chen 0005, Maobin Tang, Ping Li 0016 |
ICME | 10 |
| 2025 | SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lightingabstract3D inverse rendering in indoor scenes with strong light sources presents a significant challenge, primarily due to the substantial ambiguity in material recovery caused by the complex interaction between lighting and shadows. To address this, we propose a novel approach that integrates an implicit-explicit shadow predictor with a three-stage material estimation process. Our method enhances shadow realism by accurately predicting light interactions, while our material estimation process improves SVBRDF quality under challenging lighting conditions. Extensive experiments demonstrate the effectiveness of our method in both quantitative and qualitative metrics, enabling realistic object insertion and material replacement with proper shadow rendering under strong indoor light sources. Xiaokang Wei, Zhuoman Liu, Ping Li 0016, Yan Luximon |
ICME | 3 |
| 2025 | Face2Wear: An automatic and user-friendly facewear personalization framework with 3D symmetry-aware face registration using RGB-D selfies
Jie Zhang 0090, Luwei Chen, Yan Luximon, Ping Li 0016 |
Comput. Aided Des. | 5 |
| 2025 | Implicit-based collision-aware clothed human reconstruction from a single image
Guiqing Li, Yongwei Nie, Feiran Yu, Ping Li 0016, Tonglai Liu, Zhao Zhang 0001 |
Comput. Graph. | 5 |
| 2025 | MeshPAD: Payload-aware mesh distortion for 3D steganography based on geometric deep learning
Weilong Peng, Keke Tang, Weixuan Tang 0002, Yong Su 0003, Meie Fang, Ping Li 0016 |
Expert Syst. Appl. | 6 |
| 2025 | Occlusion-Preserved Surveillance Video Synopsis with Flexible Object Graph
Yongwei Nie, Siming Zeng, Qing Zhang 0006, Guiqing Li, Ping Li 0016, Hongmin Cai |
Int. J. Comput. Vis. | 6 |
| 2025 | FunBreath: A novel interactive nebulizer mask with gamification system for children's effective and enjoyable treatment
Qiuyu Ye, Jingyan Yang, Jun Zhang 0072, Ping Li 0016, Yan Luximon, Jie Zhang 0090 |
Int. J. Hum. Comput. Stud. | 5 |
| 2025 | Gradient amplification for gradient matching based dataset distillation
Ping Li 0016, Bin Sheng 0001 |
Neural Networks | 4 |
| 2025 | UpGen: Unleashing Potential of Foundation Models for Training-Free Camouflage Detection via Generative ModelsabstractCamouflaged Object Detection (COD) aims to segment objects resembling their environment. To address the challenges of extensive annotations and complex optimizations in supervised learning, recent prompt-based segmentation methods excavate insightful prompts from Large Vision-Language Models (LVLMs) and refine them using various foundation models. These are subsequently fed into the Segment Anything Model (SAM) for segmentation. However, due to the hallucinations of LVLMs and insufficient image-prompt interactions during the refinement stage, these prompts often struggle to capture well-established class differentiation and localization of camouflaged objects, resulting in performance degradation. To provide SAM with more informative prompts, we present UpGen, a pipeline that prompts SAM with generative prompts without requiring training, marking a novel integration of generative models with LVLMs. Specifically, we propose the Multi-Student-Single-Teacher (MSST) knowledge integration framework to alleviate hallucinations of LVLMs. This framework integrates insights from multiple sources to enhance the classification of camouflaged objects. To enhance interactions during the prompt refinement stage, we are the first to leverage generative models on real camouflage images to produce SAM-style prompts without fine-tuning. By capitalizing on the unique learning mechanism and structure of generative models, we effectively enable image-prompt interactions and generate highly informative prompts for SAM. Our extensive experiments demonstrate that UpGen outperforms weakly-supervised models and its SAM-based counterparts. We also integrate our framework into existing weakly-supervised methods to generate pseudo-labels, resulting in consistent performance gains. Moreover, with minor adjustments, UpGen shows promising results in open-vocabulary COD, referring COD, salient object detection, marine animal segmentation, and transparent object segmentation. Ji Du, Jiesheng Wu, Desheng Kong, Weiyun Liang, Fangwei Hao, Jing Xu 0008, Grace Guiling Wang, Ping Li 0016 |
IEEE Trans. Image Process. | 9 |
| 2025 | Federated Pseudo Modality Generation for Incomplete Multi-Modal MRI ReconstructionabstractWhile multi-modal learning has been widely used for MRI reconstruction, it relies on paired multi-modal data, which is difficult to acquire in real clinical scenarios. Especially in the federated setting, there is a common issue that several medical institutions suffer from missing modalities or even only have single-modal data. Therefore, it is infeasible to deploy a standard federated learning framework in such conditions. In this paper, we propose a novel communication-efficient federated learning framework (namely Fed-PMG) to address the missing modality challenge in federated multi-modal MRI reconstruction. Specifically, we utilize a pseudo modality generation mechanism to recover the missing modality for each single-modal client by sharing the distribution information of the amplitude spectrum in frequency space. However, the step of sharing the original amplitude spectrum leads to heavy communication costs. To reduce the communication cost, we introduce a clustering scheme to project the set of amplitude spectrum into a finite number of cluster centroids and share them among the clients. With such an elaborate design, our approach can effectively complete the missing modality within an acceptable communication cost. Extensive experimental results demonstrate that our proposed method can outperform state-of-the-art methods and reach a performance similar to the ideal scenario (i.e., all clients have the full set of modalities). Yunlu Yan, Chun-Mei Feng 0001, Yuexiang Li, Ping Li 0016, Rick Siow Mong Goh, Bai Ying Lei, Weiming Wang 0002, David Dagan Feng, Lei Zhu 0003 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | VGNet: Multimodal Feature Extraction and Fusion Network for 3D CAD Model RetrievalabstractThe reuse of 3D CAD models is crucial for industrial manufacturing because it shortens development cycles and reduces costs. Significant progress has been made in deep learning-based 3D model retrievals. There are many representations for 3D models, among which the multi-view representation has demonstrated a superior retrieval performance. However, directly applying these 3D model retrieval approaches to 3D CAD model retrievals may result in issues such as the loss of the engineering semantic and structural information. In this paper, we find that multiple views and B-rep can complement each other. Therefore, we propose the view graph neural network (VGNet), which effectively combines multiple views and B-rep to accomplish 3D CAD model retrieval. More specifically, based on the characteristics of the regular shape of 3D CAD models, and the richness of the attribute information in the B-rep attribute graph, we separately design two feature extraction networks for each modality. Moreover, to explore the latent relationships between the multiple views and B-rep attribute graphs, a multi-head attention enhancement module is designed. Furthermore, the multimodal fusion module is adopted to make the joint representation of the 3D CAD models more discriminative by using a correlation loss function. Experiments are carried out on a real manufacturing 3D CAD dataset and a public dataset to validate the effectiveness of the proposed approach. Fei-wei Qin, Gaoyang Zhan, Meie Fang, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Multim. | 5 |
| 2025 | Accurate-PGNet: Learning to Assemble Perceptual Body Parts for Accurate Human Skeleton EstablishmentabstractThe human skeleton establishment aims to provide accurate localization information of the human body from RGB images and establish a complete human skeleton for many applications, such as action recognition, video surveillance, and human-computer interaction. Considering the inherent human body structure, many recent methods group the relevant body parts and utilize the deep convolutional network to learn the visual context from the part groups. However, the grouping approaches used in these methods heavily rely on prior knowledge of the human body shape but lose important relationships between parts. In this paper, we introduce the Accurate Part Grouping Network (Accurate-PGNet), a novel network for hierarchically grouping body parts in a data-driven manner. In contrast to the previous methods, we use neural architecture search (NAS) to optimize the architecture of Accurate-PGNet and properly group the body parts. The part grouping respects the diverse visual patterns of parts, producing groups containing different body parts. From each group, we learn the visual feature map. It helps to capture the correlation between parts and predict their locations. The feature maps of the part groups are merged hierarchically to capture the higher-order context of parts in larger groups. We extensively evaluated our method on the challenging benchmarks, demonstrating that Accurate-PGNet effectively helps to achieve state-of-the-art results. Di Lin 0002, Xin Wang 0118, George Baciu, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Multim. | 6 |
| 2025 | Continuous Bijection Supervised Pyramid Diffeomorphic Deformation for Learning Tooth Meshes From CBCT ImagesabstractAccurate and high-quality tooth mesh generation from cone-beam computerized tomography (CBCT) is an essential computer-aided technology for digital dentistry. However, existing segmentation-based methods require complicated post-processing and significant manual correction to generate regular tooth meshes. In this paper, we propose a method of continuous bijection supervised pyramid diffeomorphic deformation (PDD) for learning tooth meshes, which could be used to directly generate high-quality tooth meshes from CBCT Images. Overall, we adopt a classic two-stage framework. In the first stage, we devise an enhanced detector to accurately locate and crop every tooth. In the second stage, a PDD network is designed to deform a sphere mesh from low resolution to high one according to pyramid flows based on diffeomorphic mesh deformations, so that the generated mesh approximates the ground truth infinitely and efficiently. To achieve that, a novel continuous bijection distance loss on the diffeomorphic sphere is also designed to supervise the deformation learning, which overcomes the shortcoming of loss based on nearest-neighbour mapping and improves the fitting precision. Experiments show that our method outperforms the state-of-the-art methods in terms of both different evaluation metrics and the geometry quality of reconstructed tooth surfaces. Zechu Zhang, Weilong Peng, Jinyu Wen, Keke Tang, Meie Fang, David Dagan Feng, Ping Li 0016 |
IEEE Trans. Multim. | 7 |
| 2025 | 3DCMM: 3D Comprehensive Morphable Models With UV-UNet for Accurate Head CreationabstractIn recent studies of 3D shape modelling and reconstruction, the focus has primarily been on the 3D face region. However, accurately creating the entire 3D head opens up a wide range of applications, including headwear design, cranial diagnosis, and avatar design. Therefore, we present our newly developed method of constructing 3D comprehensive morphable models (3DCMM) specifically tailored for human heads, along with a novel 3DCMM-based stepwise pipeline for creating accurate full 3D heads. Within our 3DCMM framework, we constructed a powerful 3D morphable face model with UV-UNet to generate the 3D face and predict the 3D scalp, resulting in a complete representation of the head. Additionally, our 3DCMM-based self-learning approach incorporates novel facial boundary-aware and structure-aware losses for highly accurate overall reconstructions of the entire facial region. Experimental evaluations demonstrate that our 3DCMM exhibits superior face representation power and achieves higher head prediction accuracy than existing models. Consequently, our 3DCMM-based 3D head creation method from a single image demonstrates outstanding performance capability on both face and head benchmarks. Jie Zhang 0090, Kangneng Zhou, Yan Luximon, Tong-Yee Lee, Ping Li 0016 |
IEEE Trans. Multim. | 5 |
| 2025 | Point-to-Set Metric-Gated Mixture of Experts for Multisource Domain Adaptation Fault DiagnosisabstractThe multisource unsupervised domain adaptation (MUDA) scenario poses a significant challenge in the field of intelligent fault diagnosis (IFD), where the goal is to transfer the knowledge learned from multiple labeled source domains to an unlabeled target domain. Existing IFD-oriented MUDA approaches frequently fail to recognize the distinct importance of each source domain relative to specific target samples, or lack flexibility in integrating diagnostic insights from multiple sources. In response, a novel MUDA approach is proposed for IFD, termed point-to-set metric-gated mixture of experts (PSMMoEs). This method leverages a mixture-of-experts (MoEs) framework to automatically integrate the complementary information from multiple source domains. It develops a deep point-to-set distance (PSD) metric learning technique within the MoE's gating mechanism, effectively fusing domain-specific features by assessing the similarity between individual target samples and each source domain. The method ensures balanced training across progressive stages, harmonizing multitask learning with joint training for the MoE framework. Furthermore, a multilayer maximum mean discrepancy (MMD) measurement is employed for domain alignment, ensuring feature alignment across different domains at multiple levels. In order to assess the efficacy of the proposed method, it is compared with several leading domain adaptation methods on publicly available and laboratory-based rotating machinery fault datasets. The experimental results demonstrate superior classification and adaptation capabilities of the proposed fault diagnosis method. Boyuan Yang 0002, Di Lin 0002, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Temporal-Interim Pose Synthesis and Distillation for Dynamic Human Pose EstimationabstractIn the task of dynamic human pose estimation (dynamic HPE), the temporal relationships between human body parts should be captured comprehensively to understand the dynamic human motions, where the correlated motion information eventually helps to recognize body parts. The popular methods are successful in terms of utilizing long-term motion information captured by low-speed cameras. Yet they neglect the underlying intermediate motions between captured frames, which comprise the temporal-interim poses lost in the video. In this article, we introduce a novel framework, temporal-interim pose synthesis and distillation, to produce and leverage the intermediate motion information for dynamic motion establishment. The pose synthesis yields the visual feature maps of the intermediate poses, which appear between the existing video frames. It allows the synthesized and current poses to form richer motion patterns. Next, the pose distillation divides the body parts into several groups, where it learns the specific part-wise relationship within each group. It degrades the complexity of learning useful part-wise relationships from rich motion patterns and extracts more detailed motion information for fine-grained part groups. We extensively evaluate our method on challenging datasets for dynamic pose estimation, achieving state-of-the-artresults. Di Lin 0002, Xin Wang 0118, Bin Sheng 0001, George Baciu, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | HRC-Net: Learning Visual Hypothesis, Representative, and Collaboration for Multi-Domain Image InpaintingabstractMulti-domain image inpainting utilizes complementary contextual information from auxiliary domain images to restore corrupted regions. While existing methods reconstruct auxiliary images to provide additional guidance, they face fundamental limitations: recovered pixels with complex patterns often lack representative details, while oversimplified patterns offer insufficient contextual information. To address these challenges, we propose HRC-Net, a novel framework incorporating three generative sub-networks for the comprehensive image inpainting task. Our architecture consists of: (1) A Hypothesis Sub-network that enables robust samplings of pixel-wise hypotheses from multi-domain inputs; (2) A Representative Sub-network that learns to score hypothesis quality based on contextual relevance; and (3) a Collaboration Sub-network that optimizes adaptive fusion kernels to integrate the most pertinent details. Together, these components model the joint distribution of representative scores and convolutional kernels, fostering a precise interaction between auxiliary hypotheses and target image corruption to meticulously repair the target image. Extensive evaluations across multiple benchmark datasets demonstrate HRC-Net's superior performance, significantly outperforming state-of-the-art methods in both quantitative metrics and visual quality. Xin Wang 0118, Di Lin 0002, Wanchao Su, Ji Du, Jie Zhang 0090, Haotian Dong, Ke Xu 0010, Qing Guo 0005, Ping Li 0016 |
ACM Trans. Graph. | 10 |
| 2025 | CCM-Net: Contrastive and Consistent Multi-Task Network for Artifact Segmentation and Quality Classification of OCTA ImagesabstractArtifacts are prevalent in Optical Coherence Tomography Angiography (OCTA) images, which probably interfere doctor’s diagnosis and greatly limit its utility. Therefore, it is desirable to segment artifacts and assess quality when using them for diagnosis. In this article, we propose an end-to-end network (named CCM-Net: C ontrastive and C onsistent M ulti-task Network) to jointly address artifact segmentation and quality classification of OCTA images. We first devise multiple Task-Specific Attention Blocks to integrate deep features at different CNN layers for segmenting artifacts and classifying the quality of the input OCTA image. In this way, the weights of different deep features can be automatically learned and are not the same for the two tasks. Moreover, we devise a contrastive loss and a consistency loss to leverage sample relations for further enhancing prediction accuracy. Specifically, given an input OCTA image, we first augment it with a color jitter and select another OCTA image with the same quality classification label. We then design a contrastive loss so that the segmentation results of the input OCTA image are similar to its enhanced OCTA image, while the segmentation results of the two selected OCTA images are not similar. Besides, we devise a consistency loss on the classification results of the three images, because we can find that these images have the same quality classification labels. Experiments on an in-house OCTA dataset (Multi-OCTA) demonstrate that the proposed CCM-Net outperforms state-of-the-art methods. Xiang-Ning Wang, Jixue Tang, Ping Li 0016, Lei Zhu 0003, Harry Qin, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Computer-Aided Colorization State-of-the-Science: A SurveyabstractThis article reviews published research in the field of computer-aided colorization technology. We argue that within this context, the colorization task can be considered to originate from computer graphics, advance by introducing computer vision, and progress towards the fusion of vision and graphics. Hence, we propose a specific taxonomy and organize the research work chronologically. We extend the existing reconstruction-based colorization evaluation techniques on the basis that aesthetic assessment should be introduced to ensure the computer-coloredimages closely satisfy human visual-related requirements. We then perform an aesthetic assessment using the proposed metric and existing evaluations, comparing the colorization performance of seven representative unconditional colorization models. Finally, we identify unresolved issues and propose fruitful areas for future research and development. Yu Cao 0019, Xin Duan, Xiangqiao Meng, P. Y. Mok 0001, Ping Li 0016, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | MSEmbGAN: Multi-Stitch Embroidery Synthesis via Region-Aware Texture GenerationabstractConvolutional neural networks (CNNs) are widely used for embroidery feature synthesis from images. However, they are still unable to predict diverse stitch types, which makes it difficult for the CNNs to effectively extract stitch features. In this paper, we propose a multi-stitch embroidery generative adversarial network (MSEmbGAN) that uses a region-aware texture generation sub-network to predict diverse embroidery features from images. To the best of our knowledge, our work is the first CNN-based generative adversarial network to succeed in this task. Our region-aware texture generation sub-network detects multiple regions in the input image using a stitch classifier and generates a stitch texture for each region based on its shape features. We also propose a colorization network with a color feature extractor, which helps achieve full image color consistency by requiring the color attributes of the output to closely resemble the input image. Because of the current lack of labeled embroidery image datasets, we provide a new multi-stitch embroidery dataset that is annotated with three single-stitch types and one multi-stitch type. Our dataset, which includes more than 30K high-quality multi-stitch embroidery images, more than 13K aligned content-embroidered images, and more than 17K unaligned images, is currently the largest embroidery dataset accessible, as far as we know. Quantitative and qualitative experimental results, including a qualitative user study, show that our MSEmbGAN outperforms current state-of-the-art embroidery synthesis and style-transfer methods on all evaluation indicators. Xinrong Hu, Ping Li 0016, Bin Sheng 0001, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | EGDNet: an efficient glomerular detection network for multiple anomalous pathological feature in glomerulonephritis
Saba Ghazanfar Ali, Ping Li 0016, Huating Li, Po Yang 0001, Younhyun Jung, Harry Qin, Jinman Kim, Bin Sheng 0001 |
Vis. Comput. | 3 |
| 2025 | A novel approach for improving open scene text translation with modified GANabstractAbstract Text, as a vital tool for communication, is playing an imperative role in modern society. Precise high-level text translation systems are essential requirements in a wide range of real-world applications, such as robot navigation, industrial automation, image search, and instant translation. Regardless of improved research, a series of grand challenges may still become upon when translating text automatically in the real-world from open scene images. The difficulties mainly stem from multiplicity and inconsistency of text in open scenes, complication and obstruction of backgrounds, and deficient imaging conditions in uncontrolled circumstances for open scene images. The existing deep learning-based text translation systems do not eliminate the text for translation, and these applications just replace text on the reconstructed scene. To address the abovementioned shortcomings, this study proposed a novel approach for open scene text translation. Our system consists of five modules including scene text detection, text recognition, text elimination, text translation, and text insertion along with scene reconstruction. The novelty presented by our model lies in the idea of first eliminating the text from the open scene for accurate translation and then reconstructs the translated text on the image for its proper alignment. We specifically modified the existing generative adversarial network (GAN) architecture for improved performance of text elimination by introducing a novel strategy of text and scene concatenation to reduce the overall loss function. For this purpose, we created a synthetic dataset to train our GAN for text elimination module. Experiments on various standard text translation systems demonstrate that our integrated system is able to outperform state-of-the-art approaches in terms of result quality. We have achieved 90.87% of precision, 83.66% of recall, 87.116% of F1-score, and reduced both losses ( $$l_1$$ l 1 and $$l_2$$ l 2 ) up to 50% which is remarkable upon state-of-the-art translation systems. Yasmeen Cheema, Muhammad Nadeem Cheema, Anam Nazir, Fahad Ahmed KhoKhar, Ping Li 0016, Ayaz Ahmed |
Vis. Comput. | 5 |
| 2025 | ZAP-2.5DSAM: zero additional parameters advancing 2.5D SAM adaptation to 3D tumor segmentation
Cai Guo, Yuxi Jin, Bishenghui Tao, Hongning Dai, Ping Li 0016 |
Vis. Comput. | 6 |
| 2025 | Graph convolutional networks for 3D skeleton-based scoliosis screening using gait sequencesabstractAbstract Adolescent idiopathic scoliosis is a significant health concern, ranked as the third most prevalent issue among adolescents after obesity and myopia. Traditional screening methods rely on the use of complex and expensive measuring instruments and expert physicians to interpret X-ray images. These methods can be both time-consuming and inaccessible for widespread screening efforts. To address these challenges, we propose a standardized protocol for the collection of scoliosis gait dataset. This protocol enables the systematic capture of relevant gait characteristics associated with scoliosis, leading to the creation of a comprehensive, annotated dataset tailored for research and diagnostic purposes. Leveraging this dataset, we developed an effective deep learning algorithm based on graph convolutional networks, which outperforms traditional CNN by effectively modeling the complex spatial and temporal dynamics of human gait and posture, leveraging skeletal structure as a graph for more accurate and robust scoliosis screening. We also explored various optimization strategies to enhance the model’s accuracy and efficiency, ensuring robust performance across diverse scenarios. Our innovative approach allows for the rapid and non-invasive recognition of scoliosis. This method is not only scalable but also eliminates the need for specialized equipment or extensive medical expertise, making it ideal for large-scale screening initiatives. By improving the accessibility and efficiency of scoliosis detection, our approach has the potential to facilitate early intervention. Zizhao Peng, Mengying Sun, Yan Wang 0116, Ping Li 0016, Fengwei An |
Vis. Comput. | 6 |
| 2025 | HRDC challenge: a public benchmark for hypertension and hypertensive retinopathy classification from fundus images
Xiangning Wang, Zhouyu Guan, An-ran Ran, Tingyao Li, Zheyuan Wang, Xinming Shu, Jinyang Xie, Shichang Liu, Guanyu Xing, Julio Silva-Rodríguez, Riadh Kobbi, Ping Li 0016, Tingli Chen, Lei Bi 0001, Jinman Kim, Weiping Jia, Huating Li, Harry Qin, Ping Zhang 0016, Ching Yu Cheng, Pheng-Ann Heng, Tien Yin Wong, Carol Y. Cheung, Nadia Magnenat-Thalmann, Bin Sheng 0001 |
Vis. Comput. | 15 |
| 2025 | CADGCL: unsupervised retrieval of CAD models via boundary representationsabstractAbstract With the widespread application of CAD technology in the industrial manufacturing sector, the efficient retrieval of target models has become a critical research topic. Despite the outstanding performance of traditional supervised retrieval methods, their reliance on large amounts of labeled data significantly limits practical applications. Data labeling is not only time-consuming and costly but also difficult to ensure accuracy and consistency. To tackle this issue, this paper introduces CADGCL, an unsupervised method for retrieving CAD models based on boundary representations. The proposed method transforms CAD models represented by boundary representations into B-rep attributed graphs that integrate geometric information and topological structures. Graph contrastive learning facilitates unsupervised CAD model retrieval. To overcome the limitations of traditional GCL methods in data augmentation and negative sampling, two novel strategies are introduced: an edge perturbation strategy based on Edge Betweenness Centrality and a negative sampling strategy based on the Beta Mixture Model. These strategies effectively improve the performance of contrastive learning. Experimental results show that the proposed method outperforms existing approaches in mAP and F1 scores under unsupervised scenarios, validating its potential for applications in industrial manufacturing. Fei-wei Qin, Liangzhe Zhu, Zijian Xu 0009, Meie Fang, Ping Li 0016 |
Vis. Comput. | 5 |
| 2025 | Msc-Net: multi-stage colorization network for real-world images with specular highlights
Meng Yang 0011, Weiliang Meng, Ping Li 0016 |
Vis. Comput. | 4 |
| 2025 | ViT-BF: vision transformer with border-aware features for visual tracking
Ping Li 0016, Jinxing Liang, Tao Peng 0006, Jia Chen 0012, Li Li 0094, Xinrong Hu, Junping Liu |
Vis. Comput. | 3 |
| 2025 | Distilling complementary information from temporal context for enhancing human appearance in human-specific NeRFabstractAbstract Reconstructing and animating digital avatars with free views from monocular videos have been an interesting research task in the computer vision field for a long time. Recently, some methods have introduced a novel category method of leveraging the neural radiance field to represent the human body in a canonical space with the help of the SMPL model. With the deformation of the points from an observation space into a canonical space, the human appearance can be learned in various poses and viewpoints. However, previous methods highly rely on pose-dependent representation learned from frame-independent optimization and ignore the temporal contexts across the continuous motion video, causing a bad influence on the dynamic appearance texture generation. To overcome these problems, we propose a novel free-viewpoint rendering framework, TMIHuman. It aims at introducing temporal information into NeRF-based rendering and distilling task-relevant information from complex pixel-wise representations. To be specific, we build a temporal fusion encoder that imports timestamps into the learning of non-rigid deformation and fuses the visual features of other frames into human representation. Then, we propose to disentangle the fused features and extract useful visual cues via mutual information objectives. We have extensively evaluated our method and achieved state-of-the-art performance on different public datasets. Xin Wang 0118, George Baciu, Ping Li 0016 |
Vis. Comput. | 4 |
| 2025 | Few-shot medical image segmentation via query transformation learning
Shihao Zheng, Huisi Wu, Zhijian Gao, Ping Li 0016 |
Vis. Comput. | 4 |
| 2025 | Adversarial relighting attacks: physically interpretable manipulation of incident light for vision model vulnerability exploration
Chengzhi Zhong, Mengda Xie, Yiling He, Ping Li 0016, Meie Fang |
Vis. Comput. | 4 |
| 2025 | Topology-guided accelerated vector field streamline visualization
Junjie Yin, Yilun Yang, Meie Fang, Ping Li 0016 |
Vis. Comput. | 5 |
| 2024 | Two-Stage Video Shadow Detection via Temporal-Spatial Adaption
Xin Duan, Yu Cao 0019, Lei Zhu 0003, Gang Fu 0003, Xin Wang 0118, Ping Li 0016 |
ECCV (48) | 7 |
| 2024 | IE-aware Consistency Losses for Detailed 3D Face Reconstruction from Multiple Images in the Wildabstract3D face reconstruction from multiple in-the-wild images in an unsupervised manner poses a significant challenge, primarily due to the pervasive presence of Intrinsic and Extrinsic inconsistencies in facial features. To tackle this, we introduce a novel set of IE-aware consistency losses designed to effectively mitigate these inconsistencies. Our Local Alignment Loss employs neighborhood search techniques to identify and optimize consistent pixel information, thereby reducing intrinsic inconsistencies. In parallel, our Region Subset Selection Loss filters out regions where significant discrepancies exist between the input and reconstructed images, effectively alleviating extrinsic inconsistencies. Extensive experimental results validate the effectiveness of our IE-aware consistency losses in reconstructing detailed 3D facial geometry from images captured in uncontrolled environments. Weilong Peng, Keke Tang, Kongyang Chen, Yangtao Wang, Ping Li 0016, Meie Fang |
ICME | 6 |
| 2024 | Timeline and Boundary Guided Diffusion Network for Video Shadow DetectionabstractVideo Shadow Detection (VSD) aims to detect the shadow masks with frame sequence. Existing works suffer from inefficient temporal learning. Moreover, few works address the VSD problem by considering the characteristic (i.e., boundary) of shadow. Motivated by this, we propose a Timeline and Boundary Guided Diffusion (TBGDiff) network for VSD where we take account of the past-future temporal guidance and boundary information jointly. In detail, we design a Dual Scale Aggregation (DSA) module for better temporal understanding by rethinking the affinity of the long-term and short-term frames for the clipped video. Next, we introduce Shadow Boundary Aware Attention (SBAA) to utilize the edge contexts for capturing the characteristics of shadows. Moreover, we are the first to introduce the Diffusion model for VSD in which we explore a Space-Time Encoded Embedding (STEE) to inject the temporal guidance for Diffusion to conduct shadow detection. Benefiting from these designs, our model can not only capture the temporal information but also the shadow property. Extensive experiments show that the performance of our approach overtakes the state-of-the-art methods, verifying the effectiveness of our components. We release the codes, weights, and results at \url{https://github.com/haipengzhou856/TBGDiff}. Haipeng Zhou, Hongqiu Wang, Tian Ye 0001, Zhaohu Xing, Jun Ma 0008, Ping Li 0016, Qiong Wang 0001, Lei Zhu 0003 |
ACM Multimedia | 6 |
| 2024 | Bi-directional Interaction and Dense Aggregation Network for RGB-D Salient Object Detection
Kang Yi, Hongyu Bai, Yinjie Wang, Jing Xu 0008, Ping Li 0016 |
MMM (1) | 6 |
| 2024 | Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene CompletionabstractSemantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accurately propagating features from related voxels, the completion likely fails while propagating features in a single pass without considering multiple potential pathways. And they are generally only suitable for static scenes and struggle to handle dynamic aspects. This paper introduces Voxel Proposal Network (VPNet) that completes scenes from 3D and Bird's-Eye-View (BEV) perspectives. It includes Confident Voxel Proposal based on voxel-wise coordinates to propose confident voxels with high reliability for completion. This method reconstructs the scene geometry and implicitly models the uncertainty of voxel-wise semantic labels by presenting multiple possibilities for voxels. VPNet employs Multi-Frame Knowledge Distillation based on the point clouds of multiple adjacent frames to accurately predict the voxel-wise labels by condensing various possibilities of voxel relationships. VPNet has shown superior performance and achieved state-of-the-art results on the SemanticKITTI and SemanticPOSS datasets. Lubo Wang, Di Lin 0002, Kairui Yang, Qing Guo 0005, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Ping Li 0016 |
NeurIPS | 10 |
| 2024 | Make static person walk again via separating pose action from shapeabstractThis paper addresses the problem of animating a person in static images, the core task of which is to infer future poses for the person. Existing approaches predict future poses in the 2D space, suffering from entanglement of pose action and shape. We propose a method that generates actions in the 3D space and then transfers them to the 2D person. We first lift the 2D pose of the person to a 3D skeleton, then propose a 3D action synthesis network predicting future skeletons, and finally devise a self-supervised action transfer network that transfers the actions of 3D skeletons to the 2D person. Actions generated in the 3D space look plausible and vivid. More importantly, self-supervised action transfer allows our method to be trained only on a 3D MoCap dataset while being able to process images in different domains. Experiments on three image datasets validate the effectiveness of our method. Yongwei Nie, Meihua Zhao, Qing Zhang 0006, Ping Li 0016, Jian Zhu 0001, Hongmin Cai |
Graph. Model. | 4 |
| 2024 | Lightweight blueprint residual network for single image super-resolution
Fangwei Hao, Jiesheng Wu, Weiyun Liang, Jing Xu 0008, Ping Li 0016 |
Expert Syst. Appl. | 5 |
| 2024 | GAN-Based Multi-Decomposition Photo CartoonizationabstractAbstract Background Cartoon images play a vital role in film production, scientific and educational animation, video games, and other fields, and are one of the key visual expressions of artistic creation. However, since hand‐crafted cartoon images often require a great deal of time and effort on the part of professional artists, it is necessary to be able to automatically transform real‐world images into different styles of cartoon images. Although cartoon images vary from artist to artist, cartoon images generally have the unique characteristics of being highly simplified and abstract, with clear edges, smooth color shading, and relatively simple textures. However, existing image cartoonization methods tend to create a number of problems when performing style transfer, which mainly include: (1) the resulting generated images do not have obvious cartoon‐style textures; and (2) the generated images are prone to structural confusion, color artifacts, and loss of the original image content. Therefore, it is also a great challenge in the field of image cartoonization to be able to make a good balance between style transfer and content keeping. Methods In this paper, we propose a GAN‐based multi‐attention mechanism for image cartoonization to address the above issues. The method combines the residual block used to extract deep network features in the generator with the attention mechanism, and further strengthens the perceptual ability of the generative model to cartoon images through the adaptive feature correction of the attention module to improve the cartoon features of the generated images. At the same time, we also introduce the attention mechanism in the convolution block of the discriminator, which is used to further reduce the image visual quality problem caused by the style transfer process. By introducing the attention mechanism into the generator and discriminator models of the generative adversarial network, our method enables the generated images to have obvious cartoon‐style features while effectively improving the image's visual quality. Results A large number of quantitative, qualitative, and ablation experiments are conducted to demonstrate the advantages of our method in the field of image cartoonization and the role of each module in the method. Jianlin Zhu, Ping Li 0016, Bin Sheng 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2024 | Learning Motion-Guided Multi-Scale Memory Features for Video Shadow DetectionabstractNatural images often contain multiple shadow regions, and existing video shadow detection methods tend to fail in fully identifying all shadow regions, since they mainly learned temporal features at single-scale and single memory. In this work, we develop a novel convolutional neural network (CNN) to learn motion-guided multi-scale memory features to obtain multi-scale temporal information based on multiple network memories for boosting video shadow detection. To do so, our network first constructs three memories (i.e., a global memory, a local memory, and a motion memory) to combine spatial context and object motion for detecting shadows. Based on these three memories, we then devise a multi-scale motion-guided long-short transformer (MMLT) module to learn multi-scale temporal and motion memory features for predicting a shadow detection map of the input video frame. Our MMLT module includes a dense-scale long transformer (DLT), a dense-scale short transformer (DST), and a dense-scale motion transformer (DMT) to read three memories for learning multi-scale transformer features. Our DLT, DST, and DMT consist of a set of memory-read pooling attention (MPA) blocks and densely connect these output features of multiple MPA blocks to learn multi-scale transformer features since the scales of these output features are varied. By doing so, we can more accurately identify multiple shadow regions with different sizes from the input video. Moreover, we devise a self-supervised pretext task to pre-training the feature encoder for enhancing the downstream video shadow detection. Experimental results on three benchmark datasets show that our video shadow detection network quantitatively and qualitatively outperforms 26 state-of-the-art methods. Jiaxing Shen, Xin Yang 0011, Huazhu Fu, Qing Zhang 0006, Ping Li 0016, Bin Sheng 0001, Liansheng Wang 0002, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | FastAL: Fast Evaluation Module for Efficient Dynamic Deep Active Learning Using Broad Learning SystemabstractState-of-the-art Active Learning (AL) methods often encounter challenges associated with a hysteretic learning process and an expensive data sampling mechanism. The former implies that data selection in the ($i+1$)-th round is solely based on the learned model’s results in the$i$-th round. The latter involves using model inference to calculate data value (e.g., uncertainty estimation based on model inference), which can be cumbersome, particularly when working with large datasets or Deep Neural Networks (DNNs). To address these challenges, we propose FastAL, an efficient and dynamic deep AL framework. Our approach includes an efficient method for calculating data value from the frequency domain perspective, generating multiple candidates. Then, we introduce the Fast Evaluation Module, which directly calculates each candidate’s contribution to future model training and selects the best options. In addition, current AL methods, particularly those based on uncertainty, are susceptible to data bias, which implies that selected data may not represent the original unlabeled data adequately. To alleviate this issue, we propose the De-similar Module, which removes partially similar data. The above three modules are model-agnostic and thus can be seamlessly integrated into any Active Learning framework. We conducted rigorous experiments on various benchmark datasets to validate our approach’s effectiveness. Our results demonstrate that FastAL outperforms other state-of-the-art methods by a significant margin, including those based on uncertainty, diversity, and expected model change. Shuzhou Sun, Huali Xu, Yan Li 0063, Ping Li 0016, Bin Sheng 0001, Xiao Lin 0012 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Eyeglass Reflection Removal With Joint Learning of Reflection Elimination and Content InpaintingabstractEyeglass reflection removal is of great importance to the portrait image processing. However, it remains a challenge to eliminate the reflections on the glass and restore the textual contents of eyes without introducing visual artifacts. Addressing this problem, in this paper, we propose an Eyeglass Reflection Removal Network (ER2Net) by learning reflection elimination and content inpainting jointly. The reflection elimination branch is effective in weak reflection regions, and the content inpainting branch is dedicated to content reasoning in strong reflection regions. We then propose a result fusion module (RFM), which adaptively fuses the elimination result and the inpainting result according to the reflection intensity of each pixel, to produce high-quality result. We also design a memory module for improving the content inpainting result, and propose an eye-symmetry loss to avoid visual artifacts. Additionally, we construct the first Real-world eyeglass Reflection (ReyeR) dataset for eyeglass reflection removal. Extensive quantitative and qualitative experiments demonstrate the superiority of the ER2Net over state-of-the-art methods for eyeglass reflection removal. Wentao Zou, Xiao Lu 0002, Zhilv Yi, Ling Zhang 0017, Gang Fu 0003, Ping Li 0016, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | High-Quality Fusion and Visualization for MR-PET Brain Tumor Images via Multi-Dimensional FeaturesabstractThe fusion of magnetic resonance imaging and positron emission tomography can combine biological anatomical information and physiological metabolic information, which is of great significance for the clinical diagnosis and localization of lesions. In this paper, we propose a novel adaptive linear fusion method for multi-dimensional features of brain magnetic resonance and positron emission tomography images based on a convolutional neural network, termed as MdAFuse. First, in the feature extraction stage, three-dimensional feature extraction modules are constructed to extract coarse, fine, and multi-scale information features from the source image. Second, at the fusion stage, the affine mapping function of multi-dimensional features is established to maintain a constant geometric relationship between the features, which can effectively utilize structural information from a feature map to achieve a better reconstruction effect. Furthermore, our MdAFuse comprises a key feature visualization enhancement algorithm designed to observe the dynamic growth of brain lesions, which can facilitate the early diagnosis and treatment of brain tumors. Extensive experimental results demonstrate that our method is superior to existing fusion methods in terms of visual perception and nine kinds of objective image fusion metrics. Specifically, in the results of MR-PET fusion, the SSIM (Structural Similarity) and VIF (Visual Information Fidelity) metrics show improvements of 5.61% and 13.76%, respectively, compared to the current state-of-the-art algorithm. Our project is publicly available at: https://github.com/22385wjy/MdAFuse. Jinyu Wen, Amei Chen, Weilong Peng, Meie Fang, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Image Process. | 7 |
| 2024 | MsgFusion: Medical Semantic Guided Two-Branch Network for Multimodal Brain Image FusionabstractMultimodal image fusion plays an essential role in medical image analysis and application, where computed tomography (CT), magnetic resonance (MR), single-photon emission computed tomography (SPECT), and positron emission tomography (PET) are commonly-used modalities, especially for brain disease diagnoses. Most existing fusion methods do not consider the characteristics of medical images, and they adopt similar strategies and assessment standards to natural image fusion. While distinctive medical semantic information (MS-Info) is hidden in different modalities, the ultimate clinical assessment of the fusion results is ignored. Our MsgFusion first builds a relationship between the key MS-Info of the MR/CT/PET/SPECT images and image features to guide the CNN feature extractions using two branches and the design of the image fusion framework. For MR images, we combine the spatial domain feature and frequency domain feature (SF) to develop one branch. For PET/SPECT/CT images, we integrate the gray color space feature and adapt the HSV color space feature (GV) to develop another branch. A classification-based hierarchical fusion strategy is also proposed to reconstruct the fusion images to persist and enhance the salient MS-Info reflecting anatomical structure and functional metabolism. Fusion experiments are carried out on many pairs of MR-PET/SPECT and MR-CT images. According to seven classical objective quality assessments and one new subjective clinical quality assessment from 30 clinical doctors, the fusion results of the proposed MsgFusion are superior to those of the existing representative methods. Jinyu Wen, Fei-wei Qin, Jiao Du, Meie Fang, Xinhua Wei, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Multim. | 7 |
| 2024 | Unsupervised Fusion Feature Matching for Data Bias in Uncertainty Active LearningabstractActive learning (AL) aims to sample the most valuable data for model improvement from the unlabeled pool. Traditional works, especially uncertainty-based methods, are prone to suffer from a data bias issue, which means that selected data cannot cover the entire unlabeled pool well. Although there have been lots of literature works focusing on this issue recently, they mainly benefit from the huge additional training costs and the artificially designed complex loss. The latter causes these methods to be redesigned when facing new models or tasks, which is very time-consuming and laborious. This article proposes a feature-matching-based uncertainty that resamples selected uncertainty data by feature matching, thus removing similar data to alleviate the data bias issue. To ensure that our proposed method does not introduce a lot of additional costs, we specially design a unsupervised fusion feature matching (UFFM), which does not require any training in our novel AL framework. Besides, we also redesign several classic uncertainty methods to be applied to more complex visual tasks. We conduct rigorous experiments on lots of standard benchmark datasets to validate our work. The experimental results show that our UFFM is better than the similar unsupervised feature matching technologies, and our proposed uncertainty calculation method outperforms random sampling, classic uncertainty approaches, and recent state-of-the-art (SOTA) uncertainty approaches. Shuzhou Sun, Xiao Lin 0012, Ping Li 0016, Lei Zhu 0003, C. L. Philip Chen, Bin Sheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | SparseVoxNet: 3-D Object Recognition With Sparsely Aggregation of 3-D Dense BlocksabstractAutomatic recognition of 3-D objects in a 3-D model by convolutional neural network (CNN) methods has been successfully applied to various tasks, e.g., robotics and augmented reality. Three-dimensional object recognition is mainly performed by analyzing the object using multi-view images, depth images, graphs, or volumetric data. In some cases, using volumetric data provides the most promising results. However, existing recognition techniques on volumetric data have many drawbacks, such as losing object details on converting points to voxels and the large size of the input volume data that leads to substantial 3-D CNNs. Using point clouds could also provide very promising results; however, point-cloud-based methods typically need sparse data entry and time-consuming training stages. Thus, using volumetric could be a more efficient and flexible recognizer for our special case in the School of Medicine, Shanghai Jiao Tong University. In this article, we propose a novel solution to 3-D object recognition from volumetric data using a combination of three compact CNN models, low-cost SparseNet, and feature representation technique. We achieve an optimized network by estimating extra geometrical information comprising the surface normal and curvature into two separated neural networks. These two models provide supplementary information to each voxel data that consequently improve the results. The primary network model takes advantage of all the predicted features and uses these features in Random Forest (RF) for recognition purposes. Our method outperforms other methods in training speed in our experiments and provides an accurate result as good as the state-of-the-art. Ahmad Karambakhsh, Bin Sheng 0001, Ping Li 0016, Huating Li, Jinman Kim, Younhyun Jung, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | A New Framework of Collaborative Learning for Adaptive Metric DistillationabstractThis article presents a new adaptive metric distillation approach that can significantly improve the student networks' backbone features, along with better classification results. Previous knowledge distillation (KD) methods usually focus on transferring the knowledge across the classifier logits or feature structure, ignoring the excessive sample relations in the feature space. We demonstrated that such a design greatly limits performance, especially for the retrieval task. The proposed collaborative adaptive metric distillation (CAMD) has three main advantages: 1) the optimization focuses on optimizing the relationship between key pairs by introducing the hard mining strategy into the distillation framework; 2) it provides an adaptive metric distillation that can explicitly optimize the student feature embeddings by applying the relation in the teacher embeddings as supervision; and 3) it employs a collaborative scheme for effective knowledge aggregation. Extensive experiments demonstrated that our approach sets a new state-of-the-art in both the classification and retrieval tasks, outperforming other cutting-edge distillers under various settings. Mang Ye, Yan Wang 0116, Sanyuan Zhao, Ping Li 0016, Jianbing Shen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Action-aware Linguistic Skeleton Optimization Network for Non-autoregressive Video CaptioningabstractNon-autoregressive video captioning methods generate visual words in parallel but often overlook semantic correlations among them, especially regarding verbs, leading to lower caption quality. To address this, we integrate action information of highlighted objects to enhance semantic connections among visual words. Our proposed Action-aware Language Skeleton Optimization Network (ALSO-Net) tackles the challenge of extracting action information across frames, improving understanding of complex context-dependent video actions and reducing sentence inconsistencies. ALSO-Net incorporates a linguistic skeleton tag generator to refine semantic correlations and a video action predictor to enhance verb prediction accuracy in video captions. We also address issues of unsatisfactory caption length and quality by jointly optimizing different levels of motion prediction loss. Experimental evaluation on prominent video captioning datasets demonstrates that ALSO-Net outperforms baseline methods by a significant margin and achieves competitive performance compared to state-of-the-art autoregressive methods with smaller model complexity and faster inference time. Shuqin Chen, Xian Zhong, Lei Zhu 0003, Ping Li 0016, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | AnimeDiffusion: Anime Diffusion ColorizationabstractBeing essential in animation creation, colorizing anime line drawings is usually a tedious and time-consuming manual task. Reference-based line drawing colorization provides an intuitive way to automatically colorize target line drawings using reference images. The prevailing approaches are based on generative adversarial networks (GANs), yet these methods still cannot generate high-quality results comparable to manually-colored ones. In this article, a new AnimeDiffusion approach is proposed via hybrid diffusions for the automatic colorization of anime face line drawings. This is the first attempt to utilize the diffusion model for reference-based colorization, which demands a high level of control over the image synthesis process. To do so, a hybrid end-to-end training strategy is designed, including phase 1 for training diffusion model with classifier-free guidance and phase 2 for efficiently updating color tone with a target reference colored image. The model learns denoising and structure-capturing ability in phase 1, and in phase 2, the model learns more accurate color information. Utilizing our hybrid training strategy, the network convergence speed is accelerated, and the colorization performance is improved. Our AnimeDiffusion generates colorization results with semantic correspondence and color consistency. In addition, the model has a certain generalization performance for line drawings of different line styles. To train and evaluate colorization methods, an anime face line drawing colorization benchmark dataset, containing 31,696 training data and 579 testing data, is introduced and shared. Extensive experiments and user studies have demonstrated that our proposed AnimeDiffusion outperforms state-of-the-art GAN-based methods and another diffusion-based model, both quantitatively and qualitatively. Yu Cao 0019, Xiangqiao Meng, P. Y. Mok 0001, Tong-Yee Lee, Xueting Liu 0001, Ping Li 0016 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | MeshWGAN: Mesh-to-Mesh Wasserstein GAN With Multi-Task Gradient Penalty for 3D Facial Geometric Age TransformationabstractAs the metaverse develops rapidly, 3D facial age transformation is attracting increasing attention, which may bring many potential benefits to a wide variety of users, e.g., 3D aging figures creation, 3D facial data augmentation and editing. Compared with 2D methods, 3D face aging is an underexplored problem. To fill this gap, we propose a new mesh-to-mesh Wasserstein generative adversarial network (MeshWGAN) with a multi-task gradient penalty to model a continuous bi-directional 3D facial geometric aging process. To the best of our knowledge, this is the first architecture to achieve 3D facial geometric age transformation via real 3D scans. As previous image-to-image translation methods cannot be directly applied to the 3D facial mesh, which is totally different from 2D images, we built a mesh encoder, decoder, and multi-task discriminator to facilitate mesh-to-mesh transformations. To mitigate the lack of 3D datasets containing children's faces, we collected scans from 765 subjects aged 5-17 in combination with existing 3D face databases, which provided a large training dataset. Experiments have shown that our architecture can predict 3D facial aging geometries with better identity preservation and age closeness compared to 3D trivial baselines. We also demonstrated the advantages of our approach via various 3D face-related graphics applications. Jie Zhang 0090, Kangneng Zhou, Yan Luximon, Tong-Yee Lee, Ping Li 0016 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | AI-enhanced digital technologies for myopia management: advancements, challenges, and future prospects
Saba Ghazanfar Ali, Zhouyu Guan, Tingli Chen, Ping Li 0016, Po Yang 0001, Zainab Ghazanfar, Younhyun Jung, Bin Sheng 0001, Xiangning Wang |
Vis. Comput. | 6 |
| 2024 | Shadow-aware image colorizationabstractAbstract Significant advancements have been made in colorization in recent years, especially with the introduction of deep learning technology. However, challenges remain in accurately colorizing images under certain lighting conditions, such as shadow. Shadows often cause distortions and inaccuracies in object recognition and visual data interpretation, impacting the reliability and effectiveness of colorization techniques. These problems often lead to unsaturated colors in shadowed images and incorrect colorization of shadows as objects. Our research proposes the first shadow-aware image colorization method, addressing two key challenges that previous studies have overlooked: integrating shadow information with general semantic understanding and preserving saturated colors while accurately colorizing shadow areas. To tackle these challenges, we develop a dual-branch shadow-aware colorization network. Additionally, we introduce our shadow-aware block, an innovative mechanism that seamlessly integrates shadow-specific information into the colorization process, distinguishing between shadow and non-shadow areas. This research significantly improves the accuracy and realism of image colorization, particularly in shadow scenarios, thereby enhancing the practical application of colorization in real-world scenarios. Xin Duan, Yu Cao 0019, Xin Wang 0118, Ping Li 0016 |
Vis. Comput. | 5 |
| 2024 | Deep choroid layer segmentation using hybrid features extraction from OCT images
Saleha Masood, Saba Ghazanfar Ali, Xiangning Wang, Afifa Masood, Ping Li 0016, Huating Li, Younhyun Jung, Bin Sheng 0001, Jinman Kim |
Vis. Comput. | 5 |
| 2023 | Reference-Based Line Drawing Colorization Through Diffusion Model
Jiaze He, Ziruo Li, Ping Li 0016, Lei Zhu 0003, Bin Sheng 0001, Subrota K. Mondal |
CGI | 5 |
| 2023 | MagicMirror: A 3-D Real-Time Virtual Try-On System Through Cloth Simulation
Zhanyi Huang, Tangsheng Guo, Ping Li 0016, Bin Sheng 0001 |
CGI | 5 |
| 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene CompletionabstractSemantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the object relationships from the complex scenes. However, the current networks lack the controllable kernels to model the object relationship across multiple views, where appropriate views provide the relevant information for suggesting the existence of the occluded objects. In this paper, we propose Cross-View Synthesis Transformer (CVSformer), which consists of Multi-View Feature Synthesis and Cross-View Transformer for learning cross-view object relationships. In the multi-view feature synthesis, we use a set of 3D convolutional kernels rotated differently to compute the multi-view features for each voxel. In the cross-view transformer, we employ the cross-view fusion to comprehensively learn the cross-view relationships, which form useful information for enhancing the features of individual views. We use the enhanced features to predict the geometric occupancies and semantic labels of all voxels. We evaluate CVSformer on public datasets, where CVS-former yields state-of-the-art results. Our code is available at https://github.com/donghaotian123/CVSformer. Haotian Dong, Enhui Ma, Lubo Wang, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016, Lingyu Liang, Kairui Yang, Di Lin 0002 |
ICCV | 7 |
| 2023 | Towards High-Quality Specular Highlight Removal by Leveraging Large-Scale Synthetic DataabstractThis paper aims to remove specular highlights from a single object-level image. Although previous methods have made some progresses, their performance remains somewhat limited, particularly for real images with complex specular highlights. To this end, we propose a three-stage network to address them. Specifically, given an input image, we first decompose it into the albedo, shading, and specular residue components to estimate a coarse specular-free image. Then, we further refine the coarse result to alleviate its visual artifacts such as color distortion. Finally, we adjust the tone of the refined result to match the tone of the input as closely as possible. In addition, to facilitate network training and quantitative evaluation, we present a large-scale synthetic dataset of object-level images, covering diverse objects and illumination conditions. Extensive experiments illustrate that our network is able to generalize well to unseen real object-level images, and even produce good results for scene-level images with multiple background objects and complex lighting. Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Chunxia Xiao, Ping Li 0016 |
ICCV | 5 |
| 2023 | 3D Statistical Head Modeling for Face/head-Related Product Design: A State-of-the-Art Review
Jie Zhang 0090, Yan Luximon, Parth B. Shah, Ping Li 0016 |
Comput. Aided Des. | 4 |
| 2023 | Capture My Head: A Convenient and Accessible Approach Combining 3D Shape Reconstruction and Size Measurement from 2D Images for Headwear Design
Jie Zhang 0090, Yan Luximon, Jingyi Wan, Ping Li 0016 |
Comput. Aided Des. | 4 |
| 2023 | Sand painting conversion based on detail preservation
Mengting Zhu, Meng Yang 0011, Weiliang Meng, Ping Li 0016 |
Comput. Graph. | 4 |
| 2023 | Stroke-GAN Painter: Learning to paint artworks using stroke-style generative adversarial networksabstractIt is a challenging task to teach machines to paint like human artists in a stroke-by-stroke fashion. Despite advances in stroke-based image rendering and deep learning-based image rendering, existing painting methods have limitations: they (i) lack flexibility to choose different art-style strokes, (ii) lose content details of images, and (iii) generate few artistic styles for paintings. In this paper, we propose a stroke-style generative adversarial network, called Stroke-GAN, to solve the first two limitations. Stroke-GAN learns styles of strokes from different stroke-style datasets, so can produce diverse stroke styles. We design three players in Stroke-GAN to generate pure-color strokes close to human artists’ strokes, thereby improving the quality of painted details. To overcome the third limitation, we have devised a neural network named Stroke-GAN Painter, based on Stroke-GAN; it can generate different artistic styles of paintings. Experiments demonstrate that our artful painter can generate various styles of paintings while well-preserving content details (such as details of human faces and building textures) and retaining high fidelity to the input images. Qian Wang 0079, Cai Guo, Hongning Dai, Ping Li 0016 |
Comput. Vis. Media | 4 |
| 2023 | Tampering localization and self-recovery using block labeling and adaptive significance
Xiaochen Yuan, Tong Liu 0021, Chan-Tong Lam, Guoheng Huang, Di Lin 0002, Ping Li 0016 |
Expert Syst. Appl. | 7 |
| 2023 | AEGAN: Generating imperceptible face synthesis via autoencoder-based generative adversarial networkabstractAbstract Face recognition (FR) systems based on convolutional neural networks have shown excellent performance in human face inference. However, some malicious users may exploit such powerful systems to identify others' face images disclosed by victims' social network accounts, consequently obtaining private information. To address this emerging issue, synthesizing face protection images with visual and protective effects is essential. However, existing face protection methods encounter three critical problems: poor visual effect, limited protective effect, and trade‐off between visual and protective effects. To address these challenges, we propose a novel face protection approach in this article. Specifically, we design a generative adversarial network (GAN) framework with an autoencoder (AEGAN) as the generator to synthesize the protection images. It is worth noting that we introduce an interpolation upsampling module in the decoder in order to let the synthesized protection images evade recognition by powerful convolution‐based FR systems. Furthermore, we introduce an attention module with a perceptual loss in AEGAN to enhance the visual effects of synthesized images by AEGAN. Extensive experiments have shown that AEGAN not only can maintain the comfortable visual quality of synthesized images but also prevent the recognition of commercial FR systems, including Baidu and iKLYTEK. Aolin Che, Cai Guo, Hongning Dai, Haoran Xie 0001, Ping Li 0016 |
Comput. Animat. Virtual Worlds | 6 |
| 2023 | Multi-stage feature-fusion dense network for motion deblurring
Cai Guo, Qian Wang 0079, Hongning Dai, Ping Li 0016 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | MNGNAS: Distilling Adaptive Combination of Multiple Searched Networks for One-Shot Neural Architecture SearchabstractRecently neural architecture (NAS) search has attracted great interest in academia and industry. It remains a challenging problem due to the huge search space and computational costs. Recent studies in NAS mainly focused on the usage of weight sharing to train a SuperNet once. However, the corresponding branch of each subnetwork is not guaranteed to be fully trained. It may not only incur huge computation costs but also affect the architecture ranking in the retraining procedure. We propose a multi-teacher-guided NAS, which proposes to use the adaptive ensemble and perturbation-aware knowledge distillation algorithm in the one-shot-based NAS algorithm. The optimization method aiming to find the optimal descent directions is used to obtain adaptive coefficients for the feature maps of the combined teacher model. Besides, we propose a specific knowledge distillation process for optimal architectures and perturbed ones in each searching process to learn better feature maps for later distillation procedures. Comprehensive experiments verify our approach is flexible and effective. We show improvement in precision and search efficiency in the standard recognition dataset. We also show improvement in correlation between the accuracy of the search algorithm and true accuracy by NAS benchmark datasets. Guhao Qiu, Ping Li 0016, Lei Zhu 0003, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Dual Multiscale Mean Teacher Network for Semi-Supervised Infection Segmentation in Chest CT Volume for COVID-19abstractAutomated detecting lung infections from computed tomography (CT) data plays an important role for combating coronavirus 2019 (COVID-19). However, there are still some challenges for developing AI system: 1) most current COVID-19 infection segmentation methods mainly relied on 2-D CT images, which lack 3-D sequential constraint; 2) existing 3-D CT segmentation methods focus on single-scale representations, which do not achieve the multiple level receptive field sizes on 3-D volume; and 3) the emergent breaking out of COVID-19 makes it hard to annotate sufficient CT volumes for training deep model. To address these issues, we first build a multiple dimensional-attention convolutional neural network (MDA-CNN) to aggregate multiscale information along different dimension of input feature maps and impose supervision on multiple predictions from different convolutional neural networks (CNNs) layers. Second, we assign this MDA-CNN as a basic network into a novel dual multiscale mean teacher network (DM [Formula: see text]-Net) for semi-supervised COVID-19 lung infection segmentation on CT volumes by leveraging unlabeled data and exploring the multiscale information. Our DM [Formula: see text]-Net encourages multiple predictions at different CNN layers from the student and teacher networks to be consistent for computing a multiscale consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from multiple predictions of MDA-CNN. Third, we collect two COVID-19 segmentation datasets to evaluate our method. The experimental results show that our network consistently outperforms the compared state-of-the-art methods. Liansheng Wang 0002, Jiacheng Wang 0002, Lei Zhu 0003, Huazhu Fu, Ping Li 0016, Gary Cheng 0001, Shuo Li 0001, Pheng-Ann Heng |
IEEE Trans. Cybern. | 5 |
| 2023 | PhotoHelper: Portrait Photographing Guidance Via Deep Feature Retrieval and FusionabstractWe introduce a new photographing guidance (PhotoHelper) for amateur photographers to enhance their portrait photo quality using deep feature retrieval and fusion. In our model, we comprehensively integrate empirical aesthetic rules, traditional machine learning algorithms and deep neural networks to extract different kinds of features in both color and space aspects. With these features, we build a modified random forest with a structured photograph collection to identify types of photos. We also define the composition matching score to measure the similarity between the given photo and the reference photo. By combining all of the above processes, a one-stop deep portrait photographing guidance is constructed to provide users with professional reference photographs that are similar to the current scene and automatically generate spatial composition guidance according to the user-selected reference photo. Experiments and evaluations show that the aesthetic quality of portrait photos can be significantly improved via the composition guidance of our photographing guidance approach. Bin Sheng 0001, Ping Li 0016, Tong-Yee Lee |
IEEE Trans. Multim. | 3 |
| 2023 | EAPT: Efficient Attention Pyramid Transformer for Image ProcessingabstractRecent transformer-based models, especially patch-based methods, have shown huge potentiality in vision tasks. However, the split fixed-size patches divide the input features into the same size patches, which ignores the fact that vision elements are often various and thus may destroy the semantic information. Also, the vanilla patch-based transformer cannot guarantee the information communication between patches, which will prevent the extraction of attention information with a global view. To circumvent those problems, we propose an Efficient Attention Pyramid Transformer (EAPT). Specifically, we first propose the Deformable Attention, which learns an offset for each position in patches. Thus, even with split fixed-size patches, our method can still obtain non-fixed attention information that can cover various vision elements. Then, we design the Encode-Decode Communication module (En-DeC module), which can obtain communication information among all patches to get more complete global attention information. Finally, we propose a position encoding specifically for vision transformers, which can be used for patches of any dimension and any length. Extensive experiments on the vision tasks of image classification, object detection, and semantic segmentation demonstrate the effectiveness of our proposed model. Furthermore, we also conduct rigorous ablation studies to evaluate the key components of the proposed structure. Xiao Lin 0012, Shuzhou Sun, Bin Sheng 0001, Ping Li 0016, David Dagan Feng |
IEEE Trans. Multim. | 5 |
| 2023 | S $^3$ Net: Self-Supervised Self-Ensembling Network for Semi-Supervised RGB-D Salient Object DetectionabstractRGB-D salient object detection aims to detect visually distinctive objects or regions from a pair of the RGB image and the depth image. State-of-the-art RGB-D saliency detectors are mainly based on convolutional neural networks but almost suffer from an intrinsic limitation relying on the labeled data, thus degrading detection accuracy in complex cases. In this work, we present a self-supervised self-ensembling network (S$^3$Net) for semi-supervised RGB-D salient object detection by leveraging the unlabeled data and exploring a self-supervised learning mechanism. To be specific, we first build a self-guided convolutional neural network (SG-CNN) as a baseline model by developing a series of three-layer cross-model feature fusion (TCF) modules to leverage complementary information among depth and RGB modalities and formulating an auxiliary task that predicts a self-supervised image rotation angle. After that, to further explore the knowledge from unlabeled data, we assign SG-CNN to a student network and a teacher network, and encourage the saliency predictions and self-supervised rotation predictions from these two networks to be consistent on the unlabeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our network quantitatively and qualitatively outperforms the state-of-the-art methods. Lei Zhu 0003, Xiaoqiang Wang 0007, Ping Li 0016, Xin Yang 0011, Qing Zhang 0006, Weiming Wang 0002, Carola-Bibiane Schönlieb, C. L. Philip Chen |
IEEE Trans. Multim. | 3 |
| 2023 | Coarse-to-Fine: Progressive Knowledge Transfer-Based Multitask Convolutional Neural Network for Intelligent Large-Scale Fault DiagnosisabstractIn modern industry, large-scale fault diagnosis of complex systems is emerging and becoming increasingly important. Most deep learning-based methods perform well on small number of fault diagnosis, but cannot converge to satisfactory results when handling large-scale fault diagnosis because the huge number of fault types will lead to the problems of intra/inter-class distance unbalance and poor local minima in neural networks. To address the above problems, a progressive knowledge transfer-based multitask convolutional neural network (PKT-MCNN) is proposed. First, to construct the coarse-to-fine knowledge structure intelligently, a structure learning algorithm is proposed via clustering fault types in different coarse-grained nodes. Thus, the intra/inter-class distance unbalance problem can be mitigated by spreading similar tasks into different nodes. Then, an MCNN architecture is designed to learn the coarse and fine-grained task simultaneously and extract more general fault information, thereby pushing the algorithm away from poor local minima. Last but not least, a PKT algorithm is proposed, which can not only transfer the coarse-grained knowledge to the fine-grained task and further alleviate the intra/inter-class distance unbalance in feature space, but also regulate different learning stages by adjusting the attention weight to each task progressively. To verify the effectiveness of the proposed method, a dataset of a nuclear power system with 66 fault types was collected and analyzed. The results demonstrate that the proposed method can be a promising tool for large-scale fault diagnosis. Yu Wang 0106, Di Lin 0002, Ping Li 0016, Qinghua Hu, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | BaGFN: Broad Attentive Graph Fusion Network for High-Order Feature InteractionsabstractModeling feature interactions is of crucial significance to high-quality feature engineering on multifiled sparse data. At present, a series of state-of-the-art methods extract cross features in a rather implicit bitwise fashion and lack enough comprehensive and flexible competence of learning sophisticated interactions among different feature fields. In this article, we propose a new broad attentive graph fusion network (BaGFN) to better model high-order feature interactions in a flexible and explicit manner. On the one hand, we design an attentive graph fusion module to strengthen high-order feature representation under graph structure. The graph-based module develops a new bilinear-cross aggregation function to aggregate the graph node information, employs the self-attention mechanism to learn the impact of neighborhood nodes, and updates the high-order representation of features by multihop fusion steps. On the other hand, we further construct a broad attentive cross module to refine high-order feature interactions at a bitwise level. The optimized module designs a new broad attention mechanism to dynamically learn the importance weights of cross features and efficiently conduct the sophisticated high-order feature interactions at the granularity of feature dimensions. The final experimental results demonstrate the effectiveness of our proposed model. Wenling Zhang, Bin Sheng 0001, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | FSAD-Net: Feedback Spatial Attention Dehazing NetworkabstractRecent dehazing networks learn more discriminative high-level features by designing deeper networks or introducing complicated structures, while ignoring inherent feature correlations in intermediate layers. In this article, we establish a novel and effective end-to-end dehazing method, named feedback spatial attention dehazing network (FSAD-Net). FSAD-Net is based on the recurrent structure and consists of four modules: a shallow feature extraction block (SFEB), a feedback block (FB), multiple advanced residual blocks (ARBs), and a reconstruction block (RB). FB is designed to handle feedback connections, and it can improve the dehazing performance by exploiting the dependencies of deep features across stages. ARB implements a novel attention-based estimation on a residual block to adapt to pixels with different distributions. Finally, RB helps restore haze-free images. It can be seen from the experimental results that FSAD-Net almost outperforms the state-of-the-arts in terms of five quantitative metrics. Moreover, the qualitatively comparisons on real-world images also demonstrate the superiority of the proposed FSAD-Net. Considering the efficiency and effectiveness of FSAD-Net, it can be expected to serve as a suitable image dehazing baseline in the future. Yu Zhou 0066, Ping Li 0016, Haitao Song 0001, C. L. Philip Chen, Bin Sheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | SThy-Net: a feature fusion-enhanced dense-branched modules network for small thyroid nodule classification from ultrasound images
Abdulrhman H. Al-Jebrni, Saba Ghazanfar Ali, Huating Li, Xiao Lin 0012, Ping Li 0016, Younhyun Jung, Jinman Kim, David Dagan Feng, Bin Sheng 0001, Lixin Jiang |
Vis. Comput. | 5 |
| 2022 | Input-Specific Robustness Certification for Randomized SmoothingabstractAlthough randomized smoothing has demonstrated high certified robustness and superior scalability to other certified defenses, the high computational overhead of the robustness certification bottlenecks the practical applicability, as it depends heavily on the large sample approximation for estimating the confidence interval. In existing works, the sample size for the confidence interval is universally set and agnostic to the input for prediction. This Input-Agnostic Sampling (IAS) scheme may yield a poor Average Certified Radius (ACR)-runtime trade-off which calls for improvement. In this paper, we propose Input-Specific Sampling (ISS) acceleration to achieve the cost-effectiveness for robustness certification, in an adaptive way of reducing the sampling size based on the input characteristic. Furthermore, our method universally controls the certified radius decline from the ISS sample size reduction. The empirical results on CIFAR-10 and ImageNet show that ISS can speed up the certification by more than three times at a limited cost of 0.05 certified radius. Meanwhile, ISS surpasses IAS on the average certified radius across the extensive hyperparameter settings. Specifically, ISS achieves ACR=0.958 on ImageNet in 250 minutes, compared to ACR=0.917 by IAS under the same condition. We release our code in https://github.com/roy-ch/Input-Specific-Certification. Ruoxin Chen, Jie Li 0002, Junchi Yan, Ping Li 0016, Bin Sheng 0001 |
AAAI | 4 |
| 2022 | SlimFliud-Net: Fast Fluid Simulation Using Admm Pruning
Songyang Yu, Ping Li 0016, Weiguang Li, Enhua Wu, Bin Sheng 0001 |
CGI | 3 |
| 2022 | MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image InpaintingabstractAlthough achieving significant progress, existing deep generative inpainting methods still show low generalization across different scenes. As a result, the generated images usually contain artifacts or the filled pixels differ greatly from the ground truth, making them far from real-world applications. Image-level predictive filtering is a widely used restoration technique by predicting suitable kernels adaptively according to different input scenes. Inspired by this inherent advantage, we explore the possibility of addressing image inpainting as a filtering task. To this end, we first study the advantages and challenges of the image-level predictive filtering for inpainting: the method can preserve local structures and avoid artifacts but fails to fill large missing areas. Then, we propose the semantic filtering by conducting filtering on deep feature level, which fills the missing semantic information but fails to recover the details. To address the issues while adopting the respective advantages, we propose a novel filtering technique, i.e., Multi-level Interactive Siamese Filtering (MISF) containing two branches: kernel prediction branch (KPB) and semantic & image filtering branch (SIFB). These two branches are interactively linked: SIFB provides multi-level features for KPB while KPB predicts dynamic kernels for SIFB. As a result, the final method takes the advantage of effective semantic & image-level filling for high-fidelity inpainting. Moreover, we discuss the relationship between MISF and the naive encoder-decoder-based inpainting, inferring that MISF provides novel dynamic convolutional operations to enhance the high generalization capability across scenes. We validate our method on three challenging datasets, i.e., Dunhuang, Places2, and CelebA. Our method outperforms state-of-the-art baselines on four metrics, i.e.,$L_{1}$, PSNR, SSIM, and LPIPS. Qing Guo 0005, Di Lin 0002, Ping Li 0016, Wei Feng 0005, Song Wang 0002 |
CVPR | 4 |
| 2022 | Generative Status Estimation and Information Decoupling for Image Rain RemovalabstractImage rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we propose SEIDNet equipped with the generative Status Estimation and Information Decoupling for rain removal. In the status estimation, we embed the pixel-wise statuses into the status space, where each status indicates a pixel of the rain or object. The status space allows sampling multiple statuses for a pixel, thus capturing the confusing rain or object. In the information decoupling, we respect the pixel-wise statuses, decoupling the appearance information of rain and object from the pixel. Based on the decoupled information, we construct the kernel space, where multiple kernels are sampled for the pixel to remove the rain and recover the object appearance. We evaluate SEIDNet on the public datasets, achieving state-of-the-art performances of image rain removal. The experimental results also demonstrate the generalization of SEIDNet, which can be easily extended to achieve state-of-the-art performances on other image restoration tasks (e.g., snow, haze, and shadow removal). Di Lin 0002, Xin Wang 0118, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016 |
NeurIPS | 9 |
| 2022 | Customize My Helmet: A Novel Algorithmic Approach Based on 3D Head Prediction
Jie Zhang 0090, Yan Luximon, Parth B. Shah, Kangneng Zhou, Ping Li 0016 |
Comput. Aided Des. | 5 |
| 2022 | Experimental protocol designed to employ Nd: YAG laser surgery for anterior chamber glaucoma detection via UBMabstractAbstract Angle closure glaucoma leads to fluid deposition in eye, and intraocular pressure occurs that damage the optic nerve, causes blindness and vision loss. Anterior chamber (AC) evaluation is imperative for determining the risk of angle‐closure. Previously, techniques were dependent on either Pentacam–Scheimpflug that interprets poor visual information, anterior segment optical coherence tomography is injurious to intercede opaque optical structures. Therefore, in this paper, an experimental protocol is designed for detailed disease analysis based on IBM SPSS statistics via ultrasound biomicroscopy which is superior in evaluating deep structures; first, the affected parameter for AC is analysed, and afterwards the direction that needs laser surgery is explored. Experiments are conducted on large‐scale clinical studies from an affiliated hospital in Shanghai, China. The dataset comprised 600 AC images in five directions of 60 subjects. The mean with standard deviation for anterior open distance is 0.158790.096779 mm, 0.158630.081435 mm, and anterior chamber angle is 18.74908.0315, 18.74108.3889 for left and right eye respectively. It is found that anterior chamber angle in the downside of the AC is wider than the upside. However, this decision is partly based on the narrowest part of the angle to widen the depth of the direction and eliminate pupil block. Saba Ghazanfar Ali, Riaz Ali, Bin Sheng 0001, Huating Li, Po Yang 0001, Ping Li 0016, Younhyun Jung, Ping Lu 0008, Jinman Kim |
IET Image Process. | 7 |
| 2022 | LNNet: Lightweight Nested Network for motion deblurring
Cai Guo, Qian Wang 0079, Hongning Dai, Hao Wang 0003, Ping Li 0016 |
J. Syst. Archit. | 5 |
| 2022 | SCPA-Net: Self-calibrated pyramid aggregation for image dehazingabstractAbstract Dehazing as an important image processing field has developed for many years, there exist many excellent methods for exploring more complex networks to solve this problem. In this paper, instead of designing a complex network structure, we propose a novel dehazing network based on the consideration of enhancing feature aggregation and feature representation abilities of dehazing architecture. Specifically, we propose a self‐calibrated pyramid aggregation network (SCPA‐Net) for image dehazing, which is based on an encoder‐decoder architecture. In the encoder, we build a self‐attention block as a unit to aggregate information from a neighborhood to adapt to its content. In the decoder, we introduce a self‐calibration block to capture long‐range spatial and channel dependencies to produce more discriminative representations. Finally, to learn the scale information, the pyramid upsampling structure is applied to aggregate the multiscale self‐calibrated attentive features. Experimental results show our SCPA‐Net can achieve impressive dehazing performance. Yu Zhou 0066, Ping Li 0016, Bin Sheng 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2022 | VDN: Variant-depth network for motion deblurringabstractAbstract Motion deblurring is a challenging task in vision and graphics. Recent researches aim to deblur by using multiple sub‐networks with multi‐scale or multi‐patch inputs. However, scaling or splitting operations on input images inevitably loses the spatial details of the images. Meanwhile, their models are usually complex and computationally expensive. To address these problems, we propose a novel variant‐depth scheme. In particular, we utilize the multiple variant‐depth sub‐networks with scale‐invariant inputs to combine into a variant‐depth network (VDN). In our design, different levels of sub‐networks accomplish progressive deblurring effects without transforming the inputs, thereby effectively reducing the computational complexity of the model. Extensive experiments have shown that our VDN outperforms the state‐of‐the‐art motion deblurring methods while maintaining a lower computational cost. The source code is publicly available at: https://github.com/CaiGuoHS/VDN . Cai Guo, Qian Wang 0079, Hongning Dai, Ping Li 0016 |
Comput. Animat. Virtual Worlds | 4 |
| 2022 | 3D-guided facial shape clustering and analysis
Jie Zhang 0090, Kangneng Zhou, Yan Luximon, Ping Li 0016, Hassan Iftikhar |
Multim. Tools Appl. | 4 |
| 2022 | Improving Video Temporal Consistency via Broad Learning SystemabstractApplying image-based processing methods to original videos on a framewise level breaks the temporal consistency between consecutive frames. Traditional video temporal consistency methods reconstruct an original frame containing flickers from corresponding nonflickering frames, but the inaccurate correspondence realized by optical flow restricts their practical use. In this article, we propose a temporally broad learning system (TBLS), an approach that enforces temporal consistency between frames. We establish the TBLS as a flat network comprising the input data, consisting of an original frame in an original video, a corresponding frame in the temporally inconsistent video on which the image-based technique was applied, and an output frame of the last original frame, as mapped features in feature nodes. Then, we refine extracted features by enhancing the mapped features as enhancement nodes with randomly generated weights. We then connect all extracted features to the output layer with a target weight vector. With the target weight vector, we can minimize the temporal information loss between consecutive frames and the video fidelity loss in the output videos. Finally, we remove the temporal inconsistency in the processed video and output a temporally consistent video. Besides, we propose an alternative incremental learning algorithm based on the increment of the mapped feature nodes, enhancement nodes, or input data to improve learning accuracy by a broad expansion. We demonstrate the superiority of our proposed TBLS by conducting extensive experiments. Bin Sheng 0001, Ping Li 0016, Riaz Ali, C. L. Philip Chen |
IEEE Trans. Cybern. | 2 |
| 2022 | Automatic Detection and Classification System of Domestic Waste via Multimodel Cascaded Convolutional Neural NetworkabstractDomestic waste classification was incorporated into legal provisions recently in China. However, relying on manpower to detect and classify domestic waste is highly inefficient. To that end, in this article, we propose a multimodel cascaded convolutional neural network (MCCNN) for domestic waste image detection and classification. MCCNN combined three subnetworks (DSSD, YOLOv4, and Faster-RCNN) to obtain the detections. Moreover, to suppress the false-positive predicts, we utilized a classification model cascaded with the detection part to judge whether the detection results are correct. To train and evaluate MCCNN, we designed a large-scale waste image dataset (LSWID), containing 30 000 domestic waste multilabeled images with 52 categories. To the best of our knowledge, the LSWID is the largest dataset on domestic waste images. Furthermore, a smart trash can is designed and applied to a Shanghai community, which helped to make waste recycling more efficient. Experimental results showed a state-of-the-art performance, with an average improvement of 10% in detection precision. Jiajia Li 0004, Jie Chen 0097, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, David Dagan Feng, Jun Qi 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | ECSU-Net: An Embedded Clustering Sliced U-Net Coupled With Fusing Strategy for Efficient Intervertebral Disc Segmentation and ClassificationabstractAutomatic vertebra segmentation from computed tomography (CT) image is the very first and a decisive stage in vertebra analysis for computer-based spinal diagnosis and therapy support system. However, automatic segmentation of vertebra remains challenging due to several reasons, including anatomic complexity of spine, unclear boundaries of the vertebrae associated with spongy and soft bones. Based on 2D U-Net, we have proposed an Embedded Clustering Sliced U-Net (ECSU-Net). ECSU-Net comprises of three modules named segmentation, intervertebral disc extraction (IDE) and fusion. The segmentation module follows an instance embedding clustering approach, where our three sliced sub-nets use axis of CT images to generate a coarse 2D segmentation along with embedding space with the same size of the input slices. Our IDE module is designed to classify vertebra and find the inter-space between two slices of segmented spine. Our fusion module takes the coarse segmentation (2D) and outputs the refined 3D results of vertebra. A novel adaptive discriminative loss (ADL) function is introduced to train the embedding space for clustering. In the fusion strategy, three modules are integrated via a learnable weight control component, which adaptively sets their contribution. We have evaluated classical and deep learning methods on Spineweb dataset-2. ECSU-Net has provided comparable performance to previous neural network based algorithms achieving the best segmentation dice score of 95.60% and classification accuracy of 96.20%, while taking less time and computation resources. Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Ping Li 0016, Huating Li, Guangtao Xue, Harry Qin, Jinman Kim, David Dagan Feng |
IEEE Trans. Image Process. | 4 |
| 2022 | Boosting RGB-D Saliency Detection by Leveraging Unlabeled RGB ImagesabstractTraining deep models for RGB-D salient object detection (SOD) often requires a large number of labeled RGB-D images. However, RGB-D data is not easily acquired, which limits the development of RGB-D SOD techniques. To alleviate this issue, we present a Dual-Semi RGB-D Salient Object Detection Network (DS-Net) to leverage unlabeled RGB images for boosting RGB-D saliency detection. We first devise a depth decoupling convolutional neural network (DDCNN), which contains a depth estimation branch and a saliency detection branch. The depth estimation branch is trained with RGB-D images and then used to estimate the pseudo depth maps for all unlabeled RGB images to form the paired data. The saliency detection branch is used to fuse the RGB feature and depth feature to predict the RGB-D saliency. Then, the whole DDCNN is assigned as the backbone in a teacher-student framework for semi-supervised learning. Moreover, we also introduce a consistency loss on the intermediate attention and saliency maps for the unlabeled data, as well as a supervised depth and saliency loss for labeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our DDCNN outperforms state-of-the-art methods both quantitatively and qualitatively. We also demonstrate that our semi-supervised DS-Net can further improve the performance, even when using an RGB image with the pseudo depth map. Xiaoqiang Wang 0007, Lei Zhu 0003, Siliang Tang, Huazhu Fu, Ping Li 0016, Fei Wu 0001, Yi Yang 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 5 |
| 2022 | Face Sketch Synthesis Using Regularized Broad Learning SystemabstractThere are two main categories of face sketch synthesis: data- and model-driven. The data-driven method synthesizes sketches from training photograph-sketch patches at the cost of detail loss. The model-driven method can preserve more details, but the mapping from photographs to sketches is a time-consuming training process, especially when the deep structures require to be refined. We propose a face sketch synthesis method via regularized broad learning system (RBLS). The broad learning-based system directly transforms photographs into sketches with rich details preserved. Also, the incremental learning scheme of broad learning system (BLS) ensures that our method easily increases feature mappings and remodels the network without retraining when the extracted feature mapping nodes are not sufficient. Besides, a Bayesian estimation-based regularization is introduced with the BLS to aid further feature selection and improve the generalization ability and robustness. Various experiments on the CUHK student data set and Aleix Robert (AR) data set demonstrated the effectiveness and efficiency of our RBLS method. Unlike existing methods, our method synthesizes high-quality face sketches much efficiently and greatly reduces computational complexity both in the training and test processes. Ping Li 0016, Bin Sheng 0001, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Deep LSAC for Fine-Grained RecognitionabstractFine-grained recognition emphasizes the identification of subtle differences among object categories given objects that appear in different shapes and poses. These variances should be reduced for reliable recognition. We propose a fine-grained recognition system that incorporates localization, segmentation, alignment, and classification in a unified deep neural network. The input to the classification module includes functions that enable backward-propagation (BP) in constructing the solver. Our major contribution is to propose a valve linkage function (VLF) for BP chaining and form our deep localization, segmentation, alignment, and classification (LSAC) system. The VLF can adaptively compromise errors of classification and alignment when training the LSAC model. It in turn helps to update the localization and segmentation. We evaluate our framework on two widely used fine-grained object data sets. The performance confirms the effectiveness of our LSAC system. Di Lin 0002, Yi Wang 0031, Lingyu Liang, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Real-time spatial normalization for dynamic gesture classification
Sofiane Zeghoud, Saba Ghazanfar Ali, Egemen Ertugrul, Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Xiaoyu Chi, Jinman Kim, Lijuan Mao |
Vis. Comput. | 6 |
| 2022 | DSD-MatchingNet: Deformable Sparse-to-Dense Feature Matching for Learning Accurate CorrespondencesabstractExploring the correspondences across multi-view images is the basis of many computer vision tasks. However, most existing methods are limited on accuracy under challenging conditions. In order to learn more robust and accurate correspondences, we propose the DSD-MatchingNet for local feature matching in this paper. First, we develop a deformable feature extraction module to obtain multi-level feature maps, which harvests contextual information from dynamic receptive fields. The dynamic receptive fields provided by deformable convolution network ensures our method to obtain dense and robust correspondences. Second, we utilize the sparse-to-dense matching with the symmetry of correspondence to implement accurate pixel-level matching, which enables our method to produce more accurate correspondences. Experiments have shown that our proposed DSD-MatchingNet achieves a better performance on image matching benchmark, as well as on visual localization benchmark. Specifically, our method achieves 91.3% mean matching accuracy on HPatches dataset and 99.3% visual localization recalls on Aachen Day-Night dataset. Yicheng Zhao, Han Zhang 0053, Ping Lu 0008, Ping Li 0016, Enhua Wu, Bin Sheng 0001 |
Virtual Real. Intell. Hardw. | 4 |
| 2021 | Dynamic Shadow Synthesis Using Silhouette Edge Optimization
Saba Ghazanfar Ali, Bin Sheng 0001, Ping Li 0016, Xiaoyu Chi, Jinman Kim, Lijuan Mao |
CGI | 5 |
| 2021 | Progressive Multi-scale Reconstruction for Guided Depth Map Super-Resolution via Deep Residual Gate Fusion Network
Bin Sheng 0001, Ping Li 0016, Xiaoyu Chi, Lijuan Mao |
CGI | 5 |
| 2021 | Multi-Stream Fusion Network for Multi-Distortion Image Super-Resolution
Yupeng Xu, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim, Xiangui He |
CGI | 4 |
| 2021 | A Multi-Task Network for Joint Specular Highlight Detection and RemovalabstractSpecular highlight detection and removal are fundamental and challenging tasks. Although recent methods have achieved promising results on the two tasks by training on synthetic training data in a supervised manner, they are typically solely designed for highlight detection or removal, and their performance usually deteriorates significantly on real-world images. In this paper, we present a novel network that aims to detect and remove highlights from natural images. To remove the domain gap between synthetic training samples and real test images, and support the investigation of learning-based approaches, we first introduce a dataset with about 16K real images, each of which has the corresponding ground truths of highlight detection and removal. Using the presented dataset, we develop a multi-task network for joint highlight detection and removal, based on a new specular highlight image formation model. Experiments on the benchmark datasets and our new dataset show that our approach clearly outperforms state-of-the-art methods for both highlight detection and removal. Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Ping Li 0016, Chunxia Xiao |
CVPR | 4 |
| 2021 | DCNet: Dual-Task Cycle Network for End-to-End Image DehazingabstractSingle image dehazing is an important technology in the field of computer vision. In this paper, we propose an image dehazing via dual learning strategy, named dual-task cycle network (DCNet). The core of DCNet is a dual learning framework, which consists of two tasks: the dehazing task and the haze generation task. The dehazing task completes the image dehazing, while the haze generation task achieves the restoration from the dehazed image to the haze image and can form a cycle to provide additional supervision. Our method uses the duality between each task as a constraint to learn and train two tasks jointly, so that the effects of the dehazing model can be improved. Since the haze generation process does not depend on clear images, the DCNet can satisfy the requirements for limited supervision. Extensive experiments demonstrate that our DCNet performs favorably on haze removal. Yu Zhou 0066, Ping Li 0016, Xiaoyu Chi, Lei Ma 0008, Bin Sheng 0001 |
ICME | 3 |
| 2021 | Deep boundary-aware semantic image segmentationabstractAbstract While extensive research efforts have been made in semantic image segmentation, the state‐of‐the‐art methods still suffer from blurry boundaries and mismatched objects due to the insufficient multiscale adaptability. In this paper, we propose a two‐branch convolutional neural network (CNN) approach to capture the multiscale context and the boundary information with the two branches, respectively. To capture the multiscale context, we propose to embed self‐attention mechanism to the atrous spatial pyramid pooling network. To capture the boundary information, we propose to fuse the low‐level features in boundary feature extraction for refining the extracted boundaries via a feature fusion layer (FFL). With FFL, our method can improve the segmentation result with clearer boundaries. A new loss function is proposed which contains a segmentation loss and a boundary loss. Experiments show that our method can predict the boundaries of objects more clearly and have better performance for small‐scale objects. Huisi Wu, Xueting Liu 0001, Ping Li 0016 |
Comput. Animat. Virtual Worlds | 5 |
| 2021 | AFF-Dehazing: Attention-based feature fusion network for low-light image DehazingabstractAbstract Images captured in haze conditions, especially at nighttime with low light, often suffer from degraded visibility, contrasts, and vividness, which makes it difficult to carry out the following vision tasks. In this article, we propose an attention‐based feature fusion network (AFF‐Dehazing) for low‐light image dehazing. Our method decomposes the low‐light image dehazing into two task‐independent streams containing four modules: image dehazing module, low‐light feature extractor module, feature fusion module, and image restoration module. The basic block of these modules is the proposed attention‐based residual dense block. Since the dual‐branch are used, AFF‐Dehazing can avoid learning the mixed degradation all‐in‐one and enhance the details of low‐light haze images. Extensive experiments show that our method surpasses previous state‐of‐the‐art image dehazing methods and low‐light enhancement methods by a very large margin both quantitatively and qualitatively. Yu Zhou 0066, Bin Sheng 0001, Ping Li 0016, Jinman Kim, Enhua Wu |
Comput. Animat. Virtual Worlds | 4 |
| 2021 | Multiview High Dynamic Range Image Synthesis Using Fuzzy Broad Learning SystemabstractCompared with the normal low dynamic range (LDR) images, the high dynamic range (HDR) images provide more dynamic range and image details. Although the existing techniques for generating the HDR images have a good effect for static scenes, they usually produce artifacts on the HDR images for dynamic scenes. In recent years, some learning-based approaches are used to synthesize the HDR images and obtain good results. However, there are also many problems, including the deficiency of explaining and the time-consuming training process. In this article, we propose a novel approach to synthesize multiview HDR images through fuzzy broad learning system (FBLS). We use a set of multiview LDR images with different exposure as input and transfer corresponding Takagi-Sugeno (TS) fuzzy subsystems; then, the structure is expanded in a wide sense in the "enhancement groups" which transfer from the TS fuzzy rules with nonlinear transformation. After integrating fuzzy subsystems and enhancement groups with the trained-well weight, the HDR image is generated. In FBLS, applying the incremental learning algorithm and the pseudoinverse method to compute the weights can greatly reduce the training time. In addition, the fuzzy system has better interpretability. In the learning process, IF-THEN fuzzy rules can effectively help the model to detect the artifacts and reject them in the final HDR result. These advantages solve the problem of existing deep-learning methods. Furthermore, we set up a new dataset of multiview LDR images with corresponding HDR ground truth to train our system. Our experimental results show that our system can synthesize high-quality multiview HDR images, which has a higher training speed than other learning methods. Bin Sheng 0001, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Cybern. | 3 |
| 2021 | GreenSea: Visual Soccer Analysis Using Broad Learning SystemabstractModern soccer increasingly places trust in visual analysis and statistics rather than only relying on the human experience. However, soccer is an extraordinarily complex game that no widely accepted quantitative analysis methods exist. The statistics collection and visualization are time consuming which result in numerous adjustments. To tackle this issue, we developed GreenSea, a visual-based assessment system designed for soccer game analysis, tactics, and training. The system uses a broad learning system (BLS) to train the model in order to avoid the time-consuming issue that traditional deep learning may suffer. Users are able to apply multiple views of a soccer game, and visual summarization of essential statistics using advanced visualization and animation that are available. A marking system trained by BLS is designed to perform quantitative analysis. A novel recurrent discriminative BLS (RDBLS) is proposed to carry out long-term tracking. In our RDBLS, the structure is adjusted to have better performance on the binary classification problem of the discriminative model. Several experiments are carried out to verify that our proposed RDBLS model can outperform the standard BLS and other methods. Two studies were conducted to verify the effectiveness of our GreenSea. The first study was on how GreenSea assists a youth training coach to assess each trainee's performance for selecting most potential players. The second study was on how GreenSea was used to help the U20 Shanghai soccer team coaching staff analyze games and make tactics during the 13th National Games. Our studies have shown the usability of GreenSea and the values of our system to both amateur and expert users. Bin Sheng 0001, Ping Li 0016, Lijuan Mao, C. L. Philip Chen |
IEEE Trans. Cybern. | 2 |
| 2021 | Automatic Symmetry Detection From Brain MRI Based on a 2-Channel Convolutional Neural NetworkabstractSymmetry detection is a method to extract the ideal mid-sagittal plane (MSP) from brain magnetic resonance (MR) images, which can significantly improve the diagnostic accuracy of brain diseases. In this article, we propose an automatic symmetry detection method for brain MR images in 2-D slices based on a 2-channel convolutional neural network (CNN). Different from the existing detection methods that mainly rely on the local image features (gradient, edge, etc.) to determine the MSP, we use a CNN-based model to implement the brain symmetry detection, which does not require any local feature detections and feature matchings. By training to learn a wide variety of benchmarks in the brain images, we can further use a 2-channel CNN to evaluate the similarity between the pairs of brain patches, which are randomly extracted from the whole brain slice based on a Poisson sampling. Finally, a scoring and ranking scheme is used to identify the optimal symmetry axis for each input brain MR slice. Our method was evaluated in 2166 artificial synthesized brain images and 3064 collected in vivo MR images, which included both healthy and pathological cases. The experimental results display that our method achieves excellent performance for symmetry detection. Comparisons with the state-of-the-art methods also demonstrate the effectiveness and advantages for our approach in achieving higher accuracy than the previous competitors. Huisi Wu, Xiujuan Chen, Ping Li 0016, Zhenkun Wen |
IEEE Trans. Cybern. | 3 |
| 2021 | Optic Disk and Cup Segmentation Through Fuzzy Broad Learning System for Glaucoma ScreeningabstractGlaucoma is an ocular disease that causes permanent blindness if not cured at an early stage. Cup-to-disk ratio (CDR), obtained by dividing the height of optic cup (OC) with the height of optic disk (OD), is a widely adopted metric used for glaucoma screening. Therefore, accurately segmenting OD and OC is crucial for calculating a CDR. Most methods have employed deep learning methods for the segmentation of OD and OC. However, these methods are very time consuming. In this article, we present a new fuzzy broad learning system-based technique for OD and OC segmentation with glaucoma screening. We comprehensively integrated extracting a region of interest from RGB images, data augmentation, extracting red and green channel images, and inputting them to the two separate fuzzy broad learning system-based neural networks for segmenting the OD and OC, respectively, and then calculated CDR. Experiments show that our fuzzy broad learning system-based technique outperforms many state-of-the-art methods. Riaz Ali, Bin Sheng 0001, Ping Li 0016, Huating Li, Po Yang 0001, Younhyun Jung, Jinman Kim, C. L. Philip Chen |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Modified GAN-CAED to Minimize Risk of Unintentional Liver Major Vessels Cutting by Controlled Segmentation Using CTA/SPET-CTabstractThis article substantially advances upon state-of-the-art to enhance liver vessels segmentation accuracy by leveraging advantages of synthetic PET-CT (SPET-CT) images in addition to computed tomography angiography (CTA) volumes. Our setup makes a hybrid solution of modified generative adversarial network-convolutional autoencoder (GAN-cAED) combining synthetic ability of GAN to deliver SPET-CT images with generative ability of cAED network in terms of latent learning to more refined segmentation of major liver vessels. We improve time complexity through a novel concept of controlled segmentation by introducing a threshold metric to stop segmentation up to a desired level. The innovative concept of controlled vessel segmentation with a stopping criterion via variant threshold levels will help surgeons to avoid unintentional major blood vessels cutting, reducing the risk of excessive blood loss. Clinically, such solutions offer computer-aided liver surgeries and drug treatment evaluation in a CTA-only environment, shorten the requirement of radioactive and expensive fused PET-CT images. Muhammad Nadeem Cheema, Anam Nazir, Po Yang 0001, Bin Sheng 0001, Ping Li 0016, Huating Li, Xiaoer Wei, Harry Qin, Jinman Kim, David Dagan Feng |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | Globally and Locally Semantic Colorization via Exemplar-Based Broad-GANabstractGiven a target grayscale image and a reference color image, exemplar-based image colorization aims to generate a visually natural-looking color image by transforming meaningful color information from the reference image to the target image. It remains a challenging problem due to the differences in semantic content between the target image and the reference image. In this paper, we present a novel globally and locally semantic colorization method called exemplar-based conditional broad-GAN, a broad generative adversarial network (GAN) framework, to deal with this limitation. Our colorization framework is composed of two sub-networks: the match sub-net and the colorization sub-net. We reconstruct the target image with a dictionary-based sparse representation in the match sub-net, where the dictionary consists of features extracted from the reference image. To enforce global-semantic and local-structure self-similarity constraints, global-local affinity energy is explored to constrain the sparse representation for matching consistency. Then, the matching information of the match sub-net is fed into the colorization sub-net as the perceptual information of the conditional broad-GAN to facilitate the personalized results. Finally, inspired by the observation that a broad learning system is able to extract semantic features efficiently, we further introduce a broad learning system into the conditional GAN and propose a novel loss, which substantially improves the training stability and the semantic similarity between the target image and the ground truth. Extensive experiments have shown that our colorization approach outperforms the state-of-the-art methods, both perceptually and semantically. Haoxuan Li 0004, Bin Sheng 0001, Ping Li 0016, Riaz Ali, C. L. Philip Chen |
IEEE Trans. Image Process. | 3 |
| 2021 | Structure-Aware Motion Deblurring Using Multi-Adversarial Optimized CycleGANabstractRecently, Convolutional Neural Networks (CNNs) have achieved great improvements in blind image motion deblurring. However, most existing image deblurring methods require a large amount of paired training data and fail to maintain satisfactory structural information, which greatly limits their application scope. In this paper, we present an unsupervised image deblurring method based on a multi-adversarial optimized cycle-consistent generative adversarial network (CycleGAN). Although original CycleGAN can handle unpaired training data well, the generated high-resolution images are probable to lose content and structure information. To solve this problem, we utilize a multi-adversarial mechanism based on CycleGAN for blind motion deblurring to generate high-resolution images iteratively. In this multi-adversarial manner, the hidden layers of the generator are gradually supervised, and the implicit refinement is carried out to generate high-resolution images continuously. Meanwhile, we also introduce the structure-aware mechanism to enhance the structure and detail retention ability of the multi-adversarial network for deblurring by taking the edge map as guidance information and adding multi-scale edge constraint functions. Our approach not only avoids the strict need for paired training data and the errors caused by blur kernel estimation, but also maintains the structural information better with multi-adversarial learning and structure-aware mechanism. Comprehensive experiments on several benchmarks have shown that our approach prevails the state-of-the-art methods for blind image motion deblurring. Jie Chen 0097, Bin Sheng 0001, Ping Li 0016, Ping Tan 0002, Tong-Yee Lee |
IEEE Trans. Image Process. | 5 |
| 2021 | NHBS-Net: A Feature Fusion Attention Network for Ultrasound Neonatal Hip Bone SegmentationabstractUltrasound is a widely used technology for diagnosing developmental dysplasia of the hip (DDH) because it does not use radiation. Due to its low cost and convenience, 2-D ultrasound is still the most common examination in DDH diagnosis. In clinical usage, the complexity of both ultrasound image standardization and measurement leads to a high error rate for sonographers. The automatic segmentation results of key structures in the hip joint can be used to develop a standard plane detection method that helps sonographers decrease the error rate. However, current automatic segmentation methods still face challenges in robustness and accuracy. Thus, we propose a neonatal hip bone segmentation network (NHBS-Net) for the first time for the segmentation of seven key structures. We design three improvements, an enhanced dual attention module, a two-class feature fusion module, and a coordinate convolution output head, to help segment different structures. Compared with current state-of-the-art networks, NHBS-Net gains outstanding performance accuracy and generalizability, as shown in the experiments. Additionally, image standardization is a common need in ultrasonography. The ability of segmentation-based standard plane detection is tested on a 50-image standard dataset. The experiments show that our method can help healthcare workers decrease their error rate from 6%-10% to 2%. In addition, the segmentation performance in another ultrasound dataset (fetal heart) demonstrates the ability of our network. Ruhan Liu, Mengyao Liu 0004, Bin Sheng 0001, Huating Li, Ping Li 0016, Haitao Song 0001, Ping Zhang 0016, Lixin Jiang, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Hybrid Refinement-Correction Heatmaps for Human Pose EstimationabstractIn this paper, we present a method (Hybrid-Pose) to improve human pose estimation in images. We adopt Stacked Hourglass Networks to design two convolutional neural network models, RNet for pose refinement and CNet for pose correction. The CNet (Correction Network) guides the pose refinement RNet (Refinement Network) to correct the joint location before generating the final pose. Each of the two models is composed of four hourglasses, and each hourglass generates a group of detection heatmaps for the joints. The RNet model hourglasses have the same structure. However, the CNet model is designed with hourglasses of different structures for pose guidance. Since the pose estimation in RGB images is very sensitive to the image scene, our proposed approach generates multiple outputs of detection heatmaps to broaden the searching scope for the correct joints locations. We use the RNet model to refine the joints locations in each hourglass stage horizontally, then the heatmaps of each stage are fused with the heatmaps of all the CNet model hourglasses vertically in a hybrid manner. Our method shows competitive results with the existing state-of-the-art approaches on MPII and FLIC benchmark datasets. Although our proposed method focuses on improving single-person pose estimation, we also show the influence of this improvement on multi-person pose estimation by detecting multiple people using SSD detector, then estimating the pose of each person individually. Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng |
IEEE Trans. Multim. | 3 |
| 2021 | Deep Texture Exemplar Extraction Based on Trimmed T-CNNabstractTexture exemplar has been widely used in synthesizing 3D movie scenes and appearances of virtual objects. Unfortunately, conventional texture synthesis methods usually only emphasized on generating optimal target textures with arbitrary sizes or diverse effects, and put little attention to automatic texture exemplar extraction. Obtaining texture exemplars is still a labor intensive task, which usually requires carefully cropping and post-processing. In this paper, we present an automatic texture exemplar extraction based on Trimmed Texture Convolutional Neural Network (Trimmed T-CNN). Specifically, our Trimmed T-CNN is filter banks for texture exemplar classification and recognition. Our Trimmed T-CNN is learned with a standard ideal exemplar dataset containing thousands of desired texture exemplars, which were collected and cropped by our invited artists. To efficiently identify the exemplar candidates from an input image, we employ a selective search algorithm to extract the potential texture exemplar patches. We then put all candidates into our Trimmed T-CNN for learning ideal texture exemplars based on our filter banks. Finally, optimal texture exemplars are identified with a scoring and ranking scheme. Our method is evaluated with various kinds of textures and user studies. Comparisons with different feature-based methods and different deep CNN architectures (AlexNet, VGG-M, Deep-TEN and FV-CNN) are also conducted to demonstrate its effectiveness. Huisi Wu, Wei Yan 0036, Ping Li 0016, Zhenkun Wen |
IEEE Trans. Multim. | 3 |
| 2021 | Broad ColorizationabstractThe scribble- and example-based colorization methods have fastidious requirements for users, and the training process of deep neural networks for colorization is quite time-consuming. We instead proposed an automatic colorization approach with no dependence on user input and no need to endure long training time, which combines local features and global features of the input gray-scale images. Low-, mid-, and high-level features are united as local features representing cues existed in the gray-scale image. The global feature is regarded as data prior to guiding the colorization process. The local broad learning system is trained for getting the chrominance value of each pixel from the local features, which could be expressed as a chrominance map according to the position of pixels. Then, the global broad learning system is trained to refine the chrominance map. There are no requirements for users in our approach, and the training time of our framework is an order of magnitude faster than the traditional methods based on deep neural networks. To increase the user's subjective initiative, our system allows users to increase training data without retraining the system. Substantial experimental results have shown that our approach outperforms state-of-the-art methods. Yuxi Jin, Bin Sheng 0001, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Efficient Body Motion Quantification and Similarity Evaluation Using 3-D Joints Skeleton CoordinatesabstractEvaluating whole-body motion is challenging because of the articulated nature of the skeleton structure. Each joint moves in an unpredictable way with uncountable possibilities of movements direction under the influence of one or many of its parent joints. This paper presents a method for human motion quantification via three-dimensional (3-D) body joints coordinates. We calculate a set of metrics that influence the joints movement considering the motion of its parent joints without requiring prior knowledge of the motion parameters. Only the raw joints coordinates data of a motion sequence are needed to automatically estimate the transformation matrix of the joints between frames. We also consider the angles between limbs as a fundamental factor to follow the joints directions. We classify the joints motion as global motion and local motion. The global motion represents the joint movement according to a fixed joint, and the local motion represents the joint movement according to its first parent joint. In order to evaluate the performance of the proposed method, we also propose a comparison algorithm between two skeletons motions based on the quantified metrics. We measured the comparative similarity between the 3-D joints coordinates on Microsoft Kinect V2 and UTD-MHAD dataset. User studies were conducted to evaluate the performance under different factors. Various results and comparisons have shown that our method effectively quantifies and evaluates the motion similarity. Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | GPSD: generative parking spot detection using multi-clue recovery model
Bin Sheng 0001, Ping Li 0016, Enhua Wu |
Vis. Comput. | 4 |
| 2020 | GARNet: Graph Attention Residual Networks Based on Adversarial Learning for 3D Human Pose Estimation
Bin Sheng 0001, Ping Li 0016 |
CGI | 4 |
| 2020 | Broad-Classifier for Remote Sensing Scene Classification with Spatial and Channel-Wise Attention
Yunna Liu, Han Zhang 0053, Bin Sheng 0001, Ping Li 0016, Guangtao Xue |
CGI | 5 |
| 2020 | FIOU Tracker: An Improved Algorithm of IOU Tracker in Video with a Lot of Background Inferences
Guhao Qiu, Han Zhang 0053, Bin Sheng 0001, Ping Li 0016 |
CGI | 5 |
| 2020 | Hierarchical Rendering System Based on Viewpoint Prediction in Virtual Reality
Ping Lu 0008, Ping Li 0016, Jinman Kim, Bin Sheng 0001, Lijuan Mao |
CGI | 3 |
| 2020 | GHand: A Graph Convolution Network for 3D Hand Pose Estimation
Pengsheng Wang, Guangtao Xue, Ping Li 0016, Jinman Kim, Bin Sheng 0001, Lijuan Mao |
CGI | 3 |
| 2020 | Preserving Temporal Consistency in Videos Through Adaptive SLIC
Han Zhang 0053, Riaz Ali, Bin Sheng 0001, Ping Li 0016, Jinman Kim |
CGI | 4 |
| 2020 | 3D Geology Scene Exploring Base on Hand-Track Somatic Interaction
Ping Lu 0008, Ping Li 0016, Bin Sheng 0001, Lijuan Mao |
CGI | 4 |
| 2020 | Gaze-Contingent Rendering in Virtual Reality
Ping Lu 0008, Ping Li 0016, Bin Sheng 0001, Lijuan Mao |
CGI | 3 |
| 2020 | RANet: Region Attention Network for Semantic SegmentationabstractRecent semantic segmentation methods model the relationship between pixels to construct the contextual representations. In this paper, we introduce the \emph{Region Attention Network} (RANet), a novel attention network for modeling the relationship between object regions. RANet divides the image into object regions, where we select representative information. In contrast to the previous methods, RANet configures the information pathways between the pixels in different regions, enabling the region interaction to exchange the regional context for enhancing all of the pixels in the image. We train the construction of object regions, the selection of the representative regional contents, the configuration of information pathways and the context exchange between pixels, jointly, to improve the segmentation accuracy. We extensively evaluate our method on the challenging segmentation benchmarks, demonstrating that RANet effectively helps to achieve the state-of-the-art results. Dingguo Shen, Yuanfeng Ji, Ping Li 0016, Yi Wang 0031, Di Lin 0002 |
NeurIPS | 3 |
| 2020 | Real-time hair simulation with heptadiagonal decomposition on mass spring system
Jianwei Jiang 0002, Bin Sheng 0001, Ping Li 0016, Lizhuang Ma, Xin Tong 0001, Enhua Wu |
Graph. Model. | 3 |
| 2020 | SPST-CNN: Spatial pyramid based searching and tagging of liver's intraoperative live views via CNN for minimal invasive surgery
Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Ping Li 0016, Huating Li, Po Yang 0001, Younhyun Jung, Harry Qin, David Dagan Feng |
J. Biomed. Informatics | 4 |
| 2020 | Automatic Video Segmentation Based on Information Centroid and Optimized SaliencyCut
Huisi Wu, Meng-Shu Liu, Lulu Yin, Ping Li 0016, Zhenkun Wen, Hon-Cheng Wong |
J. Comput. Sci. Technol. | 4 |
| 2020 | Embedding 3D models in offline physical environmentsabstractAbstract This article introduces a novel approach for embedding 3D models in offline physical environments using quick response (QR) codes. Unlike conventional methods, we consider settings where 3D models cannot be retrieved from a remote server. Our method involves generating octree models from voxelized 3D models and storing them in QR codes using a space‐efficient data structure. This allows storing 3D models that are both intelligible and purposeful on standard QR codes while addressing the major storage constraint that is present in offline situations. Furthermore, we explore 3D convolutional neural networks (CNN) and autoencoders (AE) to compress 3D models with high resolutions where using octrees alone does not suffice. To the best of our knowledge, our AE network is the first to employ octrees to further compress its encoded data. Through user‐friendly desktop and mobile applications, we allow users to encode, decode and visualize 3D models in augmented reality (AR) using QR codes, thus experiment with our methods. The proposed approach enables unique applications and future research in ubiquitous computing, 3D data compression and transmission, 3D AEs, AR and Virtual Reality, low‐cost autonomous robots, and 3D printing. Egemen Ertugrul, Han Zhang 0053, Ping Lu 0008, Ping Li 0016, Bin Sheng 0001, Enhua Wu |
Comput. Animat. Virtual Worlds | 5 |
| 2020 | Video flickering removal using temporal reconstruction optimization
Bin Sheng 0001, Ping Li 0016, Gaoqi He |
Multim. Tools Appl. | 4 |
| 2020 | Illumination-Invariant Video Cut-Out Using Octagon Sensitive OptimizationabstractThis paper presents an effective video cut-out approach, which can be utilized to segment the moving object in video shots. We first introduce the Octagon-Sensitive-Filtering (OSF) and its illumination invariant feature (IIF), which is computed on each pixel of the image via adding contributions from neighboring pixels. We integrate our IIF into the variational model and obtain the seeds during preprocessing to help address large displacement and illumination changes. An effective seed update method based on tracking-then-refinement based on IIF is presented to compensate for location ambiguities, and the strategy is effective to deal with illumination variances and objects deformation. Furthermore, we apply the IIF-based graph-cut to deal with fuzzy boundaries. Multiple experiments on quantitative challenging datasets have shown the robustness, high-quality video cut-out and efficiency of our approach to acute variances of illumination and complex motion. Jingye Wang, Bin Sheng 0001, Ping Li 0016, David Dagan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Interactive Contour Extraction via Sketch-Alike Dense-Validation OptimizationabstractWe propose an interactive contour extraction method inspired by a skill often adopted in sketching: an artist usually sketches an object by first drawing lots of short, directional, and redundant strokes, then following these small strokes to draw the final outline of the object. Our method simulates this process. To extract a contour, our method relies on user interaction, which provides us with a narrow band containing the target contour. Then, we densely sample sub-bands from the whole band, with each sub-band containing a local segment of the target contour. We design a curve-centered coordinate system in which a dynamic programming algorithm is proposed to extract the local segment in each sub-band. The local segment is guaranteed to be as evident and smooth as possible, to mimic the strokes sketched by the artist. Finally, we integrate all local segments of all sub-bands together to obtain the whole target contour based on the weighted principal component analysis. Our method can extract high-quality object contours due to the dense validations among local segments. That is, even if one segment deviates from the right location, several other segments in its local neighborhood can correct it in the integration stage. Both quantitative experiments and a user study demonstrate the effectiveness of the proposed method. Yongwei Nie, Ping Li 0016, Qing Zhang 0006, Zhensong Zhang, Guiqing Li, Hanqiu Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Depth-Aware Motion Deblurring Using Loopy Belief PropagationabstractMost motion-blurred images captured in the real world have spatially-varying point-spread functions, and some are caused by different positions and depth values, which cannot be handled by most state-of-the-art deblurring methods based on deconvolution. To overcome this problem, we propose a depth-aware motion blur model that treats a blurred image as an integration of a sequence of clear images. To restore the clear latent image, we extend the Richardson-Lucy method to incorporate our blur model with a given depth image. The empty holes in the depth image, caused by occlusion or device limitations, are fixed by PatchMatch-based depth filling. We regard the depth image as a Markov random field and select candidate labels by using belief propagation to set and smooth depth values for empty areas. Deblurring and depth filling are performed iteratively to refine the results. Our method can also be applied to real-world images with the assistance of motion estimation. The deblurring process is shown to be convergent; moreover, the number of iterations and the level of noise amplification are acceptable. The experimental results show that our method can not only handle depth-variant motion blur but also refine depth images. Bin Sheng 0001, Ping Li 0016, Xiaoxin Fang, Ping Tan 0002, Enhua Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Outdoor Shadow Estimating Using Multiclass Geometric Decomposition Based on BLSabstractIllumination is a significant component of an image, and illumination estimation of an outdoor scene from given images is still challenging yet it has wide applications. Most of the traditional illumination estimating methods require prior knowledge or fixed objects within the scene, which makes them often limited by the scene of a given image. We propose an optimization approach that integrates the multiclass cues of the image(s) [a main input image and optional auxiliary input image(s)]. First, Sun visibility is estimated by the efficient broad learning system. And then for the scene with visible Sun, we classify the information in the image by the proposed classification algorithm, which combines the geometric information and shadow information to make the most of the information. And we apply a respective algorithm for every class to estimate the illumination parameters. Finally, our approach integrates all of the estimating results by the Markov random field. We make full use of the cues in the given image instead of an extra requirement for the scene, and the qualitative results are presented and show that our approach outperformed other methods with similar conditions. Bin Sheng 0001, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Cybern. | 4 |
| 2020 | SCN: Switchable Context Network for Semantic Segmentation of RGB-D ImagesabstractContext representations have been widely used to profit semantic image segmentation. The emergence of depth data provides additional information to construct more discriminating context representations. Depth data preserves the geometric relationship of objects in a scene, which is generally hard to be inferred from RGB images. While deep convolutional neural networks (CNNs) have been successful in solving semantic segmentation, we encounter the problem of optimizing CNN training for the informative context using depth data to enhance the segmentation accuracy. In this paper, we present a novel switchable context network (SCN) to facilitate semantic segmentation of RGB-D images. Depth data is used to identify objects existing in multiple image regions. The network analyzes the information in the image regions to identify different characteristics, which are then used selectively through switching network branches. With the content extracted from the inherent image structure, we are able to generate effective context representations that are aware of both image structures and object relationships, leading to a more coherent learning of semantic segmentation network. We demonstrate that our SCN outperforms state-of-the-art methods on two public datasets. Di Lin 0002, Ruimao Zhang, Yuanfeng Ji, Ping Li 0016, Hui Huang 0004 |
IEEE Trans. Cybern. | 4 |
| 2020 | Automated Decision Support System for Lung Cancer Detection and Classification via Enhanced RFCN With Multilayer Fusion RPNabstractDetection of lung cancer at early stages is critical, in most of the cases radiologists read computed tomography (CT) images to prescribe follow-up treatment. The conventional method for detecting nodule presence in CT images is tedious. In this article, we propose an enhanced multidimensional region-based fully convolutional network (mRFCN) based automated decision support system for lung nodule detection and classification. The mRFCN is used as an image classifier backbone for feature extraction along with the novel multilayer fusion region proposal network (mLRPN) with position-sensitive score maps being explored. We applied a median intensity projection to leverage three-dimensional information from CT scans and introduced deconvolutional layer to adopt proposed mLRPN in our architecture to automatically select the potential region of interest. Our system has been trained and evaluated using LIDC dataset, and the experimental results showed promising detection performance in comparison to the state-of-the-art nodule detection/classification methods, achieving a sensitivity of 98.1% and classification accuracy of 97.91%. Anum Masood, Bin Sheng 0001, Po Yang 0001, Ping Li 0016, Huating Li, Jinman Kim, David Dagan Feng |
IEEE Trans. Ind. Informatics | 4 |
| 2020 | OFF-eNET: An Optimally Fused Fully End-to-End Network for Automatic Dense Volumetric 3D Intracranial Blood Vessels SegmentationabstractIntracranial blood vessels segmentation from computed tomography angiography (CTA) volumes is a promising biomarker for diagnosis and therapeutic treatment in cerebrovascular diseases. These segmentation outputs are a fundamental requirement in the development of automated decision support systems for preoperative assessment or intraoperative guidance in neuropathology. The state-of-the-art in medical image segmentation methods are reliant on deep learning architectures based on convolutional neural networks. However, despite their popularity, there is a research gap in the current deep learning architectures optimized to address the technical challenges in blood vessel segmentation. These challenges include: (i) the extraction of concrete brain vessels close to the skull; and (ii) the precise marking of the vessel locations. We propose an Optimally Fused Fully end-to-end Network (OFF-eNET) for automatic segmentation of the volumetric 3D intracranial vascular structures. OFF-eNET comprises of three modules. In the first module, we exploit the up-skip connections to enhance information flow, and dilated convolution for detailed preservation of spatial feature map that are designed for thin blood vessels. In the second module, we employ residual mapping along with inception module for speedy network convergence and richer visual representation. For the third module, we make use of the transferred knowledge in the form of cascaded training strategy to gradually optimize the three segmentation stages (basic, complete, and enhanced) to segment thin vessels located close to the skull. All these modules are designed to be computationally efficient. Our OFF-eNET, evaluated using 70 CTA image volumes, resulted in 90.75% performance in the segmentation of intracranial blood vessels and outperformed the state-of-the-art counterparts. Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Huating Li, Ping Li 0016, Po Yang 0001, Younhyun Jung, Harry Qin, Jinman Kim, David Dagan Feng |
IEEE Trans. Image Process. | 5 |
| 2020 | Intrinsic Image Decomposition with Step and Drift Shading SeparationabstractDecomposing an image into the shading and reflectance layers remains challenging due to its severely under-constrained nature. We present an approach based on illumination decomposition that recovers the intrinsic images without additional information, e.g., depth or user interaction. Our approach is based on the rationale that the shading component contains the step and drift channels simultaneously. We decompose the illumination into two channels: the step shading, corresponding to the sharp shading changes due to cast shadow or abrupt shape changes; the drift shading, accounting for the smooth shading variations due to gradual illumination changes or slow shape changes. Due to such transformation of turning the conventional assumption that shading has smoothness as reasonable prior, our model has the advantages in handling real images, especially with the cast shadows or strong shape edges. We also apply a much stricter edge classifier along with a reinforcement process to enhance our method. We formulate the problem using a two-parameter energy function and split it into two energy functions corresponding to the reflectance and step shading. Experiments on the MIT dataset, the IIW dataset and the MPI Sintel dataset have shown the success of our approach over the state-of-the-art methods. Bin Sheng 0001, Ping Li 0016, Yuxi Jin, Ping Tan 0002, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Depth of Field Rendering Using Multilayer-Neighborhood OptimizationabstractDepth of field (DOF) is utilized widely to deliver artistic effects in photography. However, existing post-processing techniques for rendering DOF effects introduce visual artifacts such as color leakage, blurring discontinuity, and the partial occlusion problems which limit the application of DOF. Traditionally, occluded pixels are ignored or not well estimated although they might make key contributions to images. In this paper, we propose a new filtering approach which takes approximated occluded pixels into account to synthesize the DOF effects for images. In our approach, images are separated into different layers based on depth. Besides, we utilize adaptive PatchMatch method to estimate the intensities of occluded pixels, especially in the background region. We again propose a new multilayer-neighborhood optimization to estimate occluded pixels contributions and render the images. Finally, we apply gathering filter to achieve the rendered images with elite DOF effects. Multiple experiments have shown that our approach can handle color leakage, blurring discontinuity and partial occlusion problem while providing high-quality DOF rendering effects. Benxuan Zhang, Bin Sheng 0001, Ping Li 0016, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Simplified non-locally dense network for single-image dehazing
Zhuoliang Hu, Bin Sheng 0001, Ping Li 0016, Jinman Kim, Enhua Wu |
Vis. Comput. | 4 |
| 2020 | On attaining user-friendly hand gesture interfaces to control existing GUIsabstractBackground Hand gesture interfaces are dedicated programs that principally perform hand tracking and hand gesture prediction to provide alternative controls and interaction methods. They take advantage of one of the most natural ways of interaction and communication, proposing novel input and showing great potential in the field of the human-computer interaction. Developing a flexible and rich hand gesture interface is known to be a time-consuming and arduous task. Previously published studies have demonstrated the significance of the finite-state-machine (FSM) approach when mapping detected gestures to GUI actions. Methods In our hand gesture interface, we broadened the FSM approach by utilizing gesture-specific attributes, such as distance between hands, distance from the camera, and time of occurrences, to enable users to perform unique GUI actions. These attributes are obtained from hand gestures detected by the RealSense SDK employed in our hand gesture interface. By means of these gesture-specific attributes, users can activate static gestures and perform them as dynamic gestures. We also provided supplementary features to enhance the efficiency, convenience, and user-friendliness of our hand gesture interface. Moreover, we developed a complementary application for recording hand gestures by capturing hand keypoints in depth and color images to facilitate the generation of hand gesture datasets. Results We conducted a small-scale user survey with fifteen subjects to test and evaluate our hand gesture interface. Anonymous feedback obtained from the users indicates that our hand gesture interface is adequately facile and self-explanatory to use. In addition, we received constructive feedback about minor flaws regarding the responsiveness of the interface. Conclusions We proposed a hand gesture interface along with key concepts to attain user-friendliness and effectiveness in the control of existing GUIs. Egemen Ertugrul, Ping Li 0016, Bin Sheng 0001 |
Virtual Real. Intell. Hardw. | 2 |
| 2020 | An accurate multi-modal biometric identification system for person identification via fusion of face and finger print
Sidra Aleem, Po Yang 0001, Saleha Masood, Ping Li 0016, Bin Sheng 0001 |
World Wide Web | 4 |
| 2019 | Optic Disc and Cup Segmentation Based on Enhanced SegNetabstractDue to imbalanced distributed and restricted medical resources, reliable analysis for medical images is hard to come by, and it is impractical to only rely on human beings to do all the analysis, which is time-consuming and not economic. Application of computer vision techniques in such fields emerges as the situation requires. In this paper, we use deep learning segmentation algorithm to segment the optic disc and the cup from each other and from the rest of the ophthalmoscopy photographs. For a better performance, we change the loss function and crop as a way of data augmentation. The segmentation results can be used to calculate the cup-to-disc ratio (CDR), which is further used to diagnose glaucoma. Challenges such as over-fitting, biased dataset, and poor generalization of the model exist in front of us. We illustrate our model and associated methods dealing with these challenges. Lianyi Wu, Yelin Shi, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim |
CASA | 5 |
| 2019 | Detect Glaucoma with Image Segmentation and Transfer LearningabstractIn this paper, we aim to automatically detect glaucoma via deep learning. To do that, we need to calculate the cup-to-disc ratio (CDR) on fine segmented retina images. To get precise segmentation, we implemented SegNet together with adversarial discriminative domain adaptation (ADDA), the former is a famous artificial neural network with encoder-decoder architecture used in image segmentation area and the latter is a transfer learning method for domain adaptation. We are the first to combine them together to detect glaucoma on test dataset which have different brightness from our training dataset. We thoroughly evaluated the proposed method with various loss functions, normal cross entropy loss, weighted cross entropy loss and dice coefficient loss included. And we show that dice loss is the best for this task. Last but not least, our experiments on transfer learning have shown that our ADDA method reduces the mean square error (MSE) between the CDR of our segmentation and annotations greatly. Lianyi Wu, Yelin Shi, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim |
CASA | 5 |
| 2019 | Deep Intrinsic Image Decomposition Using Joint Parallel Learning
Yuan Yuan 0022, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim, Enhua Wu |
CGI | 3 |
| 2019 | Automatic Computer Aided System for Lung Cancer in Chest CTs Using MD-RFCN Combined with Tri-Level Region Proposal NetworkabstractPulmonary cancer is one of the major causes of deaths caused by cancer around the globe. Early stage lung cancer detection can prove to be essential for the patients, for which the computed tomography (CT) images are analyzed by the radiologists to determine the presence of nodules and diagnose the disease. Conventional techniques used by the radiologists for nodule detection in CT images is time-consuming and inefficient; to assist in the diagnosis process and further enhance its efficiency and accuracy, decision support systems have been developed in the past few years. In our paper, we proposed a Multi-Dimension Region-based Fully Convolutional Network based decision support system for detection and classification of lung nodule. The Multi-Dimension RFCN serves as an image classifier backbone for our feature extraction step in addition to the proposed Tri-Level Region Proposal Network (3L-RPN) along with the position-sensitive score maps (PSSM) being explored. A novel median intensity projection method is used to leverage the multi-dimensional information from CT images and introduced an additional deconvolutional layer to adopt the proposed Tri-Level Region Proposal Network in our architecture to automatically identify the potential Region of Interest. We trained and evaluated our proposed decision support system using LIDC-IDRI dataset. The evaluation results demonstrated the high level performance of our proposed model in comparison to the state-of-the-art nodule detection and classification methods by attaining classification accuracy of 97.61% and sensitivity of 97.4%. Anum Masood, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, Jinman Kim |
INDIN | 3 |
| 2019 | An Investigation of 3D Human Pose Estimation for Learning Tai Chi: A Human Factor PerspectiveabstractIn this article, we propose a Tai Chi training system based on pose estimation using Convolutional Neural Networks (CNNs) called iTai-Chi. Our system aims to overcome the disadvantages of insufficient accurate feedback in traditional teaching methods such as one-to-many tutorial and video watching. With the specially trained neural network, our iTai-Chi system can estimate learners’ poses more accurately compared to Kinect V2. In our system, user’s motion is evaluated through comparison with the template motion. The evaluated results are presented to the user to locate the error in their motions and help their correction. To verify the effectiveness of our system, we carried out a series of user studies. Results reflect that the iTai-Chi system successfully improve users’ performance in movement accuracy. Also, our system assists elder Tai Chi practitioners and students without prior knowledge to overcome learning obstacles and improve their skills. The users agreed that our system is interesting and supportive for their Tai Chi learning. Aouaidjia Kamel, Bowen Liu 0015, Ping Li 0016, Bin Sheng 0001 |
Int. J. Hum. Comput. Interact. | 3 |
| 2019 | SRNPD: Spatial rendering network for pencil drawing stylizationabstractAbstract Pencil drawing is a simple yet effective way to depict what people see by clearly presenting details of the scene. Existing methods usually extract strokes of the input image and adjust the result image tone to make it look like a pencil drawing. However, they do not consider the quality of the stroke image and the geometry information of lines in the stroke image, which unavoidably results in the violation of original essential structures and in a flatten pencil drawing with unrealistic appearance. We put forward a spatial rendering network for pencil drawing stylization. Spatial stroke images are extracted from the image pyramid by a single‐shot bottom‐up neural network to improve the quality of these stroke images. Unlike the former tone adjustment–based methods, we analyze perceptual cues of strokes at different stroke image levels and use the obtained geometry information to constrain the stroke shading procedure. The final pencil drawing result is achieved by the stroke shading fusion of different levels' shading results. The effectiveness of our spatial rendering network for pencil drawing stylization is demonstrated by an ablation study, comparison to the state of the art, and a user study. Yuxi Jin, Ping Li 0016, Bin Sheng 0001, Yongwei Nie, Jinman Kim, Enhua Wu |
Comput. Animat. Virtual Worlds | 2 |
| 2019 | Multiview-coherent disocclusion synthesis using connected regions optimizationabstractAbstract Handling of missing areas is a key step for depth‐based rendering to synthesize virtual views. Existing methods usually consider finding candidate pixels from only one reference view to fill missing areas. However, the information provided by one reference view is restricted by the position of the view. By utilizing two reference views located on both the left and right sides of the virtual view, we propose to synthesize the missing areas at hole level with connected regions optimization. To avoid the appearance of ghost boundary, we apply morphological operations to generate a boundary band map for the depth map, which restricts the warping of the virtual view. We use a binary map to mark unknown pixels in a hole, label the connected unknown regions, and count the area of each connected region, which decide the order in our enhanced inpainting synthesis. Besides, we separate the foreground and background regions of the depth map to constrain the searching of candidate pixels. Multiple experiments on virtual view synthesis have shown the effectiveness and high quality of our multiview‐coherent disocclusion synthesis. Ping Li 0016, Yuxi Jin, Bin Sheng 0001, Di Lin 0002, Yongwei Nie, Enhua Wu |
Comput. Animat. Virtual Worlds | 1 |
| 2019 | Retinal Vessel Segmentation Using Minimum Spanning Superpixel Tree DetectorabstractThe retinal vessel is one of the determining factors in an ophthalmic examination. Automatic extraction of retinal vessels from low-quality retinal images still remains a challenging problem. In this paper, we propose a robust and effective approach that qualitatively improves the detection of low-contrast and narrow vessels. Rather than using the pixel grid, we use a superpixel as the elementary unit of our vessel segmentation scheme. We regularize this scheme by combining the geometrical structure, texture, color, and space information in the superpixel graph. And the segmentation results are then refined by employing the efficient minimum spanning superpixel tree to detect and capture both global and local structure of the retinal images. Such an effective and structure-aware tree detector significantly improves the detection around the pathologic area. Experimental results have shown that the proposed technique achieves advantageous connectivity-area-length (CAL) scores of 80.92% and 69.06% on two public datasets, namely, DRIVE and STARE, thereby outperforming state-of-the-art segmentation methods. In addition, the tests on the challenging retinal image database have further demonstrated the effectiveness of our method. Our approach achieves satisfactory segmentation performance in comparison with state-of-the-art methods. Our technique provides an automated method for effectively extracting the vessel from fundus images. Bin Sheng 0001, Ping Li 0016, Shuangjia Mo, Huating Li, Xuhong Hou, Harry Qin, Ruogu Fang, David Dagan Feng |
IEEE Trans. Cybern. | 2 |
| 2019 | Fast and Accurate Retinal Identification System: Using Retinal Blood Vasculature LandmarksabstractThe expansion of automation techniques and increased risk of identity theft have led emphasis on the tremendous need of automated identification system. Due to the high recognition accuracy and robustness to changes in human physiology, retinal biometric identification system has drawn much attention in this research field. In this paper, we aim to propose an automatic fast and accurate retinal identification system for the multisample dataset. The proposed approach uses a hybrid segmentation technique to segment out both thick/thin vessels for effectively balancing the difference of wavelet response between thick/thin blood vessels. As a result, recognition accuracy is improved. A Principle Component Analysis-based feature processing approach is proposed for efficiently reducing the dimensionality of a large number of vessels features. It significantly reduces computation time and accelerates the matching process in the retinal identification system. The proposed technique is validated on DRIVE, STARE, VARIA, RIDB, HRF, Messidor, DIARETDB0, and a large multisample per subject database created by authors using the images provided by Dr. Chen (Shanghai Jiao Tong University Affiliated Sixth People Hospital). Experimental results demonstrated that the proposed approach outperforms other existing techniques. Segmentation achieves an overall accuracy of 99.65% with the recognition rate of 99.40% on all these databases. Sidra Aleem, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, David Dagan Feng |
IEEE Trans. Ind. Informatics | 3 |
| 2019 | Illumination-Guided Video Composition via Gradient Consistency OptimizationabstractVideo composition aims at cloning a patch from the source video into the target scene to create a seamless and harmonious blending frame sequence. Previous work in video composition usually suffer from artifacts around the blending region and spatial-temporal consistency when illumination intensity varies in the input source and target video. We propose an illumination-guided video composition method via a unified spatial and temporal optimization framework. Our method can produce globally consistent composition results and maintain the temporal coherency. We first compute a spatial-temporal blending boundary iteratively. For each frame, the gradient field of the target and source frames are mixed adaptively based on gradients and inter-frame color difference. The temporal consistency is further obtained by optimizing luminance gradients throughout all the composition frames. Moreover, we extend the mean-value cloning by smoothing discrepancies between the source and target frames, then eliminate the color distribution overflow exponentially to reduce falsely blending pixels. Various experiments have shown the effectiveness and high-quality performance of our illumination-guided composition. Jingye Wang, Bin Sheng 0001, Ping Li 0016, Yuxi Jin, David Dagan Feng |
IEEE Trans. Image Process. | 3 |
| 2019 | Deep Color Guided Coarse-to-Fine Convolutional Network Cascade for Depth Image Super-ResolutionabstractDepth image super-resolution is a significant yet challenging task. In this paper, we introduce a novel deep color guided coarse-to-fine convolutional neural network (CNN) framework to address this problem. First, we present a datadriven filter method to approximate the ideal filter for depth image super-resolution instead of hand-designed filters. Based on large data samples, the filter learned is more accurate and stable for upsampling depth image. Second, we introduce a coarse-to-fine CNN to learn different sizes of filter kernels. In coarse stage, larger filter kernels are learned by CNN to achieve crude high-resolution depth image. As to fine stage, the crude high-resolution depth image is used as the input so that smaller filter kernels are learned to gain more accurate results. Benefit from this network, we can progressively recover the high frequency details. Third, we construct a color guidance strategy that fuses color difference and spatial distance for depth image upsampling. We revise the interpolated high-resolution depth image according to the corresponding pixels in highresolution color maps. Guided by color information, the depth of high-resolution image obtained can alleviate texture copying artifacts and preserve edge details effectively. Quantitative and qualitative experimental results demonstrate our state-of-the-art performance for depth map super-resolution. Bin Sheng 0001, Ping Li 0016, Weiyao Lin, David Dagan Feng |
IEEE Trans. Image Process. | 3 |
| 2019 | Deep Convolutional Neural Networks for Human Action Recognition Using Depth Maps and PosturesabstractIn this paper, we present a method (Action-Fusion) for human action recognition from depth maps and posture data using convolutional neural networks (CNNs). Two input descriptors are used for action representation. The first input is a depth motion image that accumulates consecutive depth maps of a human action, whilst the second input is a proposed moving joints descriptor which represents the motion of body joints over time. In order to maximize feature extraction for accurate action classification, three CNN channels are trained with different inputs. The first channel is trained with depth motion images (DMIs), the second channel is trained with both DMIs and moving joint descriptors together, and the third channel is trained with moving joint descriptors only. The action predictions generated from the three CNN channels are fused together for the final action classification. We propose several fusion score operations to maximize the score of the right action. The experiments show that the results of fusing the output of three channels are better than using one channel or fusing two channels only. Our proposed method was evaluated on three public datasets: 1) Microsoft action 3-D dataset (MSRAction3D); 2) University of Texas at Dallas-multimodal human action dataset; and 3) multimodal action dataset (MAD) dataset. The testing results indicate that the proposed approach outperforms most of existing state-of-the-art methods, such as histogram of oriented 4-D normals and Actionlet on MSRAction3D. Although MAD dataset contains a high number of actions (35 actions) compared to existing action RGB-D datasets, this paper surpasses a state-of-the-art method on the dataset by 6.84%. Aouaidjia Kamel, Bin Sheng 0001, Po Yang 0001, Ping Li 0016, Ruimin Shen, David Dagan Feng |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2019 | Deep Neural Representation Guided Face Sketch SynthesisabstractFace sketch synthesis shows great applications in a lot of fields such as online entertainment and suspects identification. Existing face sketch synthesis methods learn the patch-wise sketch style from the training dataset containing photo-sketch pairs. These methods manipulate the whole process directly in the field of RGB space, which unavoidably results in unsmooth noises at patch boundaries. If denoising methods are used, the sketch edges would be blurred and face structures could not be restored. Recent researches of feature maps, which are the outputs of a certain neural network layer, have achieved great success in texture synthesis and artistic image generation. In this paper, we reformulate the face sketch synthesis problem into a neural network feature maps based optimization task. Our results accurately capture the sketch drawing style and make full use of the whole stylistic information hidden in the training dataset. Unlike former feature map based methods, we utilize the Enhanced 3D PatchMatch and cross-layer cost aggregation methods to obtain the target feature maps for the final results. Multiple experiments have shown that our approach imitates hand-drawn sketch style vividly, and has high-quality visual effects on CUHK, AR, XM2VTS and CUFSF face sketch datasets. Bin Sheng 0001, Ping Li 0016, Chenhao Gao, Kwan-Liu Ma |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Structure-preserving image completion with multi-level dynamic patches
Bowen Liu 0015, Ping Li 0016, Bin Sheng 0001, Yongwei Nie, Enhua Wu |
Vis. Comput. | 2 |
| 2018 | Parallel Pencil Drawing Stylization via Structure-Aware OptimizationabstractThis paper presents a non-photorealistic rendering technique for stylizing a photograph in the pencil drawing style, which can well preserve the fine structure of the original image. We first construct a structure map from the anti-color image of the original image to model the detailed underlying fine structure of the original image, and generate a coarse pencil drawing image with line integral convolution. Then we refine the coarse pencil drawing image with the structure map for enrich the structure information. The presented algorithm is highly parallel allowing a real-time performance with GPU implementation. Experimental results show that our approach can produce more attractive and impressive pencil drawing effects with a variety of photographs. Yuxi Jin, Bin Sheng 0001, Ping Li 0016, Hanqiu Sun |
CASA | 4 |
| 2018 | Voxelized Facial Reconstruction Using Deep Neural NetworkabstractThis paper presents an approach to predicting variation tendency of human faces with regard to cranium changes based on deep learning. Our work focuses on generating individual customized facial models with high plausibility. Inspired by the performance of encoder-decoder convolutional neural network, the core trainable predicting engine of our learning network is designed for three-dimension voxelized data representation as the encoder-decoder structure and the encoder part is similar to the 7 layers of VGG16 network. To take full consideration of the cranium changes and features of original human face, a novel formation of channeled volumetric data structure is presented, and also the corresponding sub and up-sampling strategies for volume data. Our encoder-decoder neural network consumes discrete 3-channel volume data and generates 1-channel volume data as predicted post-variation human face. This framework is quantified with clinical dataset and it shows that its' performance improves in comparison with the state-of-the-art technologies. Xiaoshuang Li, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng |
CGI | 3 |
| 2018 | Accelerated robust Boolean operations based on hybrid representations
Bin Sheng 0001, Bowen Liu 0015, Ping Li 0016, Hongbo Fu 0001, Lizhuang Ma, Enhua Wu |
Comput. Aided Geom. Des. | 3 |
| 2018 | Efficient non-incremental constructive solid geometry evaluation for triangular meshes
Bin Sheng 0001, Ping Li 0016, Hongbo Fu 0001, Lizhuang Ma, Enhua Wu |
Graph. Model. | 2 |
| 2018 | Automatic choroid layer segmentation using normalized graph cutabstractOptical coherence tomography is an immersive technique for depth analysis of retinal layers. Automatic choroid layer segmentation is a challenging task because of the low contrast inputs. Existing methodologies carried choroid layer segmentation manually or semi‐automatically. The authors proposed automated choroid layer segmentation based on normalised cut algorithm, which aims at extracting the global impression of images and treats the segmentation as a graph partitioning problem. Due to the structure complexity of retinal and choroid layers, the authors employed a series of pre‐processing to make the cut more deterministic and accurate. The proposed method divided the image into several patches and ran the normalised cut algorithm on every patch separately. The aim was to avoid insignificant vertical cuts and focus on horizontal cutting. After processing every patch, the authors acquired a global cut on the original image by combining all the patches. Later the authors measured the choroidal thickness which is highly helpful in the diagnosis of several retinal diseases. The results were computed on a total of 525 images of 21 real patients. Experimental results showed that the mean relative error rate of the proposed method was around 0.4 when compared with the manual segmentation performed by the experts. Saleha Masood, Bin Sheng 0001, Ping Li 0016, Ruimin Shen, Ruogu Fang |
IET Image Process. | 3 |
| 2018 | Computer-Assisted Decision Support System in Pulmonary Cancer detection and stage classification on CT images
Anum Masood, Bin Sheng 0001, Ping Li 0016, Xuhong Hou, Xiaoer Wei, Harry Qin, David Dagan Feng |
J. Biomed. Informatics | 3 |
| 2018 | Illumination-aware live videos background replacement using antialiasing optimization
Qiaoping Hu, Hanqiu Sun, Ping Li 0016, Ruimin Shen, Bin Sheng 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Video Decolorization Using Visual Proximity Coherence OptimizationabstractVideo decolorization is to filter out the color information while preserving the perceivable content in the video as much and correct as possible. Existing methods mainly apply image decolorization strategies on videos, which may be slow and produce incoherent results. In this paper, we propose a video decolorization framework that considers frame coherence and saves decolorization time by referring to the decolorized frames. It has three main contributions. First, we define decolorization proximity to measure the similarity of adjacent frames. Second, we propose three decolorization strategies for frames with low, medium, and high proximities, to preserve the quality of these three types of frames. Third, we propose a novel decolorization Gaussian mixture model to classify the frames and assign appropriate decolorization strategies to them based on their decolorization proximity. To evaluate our results, we measure them from three aspects: 1) qualitative; 2) quantitative; and 3) user study. We apply color contrast preserving ratio and C2G-SSIM to evaluate the quality of single frame decolorization. We propose a novel temporal coherence degree metric to evaluate the temporal coherence of the decolorized video. Compared with current methods, the proposed approach shows all around better performance in time efficiency, temporal coherence, and quality preservation. Yizhang Tao, Yiyi Shen, Bin Sheng 0001, Ping Li 0016, Rynson W. H. Lau |
IEEE Trans. Cybern. | 4 |
| 2017 | Image completion with dynamic patchesabstractThis paper presents an approach of structure-preserving image completion with dynamic patches. Existing image completion methods may generate unnatural abnormal structures or structure disorders due to limited patches and patterns availability. Our structure-preserving image completion utilizes objective function minimization considering the coherence not only within the image cavity but also with global constraints. A series of dynamic patch-based optimizations are applied to fulfill the cavity. Unlike traditional fixed-size patch-based methods, our image completion with competitive dynamic patch-matching mechanism provides more effective structure restoration. Parallel searching of different-sized patches is performed to retrieve optimal patches for completing the image cavity with nice structure preservation. The experiments show that the realistic completed images by our approach are visually pleasing with nice structural coherence. Bowen Liu 0015, Ping Li 0016, Bin Sheng 0001, Enhua Wu |
CGI | 2 |
| 2017 | Dynamic RGB-to-CMYK conversion using visual contrast optimisationabstractAs the standard colour space used by printers, Cyan, Magenta, Yellow, Black (CMYK) colour model is a subtractive colour space used to describe the printing process. Existing CMYK conversion methods rely on static conversion table, which may not preserve the subtle visual structures of images, due to the local visual contrast loss caused by the static colour mapping. Therefore, the authors propose a novel dynamic Red, Green, Blue (RGB)‐to‐CMYK colour conversion, which utilises the weighted entropy to extract the pixels with filter response change dramatically. They obtain the image activity map by combining these pixels with high skin probability regions, and optimise the colour conversion of each pixel to ensure that the ink used for each pixel can be saved, while the visual contrast can be preserved with ink‐saving. In this way, their proposed technique can achieve dynamic CMYK colour conversion, in which the consumption of ink can be reduced without the loss of visual contrast. The experimental results have shown that their dynamic CMYK colour conversion saved 10–25% ink consumption compared with the static conversion method, while with high visual quality for the converted images. Zhenzhu Wang, Bin Sheng 0001, Ruimin Shen, Ping Li 0016 |
IET Image Process. | 6 |
| 2017 | Abdominal adipose tissues extraction using multi-scale deep neural network
Fei Jiang 0006, Huating Li, Xuhong Hou, Bin Sheng 0001, Ruimin Shen, Xiao-Yang Liu, Weiping Jia, Ping Li 0016, Ruogu Fang |
Neurocomputing | 8 |
| 2017 | Antialiased super-resolution with parallel high-frequency synthesis
Xudong Jiang 0003, Bin Sheng 0001, Weiyao Lin, Ping Li 0016, Lizhuang Ma, Ruimin Shen |
Multim. Tools Appl. | 4 |
| 2017 | Retinal optic disc localization using convergence tracking of blood vessels
Rui Wang 0130, Linghan Zheng, Chaoqun Xiong, Chunfang Qiu, Huating Li, Xuhong Hou, Bin Sheng 0001, Ping Li 0016 |
Multim. Tools Appl. | 8 |
| 2017 | Automatic diabetic retinopathy diagnosis using adjustable ophthalmoscope and multi-scale line operator
Meng Qu, Chun Ni, Mufan Chen, Linghan Zheng, Bin Sheng 0001, Ping Li 0016 |
Pervasive Mob. Comput. | 7 |
| 2016 | A Case Study Illustrating Coding for ComputationalThinking DevelopmentabstractPrevious studies suggested that computational thinking is an important skill that schools should start equipping children with from primary school education. This study proposes way to develop a mobile app coding curriculum in primary 4 to 6 to nurture students’ computational thinking. It will guide students to undergo learning stages from developing their own codes and combining with others’ work to create their own apps and staging their apps to gain confidence and recognition of performance in coding. An example of mathematic game in the curriculum is selected as a case study to illustrate in this study how it incorporates the elements of computational thinking in the coding activities progressively. It is believed that the proposed way of curriculum development paves a learning path which may drive students’ interest in coding and students could develop computational thinking skills progressively Siu Cheung Kong, Ping Li 0016 |
ICCE | 2 |
| 2016 | The Interest-Driven Creator Theory and Computational ThinkingabstractThere is a growing interest for Computational Thinking (CT) for the last decade. Most studies focus on teaching CT skills in K-12 level. In higher education, applying CT methods over all disciplines still needs cross institute movement, proper teaching tools and assessment procedure. In this paper we discuss the potential of applying CT methods in the mathematics courses at university level. With the inspiration of the Interest-Driven Creator (IDC) theory, we suggest that applying CT methods in mathematics benefits students’ understanding of concepts and overcomes the drawbacks of traditional pedagogy. We study the case of introducing the limit of a sequence, which is a fundamental concept in calculus. An algorithm, inspired by the ε-N definition, is designed to find suitable N given a specific ε with exhaustion methods. Based on the algorithm designed, a game that help to introduce the ε-N definition of the limit of a sequence is presented as an example. The game would be developed on mobile devices for easy accessibility and for catering the trend of mobile device. Bowen Liu 0015, Ping Li 0016, Siu Cheung Kong, Sing Kai Lo |
ICCE | 2 |
| 2016 | Density-enhanced perceptual mosaic on GPUabstractAbstract Image mosaic effects are wildly applied in print media, domestic decoration, and many image beautification applications. However, the current image mosaic methods are mostly based on fixed‐size image tiles, simple color adjustment, and irregular image segmentation, which are inaccurate and very time‐consuming. In this paper, we present a graphics processing unit‐accelerated perceptual mosaic using density tiles replacement and brightness lighting optimization, keeping original image structure details and providing more expressive visual effects. Automatic density replacement map segmentation and color‐based region tiles replacement are performed to facilitate the mosaic. Delicate brightness optimization and perceptual color correction are further applied to enhance expressive lighting effects. We also consider the salience perception of images and similarity correlation among neighboring tiles for our perceptual mosaic. The experimental results have shown the efficiency and high‐quality performance of our density‐enhanced perceptual mosaic on graphics processing unit. Copyright © 2016 John Wiley & Sons, Ltd. Ping Li 0016, Hanqiu Sun |
Comput. Animat. Virtual Worlds | 1 |
| 2016 | Accurate gaze tracking from single camera using gabor corner detector
Bin Sheng 0001, Wen Wu 0001, Lizhuang Ma, Ping Li 0016 |
Multim. Tools Appl. | 5 |
| 2015 | Enhanced Bilingual Text Analysis for BYOD with Hierarchical Visualization
Ping Li 0016, Siu Cheung Kong, Tak-Lam Wong, Chengwei Guo |
ICCE | 1 |
| 2015 | Vessel extraction from non-fluorescein fundus images using orientation-aware detector
Benjun Yin, Huating Li, Bin Sheng 0001, Xuhong Hou, Wen Wu 0001, Ping Li 0016, Ruimin Shen, Yuqian Bao, Weiping Jia |
Medical Image Anal. | 7 |
| 2014 | Detailed User Relation Visualization on MoodleabstractTo enhance the learners’ performance and learning efficiency on e-learning platform is an essential issue in current teaching and learning engagement. This paper applies IT techniques and analysis methods to visualize the learners’ user relationship on the discussion board of Moodle based on the replying order and times they made in the selected topics/forum in a Moodle course, which will be very useful in the future for embedding the visualization into large-scale systems and further development of related applications. We mainly first extract the inner relation of users’ participation in one/several selected discussion topics or a whole forum of a certain course. Then we visualize the user participation relation using Scalable Vector Graphics (SVG) provided by D3.js with aesthetic graph visualization. Our visualization with directed graph (digraph) structures provides easy understanding of users’ network connections in topics/forum, and the weights of the connections indicating the replying times are shown both numerically on edges and visually by the thickness of edges. Digraph is applied here to indicate different replying order. We propose to perform visualization construction and representation via PHP with latest advanced programming libraries of jQuery and D3.js, while easy control of tree-structured menu for user relation visualization is realized by jsTree. In the experiments, our method consistently demonstrates high-quality detailed visualization of users’ underlying relation structure they belong to, which may benefit the further study of improving the learning efficiency on e-learning platform. Ping Li 0016, Siu Cheung Kong |
ICCE | 1 |
| 2014 | Video Colorization Using Parallel Optimization in Feature SpaceabstractWe present a new scheme for video colorization using optimization in rotation-aware Gabor feature space. Most current methods of video colorization incur temporal artifacts and prohibitive processing costs, while this approach is designed in a spatiotemporal manner to preserve temporal coherence. The parallel implementation on graphics hardware is also facilitated to achieve realtime performance of color optimization. By adaptively clustering video frames and extending Gabor filtering to optical flow computation, we can achieve real-time color propagation within and between frames. Temporal coherence is further refined through user scribbles in video frames. The experimental results demonstrate that our proposed approach is efficient in producing high-quality colorized videos. Bin Sheng 0001, Hanqiu Sun, Marcus A. Magnor, Ping Li 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Object Movements Synopsis viaPart Assembling and StitchingabstractVideo synopsis aims at removing video's less important information, while preserving its key content for fast browsing, retrieving, or efficient storing. Previous video synopsis methods, including frame-based and object-based approaches that remove valueless whole frames or combine objects from time shots, cannot handle videos with redundancies existing in the movements of video object. In this paper, we present a novel part-based object movements synopsis method, which can effectively compress the redundant information of a moving video object and represent the synopsized object seamlessly. Our method works by part-based assembling and stitching. The object movement sequence is first divided into several part movement sequences. Then, we optimally assemble moving parts from different part sequences together to produce an initial synopsis result. The optimal assembling is formulated as a part movement assignment problem on a Markov Random Field (MRF), which guarantees the most important moving parts are selected while preserving both the spatial compatibility between assembled parts and the chronological order of parts. Finally, we present a non-linear spatiotemporal optimization formulation to stitch the assembled parts seamlessly, and achieve the final compact video object synopsis. The experiments on a variety of input video objects have demonstrated the effectiveness of the presented synopsis method. Yongwei Nie, Hanqiu Sun, Ping Li 0016, Chunxia Xiao, Kwan-Liu Ma |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2013 | Compact Video Synopsis via Global Spatiotemporal OptimizationabstractVideo synopsis aims at providing condensed representations of video data sets that can be easily captured from digital cameras nowadays, especially for daily surveillance videos. Previous work in video synopsis usually moves active objects along the time axis, which inevitably causes collisions among the moving objects if compressed much. In this paper, we propose a novel approach for compact video synopsis using a unified spatiotemporal optimization. Our approach globally shifts moving objects in both spatial and temporal domains, which shifting objects temporally to reduce the length of the video and shifting colliding objects spatially to avoid visible collision artifacts. Furthermore, using a multilevel patch relocation (MPR) method, the moving space of the original video is expanded into a compact background based on environmental content to fit with the shifted objects. The shifted objects are finally composited with the expanded moving space to obtain the high-quality video synopsis, which is more condensed while remaining free of collision artifacts. Our experimental results have shown that the compact video synopsis we produced can be browsed quickly, preserves relative spatiotemporal relationships, and avoids motion collisions. Yongwei Nie, Chunxia Xiao, Hanqiu Sun, Ping Li 0016 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2012 | Interactive image/video retexturing using GPU parallelism
Ping Li 0016, Hanqiu Sun, Jianbing Shen, Yongwei Nie |
Comput. Graph. | 1 |
| 2012 | Image stylization with enhanced structure on GPU
Ping Li 0016, Hanqiu Sun, Bin Sheng 0001, Jianbing Shen |
Sci. China Inf. Sci. | 1 |
| 2011 | Multi-keyframe abstraction from videosabstractThis paper presents a method for abstracting multi-keyframe from video datasets. Existing video abstraction methods focused on simple view videos, and the results will be unacceptable if applied to overlapping views directly due to limitations like unavoidable redundancy and complicated inner correlations. We propose a correlation map to naturally model the correlations with various attributes among multi-keyframe, keyframe importance and weighted correlations are then computed to construct the map. The weighted correlations, unlike the unweighted ones, not only model probabilistic relationship among keyframes but also address the temporal and visual similarity. We facilitate the abstraction process via SVM classification and keyframes reduction using rough set. The multi-keyframe correlation map, which serially assembles event-centered keyframes in temporal order, is presented for displaying the abstraction, which shows the correlations and improves the browsability of video datasets. Ping Li 0016, Yanwen Guo 0001, Hanqiu Sun |
ICIP | 1 |
| 2011 | Interactive soft-fabrics watering simulation on GPUabstractAbstract Physics‐based simulation is usually complex and time consuming, and consequently not suitable for real‐time applications. In this paper, we propose the efficient dynamics models for the real‐time simulation of soft fabrics interacting with water, including multi‐soaking effect and underwater dynamics. The multi‐soaking effect of soft fabrics is modeled based on the physics processes. Further, we develop the optimized mass–spring model that supports the large flow forces interacting with the fabrics underwater. The fabric spring forces are linearly derived and integrated with GPU–CUDA acceleration, feasible for real‐time VR applications with large set of fabric particles. Copyright © 2011 John Wiley & Sons, Ltd. Hanqiu Sun, Shiguang Liu, Ping Li 0016 |
Comput. Animat. Virtual Worlds | 4 |
| 2009 | Image-Based Material Restyling with Fast Non-local Means FilteringabstractThis paper presents a new GPU-based implementation of fast non-local means (NLM) filtering for material restyling. Our fast NLM filtering algorithm is able to achive realtime feedback of interactive image editing. Furthermore a novel material editing method based on our fast NLM filtering is proposed to change the material appearance of image-based objects. Given an input image, an alpha matte is created to differentiate the object from its background. After the automatic matting process, the object is removed from the background, and the 3D shape of the object is recovered. We use the fast non-local means (NLM) filtering to process the luminance channel of an image and obtain a pseudo-depth map that is sufficient for altering the material appearance of the observed object. We recover the gradient luminance maps for the region to be material-restyled; and change the material property in the region-of-interest by solving the Poisson equation. In this way, the new material is mapped onto the 3D shape, resulting in an object which appears to be made of a different material. Our new NLM filtering-based material editing is easy to implement in parallel with graphics hardware. The experimental results have demonstrated the satisfactory performance of our method. Bin Sheng 0001, Ping Li 0016, Hanqiu Sun |
ICIG | 2 |