EDBT 2026 Demo / reviewers in the wild / expert
Bin Sheng 0001
dblp:24/2408-1 · also Bing Sheng 0001
· DBLP profile ↗
244ranked-venue papers
20as first author
123since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 175 · 15 first-author · 83 since 2021Applied, interdisciplinary, general and emerging computing · 37 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 36 · 3 first-author · 27 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 2 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SA-Edit: Accelerating Editing Models via Test-time Spatial AccelerationabstractDiffusion-based image editing models have demonstrated remarkable capabilities for generating high-quality results. However, the iterative inference process poses a significant challenge in achieving real-time generation. Previously proposed methods, such as feature caching or model distillation, often require model-specific designs and lack flexibility. In this paper, we introduce SA-Edit, an efficient, training-free, and plug-and-play algorithm for the universal acceleration of diffusion-based image editing models. Specifically, we propose a spatial scaling strategy to reduce redundant latent tokens and enhance efficiency. To address aliasing and blurring artifacts, we introduce a score-based filter and adaptively refine high-score regions after each upsampling operation. Our method achieves at least 4.2 × faster inference for image editing while maintaining high output quality. Furthermore, our approach can be seamlessly integrated with existing acceleration techniques to achieve even greater speedups. Extensive experiments demonstrate the effectiveness and efficiency of our proposed method. The code is released at: https://github.com/ouroboros-phy/SA-Edit Yihao Song, Ran Yi 0002, Xiaoning Lei, Bin Sheng 0001 |
ICMR | 5 |
| 2026 | SynTaskNet: A synergistic multi-task network for joint segmentation and classification of small anatomical structures in ultrasound imaging
Abdulrhman H. Al-Jebrni, Saba Ghazanfar Ali, Bin Sheng 0001, Huating Li, Xiao Lin 0012, Ping Li 0016, Younhyun Jung, Jinman Kim, Lixin Jiang |
Comput. Vis. Image Underst. | 3 |
| 2026 | Slimmable neural architecture design based on cross architecture and token distillation
Guhao Qiu, Ping Li 0016, Bin Sheng 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Physical AI: Evolution, Progress, Challenges, and Prospects
Enhua Wu, You-Quan Liu, Tianchen Xu, Li-Xin Ren, Yi-Ming Qin, Ming-Yu Wei, Xiao-Wei He, Dong-Yan Yuan, Wen-Chao Hou, Zhi-Wei Ma, Bin Sheng 0001 |
J. Comput. Sci. Technol. | 11 |
| 2026 | Autorep: Automatic network search with structured reparameterized based linear operation expansion and gradient proxy guided reduction
Guhao Qiu, Ruoxin Chen, Ping Li 0016, Bin Sheng 0001 |
Neural Networks | 6 |
| 2026 | Dataset Distillation via a Noise-Unconstrained Generative ModelabstractDataset distillation (DD) aims to synthesize a more compact dataset than the original one and models trained on it are expected to have the same generalization capabilities as on the original dataset. Previous work via a generative model (GM) faces several limitations. First, GM struggles to generate representative samples due to a lack of constraints. Second, it overlooks the relationships between generated samples, limiting its effectiveness. In this paper, a new noise-unconstrained GM-based DD framework is proposed. In the distillation stage, an adaptive matching coefficient is introduced to align generated images with representative class elements and the MiniMax loss function is extended to reduce the optimization difficulty. In the deployment stage, features among each generative image are ensembled by gradient-matching based DD. Theoretical analysis based on McDiarmid's inequality demonstrates that the proposed components can reduce the generalization error of the original baseline method. We also provide insights into the potential of generated images as an effective proxy dataset for DD. For example, on the ImageWoof dataset with 50 distilled images per class using a 6-layer ConvNet for evaluation, generated images outperform 25%, 50%, and 75% original images by 8.4%, 6.3%, and 8.3% in distillation performance. Our method effectively handles both low- and high-resolution datasets, with experiments on 11 benchmarks demonstrating its efficacy. Fei Ye 0004, Ping Li 0016, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | DANIM: Domain adaptation network with intermediate domain masking for night-time scene parsing
Qijian Tian, Ran Yi 0002, Zufeng Zhang, Bin Sheng 0001, Xin Tan 0002, Lizhuang Ma |
Pattern Recognit. | 5 |
| 2026 | LODNeuS: A Flexible Lightweight Neural Implicit Surface Representation With Unconstrained Viewpoint RenderingabstractNeRF-like methods learn implicit 3D neural representations from 2D multiview images, enabling the synthesis of compelling novel views. However, to capture high-fidelity geometry, prior methods often rely on large-scale networks. This dependency hampers the potential applications of neural implicit representations, such as MR visualization. To address this, we introduce LODNeuS, an implicit surface representation based on feature voxel grids. LODNeuS captures multiple LODs of implicit geometry by maintaining voxel grids paired with a set of corresponding lightweight decoders. This allows for high-quality rendering with the ability to dynamically switch between detail levels. Another challenge is that existing methods, both volumetric and surface-based, tend to train and render their representations within a confined space, without explicitly restricting the sampling points properly. This lack of constraints can result in ambiguity, artifacts, and inefficient use of computational resources. We study this effect during free viewpoint rendering using conventional methods and develop an adaptive sampling scheme that emphasizes a valid geometric space for sampling point allocation. Our experimental results show that LODNeuS can match the visual quality of existing methods while offering flexible and lightweight inference. The benefits of adaptive sampling are also demonstrated in the free viewpoint rendering subsection. Our work extends the capabilities of neural implicit representations beyond previously defined limitations, broadening the scope of potential applications. Ping Li 0016, Lei Zhu 0003, Bin Sheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Boosting Video Object Segmentation With Discriminative Core Features and Adaptive Position Refinement
Yadang Chen, Guolong Li, Yuhui Zheng, Bin Sheng 0001, Zhi-Xin Yang 0001, Enhua Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Contrastive Decoupling: Dynamic Regularization for Enhanced Fine-Grained Image Classification
Zheyuan Wang, Tingyao Li, Bin Sheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | HPRNet: Human Parsing Reconstruction With Non-Local Multi-Scale Perception Network for Cloth-Changing Person Re-IdentificationabstractCloth-changing Person Re-Identification (CC-ReID) is a challenging data modeling task that involves identifying specific pedestrians wearing different outfits. Existing methods primarily focus on altering clothing color and directly reconstructing appearance to extract features independent of the clothes. Real pedestrians differ in height, body shape, etc. Such methods are prone to losing the intrinsic information of the original sample (i.e., the person identity) owing to the absence of contextual phenomena (e.g., texture structure and local correlation), which decreases the recognition performance. To address this problem, we propose a framework called HPRNet, or ”Human Parsing Reconstruction with Non-Local Multi-Scale Perception Network,” which includes a non-local weighted multi-scale perception (NWMP) module and a parsing reconstruction exploration (PRE) module. In particular, the proposed NWMP module effectively captures the global receptive field of a sample and obtains a contextual correlation between non-neighboring pixels within the sample image. The PRE module was used to achieve a more accurate reconstruction of human body components with a clothing parsing model to better distinguish features related to or unrelated to clothes. Extensive experiments were conducted on CC-ReID public datasets (LTCC, PRCC, and CCVID) to demonstrate the effectiveness and competitiveness of the proposed method with state-of-the-art (SOTA) baselines for this complex modeling task. Mingfu Xiong, Longlong Ge, Ruimin Hu, Khan Muhammad 0001, Sambit Bakshi, Javier Del Ser, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | DSDFormer: An Innovative Transformer-Mamba Framework for Robust High-Precision Driver Distraction IdentificationabstractDriver distraction remains a leading cause of traffic accidents, posing a critical threat to road safety globally. As intelligent transportation systems evolve, accurate and real-time identification of driver distraction has become essential. However, existing methods struggle to capture both global contextual and fine-grained local features while contending with noisy labels in training datasets. To address these challenges, we propose DSDFormer, a novel framework that integrates the strengths of Transformer and Mamba architectures through a Dual State Domain Attention (DSDA) mechanism, enabling a balance between long-range dependencies and detailed feature extraction for robust driver behavior recognition. Additionally, we introduce Temporal Reasoning Confident Learning (TRCL), an unsupervised approach that refines noisy labels by leveraging spatiotemporal correlations in video sequences. Beyond achieving state-of-the-art results on AUC-V1, AUC-V2, and 100-Driver datasets, the proposed model is deployable in real-time on embedded platforms such as NVIDIA Jetson AGX Orin and Xavier. Extensive experimental results confirm that DSDFormer and TRCL significantly improve both the accuracy and robustness of driver distraction detection, offering a scalable solution to enhance road safety. Our code has been released athttps://github.com/zhangzr23/driver-noises-learning Junzhou Chen 0001, Heqiang Huang, Xuemiao Xu, Bin Sheng 0001, Hong Yan 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2026 | M2Net: Multimodal Multitask Mutual Learning for Anti-VEGF Efficacy PredictionabstractAge-related macular degeneration with abnormal blood vessel growth (neovascular AMD) is the leading cause of vision loss in elderly populations. While anti-VEGF injections are the standard treatment, they present financial burdens for patients and vary in effectiveness. Predicting treatment efficacy is therefore crucial for patient care. Current prediction methods fail to fully integrate information from different imaging techniques, typically focusing on either forecasting vision improvements or generating post-treatment images-but not both simultaneously. This approach overlooks the important relationship between these tasks. We present M2Net, a novel joint generation and classification network based on Multimodal Multitask Mutual learning, to simultaneously predict changes in visual acuity and generate post-treatment retinal images. M2Net employs a dual-branch structure that processes both fundus photographs and Optical Coherence Tomography (OCT) scans to improve prediction accuracy. Our framework includes two key innovations: the Multimodal Collaborative Treatment Efficacy Prediction module, which interacts the features between the two modalities and provides initial visual acuity change classification to guide the generation of post-treatment images; and the Pre-Post Treatment Image Joint Analysis module, which identifies both common and changing features between pre-treatment and post-treatment images to enhance prediction accuracy. To validate our approach, we created the dataset (MMPD) containing paired multimodal retinal images with corresponding visual acuity measurements. Experiments on the dataset demonstrate that M2Net achieves superior performance compared to existing methods, with a classification accuracy of 96.03%, an SSIM of 0.6377 on the OCT modality, and an SSIM of 0.8347 on the fundus modality. Our code will be available at https://github.com/zengying123/M2Net. Lei Bi 0001, Wuzhen Shi, Huazhu Fu, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2026 | Fine-Grained Lexical-Centric Semantic Network for Coherent Video Paragraph CaptioningabstractVideo paragraph captioning (VPC) aims to generate coherent, detailed narratives that accurately reflect a video's content. However, existing methods typically depend on coarse-grained event correlations and neglect the nuanced spatio-temporal interactions critical for comprehensive understanding. Refined verbs and prepositions, encoding actions and spatial relations, are essential for clear, consistent descriptions. To address these issues, we propose the Fine-Grained Lexical-Centric Semantic Network (FLS-Net), which emphasizes verbs and prepositions linked to salient objects to improve spatio-temporal coherence across events. FLS-Net integrates a multi-lexical synergy mechanism, leveraging nouns obtained via multi-modal matching, and employs a Verb-Guided Event Consistency Module (VECM) alongside a Preposition-Driven Relation Representation Module (PRRM). A cyclic encoder-decoder architecture further enforces event consistency, significantly boosting VPC performance. Extensive experiments onActivityNet CaptionsandYouCook2demonstrate FLS-Net's superiority over state-of-the-art approaches. The source code is available athttps://github.com/yangxingrui/FLS. Shuqin Chen, Xian Zhong, Xingrui Yang 0003, Bin Sheng 0001, Alex Chichung Kot |
IEEE Trans. Multim. | 5 |
| 2026 | Spatio-Temporal Disentanglement and Constrained Self-Attention for Multi-Modal Deception DetectionabstractMulti-modal deception detection is a challenging yet important task, having pivotal applications in many fields such as business credibility assessment and multimedia anti-frauds. Previous methods either rely solely on spatial features or overemphasize only temporal information within or across modalities, which may overlook potential critical clues. Motivated by these observations, we propose a Spatio-Temporal Representation Disentanglement (STRD) framework for multi-modal deception detection, which uses a dual-encoder structure to learn spatial and temporal representations for each modality. Specifically, we introduce a pre-trained foundation model to act as the spatial encoder and design a lightweight network as the temporal encoder, extracting spatial semantics and capturing dynamic temporal patterns. Then, we propose a Constrained Self-Attention Block (CSAB), in which self-attention distribution of each head is regarded as spatial distribution and is constrained to attend a certain facial local region. Furthermore, we present a Cross-Modal Correlation Fusion Block (CCFB) to achieve temporal synchronization across modalities by measuring the correlations between visual and audio features. Extensive experiments show that our STRD outperforms the state-of-the-art methods on challenging DOLOS, BOL, BgOL, and RLtrial benchmarks. Particularly, STRD improves by 2.12% and 1.88% over the previous best results in terms of ACC on the DOLOS and BOL datasets, respectively. Additionally, STRD outperforms previous methods in cross-dataset testing, highlighting its superior generalization ability. Zhiwen Shao, Hancheng Zhu, Rui Yao 0006, Lixin Zou, Mengtian Li 0002, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2026 | Multimodal RAG for financial documents: BART-based financial named entity recognition and attention-based table parsing for financial QA enhancement
Ying Ni, Hanghang Peng, Pengle Zhang, Bin Sheng 0001 |
Vis. Comput. | 6 |
| 2025 | Gradient amplification for gradient matching based dataset distillation
Ping Li 0016, Bin Sheng 0001 |
Neural Networks | 5 |
| 2025 | Non-Rigid Point Cloud Registration via Anisotropic Hybrid Field HarmonizationabstractCurrent point cloud registration algorithms struggle to effectively handle both deformations and occlusions simultaneously. Our manifold analysis reveals this limitation arises from the inaccurate modeling of the shape's underlying manifold and the lack of an effective optimization strategy for fragmented manifold structures. In this paper, we present AniSym-Net, a novel non-rigid registration framework designed to address near-isometric deformation registration in the presence of occlusions. To encode object's coarse topological properties and local geometric information, AniSym-Net introduces a novel anisotropic hybrid shape-motion deformation field. The effectiveness of the anisotropic hybrid shape-motion fields relies on both the holonomic constraints from the symplectic structure modeling in AniSym-Net and the motion-conditional cross-attention during fusion, which calibrates geometric features using velocity-boundary constrained point motion patterns. The harmonization of correspondences derived from anisotropic hybrid fields and those from motion-shape fields significantly mitigates registration errors and occlusions. This is achieved through the optimization of loop closures of cotangent bundles within the symplectic manifold framework. We conduct comprehensive evaluation across five popular benchmarks, namely CAPE, DT4D, SAPIEN, FAUST, and DeepDeform, to demonstrate our AniSym-Net's superior performance compared to the state-of-the-art methods. Code will be publicly available. Xuequan Lu, Mohammed Bennamoun, Bin Sheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Semi-Supervised Privacy-Preserving EEG-Based Motor Imagery Classification via Self and Adversarial TrainingabstractElectroencephalogram (EEG)-based motor imagery (MI) signals are frequently used in brain-computer interfaces (BCIs) due to their wide applications in the rehabilitation field. However, cross-subject variations often result in a model trained on one participant failing when applied to another. Additionally, privacy concerns regarding sensitive health and mental information in EEG-based MI signals further complicate the situation. Source-free domain adaptation aims to address these cross-subject variations by transferring knowledge from a source domain (i.e., a previous participant) to a target domain (i.e., a new participant) without accessing sensitive source data. However, source-free unsupervised domain adaptation models often face issues with incorrect pseudo-labels, which can lead to unstable and ineffective adaptation. To address this, we propose a source-free semi-supervised domain adaptation algorithm for EEG-based MI signal classification. This algorithm tackles noise accumulation caused by incorrect pseudo-labels while effectively handling data distribution variations and privacy concerns, similar to source-free unsupervised domain adaptation models. Specifically, we train the classifier head using only a limited amount of labeled target data to prevent noise accumulation, and generate pseudo-labels for the unlabeled target data. Furthermore, we introduce an independent self-training head that learns better representations using the generated pseudo-labels, mitigating overfitting caused by the limited labeled target data. Additionally, we design an adversarial head that plays a minimax game to extract more discriminative feature representations from the unlabeled target data. Extensive experiments on three benchmark datasets, compared with eighteen state-of-the-art SFDA methods, demonstrate the superiority of our approach. Jian Zhu 0001, Ganxi Xu, Zhizhe Lin, Jinyi Long, Teng Zhou, Bin Sheng 0001, Xiaokang Yang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | MHFNet: A Multimodal Hybrid-Embedding Fusion Network for Automatic Sleep StagingabstractScoring sleep stages is essential for evaluating the status of sleep continuity and comprehending its structure. Despite previous attempts, automating sleep scoring remains challenging. First, most existing works did not fuse local and global temporal information. Second, the correlation for special waves in different signals is rarely used in sleep staging modeling. Third, the logic of scoring rules based on adjacent epochs is not considered in developing sleep staging models. This paper introduces a multimodal hybrid-embedding fusion network (MHFNet), which aims to tackle these challenges in automating sleep stage scoring. MHFNet comprises multi-stream Xception blocks to extract wave characteristics, a hybrid time-embedding module to combine local and global temporal information, a dual-path gate transformer to fuse and enhance attention features, and a refined output header to reconstruct sleep scoring. We perform experiments using three publicly available datasets (SleepEDF-ST, SleepEDF-SC, and SHHS). Experimental results indicate the superiority of MHFNet over baseline approaches in cross-validation. Moreover, at the individual level, MHFNet yielded an average $R^{2}$ score improvement of 9$\%$ in the testing dataset compared to state-of-the-art models, paving the way for its applications in real-world sleep medicine. Ruhan Liu, Jiajia Li 0004, Bin Sheng 0001, David Dagan Feng, Ping Zhang 0016 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Serp-Mamba: Advancing High-Resolution Retinal Vessel Segmentation With Selective State-Space ModelabstractUltra-Wide-Field Scanning Laser Ophthalmoscopy (UWF-SLO) images capture high-resolution views of the retina with typically spanning 200 degrees. Accurate segmentation of vessels in UWF-SLO images is essential for detecting and diagnosing fundus disease. Recent studies highlight that Mamba's selective State Space Model (SSM) excels in modeling long-range dependencies with linear computational complexity, making it highly suitable for preserving the continuity of elongated vessel structures, especially for high-resolution UWF images. Inspired by this, we propose the Serpentine Mamba (Serp-Mamba) network to address this challenging task. Specifically, we recognize the intricate, varied, and delicate nature of the tubular structure of vessels. Furthermore, the high-resolution of UWF-SLO images exacerbates the imbalance between the vessel and background categories. Based on the above observations, we first devise a Serpentine Interwoven Adaptive (SIA) scan mechanism, which scans UWF-SLO images along curved vessel structures in a snake-like crawling manner. This approach, consistent with vascular texture transformations, ensures the effective and continuous capture of curved vascular structure features. Second, we propose an Ambiguity-Driven Dual Recalibration (ADDR) module to address the category imbalance problem intensified by high-resolution images. Our ADDR module delineates pixels by two learnable thresholds and refines ambiguous pixels through a dual-driven strategy, thereby accurately distinguishing vessels and background regions. Experiment results on three datasets demonstrate the superior performance of our Serp-Mamba on high-resolution vessel segmentation. We also conduct a series of ablation studies to verify the impact of our designs. Our code will be released upon publication (https://github.com/whq-xxh/Serp-Mamba). Hongqiu Wang, Bin Sheng 0001, Huazhu Fu, Guang Yang 0006, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Identity and Modality Attributes Driven Multimodal Fusion Networks for Emotion Recognition in ConversationsabstractEmotion recognition in conversations (ERC) is a crucial aspect of human-computer interaction and plays an important role in various domains, including healthcare, entertainment, and education. Since the conversation data in the form of multimodal sequences is well suited to be constructed into graphs, the methods based on graph convolutional network (GCN) show incomparable advantages. However, existing methods attempt to model the highly uncertain emotional relationships between different speakers, which is not an easy task and may even introduce interference information. Therefore, we propose an identity and modality attributes driven multimodality fusion network (dubbed IMDNet) for emotion recognition in conversations. Specifically, we construct a speaker-centric graph that only connects nodes of the same speaker within modalities to each other, reducing the interference between the emotions of different speakers. We also introduce the attribute embedding mechanism, which facilitates the correct calculation of correlations between nodes for better multimodal feature fusion. Considering that the emotional correlation between utterances will decrease over time, we present an utterance distance attention to make the fusion network pay more attention to the adjacent utterances. Furthermore, we explore the solution to the data imbalance problem suitable for conversation scenarios. Given the presence of possible anomalous samples in the dataset, we opt for the BoundaryFocalLoss. Experiments on the IEMOCAP and MELD datasets show that our IMDNet outperforms the state-of-the-art methods. Wuzhen Shi, Xuping Chen, Biyun Yao, Bin Sheng 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | SAT-Net: Structure-Aware Transformer-Based Attention Fusion Network for Low-Quality Retinal FunduImages EnhancementabstractIn ophthalmology diagnosis, high-fidelity fundus images are essential for disease diagnosis and intervention. However, many real-world clinical conditions may degrade the quality of the acquired images and thus affect clinical diagnostic accuracy. Traditional convolutional neural network-based retinal fundus image enhancement methods cannot always capture long-range dependencies, which reduces the overall visual quality of images, especially for real retinal fundus images. Furthermore, existing enhancement methods often fail to fully utilize low-resolution structural detail information, which potentially leads to inaccurate pivotal fundus vessel topology or capillary details. In this paper, we propose a novel Structure-Aware Transformer-based attention fusion Network (SAT-Net) for low-quality retinal fundus image enhancement. First, we introduce a Transformer-based attention fusion module which incorporates window-based self-attention and channel self-attention to capture global spatial dependencies and emphasize important feature channels simultaneously. This fusion significantly improves the overall perceptual quality of the image by enhancing both the local details and the uniformity of the non-vessel background regions. Second, we introduce a cross-quality knowledge distillation technique, which bridges the quality gap between high-quality and low-quality fundus images. By designing a high-performing teacher network to guide a lightweight student network, the student network enables to capture detailed features from low-quality fundus images, further preserving critical diagnostic information and fine topology structures. Moreover, we design a structure-aware multi-scale loss function by using a trainable subnetwork to obtain the edge structure from different scales to better constrain pivotal fundus vessel structure and capillary details. Comprehensive quantitative and qualitative experiments on both synthetic and real fundus image datasets robustly validate that our proposed SAT-Net outperforms other state-of-the-art methods for fundus image enhancement. In addition, extensive comparative experiments on both the vessel segmentation and Optic Disc/Cup detection tasks further validate the effectiveness and superiority of our proposed method. Wuzhen Shi, Jianhua Ji, Wenming Cao 0001, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Adaptive Clustering and Weighted Regularization Contrastive Learning Framework for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (ReID) has recently gained significant attention from researchers. ReID matches images of the same person from different camera views in various scenes without any labels. Existing clustering methods primarily rely on a fixed threshold (the maximum distance between sample points and clustering centroids) and overlook the importance of adjusting this threshold during continuous model optimization. This mismatch between clustering thresholds and inter- or intra-class spacing reduces clustering accuracy. To address this issue, this study proposes an Adaptive Clustering and Weighted Regularization Contrastive Learning (ACWRCL) framework for unsupervised person ReID. The ACWRCL framework comprises two main components: (1) the Clustering Threshold Adaptive Adjustment (CTAA) module, and (2) the Weighted Regularization Contrastive Learning (WRCL) module. The CTAA module dynamically adjusts the clustering threshold to align with model optimization, ensuring that the threshold remains within an appropriate range to prevent under- or over-robustness in the clustering model. The WRCL module uses the similarity ratio between the query sample and the clustering centroid relative to the overall similarity of all samples with the same labels as the query sample. This ratio is used as the weight in the loss function to penalize incorrect clustering and improve pseudo-label generation accuracy. Extensive experiments on public ReID datasets—Market-1501, MSMT17, Veri776, CUHK03, and PersonX—demonstrate the effectiveness of the proposed method. Mingfu Xiong, Kaikang Hu, Zhongyuan Wang 0001, Ruimin Hu, Khan Muhammad 0001, Javier Del Ser, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Multim. | 8 |
| 2025 | SGG-Nets: Generic Rotation-Invariant Plugin Networks for Point Cloud AnalysisabstractRotation invariance is a crucial requirement for the analysis of 3D point clouds. However, current methods often achieve rotation invariance by employing specific network designs. These networks, though perform well on rotation-aware tasks, is inferior in general tasks such as classification and segmentation. On the other hand, many powerful point processing networks, such as PointNet++, DGCNN, etc., have general point processing abilities, but do not own the property of rotation invariance. In this paper, we propose a standalone rotation-invariant convolution operator called SGGConv (Spherical Geometric Graph-based Convolution) and two ways integrating it with common point-based networks. The networks equipped with SGGConvs are called SGG-Nets which promote the rotation-invariance ability of regular point networks without modifying their network architectures much. Our contributions are three-fold. First, we propose a rotation-invariant feature descriptor, namely Spherical Geometry Descriptor (SGD), which captures point-pair features in a Local Spherical Coordinate System (LSCS). Second, we propose the SGGConv based on SGD and LSCS with an efficient Graph-based Spherical Feature Passing (GSFP) mechanism. Thirdly, we define two modules S-SGGConvMdl and M-SGGConvMdl, which are used to integrate SGGConv into baseline point nets. We test SGG-Nets, such as SGG-PointNet++, SGG-DGCNN, SGG-RIConv++, on representative point cloud datasets. These models, equipped with our SGGConvs, not only enhance the rotation-invariance of the baseline network but also improve its performance on point cloud analysis tasks such as classification and part segmentation, without incurring too much computational overhead. Jian Zhu 0001, Jianrong Yan, Jiebin Huang, Yongwei Nie, Bin Sheng 0001, Tong-Yee Lee |
IEEE Trans. Multim. | 5 |
| 2025 | Temporal-Interim Pose Synthesis and Distillation for Dynamic Human Pose EstimationabstractIn the task of dynamic human pose estimation (dynamic HPE), the temporal relationships between human body parts should be captured comprehensively to understand the dynamic human motions, where the correlated motion information eventually helps to recognize body parts. The popular methods are successful in terms of utilizing long-term motion information captured by low-speed cameras. Yet they neglect the underlying intermediate motions between captured frames, which comprise the temporal-interim poses lost in the video. In this article, we introduce a novel framework, temporal-interim pose synthesis and distillation, to produce and leverage the intermediate motion information for dynamic motion establishment. The pose synthesis yields the visual feature maps of the intermediate poses, which appear between the existing video frames. It allows the synthesized and current poses to form richer motion patterns. Next, the pose distillation divides the body parts into several groups, where it learns the specific part-wise relationship within each group. It degrades the complexity of learning useful part-wise relationships from rich motion patterns and extracts more detailed motion information for fine-grained part groups. We extensively evaluate our method on challenging datasets for dynamic pose estimation, achieving state-of-the-artresults. Di Lin 0002, Xin Wang 0118, Bin Sheng 0001, George Baciu, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | A Lightweight Depthwise Separable ConvNet with Frequency-domain Enhancement for Retinal Vessel SegmentationabstractAutomatic retinal vessel segmentation is crucial in the diagnosis and treatment of various cardiovascular and eye diseases. Although current vessel segmentation methods have achieved impressive performance, some challenging issues still need to be addressed. For example, existing methods always cannot segment complex capillaries well because they may be interfered with or covered by other components in the retina, and they need to further improve the continuity and consistency of vessel segmentation results. Moreover, the excellent vessel segmentation methods are usually built on bulky and cumbersome models which greatly limit their application range. In this article, we propose a novel efficient depthwise separable convolution network with frequency-domain enhancement (dubbed RetiNeXt) for retinal vessel segmentation. Firstly, we design a lightweight vessel enhancement module to extract global fine topological structure features from the frequency domain to enhance the complex capillary vessel details. Secondly, we propose a global feature extraction block to fully capture the large-scale spatial information and global characterizations, which enables the model to maintain vessel structural coherence from a global perspective. Thirdly, we construct a local feature mixing block based on SimAM attention mechanism to highlight the tiny capillary topological structure features and optimize the segmentation of low-contrast blood vessels, thereby improving the integrity and continuity of complex capillaries. Comprehensive comparison experiments on three well-benchmarked retinal vessel segmentation datasets fully verify the effectiveness and superiority of the proposed RetiNeXt. To further demonstrate the universality of RetiNeXt for medical image segmentation, we also conduct sufficient comparative experiments on two classical coronary angiography datasets. Extensive quantitative and qualitative experiments fully show that RetiNeXt outperforms other state-of-the-art methods with only 0.4M of trainable parameters. Shunzhe Shen, Wuzhen Shi, Wenming Cao 0001, Lei Bi 0001, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | CCM-Net: Contrastive and Consistent Multi-Task Network for Artifact Segmentation and Quality Classification of OCTA ImagesabstractArtifacts are prevalent in Optical Coherence Tomography Angiography (OCTA) images, which probably interfere doctor’s diagnosis and greatly limit its utility. Therefore, it is desirable to segment artifacts and assess quality when using them for diagnosis. In this article, we propose an end-to-end network (named CCM-Net: C ontrastive and C onsistent M ulti-task Network) to jointly address artifact segmentation and quality classification of OCTA images. We first devise multiple Task-Specific Attention Blocks to integrate deep features at different CNN layers for segmenting artifacts and classifying the quality of the input OCTA image. In this way, the weights of different deep features can be automatically learned and are not the same for the two tasks. Moreover, we devise a contrastive loss and a consistency loss to leverage sample relations for further enhancing prediction accuracy. Specifically, given an input OCTA image, we first augment it with a color jitter and select another OCTA image with the same quality classification label. We then design a contrastive loss so that the segmentation results of the input OCTA image are similar to its enhanced OCTA image, while the segmentation results of the two selected OCTA images are not similar. Besides, we devise a consistency loss on the classification results of the three images, because we can find that these images have the same quality classification labels. Experiments on an in-house OCTA dataset (Multi-OCTA) demonstrate that the proposed CCM-Net outperforms state-of-the-art methods. Xiang-Ning Wang, Jixue Tang, Ping Li 0016, Lei Zhu 0003, Harry Qin, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2025 | MSEmbGAN: Multi-Stitch Embroidery Synthesis via Region-Aware Texture GenerationabstractConvolutional neural networks (CNNs) are widely used for embroidery feature synthesis from images. However, they are still unable to predict diverse stitch types, which makes it difficult for the CNNs to effectively extract stitch features. In this paper, we propose a multi-stitch embroidery generative adversarial network (MSEmbGAN) that uses a region-aware texture generation sub-network to predict diverse embroidery features from images. To the best of our knowledge, our work is the first CNN-based generative adversarial network to succeed in this task. Our region-aware texture generation sub-network detects multiple regions in the input image using a stitch classifier and generates a stitch texture for each region based on its shape features. We also propose a colorization network with a color feature extractor, which helps achieve full image color consistency by requiring the color attributes of the output to closely resemble the input image. Because of the current lack of labeled embroidery image datasets, we provide a new multi-stitch embroidery dataset that is annotated with three single-stitch types and one multi-stitch type. Our dataset, which includes more than 30K high-quality multi-stitch embroidery images, more than 13K aligned content-embroidered images, and more than 17K unaligned images, is currently the largest embroidery dataset accessible, as far as we know. Quantitative and qualitative experimental results, including a qualitative user study, show that our MSEmbGAN outperforms current state-of-the-art embroidery synthesis and style-transfer methods on all evaluation indicators. Xinrong Hu, Ping Li 0016, Bin Sheng 0001, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | EGDNet: an efficient glomerular detection network for multiple anomalous pathological feature in glomerulonephritis
Saba Ghazanfar Ali, Ping Li 0016, Huating Li, Po Yang 0001, Younhyun Jung, Harry Qin, Jinman Kim, Bin Sheng 0001 |
Vis. Comput. | 9 |
| 2025 | DSTS-GF: a dual-stream temporal-spatial transformer with gated fusion for the classification of Obstructive Sleep Apnea
Yuanqi Yao, Zhouyu Guan, Jun Pu, Ruhan Liu, Bin Sheng 0001, Shankai Yin |
Vis. Comput. | 8 |
| 2025 | Multimodal and multi-time-point fusion approach for automated diagnosis and grading of carotid atherosclerosis using bilateral ultrasound images and metadata
Pinqi Fang, Dong Lang, Zhouyu Guan, Yiting Wu, Yulian Zhang, Yuqian Bao, Huating Li, Chengxing Shen, Jun Pu, Bin Sheng 0001 |
Vis. Comput. | 12 |
| 2025 | Deep contour attention learning for scleral deformation from OCT images
Hao Chen 0011, Yupeng Xu, Huating Li, Yuan Xie 0006, David Dagan Feng, Jinman Kim, Lei Bi 0001, Xiangui He, Bin Sheng 0001 |
Vis. Comput. | 12 |
| 2025 | HRDC challenge: a public benchmark for hypertension and hypertensive retinopathy classification from fundus images
Xiangning Wang, Zhouyu Guan, An-ran Ran, Tingyao Li, Zheyuan Wang, Xinming Shu, Jinyang Xie, Shichang Liu, Guanyu Xing, Julio Silva-Rodríguez, Riadh Kobbi, Ping Li 0016, Tingli Chen, Lei Bi 0001, Jinman Kim, Weiping Jia, Huating Li, Harry Qin, Ping Zhang 0016, Ching Yu Cheng, Pheng-Ann Heng, Tien Yin Wong, Carol Y. Cheung, Nadia Magnenat-Thalmann, Bin Sheng 0001 |
Vis. Comput. | 29 |
| 2025 | Temporal goal-aware transformer assisted visual reinforcement learning for virtual table tennis agent
Haoxuan Li 0004, Xiaojun Huang, Weibing Wu, Bin Sheng 0001 |
Vis. Comput. | 8 |
| 2025 | Dost: a dual optimization method for text-guided face images style transfer
Ran Yi 0002, Bin Sheng 0001 |
Vis. Comput. | 3 |
| 2025 | GAMNet: a gated attention mechanism network for grading myopic traction maculopathy in OCT images
Tingyao Li, Shiqun Lin, Bin Sheng 0001, Ruhan Liu, Rongping Dai |
Vis. Comput. | 5 |
| 2025 | Urgent needs, opportunities and challenges of virtual reality in healthcare and medicine in the era of large language modelsabstractThe convergence of large language models (LLMs) and virtual reality (VR) technologies has led to significant breakthroughs across multiple domains, particularly in healthcare and medicine. Owing to its immersive and interactive capabilities, VR technology has demonstrated exceptional utility in surgical simulation, rehabilitation, physical therapy, mental health, and psychological treatment. By creating highly realistic and precisely controlled environments, VR not only enhances the efficiency of medical training but also enables personalized therapeutic approaches for patients. The convergence of LLMs and VR extends the potential of both technologies. LLM-empowered VR can transform medical education through interactive learning platforms and address complex healthcare challenges using comprehensive solutions. This convergence enhances the quality of training, decision-making, and patient engagement, paving the way for innovative healthcare delivery. This study aims to comprehensively review the current applications, research advancements, and challenges associated with these two technologies in healthcare and medicine. The rapid evolution of these technologies is driving the healthcare industry toward greater intelligence and precision, establishing them as critical forces in the transformation of modern medicine. Xinming Xu, Haoxuan Li 0004, Zhouyu Guan, Dian Zeng, Qingqing Zheng, Huating Li, Chwee Teck Lim, Tien Yin Wong, Enhua Wu, Weiping Jia, Bin Sheng 0001 |
Virtual Real. Intell. Hardw. | 13 |
| 2024 | Text2City: One-Stage Text-Driven Urban Layout RegenerationabstractRegenerating urban layout is an essential process for urban regeneration. In this paper, we propose a new task called text-driven urban layout regeneration, which provides an intuitive input modal - text - for users to specify the regeneration, instead of designing complex rules. Given the target region to be regenerated, we propose a one-stage text-driven urban layout regeneration model, Text2City, to jointly and progressively regenerate the urban layout (i.e., road and building layouts) based on textual layout descriptions and surrounding context (i.e., urban layouts and functions of the surrounding regions). Text2City first extracts road and building attributes from the textual layout description to guide the regeneration. It includes a novel one-stage joint regenerator network based on the conditioned denoising diffusion probabilistic models (DDPMs) and prior knowledge exchange. To harmonize the regenerated layouts through joint optimization, we propose the interactive & enhanced guidance module for self-enhancement and prior knowledge exchange between road and building layouts during the regeneration. We also design a series of constraints from attribute-, geometry- and pixel-levels to ensure rational urban layout generation. To train our model, we build a large-scale dataset containing urban layouts and layout descriptions, covering 147K regions. Qualitative and quantitative evaluations show that our proposed method outperforms the baseline methods in regenerating desirable urban layouts that meet the textual descriptions. Nanxuan Zhao, Bin Sheng 0001, Rynson W. H. Lau |
AAAI | 3 |
| 2024 | MSCE-LT: Multi-Label Supervised Contrastive Enhancement for Long-Tailed Retinal Diseases RecognitionabstractRetinal diseases are leading causes of blindness globally. In real-world clinical practice, a patient may suffer from multiple retinal diseases, and these diseases are often under a long-tailed distribution, which poses significant challenges for accurate diagnosis. In this work, we propose a novel contrastive learning(CL)-based framework for multi-label retinal disease recognition. It consists of two parallel branches, a multi-label supervised contrastive learning branch and a classifier branch. The positive sets are created by the extent of proportional label overlap between samples and the anchor in calculating contrastive loss. For minority information enhancement, we design a hybrid-proxy model to generate class-dependent proxies, which are updated alongside the network. We capture rich relations samples, proxies, and labels by introducing the hybrid sample-proxy contrastive loss. Taking both label co-occurrence and data imbalance into consideration, we further utilize Distribution Balanced (DB) binary cross-entropy loss to guide the classifier branch learning. Experimental results on four public retinal disease datasets have demonstrated the superiority and effectiveness of our method. Tingyao Li, Bin Sheng 0001 |
BIBM | 2 |
| 2024 | Context-Aware Transformer for Single Image Rain Streaks RemovalabstractDeep learning based image deraining has been widely researched. However, rain streaks are hard to differentiate with similar textures of background without context knowledge. In this paper, a novel Context-Aware Transformer (CAT) is proposed for single image deraining where both local and global context within the input rainy image are utilized for better background reconstruction performance. The proposed CAT perceives a comprehensive context view through efficient self-attention mechanism and dilated convolutions in the Context-Aware Transformer Block (CATB). The Rain-Aware Feature Selection module (RAFS) generates feature blending coefficients adaptively to filter out rain streaks components and preserves clear background in hierarchical features of CAT. Meanwhile, a High-Frequency Preserved Loss (HFPL) provides further supervision on training and promotes reserving clearer structures and sharper details. Experiments on synthesized and real-world benchmarks illustrate the outstanding performance over state-of-the-art methods and pleasing visual results in various scenes. Lei Liang 0003, Yeting Huang, Bin Sheng 0001 |
ICASSP | 6 |
| 2024 | Self-Paced Co-Training and Foundation Model for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning is an effective approach for image segmentation, especially in medical images where segmentation labels are scarce and require expertise. Currently, there is still much room for improvement in these existing semi-supervised methods. In this paper, we propose a semi-supervised learning method for image segmentation with self-paced learning and foundation model. We first synthesize virtual unlabeled images by mixing the original unlabeled images in training set with normal images without lesions through mixup operation. Then we design a novel self-paced co-training strategy that constrains the different networks to learn consistent prediction distributions for the paired virtual unlabeled images that are synthesized from different normal images but the same unlabeled images. In addition, by incorporating both pixel-level and image-level importance into the self-paced learning, the designed approach allows the network to learn both pixel-level and image-level features from easy to hard. We perform the experiment using the pre-training foundation model as the network architecture. The experimental results show the effectiveness of our method in semi-supervised segmentation and can further improve the performance of the foundation model. Bin Sheng 0001 |
ICME | 3 |
| 2024 | 3DPX: Progressive 2D-to-3D Oral Image Reconstruction with Hybrid MLP-CNN Networks
Xiaoshuang Li, Mingyuan Meng, Zimo Huang, Lei Bi 0001, Eduardo Delamare, David Dagan Feng, Bin Sheng 0001, Jinman Kim |
MICCAI (7) | 7 |
| 2024 | SSM-Net: Semi-supervised multi-task network for joint lesion segmentation and classification from pancreatic EUS images
Jiajia Li 0004, Lei Zhu 0003, Ping Zhang 0016, Ruhan Liu, Bin Sheng 0001 |
Artif. Intell. Medicine | 8 |
| 2024 | UrbanEvolver: Function-Aware Urban Layout Regeneration
Nanxuan Zhao, Jiale Yang, Siyuan Pan, Bin Sheng 0001, Rynson W. H. Lau |
Int. J. Comput. Vis. | 5 |
| 2024 | A Transfer Function Design for Medical Volume Data Using a Knowledge Database Based on Deep Image and Primitive Intensity Profile Features Retrieval
Younhyun Jung, Jim Kong, Bin Sheng 0001, Jinman Kim |
J. Comput. Sci. Technol. | 3 |
| 2024 | Soccer match broadcast video analysis method based on detection and trackingabstractAbstract We propose a comprehensive soccer match video analysis pipeline tailored for broadcast footage, which encompasses three pivotal stages: soccer field localization, player tracking, and soccer ball detection. Firstly, we introduce sports camera calibration to seamlessly map soccer field images from match videos onto a standardized two‐dimensional soccer field template. This addresses the challenge of consistent analysis across video frames amid continuous camera angle changes. Secondly, given challenges such as occlusions, high‐speed movements, and dynamic camera perspectives, obtaining accurate position data for players and the soccer ball is non‐trivial. To mitigate this, we curate a large‐scale, high‐precision soccer ball detection dataset and devise a robust detection model, which achieved the of 80.9%. Additionally, we develop a high‐speed, efficient, and lightweight tracking model to ensure precise player tracking. Through the integration of these modules, our pipeline focuses on real‐time analysis of the current camera lens content during matches, facilitating rapid and accurate computation and analysis while offering intuitive visualizations. Meng Yang 0011, Jianglang Kang, Xiang Suo, Weiliang Meng, Lijuan Mao, Bin Sheng 0001, Jun Qi 0001 |
Comput. Animat. Virtual Worlds | 9 |
| 2024 | GAN-Based Multi-Decomposition Photo CartoonizationabstractAbstract Background Cartoon images play a vital role in film production, scientific and educational animation, video games, and other fields, and are one of the key visual expressions of artistic creation. However, since hand‐crafted cartoon images often require a great deal of time and effort on the part of professional artists, it is necessary to be able to automatically transform real‐world images into different styles of cartoon images. Although cartoon images vary from artist to artist, cartoon images generally have the unique characteristics of being highly simplified and abstract, with clear edges, smooth color shading, and relatively simple textures. However, existing image cartoonization methods tend to create a number of problems when performing style transfer, which mainly include: (1) the resulting generated images do not have obvious cartoon‐style textures; and (2) the generated images are prone to structural confusion, color artifacts, and loss of the original image content. Therefore, it is also a great challenge in the field of image cartoonization to be able to make a good balance between style transfer and content keeping. Methods In this paper, we propose a GAN‐based multi‐attention mechanism for image cartoonization to address the above issues. The method combines the residual block used to extract deep network features in the generator with the attention mechanism, and further strengthens the perceptual ability of the generative model to cartoon images through the adaptive feature correction of the attention module to improve the cartoon features of the generated images. At the same time, we also introduce the attention mechanism in the convolution block of the discriminator, which is used to further reduce the image visual quality problem caused by the style transfer process. By introducing the attention mechanism into the generator and discriminator models of the generative adversarial network, our method enables the generated images to have obvious cartoon‐style features while effectively improving the image's visual quality. Results A large number of quantitative, qualitative, and ablation experiments are conducted to demonstrate the advantages of our method in the field of image cartoonization and the role of each module in the method. Jianlin Zhu, Ping Li 0016, Bin Sheng 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2024 | Framework of personalized layout for a museum exhibition hall
Meng Yang 0011, Jiaxiu Zhang, Le-Xin Guo, Zhi-Peng Yu, Bin Sheng 0001, Lizhuang Ma |
Multim. Tools Appl. | 7 |
| 2024 | AGG: attention-based gated convolutional GAN with prior guidance for image inpainting
Xiankang Yu, Bin Sheng 0001 |
Neural Comput. Appl. | 4 |
| 2024 | Correction: AGG: attention-based gated convolutional GAN with prior guidance for image inpainting
Xiankang Yu, Bin Sheng 0001 |
Neural Comput. Appl. | 4 |
| 2024 | Learning Motion-Guided Multi-Scale Memory Features for Video Shadow DetectionabstractNatural images often contain multiple shadow regions, and existing video shadow detection methods tend to fail in fully identifying all shadow regions, since they mainly learned temporal features at single-scale and single memory. In this work, we develop a novel convolutional neural network (CNN) to learn motion-guided multi-scale memory features to obtain multi-scale temporal information based on multiple network memories for boosting video shadow detection. To do so, our network first constructs three memories (i.e., a global memory, a local memory, and a motion memory) to combine spatial context and object motion for detecting shadows. Based on these three memories, we then devise a multi-scale motion-guided long-short transformer (MMLT) module to learn multi-scale temporal and motion memory features for predicting a shadow detection map of the input video frame. Our MMLT module includes a dense-scale long transformer (DLT), a dense-scale short transformer (DST), and a dense-scale motion transformer (DMT) to read three memories for learning multi-scale transformer features. Our DLT, DST, and DMT consist of a set of memory-read pooling attention (MPA) blocks and densely connect these output features of multiple MPA blocks to learn multi-scale transformer features since the scales of these output features are varied. By doing so, we can more accurately identify multiple shadow regions with different sizes from the input video. Moreover, we devise a self-supervised pretext task to pre-training the feature encoder for enhancing the downstream video shadow detection. Experimental results on three benchmark datasets show that our video shadow detection network quantitatively and qualitatively outperforms 26 state-of-the-art methods. Jiaxing Shen, Xin Yang 0011, Huazhu Fu, Qing Zhang 0006, Ping Li 0016, Bin Sheng 0001, Liansheng Wang 0002, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | FastAL: Fast Evaluation Module for Efficient Dynamic Deep Active Learning Using Broad Learning SystemabstractState-of-the-art Active Learning (AL) methods often encounter challenges associated with a hysteretic learning process and an expensive data sampling mechanism. The former implies that data selection in the ($i+1$)-th round is solely based on the learned model’s results in the$i$-th round. The latter involves using model inference to calculate data value (e.g., uncertainty estimation based on model inference), which can be cumbersome, particularly when working with large datasets or Deep Neural Networks (DNNs). To address these challenges, we propose FastAL, an efficient and dynamic deep AL framework. Our approach includes an efficient method for calculating data value from the frequency domain perspective, generating multiple candidates. Then, we introduce the Fast Evaluation Module, which directly calculates each candidate’s contribution to future model training and selects the best options. In addition, current AL methods, particularly those based on uncertainty, are susceptible to data bias, which implies that selected data may not represent the original unlabeled data adequately. To alleviate this issue, we propose the De-similar Module, which removes partially similar data. The above three modules are model-agnostic and thus can be seamlessly integrated into any Active Learning framework. We conducted rigorous experiments on various benchmark datasets to validate our approach’s effectiveness. Our results demonstrate that FastAL outperforms other state-of-the-art methods by a significant margin, including those based on uncertainty, diversity, and expected model change. Shuzhou Sun, Huali Xu, Yan Li 0063, Ping Li 0016, Bin Sheng 0001, Xiao Lin 0012 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Dual Branch Multi-Level Semantic Learning for Few-Shot SegmentationabstractFew-shot semantic segmentation aims to segment novel-class objects in a query image with only a few annotated examples in support images. Although progress has been made recently by combining prototype-based metric learning, existing methods still face two main challenges. First, various intra-class objects between the support and query images or semantically similar inter-class objects can seriously harm the segmentation performance due to their poor feature representations. Second, the latent novel classes are treated as the background in most methods, leading to a learning bias, whereby these novel classes are difficult to correctly segment as foreground. To solve these problems, we propose a dual-branch learning method. The class-specific branch encourages representations of objects to be more distinguishable by increasing the inter-class distance while decreasing the intra-class distance. In parallel, the class-agnostic branch focuses on minimizing the foreground class feature distribution and maximizing the features between the foreground and background, thus increasing the generalizability to novel classes in the test stage. Furthermore, to obtain more representative features, pixel-level and prototype-level semantic learning are both involved in the two branches. The method is evaluated on PASCAL-5i1-shot, PASCAL-5i5-shot, COCO-20i1-shot, and COCO-20i5-shot, and extensive experiments show that our approach is effective for few-shot semantic segmentation despite its simplicity. Yadang Chen, Ren Jiang, Yuhui Zheng, Bin Sheng 0001, Zhi-Xin Yang 0001, Enhua Wu |
IEEE Trans. Image Process. | 4 |
| 2024 | DSMT-Net: Dual Self-Supervised Multi-Operator Transformation for Multi-Source Endoscopic Ultrasound DiagnosisabstractPancreatic cancer has the worst prognosis of all cancers. The clinical application of endoscopic ultrasound (EUS) for the assessment of pancreatic cancer risk and of deep learning for the classification of EUS images have been hindered by inter-grader variability and labeling capability. One of the key reasons for these difficulties is that EUS images are obtained from multiple sources with varying resolutions, effective regions, and interference signals, making the distribution of the data highly variable and negatively impacting the performance of deep learning models. Additionally, manual labeling of images is time-consuming and requires significant effort, leading to the desire to effectively utilize a large amount of unlabeled data for network training. To address these challenges, this study proposes the Dual Self-supervised Multi-Operator Transformation Network (DSMT-Net) for multi-source EUS diagnosis. The DSMT-Net includes a multi-operator transformation approach to standardize the extraction of regions of interest in EUS images and eliminate irrelevant pixels. Furthermore, a transformer-based dual self-supervised network is designed to integrate unlabeled EUS images for pre-training the representation model, which can be transferred to supervised tasks such as classification, detection, and segmentation. A large-scale EUS-based pancreas image dataset (LEPset) has been collected, including 3,500 pathologically proven labeled EUS images (from pancreatic and non-pancreatic cancers) and 8,000 unlabeled EUS images for model development. The self-supervised method has also been applied to breast cancer diagnosis and was compared to state-of-the-art deep learning models on both datasets. The results demonstrate that the DSMT-Net significantly improves the accuracy of pancreatic and breast cancer diagnosis. Jiajia Li 0004, Lei Zhu 0003, Ruhan Liu, Dinggang Shen, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2024 | Multi-Label Chest X-Ray Image Classification With Single Positive LabelsabstractDeep learning approaches for multi-label Chest X-ray (CXR) images classification usually require large-scale datasets. However, acquiring such datasets with full annotations is costly, time-consuming, and prone to noisy labels. Therefore, we introduce a weakly supervised learning problem called Single Positive Multi-label Learning (SPML) into CXR images classification (abbreviated as SPML-CXR), in which only one positive label is annotated per image. A simple solution to SPML-CXR problem is to assume that all the unannotated pathological labels are negative, however, it might introduce false negative labels and decrease the model performance. To this end, we present a Multi-level Pseudo-label Consistency (MPC) framework for SPML-CXR. First, inspired by the pseudo-labeling and consistency regularization in semi-supervised learning, we construct a weak-to-strong consistency framework, where the model prediction on weakly-augmented image is treated as the pseudo label for supervising the model prediction on a strongly-augmented version of the same image, and define an Image-level Perturbation-based Consistency (IPC) regularization to recover the potential mislabeled positive labels. Besides, we incorporate Random Elastic Deformation (RED) as an additional strong augmentation to enhance the perturbation. Second, aiming to expand the perturbation space, we design a perturbation stream to the consistency framework at the feature-level and introduce a Feature-level Perturbation-based Consistency (FPC) regularization as a supplement. Third, we design a Transformer-based encoder module to explore the sample relationship within each mini-batch by a Batch-level Transformer-based Correlation (BTC) regularization. Extensive experiments on the CheXpert and MIMIC-CXR datasets have shown the effectiveness of our MPC framework for solving the SPML-CXR problem. Jiayin Xiao, Si Li 0005, Tongxu Lin, Jian Zhu 0001, Xiaochen Yuan, David Dagan Feng, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Unsupervised Fusion Feature Matching for Data Bias in Uncertainty Active LearningabstractActive learning (AL) aims to sample the most valuable data for model improvement from the unlabeled pool. Traditional works, especially uncertainty-based methods, are prone to suffer from a data bias issue, which means that selected data cannot cover the entire unlabeled pool well. Although there have been lots of literature works focusing on this issue recently, they mainly benefit from the huge additional training costs and the artificially designed complex loss. The latter causes these methods to be redesigned when facing new models or tasks, which is very time-consuming and laborious. This article proposes a feature-matching-based uncertainty that resamples selected uncertainty data by feature matching, thus removing similar data to alleviate the data bias issue. To ensure that our proposed method does not introduce a lot of additional costs, we specially design a unsupervised fusion feature matching (UFFM), which does not require any training in our novel AL framework. Besides, we also redesign several classic uncertainty methods to be applied to more complex visual tasks. We conduct rigorous experiments on lots of standard benchmark datasets to validate our work. The experimental results show that our UFFM is better than the similar unsupervised feature matching technologies, and our proposed uncertainty calculation method outperforms random sampling, classic uncertainty approaches, and recent state-of-the-art (SOTA) uncertainty approaches. Shuzhou Sun, Xiao Lin 0012, Ping Li 0016, Lei Zhu 0003, C. L. Philip Chen, Bin Sheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | SparseVoxNet: 3-D Object Recognition With Sparsely Aggregation of 3-D Dense BlocksabstractAutomatic recognition of 3-D objects in a 3-D model by convolutional neural network (CNN) methods has been successfully applied to various tasks, e.g., robotics and augmented reality. Three-dimensional object recognition is mainly performed by analyzing the object using multi-view images, depth images, graphs, or volumetric data. In some cases, using volumetric data provides the most promising results. However, existing recognition techniques on volumetric data have many drawbacks, such as losing object details on converting points to voxels and the large size of the input volume data that leads to substantial 3-D CNNs. Using point clouds could also provide very promising results; however, point-cloud-based methods typically need sparse data entry and time-consuming training stages. Thus, using volumetric could be a more efficient and flexible recognizer for our special case in the School of Medicine, Shanghai Jiao Tong University. In this article, we propose a novel solution to 3-D object recognition from volumetric data using a combination of three compact CNN models, low-cost SparseNet, and feature representation technique. We achieve an optimized network by estimating extra geometrical information comprising the surface normal and curvature into two separated neural networks. These two models provide supplementary information to each voxel data that consequently improve the results. The primary network model takes advantage of all the predicted features and uses these features in Random Forest (RF) for recognition purposes. Our method outperforms other methods in training speed in our experiments and provides an accurate result as good as the state-of-the-art. Ahmad Karambakhsh, Bin Sheng 0001, Ping Li 0016, Huating Li, Jinman Kim, Younhyun Jung, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Action-aware Linguistic Skeleton Optimization Network for Non-autoregressive Video CaptioningabstractNon-autoregressive video captioning methods generate visual words in parallel but often overlook semantic correlations among them, especially regarding verbs, leading to lower caption quality. To address this, we integrate action information of highlighted objects to enhance semantic connections among visual words. Our proposed Action-aware Language Skeleton Optimization Network (ALSO-Net) tackles the challenge of extracting action information across frames, improving understanding of complex context-dependent video actions and reducing sentence inconsistencies. ALSO-Net incorporates a linguistic skeleton tag generator to refine semantic correlations and a video action predictor to enhance verb prediction accuracy in video captions. We also address issues of unsatisfactory caption length and quality by jointly optimizing different levels of motion prediction loss. Experimental evaluation on prominent video captioning datasets demonstrates that ALSO-Net outperforms baseline methods by a significant margin and achieves competitive performance compared to state-of-the-art autoregressive methods with smaller model complexity and faster inference time. Shuqin Chen, Xian Zhong, Lei Zhu 0003, Ping Li 0016, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | Clustering Environment Aware Learning for Active Domain AdaptationabstractDespite the significant progress in unsupervised domain adaptation (UDA), the performance of UDA methods is still far inferior to that of the fully supervised ones. In practical scenarios, it is usually feasible to acquire labels on a small portion of the target data through active learning (AL), which aims to train an effective model with as few queried instances as possible. However, due to the domain shift, the instances selected by existing AL algorithms can be uninformative, redundant, or outlying. To address this issue, we propose a novel approach, namely, clustering environment-aware learning (CEAL), for active domain adaptation (ADA). CEAL selects potentially the most valuable instances under domain shift by exploring the informativeness and representativeness of target samples in a clustering environment-aware manner. Specifically, for the informativeness, we not only leverage the knowledge of individual points but also their nearby neighbors, by measuring the proposed clustering environment aware informativeness score (CEAIS), thus ensuring that the selected samples are highly informative. For the representativeness, we design two schemes called point distance release (PDR) and informativeness score difference exclusion (ISDE) to guarantee the diversity and validity of the selected samples. Furthermore, we fully utilize the large amount of unlabeled data from target domain via pseudo labeling and adopt information maximization to improve the reliability of the target pseudo labels, thereby further improving the performance of the model. The effectiveness of our method is empirically verified on various benchmark datasets against recent state-of-the-art algorithms. Jian Zhu 0001, Qintai Hu, Yutang Xiao, Boyu Wang 0004, Bin Sheng 0001, C. L. Philip Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2024 | Efficient Binocular Rendering of Volumetric Density Fields With Coupled Adaptive Cube-Map Ray Marching for Virtual RealityabstractCreating visualizations of multiple volumetric density fields is demanding in virtual reality (VR) applications, which often include divergent volumetric density distributions mixed with geometric models and physics-based simulations. Real-time rendering of such complex environments poses significant challenges for rendering quality and performance. This article presents a novel scheme for efficient real-time rendering of varying translucent volumetric density fields with global illumination (GI) effects on high-resolution binocular VR displays. Our scheme proposes creative solutions to address three challenges involved in the target problem. First, to tackle the doubled heavy workloads of binocular ray marching, we explore the anti-aliasing principles and more advanced potentials of ray marching on interior cube-map faces, and propose a coupled ray-marching technique that converges to multi-resolution cube maps with interleaved adaptive sampling. Second, we devise a fully dynamic ambient GI approximation method that leverages spherical-harmonics (SH) transform information of the phase function to reduce the huge amount of ray sampling required for GI while ensuring fidelity. The method catalyzes spatial ray-marching reuse and adaptive temporal accumulation. Third, we deploy a two-phase ray-tracing algorithm with a tiled k-buffer to achieve fast processing of order-independent transparency (OIT) for multiple volume instances. Consequently, high-quality and high-performance real-time dynamic volume rendering can be achieved under constrained budgets controlled by developers. As our solution supports mixed mesh-volume rendering, the test results prove the practical usefulness of our approach for high-resolution binocular VR rendering on hybrid multi-volumetric and geometric environments. Tianchen Xu, Xiaohua Ren, Jiale Yang, Bin Sheng 0001, Enhua Wu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | AI-enhanced digital technologies for myopia management: advancements, challenges, and future prospects
Saba Ghazanfar Ali, Zhouyu Guan, Tingli Chen, Ping Li 0016, Po Yang 0001, Zainab Ghazanfar, Younhyun Jung, Bin Sheng 0001, Xiangning Wang |
Vis. Comput. | 11 |
| 2024 | Attention-based multi-scale feature fusion network for myopia grading using optical coherence tomography images
Gengyou Huang, Lei Bi 0001, Tingli Chen, Bin Sheng 0001 |
Vis. Comput. | 6 |
| 2024 | Deep choroid layer segmentation using hybrid features extraction from OCT images
Saleha Masood, Saba Ghazanfar Ali, Xiangning Wang, Afifa Masood, Ping Li 0016, Huating Li, Younhyun Jung, Bin Sheng 0001, Jinman Kim |
Vis. Comput. | 8 |
| 2024 | TSNet: Task-specific network for joint diabetic retinopathy grading and lesion segmentation of ultra-wide optical coherence tomography angiography images
Jixue Tang, Xiang-ning Wang, Tingli Chen, Bin Sheng 0001 |
Vis. Comput. | 7 |
| 2023 | AsT: An Asymmetric-Sensitive Transformer for Osteonecrosis of the Femoral Head Detection (Student Abstract)abstractEarly diagnosis of osteonecrosis of the femoral head (ONFH) can inhibit the progression and improve femoral head preservation. The radiograph difference between early ONFH and healthy ones is not apparent to the naked eye. It is also hard to produce a large dataset to train the classification model. In this paper, we propose Asymmetric-Sensitive Transformer (AsT) to capture the uneven development of the bilateral femoral head to enable robust ONFH detection. Our ONFH detection is realized using the self-attention mechanism to femoral head regions while conferring sensitivity to the uneven development by the attention-shared transformer. The real-world experiment studies show that AsT achieves the best performance of AUC 0.9313 in the early diagnosis of ONFH and can find out misdiagnosis cases firmly. Feng Lu 0003, Wei Li 0058, Bin Sheng 0001, Hai Jin 0001, Albert Y. Zomaya |
AAAI | 5 |
| 2023 | Reference-Based Line Drawing Colorization Through Diffusion Model
Jiaze He, Ziruo Li, Ping Li 0016, Lei Zhu 0003, Bin Sheng 0001, Subrota K. Mondal |
CGI | 7 |
| 2023 | MagicMirror: A 3-D Real-Time Virtual Try-On System Through Cloth Simulation
Zhanyi Huang, Tangsheng Guo, Ping Li 0016, Bin Sheng 0001 |
CGI | 6 |
| 2023 | Scene-aware Human Pose Generation using TransformerabstractAffordance learning considers the interaction opportunities for an actor in the scene and thus has wide application in scene understanding and intelligent robotics. In this paper, we focus on contextual affordance learning, i.e., using affordance as context to generate a reasonable human pose in a scene. Existing scene-aware human pose generation methods could be divided into two categories depending on whether using pose templates. Our proposed method belongs to the template-based category, which benefits from the representative pose templates. Moreover, inspired by recent transformer-based methods, we associate each query embedding with a pose template, and use the interaction between query embeddings and scene feature map to effectively predict the scale and offsets for each pose template. In addition, we employ knowledge distillation to facilitate the offset learning given the predicted scale. Comprehensive experiments on Sitcom dataset demonstrate the effectiveness of our method. Jieteng Yao, Junjie Chen 0008, Li Niu 0002, Bin Sheng 0001 |
ACM Multimedia | 4 |
| 2023 | Global-and-local aware network for low-light image enhancement
Xufeng He, Lei Liang 0003, Jianfa Wu, Bin Sheng 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | EditorialabstractThis special issue is the third issue dedicated to the best papers of the second Call of the CGI 2022 conference. In 2022, CGI (Computer Graphics International) was organized by MIRALab at the University of Geneva. It was supported by the Computer Graphics Society. The conference was very successful and attracted more than 150 online participants through Zoom. Twenty-six papers were selected from more than 100 papers submitted to the second call of the conference. This issue contains the last eight full papers. Nine papers were already published in the 33.5 issue and nine papers in the 33.6 issue. All papers have been reviewed by two or three reviewers of the CGI 2022 Program Committee, revised according to the reviewers' comments, and checked and reviewed again by the CAVW Editorial Board. Data-driven based double-layer bicycle simulation model by Tianlu Mao, Zhong Fang, Qinyuan Yan, and Zhaoqi Wang, from Academy of Sciences, Beijing, and Ruoyu Meng and Shaohua Liu, all in China. This special issue is edited by the program-co-chairs of CGI2022 conference. Jinman Kim, George Papagiannakis, Bin Sheng 0001, Daniel Thalmann |
Comput. Animat. Virtual Worlds | 3 |
| 2023 | MNGNAS: Distilling Adaptive Combination of Multiple Searched Networks for One-Shot Neural Architecture SearchabstractRecently neural architecture (NAS) search has attracted great interest in academia and industry. It remains a challenging problem due to the huge search space and computational costs. Recent studies in NAS mainly focused on the usage of weight sharing to train a SuperNet once. However, the corresponding branch of each subnetwork is not guaranteed to be fully trained. It may not only incur huge computation costs but also affect the architecture ranking in the retraining procedure. We propose a multi-teacher-guided NAS, which proposes to use the adaptive ensemble and perturbation-aware knowledge distillation algorithm in the one-shot-based NAS algorithm. The optimization method aiming to find the optimal descent directions is used to obtain adaptive coefficients for the feature maps of the combined teacher model. Besides, we propose a specific knowledge distillation process for optimal architectures and perturbed ones in each searching process to learn better feature maps for later distillation procedures. Comprehensive experiments verify our approach is flexible and effective. We show improvement in precision and search efficiency in the standard recognition dataset. We also show improvement in correlation between the accuracy of the search algorithm and true accuracy by NAS benchmark datasets. Guhao Qiu, Ping Li 0016, Lei Zhu 0003, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | TMM-Nets: Transferred Multi- to Mono-Modal Generation for Lupus Retinopathy DiagnosisabstractRare diseases, which are severely underrepresented in basic and clinical research, can particularly benefit from machine learning techniques. However, current learning-based approaches usually focus on either mono-modal image data or matched multi-modal data, whereas the diagnosis of rare diseases necessitates the aggregation of unstructured and unmatched multi-modal image data due to their rare and diverse nature. In this study, we therefore propose diagnosis-guided multi-to-mono modal generation networks (TMM-Nets) along with training and testing procedures. TMM-Nets can transfer data from multiple sources to a single modality for diagnostic data structurization. To demonstrate their potential in the context of rare diseases, TMM-Nets were deployed to diagnose the lupus retinopathy (LR-SLE), leveraging unmatched regular and ultra-wide-field fundus images for transfer learning. The TMM-Nets encoded the transfer learning from diabetic retinopathy to LR-SLE based on the similarity of the fundus lesions. In addition, a lesion-aware multi-scale attention mechanism was developed for clinical alerts, enabling TMM-Nets not only to inform patient care, but also to provide insights consistent with those of clinicians. An adversarial strategy was also developed to refine multi- to mono-modal image generation based on diagnostic results and the data distribution to enhance the data augmentation performance. Compared to the baseline model, the TMM-Nets showed 35.19% and 33.56% F1 score improvements on the test and external validation sets, respectively. In addition, the TMM-Nets can be used to develop diagnostic models for other rare diseases. Ruhan Liu, Tianqin Wang, Huating Li, Ping Zhang 0016, Xiaokang Yang 0001, Dinggang Shen, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2023 | PhotoHelper: Portrait Photographing Guidance Via Deep Feature Retrieval and FusionabstractWe introduce a new photographing guidance (PhotoHelper) for amateur photographers to enhance their portrait photo quality using deep feature retrieval and fusion. In our model, we comprehensively integrate empirical aesthetic rules, traditional machine learning algorithms and deep neural networks to extract different kinds of features in both color and space aspects. With these features, we build a modified random forest with a structured photograph collection to identify types of photos. We also define the composition matching score to measure the similarity between the given photo and the reference photo. By combining all of the above processes, a one-stop deep portrait photographing guidance is constructed to provide users with professional reference photographs that are similar to the current scene and automatically generate spatial composition guidance according to the user-selected reference photo. Experiments and evaluations show that the aesthetic quality of portrait photos can be significantly improved via the composition guidance of our photographing guidance approach. Bin Sheng 0001, Ping Li 0016, Tong-Yee Lee |
IEEE Trans. Multim. | 2 |
| 2023 | EAPT: Efficient Attention Pyramid Transformer for Image ProcessingabstractRecent transformer-based models, especially patch-based methods, have shown huge potentiality in vision tasks. However, the split fixed-size patches divide the input features into the same size patches, which ignores the fact that vision elements are often various and thus may destroy the semantic information. Also, the vanilla patch-based transformer cannot guarantee the information communication between patches, which will prevent the extraction of attention information with a global view. To circumvent those problems, we propose an Efficient Attention Pyramid Transformer (EAPT). Specifically, we first propose the Deformable Attention, which learns an offset for each position in patches. Thus, even with split fixed-size patches, our method can still obtain non-fixed attention information that can cover various vision elements. Then, we design the Encode-Decode Communication module (En-DeC module), which can obtain communication information among all patches to get more complete global attention information. Finally, we propose a position encoding specifically for vision transformers, which can be used for patches of any dimension and any length. Extensive experiments on the vision tasks of image classification, object detection, and semantic segmentation demonstrate the effectiveness of our proposed model. Furthermore, we also conduct rigorous ablation studies to evaluate the key components of the proposed structure. Xiao Lin 0012, Shuzhou Sun, Bin Sheng 0001, Ping Li 0016, David Dagan Feng |
IEEE Trans. Multim. | 4 |
| 2023 | FFFN: Frame-By-Frame Feedback Fusion Network for Video Super-ResolutionabstractVideo super-resolution (VSR) is a fundamental and challenging task in computer vision. Many of the existing VSR works focus on how to effectively align neighboring frames to better incorporate temporal information, while little work is devoted to the important subsequent step of inter-frame information fusion, and the existing methods on frame fusion have shortcomings such as not being able to make full use of spatio-temporal information. In this work, we propose a Frame-by-frame Feedback Fusion Network (FFFN) for VSR tasks. By applying the feedback learning mechanism commonly existing in the human cognitive system to the frame fusion stage, FFFN can refine low-level representation of the fused frames with high-level information in a coarse-to-fine manner. Specifically, after the neighboring frames are aligned, we first rearrange them from near to far according to the distance from the reference frame in the temporal space, and then feed them one-by-one into a proposed recurrent structure called Feedback Fusion Module (FFM), which is then able to iteratively generate high-level representation of the fused frames with several Feature Refinement Groups (FRGs) and feedback connections. Finally, we design a Dual-path Residual Reconstruction Module (DRRM) to reconstruct the final high-resolution image. The proposed FFFN comes with a strong frame fusion and reconstruction ability, and extensive experiments on several benchmark data sets show that it achieves favorable performance against state-of-the-art methods. Jian Zhu 0001, Qingwu Zhang, Lunke Fei, Ruichu Cai, Yuan Xie 0006, Bin Sheng 0001, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | BaGFN: Broad Attentive Graph Fusion Network for High-Order Feature InteractionsabstractModeling feature interactions is of crucial significance to high-quality feature engineering on multifiled sparse data. At present, a series of state-of-the-art methods extract cross features in a rather implicit bitwise fashion and lack enough comprehensive and flexible competence of learning sophisticated interactions among different feature fields. In this article, we propose a new broad attentive graph fusion network (BaGFN) to better model high-order feature interactions in a flexible and explicit manner. On the one hand, we design an attentive graph fusion module to strengthen high-order feature representation under graph structure. The graph-based module develops a new bilinear-cross aggregation function to aggregate the graph node information, employs the self-attention mechanism to learn the impact of neighborhood nodes, and updates the high-order representation of features by multihop fusion steps. On the other hand, we further construct a broad attentive cross module to refine high-order feature interactions at a bitwise level. The optimized module designs a new broad attention mechanism to dynamically learn the importance weights of cross features and efficiently conduct the sophisticated high-order feature interactions at the granularity of feature dimensions. The final experimental results demonstrate the effectiveness of our proposed model. Wenling Zhang, Bin Sheng 0001, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | FSAD-Net: Feedback Spatial Attention Dehazing NetworkabstractRecent dehazing networks learn more discriminative high-level features by designing deeper networks or introducing complicated structures, while ignoring inherent feature correlations in intermediate layers. In this article, we establish a novel and effective end-to-end dehazing method, named feedback spatial attention dehazing network (FSAD-Net). FSAD-Net is based on the recurrent structure and consists of four modules: a shallow feature extraction block (SFEB), a feedback block (FB), multiple advanced residual blocks (ARBs), and a reconstruction block (RB). FB is designed to handle feedback connections, and it can improve the dehazing performance by exploiting the dependencies of deep features across stages. ARB implements a novel attention-based estimation on a residual block to adapt to pixels with different distributions. Finally, RB helps restore haze-free images. It can be seen from the experimental results that FSAD-Net almost outperforms the state-of-the-arts in terms of five quantitative metrics. Moreover, the qualitatively comparisons on real-world images also demonstrate the superiority of the proposed FSAD-Net. Considering the efficiency and effectiveness of FSAD-Net, it can be expected to serve as a suitable image dehazing baseline in the future. Yu Zhou 0066, Ping Li 0016, Haitao Song 0001, C. L. Philip Chen, Bin Sheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | SThy-Net: a feature fusion-enhanced dense-branched modules network for small thyroid nodule classification from ultrasound images
Abdulrhman H. Al-Jebrni, Saba Ghazanfar Ali, Huating Li, Xiao Lin 0012, Ping Li 0016, Younhyun Jung, Jinman Kim, David Dagan Feng, Bin Sheng 0001, Lixin Jiang |
Vis. Comput. | 9 |
| 2023 | TransMRSR: transformer-based self-distilled generative prior for brain MRI super-resolution
Xiaohong Liu 0001, Tao Tan 0002, Menghan Hu, Xiaoer Wei, Tingli Chen, Bin Sheng 0001 |
Vis. Comput. | 7 |
| 2023 | GuideRender: large-scale scene navigation based on multi-modal view frustum movement prediction
Xiaoyu Chi, Bin Sheng 0001, Rynson W. H. Lau |
Vis. Comput. | 3 |
| 2022 | Input-Specific Robustness Certification for Randomized SmoothingabstractAlthough randomized smoothing has demonstrated high certified robustness and superior scalability to other certified defenses, the high computational overhead of the robustness certification bottlenecks the practical applicability, as it depends heavily on the large sample approximation for estimating the confidence interval. In existing works, the sample size for the confidence interval is universally set and agnostic to the input for prediction. This Input-Agnostic Sampling (IAS) scheme may yield a poor Average Certified Radius (ACR)-runtime trade-off which calls for improvement. In this paper, we propose Input-Specific Sampling (ISS) acceleration to achieve the cost-effectiveness for robustness certification, in an adaptive way of reducing the sampling size based on the input characteristic. Furthermore, our method universally controls the certified radius decline from the ISS sample size reduction. The empirical results on CIFAR-10 and ImageNet show that ISS can speed up the certification by more than three times at a limited cost of 0.05 certified radius. Meanwhile, ISS surpasses IAS on the average certified radius across the extensive hyperparameter settings. Specifically, ISS achieves ACR=0.958 on ImageNet in 250 minutes, compared to ACR=0.917 by IAS under the same condition. We release our code in https://github.com/roy-ch/Input-Specific-Certification. Ruoxin Chen, Jie Li 0002, Junchi Yan, Ping Li 0016, Bin Sheng 0001 |
AAAI | 5 |
| 2022 | SlimFliud-Net: Fast Fluid Simulation Using Admm Pruning
Songyang Yu, Ping Li 0016, Weiguang Li, Enhua Wu, Bin Sheng 0001 |
CGI | 6 |
| 2022 | Graph Adversarial Network with Bottleneck Adapter Tuning for Sign Language Production
Chunpeng Yu, Jiajia Liang, Yihui Liao, Bin Sheng 0001 |
CGI | 5 |
| 2022 | Visual Indoor Navigation Using Mobile Augmented Reality
Han Zhang 0053, Mengsi Guo, Ping Lu 0008, Liu Sen, Bin Sheng 0001 |
CGI | 8 |
| 2022 | Optimal Transport for Label-Efficient Visible-Infrared Person Re-Identification
Jiangming Wang, Zhizhong Zhang 0001, Mingang Chen, Cong Wang 0039, Bin Sheng 0001, Yanyun Qu, Yuan Xie 0006 |
ECCV (24) | 6 |
| 2022 | Experimental protocol designed to employ Nd: YAG laser surgery for anterior chamber glaucoma detection via UBMabstractAbstract Angle closure glaucoma leads to fluid deposition in eye, and intraocular pressure occurs that damage the optic nerve, causes blindness and vision loss. Anterior chamber (AC) evaluation is imperative for determining the risk of angle‐closure. Previously, techniques were dependent on either Pentacam–Scheimpflug that interprets poor visual information, anterior segment optical coherence tomography is injurious to intercede opaque optical structures. Therefore, in this paper, an experimental protocol is designed for detailed disease analysis based on IBM SPSS statistics via ultrasound biomicroscopy which is superior in evaluating deep structures; first, the affected parameter for AC is analysed, and afterwards the direction that needs laser surgery is explored. Experiments are conducted on large‐scale clinical studies from an affiliated hospital in Shanghai, China. The dataset comprised 600 AC images in five directions of 60 subjects. The mean with standard deviation for anterior open distance is 0.158790.096779 mm, 0.158630.081435 mm, and anterior chamber angle is 18.74908.0315, 18.74108.3889 for left and right eye respectively. It is found that anterior chamber angle in the downside of the AC is wider than the upside. However, this decision is partly based on the narrowest part of the angle to widen the depth of the direction and eliminate pupil block. Saba Ghazanfar Ali, Riaz Ali, Bin Sheng 0001, Huating Li, Po Yang 0001, Ping Li 0016, Younhyun Jung, Ping Lu 0008, Jinman Kim |
IET Image Process. | 3 |
| 2022 | SCPA-Net: Self-calibrated pyramid aggregation for image dehazingabstractAbstract Dehazing as an important image processing field has developed for many years, there exist many excellent methods for exploring more complex networks to solve this problem. In this paper, instead of designing a complex network structure, we propose a novel dehazing network based on the consideration of enhancing feature aggregation and feature representation abilities of dehazing architecture. Specifically, we propose a self‐calibrated pyramid aggregation network (SCPA‐Net) for image dehazing, which is based on an encoder‐decoder architecture. In the encoder, we build a self‐attention block as a unit to aggregate information from a neighborhood to adapt to its content. In the decoder, we introduce a self‐calibration block to capture long‐range spatial and channel dependencies to produce more discriminative representations. Finally, to learn the scale information, the pyramid upsampling structure is applied to aggregate the multiscale self‐calibrated attentive features. Experimental results show our SCPA‐Net can achieve impressive dehazing performance. Yu Zhou 0066, Ping Li 0016, Bin Sheng 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2022 | Special issue on computer graphics international 2022 part 1abstractThis special issue is dedicated to the best papers of the second Call of the CGI 2022 conference. This year CGI (Computer Graphics International) was organized by MIRALab at the University of Geneva. It was supported by the Computer Graphics Society. The conference was very successful and attracted more than 150 online participants through Zoom. Twenty-six papers were selected from more than 100 papers submitted to the second call of the conference. This issue contains nine full papers. The other selected papers will be published in the two next issues respectively 33.6 and 34.1. All papers have been reviewed by two or three reviewers of the CGI 2022 Program Committee, revised according to the reviewers' comments, and checked and reviewed again by the CAVW Editorial Board. The first paper in this issue is the recipient of the CAVW Best Paper Award given at CGI 2022 by the Award Committee chaired by Professor Nadia Magnenat Thalmann, MIRALab-University of Geneva, Switzerland, and Professor Constantine Stephanidis, ICS—FORTH, Greece. This awarded paper is: Simulation of collective pursuit-evasion behavior with runtime situational awareness by Zhenjing Yu, Tan Junyin, and Sheng Li, from Peking University, China. Jinman Kim, George Papagiannakis, Bin Sheng 0001, Daniel Thalmann |
Comput. Animat. Virtual Worlds | 3 |
| 2022 | EditorialabstractThis special issue is the second issue dedicated to the best papers of the second Call of the CGI 2022 conference. This year CGI (Computer Graphics International) was organized by MIRALab at the University of Geneva. It was supported by the Computer Graphics Society. The conference was very successful and attracted more than 150 online participants through Zoom. Twenty-six papers were selected from more than 100 papers submitted to the second call of the conference. This issue contains nine full papers. Nine papers were already published in the 33.5 issue and the last papers will appear in the issue 34.1. All papers have been reviewed by two or three reviewers of the CGI 2022 Program Committee, revised according to the reviewers' comments, and checked and reviewed again by the CAVW Editorial Board. This special issue is edited by the program co-chairs of CGI2022 conference: Jinman Kim, Sydney University, Australia, George Papagiannakis, University of Crete, Greece, Bin Sheng, Shanghai Jiao Tong University, China, Daniel Thalmann, EPFL, Switzerland. Jinman Kim, George Papagiannakis, Bin Sheng 0001, Daniel Thalmann |
Comput. Animat. Virtual Worlds | 3 |
| 2022 | BAW: learning from class imbalance and noisy labels with batch adaptation weighted loss
Siyuan Pan, Bin Sheng 0001, Gaoqi He, Huating Li, Guangtao Xue |
Multim. Tools Appl. | 2 |
| 2022 | Improving Video Temporal Consistency via Broad Learning SystemabstractApplying image-based processing methods to original videos on a framewise level breaks the temporal consistency between consecutive frames. Traditional video temporal consistency methods reconstruct an original frame containing flickers from corresponding nonflickering frames, but the inaccurate correspondence realized by optical flow restricts their practical use. In this article, we propose a temporally broad learning system (TBLS), an approach that enforces temporal consistency between frames. We establish the TBLS as a flat network comprising the input data, consisting of an original frame in an original video, a corresponding frame in the temporally inconsistent video on which the image-based technique was applied, and an output frame of the last original frame, as mapped features in feature nodes. Then, we refine extracted features by enhancing the mapped features as enhancement nodes with randomly generated weights. We then connect all extracted features to the output layer with a target weight vector. With the target weight vector, we can minimize the temporal information loss between consecutive frames and the video fidelity loss in the output videos. Finally, we remove the temporal inconsistency in the processed video and output a temporally consistent video. Besides, we propose an alternative incremental learning algorithm based on the increment of the mapped feature nodes, enhancement nodes, or input data to improve learning accuracy by a broad expansion. We demonstrate the superiority of our proposed TBLS by conducting extensive experiments. Bin Sheng 0001, Ping Li 0016, Riaz Ali, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2022 | Automatic Detection and Classification System of Domestic Waste via Multimodel Cascaded Convolutional Neural NetworkabstractDomestic waste classification was incorporated into legal provisions recently in China. However, relying on manpower to detect and classify domestic waste is highly inefficient. To that end, in this article, we propose a multimodel cascaded convolutional neural network (MCCNN) for domestic waste image detection and classification. MCCNN combined three subnetworks (DSSD, YOLOv4, and Faster-RCNN) to obtain the detections. Moreover, to suppress the false-positive predicts, we utilized a classification model cascaded with the detection part to judge whether the detection results are correct. To train and evaluate MCCNN, we designed a large-scale waste image dataset (LSWID), containing 30 000 domestic waste multilabeled images with 52 categories. To the best of our knowledge, the LSWID is the largest dataset on domestic waste images. Furthermore, a smart trash can is designed and applied to a Shanghai community, which helped to make waste recycling more efficient. Experimental results showed a state-of-the-art performance, with an average improvement of 10% in detection precision. Jiajia Li 0004, Jie Chen 0097, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, David Dagan Feng, Jun Qi 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | ECSU-Net: An Embedded Clustering Sliced U-Net Coupled With Fusing Strategy for Efficient Intervertebral Disc Segmentation and ClassificationabstractAutomatic vertebra segmentation from computed tomography (CT) image is the very first and a decisive stage in vertebra analysis for computer-based spinal diagnosis and therapy support system. However, automatic segmentation of vertebra remains challenging due to several reasons, including anatomic complexity of spine, unclear boundaries of the vertebrae associated with spongy and soft bones. Based on 2D U-Net, we have proposed an Embedded Clustering Sliced U-Net (ECSU-Net). ECSU-Net comprises of three modules named segmentation, intervertebral disc extraction (IDE) and fusion. The segmentation module follows an instance embedding clustering approach, where our three sliced sub-nets use axis of CT images to generate a coarse 2D segmentation along with embedding space with the same size of the input slices. Our IDE module is designed to classify vertebra and find the inter-space between two slices of segmented spine. Our fusion module takes the coarse segmentation (2D) and outputs the refined 3D results of vertebra. A novel adaptive discriminative loss (ADL) function is introduced to train the embedding space for clustering. In the fusion strategy, three modules are integrated via a learnable weight control component, which adaptively sets their contribution. We have evaluated classical and deep learning methods on Spineweb dataset-2. ECSU-Net has provided comparable performance to previous neural network based algorithms achieving the best segmentation dice score of 95.60% and classification accuracy of 96.20%, while taking less time and computation resources. Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Ping Li 0016, Huating Li, Guangtao Xue, Harry Qin, Jinman Kim, David Dagan Feng |
IEEE Trans. Image Process. | 3 |
| 2022 | Face Sketch Synthesis Using Regularized Broad Learning SystemabstractThere are two main categories of face sketch synthesis: data- and model-driven. The data-driven method synthesizes sketches from training photograph-sketch patches at the cost of detail loss. The model-driven method can preserve more details, but the mapping from photographs to sketches is a time-consuming training process, especially when the deep structures require to be refined. We propose a face sketch synthesis method via regularized broad learning system (RBLS). The broad learning-based system directly transforms photographs into sketches with rich details preserved. Also, the incremental learning scheme of broad learning system (BLS) ensures that our method easily increases feature mappings and remodels the network without retraining when the extracted feature mapping nodes are not sufficient. Besides, a Bayesian estimation-based regularization is introduced with the BLS to aid further feature selection and improve the generalization ability and robustness. Various experiments on the CUHK student data set and Aleix Robert (AR) data set demonstrated the effectiveness and efficiency of our RBLS method. Unlike existing methods, our method synthesizes high-quality face sketches much efficiently and greatly reduces computational complexity both in the training and test processes. Ping Li 0016, Bin Sheng 0001, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | PointALCR: adversarial latent GAN and contrastive regularization for point cloud completion
Changjie Cheng, Bin Sheng 0001, Lizhuang Ma |
Vis. Comput. | 4 |
| 2022 | Joint feedback and recurrent deraining network with ensemble learning
Yu Luo 0004, Menghua Wu, Qingdong Huang, Jian Zhu 0001, Jie Ling 0002, Bin Sheng 0001 |
Vis. Comput. | 6 |
| 2022 | Real-time spatial normalization for dynamic gesture classification
Sofiane Zeghoud, Saba Ghazanfar Ali, Egemen Ertugrul, Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Xiaoyu Chi, Jinman Kim, Lijuan Mao |
Vis. Comput. | 5 |
| 2022 | RADepthNet: Reflectance-Aware Monocular Depth EstimationabstractMonocular depth estimation aims to predict the dense depth map from a single RGB image, which has important applications in 3D reconstruction, automatic driving, and augmented reality. However, existing methods directly feed the original RGB image into the model to extract depth features without avoiding the interference of depth-irrelevant information on depth estimation accuracy, which leads to inferior performance. To remove the influence of depth-irrelevant information and improve depth prediction accuracy, we propose RADepthNet, a novel reflectance-guided network fusing boundary features. Specifically, our method predicts depth maps using three steps: 1) Intrinsic Image Decomposition. We propose a Reflectance extraction module consisting of an encoder-decoder structure to extract depth-related reflectance. We demonstrate that the module can reduce the influence of illumination on depth estimation through an ablation study. 2) Boundary Detection. Boundary extraction module, consisting of an encoder, a refinement block, and an upsample block, is proposed to better predict depth at object boundaries utilizing gradient constraints. 3) Depth Prediction Module. Use a different encoder from 2) to obtain depth features from the reflectance map and fuse boundary features to predict depth. Besides, we proposed FIFADataset, a depth estimation dataset applied in soccer scenarios. Extensive experiments on the public dataset and our proposed FIFADataset show that our method achieves state-of-the-art performance. Chuxuan Li, Ran Yi 0002, Saba Ghazanfar Ali, Lizhuang Ma, Enhua Wu, Lijuan Mao, Bin Sheng 0001 |
Virtual Real. Intell. Hardw. | 8 |
| 2022 | Computer graphics for metaverseabstractCGI is one of the oldest international conferences in Computer Graphics in the world.It is the official conference of the Computer Graphics Society (CGS), a long-standing international computer graphics organization.CGI conference has been held annually in many different countries across the world and has gained a reputation as one of the key conferences for researchers and practitioners to share their achievements and discover the latest advances in Computer Graphics.With the change in the form of networking and intelligence in industry, manufacturing and all aspects of society, and the development of technology, we are aware of the increasingly obvious trend of evolution of intelligence in human society.Among them, metaverse is increasingly becoming a hot spot for research in various industries and has broad application prospects.It absorbs the results of the information revolution, the Internet revolution, the artificial intelligence revolution, and the virtual reality technology revolution including VR, AR, MR, and especially game engines, showing mankind the possibility of building a holographic digital world parallel to the traditional physical world.The core of the metaverse lies in the hosting of virtual assets and virtual identities.Unlike traditional games, users can experience different content, make different friends, create their own creations, and perform a series of virtual activities in the metaverse.With the popularity of smart terminals and the rise of applications such as e-commerce/short videos/games, "metaverse" has become an inevitable trend in the development of digital society.In a broad sense, the "metaverse" is a virtual space-time consisting of a series of augmented reality (AR), virtual reality (VR) and the Internet; in a narrow sense, the "metaverse" is a virtual world parallel to the real world.By wearing a helmet and headset device, one can enter a three-dimensional world constructed by computer simulation through a terminal connection".The new mode of "virtual reality" presentation and scene interaction for scene visualization will be more conducive to better visual effects and interactive operations in the digital world.The metaverse becomes the best track and new growth point for AI applications because of its huge imagination, close social attention and rich landing scenes, while AI and related arithmetic, big data and other technical fields are the technical base for the metaverse to become a kind of concrete expression in the future.Overall, with the further development of human technology and the improvement of hardware level, it becomes possible for humans to build a "meta" world.This year, CGI 2022 is still online as the pandemic prevents many researchers to come to Geneva.The conference CGI is organized from September 12 to September 16, 2022, by MIRALab at the Computer Research Centre (CUI) of the University of Geneva, in Switzerland.All presentations are online.In addition to the Visual Computer journal published by Springer, and the CAVW journal (Computer Animation and Virtual Worlds) published by Wiley, we have also included the twenty-three accepted papers in the VRIH journal (Virtual Reality and Intelligent Hardware journal published by Science Press).This special issue is composed of the six papers related to the topic of metaverse from these twenty-three accepted papers. Nadia Magnenat-Thalmann, Jinman Kim, George Papagiannakis, Daniel Thalmann, Bin Sheng 0001 |
Virtual Real. Intell. Hardw. | 5 |
| 2022 | DSD-MatchingNet: Deformable Sparse-to-Dense Feature Matching for Learning Accurate CorrespondencesabstractExploring the correspondences across multi-view images is the basis of many computer vision tasks. However, most existing methods are limited on accuracy under challenging conditions. In order to learn more robust and accurate correspondences, we propose the DSD-MatchingNet for local feature matching in this paper. First, we develop a deformable feature extraction module to obtain multi-level feature maps, which harvests contextual information from dynamic receptive fields. The dynamic receptive fields provided by deformable convolution network ensures our method to obtain dense and robust correspondences. Second, we utilize the sparse-to-dense matching with the symmetry of correspondence to implement accurate pixel-level matching, which enables our method to produce more accurate correspondences. Experiments have shown that our proposed DSD-MatchingNet achieves a better performance on image matching benchmark, as well as on visual localization benchmark. Specifically, our method achieves 91.3% mean matching accuracy on HPatches dataset and 99.3% visual localization recalls on Aachen Day-Night dataset. Yicheng Zhao, Han Zhang 0053, Ping Lu 0008, Ping Li 0016, Enhua Wu, Bin Sheng 0001 |
Virtual Real. Intell. Hardw. | 6 |
| 2021 | Dynamic Shadow Synthesis Using Silhouette Edge Optimization
Saba Ghazanfar Ali, Bin Sheng 0001, Ping Li 0016, Xiaoyu Chi, Jinman Kim, Lijuan Mao |
CGI | 4 |
| 2021 | Progressive Multi-scale Reconstruction for Guided Depth Map Super-Resolution via Deep Residual Gate Fusion Network
Bin Sheng 0001, Ping Li 0016, Xiaoyu Chi, Lijuan Mao |
CGI | 4 |
| 2021 | A Classification Network for Ocular Diseases Based on Structure Feature and Visual Attention
Yupeng Xu, Bin Sheng 0001, Lei Bi 0001, Jinman Kim, Xiangui He |
CGI | 4 |
| 2021 | Multi-Stream Fusion Network for Multi-Distortion Image Super-Resolution
Yupeng Xu, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim, Xiangui He |
CGI | 3 |
| 2021 | CFMNet: Coarse-to-Fine Cascaded Feature Mapping Network for Hair Attribute Transfer
Guisong Zhang, Chunpeng Yu, Jiaheng Zheng, Bin Sheng 0001 |
CGI | 5 |
| 2021 | DCNet: Dual-Task Cycle Network for End-to-End Image DehazingabstractSingle image dehazing is an important technology in the field of computer vision. In this paper, we propose an image dehazing via dual learning strategy, named dual-task cycle network (DCNet). The core of DCNet is a dual learning framework, which consists of two tasks: the dehazing task and the haze generation task. The dehazing task completes the image dehazing, while the haze generation task achieves the restoration from the dehazed image to the haze image and can form a cycle to provide additional supervision. Our method uses the duality between each task as a constraint to learn and train two tasks jointly, so that the effects of the dehazing model can be improved. Since the haze generation process does not depend on clear images, the DCNet can satisfy the requirements for limited supervision. Extensive experiments demonstrate that our DCNet performs favorably on haze removal. Yu Zhou 0066, Ping Li 0016, Xiaoyu Chi, Lei Ma 0008, Bin Sheng 0001 |
ICME | 6 |
| 2021 | AFF-Dehazing: Attention-based feature fusion network for low-light image DehazingabstractAbstract Images captured in haze conditions, especially at nighttime with low light, often suffer from degraded visibility, contrasts, and vividness, which makes it difficult to carry out the following vision tasks. In this article, we propose an attention‐based feature fusion network (AFF‐Dehazing) for low‐light image dehazing. Our method decomposes the low‐light image dehazing into two task‐independent streams containing four modules: image dehazing module, low‐light feature extractor module, feature fusion module, and image restoration module. The basic block of these modules is the proposed attention‐based residual dense block. Since the dual‐branch are used, AFF‐Dehazing can avoid learning the mixed degradation all‐in‐one and enhance the details of low‐light haze images. Extensive experiments show that our method surpasses previous state‐of‐the‐art image dehazing methods and low‐light enhancement methods by a very large margin both quantitatively and qualitatively. Yu Zhou 0066, Bin Sheng 0001, Ping Li 0016, Jinman Kim, Enhua Wu |
Comput. Animat. Virtual Worlds | 3 |
| 2021 | Compensating the vorticity loss during advection with an adaptive vorticity confinement forceabstractAbstract The advection step in grid‐based fluid simulation is prone to numerical dissipation, which results in loss of detail. How to improve the advection accuracy to preserve more fluid details is still challenging. On the other hand, a common way to enhance smoke details is to use vorticity confinement. However, most of the previous methods simply used a fine‐tuned scale factor ε to adjust the strength of the confinement force, which can only amplify existing vortex details and is easy to cause instability when ε is large. In this article, we proposed an adaptive vorticity confinement method, which does not suffer from the above problems, to compensate the vorticity loss during advection with little extra cost. The main idea is to first calculate a scale factor whose value depends on the vorticity loss during advection, and then use it to adaptively control the vorticity confinement force for vorticity compensation with high stability. The experiment results show the effectiveness and efficiency of our method. Jian Zhu 0001, Silong Li, Ruichu Cai, Guoheng Huang, Bin Sheng 0001, Enhua Wu |
Comput. Animat. Virtual Worlds | 6 |
| 2021 | Cost-effective broad learning-based ultrasound biomicroscopy with 3D reconstruction for ocular anterior segmentation
Saba Ghazanfar Ali, Bin Sheng 0001, Huating Li, Po Yang 0001, Khan Muhammad 0001, Geng Yang 0003 |
Multim. Tools Appl. | 3 |
| 2021 | Multiview High Dynamic Range Image Synthesis Using Fuzzy Broad Learning SystemabstractCompared with the normal low dynamic range (LDR) images, the high dynamic range (HDR) images provide more dynamic range and image details. Although the existing techniques for generating the HDR images have a good effect for static scenes, they usually produce artifacts on the HDR images for dynamic scenes. In recent years, some learning-based approaches are used to synthesize the HDR images and obtain good results. However, there are also many problems, including the deficiency of explaining and the time-consuming training process. In this article, we propose a novel approach to synthesize multiview HDR images through fuzzy broad learning system (FBLS). We use a set of multiview LDR images with different exposure as input and transfer corresponding Takagi-Sugeno (TS) fuzzy subsystems; then, the structure is expanded in a wide sense in the "enhancement groups" which transfer from the TS fuzzy rules with nonlinear transformation. After integrating fuzzy subsystems and enhancement groups with the trained-well weight, the HDR image is generated. In FBLS, applying the incremental learning algorithm and the pseudoinverse method to compute the weights can greatly reduce the training time. In addition, the fuzzy system has better interpretability. In the learning process, IF-THEN fuzzy rules can effectively help the model to detect the artifacts and reject them in the final HDR result. These advantages solve the problem of existing deep-learning methods. Furthermore, we set up a new dataset of multiview LDR images with corresponding HDR ground truth to train our system. Our experimental results show that our system can synthesize high-quality multiview HDR images, which has a higher training speed than other learning methods. Bin Sheng 0001, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Cybern. | 2 |
| 2021 | GreenSea: Visual Soccer Analysis Using Broad Learning SystemabstractModern soccer increasingly places trust in visual analysis and statistics rather than only relying on the human experience. However, soccer is an extraordinarily complex game that no widely accepted quantitative analysis methods exist. The statistics collection and visualization are time consuming which result in numerous adjustments. To tackle this issue, we developed GreenSea, a visual-based assessment system designed for soccer game analysis, tactics, and training. The system uses a broad learning system (BLS) to train the model in order to avoid the time-consuming issue that traditional deep learning may suffer. Users are able to apply multiple views of a soccer game, and visual summarization of essential statistics using advanced visualization and animation that are available. A marking system trained by BLS is designed to perform quantitative analysis. A novel recurrent discriminative BLS (RDBLS) is proposed to carry out long-term tracking. In our RDBLS, the structure is adjusted to have better performance on the binary classification problem of the discriminative model. Several experiments are carried out to verify that our proposed RDBLS model can outperform the standard BLS and other methods. Two studies were conducted to verify the effectiveness of our GreenSea. The first study was on how GreenSea assists a youth training coach to assess each trainee's performance for selecting most potential players. The second study was on how GreenSea was used to help the U20 Shanghai soccer team coaching staff analyze games and make tactics during the 13th National Games. Our studies have shown the usability of GreenSea and the values of our system to both amateur and expert users. Bin Sheng 0001, Ping Li 0016, Lijuan Mao, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2021 | Optic Disk and Cup Segmentation Through Fuzzy Broad Learning System for Glaucoma ScreeningabstractGlaucoma is an ocular disease that causes permanent blindness if not cured at an early stage. Cup-to-disk ratio (CDR), obtained by dividing the height of optic cup (OC) with the height of optic disk (OD), is a widely adopted metric used for glaucoma screening. Therefore, accurately segmenting OD and OC is crucial for calculating a CDR. Most methods have employed deep learning methods for the segmentation of OD and OC. However, these methods are very time consuming. In this article, we present a new fuzzy broad learning system-based technique for OD and OC segmentation with glaucoma screening. We comprehensively integrated extracting a region of interest from RGB images, data augmentation, extracting red and green channel images, and inputting them to the two separate fuzzy broad learning system-based neural networks for segmenting the OD and OC, respectively, and then calculated CDR. Experiments show that our fuzzy broad learning system-based technique outperforms many state-of-the-art methods. Riaz Ali, Bin Sheng 0001, Ping Li 0016, Huating Li, Po Yang 0001, Younhyun Jung, Jinman Kim, C. L. Philip Chen |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Modified GAN-CAED to Minimize Risk of Unintentional Liver Major Vessels Cutting by Controlled Segmentation Using CTA/SPET-CTabstractThis article substantially advances upon state-of-the-art to enhance liver vessels segmentation accuracy by leveraging advantages of synthetic PET-CT (SPET-CT) images in addition to computed tomography angiography (CTA) volumes. Our setup makes a hybrid solution of modified generative adversarial network-convolutional autoencoder (GAN-cAED) combining synthetic ability of GAN to deliver SPET-CT images with generative ability of cAED network in terms of latent learning to more refined segmentation of major liver vessels. We improve time complexity through a novel concept of controlled segmentation by introducing a threshold metric to stop segmentation up to a desired level. The innovative concept of controlled vessel segmentation with a stopping criterion via variant threshold levels will help surgeons to avoid unintentional major blood vessels cutting, reducing the risk of excessive blood loss. Clinically, such solutions offer computer-aided liver surgeries and drug treatment evaluation in a CTA-only environment, shorten the requirement of radioactive and expensive fused PET-CT images. Muhammad Nadeem Cheema, Anam Nazir, Po Yang 0001, Bin Sheng 0001, Ping Li 0016, Huating Li, Xiaoer Wei, Harry Qin, Jinman Kim, David Dagan Feng |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Globally and Locally Semantic Colorization via Exemplar-Based Broad-GANabstractGiven a target grayscale image and a reference color image, exemplar-based image colorization aims to generate a visually natural-looking color image by transforming meaningful color information from the reference image to the target image. It remains a challenging problem due to the differences in semantic content between the target image and the reference image. In this paper, we present a novel globally and locally semantic colorization method called exemplar-based conditional broad-GAN, a broad generative adversarial network (GAN) framework, to deal with this limitation. Our colorization framework is composed of two sub-networks: the match sub-net and the colorization sub-net. We reconstruct the target image with a dictionary-based sparse representation in the match sub-net, where the dictionary consists of features extracted from the reference image. To enforce global-semantic and local-structure self-similarity constraints, global-local affinity energy is explored to constrain the sparse representation for matching consistency. Then, the matching information of the match sub-net is fed into the colorization sub-net as the perceptual information of the conditional broad-GAN to facilitate the personalized results. Finally, inspired by the observation that a broad learning system is able to extract semantic features efficiently, we further introduce a broad learning system into the conditional GAN and propose a novel loss, which substantially improves the training stability and the semantic similarity between the target image and the ground truth. Extensive experiments have shown that our colorization approach outperforms the state-of-the-art methods, both perceptually and semantically. Haoxuan Li 0004, Bin Sheng 0001, Ping Li 0016, Riaz Ali, C. L. Philip Chen |
IEEE Trans. Image Process. | 2 |
| 2021 | Structure-Aware Motion Deblurring Using Multi-Adversarial Optimized CycleGANabstractRecently, Convolutional Neural Networks (CNNs) have achieved great improvements in blind image motion deblurring. However, most existing image deblurring methods require a large amount of paired training data and fail to maintain satisfactory structural information, which greatly limits their application scope. In this paper, we present an unsupervised image deblurring method based on a multi-adversarial optimized cycle-consistent generative adversarial network (CycleGAN). Although original CycleGAN can handle unpaired training data well, the generated high-resolution images are probable to lose content and structure information. To solve this problem, we utilize a multi-adversarial mechanism based on CycleGAN for blind motion deblurring to generate high-resolution images iteratively. In this multi-adversarial manner, the hidden layers of the generator are gradually supervised, and the implicit refinement is carried out to generate high-resolution images continuously. Meanwhile, we also introduce the structure-aware mechanism to enhance the structure and detail retention ability of the multi-adversarial network for deblurring by taking the edge map as guidance information and adding multi-scale edge constraint functions. Our approach not only avoids the strict need for paired training data and the errors caused by blur kernel estimation, but also maintains the structural information better with multi-adversarial learning and structure-aware mechanism. Comprehensive experiments on several benchmarks have shown that our approach prevails the state-of-the-art methods for blind image motion deblurring. Jie Chen 0097, Bin Sheng 0001, Ping Li 0016, Ping Tan 0002, Tong-Yee Lee |
IEEE Trans. Image Process. | 3 |
| 2021 | NHBS-Net: A Feature Fusion Attention Network for Ultrasound Neonatal Hip Bone SegmentationabstractUltrasound is a widely used technology for diagnosing developmental dysplasia of the hip (DDH) because it does not use radiation. Due to its low cost and convenience, 2-D ultrasound is still the most common examination in DDH diagnosis. In clinical usage, the complexity of both ultrasound image standardization and measurement leads to a high error rate for sonographers. The automatic segmentation results of key structures in the hip joint can be used to develop a standard plane detection method that helps sonographers decrease the error rate. However, current automatic segmentation methods still face challenges in robustness and accuracy. Thus, we propose a neonatal hip bone segmentation network (NHBS-Net) for the first time for the segmentation of seven key structures. We design three improvements, an enhanced dual attention module, a two-class feature fusion module, and a coordinate convolution output head, to help segment different structures. Compared with current state-of-the-art networks, NHBS-Net gains outstanding performance accuracy and generalizability, as shown in the experiments. Additionally, image standardization is a common need in ultrasonography. The ability of segmentation-based standard plane detection is tested on a 50-image standard dataset. The experiments show that our method can help healthcare workers decrease their error rate from 6%-10% to 2%. In addition, the segmentation performance in another ultrasound dataset (fetal heart) demonstrates the ability of our network. Ruhan Liu, Mengyao Liu 0004, Bin Sheng 0001, Huating Li, Ping Li 0016, Haitao Song 0001, Ping Zhang 0016, Lixin Jiang, Dinggang Shen |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Hybrid Refinement-Correction Heatmaps for Human Pose EstimationabstractIn this paper, we present a method (Hybrid-Pose) to improve human pose estimation in images. We adopt Stacked Hourglass Networks to design two convolutional neural network models, RNet for pose refinement and CNet for pose correction. The CNet (Correction Network) guides the pose refinement RNet (Refinement Network) to correct the joint location before generating the final pose. Each of the two models is composed of four hourglasses, and each hourglass generates a group of detection heatmaps for the joints. The RNet model hourglasses have the same structure. However, the CNet model is designed with hourglasses of different structures for pose guidance. Since the pose estimation in RGB images is very sensitive to the image scene, our proposed approach generates multiple outputs of detection heatmaps to broaden the searching scope for the correct joints locations. We use the RNet model to refine the joints locations in each hourglass stage horizontally, then the heatmaps of each stage are fused with the heatmaps of all the CNet model hourglasses vertically in a hybrid manner. Our method shows competitive results with the existing state-of-the-art approaches on MPII and FLIC benchmark datasets. Although our proposed method focuses on improving single-person pose estimation, we also show the influence of this improvement on multi-person pose estimation by detecting multiple people using SSD detector, then estimating the pose of each person individually. Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng |
IEEE Trans. Multim. | 2 |
| 2021 | Utilizing Two-Phase Processing With FBLS for Single Image DerainingabstractRain removal from a single image is a challenging problem and has attracted much attention in recent years. In this paper, we revisit the single image deraining problem, and present a novel solution. The central idea of our solution is to merge the merits of two-phase processing methods and the Fuzzy Broad Learning System (FBLS). Specifically, our solution first uses the dehazing algorithm to preprocess the input rainy image and separates it into the detail layer and the base layer. After that, it puts the Y-channel image of the detail layer into the FBLS to obtain the derained Y channel image, which is then combined with the Cb and Cr channel images to obtain the derained detail layer. Later, it fuses the derained detail layer and the base layer to get a preliminary derained image. Finally, it superimposes the details extracted from the dehazed image with some transparency on the preliminary result, obtaining the final result. Experimental results based on both real and synthetic rainy images demonstrate that our proposed solution can outperform several state-of-the-art algorithms, while it consumes much less running time and training time, compared against the competitors. Xiao Lin 0012, Lizhuang Ma, Bin Sheng 0001, Zhi-Jie Wang 0009, Wansheng Chen |
IEEE Trans. Multim. | 3 |
| 2021 | Broad ColorizationabstractThe scribble- and example-based colorization methods have fastidious requirements for users, and the training process of deep neural networks for colorization is quite time-consuming. We instead proposed an automatic colorization approach with no dependence on user input and no need to endure long training time, which combines local features and global features of the input gray-scale images. Low-, mid-, and high-level features are united as local features representing cues existed in the gray-scale image. The global feature is regarded as data prior to guiding the colorization process. The local broad learning system is trained for getting the chrominance value of each pixel from the local features, which could be expressed as a chrominance map according to the position of pixels. Then, the global broad learning system is trained to refine the chrominance map. There are no requirements for users in our approach, and the training time of our framework is an order of magnitude faster than the traditional methods based on deep neural networks. To increase the user's subjective initiative, our system allows users to increase training data without retraining the system. Substantial experimental results have shown that our approach outperforms state-of-the-art methods. Yuxi Jin, Bin Sheng 0001, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Efficient Body Motion Quantification and Similarity Evaluation Using 3-D Joints Skeleton CoordinatesabstractEvaluating whole-body motion is challenging because of the articulated nature of the skeleton structure. Each joint moves in an unpredictable way with uncountable possibilities of movements direction under the influence of one or many of its parent joints. This paper presents a method for human motion quantification via three-dimensional (3-D) body joints coordinates. We calculate a set of metrics that influence the joints movement considering the motion of its parent joints without requiring prior knowledge of the motion parameters. Only the raw joints coordinates data of a motion sequence are needed to automatically estimate the transformation matrix of the joints between frames. We also consider the angles between limbs as a fundamental factor to follow the joints directions. We classify the joints motion as global motion and local motion. The global motion represents the joint movement according to a fixed joint, and the local motion represents the joint movement according to its first parent joint. In order to evaluate the performance of the proposed method, we also propose a comparison algorithm between two skeletons motions based on the quantified metrics. We measured the comparative similarity between the 3-D joints coordinates on Microsoft Kinect V2 and UTD-MHAD dataset. User studies were conducted to evaluate the performance under different factors. Various results and comparisons have shown that our method effectively quantifies and evaluates the motion similarity. Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | GPSD: generative parking spot detection using multi-clue recovery model
Bin Sheng 0001, Ping Li 0016, Enhua Wu |
Vis. Comput. | 3 |
| 2021 | Unsupervised face super-resolution via gradient enhancement and semantic guidance
Junshu Tang, Bin Sheng 0001, Lijuan Mao, Lizhuang Ma |
Vis. Comput. | 4 |
| 2020 | GARNet: Graph Attention Residual Networks Based on Adversarial Learning for 3D Human Pose Estimation
Bin Sheng 0001, Ping Li 0016 |
CGI | 3 |
| 2020 | Broad-Classifier for Remote Sensing Scene Classification with Spatial and Channel-Wise Attention
Yunna Liu, Han Zhang 0053, Bin Sheng 0001, Ping Li 0016, Guangtao Xue |
CGI | 4 |
| 2020 | FIOU Tracker: An Improved Algorithm of IOU Tracker in Video with a Lot of Background Inferences
Guhao Qiu, Han Zhang 0053, Bin Sheng 0001, Ping Li 0016 |
CGI | 4 |
| 2020 | Dynamic Shadow Rendering with Shadow Volume Optimization
Zhibo Fu, Han Zhang 0053, Po Yang 0001, Bin Sheng 0001, Lijuan Mao |
CGI | 6 |
| 2020 | Hierarchical Rendering System Based on Viewpoint Prediction in Virtual Reality
Ping Lu 0008, Ping Li 0016, Jinman Kim, Bin Sheng 0001, Lijuan Mao |
CGI | 5 |
| 2020 | GPU-based Grass Simulation with Accurate Blade Reconstruction
Saba Ghazanfar Ali, Ping Lu 0008, Po Yang 0001, Bin Sheng 0001, Lijuan Mao |
CGI | 6 |
| 2020 | GHand: A Graph Convolution Network for 3D Hand Pose Estimation
Pengsheng Wang, Guangtao Xue, Ping Li 0016, Jinman Kim, Bin Sheng 0001, Lijuan Mao |
CGI | 5 |
| 2020 | Preserving Temporal Consistency in Videos Through Adaptive SLIC
Han Zhang 0053, Riaz Ali, Bin Sheng 0001, Ping Li 0016, Jinman Kim |
CGI | 3 |
| 2020 | 3D Geology Scene Exploring Base on Hand-Track Somatic Interaction
Ping Lu 0008, Ping Li 0016, Bin Sheng 0001, Lijuan Mao |
CGI | 5 |
| 2020 | Gaze-Contingent Rendering in Virtual Reality
Ping Lu 0008, Ping Li 0016, Bin Sheng 0001, Lijuan Mao |
CGI | 4 |
| 2020 | Malocclusion Treatment Planning via PointNet Based Spatial Transformation Network
Xiaoshuang Li, Lei Bi 0001, Jinman Kim, Tingyao Li, Peng Li 0079, Bin Sheng 0001, David Dagan Feng |
MICCAI (3) | 7 |
| 2020 | Shape Mask Generator: Learning to Refine Shape Priors for Segmenting Overlapping Cervical Cytoplasms
Youyi Song, Lei Zhu 0003, Bai Ying Lei, Bin Sheng 0001, Qi Dou 0001, Harry Qin, Kup-Sze Choi |
MICCAI (4) | 4 |
| 2020 | ChefGAN: Food Image Generation from RecipesabstractAlthough significant progress has been made in generating images from the text by using generative adversarial networks (GANs), it is still challenging to deal with long text, which contains complex semantic information like recipes. This paper focuses on generating images with high visual realism and semantic consistency from the complex text of recipes. To achieve this, we propose a GANs based method termed ChefGAN. The critical concept of ChefGAN is that a joint image-recipe embedding model is used before the generation task to provide high-quality representations of recipes, and it acts as an extra regularization during the generation to improve semantic consistency. Two modules are designed for this image text embedding module (ITEM) and a cascaded image generation module (CIGM). The generation process is carried out in 3 steps: (1) Two encoders in ITEM are trained simultaneously to generate similar representations for each image-recipe pair. (2) CIGM generates images according to the representations from ITEM's text encoder. (3) The generated image is fed into ITEM's image encoder to calculate the similarity with the given recipe. This process can provide additional regularization effect other than the impact of a discriminator. To facilitate convergence, we applied a two-stage training strategy, which generates an image with low resolution and then one with high resolution in the CIGM module. Compared with other representative state-of-the-art methods, ChefGAN demonstrates better performance both in visual realism and semantic consistency. Siyuan Pan, Xuhong Hou, Huating Li, Bin Sheng 0001 |
ACM Multimedia | 5 |
| 2020 | Real-time hair simulation with heptadiagonal decomposition on mass spring system
Jianwei Jiang 0002, Bin Sheng 0001, Ping Li 0016, Lizhuang Ma, Xin Tong 0001, Enhua Wu |
Graph. Model. | 2 |
| 2020 | SPST-CNN: Spatial pyramid based searching and tagging of liver's intraoperative live views via CNN for minimal invasive surgery
Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Ping Li 0016, Huating Li, Po Yang 0001, Younhyun Jung, Harry Qin, David Dagan Feng |
J. Biomed. Informatics | 3 |
| 2020 | Embedding 3D models in offline physical environmentsabstractAbstract This article introduces a novel approach for embedding 3D models in offline physical environments using quick response (QR) codes. Unlike conventional methods, we consider settings where 3D models cannot be retrieved from a remote server. Our method involves generating octree models from voxelized 3D models and storing them in QR codes using a space‐efficient data structure. This allows storing 3D models that are both intelligible and purposeful on standard QR codes while addressing the major storage constraint that is present in offline situations. Furthermore, we explore 3D convolutional neural networks (CNN) and autoencoders (AE) to compress 3D models with high resolutions where using octrees alone does not suffice. To the best of our knowledge, our AE network is the first to employ octrees to further compress its encoded data. Through user‐friendly desktop and mobile applications, we allow users to encode, decode and visualize 3D models in augmented reality (AR) using QR codes, thus experiment with our methods. The proposed approach enables unique applications and future research in ubiquitous computing, 3D data compression and transmission, 3D AEs, AR and Virtual Reality, low‐cost autonomous robots, and 3D printing. Egemen Ertugrul, Han Zhang 0053, Ping Lu 0008, Ping Li 0016, Bin Sheng 0001, Enhua Wu |
Comput. Animat. Virtual Worlds | 6 |
| 2020 | Animating turbulent fluid with a robust and efficient high-order advection methodabstractAbstract The accuracy of advection has a great influence on the visual effect of fluid simulation. Constrained interpolation profile (CIP) method has been an important advection scheme because of its third‐order accuracy and the fact that it only needs to be performed over a compact stencil, but extending it to high‐dimensional advection equations is not easy, because it involves complex calculations and large memory overheads, and is usually unstable. In this article, we propose a stable and efficient three‐dimensional (3D) CIP scheme which can maintain high accuracy but requires low computation and memory cost. We first construct an efficient two‐dimensional (2D) CIP scheme based on dimensional splitting and local Taylor expansions, and then propose an effective way to extend it for 3D applications without decreasing the computational accuracy or affecting the stability. The experimental results show the advantages of our method over the state‐of‐the‐art advection schemes. Jian Zhu 0001, Silong Li, Ruichu Cai, Guoheng Huang, Bin Sheng 0001, Enhua Wu |
Comput. Animat. Virtual Worlds | 6 |
| 2020 | Domain-invariant interpretable fundus image quality assessment
Yaxin Shen, Bin Sheng 0001, Ruogu Fang, Huating Li, Skylar E. Stolte, Harry Qin, Weiping Jia, Dinggang Shen |
Medical Image Anal. | 2 |
| 2020 | Video flickering removal using temporal reconstruction optimization
Bin Sheng 0001, Ping Li 0016, Gaoqi He |
Multim. Tools Appl. | 3 |
| 2020 | Illumination-Invariant Video Cut-Out Using Octagon Sensitive OptimizationabstractThis paper presents an effective video cut-out approach, which can be utilized to segment the moving object in video shots. We first introduce the Octagon-Sensitive-Filtering (OSF) and its illumination invariant feature (IIF), which is computed on each pixel of the image via adding contributions from neighboring pixels. We integrate our IIF into the variational model and obtain the seeds during preprocessing to help address large displacement and illumination changes. An effective seed update method based on tracking-then-refinement based on IIF is presented to compensate for location ambiguities, and the strategy is effective to deal with illumination variances and objects deformation. Furthermore, we apply the IIF-based graph-cut to deal with fuzzy boundaries. Multiple experiments on quantitative challenging datasets have shown the robustness, high-quality video cut-out and efficiency of our approach to acute variances of illumination and complex motion. Jingye Wang, Bin Sheng 0001, Ping Li 0016, David Dagan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Depth-Aware Motion Deblurring Using Loopy Belief PropagationabstractMost motion-blurred images captured in the real world have spatially-varying point-spread functions, and some are caused by different positions and depth values, which cannot be handled by most state-of-the-art deblurring methods based on deconvolution. To overcome this problem, we propose a depth-aware motion blur model that treats a blurred image as an integration of a sequence of clear images. To restore the clear latent image, we extend the Richardson-Lucy method to incorporate our blur model with a given depth image. The empty holes in the depth image, caused by occlusion or device limitations, are fixed by PatchMatch-based depth filling. We regard the depth image as a Markov random field and select candidate labels by using belief propagation to set and smooth depth values for empty areas. Deblurring and depth filling are performed iteratively to refine the results. Our method can also be applied to real-world images with the assistance of motion estimation. The deblurring process is shown to be convergent; moreover, the number of iterations and the level of noise amplification are acceptable. The experimental results show that our method can not only handle depth-variant motion blur but also refine depth images. Bin Sheng 0001, Ping Li 0016, Xiaoxin Fang, Ping Tan 0002, Enhua Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Outdoor Shadow Estimating Using Multiclass Geometric Decomposition Based on BLSabstractIllumination is a significant component of an image, and illumination estimation of an outdoor scene from given images is still challenging yet it has wide applications. Most of the traditional illumination estimating methods require prior knowledge or fixed objects within the scene, which makes them often limited by the scene of a given image. We propose an optimization approach that integrates the multiclass cues of the image(s) [a main input image and optional auxiliary input image(s)]. First, Sun visibility is estimated by the efficient broad learning system. And then for the scene with visible Sun, we classify the information in the image by the proposed classification algorithm, which combines the geometric information and shadow information to make the most of the information. And we apply a respective algorithm for every class to estimate the illumination parameters. Finally, our approach integrates all of the estimating results by the Markov random field. We make full use of the cues in the given image instead of an extra requirement for the scene, and the qualitative results are presented and show that our approach outperformed other methods with similar conditions. Bin Sheng 0001, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Cybern. | 3 |
| 2020 | Automated Decision Support System for Lung Cancer Detection and Classification via Enhanced RFCN With Multilayer Fusion RPNabstractDetection of lung cancer at early stages is critical, in most of the cases radiologists read computed tomography (CT) images to prescribe follow-up treatment. The conventional method for detecting nodule presence in CT images is tedious. In this article, we propose an enhanced multidimensional region-based fully convolutional network (mRFCN) based automated decision support system for lung nodule detection and classification. The mRFCN is used as an image classifier backbone for feature extraction along with the novel multilayer fusion region proposal network (mLRPN) with position-sensitive score maps being explored. We applied a median intensity projection to leverage three-dimensional information from CT scans and introduced deconvolutional layer to adopt proposed mLRPN in our architecture to automatically select the potential region of interest. Our system has been trained and evaluated using LIDC dataset, and the experimental results showed promising detection performance in comparison to the state-of-the-art nodule detection/classification methods, achieving a sensitivity of 98.1% and classification accuracy of 97.91%. Anum Masood, Bin Sheng 0001, Po Yang 0001, Ping Li 0016, Huating Li, Jinman Kim, David Dagan Feng |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | OFF-eNET: An Optimally Fused Fully End-to-End Network for Automatic Dense Volumetric 3D Intracranial Blood Vessels SegmentationabstractIntracranial blood vessels segmentation from computed tomography angiography (CTA) volumes is a promising biomarker for diagnosis and therapeutic treatment in cerebrovascular diseases. These segmentation outputs are a fundamental requirement in the development of automated decision support systems for preoperative assessment or intraoperative guidance in neuropathology. The state-of-the-art in medical image segmentation methods are reliant on deep learning architectures based on convolutional neural networks. However, despite their popularity, there is a research gap in the current deep learning architectures optimized to address the technical challenges in blood vessel segmentation. These challenges include: (i) the extraction of concrete brain vessels close to the skull; and (ii) the precise marking of the vessel locations. We propose an Optimally Fused Fully end-to-end Network (OFF-eNET) for automatic segmentation of the volumetric 3D intracranial vascular structures. OFF-eNET comprises of three modules. In the first module, we exploit the up-skip connections to enhance information flow, and dilated convolution for detailed preservation of spatial feature map that are designed for thin blood vessels. In the second module, we employ residual mapping along with inception module for speedy network convergence and richer visual representation. For the third module, we make use of the transferred knowledge in the form of cascaded training strategy to gradually optimize the three segmentation stages (basic, complete, and enhanced) to segment thin vessels located close to the skull. All these modules are designed to be computationally efficient. Our OFF-eNET, evaluated using 70 CTA image volumes, resulted in 90.75% performance in the segmentation of intracranial blood vessels and outperformed the state-of-the-art counterparts. Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Huating Li, Ping Li 0016, Po Yang 0001, Younhyun Jung, Harry Qin, Jinman Kim, David Dagan Feng |
IEEE Trans. Image Process. | 3 |
| 2020 | Intrinsic Image Decomposition with Step and Drift Shading SeparationabstractDecomposing an image into the shading and reflectance layers remains challenging due to its severely under-constrained nature. We present an approach based on illumination decomposition that recovers the intrinsic images without additional information, e.g., depth or user interaction. Our approach is based on the rationale that the shading component contains the step and drift channels simultaneously. We decompose the illumination into two channels: the step shading, corresponding to the sharp shading changes due to cast shadow or abrupt shape changes; the drift shading, accounting for the smooth shading variations due to gradual illumination changes or slow shape changes. Due to such transformation of turning the conventional assumption that shading has smoothness as reasonable prior, our model has the advantages in handling real images, especially with the cast shadows or strong shape edges. We also apply a much stricter edge classifier along with a reinforcement process to enhance our method. We formulate the problem using a two-parameter energy function and split it into two energy functions corresponding to the reflectance and step shading. Experiments on the MIT dataset, the IIW dataset and the MPI Sintel dataset have shown the success of our approach over the state-of-the-art methods. Bin Sheng 0001, Ping Li 0016, Yuxi Jin, Ping Tan 0002, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Depth of Field Rendering Using Multilayer-Neighborhood OptimizationabstractDepth of field (DOF) is utilized widely to deliver artistic effects in photography. However, existing post-processing techniques for rendering DOF effects introduce visual artifacts such as color leakage, blurring discontinuity, and the partial occlusion problems which limit the application of DOF. Traditionally, occluded pixels are ignored or not well estimated although they might make key contributions to images. In this paper, we propose a new filtering approach which takes approximated occluded pixels into account to synthesize the DOF effects for images. In our approach, images are separated into different layers based on depth. Besides, we utilize adaptive PatchMatch method to estimate the intensities of occluded pixels, especially in the background region. We again propose a new multilayer-neighborhood optimization to estimate occluded pixels contributions and render the images. Finally, we apply gathering filter to achieve the rendered images with elite DOF effects. Multiple experiments have shown that our approach can handle color leakage, blurring discontinuity and partial occlusion problem while providing high-quality DOF rendering effects. Benxuan Zhang, Bin Sheng 0001, Ping Li 0016, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Simplified non-locally dense network for single-image dehazing
Zhuoliang Hu, Bin Sheng 0001, Ping Li 0016, Jinman Kim, Enhua Wu |
Vis. Comput. | 3 |
| 2020 | Simulation of multi-solvent stains on textile
Lei Ma 0008, Yanyun Chen, Guangzheng Fei, Bin Sheng 0001, Enhua Wu |
Vis. Comput. | 5 |
| 2020 | On attaining user-friendly hand gesture interfaces to control existing GUIsabstractBackground Hand gesture interfaces are dedicated programs that principally perform hand tracking and hand gesture prediction to provide alternative controls and interaction methods. They take advantage of one of the most natural ways of interaction and communication, proposing novel input and showing great potential in the field of the human-computer interaction. Developing a flexible and rich hand gesture interface is known to be a time-consuming and arduous task. Previously published studies have demonstrated the significance of the finite-state-machine (FSM) approach when mapping detected gestures to GUI actions. Methods In our hand gesture interface, we broadened the FSM approach by utilizing gesture-specific attributes, such as distance between hands, distance from the camera, and time of occurrences, to enable users to perform unique GUI actions. These attributes are obtained from hand gestures detected by the RealSense SDK employed in our hand gesture interface. By means of these gesture-specific attributes, users can activate static gestures and perform them as dynamic gestures. We also provided supplementary features to enhance the efficiency, convenience, and user-friendliness of our hand gesture interface. Moreover, we developed a complementary application for recording hand gestures by capturing hand keypoints in depth and color images to facilitate the generation of hand gesture datasets. Results We conducted a small-scale user survey with fifteen subjects to test and evaluate our hand gesture interface. Anonymous feedback obtained from the users indicates that our hand gesture interface is adequately facile and self-explanatory to use. In addition, we received constructive feedback about minor flaws regarding the responsiveness of the interface. Conclusions We proposed a hand gesture interface along with key concepts to attain user-friendliness and effectiveness in the control of existing GUIs. Egemen Ertugrul, Ping Li 0016, Bin Sheng 0001 |
Virtual Real. Intell. Hardw. | 3 |
| 2020 | An accurate multi-modal biometric identification system for person identification via fusion of face and finger print
Sidra Aleem, Po Yang 0001, Saleha Masood, Ping Li 0016, Bin Sheng 0001 |
World Wide Web | 5 |
| 2019 | Optic Disc and Cup Segmentation Based on Enhanced SegNetabstractDue to imbalanced distributed and restricted medical resources, reliable analysis for medical images is hard to come by, and it is impractical to only rely on human beings to do all the analysis, which is time-consuming and not economic. Application of computer vision techniques in such fields emerges as the situation requires. In this paper, we use deep learning segmentation algorithm to segment the optic disc and the cup from each other and from the rest of the ophthalmoscopy photographs. For a better performance, we change the loss function and crop as a way of data augmentation. The segmentation results can be used to calculate the cup-to-disc ratio (CDR), which is further used to diagnose glaucoma. Challenges such as over-fitting, biased dataset, and poor generalization of the model exist in front of us. We illustrate our model and associated methods dealing with these challenges. Lianyi Wu, Yelin Shi, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim |
CASA | 4 |
| 2019 | Detect Glaucoma with Image Segmentation and Transfer LearningabstractIn this paper, we aim to automatically detect glaucoma via deep learning. To do that, we need to calculate the cup-to-disc ratio (CDR) on fine segmented retina images. To get precise segmentation, we implemented SegNet together with adversarial discriminative domain adaptation (ADDA), the former is a famous artificial neural network with encoder-decoder architecture used in image segmentation area and the latter is a transfer learning method for domain adaptation. We are the first to combine them together to detect glaucoma on test dataset which have different brightness from our training dataset. We thoroughly evaluated the proposed method with various loss functions, normal cross entropy loss, weighted cross entropy loss and dice coefficient loss included. And we show that dice loss is the best for this task. Last but not least, our experiments on transfer learning have shown that our ADDA method reduces the mean square error (MSE) between the CDR of our segmentation and annotations greatly. Lianyi Wu, Yelin Shi, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim |
CASA | 4 |
| 2019 | Deep Intrinsic Image Decomposition Using Joint Parallel Learning
Yuan Yuan 0022, Bin Sheng 0001, Ping Li 0016, Lei Bi 0001, Jinman Kim, Enhua Wu |
CGI | 2 |
| 2019 | Dynamic Region Division for Adaptive Learning Pedestrian CountingabstractAccurate pedestrian counting algorithm is critical to eliminate insecurity in the congested public scenes. However, counting pedestrians in crowded scenes often suffer from severe perspective distortion. In this paper, basing on the straightline double region pedestrian counting method, we propose a dynamic region division algorithm to keep the completeness of counting objects. Utilizing the object bounding boxes obtained by YoloV3 and expectation division line of the scene, the boundary for nearby region and distant one is generated under the premise of retaining whole head. Ulteriorly, appropriate learning models are applied to count pedestrians in each obtained region. In the distant region, a novel inception dilated convolutional neural network is proposed to solve the problem of choosing dilation rate. In the nearby region, YoloV3 is used for detecting the pedestrian in multi-scale. Accordingly, the total number of pedestrians in each frame is obtained by fusing the result in nearby and distant regions. A typical subway pedestrian video dataset is chosen to conduct experiment in this paper. The result demonstrate that proposed algorithm is superior to existing machine learning based methods in general performance. Gaoqi He, Zhenwei Ma, Binhao Huang, Bin Sheng 0001, Yubo Yuan 0001 |
ICME | 4 |
| 2019 | Automatic Computer Aided System for Lung Cancer in Chest CTs Using MD-RFCN Combined with Tri-Level Region Proposal NetworkabstractPulmonary cancer is one of the major causes of deaths caused by cancer around the globe. Early stage lung cancer detection can prove to be essential for the patients, for which the computed tomography (CT) images are analyzed by the radiologists to determine the presence of nodules and diagnose the disease. Conventional techniques used by the radiologists for nodule detection in CT images is time-consuming and inefficient; to assist in the diagnosis process and further enhance its efficiency and accuracy, decision support systems have been developed in the past few years. In our paper, we proposed a Multi-Dimension Region-based Fully Convolutional Network based decision support system for detection and classification of lung nodule. The Multi-Dimension RFCN serves as an image classifier backbone for our feature extraction step in addition to the proposed Tri-Level Region Proposal Network (3L-RPN) along with the position-sensitive score maps (PSSM) being explored. A novel median intensity projection method is used to leverage the multi-dimensional information from CT images and introduced an additional deconvolutional layer to adopt the proposed Tri-Level Region Proposal Network in our architecture to automatically identify the potential Region of Interest. We trained and evaluated our proposed decision support system using LIDC-IDRI dataset. The evaluation results demonstrated the high level performance of our proposed model in comparison to the state-of-the-art nodule detection and classification methods by attaining classification accuracy of 97.61% and sensitivity of 97.4%. Anum Masood, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, Jinman Kim |
INDIN | 2 |
| 2019 | Abdominal Adipose Tissue Segmentation in MRI with Double Loss Function Collaborative Learning
Siyuan Pan, Xuhong Hou, Huating Li, Bin Sheng 0001, Ruogu Fang, Yuxin Xue, Weiping Jia, Harry Qin |
MICCAI (6) | 4 |
| 2019 | ADMM-Based Decentralized Electric Vehicle Charging with Trip Duration LimitsabstractWith the large-scale deployment of Electric Vehicles (EVs), the unbalanced distribution of charging needs and random charging behaviors cause charging stations (CSs) congestion. This degrades EV drivers' quality of experience by extending charging waiting time and increasing charging fee. Thus, EV owners are facing a critical issue on how to decrease the cost of charging, which consists of two parts: charging duration and charging fee. A great deal of existing work is confined to finding CSs to optimize the two parts individually. However, it still remains unexplored how to jointly minimize charging duration and charging fee under an overall time limit (i.e., deadline) of a scheduled trip. The problem is the focus of this paper. First, we formulate this problem as a 0-1 Integer Linear Programming problem and show its NP-Hardness. Then, we propose an efficient distributed algorithm based on the Alternating Direction Method of Multipliers (ADMM). The algorithm decomposes the original problem into sub-problems that can be solved locally and in parallel between charging stations and the global coordinator. Finally, we carry out extensive simulations based on real-life transport network data, and the results show that the proposed approach brings significant cost savings over existing ones. Gaoqi He, Zhifu Chai, Xingjian Lu, Fanxin Kong, Bin Sheng 0001 |
RTSS | 5 |
| 2019 | An Investigation of 3D Human Pose Estimation for Learning Tai Chi: A Human Factor PerspectiveabstractIn this article, we propose a Tai Chi training system based on pose estimation using Convolutional Neural Networks (CNNs) called iTai-Chi. Our system aims to overcome the disadvantages of insufficient accurate feedback in traditional teaching methods such as one-to-many tutorial and video watching. With the specially trained neural network, our iTai-Chi system can estimate learners’ poses more accurately compared to Kinect V2. In our system, user’s motion is evaluated through comparison with the template motion. The evaluated results are presented to the user to locate the error in their motions and help their correction. To verify the effectiveness of our system, we carried out a series of user studies. Results reflect that the iTai-Chi system successfully improve users’ performance in movement accuracy. Also, our system assists elder Tai Chi practitioners and students without prior knowledge to overcome learning obstacles and improve their skills. The users agreed that our system is interesting and supportive for their Tai Chi learning. Aouaidjia Kamel, Bowen Liu 0015, Ping Li 0016, Bin Sheng 0001 |
Int. J. Hum. Comput. Interact. | 4 |
| 2019 | SRNPD: Spatial rendering network for pencil drawing stylizationabstractAbstract Pencil drawing is a simple yet effective way to depict what people see by clearly presenting details of the scene. Existing methods usually extract strokes of the input image and adjust the result image tone to make it look like a pencil drawing. However, they do not consider the quality of the stroke image and the geometry information of lines in the stroke image, which unavoidably results in the violation of original essential structures and in a flatten pencil drawing with unrealistic appearance. We put forward a spatial rendering network for pencil drawing stylization. Spatial stroke images are extracted from the image pyramid by a single‐shot bottom‐up neural network to improve the quality of these stroke images. Unlike the former tone adjustment–based methods, we analyze perceptual cues of strokes at different stroke image levels and use the obtained geometry information to constrain the stroke shading procedure. The final pencil drawing result is achieved by the stroke shading fusion of different levels' shading results. The effectiveness of our spatial rendering network for pencil drawing stylization is demonstrated by an ablation study, comparison to the state of the art, and a user study. Yuxi Jin, Ping Li 0016, Bin Sheng 0001, Yongwei Nie, Jinman Kim, Enhua Wu |
Comput. Animat. Virtual Worlds | 3 |
| 2019 | Multiview-coherent disocclusion synthesis using connected regions optimizationabstractAbstract Handling of missing areas is a key step for depth‐based rendering to synthesize virtual views. Existing methods usually consider finding candidate pixels from only one reference view to fill missing areas. However, the information provided by one reference view is restricted by the position of the view. By utilizing two reference views located on both the left and right sides of the virtual view, we propose to synthesize the missing areas at hole level with connected regions optimization. To avoid the appearance of ghost boundary, we apply morphological operations to generate a boundary band map for the depth map, which restricts the warping of the virtual view. We use a binary map to mark unknown pixels in a hole, label the connected unknown regions, and count the area of each connected region, which decide the order in our enhanced inpainting synthesis. Besides, we separate the foreground and background regions of the depth map to constrain the searching of candidate pixels. Multiple experiments on virtual view synthesis have shown the effectiveness and high quality of our multiview‐coherent disocclusion synthesis. Ping Li 0016, Yuxi Jin, Bin Sheng 0001, Di Lin 0002, Yongwei Nie, Enhua Wu |
Comput. Animat. Virtual Worlds | 3 |
| 2019 | Retinal Vessel Segmentation Using Minimum Spanning Superpixel Tree DetectorabstractThe retinal vessel is one of the determining factors in an ophthalmic examination. Automatic extraction of retinal vessels from low-quality retinal images still remains a challenging problem. In this paper, we propose a robust and effective approach that qualitatively improves the detection of low-contrast and narrow vessels. Rather than using the pixel grid, we use a superpixel as the elementary unit of our vessel segmentation scheme. We regularize this scheme by combining the geometrical structure, texture, color, and space information in the superpixel graph. And the segmentation results are then refined by employing the efficient minimum spanning superpixel tree to detect and capture both global and local structure of the retinal images. Such an effective and structure-aware tree detector significantly improves the detection around the pathologic area. Experimental results have shown that the proposed technique achieves advantageous connectivity-area-length (CAL) scores of 80.92% and 69.06% on two public datasets, namely, DRIVE and STARE, thereby outperforming state-of-the-art segmentation methods. In addition, the tests on the challenging retinal image database have further demonstrated the effectiveness of our method. Our approach achieves satisfactory segmentation performance in comparison with state-of-the-art methods. Our technique provides an automated method for effectively extracting the vessel from fundus images. Bin Sheng 0001, Ping Li 0016, Shuangjia Mo, Huating Li, Xuhong Hou, Harry Qin, Ruogu Fang, David Dagan Feng |
IEEE Trans. Cybern. | 1 |
| 2019 | Fast and Accurate Retinal Identification System: Using Retinal Blood Vasculature LandmarksabstractThe expansion of automation techniques and increased risk of identity theft have led emphasis on the tremendous need of automated identification system. Due to the high recognition accuracy and robustness to changes in human physiology, retinal biometric identification system has drawn much attention in this research field. In this paper, we aim to propose an automatic fast and accurate retinal identification system for the multisample dataset. The proposed approach uses a hybrid segmentation technique to segment out both thick/thin vessels for effectively balancing the difference of wavelet response between thick/thin blood vessels. As a result, recognition accuracy is improved. A Principle Component Analysis-based feature processing approach is proposed for efficiently reducing the dimensionality of a large number of vessels features. It significantly reduces computation time and accelerates the matching process in the retinal identification system. The proposed technique is validated on DRIVE, STARE, VARIA, RIDB, HRF, Messidor, DIARETDB0, and a large multisample per subject database created by authors using the images provided by Dr. Chen (Shanghai Jiao Tong University Affiliated Sixth People Hospital). Experimental results demonstrated that the proposed approach outperforms other existing techniques. Segmentation achieves an overall accuracy of 99.65% with the recognition rate of 99.40% on all these databases. Sidra Aleem, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, David Dagan Feng |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Illumination-Guided Video Composition via Gradient Consistency OptimizationabstractVideo composition aims at cloning a patch from the source video into the target scene to create a seamless and harmonious blending frame sequence. Previous work in video composition usually suffer from artifacts around the blending region and spatial-temporal consistency when illumination intensity varies in the input source and target video. We propose an illumination-guided video composition method via a unified spatial and temporal optimization framework. Our method can produce globally consistent composition results and maintain the temporal coherency. We first compute a spatial-temporal blending boundary iteratively. For each frame, the gradient field of the target and source frames are mixed adaptively based on gradients and inter-frame color difference. The temporal consistency is further obtained by optimizing luminance gradients throughout all the composition frames. Moreover, we extend the mean-value cloning by smoothing discrepancies between the source and target frames, then eliminate the color distribution overflow exponentially to reduce falsely blending pixels. Various experiments have shown the effectiveness and high-quality performance of our illumination-guided composition. Jingye Wang, Bin Sheng 0001, Ping Li 0016, Yuxi Jin, David Dagan Feng |
IEEE Trans. Image Process. | 2 |
| 2019 | Deep Color Guided Coarse-to-Fine Convolutional Network Cascade for Depth Image Super-ResolutionabstractDepth image super-resolution is a significant yet challenging task. In this paper, we introduce a novel deep color guided coarse-to-fine convolutional neural network (CNN) framework to address this problem. First, we present a datadriven filter method to approximate the ideal filter for depth image super-resolution instead of hand-designed filters. Based on large data samples, the filter learned is more accurate and stable for upsampling depth image. Second, we introduce a coarse-to-fine CNN to learn different sizes of filter kernels. In coarse stage, larger filter kernels are learned by CNN to achieve crude high-resolution depth image. As to fine stage, the crude high-resolution depth image is used as the input so that smaller filter kernels are learned to gain more accurate results. Benefit from this network, we can progressively recover the high frequency details. Third, we construct a color guidance strategy that fuses color difference and spatial distance for depth image upsampling. We revise the interpolated high-resolution depth image according to the corresponding pixels in highresolution color maps. Guided by color information, the depth of high-resolution image obtained can alleviate texture copying artifacts and preserve edge details effectively. Quantitative and qualitative experimental results demonstrate our state-of-the-art performance for depth map super-resolution. Bin Sheng 0001, Ping Li 0016, Weiyao Lin, David Dagan Feng |
IEEE Trans. Image Process. | 2 |
| 2019 | Segmentation of Overlapping Cytoplasm in Cervical Smear Images via Adaptive Shape Priors Extracted From Contour FragmentsabstractWe present a novel approach for segmenting overlapping cytoplasm of cells in cervical smear images by leveraging the adaptive shape priors extracted from cytoplasm's contour fragments and shape statistics. The main challenge of this task is that many occluded boundaries in cytoplasm clumps are extremely difficult to be identified and, sometimes, even visually indistinguishable. Given a clump where multiple cytoplasms overlap, our method starts by cutting its contour into a set of contour fragments. We then locate the corresponding contour fragments of each cytoplasm by a grouping process. For each cytoplasm, according to the grouped fragments and a set of known shape references, we construct its shape and, then, connect the fragments to form a closed contour as the segmentation result, which is explicitly constrained by the constructed shape. We further integrate the intensity and curvature information, which is complementary to the shape priors extracted from contour fragments, into our framework to improve the segmentation accuracy. We propose to iteratively conduct fragments grouping, shape constructing, and fragments connecting for progressively refining the shape priors and improving the segmentation results. We extensively evaluate the effectiveness of our method on two typical cervical smear datasets. The experimental results demonstrate that our approach is highly effective and consistently outperforms the state-of-the-art approaches. The proposed method is general enough to be applied to other similar microscopic image segmentation tasks, where heavily overlapped objects exist. Youyi Song, Lei Zhu 0003, Harry Qin, Bai Ying Lei, Bin Sheng 0001, Kup-Sze Choi |
IEEE Trans. Medical Imaging | 5 |
| 2019 | Deep Convolutional Neural Networks for Human Action Recognition Using Depth Maps and PosturesabstractIn this paper, we present a method (Action-Fusion) for human action recognition from depth maps and posture data using convolutional neural networks (CNNs). Two input descriptors are used for action representation. The first input is a depth motion image that accumulates consecutive depth maps of a human action, whilst the second input is a proposed moving joints descriptor which represents the motion of body joints over time. In order to maximize feature extraction for accurate action classification, three CNN channels are trained with different inputs. The first channel is trained with depth motion images (DMIs), the second channel is trained with both DMIs and moving joint descriptors together, and the third channel is trained with moving joint descriptors only. The action predictions generated from the three CNN channels are fused together for the final action classification. We propose several fusion score operations to maximize the score of the right action. The experiments show that the results of fusing the output of three channels are better than using one channel or fusing two channels only. Our proposed method was evaluated on three public datasets: 1) Microsoft action 3-D dataset (MSRAction3D); 2) University of Texas at Dallas-multimodal human action dataset; and 3) multimodal action dataset (MAD) dataset. The testing results indicate that the proposed approach outperforms most of existing state-of-the-art methods, such as histogram of oriented 4-D normals and Actionlet on MSRAction3D. Although MAD dataset contains a high number of actions (35 actions) compared to existing action RGB-D datasets, this paper surpasses a state-of-the-art method on the dataset by 6.84%. Aouaidjia Kamel, Bin Sheng 0001, Po Yang 0001, Ping Li 0016, Ruimin Shen, David Dagan Feng |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2019 | Deep Neural Representation Guided Face Sketch SynthesisabstractFace sketch synthesis shows great applications in a lot of fields such as online entertainment and suspects identification. Existing face sketch synthesis methods learn the patch-wise sketch style from the training dataset containing photo-sketch pairs. These methods manipulate the whole process directly in the field of RGB space, which unavoidably results in unsmooth noises at patch boundaries. If denoising methods are used, the sketch edges would be blurred and face structures could not be restored. Recent researches of feature maps, which are the outputs of a certain neural network layer, have achieved great success in texture synthesis and artistic image generation. In this paper, we reformulate the face sketch synthesis problem into a neural network feature maps based optimization task. Our results accurately capture the sketch drawing style and make full use of the whole stylistic information hidden in the training dataset. Unlike former feature map based methods, we utilize the Enhanced 3D PatchMatch and cross-layer cost aggregation methods to obtain the target feature maps for the final results. Multiple experiments have shown that our approach imitates hand-drawn sketch style vividly, and has high-quality visual effects on CUHK, AR, XM2VTS and CUFSF face sketch datasets. Bin Sheng 0001, Ping Li 0016, Chenhao Gao, Kwan-Liu Ma |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Structure-preserving image completion with multi-level dynamic patches
Bowen Liu 0015, Ping Li 0016, Bin Sheng 0001, Yongwei Nie, Enhua Wu |
Vis. Comput. | 3 |
| 2018 | Action Recognition With Coarse-to-Fine Deep Feature Integration and Asynchronous FusionabstractAction recognition is an important yet challenging task in computer vision. In this paper, we propose a novel deep-based framework for action recognition, which improves the recognition accuracy by: 1) deriving more precise features for representing actions, and 2) reducing the asynchrony between different information streams. We first introduce a coarse-to-fine network which extracts shared deep features at different action class granularities and progressively integrates them to obtain a more accurate feature representation for input actions. We further introduce an asynchronous fusion network. It fuses information from different streams by asynchronously integrating stream-wise features at different time points, hence better leveraging the complementary information in different streams. Experimental results on action recognition benchmarks demonstrate that our approach achieves the state-of-the-art performance. Weiyao Lin, Ke Lu 0002, Bin Sheng 0001, Jianxin Wu 0001, Bingbing Ni, Hongkai Xiong |
AAAI | 4 |
| 2018 | Parallel Pencil Drawing Stylization via Structure-Aware OptimizationabstractThis paper presents a non-photorealistic rendering technique for stylizing a photograph in the pencil drawing style, which can well preserve the fine structure of the original image. We first construct a structure map from the anti-color image of the original image to model the detailed underlying fine structure of the original image, and generate a coarse pencil drawing image with line integral convolution. Then we refine the coarse pencil drawing image with the structure map for enrich the structure information. The presented algorithm is highly parallel allowing a real-time performance with GPU implementation. Experimental results show that our approach can produce more attractive and impressive pencil drawing effects with a variety of photographs. Yuxi Jin, Bin Sheng 0001, Ping Li 0016, Hanqiu Sun |
CASA | 3 |
| 2018 | Voxelized Facial Reconstruction Using Deep Neural NetworkabstractThis paper presents an approach to predicting variation tendency of human faces with regard to cranium changes based on deep learning. Our work focuses on generating individual customized facial models with high plausibility. Inspired by the performance of encoder-decoder convolutional neural network, the core trainable predicting engine of our learning network is designed for three-dimension voxelized data representation as the encoder-decoder structure and the encoder part is similar to the 7 layers of VGG16 network. To take full consideration of the cranium changes and features of original human face, a novel formation of channeled volumetric data structure is presented, and also the corresponding sub and up-sampling strategies for volume data. Our encoder-decoder neural network consumes discrete 3-channel volume data and generates 1-channel volume data as predicted post-variation human face. This framework is quantified with clinical dataset and it shows that its' performance improves in comparison with the state-of-the-art technologies. Xiaoshuang Li, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng |
CGI | 2 |
| 2018 | Group Re-Identification: Leveraging and Integrating Multi-Grain InformationabstractThis paper addresses an important yet less-studied problem: re-identifying groups of people in different camera views. Group re-identification (Re-ID) is very challenging since it is not only interfered by view-point and human pose variations in the traditional single-object Re-ID tasks, but also suffers from group layout and group member variations. To handle these issues, we propose to leverage the information of multi-grain objects: individual person and subgroups of two and three people inside a group image. We compute multi-grain representations to characterize the appearance and spatial features of multi-grain objects and evaluate the importance weight of each object for group Re-ID, so as to handle the interferences from group dynamics. We compute the optimal group-wise matching by using a multi-order matching process based on the multi-grain representation and importance weights. Furthermore, we dynamically update the importance weights according to the current matching results and then compute a new optimal group-wise matching. The two steps are iteratively conducted, yielding the final matching results.Experimental results on various datasets demonstrate the effectiveness of our approach. Weiyao Lin, Bin Sheng 0001, Ke Lu 0002, Junchi Yan, Jingdong Wang 0001, Errui Ding, Hongkai Xiong |
ACM Multimedia | 3 |
| 2018 | Accelerated robust Boolean operations based on hybrid representations
Bin Sheng 0001, Bowen Liu 0015, Ping Li 0016, Hongbo Fu 0001, Lizhuang Ma, Enhua Wu |
Comput. Aided Geom. Des. | 1 |
| 2018 | Better initialization for regression-based face alignment
Hengliang Zhu, Bin Sheng 0001, Zhiwen Shao, Yangyang Hao, Xiao-Nan Hou, Lizhuang Ma |
Comput. Graph. | 2 |
| 2018 | Biorthogonal Wavelet Surface Reconstruction Using Partial IntegrationsabstractAbstract We introduce a new biorthogonal wavelet approach to creating a water‐tight surface defined by an implicit function, from a finite set of oriented points. Our approach aims at addressing problems with previous wavelet methods which are not resilient to missing or nonuniformly sampled data. To address the problems, our approach has two key elements. First, by applying a three‐dimensional partial integration, we derive a new integral formula to compute the wavelet coefficients without requiring the implicit function to be an indicator function. It can be shown that the previously used formula is a special case of our formula when the integrated function is an indicator function. Second, a simple yet general method is proposed to construct smooth wavelets with small support. With our method, a family of wavelets can be constructed with the same support size as previously used wavelets while having one more degree of continuity. Experiments show that our approach can robustly produce results comparable to those produced by the Fourier and Poisson methods, regardless of the input data being noisy, missing or nonuniform. Moreover, our approach does not need to compute global integrals or solve large linear systems. Xiaohua Ren, Luan Lyu, Xiaowei He 0004, Wei Cao 0008, Zhi-Xin Yang 0001, Bin Sheng 0001, Yanci Zhang, Enhua Wu |
Comput. Graph. Forum | 6 |
| 2018 | Efficient non-incremental constructive solid geometry evaluation for triangular meshes
Bin Sheng 0001, Ping Li 0016, Hongbo Fu 0001, Lizhuang Ma, Enhua Wu |
Graph. Model. | 1 |
| 2018 | Automatic choroid layer segmentation using normalized graph cutabstractOptical coherence tomography is an immersive technique for depth analysis of retinal layers. Automatic choroid layer segmentation is a challenging task because of the low contrast inputs. Existing methodologies carried choroid layer segmentation manually or semi‐automatically. The authors proposed automated choroid layer segmentation based on normalised cut algorithm, which aims at extracting the global impression of images and treats the segmentation as a graph partitioning problem. Due to the structure complexity of retinal and choroid layers, the authors employed a series of pre‐processing to make the cut more deterministic and accurate. The proposed method divided the image into several patches and ran the normalised cut algorithm on every patch separately. The aim was to avoid insignificant vertical cuts and focus on horizontal cutting. After processing every patch, the authors acquired a global cut on the original image by combining all the patches. Later the authors measured the choroidal thickness which is highly helpful in the diagnosis of several retinal diseases. The results were computed on a total of 525 images of 21 real patients. Experimental results showed that the mean relative error rate of the proposed method was around 0.4 when compared with the manual segmentation performed by the experts. Saleha Masood, Bin Sheng 0001, Ping Li 0016, Ruimin Shen, Ruogu Fang |
IET Image Process. | 2 |
| 2018 | Computer-Assisted Decision Support System in Pulmonary Cancer detection and stage classification on CT images
Anum Masood, Bin Sheng 0001, Ping Li 0016, Xuhong Hou, Xiaoer Wei, Harry Qin, David Dagan Feng |
J. Biomed. Informatics | 2 |
| 2018 | Illumination-aware live videos background replacement using antialiasing optimization
Qiaoping Hu, Hanqiu Sun, Ping Li 0016, Ruimin Shen, Bin Sheng 0001 |
Multim. Tools Appl. | 5 |
| 2018 | Tracking soccer players using spatio-temporal context learning under multiple views
Linghan Zheng, Lijuan Mao, Bin Sheng 0001 |
Multim. Tools Appl. | 6 |
| 2018 | Video Decolorization Using Visual Proximity Coherence OptimizationabstractVideo decolorization is to filter out the color information while preserving the perceivable content in the video as much and correct as possible. Existing methods mainly apply image decolorization strategies on videos, which may be slow and produce incoherent results. In this paper, we propose a video decolorization framework that considers frame coherence and saves decolorization time by referring to the decolorized frames. It has three main contributions. First, we define decolorization proximity to measure the similarity of adjacent frames. Second, we propose three decolorization strategies for frames with low, medium, and high proximities, to preserve the quality of these three types of frames. Third, we propose a novel decolorization Gaussian mixture model to classify the frames and assign appropriate decolorization strategies to them based on their decolorization proximity. To evaluate our results, we measure them from three aspects: 1) qualitative; 2) quantitative; and 3) user study. We apply color contrast preserving ratio and C2G-SSIM to evaluate the quality of single frame decolorization. We propose a novel temporal coherence degree metric to evaluate the temporal coherence of the decolorized video. Compared with current methods, the proposed approach shows all around better performance in time efficiency, temporal coherence, and quality preservation. Yizhang Tao, Yiyi Shen, Bin Sheng 0001, Ping Li 0016, Rynson W. H. Lau |
IEEE Trans. Cybern. | 3 |
| 2018 | Clinical Report Guided Retinal Microaneurysm Detection With Multi-Sieving Deep LearningabstractNotice of Violation of IEEE Publication Principles"Clinical Report Guided Retinal Microaneurysm Detection With Multi-Sieving Deep Learning,"by Ling Dai, Ruogu Fang, Huating Li, Xuhong Hou, Bin Sheng, Qiang Wu, and Weiping Jiain the IEEE Transactions on Medical Imaging, vol. 37, no. 5, May 2018, pp. 1149-1161After careful and considered review of the content and authorship of this paper by a duly constituted expert committee, this paper has been found to be in violation of IEEE's Publication Principles.This paper contains significant portions of original text from the paper cited below. The original text was copied without attribution (including appropriate references to the original author(s) and/or paper title) and without permission."Mapping Visual Features to Semantic Profiles for Retrieval in Medical Imaging,"by Johannes Hofmanninger ; Georg Langsin the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015, pp. 457-465.Timely detection and treatment of microaneurysms is a critical step to prevent the development of vision-threatening eye diseases such as diabetic retinopathy. However, detecting microaneurysms in fundus images is a highly challenging task due to the low image contrast, misleading cues of other red lesions, and the large variation of imaging conditions. Existing methods tend to fail in face of the large intra-class variation and small inter-class variations for microaneurysm detection in fundus images. Recently, hybrid text/image mining computer-aided diagnosis systems have emerged to offer a promise of bridging the semantic gap between images and diagnostic information. In this paper, we focus on developing an interleaved deep mining technique to cope intelligently with the unbalanced microaneurysm detection problem. Specifically, we present a clinical report guided multi-sieving convolutional neural network, which leverages a small amount of supervised information in clinical reports to identify the potential microaneurysm regions via the image-to-text mapping in the feature space. These potential microaneurysm regions are then interleaved with fundus image information for multi-sieving deep mining in a highly unbalanced classification problem. Critically, the clinical reports are employed to bridge the semantic gap between low-level image features and high-level diagnostic information. We build an efficient microaneurysm detection framework based on the hybrid text/image interleaving and validate its performance on challenging clinical data sets acquired from diabetic retinopathy patients. Extensive evaluations are carried out in terms of fundus detection and classification. Experimental results show that our framework achieves 99.7% precision and 87.8% recall, comparing favorably with the state-of-the-art algorithms. Integration of expert domain knowledge and image information demonstrates the feasibility of reducing the difficulty of training classifiers under extremely unbalanced data distributions. Ruogu Fang, Huating Li, Xuhong Hou, Bin Sheng 0001, Weiping Jia |
IEEE Trans. Medical Imaging | 5 |
| 2017 | Image completion with dynamic patchesabstractThis paper presents an approach of structure-preserving image completion with dynamic patches. Existing image completion methods may generate unnatural abnormal structures or structure disorders due to limited patches and patterns availability. Our structure-preserving image completion utilizes objective function minimization considering the coherence not only within the image cavity but also with global constraints. A series of dynamic patch-based optimizations are applied to fulfill the cavity. Unlike traditional fixed-size patch-based methods, our image completion with competitive dynamic patch-matching mechanism provides more effective structure restoration. Parallel searching of different-sized patches is performed to retrieve optimal patches for completing the image cavity with nice structure preservation. The experiments show that the realistic completed images by our approach are visually pleasing with nice structural coherence. Bowen Liu 0015, Ping Li 0016, Bin Sheng 0001, Enhua Wu |
CGI | 3 |
| 2017 | Structure-Preserved Face Cartoonization
Chenhao Gao, Bin Sheng 0001, Ruimin Shen |
ICONIP (3) | 2 |
| 2017 | Retinal Microaneurysm Detection Using Clinical Report Guided Multi-sieving CNN
Bin Sheng 0001, Huating Li, Xuhong Hou, Weiping Jia, Ruogu Fang |
MICCAI (3) | 2 |
| 2017 | Intrinsic Image Decomposition Using Multi-Scale Measurements and SparsityabstractAbstract Automatic decomposition of intrinsic images, especially for complex real‐world images, is a challenging under‐constrained problem. Thus, we propose a new algorithm that generates and combines multi‐scale properties of chromaticity differences and intensity contrast. The key observation is that the estimation of image reflectance, which is neither a pixel‐based nor a region‐based property, can be improved by using multi‐scale measurements of image content. The new algorithm iteratively coarsens a graph reflecting the reflectance similarity between neighbouring pixels. Then multi‐scale reflectance properties are aggregated so that the graph reflects the reflectance property at different scales. This is followed by a L0 sparse regularization on the whole reflectance image, which enforces the variation in reflectance images to be high‐frequency and sparse. We formulate this problem through energy minimization which can be solved efficiently within a few iterations. The effectiveness of the new algorithm is tested with the Massachusetts Institute of Technology (MIT) dataset, the Intrinsic Images in the Wild (IIW) dataset, and various natural images. Shouhong Ding, Bin Sheng 0001, Xiao-Nan Hou, Lizhuang Ma |
Comput. Graph. Forum | 2 |
| 2017 | Dynamic RGB-to-CMYK conversion using visual contrast optimisationabstractAs the standard colour space used by printers, Cyan, Magenta, Yellow, Black (CMYK) colour model is a subtractive colour space used to describe the printing process. Existing CMYK conversion methods rely on static conversion table, which may not preserve the subtle visual structures of images, due to the local visual contrast loss caused by the static colour mapping. Therefore, the authors propose a novel dynamic Red, Green, Blue (RGB)‐to‐CMYK colour conversion, which utilises the weighted entropy to extract the pixels with filter response change dramatically. They obtain the image activity map by combining these pixels with high skin probability regions, and optimise the colour conversion of each pixel to ensure that the ink used for each pixel can be saved, while the visual contrast can be preserved with ink‐saving. In this way, their proposed technique can achieve dynamic CMYK colour conversion, in which the consumption of ink can be reduced without the loss of visual contrast. The experimental results have shown that their dynamic CMYK colour conversion saved 10–25% ink consumption compared with the static conversion method, while with high visual quality for the converted images. Zhenzhu Wang, Bin Sheng 0001, Ruimin Shen, Ping Li 0016 |
IET Image Process. | 3 |
| 2017 | Integrated tone and structure refinement for high-fidelity colour transferabstractA high‐fidelity colour transfer should align the colour distributions between images and meanwhile avoid the damage to the original structure. However, the traditional methods often fail to yield high‐fidelity transfer results due to some existing tone and structure artefacts. In this study, the authors propose a new framework to effectively integrate the tone and structure refinements of colour transfer. They develop the ideas of image decomposition and gradient guidance to perform tone reconstruction while protecting original structure. Its overall flow includes the five key steps: tone clustering, structure extraction, structure optimisation, gradient‐guided tone reconstruction, and structure restoration. Moreover, they propose an evaluation metric to measure the differences of tone and structure between images. They demonstrate the performance of the proposed method through a number of experiments in visual comparison and objective evaluation. Shouhong Ding, Bin Sheng 0001, Lizhuang Ma |
IET Image Process. | 3 |
| 2017 | Abdominal adipose tissues extraction using multi-scale deep neural network
Fei Jiang 0006, Huating Li, Xuhong Hou, Bin Sheng 0001, Ruimin Shen, Xiao-Yang Liu, Weiping Jia, Ping Li 0016, Ruogu Fang |
Neurocomputing | 4 |
| 2017 | Antialiased super-resolution with parallel high-frequency synthesis
Xudong Jiang 0003, Bin Sheng 0001, Weiyao Lin, Ping Li 0016, Lizhuang Ma, Ruimin Shen |
Multim. Tools Appl. | 2 |
| 2017 | Retinal optic disc localization using convergence tracking of blood vessels
Rui Wang 0130, Linghan Zheng, Chaoqun Xiong, Chunfang Qiu, Huating Li, Xuhong Hou, Bin Sheng 0001, Ping Li 0016 |
Multim. Tools Appl. | 7 |
| 2017 | Automatic diabetic retinopathy diagnosis using adjustable ophthalmoscope and multi-scale line operator
Meng Qu, Chun Ni, Mufan Chen, Linghan Zheng, Bin Sheng 0001, Ping Li 0016 |
Pervasive Mob. Comput. | 6 |
| 2017 | Colorization Using Neural Network EnsembleabstractThis paper investigates into the colorization problem, which converts a grayscale image to a colorful version. This is a difficult problem and normally requires manual adjustment to achieve artifact-free quality. For instance, it normally requires human-labeled color scribbles on the grayscale target image or a careful selection of colorful reference images. The recent learning-based colorization techniques automatically colorize a grayscale image using a single neural network. Since different scenes usually have distinct color styles, it is difficult to accurately capture the color characteristics using a single neural network. We propose a mixture learning model representing the presence of sub-color-style within an overall image data set. We, therefore, ensemble multiple neural networks to obtain better color estimation performance than could be obtained from any of the constituent neural network alone. A two-step colorization strategy is utilized as an adaptive color style clustering followed by a neural network ensemble. To ensure artifact-free quality, a joint bilateral filtering-based post-processing step is proposed. Numerous experiments demonstrate that our method generates high-quality results comparable with state-of-the-art algorithms. Zezhou Cheng, Qingxiong Yang, Bin Sheng 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Temporal Coherence-Based Deblurring Using Non-Uniform Motion OptimizationabstractNon-uniform motion blur due to object movement or camera jitter is a common phenomenon in videos. However, the state-of-the-art video deblurring methods used to deal with this problem can introduce artifacts, and may sometimes fail to handle motion blur due to the movements of the object or the camera. In this paper, we propose a non-uniform motion model to deblur video frames. The proposed method is based on superpixel matching in the video sequence to reconstruct sharp frames from blurry ones. To identify a suitable sharp superpixel to replace a blurry one, we enrich the search space with a non-uniform motion blur kernel, and use a generalized PatchMatch algorithm to handle rotation, scale, and blur differences in the matching step. Instead of using pixel-based or regular patch-based representation, we adopt a superpixel-based representation, and use color and motion to gather similar pixels. Our non-uniform motion blur kernels are estimated from the motion field of these superpixels, and our spatially varying motion model considers spatial and temporal coherence to find sharp superpixels. Experimental results showed that the proposed method can reconstruct sharp video frames from blurred frames caused by complex object and camera movements, and performs better than the state-of-the-art methods. Congbin Qiao, Rynson W. H. Lau, Bin Sheng 0001, Benxuan Zhang, Enhua Wu |
IEEE Trans. Image Process. | 3 |
| 2017 | Intrinsic image estimation using near-L0 sparse optimization
Shouhong Ding, Bin Sheng 0001, Lizhuang Ma |
Vis. Comput. | 2 |
| 2017 | Incremental collision-free feathering for animated surfaces
Le Liu 0004, Xuehui Liu, Bin Sheng 0001, Yanyun Chen, Enhua Wu |
Vis. Comput. | 3 |
| 2016 | Real-Time Cloud Simulation Using Lennard-Jones ApproximationabstractCloud simulation is important for creating images of outdoor scenes. However, the complexity of this natural phenomenon makes the simulation of large-scale clouds difficult in real time. In this paper, we present a new method for 3D cloud simulation in which cloud animation is simplified and simulated by approximating Lennard-Jones Potential. To solve the N-body problem in Lennard-Jones Potential, we minimized the interaction between particles by dividing the simulation space into many cells and we defined a cutoff distance to perform calculation between neighboring particles. Additionally, a separate distance is introduced between particles to maintain the stability in the Lennard-Jones system. Our experimental results demonstrate that our method is computationally inexpensive and suitable for real time applications where large-scale simulation of clouds is required. Akila Elhaddad, Feriel Elhaddad, Bin Sheng 0001, Hanqiu Sun, Enhua Wu |
CASA | 3 |
| 2016 | Foreground Object Sensing for Saliency DetectionabstractMany state-of-the-art saliency detection algorithms rely on the boundary prior, but these algorithms simply suppose the boundaries around an image as background regions. Here we propose a fast and effective algorithm for salient object detection. First, a novel method is proposed to approximately locate the foreground object by using the convex hull from Harris corner. On this basis, we divide the saliency values of different regions into two parts and generate the corresponding cue maps (foreground and background), which are combined into a convex hull prior map. Then a new prior based on distance to the convex hull center is proposed to replace the center prior. Finally, the convex hull prior map and the convex hull center-biased map are combined to be the saliency map, which is then optimized to get the final result. Compared with eighteen existing algorithms and tested on several datasets, the present algorithm performs well in terms of precision and recall. Hengliang Zhu, Bin Sheng 0001, Xiao Lin 0012, Yangyang Hao, Lizhuang Ma |
ICMR | 2 |
| 2016 | Structure-aware image inpainting using patch scale optimization
Chao Dai, Bin Sheng 0001, Jing Zhang 0041, Weiyao Lin, Yubo Yuan 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Image saliency detection using Gabor texture cues
Bin Sheng 0001, Jian-ning Liang, Jing Zhang 0041, Yubo Yuan 0001 |
Multim. Tools Appl. | 3 |
| 2016 | Accurate gaze tracking from single camera using gabor corner detector
Bin Sheng 0001, Wen Wu 0001, Lizhuang Ma, Ping Li 0016 |
Multim. Tools Appl. | 2 |
| 2016 | Image saliency detection based on rectangular-wave spectrum analysis
Bin Sheng 0001, Wen Wu 0001, Zezhou Cheng, Ruimin Shen |
Multim. Tools Appl. | 2 |
| 2016 | Tree-Based Visualization and Optimization for Image CollectionabstractThe visualization of an image collection is the process of displaying a collection of images on a screen under some specific layout requirements. This paper focuses on an important problem that is not well addressed by the previous methods: visualizing image collections into arbitrary layout shapes while arranging images according to user-defined semantic or visual correlations (e.g., color or object category). To this end, we first propose a property-based tree construction scheme to organize images of a collection into a tree structure according to user-defined properties. In this way, images can be adaptively placed with the desired semantic or visual correlations in the final visualization layout. Then, we design a two-step visualization optimization scheme to further optimize image layouts. As a result, multiple layout effects including layout shape and image overlap ratio can be effectively controlled to guarantee a satisfactory visualization. Finally, we also propose a tree-transfer scheme such that visualization layouts can be adaptively changed when users select different "images of interest." We demonstrate the effectiveness of our proposed approach through the comparisons with state-of-the-art visualization techniques. Xintong Han, Weiyao Lin, Mingliang Xu 0001, Bin Sheng 0001, Tao Mei 0001 |
IEEE Trans. Cybern. | 5 |
| 2015 | Real Time Learning Evaluation Based on Gaze TrackingabstractIn this paper, we present a system that extracts the information implied by eye movements and use this information to analyze students' learning behavior. Our system uses a common webcam to capture students' facial image sequences when they are learning in front of monitors. We then process these images and establish sequences of changing location of iris, which represent the movements of eyes. With the eye movement sequences, we train a HMM classifier that can analyze their pattern and generate learning status for any given moment in the lesson. These statuses could help the computers to get a better understanding about the students' intention and behavior during online learning. The status sequences of those who view the same lesson could also be used as a reference for teaching quality assessment. Jiayue Yi, Bin Sheng 0001, Ruimin Shen, Weiyao Lin, Enhua Wu |
CAD/Graphics | 2 |
| 2015 | Deep ColorizationabstractThis paper investigates into the colorization problem which converts a grayscale image to a colorful version. This is a very difficult problem and normally requires manual adjustment to achieve artifact-free quality. For instance, it normally requires human-labelled color scribbles on the grayscale target image or a careful selection of colorful reference images (e.g., capturing the same scene in the grayscale target image). Unlike the previous methods, this paper aims at a high-quality fully-automatic colorization method. With the assumption of a perfect patch matching technique, the use of an extremely large-scale reference database (that contains sufficient color images) is the most reliable solution to the colorization problem. However, patch matching noise will increase with respect to the size of the reference database in practice. Inspired by the recent success in deep learning techniques which provide amazing modeling of large-scale data, this paper re-formulates the colorization problem so that deep learning techniques can be directly employed. To ensure artifact-free quality, a joint bilateral filtering based post-processing step is proposed. Numerous experiments demonstrate that our method outperforms the state-of-art algorithms both in terms of quality and speed. Zezhou Cheng, Qingxiong Yang, Bin Sheng 0001 |
ICCV | 3 |
| 2015 | GPU-Accelerated Video Background Subtraction Using Gabor Detector
Lixia Qin, Bin Sheng 0001, Weiyao Lin, Wen Wu 0001, Ruimin Shen |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Vessel extraction from non-fluorescein fundus images using orientation-aware detector
Benjun Yin, Huating Li, Bin Sheng 0001, Xuhong Hou, Wen Wu 0001, Ping Li 0016, Ruimin Shen, Yuqian Bao, Weiping Jia |
Medical Image Anal. | 3 |
| 2015 | Image-based non-photorealistic rendering for realtime virtual sculpting
Ping Lu 0008, Bin Sheng 0001, Shengmei Luo, Xia Jia, Wen Wu 0001 |
Multim. Tools Appl. | 2 |
| 2015 | Saliency-Guided Color-to-Gray Conversion Using Region-Based OptimizationabstractImage decolorization is a fundamental problem for many real-world applications, including monochrome printing and photograph rendering. In this paper, we propose a new color-to-gray conversion method that is based on a region-based saliency model. First, we construct a parametric color-to-gray mapping function based on global color information as well as local contrast. Second, we propose a region-based saliency model that computes visual contrast among pixel regions. Third, we minimize the salience difference between the original color image and the output grayscale image in order to preserve contrast discrimination. To evaluate the performance of the proposed method in preserving contrast in complex scenarios, we have constructed a new decolorization data set with 22 images, each of which contains abundant colors and patterns. Extensive experimental evaluations on the existing and the new data sets show that the proposed method outperforms the state-of-the-art methods quantitatively and qualitatively. Shengfeng He, Bin Sheng 0001, Lizhuang Ma, Rynson W. H. Lau |
IEEE Trans. Image Process. | 3 |
| 2015 | Structure-aware QR Code abstraction
Siyuan Qiao, Xiaoxin Fang, Bin Sheng 0001, Wen Wu 0001, Enhua Wu |
Vis. Comput. | 3 |
| 2014 | Finding Coherent Motions and Semantic Regions in Crowd Scenes: A Diffusion and Clustering Approach
Weiyue Wang 0002, Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Jingdong Wang 0001, Bin Sheng 0001 |
ECCV (1) | 6 |
| 2014 | Saliency preserving decolorizationabstractThis paper presents a new decolorization approach using saliency including color and position information to maintain original contrast. Our approach includes three steps: (1) we propose a visual contrast model which quantifies saliency via pixel based contrast along with area based contrast, using both color difference and spatial relationship, (2) we put forward an energy function with the purpose of minimizing the gap between color image and corresponding grayscale's saliency, where grayscale is represented by a mapping function of color channels, (3) we accelerate the calculation of energy function by loosening or strengthening constraint. Experimental results on public benchmark images show that the use of saliency can preserve the contrast of the original color image better than state of the art decolorization methods. Mingqi Zhou, Bin Sheng 0001, Lizhuang Ma |
ICME | 2 |
| 2014 | Real-time depth-of-field rendering using single-layer compositionabstractABSTRACT In this paper, we propose a single‐layer post‐processing method for real‐time depth‐of‐field rendering that uses single‐layer composition. In the proposed method, blurring is achieved by gathering background pixels and scattering foreground pixels. Major artifacts in post‐filtering techniques such as intensity leakage and blurring discontinuity are reduced by using two different blurring functions and the controllable parameter in the gathering process. The method can be entirely implemented in GPU parallelization to achieve the real‐time performance required for virtual reality. The results of comparisons of our method with recent post‐processing methods in terms of rendering quality and rendering performance indicate that our method generates realistic natural images and is also the fastest in terms of frames per second. Copyright © 2014 John Wiley & Sons, Ltd. Xiaoxin Fang, Bin Sheng 0001, Wen Wu 0001, Zengzhi Fan, Lizhuang Ma |
Comput. Animat. Virtual Worlds | 2 |
| 2014 | Facial expression cloning with elastic and muscle models
Weiyao Lin, Bing Zhou 0003, Zhenzhong Chen 0001, Bin Sheng 0001, Jianxin Wu 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2014 | Image anti-aliasing techniques for Internet visual media processing: a reviewabstractAnti-aliasing is a well-established technique in computer graphics that reduces the blocky or stair-wise appearance of pixels. This paper provides a comprehensive overview of the anti-aliasing techniques used in computer graphics, which can be classified into two categories: post-filtering based anti-aliasing and pre-filtering based anti-aliasing. We discuss post-filtering based anti-aliasing algorithms through classifying them into hardware anti-aliasing techniques and post-process techniques for deferred rendering. Comparisons are made among different methods to illustrate the strengths and weaknesses of every category. We also review the utilization of anti-aliasing techniques from the first category in different graphic processing units, i.e., different NVIDIA and AMD series. This review provides a guide that should allow researchers to position their work in this important research area, and new research problems are identified. Xudong Jiang 0003, Bin Sheng 0001, Weiyao Lin, Lizhuang Ma |
J. Zhejiang Univ. Sci. C | 2 |
| 2014 | Perception-motivated multiresolution rendering on sole-cube maps
Bin Sheng 0001, Weiliang Meng, Hanqiu Sun, Wen Wu 0001, Enhua Wu |
Multim. Tools Appl. | 1 |
| 2014 | A New Network-Based Algorithm for Human Activity Recognition in VideosabstractIn this paper, a new network-transmission-based (NTB) algorithm is proposed for human activity recognition in videos. The proposed NTB algorithm models the entire scene as an error-free network. In this network, each node corresponds to a patch of the scene and each edge represents the activity correlation between the corresponding patches. Based on this network, we further model people in the scene as packages, while human activities can be modeled as the process of package transmission in the network. By analyzing these specific package transmission processes, various activities can be effectively detected. The implementation of our NTB algorithm into abnormal activity detection and group activity recognition are described in detail in this paper. Experimental results demonstrate the effectiveness of our proposed algorithm. Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Hanli Wang, Bin Sheng 0001, Hongxiang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2014 | Video Colorization Using Parallel Optimization in Feature SpaceabstractWe present a new scheme for video colorization using optimization in rotation-aware Gabor feature space. Most current methods of video colorization incur temporal artifacts and prohibitive processing costs, while this approach is designed in a spatiotemporal manner to preserve temporal coherence. The parallel implementation on graphics hardware is also facilitated to achieve realtime performance of color optimization. By adaptively clustering video frames and extending Gabor filtering to optical flow computation, we can achieve real-time color propagation within and between frames. Temporal coherence is further refined through user scribbles in video frames. The experimental results demonstrate that our proposed approach is efficient in producing high-quality colorized videos. Bin Sheng 0001, Hanqiu Sun, Marcus A. Magnor, Ping Li 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Depth-of-Field Rendering with Saliency-Based Bilateral FilteringabstractDepth of Field (DoF) is an indispensable feature of photo realistic rendering and photography retouching. In this paper, we propose an image-based rendering technique which can simulate the depth-of-field effect. The proposed technique can render the depth-of-field effect automatically without any interactions. Compared to the ordinary depth-of-field rendering technique, our algorithm is less time-consuming and needs no additional depth maps to assist the depth-of-field rendering. In our proposed algorithm, the saliency detection technique is employed to simulate the depth information. The flash-based technique is also introduced to promote the final depth-of-field rendering visual effect. Weichen Xue, Dong Xing, Bin Sheng 0001, Lizhuang Ma |
CAD/Graphics | 5 |
| 2013 | Fast vehicle detection based on feature and real-time predictionabstractThe vehicle identification is a key technology of vehicle automatic driving and assistance systems. This paper proposes a new fast vehicle detection method based on feature learning and real-time prediction by combining ARMA model and AdaBoost algorithm, which can be applied in car driver assistance systems for road detection and vehicle identification with a monocular camera. Experimental results show that our proposed algorithm can take the target's prior information into account, and extend AdaBoost algorithm in the time dimension that improve the accuracy of real-time detection to be faster and more accurate than the existing methods. Bin Sheng 0001, Lizhuang Ma |
ISCAS | 3 |
| 2013 | Sketch-based design for green geometry and image deformation
Bin Sheng 0001, Weiliang Meng, Hanqiu Sun, Enhua Wu |
Multim. Tools Appl. | 1 |
| 2013 | Temporally Coherent Video Saliency Using Regional Dynamic ContrastabstractSaliency detection for images and videos has become increasingly popular due to its wide applicability. In this paper, we present a new method that takes advantage of region-based visual dynamic contrast to generate temporally coherent video saliency maps. The concept of visual dynamics is formulated to represent both visual and motional variabilities of video content. Moreover, the regions are regarded as primitives for saliency computation by using spatiotemporal appearance contrasts. Then, region matching is performed across successive video frames to form temporally coherent regions, which are computed on the basis of spatiotemporal similarity in the visual dynamics of the different regions along the optical flow in the video. The region matching can effectively eliminate saliency discontinuities, particularly in the areas of oversegmentation that are otherwise highly problematic. The proposed approach is tested on a challenging set of video sequences and is compared with contemporary methods to demonstrate its superior performance in terms of its computational efficiency and ability to detect salient video content. Bin Sheng 0001, Lizhuang Ma, Wen Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | A Heat-Map-Based Algorithm for Recognizing Group Activities in VideosabstractIn this paper, a new heat-map-based algorithm is proposed for group activity recognition. The proposed algorithm first models human trajectories as series of heat sources and then applies a thermal diffusion process to create a heat map (HM) for representing the group activities. Based on this HM, a new key-point-based (KPB) method is used for handling the alignments among HMs with different scales and rotations. A surface-fitting (SF) method is also proposed for recognizing group activities. Our proposed HM feature can efficiently embed the temporal motion information of the group activities while the proposed KPB and SF methods can effectively utilize the characteristics of the HM for activity recognition. Section IV demonstrates the effectiveness of our proposed algorithms. Weiyao Lin, Hang Chu, Jianxin Wu 0001, Bin Sheng 0001, Zhenzhong Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Accurate Depth-of-Field Rendering Using Adaptive Bilateral Depth Filtering
Shang Wu 0001, Bin Sheng 0001, Feiyue Huang, Lizhuang Ma |
CVM | 3 |
| 2012 | Facial expression mapping based on elastic and muscle-distribution-based modelsabstractIn this paper, a new algorithm is proposed for facial expression mapping. The proposed algorithm first introduces a new elastic model to balance the global and local warping effects such that the impacts from facial feature differences between people can be avoided, thus more reasonable geometric warping results can be created. Furthermore, a muscle-distribution-based (MD) model is also proposed. The proposed MD model utilizes the muscle distribution information of the human face to evaluate and strengthen the facial illumination details. By this way, the impacts from human face difference as well as the effects of unsuitable noise filtering can be effectively alleviated. Experimental results show that our proposed algorithm can create obviously better facial expression results than the existing methods. Weiyao Lin, Bin Sheng 0001, Jianxin Wu 0001, Hongxiang Li 0001 |
ISCAS | 3 |
| 2012 | Image stylization with enhanced structure on GPU
Ping Li 0016, Hanqiu Sun, Bin Sheng 0001, Jianbing Shen |
Sci. China Inf. Sci. | 3 |
| 2012 | Seamlet carving for shape-aware image resizing
Xiao Lin 0012, Bin Sheng 0001, Lizhuang Ma, Yang Shen 0011 |
Sci. China Inf. Sci. | 2 |
| 2012 | A Customized Framework to Recompress Massive Internet Images
Shouhong Ding, Feiyue Huang, Yongjian Wu 0001, Bin Sheng 0001, Lizhuang Ma |
J. Comput. Sci. Technol. | 5 |
| 2012 | Fast Image Correspondence with Global Structure Projection
Qing-Liang Lin, Bin Sheng 0001, Yang Shen 0011, Lizhuang Ma |
J. Comput. Sci. Technol. | 2 |
| 2012 | Video composition by optimized 3D mean-value coordinatesabstractABSTRACT In this paper, we propose a new video composition method by 3D mean‐value coordinate (MVC). 2D MVCs have been widely used in image composition; however, when 2D MVC is applied to a video sequence directly, because of over‐blending and the lack of temporal consistency, some unnatural effects may appear in the final composite results. Although 3D Poisson editing can maintain spatial and temporal consistency, it also leads to high algorithm complexity. Instead of 3D Poisson editing, we use the 3D MVC to seamlessly blend a given source video patch into a target video sequence; this approach is able to achieve high‐performance blending with less computation. We show that the combination of alpha matte‐based approaches and our method can further refine the produced video when the boundaries of the source object and the target object are very different. Our algorithm can be paralleled and run on a graphics processing unit. The experimental results show that our method is effective and efficient. Copyright © 2012 John Wiley & Sons, Ltd. Yang Shen 0011, Xiao Lin 0012, Yan Gao 0004, Bin Sheng 0001, Qisong Liu |
Comput. Animat. Virtual Worlds | 4 |
| 2011 | MCGIM-Based Model Streaming for Realtime Progressive Rendering
Bin Sheng 0001, Weiliang Meng, Hanqiu Sun, Enhua Wu |
J. Comput. Sci. Technol. | 1 |
| 2010 | Efficient deformable geometry image-mapsabstractMultiresolution rendering of deformable models, performing fast rendering of global deformations while preserving local surface details, is usually a computation-costly and time-consuming process, because two non-trivial operators are involved. We propose a novel GPU-based adaptive rendering of deforming mesh sequence on sole-cube maps (SCM), which is a variant of geometry images built upon spherical parameterizations. We also introduce the differential coordinates to bound the local resampling error for supporting details preservation and view-dependent visualization. By precomputing the adaptive SCM texture atlas of deforming mesh sequences as well as their differential coordinates, we can map both the deformation and the level-of-details (LOD) operator to the GPU. The proposed algorithm enables us to reconstruct the deformed positions and sufficiently fine-scale approximations of deforming mesh sequences for efficient GPU processing. The GPU-friendly data structure and process allow us to render dynamically deforming 3D models with GPU parallelization, also our system improves the efficiency of the manipulations of node selection and boundary stitching, significantly alleviating the computing load on CPU. Bin Sheng 0001, Hanqiu Sun |
VRST | 1 |
| 2010 | Differential geometry images: remeshing and morphing with local shape preservation
Weiliang Meng, Bin Sheng 0001, Weiwei Lv, Hanqiu Sun, Enhua Wu |
Vis. Comput. | 2 |
| 2009 | Multi-level tree branch modeling and animationabstractWe present a new approach for quickly designing 3D models of botanical trees using an iterative addition of new nodes to the tree branch structure. This process is guided by the proximity of points marking volume density data captured from photographs. Numerical parameters provide the user controls that are consistent with the characteristics of trees in landscaping and make it possible to generate a wide variety of tree styles. Meanwhile we synthesize visually believable motions for the generated tree models affected by a wind field. Our system enables the simulation of tree animation, by introducing physically-based transformation matrix calculations for hierarchical branch patterns. The system also supports the tree-shaping modes in which many branches and leaves are generated by interactively designed their distribution density. Experimental results show that our approach can design a variety of reasonably natural-looking trees and their motions. Meng Yang 0011, Bin Sheng 0001, Enhua Wu, Hanqiu Sun |
CAD/Graphics | 2 |
| 2009 | Image-Based Material Restyling with Fast Non-local Means FilteringabstractThis paper presents a new GPU-based implementation of fast non-local means (NLM) filtering for material restyling. Our fast NLM filtering algorithm is able to achive realtime feedback of interactive image editing. Furthermore a novel material editing method based on our fast NLM filtering is proposed to change the material appearance of image-based objects. Given an input image, an alpha matte is created to differentiate the object from its background. After the automatic matting process, the object is removed from the background, and the 3D shape of the object is recovered. We use the fast non-local means (NLM) filtering to process the luminance channel of an image and obtain a pseudo-depth map that is sufficient for altering the material appearance of the observed object. We recover the gradient luminance maps for the region to be material-restyled; and change the material property in the region-of-interest by solving the Poisson equation. In this way, the new material is mapped onto the 3D shape, resulting in an object which appears to be made of a different material. Our new NLM filtering-based material editing is easy to implement in parallel with graphics hardware. The experimental results have demonstrated the satisfactory performance of our method. Bin Sheng 0001, Ping Li 0016, Hanqiu Sun |
ICIG | 1 |
| 2009 | Lumiproxy: A Hybrid Representation of Image-Based Models
Bin Sheng 0001, Jian Zhu 0001, Enhua Wu, Yanci Zhang |
J. Comput. Sci. Technol. | 1 |
| 2009 | Furstyling on angle-split shell texturesabstractAbstract This paper presents a new method for modeling and rendering fur with a wide variety of furstyles. We simulate virtual fur using shell textures—a multiple layers of textured slices for its generality and efficiency. As shell textures usually suffer from the inherent visual gap errors due to the uniform discretization nature, we present theangle‐split shell textures(ASST) approach, which classifies the shell textures into different types with different numbers of texture layers, by splitting the angle space of the viewing angles between fur orientation and view direction. Our system can render the fur with biological patterns, and utilizes vector field and scalar field on ASST to control the geometric variations of the furry shape. Users can intuitively shape the fur by applying the combing, blowing, and interpolating effects in real time. Our approach is intuitive to implement without using complex data structures, with real‐time performance for dynamic fur appearances. Copyright © 2009 John Wiley & Sons, Ltd. Bin Sheng 0001, Hanqiu Sun, Gang Yang 0007, Enhua Wu |
Comput. Animat. Virtual Worlds | 1 |
| 2008 | Sketching freeform meshes using graph rotation functions
Bin Sheng 0001, Enhua Wu, Hanqiu Sun |
Vis. Comput. | 1 |
| 2007 | Topology-Consistent Design for 3D Freeform Meshes with Harmonic Interpolation
Bin Sheng 0001, Enhua Wu |
ICEC | 1 |
| 2007 | Walking into Images: Virtual Plane Mosaics for Plenoptic ModelingabstractAn effective method of depth image based rendering is proposed by applying texture mapping onto virtual but pre-defined multiple planes in the scene, called virtual planes. The method allows the viewpoint either static or moving around, including crossing the plane of the source images. By this approach, the relief textures from depth images are mapped onto the virtual planes, and through the pre-warping process, the virtual planes are converted into standard polygonal textures. After the virtual plane mosaics, the resultant image that supports 3D objects and immersive scenes can be generated by polygonal texture mapping. In addition, both hardware and software implementation of the method can increase the power of conventional texture mapping in image based rendering. In particular, the scope of the viewpoint could be extended into the inner space of depth images, and as a result, a novel solution is provided for constructing real-time walkthrough systems as well as for panoramic modeling from an arbitrary viewpoint in the depth image space Bin Sheng 0001, Enhua Wu |
VR | 1 |
| 2007 | View-dependent mesh streaming using multi-chart geometry imagesabstractMany mesh streaming algorithms have focused on the transmission order of the polygon data with respect to the current viewpoint. In contrast to the conventional progressive streaming where the resolution of a model changes in the geometry space, we present an new approach which firstly partitions a mesh into several patches, then converts these patch into multi-chart geometry images(MCGIM). After all the MCGIM and normal map atlas are obtained by regular re-sampling, we could construct the regular quadtree-based hierarchical representation based on MCGIM. Experimental results have shown the effectiveness of our approach where one server streams the MCGIM texture atlas to the clients. Bin Sheng 0001, Enhua Wu |
VRST | 1 |