VLDB 2026 Research / reviewers in the wild / expert
Songhua Xu
dblp:88/3319
· DBLP profile ↗
110ranked-venue papers
27as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 8 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 15 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 8 first-author · 15 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AAHL: Attention-Guided Adaptive Hypergraph Learning for Multi-Label Image Classification
Zhengshen Gu, Songhua Xu |
ICPRAM | 5 |
| 2026 | Cross-scene hyperspectral image classification based on cross-domain feature extraction and category decision collaborative optimizationabstractCross-scene hyperspectral image classification aims to enable the model to complete the classification of unlabeled target domain data by learning from labeled source domain data. Aiming at the problem that most current cross-scene hyperspectral image classification algorithms do not fully consider the cross-domain feature representation and category decision boundary optimization, a cross-domain Feature Extraction and Category Decision collaborative optimization (FECD) network is proposed. First, an adaptive feature discovery based on dynamic masks is designed. In this mechanism, the dynamically scaled masks are applied to the 3D representation of source and target domain data to generate an informative feature space and enhance the cross-scene discrimination potential of the model. Second, a dual-stream convolutional cross-domain feature extraction based on Mamba stream and ViT stream is constructed. Long sequence modeling and convolutional attention mechanisms are used to capture cross-domain spectral features between pixel, and self-attention mechanisms and multi-scale convolution are used to excavate cross-domain space patterns of pixel. Finally, a category decision based on the co-optimization of dual-stream classifiers is implemented. The spectral and spatial boundaries learned by the dual streams are fused to optimize the category decision. Therefore, the risk of false labeling is avoided while obtaining more accurate category boundaries. Compared with seven state-of-the-art algorithms on three widely used datasets, FECD obtains better categorization results on three categorization metrics: OA, AA, and Kappa. Ronghua Shang, Yangyang Li 0001, Jie Feng 0003, Songhua Xu |
Expert Syst. Appl. | 5 |
| 2026 | Unsupervised feature selection based on adaptive latent representation learning and multi-group data similarity
Lizhuo Gao, Lei Liu 0014, Ronghua Shang, Dongzhu Feng, Yangyang Li 0001, Songhua Xu |
Neurocomputing | 7 |
| 2026 | Domain-consistent networks for cross-scene hyperspectral image classification
Ronghua Shang, Yangyang Li 0001, Jie Feng 0003, Songhua Xu |
Neurocomputing | 5 |
| 2026 | Structurally and attention-adaptive explainable deep learning for multimodal medical diagnosis using the unified fuzzy membership function
Songhua Xu, Muhammad Usman Aslam |
Inf. Sci. | 3 |
| 2026 | Improving the performance of medical image segmentation with instructive feature learning
Duwei Dai, Caixia Dong, Haolin Huang, Zongfang Li, Songhua Xu |
Medical Image Anal. | 6 |
| 2026 | Corrigendum to "A novel multi-attention, multi-scale 3D deep network for coronary artery segmentation" [Medical Image Analysis 85 (2023) 102745]
Caixia Dong, Songhua Xu, Duwei Dai, Yizhi Zhang, Zongfang Li |
Medical Image Anal. | 2 |
| 2026 | High-quality coronary artery segmentation via fuzzy logic modeling coupled with dynamic graph convolutional network
Caixia Dong, Duwei Dai, Yang Li 0104, Songhua Xu |
Pattern Recognit. | 4 |
| 2026 | Unsupervised feature selection based on dual-graph clustering learning and adaptive weighting
Ronghua Shang, Yangyang Li 0001, Songhua Xu |
Pattern Recognit. | 5 |
| 2026 | Large-Scale Multiview Clustering via Joint Learning of Anchor Representation and Multigraph AlignmentabstractThe anchor-based clustering method is currently a predominant technique for handling large-scale data. However, in multiview data, existing anchor-based methods face a key challenge: balancing individual anchor graph distinctiveness with final consistency. To address this challenge, we propose a large-scale multiview clustering (MVC) method via joint learning of anchor representation and multigraph alignment (ARMGA). Specifically, ARMGA introduces a unified framework that facilitates the concurrent learning of single-view anchor representations and virtual graph-based multigraph alignment. The approach aims to preserve the adaptability of anchor learning across different views, while ensuring the ultimate consistency of the merged anchor graph. Furthermore, ARMGA employs Schatten- $\boldsymbol {p}$ norm on the tensor formed by the adaptive anchor representation, originating from multigraph alignment, to reinforce cross-view consistency. This technique effectively leverages complementary information preserved across views to bolster the overall structure and consensus information. Ultimately, to attenuate the noise impact on the anchor representation matrix, ARMGA capitalizes on the cosine angle information from the low-rank representation as coefficients within the relationship matrix and efficiently reduces computational complexity through deductions. On nine datasets, ARMGA has exhibited a notable improvement in clustering performance indicators by 2%-10% over other algorithms, while also maintaining lower time complexity. Ronghua Shang, Jingya Liu, Jingyu Zhong, Songhua Xu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Unleashing Vision Foundation Models for Coronary Artery Segmentation: Parallel ViT-CNN Encoding and Variational Fusion
Caixia Dong, Duwei Dai, Xinyi Han, Zongfang Li, Songhua Xu |
MICCAI (5) | 7 |
| 2025 | Improved fuzzy control charts for monitoring defined health ranges using trapezoidal fuzzy numbers
Muhammad Usman Aslam, Songhua Xu, Zahid Rasheed, Muhammad Noorulamin |
Expert Syst. Appl. | 2 |
| 2025 | Robust multi-view subspace clustering via neighbor embedding on manifold and low-rank representation learning
Jiarui Kong, Jingya Liu, Ronghua Shang, Songhua Xu, Yangyang Li 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Group-spectral superposition and position self-attention transformer for hyperspectral image classification
Mingwei Hu, Sihan Hou, Ronghua Shang, Jie Feng 0003, Songhua Xu |
Expert Syst. Appl. | 6 |
| 2025 | Node classification based on structure migration and graph attention convolutional crossover network
Ruolin Li, Ronghua Shang, Songhua Xu |
Knowl. Based Syst. | 5 |
| 2025 | Few-Shot Learning Based on Embedded Self-Distillation and Adaptive Wasserstein Distance for Hyperspectral Image ClassificationabstractDue to the domain shift, it is challenging to achieve ideal experimental results for cross-domain few-shot learning (FSL) in hyperspectral image (HSI) classification. Most existing FSL algorithms are impacted by the limited samples, and they do not effectively leverage the representations from different layers of the network. Therefore, this article proposes an FSL based on embedded self-distillation and adaptive Wasserstein (ESAW-FSL) distance for HSI classification. First, the embedding self-distillation network is proposed in the feature extraction process of the source domain (SD) and the target domain (TD). The embedding self-distillation network utilizes self-distillation from different perspectives to get discriminative features. In the SD, the mask evaluation of embedded features is employed to guarantee the learning of guiding features. Second, a domain adaptation based on adaptive Wasserstein distance is designed to alleviate the domain shift problem between the domains. A lightweight feature correlation network learns the comprehensive cost matrix in the Wasserstein distance adaptively, and the obtained cost matrix helps achieve domain adaptation by an iterative algorithm. Finally, a focal loss based on double softening is adopted in the process of FSL. The probability is double softened to improve the ratio of correctly classifying hard samples. Experiments are conducted on three widely used hyperspectral datasets and compared with six state-of-the-art algorithms. The overall accuracy (OA) and average accuracy (AA) are achieved in multiple experiments, demonstrating the effectiveness of ESAW-FSL. Shizhe Shang, Ronghua Shang, Dongzhu Feng, Chao Wang 0099, Jie Feng 0003, Songhua Xu |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Edge-Enhanced Cascaded MRF for SAR Image SegmentationabstractMarkov Random Fields (MRF) effectively capture local contextual information by modeling the spatial dependencies between pixels, which helps highlight details and enhances segmentation smoothness. To fully exploit MRF for synthetic aperture radar (SAR) image segmentation, we propose a novel edge-enhanced cascaded MRF (ECMRF) approach. Specifically, we introduce multiple edge-constrained filters to emphasize SAR image boundaries and provide relatively clean features. Building on this, we present a cascaded MRF framework that sequentially integrates region-level and pixel-level segmentation with feature perturbation and fusion to generate the final segmentation output. The framework comprises four key components: (1) a region-level MRF, regulated by edge features, to achieve precise region segmentation; (2) a pixel-level MRF with selective label smoothing to refine edges and reduce noise clusters; (3) equal-channel feature perturbation to increase feature diversity; and (4) a random probability-based feature fusion scheme to merge the input features. Experimental results demonstrate that our ECMRF outperforms six state-of-the-art comparable methods, underscoring its competitive performance. Ronghua Shang, Kang Liu 0025, Jie Feng 0003, Chao Wang 0099, Songhua Xu, Yangyang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Knowledge Distillation Based on Adaptive Learning and Channel Amplification Features for PolSAR Image ClassificationabstractThe models currently used for Polarimetric Synthetic Aperture Radar (PolSAR) image classification tasks have problems such as complex network structures, poor distinction of detailed features, and fixed loss weights during the training process. In response to these problems, this paper proposes a PolSAR image classification method based on knowledge distillation using adaptive learning and channel amplification features. Firstly, this paper builds a knowledge distillation framework for PolSAR. Using a teacher network trained in advance that can acquire global knowledge to guide the student. This framework reduces the computational complexity and improves the classification accuracy of the student. Then, an adaptive loss weight learning mechanism is designed, which sets the weight of the Kullback-Leibler divergence loss during training into a learnable mode. The weight can be automatically adjusted according to the actual training situation of the student. Finally, a scheme for channel amplification to enhance features is proposed. This scheme obtains channel weights based on the student’s feature map information. These weights are amplified, strengthening the network’s ability to obtain feature information. Compared with the five PolSAR image classification algorithms, the method proposed in this paper uses lower computational complexity to obtain higher classification accuracy on the Flevoland, San Francisco, and Xi’an datasets. Ronghua Shang, Mingwei Hu, Lei Liu 0014, Jie Feng 0003, Songhua Xu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Multilabel Feature Selection via Shared Latent Sublabel Structure and Simultaneous Orthogonal Basis ClusteringabstractMultilabel feature selection solves the dimension distress of high-dimensional multilabel data by selecting the optimal subset of features. Noisy and incomplete labels of raw multilabel data hinder the acquisition of label-guided information. In existing approaches, mapping the label space to a low-dimensional latent space by semantic decomposition to mitigate label noise is considered an effective strategy. However, the decomposed latent label space contains redundant label information, which misleads the capture of potential label relevance. To eliminate the effect of redundant information on the extraction of latent label correlations, a novel method named SLOFS via shared latent sublabel structure and simultaneous orthogonal basis clustering for multilabel feature selection is proposed. First, a latent orthogonal base structure shared (LOBSS) term is engineered to guide the construction of a redundancy-free latent sublabel space via the separated latent clustering center structure. The LOBSS term simultaneously retains latent sublabel information and latent clustering center structure. Moreover, the structure and relevance information of nonredundant latent sublabels are fully explored. The introduction of graph regularization ensures structural consistency in the data space and latent sublabels, thus helping the feature selection process. SLOFS employs a dynamic sublabel graph to obtain a high-quality sublabel space and uses regularization to constrain label correlations on dynamic sublabel projections. Finally, an effective convergence provable optimization scheme is proposed to solve the SLOFS method. The experimental studies on the 18 datasets demonstrate that the presented method performs consistently better than previous feature selection methods. Ronghua Shang, Jingyu Zhong, Songhua Xu, Yangyang Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Attribute community detection based on attribute edges weights fusion and graph embedding factorization
Shuaize Yang, Ronghua Shang, Songhua Xu, Chao Wang 0099 |
Appl. Intell. | 4 |
| 2024 | Non-convex feature selection based on feature correlation representation and dual manifold optimization
Ronghua Shang, Lizhuo Gao, Haijing Chi, Jiarui Kong, Songhua Xu |
Expert Syst. Appl. | 6 |
| 2024 | Unsupervised feature selection method based on dual manifold learning and dual spatial latent representation
Ronghua Shang, Yangyang Li 0001, Songhua Xu |
Expert Syst. Appl. | 5 |
| 2024 | Double-dictionary learning unsupervised feature selection cooperating with low-rank and sparsity
Ronghua Shang, Jiuzheng Song, Lizhuo Gao, Mengyao Lu, Licheng Jiao, Songhua Xu, Yangyang Li 0001 |
Knowl. Based Syst. | 6 |
| 2024 | Ellipse IoU Loss: Better Learning for Rotated Bounding Box RegressionabstractRotated object detection is an important research content in the field of remote-sensing images. However, in the rotated object detection, the inconsistency between the loss function and the final detection metric has become an important factor restricting the improvement of detection accuracy. So, in this letter, an ellipse intersection over union (IoU) loss (EPIoU loss) is proposed to solve these problems. The EPIoU loss uses IoU between the bounding boxes’ inscribed ellipses, which is approximate to the original bounding box IoU. This loss function can jointly optimize the prediction box parameters and promote the model to locate the object better. Compared to the complex intersection of rotated rectangles, the intersection calculation of two rotated ellipses is simple. A unified and differentiable process is also designed to calculate EPIoU, which avoids the complexity of the original bounding box IoU calculation. The experiments on DOTA, DIOR, and HRSC datasets verify that the proposed loss function can effectively improve the accuracy of the model. Ronghua Shang, Zihan Ju, Jie Feng 0003, Songhua Xu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | I2U-Net: A dual-path U-Net with rich information interaction for medical image segmentation
Duwei Dai, Caixia Dong, Qingsen Yan, Yongheng Sun, Zongfang Li, Songhua Xu |
Medical Image Anal. | 7 |
| 2024 | Robust feature selection via central point link information and sparse latent representation
Jiarui Kong, Ronghua Shang, Chao Wang 0099, Songhua Xu |
Pattern Recognit. | 5 |
| 2024 | Graph embedding orthogonal decomposition: A synchronous feature selection technique based on collaborative particle swarm optimization
Jingyu Zhong, Ronghua Shang, Songhua Xu, Yangyang Li 0001 |
Pattern Recognit. | 3 |
| 2024 | SAR Image Segmentation Based on Complicated Region-Sensitive Adaptive Superpixel Generation and Hybrid Edge CorrectionabstractSuperpixel segmentation algorithms are predominently based on simple linear iterative clustering (SLIC), and treat homogeneous and complex regions equally. This can lead to suboptimal segmentation results, especially in complex images with multiple objects. We address this problem by proposing an SAR image segmentation algorithm based on complicated region-sensitive adaptive superpixel generation and hybrid edge correction (RSASGEC). First, a dynamic initialization algorithm for superpixel seeds based on region complexity is designed. Specifically, a new superpixel representation structure for superpixel seeds is constructed by combining superpixel complexity and the number of contained pixels. The algorithm gives priority to regions with high complexity, dynamically selecting the region with the highest complexity for further partitioning. This results in a dense distribution of superpixel seeds in complex regions, and sparse distributions in homogeneous regions with low complexity. Second, an iterative superpixel segmentation process based on an adaptive energy function is proposed. The Lagrange multiplier mathematical strategy is employed to optimize the adaptive energy function within an adjustable search window, resulting in more compact superpixel segmentation. Finally, a label correction method, based on edge mixture model constraints, is proposed for postprocessing. By integrating edge information from the Gaussian edge detector and the Canny algorithm as constraints, this method leverages majority voting and region growth methods to mitigate edge noise and outliers, refining the superpixel labels. The RSASGEC algorithm is verified in experiments, using one simulated image and six real SAR images. The results indicate that RSASGEC outperforms six representative algorithms, achieving more satisfactory segmentation performance. Jinhong Ren, Ronghua Shang, Jiansheng Chen 0004, Jie Feng 0003, Chao Wang 0099, Songhua Xu, Rustam Stolkin |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Joint Adversarial Network With Semantic and Topology Fusion for Cross-Scene Hyperspectral Image ClassificationabstractHyperspectral image cross-scene classification (HSICC) poses a significant challenge due to distribution variations between source and target domains. Existing unsupervised domain adaptation methods primarily focus on local knowledge transfer, often neglecting the critical semantic information and sample topological structure inherent in hyperspectral images (HSIs). To address these limitations, this article introduces an end-to-end joint adversarial network with semantic and topology fusion (JAN-STF). This network liberates from the constraints of local perception by integrating semantic and topological information into both domain- and class-level adversarial learning processes. First, the network constructs a semantic-guided cross-domain graph structure to obtain cross-domain features. Subsequently, domain-level adversarial learning is conducted using these features to achieve domain-invariant representation with robust transferability. Moreover, to bolster stability in the ensuing class-level adversarial procedure, the network dynamically computes cross-domain category center distance loss utilizing an intra-domain topological semantic attention mechanism, thereby mapping features to proximate spaces. Finally, class-level adversarial learning is performed by leveraging the prediction discrepancy between the local classifier and the topological classifier, thus enhancing the discriminative performance of the domain-invariant representation. Extensive experiments on three broadly utilized HSICC datasets demonstrate JAN-STF’s superiority in accuracy and Kappa coefficient (KC) metrics over nine leading algorithms. Ronghua Shang, Jie Feng 0003, Songhua Xu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Negative Label and Noise Information Guided Disambiguation for Partial Multi-Label LearningabstractPartial multi-label learning (PML) is defined as the construction of robust multi-label classification models from a training set where all instances are correlated with a corresponding group of candidate labels that are only partially accurate. Existing PML approaches have attempted to elicit reliable labels by parsing the guideline information of candidate labels. However, the other side information in the labels that describes what the sample does not contain is largely ignored. Moreover, existing PML approaches only focus on distinguishing the noise information and lack the effective use of noise information. To this end, a partial multi-label learning disambiguation approach guided by negative labels and noise information is proposed. Specifically, a negative label information-inducing paradigm is established based on the constructed negative label encoding matrix. Meanwhile, the negative correlation information guides an iterative label propagation process to induce ground-truth labels with high credibility. In addition, the truth and noise labels are formalized in a unified framework by constructing a regularizer. Moreover, the multi-label predictor is induced by discriminating regularization and disambiguation of label-specific features using the identified noise feature information. Extensive experiments on existing and constructed datasets have demonstrated that the negative label information bootstrapping strategy can be more effective in finding truth labels hidden in candidate labels. Moreover, noisy feature information-induced multi-label prediction outperforms state-of-the-art approaches. Jingyu Zhong, Ronghua Shang, Songhua Xu |
IEEE Trans. Multim. | 5 |
| 2023 | Twin Pseudo-training for semi-supervised semantic segmentation
Huiwen Huang, Songhua Xu, Youxing Li |
Comput. Graph. | 3 |
| 2023 | Effectively fusing clinical knowledge and AI knowledge for reliable lung nodule diagnosis
Duwei Dai, Yongheng Sun, Caixia Dong, Qingsen Yan, Zongfang Li, Songhua Xu |
Expert Syst. Appl. | 6 |
| 2023 | A Novel Three-Staged Generative Model for Skeletonizing Chinese Characters with Versatile Styles
Ye-Chuan Tian, Songhua Xu, Cheickna Sylla |
J. Comput. Sci. Technol. | 2 |
| 2023 | BookKD: A novel knowledge distillation for reducing distillation costs by decoupling knowledge generation and learning
Songling Zhu, Ronghua Shang, Songhua Xu, Yangyang Li 0001 |
Knowl. Based Syst. | 4 |
| 2023 | A novel multi-attention, multi-scale 3D deep network for coronary artery segmentation
Caixia Dong, Songhua Xu, Duwei Dai, Yizhi Zhang, Zongfang Li |
Medical Image Anal. | 2 |
| 2023 | MSCA-Net: Multi-scale contextual attention network for skin lesion segmentationabstractLesion segmentation algorithms automatically outline lesion areas in medical images, facilitating more effective identification and assessment of the clinically relevant features, and improving the efficacy and diagnosis accuracy. However, most fully convolutional network based segmentation methods suffer from spatial and contextual information loss when decreasing image resolution. To overcome this shortcoming, this paper proposes a skin lesion segmentation model , namely, the Multi-Scale Contextual Attention Network (MSCA-Net), which can exploit the multi-scale contextual information in images. Inspired by the skip connection of U-Net, we design a multi-scale bridge (MSB) module which interacts with multi-scale features to effectively fuse the multi-scale contextual information of the encoder and decoder path features. We further propose a global-local channel spatial attention module (GL-CSAM), aiming at capturing global contextual information. In addition, to take full advantage of the multi-scale features of the decoder, we propose a scale-aware deep supervision (SADS) module to achieve hierarchical iterative deep supervision. Comprehensive experimental results on the public dataset of ISIC 2017, ISIC 2018, and PH 2 show that our proposed method outperforms other state-of-the-art methods, demonstrating the efficacy of our method in skin lesion segmentation. Our code is available at https://github.com/YonghengSun1997/MSCA-Net . Yongheng Sun, Duwei Dai, Qianni Zhang, Yaqi Wang 0002, Songhua Xu, Chunfeng Lian |
Pattern Recognit. | 5 |
| 2023 | 3D Medical image segmentation using parallel transformers
Qingsen Yan, Shengqiang Liu, Songhua Xu, Caixia Dong, Zongfang Li, Qinfeng Shi, Yanning Zhang 0001, Duwei Dai |
Pattern Recognit. | 3 |
| 2022 | Ms RED: A novel multi-scale residual encoding and decoding network for skin lesion segmentation
Duwei Dai, Caixia Dong, Songhua Xu, Qingsen Yan, Zongfang Li, Nana Luo |
Medical Image Anal. | 3 |
| 2022 | Rethinking adversarial domain adaptation: Orthogonal decomposition for unsupervised domain adaptation in medical image segmentation
Yongheng Sun, Duwei Dai, Songhua Xu |
Medical Image Anal. | 3 |
| 2022 | A hierarchical model for learning to understand head gesture videos
Songhua Xu, Xueying Qin |
Pattern Recognit. | 2 |
| 2022 | Real-Time Shadow Detection From Live Outdoor Videos for Augmented RealityabstractSimulating shadow interactions between real and virtual objects is important for augmented reality (AR), in which accurately and efficiently detecting real shadows from live videos is a crucial step. Most of the existing methods are capable of processing only scenes captured under a fixed viewpoint. In contrast, this article proposes a new framework for shadow detection in live outdoor videos captured under moving viewpoints. The framework splits each frame into a tracked region, which is the region tracked from the previous video frame through optical flow analysis, and an emerging region, which is newly introduced into the scene due to the moving viewpoint. The framework subsequently extracts features based on the intensity profiles surrounding the boundaries of candidate shadow regions. These features are then utilized to both correct erroneous shadow boundaries for the tracked region and to detect shadow boundaries for the emerging region by a Bayesian learning module. To remove spurious shadows, spatial layout constraints are further considered for emerging regions. The experimental results demonstrate that the proposed framework outperforms the state-of-the-art shadow tracking and detection algorithms on a variety of challenging cases in real time, including shadows on backgrounds with complex textures, nonplanar shadows, fast-moving shadows with changing typologies, and shadows cast by nonrigid objects. The quantitative experiments show that our method outperforms the best existing method, achieving a 33.3% increase in the average$F_{measure}$on a self-collected database. Coupled with an image-based shadow-casting method, the proposed framework generates realistic shadow interaction results. This capability will be particularly beneficial for supporting AR applications. Yanli Liu 0002, Xingming Zou, Songhua Xu, Guanyu Xing, Housheng Wei, Yanci Zhang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Fusing Temporally Distributed Multi-Modal Semantic Clues for Video Question AnsweringabstractVideo Question Answering (VideoQA) is an intriguing topic, attracting increasing interest among the broad AI community. Yet videoQA is a difficult task. An algorithm competently tackle this task that needs to be able to: 1) extract rich semantics supplied in each modality of a video and incorporate them across modalities, and 2) identify and integrate such multimodal semantics from pertinent moments of a video, which may or may not be temporally adjacent or nearby, while filtering away irrelevant or even detractive portions of the video, to yield the most precise and sensible semantic context for executing the QA task. In response to the above requirements, a novel deep VideoQA solution is proposed in this paper, which comprises a multi-modal semantic clue extraction module, driven by a series of deep networks, each dedicated to digesting signals of a distinct modality type, to develop the first algorithmic QA capability, and a multi-modal temporal QA module empowered by a deep graph attention network to build the second algorithmic QA capability. Comprehensive experiments are conducted on publicly available benchmark data to validate advantages of the new solution in the end. Ruomei Wang 0001, Songhua Xu, Fan Zhou 0001 |
ICME | 3 |
| 2021 | Hierarchical Temporal Multi-Instance Learning for Video-based Student Learning Engagement AssessmentabstractVideo-based automatic assessment of a student's learning engagement on the fly can provide immense values for delivering personalized instructional services, a vehicle particularly important for massive online education. To train such an assessor, a major challenge lies in the collection of sufficient labels at the appropriate temporal granularity since a learner's engagement status may continuously change throughout a study session. Supplying labels at either frame or clip level incurs a high annotation cost. To overcome such a challenge, this paper proposes a novel hierarchical multiple instance learning (MIL) solution, which only requires labels anchored on full-length videos to learn to assess student engagement at an arbitrary temporal granularity and for an arbitrary duration in a study session. The hierarchical model mainly comprises a bottom module and a top module, respectively dedicated to learning the latent relationship between a clip and its constituent frames and that between a video and its constituent clips, with the constraints on the training stage that the average engagements of local clips is that of the video label. To verify the effectiveness of our method, we compare the performance of the proposed approach with that of several state-of-the-art peer solutions through extensive experiments. Jiayao Ma 0001, Xinbo Jiang, Songhua Xu, Xueying Qin |
IJCAI | 3 |
| 2021 | Shadow Detection via Predicting the Confidence Maps of Shadow Detection MethodsabstractToday's mainstream shadow detection methods are manually designed via a case-by-case approach. Accordingly, these methods may only be able to detect shadows for specific scenes. Given the complex and diverse shadow scenes in reality, none of the existing methods can provide a one-size-fits-all solution with satisfactory performance. To address this problem, this paper introduces a new concept, named shadow detection confidence, which can be used to evaluate the effect of any shadow detection method for any given scene. The best detection effect for a scene is achieved by combining prediction results by multiple methods. To measure the shadow detection confidence characteristics of an image, a novel relative confidence map prediction network (RCMPNet) is proposed. Experimental results show that the proposed method outperforms multiple state-of-the-art shadow detection methods on four shadow detection benchmark datasets. Jingwei Liao, Yanli Liu 0002, Guanyu Xing, Housheng Wei, Jueyu Chen, Songhua Xu |
ACM Multimedia | 6 |
| 2021 | 3D Object Tracking with Adaptively Weighted Local Bundles
Jia-Chen Li, Fan Zhong 0001, Songhua Xu, Xueying Qin |
J. Comput. Sci. Technol. | 3 |
| 2021 | Orthogonal Nonnegative Matrix Factorization using a novel deep Autoencoder Network
Songhua Xu |
Knowl. Based Syst. | 2 |
| 2021 | A novel deep quantile matrix completion model for top-N recommendation
Songhua Xu |
Knowl. Based Syst. | 2 |
| 2021 | A Triplet network framework based automatic assessment of simulation quality for respiratory droplet propagation
Songhua Xu, Xiangdong Ding |
Pattern Recognit. | 2 |
| 2020 | PanelNet: A Novel Deep Neural Network for Predicting Collective Diagnostic Ratings by a Panel of Radiologists for Pulmonary NodulesabstractReducing misdiagnosis rate is a central concern in modern medicine. In clinical practice, group-based collective diagnosis is frequently exercised to curb the misdiagnosis rate. However, little effort has been dedicated to emulating the collective intelligence behind the group-based decision making practice in computer-aided diagnosis research to this day. To fill the overlooked gap, this study introduces a novel deep neural network, titled PanelNet, that is able to computationally model and reproduce the aforesaid collective diagnosis capability demonstrated by a group of medical experts. To experimentally explore the validity of the new solution, we apply the proposed PanelNet to one of the key tasks in radiology---assessing malignant ratings of pulmonary nodules. For each nodule and a given panel, PanelNet is able to predict statistical distribution of malignant ratings collectively judged by the panel of radiologists. Extensive experimental results consistently demonstrate PanelNet outperforms multiple state-of-the-art computer-aided diagnosis methods applicable to the collective diagnostic task. To our best knowledge, no other collective computer-aided diagnosis method grounded on modern machine learning technologies has been previously proposed. By its design, PanelNet can also be easily applied to model collective diagnosis processes employed for other diseases. Songhua Xu, Zongfang Li |
ACM Multimedia | 2 |
| 2020 | A novel patch-based nonlinear matrix completion algorithm for image analysis through convolutional neural network
Songhua Xu |
Neurocomputing | 2 |
| 2020 | A Deep Learning Model for Transportation Mode Detection Based on Smartphone Sensing DataabstractUnderstanding people's transportation modes is beneficial for empowering many intelligent transportation systems, such as supporting urban transportation planning. Yet, current methodologies in collecting travelers' transportation modes are costly and inaccurate. Fortunately, the increasing sensing and computing capabilities of smartphones and their high penetration rate offer a promising approach to automatic transportation mode detection via mobile computation. This paper introduces a light-weighted and energy-efficient transportation mode detection system using only accelerometer sensors in smartphones. The system collects accelerometer data in an efficient way and leverages a deep learning model to determine transportation modes. Different architectures and classification methods are tested with the proposed deep learning model to optimize the system design. Performance evaluation shows that the proposed new approach achieves a better accuracy than existing work in detecting people's transportation modes. Xiaoyuan Liang, Yuchuan Zhang, Grace Guiling Wang, Songhua Xu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | A New Visual Interface for Searching and Navigating Slide-Based Lecture VideosabstractThe rapid development of distance education technologies, e.g. MOOCs, provide learners unprecedented access to high-quality online lecture videos at scale, anytime and anywhere. Unfortunately, these valuable resources are often underutilized by online learners. One prevailing reason is the lack of support for and the resulting difficulty of exploring and locating content of interest among lengthy recordings of course lectures. To address this deficiency, we introduce a novel visual interface that supports efficient search and navigation of video content at fine granularities and with rich semantic clues. The interface is particularly designed for slide-based lecture videos (SBLV), which represent a significant portion of online lecture videos. The interface comprehensively derives versatile semantic clues for video content indexing and visual aid generation according to visual elements, text, and mathematical expressions included on lecture slides, speeches recorded, as well as mouse and cursor pointing actions captured during a lecture. Empowered by such semantically revealing indices and visual assistance, the interface is able to noticeably enhance online learners' capabilities in searching and browsing of content in need from SBLVs. The advantages of the new interface are demonstrated through benchmarked experimental results in comparison with peer methods. Baoquan Zhao, Songhua Xu, Shujin Lin, Ruomei Wang 0001 |
ICME | 2 |
| 2018 | Sparsely Grouped Multi-Task Generative Adversarial Networks for Facial Attribute ManipulationabstractRecently, Image-to-Image Translation (IIT) has achieved great progress in image style transfer and semantic context manipulation for images. However, existing approaches require exhaustively labelling training data, which is labor demanding, difficult to scale up, and hard to adapt to a new domain. To overcome such a key limitation, we propose Sparsely Grouped Generative Adversarial Networks (SG-GAN) as a novel approach that can translate images in sparsely grouped datasets where only a few train samples are labelled. Using a one-input multi-output architecture, SG-GAN is well-suited for tackling multi-task learning and sparsely grouped learning tasks. The new model is able to translate images among multiple groups using only a single trained model. To experimentally validate the advantages of the new model, we apply the proposed method to tackle a series of attribute manipulation tasks for facial images as a case study. Experimental results show that SG-GAN can achieve comparable results with state-of-the-art methods on adequately labelled datasets while attaining a superior image translation quality on sparsely grouped datasets~\footnoteCode is available at https://github.com/zhangqianhui/SGGAN-tensorflow.. Jichao Zhang, Yezhi Shu, Songhua Xu, Gongze Cao, Fan Zhong 0001, Meng Liu 0006, Xueying Qin |
ACM Multimedia | 3 |
| 2018 | An advanced operating environment for mathematics education resources
Yongsheng Rao, Jingzhong Zhang, Yanchun Sun, Xiangping Chen, Songhua Xu |
Sci. China Inf. Sci. | 6 |
| 2018 | A new method for retrieving batik shape patternsabstractBatik as a traditional art is well regarded due to its high aesthetic quality and cultural heritage values. It is not uncommon to reuse versatile decorative shape patterns across batiks. General‐purpose image retrieval methods often fail to pay sufficient attention to such a frequent reuse of shape patterns in the graphical compositions of batiks, leading to suboptimal retrieval results, in particular for identifying batiks that use copyrighted shape patterns without proper authorization for law‐enforcement purposes. To address the lack of an optimized image retrieval method suited for batiks, this study proposes a new method for retrieving salient shape patterns in batiks using a rich combination of global and local features. The global features deployed were extracted according to the Zernike moments (ZMs); the local features adopted were extracted through curvelet transformations that characterize shape contours embedded in batiks. The method subsequently incorporated both types of features via matching a weighted bipartite graph to measure the visual similarity between any pair of batik shape patterns through supervised distance metric learning. The derived similarity metric can then be used to detect and retrieve similar shape patterns appearing across batiks, which in turn can be employed as a reliable similarity metric for retrieving batiks. To explore the usefulness of the proposed method, the performance of the new retrieval method is compared against that of three peer methods as well as two variants of the proposed method. The experimental results consistently and convincingly demonstrate that the new method indeed outperforms the state‐of‐the‐art methods in retrieving salient shape patterns in batiks. Qingni Yuan, Songhua Xu, Lv Jian |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2018 | Learning to sketch human facial portraits using personal styles by case-based reasoning
Bingwen Jin, Songhua Xu, Weidong Geng |
Multim. Tools Appl. | 2 |
| 2017 | A Novel System for Visual Navigation of Educational Videos Using Multimodal CuesabstractWith recent developments and advances in distance learning and MOOCs, the amount of open educational videos on the Internet has grown dramatically in the past decade. However, most of these videos are lengthy and lack of high-quality indexing and annotations, which triggers an urgent demand for efficient and effective tools that facilitate video content navigation and exploration. In this paper, we propose a novel visual navigation system for exploring open educational videos. The system tightly integrates multimodal cues obtained from the visual, audio and textual channels of the video and presents them with a series of interactive visualization components. With the help of this system, users can explore the video content using multiple levels of details to identify content of interest with ease. Extensive experiments and comparisons against previous studies demonstrate the effectiveness of the proposed system. Baoquan Zhao, Shujin Lin, Songhua Xu, Ruomei Wang 0001 |
ACM Multimedia | 4 |
| 2017 | A local context-aware LDA model for topic modeling in a document networkabstractWith the rapid development of the Internet and its applications, growing volumes of documents increasingly become interconnected to form large‐scale document networks. Accordingly, topic modeling in a network of documents has been attracting continuous research attention. Most of the existing network‐based topic models assume that topics in a document are influenced by its directly linked neighbouring documents in a document network and overlook the potential influence from indirectly linked ones. The existing work also has not carefully modeled variations of such influence among neighboring documents. Recognizing these modeling limitations, this paper introduces a novel Local Context‐Aware LDA Model (LC‐LDA), which is capable of observing a local context comprising a rich collection of documents that may directly or indirectly influence the topic distributions of a target document. The proposed model can also differentiate the respective influence of each document in the local context on the target document according to both structural and temporal relationships between the two documents. The proposed model is extensively evaluated through multiple document clustering and classification tasks conducted over several large‐scale document sets. Evaluation results clearly and consistently demonstrate the effectiveness and superiority of the new model with respect to several state‐of‐the‐art peer models. Yang Liu 0046, Songhua Xu |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Intelligent bus routing with heterogeneous human mobility patterns
Yanchi Liu, Chuanren Liu, Nicholas Jing Yuan, Yanjie Fu, Hui Xiong 0001, Songhua Xu, Junjie Wu 0002 |
Knowl. Inf. Syst. | 7 |
| 2016 | The utility of web mining for epidemiological research: studying the association between parity and cancer riskabstractBACKGROUND: The World Wide Web has emerged as a powerful data source for epidemiological studies related to infectious disease surveillance. However, its potential for cancer-related epidemiological discoveries is largely unexplored. METHODS: Using advanced web crawling and tailored information extraction procedures, the authors automatically collected and analyzed the text content of 79 394 online obituary articles published between 1998 and 2014. The collected data included 51 911 cancer (27 330 breast; 9470 lung; 6496 pancreatic; 6342 ovarian; 2273 colon) and 27 483 non-cancer cases. With the derived information, the authors replicated a case-control study design to investigate the association between parity (i.e., childbearing) and cancer risk. Age-adjusted odds ratios (ORs) with 95% confidence intervals (CIs) were calculated for each cancer type and compared to those reported in large-scale epidemiological studies. RESULTS: Parity was found to be associated with a significantly reduced risk of breast cancer (OR = 0.78, 95% CI, 0.75-0.82), pancreatic cancer (OR = 0.78, 95% CI, 0.72-0.83), colon cancer (OR = 0.67, 95% CI, 0.60-0.74), and ovarian cancer (OR = 0.58, 95% CI, 0.54-0.62). Marginal association was found for lung cancer risk (OR = 0.87, 95% CI, 0.81-0.92). The linear trend between increased parity and reduced cancer risk was dramatically more pronounced for breast and ovarian cancer than the other cancers included in the analysis. CONCLUSION: This large web-mining study on parity and cancer risk produced findings very similar to those reported with traditional observational studies. It may be used as a promising strategy to generate study hypotheses for guiding and prioritizing future epidemiological studies. Georgia D. Tourassi, Hong-Jun Yoon, Songhua Xu, Xuesong Han |
J. Am. Medical Informatics Assoc. | 3 |
| 2016 | A new visual navigation system for exploring biomedical Open Educational Resource (OER) videosabstractOBJECTIVE: Biomedical videos as open educational resources (OERs) are increasingly proliferating on the Internet. Unfortunately, seeking personally valuable content from among the vast corpus of quality yet diverse OER videos is nontrivial due to limitations of today's keyword- and content-based video retrieval techniques. To address this need, this study introduces a novel visual navigation system that facilitates users' information seeking from biomedical OER videos in mass quantity by interactively offering visual and textual navigational clues that are both semantically revealing and user-friendly. MATERIALS AND METHODS: The authors collected and processed around 25 000 YouTube videos, which collectively last for a total length of about 4000 h, in the broad field of biomedical sciences for our experiment. For each video, its semantic clues are first extracted automatically through computationally analyzing audio and visual signals, as well as text either accompanying or embedded in the video. These extracted clues are subsequently stored in a metadata database and indexed by a high-performance text search engine. During the online retrieval stage, the system renders video search results as dynamic web pages using a JavaScript library that allows users to interactively and intuitively explore video content both efficiently and effectively.ResultsThe authors produced a prototype implementation of the proposed system, which is publicly accessible athttps://patentq.njit.edu/oer To examine the overall advantage of the proposed system for exploring biomedical OER videos, the authors further conducted a user study of a modest scale. The study results encouragingly demonstrate the functional effectiveness and user-friendliness of the new system for facilitating information seeking from and content exploration among massive biomedical OER videos. CONCLUSION: Using the proposed tool, users can efficiently and effectively find videos of interest, precisely locate video segments delivering personally valuable information, as well as intuitively and conveniently preview essential content of a single or a collection of videos. Baoquan Zhao, Songhua Xu, Shujin Lin |
J. Am. Medical Informatics Assoc. | 2 |
| 2016 | A novel web informatics approach for automated surveillance of cancer mortality trends
Georgia D. Tourassi, Hong-Jun Yoon, Songhua Xu |
J. Biomed. Informatics | 3 |
| 2016 | Detecting Rumors Through Modeling Information Propagation Networks in a Social Media EnvironmentabstractIn the midst of today's pervasive influence of social media content and activities, information credibility has increasingly become a major issue. Accordingly, identifying false information, e.g., rumors circulated in social media environments, attracts expanding research attention and growing interests. Many previous studies have exploited user-independent features for rumor detection. These prior investigations uniformly treat all users relevant to the propagation of a social media message as instances of a generic entity. Such a modeling approach usually adopts a homogeneous network to represent all users, the practice of which ignores the variety across an entire user population in a social media environment. Recognizing this limitation in modeling methodologies, this paper explores user-specific features in a social media environment for rumor detection. The new approach hypothesizes whether a user tending to spread a rumor message is dependent on specific attributes of the user in addition to content characteristics of the message itself. Under this hypothesis, the information propagation patterns of rumors versus those of credible messages in a social media environment are differentiable. To explore and exploit this hypothesis, we develop a new information propagation model based on a heterogeneous user representation and modeling approach. By applying the new approach, we are able to differentiate rumors from credible messages through observing distinctions in their respective propagation patterns in social media. The experimental results show that the new information propagation model based on heterogeneous user representation can effectively distinguish rumors from credible social media content. Our experimental findings further show that rumors are more likely to spread among certain user groups. Yang Liu 0046, Songhua Xu |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2014 | Extracting Patient Demographics and Personal Medical Information from Online Health Forums
Yang Liu 0046, Songhua Xu, Hong-Jun Yoon, Georgia D. Tourassi |
AMIA | 2 |
| 2014 | A New Visual Navigation System for Exploring Biomedical Patents
Christopher Markson, Songhua Xu |
AMIA | 2 |
| 2014 | Relationship Emergence Prediction in Heterogeneous Networks through Dynamic Frequent Subgraph MiningabstractWith the rapid development of Web 2.0 and the Internet of things, predicting relationships in heterogeneous networks has evolved as a heated research topic. Traditionally, people analyze existing relationships in heterogeneous networks that relate in a particular way to a target relationship of interest to predict the emergence of the target relationship. However most existing methods are incapable of systematically identifying relevant relationships useful for the prediction task, especially those relationships involving multiple objects of heterogeneous types, which may not rest on a simple path in the concerned heterogeneous network. Another problem with the current practice is that the existing solutions often ignore the dynamic evolution of the network structure after the introduction of newly emerged relationships. To overcome the first limitation, we propose a new algorithm that can systematically and comprehensively detect relevant relationships useful for the prediction of an arbitrarily given target relationship through a disciplined graph searching process. To address the second limitation, the new algorithm leverages a series of temporally-sensitive features for the relationship occurrence prediction via a supervised learning approach. To explore the effectiveness of the new algorithm, we apply the prototype implementation of the algorithm on the DBLP bibliographic network to predict the author citation relationships and compare the algorithm performance with that of a state-of-the-art peer method and a series of baseline methods. The comparison shows consistently higher prediction accuracy under a range of prediction scenarios. Yang Liu 0046, Songhua Xu |
CIKM | 2 |
| 2014 | Exploiting Heterogeneous Human Mobility Patterns for Intelligent Bus RoutingabstractOptimal planning for public transportation is one of the keys to sustainable development and better quality of life in urban areas. Compared to private transportation, public transportation uses road space more efficiently and produces fewer accidents and emissions. In this paper, we focus on the identification and optimization of flawed bus routes to improve utilization efficiency of public transportation services, according to people's real demand for public transportation. To this end, we first provide an integrated mobility pattern analysis between the location traces of taxicabs and the mobility records in bus transactions. Based on mobility patterns, we propose a localized transportation mode choice model, with which we can accurately predict the bus travel demand for different bus routing. This model is then used for bus routing optimization which aims to convert as many people from private transportation to public transportation as possible given budget constraints on the bus route modification. We also leverage the model to identify region pairs with flawed bus routes, which are effectively optimized using our approach. To validate the effectiveness of the proposed methods, extensive studies are performed on real world data collected in Beijing which contains 19 million taxi trips and 10 million bus trips. Yanchi Liu, Chuanren Liu, Nicholas Jing Yuan, Yanjie Fu, Hui Xiong 0001, Songhua Xu, Junjie Wu 0002 |
ICDM | 7 |
| 2014 | A user-oriented web crawler for selectively acquiring online content in e-health researchabstractMOTIVATION: Life stories of diseased and healthy individuals are abundantly available on the Internet. Collecting and mining such online content can offer many valuable insights into patients' physical and emotional states throughout the pre-diagnosis, diagnosis, treatment and post-treatment stages of the disease compared with those of healthy subjects. However, such content is widely dispersed across the web. Using traditional query-based search engines to manually collect relevant materials is rather labor intensive and often incomplete due to resource constraints in terms of human query composition and result parsing efforts. The alternative option, blindly crawling the whole web, has proven inefficient and unaffordable for e-health researchers. RESULTS: We propose a user-oriented web crawler that adaptively acquires user-desired content on the Internet to meet the specific online data source acquisition needs of e-health researchers. Experimental results on two cancer-related case studies show that the new crawler can substantially accelerate the acquisition of highly relevant online content compared with the existing state-of-the-art adaptive web crawling technology. For the breast cancer case study using the full training set, the new method achieves a cumulative precision between 74.7 and 79.4% after 5 h of execution till the end of the 20-h long crawling session as compared with the cumulative precision between 32.8 and 37.0% using the peer method for the same time period. For the lung cancer case study using the full training set, the new method achieves a cumulative precision between 56.7 and 61.2% after 5 h of execution till the end of the 20-h long crawling session as compared with the cumulative precision between 29.3 and 32.4% using the peer method. Using the reduced training set in the breast cancer case study, the cumulative precision of our method is between 44.6 and 54.9%, whereas the cumulative precision of the peer method is between 24.3 and 26.3%; for the lung cancer case study using the reduced training set, the cumulative precisions of our method and the peer method are, respectively, between 35.7 and 46.7% versus between 24.1 and 29.6%. These numbers clearly show a consistently superior accuracy of our method in discovering and acquiring user-desired online content for e-health research. AVAILABILITY AND IMPLEMENTATION: The implementation of our user-oriented web crawler is freely available to non-commercial users via the following Web site: http://bsec.ornl.gov/AdaptiveCrawler.shtml. The Web site provides a step-by-step guide on how to execute the web crawler implementation. In addition, the Web site provides the two study datasets including manually labeled ground truth, initial seeds and the crawling results reported in this article. Songhua Xu, Hong-Jun Yoon, Georgia D. Tourassi |
Bioinform. | 1 |
| 2014 | A new algorithm for product image search based on salient edge characterizationabstractVisually assisted product image search has gained increasing popularity because of its capability to greatly improve end users' e‐commerce shopping experiences. Different from general‐purpose content‐based image retrieval (CBIR) applications, the specific goal of product image search is to retrieve and rank relevant products from a large‐scale product database to visually assist a user's online shopping experience. In this paper, we explore the problem of product image search through salient edge characterization and analysis, for which we propose a novel image search method coupled with an interactive user region‐of‐interest indication function. Given a product image, the proposed approach first extracts an edge map, based on which contour curves are further extracted. We then segment the extracted contours into fragments according to the detected contour corners. After that, a set of salient edge elements is extracted from each product image. Based on salient edge elements matching and similarity evaluation, the method derives a new pairwise image similarity estimate. Using the new image similarity, we can then retrieve product images. To evaluate the performance of our algorithm, we conducted 120 sessions of querying experiments on a data set comprised of around 13k product images collected from multiple, real‐world e‐commerce websites. We compared the performance of the proposed method with that of a bag‐of‐words method (Philbin, Chum, Isard, Sivic, & Zisserman, 2008) and a Pyramid Histogram of Orientated Gradients (PHOG) method (Bosch, Zisserman, & Munoz, 2007). Experimental results demonstrate that the proposed method improves the performance of example‐based product image retrieval. Yuhua Li 0002, Songhua Xu, Shujin Lin |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2014 | Selecting the Right Correlation Measure for Binary DataabstractFinding the most interesting correlations among items is essential for problems in many commercial, medical, and scientific domains. Although there are numerous measures available for evaluating correlations, different correlation measures provide drastically different results. Piatetsky-Shapiro provided three mandatory properties for any reasonable correlation measure, and Tan et al. proposed several properties to categorize correlation measures; however, it is still hard for users to choose the desirable correlation measures according to their needs. In order to solve this problem, we explore the effectiveness problem in three ways. First, we propose two desirable properties and two optional properties for correlation measure selection and study the property satisfaction for different correlation measures. Second, we study different techniques to adjust correlation measures and propose two new correlation measures: the Simplified χ2with Continuity Correction and the Simplified χ2with Support. Third, we analyze the upper and lower bounds of different measures and categorize them by the bound differences. Combining these three directions, we provide guidelines for users to choose the proper measure according to their needs. W. Nick Street, Yanchi Liu, Songhua Xu, Yi-fang Brook Wu |
ACM Trans. Knowl. Discov. Data | 4 |
| 2013 | A new algorithm for context-based biomedical diagram similarity estimationabstractMOTIVATION: Diagrams embedded in the biomedical literature convey rich contents, which often concisely and intuitively highlight key thesis of a research article. Despite their vital importance and informative clues for biomedical literature navigation and retrieval; currently, we miss an effective computational method for automatically understanding and accessing these valuable resources. PROPOSED METHOD: To address the aforementioned gap, we propose a novel context-based algorithm for estimating the similarity between a pair of biomedical diagrams. The main difference of the proposed algorithm with respect to the existing methods lies in the new algorithm's incorporation of the semantic context associated with diagrams in their source documents into the diagram similarity estimation process. In addition, the new approach also performs a series of advanced image processing and text mining operations to comprehensively extract the semantic content graphically encoded inside diagram images. RESULTS: The new algorithm can be deployed as a reusable component providing a fundamental function for building many advanced, semantic-aware applications on biomedical diagram processing. As a case study, in our experiments, we demonstrate the advantage of the new algorithm for diagram retrieval. A set of biomedical diagram search and ranking experiments were conducted, where the performance of the new method was compared with that of five peer methods. The comparison results demonstrate the performance superiority of the new algorithm with all peer methods with statistical significance. Songhua Xu, Jianqiang Sheng |
Bioinform. | 1 |
| 2013 | A new interpolation subdivision scheme for triangle/quad mesh
Shujin Lin, Songhua Xu, Jianmin Wang 0013 |
Graph. Model. | 3 |
| 2013 | Cylindrical panoramic mosaicing from a pipeline video through MRF based optimization
Chuan Niu, Fan Zhong 0001, Songhua Xu, Chenglei Yang, Xueying Qin |
Vis. Comput. | 3 |
| 2012 | Novel image features for categorizing biomedical imagesabstractImages embedded in biomedical publications are richly informative. For example, they often concisely summarize key hypotheses, illustrate new methods, and highlight major experimental findings in a research article. Prior studies [1] suggested that images embedded in biomedical publications offer effective clues for retrieving and mining their source documents. To facilitate accessing such valuable imagery resources, image categorization can be helpful. Like many other image processing tasks, extracting discriminative image features is critical for the success of image categorization. For biomedical images, we notice that many of them are embedded with abundant annotation text. Observing this property, we introduce a set of novel image features that exploit the spatial distribution of text information inside an image as essential clues for categorizing biomedical images. Through results of our evaluation experiments, this paper demonstrates the effectiveness of the proposed novel features - compared with conventional image features, our new features can help categorize biomedical images with superior performance using a standard supervised learning based approach. Jianqiang Sheng, Songhua Xu, Weicai Deng |
BIBM | 2 |
| 2012 | A semantic relatedness approach for traceability link recoveryabstractHuman analysts working with automated tracing tools need to directly vet candidate traceability links in order to determine the true traceability information. Currently, human intervention happens at the end of the traceability process, after candidate traceability links have already been generated. This often leads to a decline in the results' accuracy. In this paper, we propose an approach, based on semantic relatedness (SR), which brings human judgment to an earlier stage of the tracing process by integrating it into the underlying retrieval mechanism. SR tries to mimic human mental model of relevance by considering a broad range of semantic relations, hence producing more semantically meaningful results. We evaluated our approach using three datasets from different application domains, and assessed the tracing results via six different performance measures concerning both result quality and browsability. The empirical evaluation results show that our SR approach achieves a significantly better performance in recovering true links than a standard Vector Space Model (VSM) in all datasets. Our approach also achieves a significantly better precision than Latent Semantic Indexing (LSI) in two of our datasets. Anas Mahmoud 0001, Nan Niu, Songhua Xu |
ICPC | 3 |
| 2012 | Example-Based Automatic Music-Driven Conventional Dance Motion SynthesisabstractWe introduce a novel method for synthesizing dance motions that follow the emotions and contents of a piece of music. Our method employs a learning-based approach to model the music to motion mapping relationship embodied in example dance motions along with those motions' accompanying background music. A key step in our method is to train a music to motion matching quality rating function through learning the music to motion mapping relationship exhibited in synchronized music and dance motion data, which were captured from professional human dance performance. To generate an optimal sequence of dance motion segments to match with a piece of music, we introduce a constraint-based dynamic programming procedure. This procedure considers both music to motion matching quality and visual smoothness of a resultant dance motion sequence. We also introduce a two-way evaluation strategy, coupled with a GPU-based implementation, through which we can execute the dynamic programming process in parallel, resulting in significant speedup. To evaluate the effectiveness of our method, we quantitatively compare the dance motions synthesized by our method with motion synthesis results by several peer methods using the motions captured from professional human dancers' performance as the gold standard. We also conducted several medium-scale user studies to explore how perceptually our dance motion synthesis method can outperform existing methods in synthesizing dance motions to match with a piece of music. These user studies produced very positive results on our music-driven dance motion synthesis experiments for several Asian dance genres, confirming the advantages of our method. Rukun Fan, Songhua Xu, Weidong Geng |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2011 | Creating Cylindrical Panoramic Mosaic from a Pipeline VideoabstractIn geological engineering, stratum structure detection is a fundamental problem in project planning and implementation. One of the most commonly employed detection technologies is to take videos of borehole using a forward moving camera. Following this approach, the problem of stratum structure detection is transformed into the problem of constructing a panoramic image from the taken video sequences, which are typically in low quality. In this paper, we propose a novel method to create a panoramic image of the borehole from the video sequence without camera calibration and tracking. To stitch together pixels of neighboring frame images, our camera model is designed with a focal length changing feature, along with a small rotation freedom in the two-dimensional image space. Essentially, our camera model assumes target objects lie on a cylindrical wall and the camera moves forward along the central axis of the cylindrical wall. Our method robustly resolves these two degrees-of-freedoms through KLT feature tracking and constructs a panoramic image by stitching strips. Experiment results show that our method could efficiently generate high-quality panoramas for very long video sequences. Chuan Niu, Fan Zhong 0001, Songhua Xu, Chenglei Yang, Xueying Qin |
CAD/Graphics | 3 |
| 2011 | Retrieving and ranking unannotated images through collaboratively mining online search resultsabstractWe present a new image search and ranking algorithm for retrieving unannotated images by collaboratively mining online search results which consist of online image and text search results. The online image search results are leveraged as reference examples to perform content-based image search over unannotated images. The online text search results are utilized to estimate the reference images' relevance to the search query. The key feature of our method is its capability to deal with unreliable online image search results through jointly mining visual and textual aspects of online search results. Through such collaborative mining, our algorithm infers the relevance of an online search result image to a text query. Once we obtain the estimate of query relevance score for each online image search result, we can selectively use query specific online search result images as reference examples for retrieving and ranking unannotated images. We tested our algorithm both on the standard public image datasets and several modestly sized personal photo collections. We also compared our method with two well-known peer methods. The results indicate that our algorithm is superior to existing content-based image search algorithms for retrieving and ranking unannotated images. Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001 |
CIKM | 1 |
| 2011 | Mining User Dwell Time for Personalized Web Search Re-RankingabstractWe propose a personalized re-ranking algorithm through mining user dwell times derived from a user’s previously online reading or browsing activities. We acquire document level user dwell times via a customized web browser, from which we then infer concept word level user dwell times in order to understand a user’s personal interest. According to the estimated concept word level user dwell times, our algorithm can estimate a user’s potential dwell time over a new document, based on which personalized webpage re-ranking can be carried out. We compare the rankings produced by our algorithm with rankings generated by popular commercial search engines and a recently proposed personalized Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001 |
IJCAI | 1 |
| 2011 | Capturing user reading behaviors for personalized document summarizationabstractWe propose a new personalized document summarization method that observes a user's personal reading preferences. These preferences are inferred from the user's reading behaviors, including facial expressions, gaze positions, and reading durations that were captured during the user's past reading activities. We compare the performance of our algorithm with that of a few peer algorithms and software packages. The results of our comparative study show that our algorithm can produce more superior personalized document summaries than all the other methods in that the summaries generated by our algorithm can better satisfy a user's personal preferences. Hao Jiang 0011, Songhua Xu, Francis C. M. Lau 0001 |
IUI | 2 |
| 2010 | Observing facial expressions and gaze positions for personalized webpage recommendationabstractWe propose a new method for personalized webpage recommendation. The method is capable of inferring a user's personal reading interest distribution according to implicit user feedbacks coming from the user's past online reading activities. With the inferred user reading interest distribution, we can recommend webpages in a search result set to a user in a personalized way. Our method is featured by its novel approach to observe the facial expressions and gaze positions of a user during the user's online reading activities as two types of implicit user feedbacks for estimating the user's reading interest distribution. To capture these implicit user feedbacks, we use an ordinary web camera and a customized web browser in the setup. The setup allows us to measure the distribution of the reading time a user spends in his or her reading activities over materials of different contents. With all the captured information, our method then estimates a user's reading interest distribution by finding correlations between the implicit feedbacks of a user with the contents of the read materials. Given the estimated user reading interest distribution, our algorithm can further predict the user's potential reading interest in any new webpage. Consequently, our algorithm can produce a personalized webpage recommendation for all the result webpages in an online search session. We compared the performance of our method with that of several mainstream commercial search engines as well as a recent personalized webpage ranking algorithm. The comparison results clearly show the superiority of our new method for personalized webpage recommendation. Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001 |
ICEC | 1 |
| 2010 | Keyword Extraction and Headline Generation Using Novel Word FeaturesabstractWe introduce several novel word features for keyword extraction and headline generation. These new word features are derived according to the background knowledge of a document as supplied by Wikipedia. Given a document, to acquire its background knowledge from Wikipedia, we first generate a query for searching the Wikipedia corpus based on the key facts present in the document. We then use the query to find articles in the Wikipedia corpus that are closely related to the contents of the document. With the Wikipedia search result article set, we extract the inlink, outlink, category and infobox information in each article to derive a set of novel word features which reflect the document's background knowledge. These newly introduced word features offer valuable indications on individual words' importance in the input document. They serve as nice complements to the traditional word features derivable from explicit information of a document. In addition, we also introduce a word-document fitness feature to charcterize the influence of a document's genre on the keyword extraction and headline generation process. We study the effectiveness of these novel word features for keyword extraction and headline generation by experiments and have obtained very encouraging results. Songhua Xu, Shaohui Yang, Francis C. M. Lau 0001 |
AAAI | 1 |
| 2010 | A Force Field Driven SOM for boundary detectionabstractWe will introduce a method to extract object boundaries from an image. This method utilizes a deformable curve based on the Self Organizing Map algorithm. The proposed SOM has some unique properties such as batch update and neuron insertion/deletion. These properties can make the SOM converge to object concavities as well as maintain a uniform distribution of neurons along the SOM. In comparison with other traditional active contour methods, this algorithm is less sensitive to initialization more flexible in noisy conditions. It outperforms the Gradient Vector Flow method. Songhua Xu, Willard L. Miranker |
IJCNN | 2 |
| 2010 | Kernel based principal component for recognizing handwritten numbersabstractWe implement kernel-based principal component analysis to recognize handwritten numbers. Then we analyze the relationship of each numeral type's eigenvalue and eigenvector with its recognition rate. We also present a modified algorithm to improve the robustness as well as the efficiency of the recognition method by employing the secondary training and detection methods from the perspective of nature of kernel function. This method can solve the problem of low recognition rate of a small number of scribbled characters at both low time cost and space complexity. Experiments using 1000 to 5000 test samples all show that our method can achieve 97.8% to 99.0% recognition accuracy. Songhua Xu, Willard L. Miranker |
IJCNN | 2 |
| 2010 | A new pivoting and iterative text detection algorithm for biomedical images
Songhua Xu, Michael Krauthammer |
J. Biomed. Informatics | 1 |
| 2009 | Personalized web content provider recommendation through mining individual users' QoSabstractWe propose an optimal web content provider recommendation algorithm based on mining QoS (quality of service) information of the Internet. The QoS refers principally to the network bandwidth and waiting time (for a connection to be established). For contents replicated over multiple sites, our algorithm recommends a list of webpages having the desired content and ranked according to their QoSs for any specific user. The recommendation is generated through a data mining procedure based on known QoSs of connections between pairs of computers. Our user QoS mining procedure incrementally constructs a neural network group for QoS prediction based on clustering over the prediction errors. An accompanying decision tree algorithm is then used to select the most appropriate neural network among the neural network group to predict the QoS for a particular user connection. Based on our proposed recommendation algorithm, we have implemented a user-oriented search engine which can identify similar web content providers and make a ranked recommendation based on the prediction over the QoS experienced by individual users. Experiment results have verified that our QoS-based personal web content provider ranking algorithm can indeed produce a recommendation that improves the QoS experienced by individual users. Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001 |
ICEC | 1 |
| 2009 | Automatic Generation of Personal Chinese Handwriting by Capturing the Characteristics of Personal Handwriting
Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001 |
IAAI | 1 |
| 2009 | User-oriented document summarization through vision-based eye-trackingabstractWe propose a new document summarization algorithm which is personalized. The key idea is to rely on the attention (reading) time of individual users spent on single words in a document as the essential clue. The prediction of user attention over every word in a document is based on the user's attention during his previous reads, which is acquired via a vision-based commodity eye-tracking mechanism. Once the user's attentions over a small collection of words are known, our algorithm can predict the user's attention over every word in the document through word semantics analysis. Our algorithm then summarizes the document according to user attention on every individual word in the document. With our algorithm, we have developed a document summarization prototype system. Experiment results produced by our algorithm are compared with the ones manually summarized by users as well as by commercial summarization software, which clearly demonstrates the advantages of our new algorithm for user-oriented document summarization. Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001 |
IUI | 1 |
| 2009 | A new visual search interface for web browsingabstractWe introduce a new visual search interface for search engines. The interface is a user-friendly and informative graphical front-end for organizing and presenting search results in the form of topic groups. Such a semantics-oriented search result presentation is in contrast with conventional search interfaces which present search results according to the physical structures of the information. Given a user query, our interface first retrieves relevant online materials via a third-party search engine. And then we analyze the semantics of search results to detect latent topics in the result set. Once the topics are detected, we map the search result pages into topic clusters. According to the topic clustering result, we divide the available screen space for our visual interface into multiple topic displaying regions, one for each topic. For each topic's displaying region, we summarize the information contained in the search results under the corresponding topic so that only key messages will be displayed. With this new visual search interface, users are conveyed the key information in the search results expediently. With the key information, users can navigate to the final, desired results with less effort and time than conventional searching. Supplementary materials for this paper are available at http://www.cs.hku.hk/~songhua/visualsearch/. Songhua Xu, Francis C. M. Lau 0001 |
WSDM | 1 |
| 2009 | A Visualization Based Approach for Digital Signature AuthenticationabstractAbstract We propose a visualization based approach for digital signature authentication. Using our method, the speed and pressure aspects of a digital signature process can be clearly and intuitively conveyed to the user for digital signature authentication. Our design takes into account both the expressiveness and aesthetics of the derived visual patterns. With the visual aid provided by our method, digital signatures can be authenticated with better accuracy than using existing methods—even novices can examine the authenticity of a digital signature in most situations using our method. To validate the effectiveness of our method, we conducted a comprehensive user study which confirms positively the advantages of our approach. Our method can be employed as a new security enhancement measure for a range of business and legal applications in reality which involve digital signature authorization and authentication. Songhua Xu, Wenxia Yang, Francis C. M. Lau 0001 |
Comput. Graph. Forum | 1 |
| 2009 | Light source estimation of outdoor scenes for mixed reality
Yanli Liu 0006, Xueying Qin, Songhua Xu, Eihachiro Nakamae, Qunsheng Peng 0001 |
Vis. Comput. | 3 |
| 2008 | A User-Oriented Webpage Ranking Algorithm Based on User Attention Time
Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001 |
AAAI | 1 |
| 2008 | Introducing Affect into Competitive Game PlayabstractA dynamic notion of affect (degree of satisfaction) that an agent acquires in competitive iterated game play is developed. Simulated play against both a random environment and a competitor agent is used to study the impact of affect on game playing strategies. For definiteness, the formulation is framed in terms of stock trading. Comments on how affect in game play informs a notion of consciousness along with simulations are given. Maha Alabduljalil, Songhua Xu, Willard L. Miranker |
ICTAI (2) | 2 |
| 2008 | Constraint-Based Placement and Routing for FPGAs Using Self-Organizing MapsabstractField-programmable gate arrays (FPGAs) are becoming increasingly popular due to low design times, easy testing and implementation procedures and low costs. FPGAs placement and routing are NP-complete problems dealt well with modern tools using heuristic algorithms. As modern FPGAs increase in size and also new capabilities, such as run-time reconfiguration (RTR), are introduced, the complexity of these problems is greatly increased. In this paper we approach both problems using a modified version of Kohonen self-organizing map. The algorithm, consisting of four phases, takes into consideration constraints that may apply to the FPGA design (such as I/O pins, resource constraints like global clock etc). The modified algorithm yields a good topological map of the design to be placed, minimizing the average distance between connecting logic blocks. Michail Maniatakos, Songhua Xu, Willard L. Miranker |
ICTAI (2) | 2 |
| 2008 | A Two-Stage Audio Retrieval Method for Searching Unannotated Audio ClipsabstractTraditional audio retrieval systems deal principally with audio clips having text descriptions. To retrieve unannotated audio clips is cumbersome because of the immaturity of content-based analysis and retrieval techniques. In this paper, we propose a two-stage audio retrieval method, consisting of a first stage of text-based retrieval and a second stage of content-based retrieval. This new retrieval method can be employed to retrieve audio clips from an audio collection having only partial text annotations, which is true of many online audio datasets. We have developed a prototype audio retrieval system based on our algorithm and carefully evaluated its performance. The results demonstrate the effectiveness of our new audio retrieval method. Our method can be generalized and applied to other kinds of non-textual data such as images and videos. Songhua Xu, Suchao Chen, Kevin Y. Yip, Francis C. M. Lau 0001, Xueying Qin |
ISM | 1 |
| 2008 | Automatic Generation of Music Slide Show Using Personal PhotosabstractWe present an algorithmic system capable of automatically generating a music slide show given a piece of music with lyrics. Different from previous approaches, our method generates slide shows using personal photos which are without annotation. We introduce a novel algorithm to infer the relevance of personal photos to the lyrics, based on which personal photos are optimally selected to match the music. The proposed system first detects the keyframes of the input music. For each music keyframe, it optimally selects an image from the personal photo collection via our image content analysis procedure. Once the keyframe images have been selected, the in-between frames are then generated via an image morphing process. Experiment results have shown that our method can successfully generate music slide shows which follow the rhythms of the music and at the same time match the lyrics. Songhua Xu, Francis C. M. Lau 0001 |
ISM | 1 |
| 2008 | Personalized online document, image and video recommendation via commodity eye-trackingabstractWe propose a new recommendation algorithm for online documents, images and videos, which is personalized. Our idea is to rely on the attention time of individual users captured through commodity eye-tracking as the essential clue. The prediction of user interest over a certain online item (a document, image or video) is based on the user's attention time acquired using vision-based commodity eye-tracking during his previous reading, browsing or video watching sessions over the same type of online materials. After acquiring a user's attention times over a collection of online materials, our algorithm can predict the user's probable attention time over a new online item through data mining. Based on our proposed algorithm, we have developed a new online content recommender system for documents, images and videos. The recommendation results produced by our algorithm are evaluated by comparing with those manually labeled by users as well as by commercial search engines including Google (Web) Search, Google Image Search and YouTube. © 2008 ACM. Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001 |
RecSys | 1 |
| 2008 | Yale Image Finder (YIF): a new search engine for retrieving biomedical imagesabstractUNLABELLED: Yale Image Finder (YIF) is a publicly accessible search engine featuring a new way of retrieving biomedical images and associated papers based on the text carried inside the images. Image queries can also be issued against the image caption, as well as words in the associated paper abstract and title. A typical search scenario using YIF is as follows: a user provides few search keywords and the most relevant images are returned and presented in the form of thumbnails. Users can click on the image of interest to retrieve the high resolution image. In addition, the search engine will provide two types of related images: those that appear in the same paper, and those from other papers with similar image content. Retrieved images link back to their source papers, allowing users to find related papers starting with an image of interest. Currently, YIF has indexed over 140 000 images from over 34 000 open access biomedical journal papers. AVAILABILITY: http://krauthammerlab.med.yale.edu/imagefinder/ Songhua Xu, Jamie P. McCusker, Michael Krauthammer |
Bioinform. | 1 |
| 2008 | Automatic Facsimile of Chinese Calligraphic WritingsabstractAbstract To imitate personal handwritings is non‐trivial. In this paper, we attempt to address the challenging problem of automatic handwriting facsimile. We focus on Chinese calligraphic writings due to their rich variation in style, high artistic values and also the fact that they are among the most difficult candidates for the problem. We first analyze the structures and shapes of the constituent components, i.e., strokes and radicals, of characters in sample calligraphic writings by the same writer. To generate calligraphic writing in the style of the writer, we facsimile the individual character elements as well as the layout relationships used to compose the character, both in the writer's personal writing style. To test our algorithm, we compare our facsimileing results of Chinese calligraphic writings with the original writings. Our results are found to be acceptable for most cases, some of which are difficult to differentiate from the real ones. More results and supplementary materials are provided in our project website at http://www.cs.hku.hk/~songhua/facsimile/ . Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001, Yunhe Pan |
Comput. Graph. Forum | 1 |
| 2007 | An Intelligent System for Chinese Calligraphy
Songhua Xu, Hao Jiang 0011, Francis C. M. Lau 0001, Yunhe Pan |
AAAI | 1 |
| 2007 | The Mental Canvas: A Tool for Conceptual Architectural Design and AnalysisabstractWe describe a computer graphics system that supports conceptual architectural design and analysis. We use as a starting point the traditional sketchbook drawings that architects use to experiment with various views, sections, and details. Rather than interpret or infer 3D structure from drawings, our system is designed to allow the designer to organize concept drawings in 3D, and gradually fuse a series of possibly geometrically-inconsistent sketches into a set of 3D strokes. Our system uses strokes and planar "canvases" as basic primitives; the basic mode of input is traditional 2D drawing. We introduce methods for the user to control stroke visibility and transfer strokes between canvases. We also introduce methods for the user to position and orient the canvases that have infinite extent. We demonstrate the use of the system to analyze existing structures and conceive new designs. Julie Dorsey, Songhua Xu, Gabe Smedresman, Holly E. Rushmeier, Leonard McMillan |
PG | 2 |
| 2007 | A Generic Pigment Model for Digital PaintingabstractAbstract We propose a generic pigment model suitable for digital painting in a wide range of genres including traditional Chinese painting and water‐based painting. The model embodies a simulation of the pigment‐water solution and its interaction with the brush and the paper at the level of pigment particles; such a level of detail is needed for achieving highly intricate effects by the artist. The simulation covers pigment diffusion and sorption processes at the paper surface, and aspects of pigment particle deposition on the paper. We follow rules and formulations from quantitative studies of adsorption and diffusion processes in surface chemistry and the textile industry. The result is a pigment model that spans a continuum from the very wet to the very dry brush stroke effects. We also propose a new pigment mixing method based on machine learning techniques to emulate pigment mixing in real life as well as to support the creation of new artificial pigments. To experiment with the proposed model, we embedded the model in a sophisticated digital brush system. The combined system exhibits interactive speed on a modest PC platform. http://www.cs.hku.hk/~songhua/pigment provides supplementary materials for this paper. Songhua Xu, Haisheng Tan, Xiantao Jiao, Francis C. M. Lau 0001, Yunhe Pan |
Comput. Graph. Forum | 1 |
| 2006 | A Novel Method for Fast and High-Quality Rendering of Hair
Songhua Xu, Francis C. M. Lau 0001, Hao Jiang 0011, Yunhe Pan |
Rendering Techniques | 1 |
| 2006 | Animating Chinese paintings through stroke-based decompositionabstractThis article proposes a technique to animate a Chinese style painting given its image. We first extract descriptions of the brush strokes that hypothetically produced it. The key to the extraction process is the use of a brush stroke library, which is obtained by digitizing single brush strokes drawn by an experienced artist. The steps in our extraction technique are first to segment the input image, then to find the best set of brush strokes that fit the regions, and, finally, to refine these strokes to account for local appearance. We model a single brush stroke using its skeleton and contour, and we characterize texture variation within each stroke by sampling perpendicularly along its skeleton. Once these brush descriptions have been obtained, the painting can be animated at the brush stroke level. In this article, we focus on Chinese paintings with relatively sparse strokes. The animation is produced using a graphical application we developed. We present several animations of real paintings using our technique. Songhua Xu, Ying-Qing Xu, Sing Bing Kang, David Salesin, Yunhe Pan, Harry Shum |
ACM Trans. Graph. | 1 |
| 2005 | Virtual hairy brush for digital painting and calligraphyabstractThe design of user friendly and expressive virtual brush systems for interactive digital painting and calligraphy has attracted a lot of attention and effort in both computer graphics and human-computer interaction circles for a long time. Providing a digital environment for paper-less artwork creation is not only challenging in terms of algorithmic design, but also promising for its potential market values. This paper proposes a novel algorithmic framework for interactive digital painting and calligraphy based a novel virtual hairy brush model. The algorithms in the kernel of our simulation framework are built upon solid modeling techniques. Implementing the algorithms, we have developed a virtual hairy brush prototype system with which end users can interactively produce high-quality digital paintings and calligraphic artwork. (The latest progress of our virtual brush project is reported at the website "http://www.cs.hku.hk/~songhua/e-brush/".) Copyright by Science in China Press 2005. Songhua Xu, Francis C. M. Lau 0001, Congfu Xu, Yunhe Pan |
Sci. China Ser. F Inf. Sci. | 1 |
| 2004 | Automatic Generation of Artistic Chinese Calligraphy
Songhua Xu, Francis C. M. Lau 0001, William Kwok-Wai Cheung, Yunhe Pan |
AAAI | 1 |
| 2004 | Virtual hairy brush for painterly rendering
Songhua Xu, Min Tang 0001, Francis C. M. Lau 0001, Yunhe Pan |
Graph. Model. | 1 |
| 2003 | Advanced Design for a Realistic Virtual BrushabstractAbstract This paper proposes a novel algorithmic framework for an advanced virtual brush to be used in interactive digitalpainting. The framework comprises the following components: a geometry model of the brush using a hierarchicalrepresentation that leads to substantial savings in every step of the painting process; fast online brush motionsimulation assisted by offline calibration that guarantees an accurate and stable simulation of the brush's dynamicbehavior; a new pigment model based on a diffusion process of random molecules that considers delicateand complex pigment behavior at dipping time as well as during painting; and a user‐adaptation component thatenables the system to cater for the personal painting habits of different users. A prototype system has been implementedbased on this framework. Compared with other virtual brushes, this new system is designed to presenta realistic brush in the sense that the system accurately and stably simulates the complex painting functionalityof a running brush, and therefore is capable of creating high‐quality digital paintings with minute aesthetic detailsthat can rival the real artwork. The advanced features also give rise to a high degree of expressiveness ofthe virtual brush that the user can comfortably manipulate. http://www.csis.hku.hk/songhuale‐brush/ providessupplementary materials for this paper. Categories and Subject Descriptors (according to ACM CCS): I.3.6 [Methodology and Techniques]: Interactiontechniques; I.3.5 [Computational Geometry and Object Modeling]: Physically based modeling; I.3.4 [GraphicsUtilities]: Paint systems; Songhua Xu, Francis C. M. Lau 0001, Yunhe Pan |
Comput. Graph. Forum | 1 |
| 2002 | A Solid Model Based Virtual Hairy BrushabstractWe present the detailed modeling of the hairy brush used typically in Chinese calligraphy. The complex model, which includes also a model for the ink and the paper, covers the various stages of the brush going through a calligraphy process. The model relies on the concept of writing primitives, which are the smallest units of hair clusters, to reduce the load on the simulation. Each such primitive is constructed through the general sweeping operation in CAD and described by a NURBS surface. The writing primitives dynamically adjust themselves during the virtual writing process, leaving an imprint on the virtual paper as they move. The behavior of the brush is an aggregation of the behavior of all the writing primitives. A software system based on the model has been built and tested. Samples of imitation artwork from using the system were obtained and found to be nearly indistinguishable from the real artwork. Categories and Subject Descriptors (according to ACM CCS): I.3.6 [Methodology and Techniques]: Interaction techniques I.3.5 [Computational Geometry and Object Modeling]: Physically based modeling I.3.4 [Graphics Utilities]: Paint systems Songhua Xu, Min Tang 0001, Francis C. M. Lau 0001, Yunhe Pan |
Comput. Graph. Forum | 1 |
| 2001 | AI Supported Computer-Generated Pen-and-Ink IllustrationabstractIn the field of computer graphics there is an increasing demand for non-photorealistic effects. We add the idea of pattern recognition guided by theoretical rules into a traditional NPR system and create non-photorealistic drawings intelligently. The Style Sample Library functions as an expert library and makes it easier to render a picture with nonphotorealistic effects even by an amateur user. Yan Gu 0005, Songhua Xu, Min Tang 0001, Jinxiang Dong |
CSCWD | 2 |