VLDB 2026 Research / reviewers in the wild / expert
Shujin Lin
dblp:56/461
· DBLP profile ↗
35ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-9871-4014ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | QTG-VQA: Question-Type-Guided Architectural for VideoQA SystemsabstractIn the domain of video question answering (VideoQA), the impact of question types on VQA systems, despite its critical importance, has been relatively under-explored to date. However, the richness of question types directly determines the range of concepts a model needs to learn, thereby affecting the upper limit of its learning capability. This paper investigates the significance of different question types for VQA systems and their impact on performance, revealing issues such as insufficient learning and model degradation caused by the uneven distribution of question types. To address these issues, we propose QTG- VQA, a novel architecture that integrates question-type-guided attention mechanisms with an adaptive learning mechanism. For temporal-type questions, we further design the Masking Frame Modeling technique to enhance temporal modeling. Finally, a new evaluation metric tailored to question types is introduced. Experimental results confirm the effectiveness of the proposed method. Zhixian He, Shujin Lin |
ICME | 3 |
| 2025 | VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual AugmentationsabstractThe widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and talks, have led to a dramatic increase in speech-rich video content. In contrast to other video types, such as surveillance footage, which typically contain abundant visual cues, speech-rich videos convey most of their meaningful information through the audio channel. This poses challenges for improving content consumption using existing visual-based video summarization, navigation, and exploration systems. In this paper, we present VisAug, a novel interactive system designed to enhance speech-rich video navigation and engagement by automatically generating informative and expressive visual augmentations based on the speech content of videos. Our findings suggest that this system has the potential to significantly enhance the consumption and engagement of information in an increasingly video-driven digital landscape. Baoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang 0001, Fan Zhou 0001, Shujin Lin |
ACM Multimedia | 6 |
| 2025 | MGQA: Mixture Gaussian for Video Grounded Question Answering via VLMsabstractVideo question answering has become a cornerstone task for evaluating vision language models. However, existing models often fail to ground their answers in relevant visual evidence or incorrectly model distributions during localization. To address this limitation, we propose MGQA, which models videos as a sequence of discrete events using mixture Gaussian distributions, with each Gaussian characterized by its center, range, and weight. MGQA leverages question-answering accuracy as a weak supervision signal and incorporates two additional Gaussian-related loss functions. The method can be easily integrated into existing models with negligible parameter overhead. Experiments conducted on the NExT-GQA and ReX-Time datasets demonstrate the effectiveness of our proposed method. Zhixian He, Xiaofan Ma, Qiushi Li 0004, Shujin Lin |
SMC | 4 |
| 2024 | Hierarchical Attention Feature Fusion and Refinement Network for Point Cloud UpsamplingabstractThis paper presents a novel hierarchical attention feature fusion and refinement network designed to address challenges in existing deep learning based point cloud upsampling methods. The network combines self-attention layers with a multi-level feature extraction architecture, effectively integrating local and global features, thereby enhancing the robustness and uniformity of the point cloud. Furthermore, a spatial refinement module is employed to predict the offset between the generated coarse dense point clouds and real point clouds, thereby enhancing consistency with the ground truth. Concurrently, a filter function is applied in the loss function to handle outliers of generated point clouds. Extensive experimental results across multiple datasets indicate that our method outperforms existing approaches. Yaori Zhang, Shujin Lin, Fan Zhou 0001, Ruomei Wang 0001 |
ICME | 2 |
| 2024 | Point Cloud Completion Method Assisted by Projected ImageabstractImage-guided point cloud completion task aims to utilize image information to address the uncertainties in point cloud completion inference. Although acquiring 2D image data is relatively simpler than 3D data, it is still ineffective in scenarios with occlusions where image data cannot be reliably obtained as a reference. Therefore, we propose a point cloud completion model assisted by projected image data, which addresses the limitations of acquiring 2D images by constructing projected images of the point cloud. Extensive experiments demonstrate that our proposed method enhances the quality of point cloud completion and outperforms other advanced methods. Shujin Lin, Zhaowen Li, Runxun Wu, Fan Zhou 0001 |
SMC | 1 |
| 2023 | MIM: Lightweight Multi-Modal Interaction Model for Joint Video Moment Retrieval and Highlight DetectionabstractJoint video moment retrieval and highlight detection aims to find the relevant moments and highlight clips in a video with natural language. It is an emerging task though its individual problems have been studied for a while. The current methods utilize transformer to interact between modals, which leads to a huge cost of parameters and computation in spite of great performance. To address this problem, we present a cross-modal attention mechanism to capture related features from different modalities in a few-parameter way. Furthermore, a lightweight multi-modal interaction model (MIM) is proposed to solve video moment retrieval and highlight detection jointly. In the case of greatly reducing the number of parameters, we achieve competitive performance and faster convergence speed compared to previous method. Extensive experiments on four datasets demonstrate the effectiveness of our method. Shujin Lin, Fan Zhou 0001, Ruomei Wang 0001 |
ICME | 3 |
| 2023 | Image-Guided Point Cloud Completion with Multi-modal Fusion TransformersabstractThe task of image-guided point cloud completion aims to leverage information from images to address uncertainty issues in the completion inference of point clouds. The key challenge in this setting lies in how to effectively combine features extracted from both modalities. Due to the large domain discrepancy between the image and point cloud, existing methods that use cross-modal attention to directly fuse features have increased attention on redundant information and noise from different modalities, resulting in poor feature fusion performance. Hence, by introducing multi-modal fusion transformers that use bottleneck tokens, we enabled point cloud feature to learn image feature through information bridges, leading to improved point cloud completion performance. Our method can not only benefit from RGB images, but also from sketches with less feature information but more emphasis on edge information. Extensive experiments demonstrate that our proposed method enhances the quality of point cloud completion and outperforms other state-of-the-art methods. Zhaowen Li, Shujin Lin, Fan Zhou 0001 |
SMC | 2 |
| 2022 | NewsThumbnail: Automatic Generation of News Video ThumbnailabstractReading news is an important way for people to obtain information. People can quickly sort out the context of events through a short news video. However, there are numerous news generated around the world every day. It’s challenging to locate the interesting video. Thumbnails are often used as video covers and play an important role in displaying video content and driving views. Video owners can choose from individual images or elaborate thumbnails to upload to the site. But manually selecting from a large number of frames is time-consuming, and customizing thumbnails requires a high degree of expertise. Therefore, this paper proposes an automatic generation method of news video thumbnail, which can screen out semantically similar contents according to user query and combine them into a thumbnail. In order to facilitate the screening of graphic materials, we also propose a video content structuring method based on multiple cues, which can accurately segment the video into theme units. At the same time, we designed a visual system to display thumbnails and designed a user survey to investigate the performance of this method in news retrieval and understanding. Compared with peer methods, the thumbnails generated by our method can help users better understand the video content and locate the videos they are interested in. Shujin Lin, Fan Zhou 0001, Ruomei Wang 0001 |
SMC | 2 |
| 2022 | Temporal-aware Mechanism with Bidirectional Complementarity for Video Q&AabstractVideo question answering (Video Q&A) is a challenging task as it requires a sufficient understanding of the video and question information. Video is composed of frame sequence, which contains multi-scale temporal relationships and corresponding contextual information. A model competently tackle Video Q&A task that needs to be able to: 1) construct long-term and neighborhood dependencies in frame sequences to extract global and local contextual features that can reflect multi-scale temporal dependencies, and deduce the temporal-aware refined features, and 2) identify static and dynamic features from pertinent moments of a video, while filtering away question-irrelated dependencies of feature sequences, to yield the most precise and reasonable temporal-aware overall contextual features. In response to the above requirements, we propose a novel Video Q&A mechanism which consists of Bidirectional Complementary Attention(BCA) module and Adaptive Temporal-aware(ATA) module. Bidirectional complementary attention module stacks multi-head self-attention layer and convolutional layer in different orders to designed two kinds of attention units, which is able to make bidirectional multi-step reasoning based on complete global information and accurate local information to obtain temporal-aware refined features. Adaptive temporal-aware module is used to filter away question-irrelated dependencies in the feature sequence to yield the most precise and reasonable temporal-aware overall contextual features. Comprehensive comparative experiments are conducted on publicly available benchmark datasets. An extended ablation study is further conducted to show the usefulness of each module of the solution in acquiring its computational Q&A capabilities. Yuanmao Luo, Ruomei Wang 0001, Fan Zhou 0001, Shujin Lin |
SMC | 5 |
| 2021 | News2Mapping: A news events correlation model for news videosabstractNews video is an important way of news communication, and people can easily get news from all over the world through the Internet. However, it lacks in the organization of news video content about temporality, presentation and relevance and fails to express the correlation between news events. In this paper, we present an event correlation model of news videos to organize news content. News content is clustered through topics, and the relationship between events is illustrated through relationship mappings. To achieve this purpose, an XLNet-based language model is presented to extract news keywords and their relationships. The clustering algorithm is designed to obtain news event topic clustering and named entity clustering. At the same time, we build the relationship mappings in news events to visualize the correlation between news events better. The user study is also designed to investigate the performance of our method in news reading and understanding. Compared with peer methods, the news information organized by our method achieves a higher user satisfaction level. Mingjie Zhou, Ruomei Wang 0001, Shujin Lin, Fan Zhou 0001, Shirou Ou |
SMC | 3 |
| 2021 | LGCPNet : Local-global combined point-based network for shape segmentation
Boliang Guan, Fan Zhou 0001, Shujin Lin, Ruomei Wang 0001 |
Comput. Graph. | 4 |
| 2020 | Multi-column point-CNN for sketch segmentation
Fei Wang 0056, Shujin Lin, Hefeng Wu, Tie Cai, Ruomei Wang 0001 |
Neurocomputing | 2 |
| 2020 | Voxel-based quadrilateral mesh generation from point cloud
Boliang Guan, Shujin Lin, Ruomei Wang 0001, Fan Zhou 0001, Yongchuan Zheng |
Multim. Tools Appl. | 2 |
| 2019 | SPFusionNet: Sketch Segmentation Using Multi-modal Data FusionabstractThe sketch segmentation problem remains largely unsolved because conventional methods are greatly challenged by the highly abstract appearances of freehand sketches and their numerous shape variations. In this work, we tackle such challenges by exploiting different modes of sketch data in a unified framework. Specifically, we propose a deep neural network SPFusionNet to capture the characteristic of sketch by fusing from its image and point set modes. The image modal component SketchNet learns hierarchically abstract ro-bust features and utilizes multi-level representations to produce pixel-wise feature maps, while the point set-modal component SPointNet captures local and global contexts of the sampled point set to produce point-wise feature maps. Then our framework aggregates these feature maps by a fusion network component to generate the sketch segmentation result. The extensive experimental evaluation and comparison with peer methods on our large SketchSeg dataset verify the effectiveness of the proposed framework. Fei Wang 0056, Shujin Lin, Hefeng Wu, Ruomei Wang 0001, Xiangjian He |
ICME | 2 |
| 2019 | Lecture2Note: Automatic Generation of Lecture Notes from Slide-Based Educational VideosabstractGiven rapid development witnessed by open educational resources (OER) in the past few decades, a considerable number of online educational videos emerge on various MOOC platforms such as Coursera and YouTube. Nevertheless, most educational videos on the internet are lengthy and lack of elaborate annotations, which poses a challenge for learners to explore and locate content of interest efficiently. To address this, we present an automatic note-generating method to establish correspondences between visual entities in the slide-based lecture video and their descriptive speech texts by evaluating the semantic relationship. Firstly, the visual entities are extracted and recognised from the presentation slides. Then, each of visual entities is associated with its corresponding descriptive speech text. Finally, a placement optimisation scheme is put forward to pack the visual entities and speech texts into a note-like layout in a compact fashion, which can help learners to improve their learning efficiency. The experimental results show that the efficient performances about visual entity extraction and correspondence matching are efficient. The user study is also designed to investigate the performance of Lecture2Note in facilitating learning. Compared with peer methods, the auto-generated note created by our method achieves a higher user satisfaction level regarding a properly structured layout as well as efficient content navigation and exploration. Chengpei Xu, Ruomei Wang 0001, Shujin Lin, Baoquan Zhao, Lijie Shao, Mengqiu Hu |
ICME | 3 |
| 2019 | A New Visual Interface for Searching and Navigating Slide-Based Lecture VideosabstractThe rapid development of distance education technologies, e.g. MOOCs, provide learners unprecedented access to high-quality online lecture videos at scale, anytime and anywhere. Unfortunately, these valuable resources are often underutilized by online learners. One prevailing reason is the lack of support for and the resulting difficulty of exploring and locating content of interest among lengthy recordings of course lectures. To address this deficiency, we introduce a novel visual interface that supports efficient search and navigation of video content at fine granularities and with rich semantic clues. The interface is particularly designed for slide-based lecture videos (SBLV), which represent a significant portion of online lecture videos. The interface comprehensively derives versatile semantic clues for video content indexing and visual aid generation according to visual elements, text, and mathematical expressions included on lecture slides, speeches recorded, as well as mouse and cursor pointing actions captured during a lecture. Empowered by such semantically revealing indices and visual assistance, the interface is able to noticeably enhance online learners' capabilities in searching and browsing of content in need from SBLVs. The advantages of the new interface are demonstrated through benchmarked experimental results in comparison with peer methods. Baoquan Zhao, Songhua Xu, Shujin Lin, Ruomei Wang 0001 |
ICME | 3 |
| 2019 | SFSegNet: Parse Freehand Sketches using Deep Fully Convolutional NetworksabstractParsing sketches via semantic segmentation is attractive but challenging, because (i) free-hand drawings are abstract with large variances in depicting objects due to different drawing styles and skills; (ii) distorting lines drawn on the touchpad make sketches more difficult to be recognized; (iii) the high-performance image segmentation via deep learning technologies needs enormous annotated sketch datasets during the training stage. In this paper, we propose a Sketch-target deep FCN Segmentation Network(SFSegNet) for automatic free-hand sketch segmentation, labeling each sketch in a single object with multiple parts. SFSegNet has an end-to-end network process between the input sketches and the segmentation results, composed of 2 parts: (i) a modified deep Fully Convolutional Network(FCN) using a reweighting strategy to ignore background pixels and classify which part each pixel belongs to; (ii) affine transform encoders that attempt to canonicalize the shaking strokes. We train our network with the dataset that consists of 10,000 annotated sketches, to find an extensively applicable model to segment stokes semantically in one ground truth. Extensive experiments are carried out and segmentation results show that our method outperforms other state-of-the-art networks. Junkun Jiang, Ruomei Wang 0001, Shujin Lin, Fei Wang 0056 |
IJCNN | 3 |
| 2018 | 3D medical model low-pass filtering based on non-uniform spectral synthesis
Yihui Guo, Zhuo Su 0001, Shujin Lin, Jiyuan Lu, Xueling Zhong |
Comput. Aided Des. | 3 |
| 2018 | PSI: A probabilistic semantic interpretable framework for fine-grained image rankingabstractImage Ranking is one of the key problems in information science research area. However, most current methods focus on increasing the performance, leaving the semantic gap problem, which refers to the learned ranking models are hard to be understood, remaining intact. Therefore, in this article, we aim at learning an interpretable ranking model to tackle the semantic gap in fine‐grained image ranking. We propose to combine attribute‐based representation and online passive‐aggressive (PA) learning based ranking models to achieve this goal. Besides, considering the highly localized instances in fine‐grained image ranking, we introduce a supervised constrained clustering method to gather class‐balanced training instances for local PA‐based models, and incorporate the learned local models into a unified probabilistic framework. Extensive experiments on the benchmark demonstrate that the proposed framework outperforms state‐of‐the‐art methods in terms of accuracy and speed. Hefeng Wu, Shujin Lin, Zhuo Su 0001 |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2018 | A new sketch-based 3D model retrieval method by using composite features
Yuhua Li 0002, Hao-Peng Lei, Shujin Lin, Guoliang Luo |
Multim. Tools Appl. | 3 |
| 2017 | Multi-view pairwise relationship learning for sketch based 3D shape retrievalabstractRecent progress in sketch-based 3D shape retrieval creates a novel and user-friendly way to explore massive 3D shapes on the Internet. However, current methods on this topic rely on designing invariant features for both sketches and 3D shapes, or complex matching strategies. Therefore, they suffer from problems like arbitrary drawings and inconsistent viewpoints. To tackle this problem, we propose a probabilistic framework based on Multi-View Pairwise Relationship (MVPR) learning. Our framework includes multiple views of 3D shapes as the intermediate layer between sketches and 3D shapes, and transforms the original retrieval problem into the form of inferring pairwise relationship between sketches and views. We accomplish pairwise relationship inference by a novel MVPR net, which can automatically predict and merge the pairwise relationships between a sketch and multiple views, thus freeing us from exhaustively selecting the best view of 3D shapes. We also propose to learn robust features for sketches and views via fine-tuning pre-trained networks. Extensive experiments on a large dataset demonstrate that the proposed method can outperform state-of-the-art methods significantly. Hefeng Wu, Xiangjian He, Shujin Lin, Ruomei Wang 0001 |
ICME | 4 |
| 2017 | A Novel System for Visual Navigation of Educational Videos Using Multimodal CuesabstractWith recent developments and advances in distance learning and MOOCs, the amount of open educational videos on the Internet has grown dramatically in the past decade. However, most of these videos are lengthy and lack of high-quality indexing and annotations, which triggers an urgent demand for efficient and effective tools that facilitate video content navigation and exploration. In this paper, we propose a novel visual navigation system for exploring open educational videos. The system tightly integrates multimodal cues obtained from the visual, audio and textual channels of the video and presents them with a series of interactive visualization components. With the help of this system, users can explore the video content using multiple levels of details to identify content of interest with ease. Extensive experiments and comparisons against previous studies demonstrate the effectiveness of the proposed system. Baoquan Zhao, Shujin Lin, Songhua Xu, Ruomei Wang 0001 |
ACM Multimedia | 2 |
| 2017 | A Data-Driven Approach for Sketch-Based 3D Shape Retrieval via Similar Drawing-Style RecommendationabstractAbstract Sketching is a simple and natural way of expression and communication for humans. For this reason, it gains increasing popularity in human computer interaction, with the emergence of multitouch tablets and styluses. In recent years, sketch‐based interactive methods are widely used in many retrieval systems. In particular, a variety of sketch‐based 3D model retrieval works have been presented. However, almost all of these works focus on directly matching sketches with the projection views of 3D models, and they suffer from the large differences between the sketch drawing and the views of 3D models, leading to unsatisfying retrieval results. Therefore, in this paper, during the matching procedure in the retrieval, we propose to match the sketch with each 3D model from historical users instead of projection views. Yet since the sketches between the current user and the historical users can have big difference, we also aim to handle users' personalized deviations and differences. To this end, we leverage recommendation algorithms to estimate the drawing style characteristic similarity between the current user and historical users. Experimental results on the Large Scale Sketch Track Benchmark(SHREC14LSSTB) demonstrate that our method outperforms several state‐of‐the‐art methods. Fei Wang 0056, Shujin Lin, Hefeng Wu, Ruomei Wang 0001, Fan Zhou 0001 |
Comput. Graph. Forum | 2 |
| 2017 | Distortion-Aware Correlation TrackingabstractRecently, correlation filter (CF)-based tracking methods have attracted considerable attention because of their high-speed performance. However, distortion, which refers to the phenomenon that the correlation outputs of CF-based trackers are distorted, remains a major obstacle for these methods. In this paper, we propose a distortion-aware correlation filter framework, which can detect distortions and recover from tracking failures. Our framework employs a simple yet effective feature termed normed correlation response to detect distortions. Meanwhile, we introduce a competition mechanism to handle distortions, in which we build a specialized graph to formulate and handle tracking under distortion as a maximum multi clique problem. Furthermore, a global-local context model is exploited to alleviate underlying distortions during the tracking process. Extensive experiments on the Online Tracking Benchmark show that our tracker can find the optimal target trajectory during the distortion period and retrieve the possibly missing target, consequently outperforms the state-of-the-art methods and improves the performance of CF-based trackers favorably. Hefeng Wu, Huifang Zhang, Shujin Lin, Ruomei Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Boosting Zero-Shot Image Classification via Pairwise Relationship Learning
Hefeng Wu, Shujin Lin, Ebroul Izquierdo |
ACCV (1) | 3 |
| 2016 | A new visual navigation system for exploring biomedical Open Educational Resource (OER) videosabstractOBJECTIVE: Biomedical videos as open educational resources (OERs) are increasingly proliferating on the Internet. Unfortunately, seeking personally valuable content from among the vast corpus of quality yet diverse OER videos is nontrivial due to limitations of today's keyword- and content-based video retrieval techniques. To address this need, this study introduces a novel visual navigation system that facilitates users' information seeking from biomedical OER videos in mass quantity by interactively offering visual and textual navigational clues that are both semantically revealing and user-friendly. MATERIALS AND METHODS: The authors collected and processed around 25 000 YouTube videos, which collectively last for a total length of about 4000 h, in the broad field of biomedical sciences for our experiment. For each video, its semantic clues are first extracted automatically through computationally analyzing audio and visual signals, as well as text either accompanying or embedded in the video. These extracted clues are subsequently stored in a metadata database and indexed by a high-performance text search engine. During the online retrieval stage, the system renders video search results as dynamic web pages using a JavaScript library that allows users to interactively and intuitively explore video content both efficiently and effectively.ResultsThe authors produced a prototype implementation of the proposed system, which is publicly accessible athttps://patentq.njit.edu/oer To examine the overall advantage of the proposed system for exploring biomedical OER videos, the authors further conducted a user study of a modest scale. The study results encouragingly demonstrate the functional effectiveness and user-friendliness of the new system for facilitating information seeking from and content exploration among massive biomedical OER videos. CONCLUSION: Using the proposed tool, users can efficiently and effectively find videos of interest, precisely locate video segments delivering personally valuable information, as well as intuitively and conveniently preview essential content of a single or a collection of videos. Baoquan Zhao, Songhua Xu, Shujin Lin |
J. Am. Medical Informatics Assoc. | 3 |
| 2016 | A 3D model perceptual feature metric based on global height field
Yihui Guo, Shujin Lin, Zhuo Su 0001, Ruomei Wang 0001, Yang Kang |
Vis. Comput. | 2 |
| 2015 | 3D Model retrieval based on skeletonabstractWe proposed a new method for 3D model retrieval based on skeletons. As we know a sketch is drawn on a two dimensional plane while models are three dimensional. For better comparison we choose the same dimension for them. For simpler user input we use the front view of the skeleton to represent 3D models, in this way 3D models can be represented in 2D form. We get the skeleton of model through Skeleton Extraction algorithm based on mesh simplification and mesh contraction and get the front view of the skeleton by mapping to two dimensional spaces. The model compared with the Model libraries by Feature description and matching to find similar models on shape and topology structure after getting the feature extraction of the model by the gradient histogram matching algorithm. Experiments showed that most users can find the 3D models they want with our system. Shujin Lin, Yihui Guo, Yun Liang 0003, Yanhua Wu |
NAS | 1 |
| 2015 | Competitive and cooperative particle swarm optimization with information sharing mechanism for global optimization problems
Yuhua Li 0002, Zhi-hui Zhan, Shujin Lin, Jun Zhang 0003 |
Inf. Sci. | 3 |
| 2014 | 3D Model Editing from Contour Drawings on Orthographic Projection Views
Yuhui Hu, Xuliang Guo, Baoquan Zhao, Shujin Lin |
ICISP | 4 |
| 2014 | A new algorithm for product image search based on salient edge characterizationabstractVisually assisted product image search has gained increasing popularity because of its capability to greatly improve end users' e‐commerce shopping experiences. Different from general‐purpose content‐based image retrieval (CBIR) applications, the specific goal of product image search is to retrieve and rank relevant products from a large‐scale product database to visually assist a user's online shopping experience. In this paper, we explore the problem of product image search through salient edge characterization and analysis, for which we propose a novel image search method coupled with an interactive user region‐of‐interest indication function. Given a product image, the proposed approach first extracts an edge map, based on which contour curves are further extracted. We then segment the extracted contours into fragments according to the detected contour corners. After that, a set of salient edge elements is extracted from each product image. Based on salient edge elements matching and similarity evaluation, the method derives a new pairwise image similarity estimate. Using the new image similarity, we can then retrieve product images. To evaluate the performance of our algorithm, we conducted 120 sessions of querying experiments on a data set comprised of around 13k product images collected from multiple, real‐world e‐commerce websites. We compared the performance of the proposed method with that of a bag‐of‐words method (Philbin, Chum, Isard, Sivic, & Zisserman, 2008) and a Pyramid Histogram of Orientated Gradients (PHOG) method (Bosch, Zisserman, & Munoz, 2007). Experimental results demonstrate that the proposed method improves the performance of example‐based product image retrieval. Yuhua Li 0002, Songhua Xu, Shujin Lin |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2013 | A new interpolation subdivision scheme for triangle/quad mesh
Shujin Lin, Songhua Xu, Jianmin Wang 0013 |
Graph. Model. | 1 |
| 2013 | Automatic 3D Shape Co-Segmentation Using Spectral Graph Method
Hao-Peng Lei, Shujin Lin, Jian-Qiang Sheng |
J. Comput. Sci. Technol. | 3 |
| 2008 | Interpolatory and Mixed Loop SchemesabstractAbstract This paper presents a new interpolatory Loop scheme and an unified and mixed interpolatory and approximation subdivision scheme for triangular meshes. The former which is C1 continuous as same as the modified Butterfly scheme has better effect in some complex models. The latter can be used to solve the “popping effect” problem when switching between meshes at different levels of resolution. The scheme generates surfaces coincident with the Loop subdivision scheme in the limit condition having the coefficient k equal 0. When k equal 1, it will be changed into a new interpolatory subdivision scheme. Eigen‐structure analysis demonstrates that subdivision surfaces generated using the new scheme are C1 continuous. All these are achieved only by changing the value of a parameter k. The method is a completely simple one without constructing and solving equations. It can achieve local interpolation and solve the “popping effect” problem which are the method's advantages over the modified Butterfly scheme. Zhuo Shi, Shujin Lin, Renhong Wang 0001 |
Comput. Graph. Forum | 2 |
| 2008 | Deducing interpolating subdivision schemes from approximating subdivision schemesabstractIn this paper we describe a method for directly deducing new interpolating subdivision masks for meshes from corresponding approximating subdivision masks. The purpose is to avoid complex computation for producing interpolating subdivision masks on extraordinary vertices. The method can be applied to produce new interpolating subdivision schemes, solve some limitations in existing interpolating subdivision schemes and satisfy some application needs. As cases, in this paper a new interpolating subdivision scheme for polygonal meshes is produced by deducing from the Catmull-Clark subdivision scheme. It can directly operate on polygonal meshes, which solves the limitation of Kobbelt's interpolating subdivision scheme. A new √3 interpolating subdivision scheme for triangle meshes and a new √2 interpolating subdivision scheme for quadrilateral meshes are also presented in the paper by deducing from √3 subdivision schemes and 4-8 subdivision schemes respectively. They both produce C 1 continuous limit surfaces and avoid the blemish in the existing interpolating √3 and √2 subdivision masks where the weight coefficients on extraordinary vertices can not be described by formulation explicitly. In addition, by adding a parameter to control the transition from approximation to interpolation, they can produce surfaces intervening between approximating and interpolating which can be used to solve the "popping effect" problem when switching between meshes at different levels of resolution. They can also force surfaces to interpolate chosen vertices. Shujin Lin, Fang You |
ACM Trans. Graph. | 1 |