VLDB 2026 Research / reviewers in the wild / expert
Qian Lin 0001
dblp:79/3108-1
· DBLP profile ↗
33ranked-venue papers
2as first author
11since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 3 since 2021Systems, architecture and hardware · 8 · 8 since 2021Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | FedTR: Federated Learning Framework with Transfer Learning for Industrial Visual InspectionabstractFederated learning (FL) is a collaborative learning scheme to train deep learning models, where collaborating parties can consolidate their models without sharing local data with other parties, hence preserving data privacy. Nevertheless, when implementing FL in Industrial visual inspection (IVI), the constraints posed by limited data availability and the intricate nature of the inspection tasks significantly impact the performance of the resulting model. This paper introduces FedTR, a novel FL framework incorporating transfer learning designed for Autonomous IVI, focusing on the challenging task of identifying label defects through end-to-end text recognition. Transfer learning is a method that leverages the knowledge of a pre-trained model to adapt to a different dataset. FedTR initially trains the model using a publicly available dataset, after which performs the essential federated learning process with model fine-tuning on the distributed and limited private data. Extensive experiment results demonstrate the effectiveness and feasibility of FedTR on private ink cartridge datasets for label defect identification. FedTR achieves an end-to-end text recognition word-level accuracy of 95.5% and 94.2% on homogeneous and heterogeneous data respectively. Additionally, it attains performance levels that are on par with those achieved through centralized training. Vikash Sathiamoorthy, Shuo Huai, Hao Kong 0001, Di Liu 0002, Wendy Yong Yi Loy, Christian Makaya, Daren Ho, Ravi Subramaniam, Qian Lin 0001, Weichen Liu 0001 |
ACM Great Lakes Symposium on VLSI | 9 |
| 2024 | Pose Guided Portrait View Interpolation from Dual Cameras with a Long BaselineabstractWe introduce a novel method for interpolating views between two fixed cameras to create a free-viewpoint experience comparable to advanced telepresence systems. Inspired by Video Frame Interpolation (VFI) works, our proposed method spatially interpolates between camera views, treating our setup as a unique frame interpolation scenario. Due to the occlusion that exists in areas like the face contour from two cameras with a long baseline, which results in serious artifacts with previous VFI methods. To mitigate this issue, we integrate face pose information of pitch, yaw and roll angles in our method as a robust guidance prior to improve pixel mapping in occluded regions. Furthermore, we synthesize a large portrait-focused multi-view dataset facilitating training of our proposed model. Our multi-stage flow refinement model progressively refines the quality of the bi-directional flows which eventually results in improved interpolated view details. With the flexible setup and fast inference speed of our proposed method, a practical and cost-effective solution can be implemented for telepresence systems without the need for expensive hardware or extensive computational resources. Weichen Xu 0002, Yezhi Shen, Qian Lin 0001, Jan P. Allebach, Fengqing Zhu 0001 |
MMSP | 3 |
| 2023 | Towards Efficient Convolutional Neural Network for Embedded Hardware via Multi-Dimensional PruningabstractIn this paper, we propose TECO, a multi-dimensional pruning framework to collaboratively prune the three dimensions (depth, width, and resolution) of convolutional neural networks (CNNs) for better execution efficiency on embedded hardware. In TECO, we first introduce a two-stage importance evaluation framework, which efficiently and comprehensively evaluates each pruning unit according to both the local importance inside each dimension and the global importance across different dimensions. Based on the evaluation framework, we present a heuristic pruning algorithm to progressively prune the three dimensions of CNNs towards the optimal trade-off between accuracy and efficiency. Experiments on multiple benchmarks validate the advantages of TECO over existing state-of-the-art (SOTA) approaches. The code and pre-trained models are available anonymously at https://github.com/ntuliuteam/Teco. Hao Kong 0001, Di Liu 0002, Shuo Huai, Ravi Subramaniam, Christian Makaya, Qian Lin 0001, Weichen Liu 0001 |
DAC | 7 |
| 2023 | EMNAPE: Efficient Multi-Dimensional Neural Architecture Pruning for EdgeAIabstractIn this paper, we propose a multi-dimensional pruning framework, EMNAPE, to jointly prune the three dimensions (depth, width, and resolution) of convolutional neural networks (CNNs) for better execution efficiency on embedded hardware. In EMNAPE, we introduce a two-stage evaluation strategy to evaluate the importance of each pruning unit and identify the computational redundancy in the three dimensions. Based on the evaluation strategy, we further present a heuristic pruning algorithm to progressively prune redundant units from the three dimensions for better accuracy and efficiency. Experiments demonstrate the superiority of EMNAPE over existing methods. Hao Kong 0001, Shuo Huai, Di Liu 0002, Ravi Subramaniam, Christian Makaya, Qian Lin 0001, Weichen Liu 0001 |
DATE | 7 |
| 2023 | Efficient Joint Video Denoising and Super-ResolutionabstractDenoising and super-resolution are two important tasks for video enhancement. Despite recent progress for each task, there are very few works that target both tasks simultaneously. In this paper, we propose an efficient noise-robust video super-resolution method that is trained end-to-end for an input video containing observable noises. We investigate current approaches to address this joint denoising and super-resolution task and compare them to our proposed method. Experimental results show that our method achieves competitive reconstruction performance with existing solutions on various datasets while maintaining a low computation cost and a small model size which prove the effectiveness of our joint model design and training. Our code is available at "https://github.com/Eventhyn/EVDSRNet.". Yuning Huang, Qian Lin 0001, Jan P. Allebach, Fengqing Zhu 0001 |
ICIP | 3 |
| 2023 | Latency-constrained DNN architecture learning for edge systems using zerorized batch normalization
Shuo Huai, Di Liu 0002, Hao Kong 0001, Weichen Liu 0001, Ravi Subramaniam, Christian Makaya, Qian Lin 0001 |
Future Gener. Comput. Syst. | 7 |
| 2023 | EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAIabstractConvolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of CNNs hinders the deployment of CNNs onto resource-constrained embedded devices. To address this issue, we propose EdgeCompress, a comprehensive compression framework to reduce the computational overhead of CNNs. In EdgeCompress, we first introduce dynamic image cropping (DIC), where we design a lightweight foreground predictor to accurately crop the most informative foreground object of input images for inference, which avoids redundant computation on background regions. Subsequently, we present compound shrinking (CS) to collaboratively compress the three dimensions (depth, width, and resolution) of CNNs according to their contribution to accuracy and model computation. DIC and CS together constitute a multidimensional CNN compression framework, which is able to comprehensively reduce the computational redundancy in both input images and neural network architectures, thereby improving the inference efficiency of CNNs. Further, we present a dynamic inference framework to efficiently process input images with different recognition difficulties, where we cascade multiple models with different complexities from our compression framework and dynamically adopt different models for different input images, which further compresses the computational redundancy and improves the inference efficiency of CNNs, facilitating the deployment of advanced CNNs onto embedded hardware. Experiments on ImageNet-1K demonstrate that EdgeCompress reduces the computation of ResNet-50 by 48.8% while improving the top-1 accuracy by 0.8%. Meanwhile, we improve the accuracy by 4.1% with similar computation compared to HRank. The state-of-the-art compression framework. The source code and models are available athttps://github.com/ntuliuteam/edge-compress. Hao Kong 0001, Di Liu 0002, Shuo Huai, Ravi Subramaniam, Christian Makaya, Qian Lin 0001, Weichen Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | CRIMP: Compact & Reliable DNN Inference on In-Memory Processing via Crossbar-Aligned Compression and Non-ideality AdaptationabstractCrossbar-based In-Memory Processing (IMP) accelerators have been widely adopted to achieve high-speed and low-power computing, especially for deep neural network (DNN) models with numerous weights and high computational complexity. However, the floating-point (FP) arithmetic is not compatible with crossbar architectures. Also, redundant weights of current DNN models occupy too many crossbars, limiting the efficiency of crossbar accelerators. Meanwhile, due to the inherent non-ideal behavior of crossbar devices, like write variations, pre-trained DNN models suffer from accuracy degradation when it is deployed on a crossbar-based IMP accelerator for inference. Although some approaches are proposed to address these issues, they often fail to consider the interaction among these issues, and introduce significant hardware overhead for solving each issue. To deploy complex models on IMP accelerators, we should compact the model and mitigate the influence of device non-ideal behaviors without introducing significant overhead from each technique. In this paper, we first propose to reuse bit-shift units in crossbars for approximately multiplying scaling factors in our quantization scheme to avoid using FP processors. Second, we propose to apply kernel-group pruning and crossbar pruning to eliminate the hardware units for data aligning. We also design a zerorize-recover training process for our pruning method to achieve higher accuracy. Third, we adopt the runtime-aware non-ideality adaptation with a self-compensation scheme to relieve the impact of non-ideality by exploiting the feature of crossbars. Finally, we integrate these three optimization procedures into one training process to form a comprehensive learning framework for co-optimization, which can achieve higher accuracy. The experimental results indicate that our comprehensive learning framework can obtain significant improvements over the original model when inferring on the crossbar-based IMP accelerator, with an average reduction of computing power and computing area by 100.02× and 17.37×, respectively. Furthermore, we can obtain totally integer-only, pruned, and reliable VGG-16 and ResNet-56 models for the Cifar-10 dataset on IMP accelerators, with accuracy drops of only 2.19% and 1.26%, respectively, without any hardware overhead. Shuo Huai, Hao Kong 0001, Shiqing Li, Ravi Subramaniam, Christian Makaya, Qian Lin 0001, Weichen Liu 0001 |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2022 | Smart Scissor: Coupling Spatial Redundancy Reduction and CNN Compression for Embedded HardwareabstractScaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising for edge AI. However, as an image usually contains much spatial redundancy, e.g., background pixels, directly shrinking the whole image will lose important features of the foreground object and lead to severe accuracy degradation. In this paper, we propose a dynamic image cropping framework to reduce the spatial redundancy by accurately cropping the foreground object from images. To achieve the instance-aware fine cropping, we introduce a lightweight foreground predictor to efficiently localize and crop the foreground of an image. The finely cropped images can be correctly recognized even at a small resolution. Meanwhile, computational redundancy also exists in CNN architectures. To pursue higher execution efficiency on resource-constrained embedded devices, we also propose a compound shrinking strategy to coordinately compress the three dimensions (depth, width, resolution) of CNNs. Eventually, we seamlessly combine the proposed dynamic image cropping and compound shrinking into a unified compression framework, Smart Scissor, which is expected to significantly reduce the computational overhead of CNNs while still maintaining high accuracy. Experiments on ImageNet-1K demonstrate that our method reduces the computational cost of ResNet50 by 41.5% while improving the top-1 accuracy by 0.3%. Moreover, compared to HRank, the state-of-the-art CNN compression framework, our method achieves 4.1% higher top-1 accuracy at the same computational cost. The codes and data are available at https://github.com/ntuliuteam/smart-scissor Hao Kong 0001, Di Liu 0002, Shuo Huai, Weichen Liu 0001, Ravi Subramaniam, Christian Makaya, Qian Lin 0001 |
ICCAD | 8 |
| 2022 | Collate: Collaborative Neural Network Learning for Latency-Critical Edge SystemsabstractFederated Learning (FL) empowers multiple clients to collaboratively learn a model, enlarging the training data of each client for high accuracy while protecting data privacy. However, when deploying FL in real-time edge systems, the heterogeneity of devices among systems has a severe impact on the performance of the inferred model. Existing optimizations on FL focus on improving the training efficiency but fail to speed up inference, especially when there is a latency constraint. In this work, we propose Collate, a novel training framework that collaboratively learns heterogeneous models to meet the latency constraints of multiple edge systems simultaneously. We design a dynamic zeroizing-recovering method to adjust each local model architecture for high accuracy under its latency constraint. A proto-corrected federated aggregation scheme is also introduced to aggregate all heterogeneous local models, satisfying the latency constraint of different systems with only one training process and maintaining high accuracy. Extensive experiments indicate that, compared to state-of-the-art methods and under a latency constraint, our extended models can improve the accuracy by 1.96% on average, and our shrunk models can also obtain a 3.09% accuracy improvement on average, with almost no extra training overhead. The related codes and data will be available at https://github.com/ntuliuteam/Collate. Shuo Huai, Di Liu 0002, Hao Kong 0001, Weichen Liu 0001, Ravi Subramaniam, Christian Makaya, Qian Lin 0001 |
ICCD | 8 |
| 2022 | Re-Compose the Image by Evaluating the Crop on More Than Just a ScoreabstractImage re-composition has always been regarded as one of the most important steps during the post-processing of a photo. The quality of an image re-composition mainly depends on a person’s taste in aesthetics, which is not an effortless task for those who have no abundant experience in photography. Besides, while re-composing one image does not require much of a person’s time, it could be quite time-consuming when there are hundreds of images to be recomposed. To solve these problems, we propose a method that automates the process of re-composing an image to the desired aspect ratio. Although there already exist many image re-composition methods, they only provide a score to their predicted best crop but fail to explain why the score is high or low. Conversely, we succeed in designing an explainable method by introducing a novel 10-layer aesthetic score map, which represents how the position of the saliency in the original uncropped image, relative to that of the crop region, contributes to the overall score of the crop, so that the crop is not just represented by a single score. We conducted experiments to show that the proposed score map boosts the performance of our algorithm, which achieves a state-of-the-art performance on both public and our own datasets. Qian Lin 0001, Jan P. Allebach |
WACV | 2 |
| 2020 | Boosting High-Level Vision with Joint Compression Artifacts Reduction and Super-ResolutionabstractDue to the limits of bandwidth and storage space, digital images are usually down-scaled and compressed when transmitted over networks, resulting in loss of details and jarring artifacts that can lower the performance of high-level visual tasks. In this paper, we aim to generate an artifact-free high-resolution image from a low-resolution one compressed with an arbitrary quality factor by exploring joint compression artifacts reduction (CAR) and super-resolution (SR) tasks. First, we propose a context-aware joint CAR and SR neural network (CAJNN) that integrates both local and non-local features to solve CAR and SR in one-stage. Finally, a deep reconstruction network is adopted to predict high quality and high-resolution images. Evaluation on CAR and SR benchmark datasets shows that our CAJNN model outperforms previous methods and also takes 26.2% shorter runtime. Based on this model, we explore addressing two critical challenges in high-level computer vision: optical character recognition of low-resolution texts, and extremely tiny face detection. We demonstrate that CAJNN can serve as an effective image preprocessing method and improve the accuracy for real-scene text recognition (from 85.30% to 85.75%) and the average precision for tiny face detection (from 0.317 to 0.611). Xiaoyu Xiang, Qian Lin 0001, Jan P. Allebach |
ICPR | 2 |
| 2020 | Print Defect Mapping with Semantic SegmentationabstractEfficient automated print defect mapping is valuable to the printing industry since such defects directly influence customer-perceived printer quality and manually mapping them is cost-ineffective. Conventional methods consist of complicated and hand-crafted feature engineering techniques, usually targeting only one type of defect. In this paper, we propose the first end-to-end framework to map print defects at pixel level, adopting an approach based on semantic segmentation. Our framework uses Convolutional Neural Networks, specifically DeepLab-v3+, and achieves promising results in the identification of defects in printed images. We use synthetic training data by simulating two types of print defects and a print-scan effect with image processing and computer graphic techniques. Compared with conventional methods, our framework is versatile, allowing two inference strategies, one being near real-time and providing coarser results, and the other focusing on offline processing with more fine-grained detection. Our model is evaluated on a dataset of real printed images. Augusto C. Valente, Cristina Wada, Deangela Neves, Deangeli Neves, Fábio V. M. Perez, Guilherme A. S. Megeto, Marcos H. Cascone, Otavio Gomes, Qian Lin 0001 |
WACV | 9 |
| 2019 | High-Accuracy Automatic Person Segmentation with Novel Spatial Saliency MapabstractIn this work, we propose a high-efficiency person segmentation system that achieves high segmentation accuracy with a much smaller CNN network. In this approach, key-point detection annotation is incorporated for the first time and a novel spatial saliency map, in which the intensity of each pixel indicates the likelihood of forming a part of the human and reflects the distance from the body, is generated to provide more spatial information. Additionally, a lightweight automatic person segmentation network is proposed, which is small and efficient for person segmentation by leveraging atrous convolution. The experimental results prove that an image pyramid resizing augmentation can also improve efficiency. Our proposed segmentation method achieves an accuracy of 94.06% on the person segmentation dataset built in this work, which exceeds the results of previous state-of-the-art methods in accuracy and efficiency. Weijuan Xi, Jianhang Chen, Qian Lin 0001, Jan P. Allebach |
ICIP | 3 |
| 2017 | A Study of Web Print: What People Print in the Digital EraabstractThis article analyzes a proprietary log of printed web pages and aims at answering questions regarding the content people print (what), the reasons they print (why), as well as attributes of their print profile (who). We present a classification of pages printed based on their print intent and we describe our methodology for processing the print dataset used in this study. In our analysis, we study the web sites, topics, and print intent of the pages printed along the following aspects: popularity, trends, activity, user diversity, and consistency. We present several findings that reveal interesting insights into printing. We analyze our findings and discuss their impact and directions for future work. Georgia Koutrika, Qian Lin 0001 |
ACM Trans. Web | 2 |
| 2013 | Feature design for aesthetic inference on photos with facesabstractDetermining the aesthetics of photographs has recently become a research topic of considerable interest. In this project, we focus on constructing meaningful features to model the aesthetic quality of photos with faces. Utilizing face information, color, composition features, as well as novel saliency-based spatial features, we construct an aesthetic inference model, which is more accurate than a state-of-the-art method. Further, we show that this model can be improved by applying different sets of features for single-face and multiple-face photos. Third, we demonstrate by combining low-level generic features with handcrafted features, that the model can be made to achieve even lower error rates. Shao-Fu Xue, Henry Tang, Daniel Tretter, Qian Lin 0001, Jan P. Allebach |
ICIP | 4 |
| 2013 | Recommendation system for automatic design of magazine coversabstractIn this paper, we present a recommendation system for the automatic design of magazine covers. Our users are non-designer designers: individuals or small and medium businesses who want to design without hiring a professional designer while still wanting to create aesthetically compelling designs. Because a design should have a purpose, we suggest a number of semantic features to the user, e.g., "clean and clear," "dynamic and active," or "formal," to describe the color mood for the purpose of his/her design. Based on these high level features and a number of low level features, such as the complexity of the visual balance in a photo, our system selects the best photos from the user's album for his/her design. Our system then generates several alternative designs that can be rated by the user. Consequently, our system generates future designs based on the user's style. In this fashion, our system personalizes the designs of a user based on his/her preferences. Ali Jahanian 0002, Jerry Liu, Qian Lin 0001, Daniel Tretter, Eamonn O'Brien-Strain, Seungyon Claire Lee, Nic Lyons, Jan P. Allebach |
IUI | 3 |
| 2013 | Cloud based multimedia analytic platformabstractMultimedia Analytic Platform is a cloud based service to expose state-of-the-art multimedia technologies for mobile and web application development. As a product-quality service platform, it offers comprehensive API documentation, code example, service description and sandbox for trial for each multimedia technology. The utilization of the cloud storage and distributed computing framework allows the service platform to run with robustness and efficiency. The current technologies supported by the platform include face detection, face verification, face demographic estimation, feature extraction, image matching, and image collage. Since its initial public launch in October 2012, it has been adopted by universities and third party companies for course support and application development. Rares Vernica, Qian Lin 0001 |
ACM Multimedia | 3 |
| 2013 | User Analytics with UbeOne: Insights into Web PrintingabstractAs web and mobile applications become more sensitive to the user context, there is a shift from purely off-line processing of user actions (log analysis) to real-time user analytics that can generate information about the user context to be instantly leveraged by the application. Ubeone is a system that enables both real-time and aggregate analytics from user data. The system is designed as a set of lightweight, composeable mechanisms that can progressively and collectively analyze a user action, such as pinning, saving or printing a web page. We will demonstrate the system capabilities on analyzing a live feed of URLs printed through a proprietary, web browser plug-in. This is in fact the first analysis of web printing activity. We will also give a taste of how the system can enable instant recommendations based on the user context. Georgia Koutrika, Qian Lin 0001, Jerry Liu |
Proc. VLDB Endow. | 2 |
| 2013 | Personal Clothing Retrieval on Photo Collections by Color and AttributesabstractAutomatic personal clothing retrieval on photo collections, i.e., searching the same clothes worn by the same person, is not a trivial problem as photos are usually taken under completely uncontrolled realistic imaging conditions. Typically, the captured clothing images have large variations due to geometric deformation, occlusion, cluttered background, and photometric variability from illumination and viewpoint, which pose significant challenges to text-based or reranking-based visual search methods. In this paper, a novel framework is presented to tackle these issues by leveraging low-level features (e.g., color) and high-level features (attributes) of clothing. First, a content-based image retrieval (CBIR) approach based on the bag-of-visual-words (BOW) model is developed as our baseline system, in which a codebook is constructed from extracted dominant color patches. A reranking approach is then proposed to improve search quality by exploiting clothing attributes, including the type of clothing, sleeves, patterns, etc. Compared to low-level features, the attributes have better robustness to clothing variations, and carry semantic meanings as high-level image representations. Different visual attribute detectors are learned from large amounts of training data to extract the corresponding attributes. The construction of codebook and building of attribute classifiers are conducted offline, which leads to fast online search performance. Extensive experiments on photo collections show that the reranking algorithm based on attribute learning significantly improves retrieval performance in combination with the proposed baseline. Even our color-based baseline alone outperforms the previous CBIR-based search approaches. The experiments also demonstrate that our approach is robust to large variations of images taken in unconstrained environment. Xianwang Wang, Tong Zhang 0007, Daniel Tretter, Qian Lin 0001 |
IEEE Trans. Multim. | 4 |
| 2012 | Face Swapping under Large Pose Variations: A 3D Model Based ApproachabstractTraditional face swapping technologies require the faces of source images and target images have similar pose and appearance (usually frontal). This limits its applications. This paper presents a method for face swapping based on personalized 3D head models. This framework builds a personalized 3D head model from a frontal face and can be rendered at any pose to match the characters in the image we want to swap. The 3D head model is constructed by a user uploaded frontal view face image. This construction process goes through face alignment and feature point matching. The final personalized 3D head is built by deforming a standard 3D head model using radial basis function to match the specific person. To make the synthesized face seamlessly blended into the image, color transfer and multi-resolution spline technique are applied. We use the proposed technique to create personalized storybook where the characters are replaced with a user's face and promising results are obtained. The system can be used in face de-identification as well. Shengjin Wang, Qian Lin 0001 |
ICME | 3 |
| 2012 | Face replacement with large-pose differencesabstractIn this paper, we present a novel face replacement system exchanging faces with large-pose differences. Traditional 2D image based face replacement can only replace faces with similar pose and appearance. This significantly limits the application of face replacement. In this paper, we propose to build a 3D head model from a single frontal face photo. The automatically constructed 3D head can be rendered under arbitrary poses and illuminations. This makes it possible to do swapping for faces with large pose variations. In the demo, the user captures a frontal face image using a capture device such as a webcam or a smartphone, and then the algorithm can automatically build the 3D model using feature detection, face alignment and reconstruction. This 3D model is used to swap to any other target face photo the user selects. While our system is automatic, we also provide interactive tools for the user to adjust the feature detection to enhance the results. Qian Lin 0001, Shengjin Wang |
ACM Multimedia | 2 |
| 2012 | TouchPaper: making print interactiveabstractTraditional printed materials such as photobooks and collages are static in that the information conveyed is unchanging and limited to what was printed. In this paper, we describe an online service for augmenting prints with rich media by using image recognition. A print can have multiple hotspots with each one linking to different content on the web (e.g. Facebook, Youtube video, animations, etc.). When a print is viewed through a mobile device, it is automatically recognized and the interactive hot regions are highlighted on the screen. The user can click regions of interest to link to relevant content on the web. This technology significantly enhances the potential for personalization and interaction and ultimately the end user experience and the value of the prints. As an example application, we leverage our Facebook auto-photobook application to automatically create interactive links for each photo in a page. The user can easily view additional content such as photo comments for each photo through the augmentation mechanism. Hao Tang 0001, Daniel Tretter, Qian Lin 0001 |
ACM Multimedia | 4 |
| 2011 | Document visual similarity measure for document searchabstractManaging large document databases has become an important task. Being able to automatically compare document layouts and classify and search documents with respect to their visual appearance proves to be desirable in many applications. We propose a new algorithm that approximates a metric function between documents based on their visual similarity. The comparison is based only on the visual appearance of the document without taking into consideration its text content. We measure the similarity of single page documents with respect to distance functions between three document components: background, text, and saliency. Each document component is represented as a Gaussian mixture distribution; and distances between the components of different documents are calculated as an approximation of the Hellinger distance between corresponding distributions. Since the Hellinger distance obeys the triangle inequality, it proves to be favorable in the task of nearest neighbor search in a document database. Thus, the computation required to find similar documents in a document database can be significantly reduced. Ildus Ahmadullin, Jan P. Allebach, Niranjan Damera-Venkata, Jian Fan, Seungyon Claire Lee, Qian Lin 0001, Jerry Liu, Eamonn O'Brien-Strain |
ACM Symposium on Document Engineering | 6 |
| 2010 | Automatic creation of face composite images for consumer applicationsabstractThis paper presents a method to perform automatic extraction, manipulation and blending of human face and hair images into other media such as licensed commercial content. We describe automatic image processing algorithms to perform such tasks, and discuss applications in personalized publishing and merchandising. We demonstrate our work through automatically segmenting out the face and hair image in a consumer photo and embedding it in another photo or document. The goal is to perform these steps seamlessly with natural look and minimal manual intervention. The primary application of this technology is the personalization of various documents and products such as licensed merchandise, children's storybook, puzzles and marketing collaterals. Suk Hwan Lim, Qian Lin 0001, Adam Petruszka |
ICASSP | 2 |
| 2010 | Mobile document scanning and copyingabstractIn this paper, we show a multimedia system for processing mobile camera captured documents. Using a client application on a mobile phone, a user can capture a document image, and send the image to a processing server so that the document image can be restored using automatic perspective and illumination corrections. The restored document can then be sent to a web-connected printer to complete the copying task. Jian Fan, Qian Lin 0001, Jerry Liu |
ACM Multimedia | 2 |
| 2009 | Mobile media search: has media search finally found its perfect platform? part IIabstractRecently, many exciting media search applications have been introduced to take advantage of smart phones' audiovisual capture capabilities and their being always on and connected. These applications address a real pain point for most mobile users and allow them to search with minimal text entry, if any. Is the mobile platform an ideal fit for media search? Are audio and visual signal processing technologies sufficiently accurate to support most mobile search applications? What are the killer applications of mobile media search? Earlier in 2009 at ICASSP, a panel on this topic stirred up great interest and enthusiasm while leaving many questions untouched due to the limited time. Berna Erol, Jiebo Luo 0001, Shih-Fu Chang, Minoru Etoh, Hsiao-Wuen Hon, Qian Lin 0001, Vidya Setlur |
ACM Multimedia | 6 |
| 1999 | A Web-Based Secure System for the Distributed Printing of Documents and Images
Ping Wah Wong, Daniel Tretter, Thomas Kite, Qian Lin 0001, Hugh Nguyen |
J. Vis. Commun. Image Represent. | 4 |
| 1998 | Automatic Digital Redeye Reduction
Andrew J. Patti, Konstantinos Konstantinides, Daniel Tretter, Qian Lin 0001 |
ICIP (3) | 4 |
| 1998 | A Web-based Secure System for the Distributed Printing of Documents and Images
Ping Wah Wong, Daniel Tretter, Thomas Kite, Qian Lin 0001, Hugh Nguyen |
ICIP (3) | 4 |
| 1996 | FM screen design using DBS algorithmabstractWe describe an algorithm to design a frequency modulated screen using the direct binary search algorithm. Compared with the direct binary search algorithm itself, we show that we can maintain halftone image quality while significantly reducing the required computation. Jan P. Allebach, Qian Lin 0001 |
ICIP (1) | 2 |
| 1995 | Screen design for printingabstractTo display or print a continuous tone image on a bi-level device, the image goes through a halftoning process. For best rendition, the halftoning algorithm should adapt to the output device. We review printer dot models, and the utilization of printer dot model in generating halftone screens, especially frequency-modulated (FM) screens. We also show the possibility of accurate control over tone reproduction using density measurement data without quantization errors. Qian Lin 0001 |
ICIP | 1 |
| 1992 | New approaches in interferometric SAR data processingabstractIt is known that interferometric synthetic-aperture radar (SAR) images can be inverted to perform surface elevation mapping. Among the factors critical to the mapping accuracy are registration of the interfering SAR images and phase unwrapping. A registration algorithm is presented that determines the registration parameters through optimization. A figure of merit is proposed that evaluates the registration result during the optimization. The phase unwrapping problem is approached through a new method involving fringe line detection. The algorithms are tested with two SEASAT SAR images of terrain near Yellowstone National Park. These images were collected on SEASAT orbits 1334 and 1420, which were very close together in space, i.e. less than 100 m.> Qian Lin 0001, John F. Vesecky, Howard A. Zebker |
IEEE Trans. Geosci. Remote. Sens. | 1 |