Francine Chen 0001

dblp:95/4481 · also Francine R. Chen · DBLP profile ↗
← Back
58ranked-venue papers
18as first author
4since 2021 · last 2025
0000-0002-0723-5609ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 9 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 17 · 4 first-authorHuman-computer interaction and ubiquitous computing · 6 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 50% Information extraction and text analysis · 26% Transfer learning and domain adaptation · 11%
Computer graphics and multimedia
5 papers
Visualization and visual analytics · 79% Multimedia analysis and retrieval · 17% Audio and music processing · 4%
Databases, data mining, and information retrieval
6 papers
Data mining · 55% Web and social media mining · 28% Information retrieval · 16%
Human-computer interaction and pervasive computing
4 papers
Learning and educational technologies · 42% Interaction techniques and input · 40% Personal fabrication and tangible interfaces · 12%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics › graph visualization
bipartite graph visualization
0.612022
Understanding Missing Links in Bipartite Networks With MissBiN · IEEE Trans. Vis. Comput. Graph. 2022
Visualization and visual analytics › visual analytics
visual analytics for machine learning
0.612022
Understanding Missing Links in Bipartite Networks With MissBiN · IEEE Trans. Vis. Comput. Graph. 2022
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.412019
Adversarial Domain Adaptation Using Artificial Titles for Abstractive Title Generation · ACL (1) 2019
Machine learning › Transfer learning and domain adaptation › domain adaptation › distribution adaptation
adversarial domain adaptation
0.412019
Adversarial Domain Adaptation Using Artificial Titles for Abstractive Title Generation · ACL (1) 2019
Natural language and speech › Language models and text generation › text summarization
title generation
0.412019
Adversarial Domain Adaptation Using Artificial Titles for Abstractive Title Generation · ACL (1) 2019
Learning and educational technologies › student knowledge modeling
knowledge tracing
0.412019
Augmenting Knowledge Tracing by Considering Forgetting Behavior · WWW 2019
Natural language and speech › Language models and text generation › text summarization
extractive summarization
0.312018
Harnessing Popularity in Social Media for Extractive Summarization of Online Conversations · EMNLP 2018
Natural language and speech › Language models and text generation
text summarization
0.312018
Harnessing Popularity in Social Media for Extractive Summarization of Online Conversations · EMNLP 2018
Visualization and visual analytics
high-dimensional data visualization
0.312018
BiDots: Visual Exploration of Weighted Biclusters · IEEE Trans. Vis. Comput. Graph. 2018
Visualization and visual analytics
interactive data exploration
0.312018
BiDots: Visual Exploration of Weighted Biclusters · IEEE Trans. Vis. Comput. Graph. 2018
Natural language and speech › Language models and text generation
large language model
0.312025
Empathy Prediction from Diverse Perspectives · ACL (1) 2025
Data mining
business intelligence
0.212016
Tweetviz: Visualizing Tweets for Business Intelligence · SIGIR 2016
Web and social media mining
social media analysis
0.212016
Tweetviz: Visualizing Tweets for Business Intelligence · SIGIR 2016
Data mining › pattern mining › matrix pattern mining
biclique mining
0.212022
Understanding Missing Links in Bipartite Networks With MissBiN · IEEE Trans. Vis. Comput. Graph. 2022
Data mining
pattern mining
0.212022
Understanding Missing Links in Bipartite Networks With MissBiN · IEEE Trans. Vis. Comput. Graph. 2022
Interaction techniques and input
gesture input
0.112011
Multi-touch document folding: gesture models, fold directions and symmetries · CHI 2011
Interaction techniques and input › touch interaction
multi-touch interaction
0.112011
Multi-touch document folding: gesture models, fold directions and symmetries · CHI 2011
Machine learning › Deep learning architectures and training
sequence modeling
0.112019
Augmenting Knowledge Tracing by Considering Forgetting Behavior · WWW 2019
Personal fabrication and tangible interfaces
paper-based interaction
0.112010
FACT: fine-grained cross-media interaction with documents via a portable hybrid paper-laptop interface · ACM Multimedia 2010
Interaction techniques and input › input modality › multimodal input
pen and gesture input
0.112010
FACT: fine-grained cross-media interaction with documents via a portable hybrid paper-laptop interface · ACM Multimedia 2010
Bioinformatics and computational biology
gene expression analysis
0.112018
BiDots: Visual Exploration of Weighted Biclusters · IEEE Trans. Vis. Comput. Graph. 2018
Computer vision › Face, body and person analysis
human pose estimation
0.112008
Context and observation driven latent variable model for human pose estimation · CVPR 2008
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.112008
Context and observation driven latent variable model for human pose estimation · CVPR 2008
Information retrieval › document retrieval › temporal information retrieval
new event detection
0.122003
A System for new event detection · SIGIR 2003
Optimizing Story Link Detection is not Equivalent to Optimizing New Event Detection · ACL 2003
Information retrieval › text analysis › topic analysis
topic detection and tracking
0.122003
A System for new event detection · SIGIR 2003
Optimizing Story Link Detection is not Equivalent to Optimizing New Event Detection · ACL 2003
Audio and music processing › speech processing
speech privacy
0.112008
Audio privacy: reducing speech intelligibility while preserving environmental sounds · ACM Multimedia 2008
Data mining › text mining
sentiment analysis
0.112016
Tweetviz: Visualizing Tweets for Business Intelligence · SIGIR 2016
Computer vision › Video understanding and tracking › multi-object tracking
multi-person tracking
0.112007
DOTS: support for effective video surveillance · ACM Multimedia 2007
Computer vision › Video understanding and tracking › object tracking
occlusion handling
0.112007
DOTS: support for effective video surveillance · ACM Multimedia 2007
Multimedia analysis and retrieval › video surveillance
multi-camera tracking
0.112007
DOTS: support for effective video surveillance · ACM Multimedia 2007

Methods — techniques the papers use, named apart from their topics

quantitative evaluation · 1.1link prediction · 1.1biclique analysis · 1.1perspective-based modeling · 0.9deep knowledge tracing · 0.8visual encoding · 0.7interaction design · 0.7distant labeling · 0.7disjunctive model · 0.7biclustering · 0.7sequential training · 0.4adversarial domain adaptation · 0.4visualization · 0.2sentiment analysis · 0.2multimodal feature representation · 0.2convolutional neural network · 0.2vocal tract transfer function replacement · 0.2audio resynthesis · 0.2
YearPublicationVenuePosition
2025 Empathy Prediction from Diverse Perspectives
abstract
Francine Chen, Scott Carter, Tatiana Lau, Nayeli Suseth Bravo, Sumanta Bhattacharyya, Kate Sieck, Charlene C. Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Francine Chen 0001, Scott A. Carter, Tatiana Lau, Nayeli Bravo, Sumanta Bhattacharyya, Katharine Sieck, Charlene C. Wu
ACL (1)1
2023 Can Behavioral Experts Predict Outcome Heterogeneity?
Rumen Iliev, Alex Filipowicz, Emily S. Sumner, Francine Chen 0001, Nikos Aréchiga, Scott A. Carter, Totte Harinen, Katharine Sieck, Charlene C. Wu
CogSci4
2022 Understanding Missing Links in Bipartite Networks With MissBiN
abstract
The analysis of bipartite networks is critical in a variety of application domains, such as exploring entity co-occurrences in intelligence analysis and investigating gene expression in bio-informatics. One important task is missing link prediction, which infers the existence of unseen links based on currently observed ones. In this article, we propose a visual analysis system, MissBiN, to involve analysts in the loop for making sense of link prediction results. MissBiN equips a novel method for link prediction in a bipartite network by leveraging the information of bi-cliques in the network. It also provides an interactive visualization for understanding the algorithm outputs. The design of MissBiN is based on three high-level analysis questions (what, why, and how) regarding missing links, which are distilled from the literature and expert interviews. We conducted quantitative experiments to assess the performance of the proposed link prediction algorithm, and interviewed two experts from different domains to demonstrate the effectiveness of MissBiN as a whole. We also provide a comprehensive usage scenario to illustrate the usefulness of the tool in an application of intelligence analysis.
Jian Zhao 0010, Maoyuan Sun, Francine Chen 0001, Patrick Chiu
IEEE Trans. Vis. Comput. Graph.3
2021 Know-What and Know-Who: Document Searching and Exploration using Topic-Based Two-Mode Networks
abstract
This paper proposes a novel approach for analyzing search results of a document collection, which can help support know-what and know-who information seeking questions. Search results are grouped by topics, and each topic is represented by a two-mode network composed of related documents and authors (i.e., biclusters). We visualize these biclusters in a 2D layout to support interactive visual exploration of the analyzed search results, which highlights a novel way of organizing entities of biclusters. We evaluated our approach using a large academic publication corpus, by testing the distribution of the relevant documents and of lead and prolific authors. The results indicate the effectiveness of our approach compared to traditional 1D ranked lists. Moreover, a user study with 12 participants was conducted to compare our proposed visualization, a simplified variation without topics, and a text-based interface. We report on participants' task performance, their preference of the three interfaces, and the different strategies used in information seeking.
Jian Zhao 0010, Maoyuan Sun, Patrick Chiu, Francine Chen 0001, Bee Liew
PacificVis4
2020 Thoracic Disease Identification and Localization using Distance Learning and Region Verification
Cheng Zhang 0014, Francine Chen 0001, Yan-Ying Chen
BMVC2
2020 Tackling challenges of neural purchase stage identification from imbalanced twitter data
abstract
Abstract Twitter and other social media platforms are often used for sharing interest in products. The identification of purchase decision stages, such as in the AIDA model (Awareness, Interest, Desire, and Action), can enable more personalized e-commerce services and a finer-grained targeting of advertisements than predicting purchase intent only. In this paper, we propose and analyze neural models for identifying the purchase stage of single tweets in a user’s tweet sequence. In particular, we identify three challenges of purchase stage identification: imbalanced label distribution with a high number of non-purchase-stage instances, limited amount of training data, and domain adaptation with no or only little target domain data. Our experiments reveal that the imbalanced label distribution is the main challenge for our models. We address it with ranking loss and perform detailed investigations of the performance of our models on the different output classes. In order to improve the generalization of the models and augment the limited amount of training data, we examine the use of sentiment analysis as a complementary, secondary task in a multitask framework. For applying our models to tweets from another product domain, we consider two scenarios: for the first scenario without any labeled data in the target product domain, we show that learning domain-invariant representations with adversarial training is most promising, while for the second scenario with a small number of labeled target examples, fine-tuning the source model weights performs best. Finally, we conduct several analyses, including extracting attention weights and representative phrases for the different purchase stages. The results suggest that the model is learning features indicative of purchase stages and that the confusion errors are sensible.
Heike Adel, Francine Chen 0001, Yan-Ying Chen
Nat. Lang. Eng.2
2019 Adversarial Domain Adaptation Using Artificial Titles for Abstractive Title Generation
abstract
A common issue in training a deep learning, abstractive summarization model is lack of a large set of training summaries.This paper examines techniques for adapting from a labeled source domain to an unlabeled target domain in the context of an encoder-decoder model for text generation.In addition to adversarial domain adaptation (ADA), we introduce the use of artificial titles and sequential training to capture the grammatical style of the unlabeled target domain.Evaluation on adapting to/from news articles and Stack Exchange posts indicates that the use of these techniques can boost performance for both unsupervised adaptation as well as fine-tuning with limited target data.
Francine Chen 0001, Yan-Ying Chen
ACL (1)1
2019 Addressing Data Bias Problems for Chest X-ray Image Report Generation
Philipp Harzig, Yan-Ying Chen, Francine Chen 0001, Rainer Lienhart
BMVC3
2019 Augmenting Knowledge Tracing by Considering Forgetting Behavior
abstract
Computer-aided education systems are now seeking to provide each student with personalized materials based on a student's individual knowledge. To provide suitable learning materials, tracing each student's knowledge over a period of time is important. However, predicting each student's knowledge is difficult because students tend to forget. The forgetting behavior is mainly because of two reasons: the lag time from the previous interaction, and the number of past trials on a question. Although there are a few studies that consider forgetting while modeling a student's knowledge, some models consider only partial information about forgetting, whereas others consider multiple features about forgetting, ignoring a student's learning sequence. In this paper, we focus on modeling and predicting a student's knowledge by considering their forgetting behavior. We extend the deep knowledge tracing model [17], which is a state-of-the-art sequential model for knowledge tracing, to consider forgetting by incorporating multiple types of information related to forgetting. Experiments on knowledge tracing datasets show that our proposed model improves the predictive performance as compared to baselines. Moreover, we also examine that the combination of multiple types of information that affect the behavior of forgetting results in performance improvement.
Koki Nagatani, Qian Zhang 0061, Masahiro Sato, Yan-Ying Chen, Francine Chen 0001, Tomoko Ohkuma
WWW5
2018 Harnessing Popularity in Social Media for Extractive Summarization of Online Conversations
abstract
We leverage a popularity measure in social media as a distant label for extractive summarization of online conversations.In social media, users can vote, share, or bookmark a post they prefer.The number of these actions is regarded as a measure of popularity.However, popularity is not determined solely by content of a post, e.g., a text or an image it contains, but is highly based on its contexts, e.g., timing, and authority.We propose Disjunctive model that computes the contribution of content and context separately.For evaluation, we build a dataset where the informativeness of comments is annotated.We evaluate the results with ranking metrics, and show that our model outperforms the baseline models which directly use popularity as a measure of informativeness.
Ryuji Kano, Yasuhide Miura, Motoki Taniguchi, Yan-Ying Chen, Francine Chen 0001, Tomoko Ohkuma
EMNLP5
2018 Learning to Disentangle Interleaved Conversational Threads with a Siamese Hierarchical Network and Similarity Ranking
abstract
Jyun-Yu Jiang, Francine Chen, Yan-Ying Chen, Wei Wang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Jyun-Yu Jiang, Francine Chen 0001, Yan-Ying Chen, Wei Wang 0010
NAACL-HLT2
2018 BiDots: Visual Exploration of Weighted Biclusters
abstract
Discovering and analyzing biclusters, i.e., two sets of related entities with close relationships, is a critical task in many real-world applications, such as exploring entity co-occurrences in intelligence analysis, and studying gene expression in bio-informatics. While the output of biclustering techniques can offer some initial low-level insights, visual approaches are required on top of that due to the algorithmic output complexity. This paper proposes a visualization technique, called BiDots, that allows analysts to interactively explore biclusters over multiple domains. BiDots overcomes several limitations of existing bicluster visualizations by encoding biclusters in a more compact and cluster-driven manner. A set of handy interactions is incorporated to support flexible analysis of biclustering results. More importantly, BiDots addresses the cases of weighted biclusters, which has been underexploited in the literature. The design of BiDots is grounded by a set of analytical tasks derived from previous work. We demonstrate its usefulness and effectiveness for exploring computed biclusters with an investigative document analysis task, in which suspicious people and activities are identified from a text corpus.
Jian Zhao 0010, Maoyuan Sun, Francine Chen 0001, Patrick Chiu
IEEE Trans. Vis. Comput. Graph.3
2017 Video to Text Summary: Joint Video Summarization and Captioning with Recurrent Neural Networks
Bor-Chun Chen, Yan-Ying Chen, Francine Chen 0001
BMVC3
2017 Image-based user profiling of frequent and regular venue categories
abstract
The availability of mobile access has shifted social media use. With that phenomenon, what users shared on social media and where they visited is naturally an excellent resource to learn their visiting behavior. Knowing visit behaviors would help market survey and customer relationship management, e.g., sending customers coupons of the businesses that they visit frequently. Most prior studies leverage meta-data e.g., check-in locations to profile visiting behavior but neglect important information from user-contributed content, e.g., images. This work addresses a novel use of image content for predicting the user visit behavior, i.e., the frequent and regular business venue categories that the content owner would visit. To collect training data, we propose a strategy to use geo-metadata associated with images for deriving the labels of an image owner's visit behavior. Moreover, we model a user's sequential images by using an end-to-end learning framework to reduce the optimization loss. That helps improve the prediction accuracy against the baseline as demonstrated in our experiments. The prediction is completely based on image content that is more available in social media than geo-metadata, and thus allows coverage in profiling a wider set of users.
Ryosuke Shigenaka, Yan-Ying Chen, Francine Chen 0001, Dhiraj Joshi, Yukihiro Tsuboshita
ICME3
2016 Business-Aware Visual Concept Discovery from Social Media for Multimodal Business Venue Recognition
abstract
Image localization is important for marketing and recommendation of local business; however, the level of granularity is still a critical issue. Given a consumer photo and its rough GPS information, we are interested in extracting the fine-grained location information, i.e. business venues, of the image. To this end, we propose a novel framework for business venue recognition. The framework mainly contains three parts. First, business-aware visual concept discovery: we mine a set of concepts that are useful for business venue recognition based on three guidelines including business awareness, visually detectable, and discriminative power. We define concepts that satisfy all of these three criteria as business-aware visual concept. Second, business-aware concept detection by convolutional neural networks (BA-CNN): we propose a new network configuration that can incorporate semantic signals mined from business reviews for extracting semantic concept features from a query image. Third, multimodal business venue recognition: we extend visually detected concepts to multimodal feature representations that allow a test image to be associated with business reviews and images from social media for business venue recognition. The experiments results show the visual concepts detected by BA-CNN can achieve up to 22.5% relative improvement for business venue recognition compared to the state-of-the-art convolutional neural network features. Experiments also show that by leveraging multimodal information from social media we can further boost the performance, especially when the database images belonging to each business venue are scarce.
Bor-Chun Chen, Yan-Ying Chen, Francine Chen 0001, Dhiraj Joshi
AAAI3
2016 Using business-aware latent topics for image captioning in social media
abstract
Captions are a central component in image posts that communicate the background story behind photos. Captions can enhance the engagement with audiences and are therefore critical to campaigns or advertisement. Previous studies in image captioning either rely solely on image content or summarize multiple web documents related to image's location; both neglect users' activities. We propose business-aware latent topics as a new contextual cue for image captioning that represent user activities at business venues. The idea is to learn the typical activities of people who posted images from business venues with similar categories (e.g., fast food restaurants) to provide appropriate context for similar topics (e.g., burger) in new posts. User activities at businesses are modeled via a latent topic representation. In turn, the image captioning model can generate sentences that better reflect user activities at business venues. In our experiments, the business-aware latent topics are effective for adapting to captions to images captured in various businesses than the existing baselines. Moreover, they complement other contextual cues (image, time) in a multi-modal framework.
Francine Chen 0001, Matthew Cooper 0002, Dhiraj Joshi
ICME2
2016 Topic Modeling of Document Metadata for Visualizing Collaborations over Time
abstract
We describe methods for analyzing and visualizing document metadata to provide insights about collaborations over time. We investigate the use of Latent Dirichlet Allocation (LDA) based topic modeling to compute areas of interest on which people collaborate. The topics are represented in a node-link force directed graph by persistent fixed nodes laid out with multidimensional scaling (MDS), and the people by transient movable nodes. The topics are also analyzed to detect bursts to highlight "hot" topics during a time interval. As the user manipulates a time interval slider, the people nodes and links are dynamically updated. We evaluate the results of LDA topic modeling for the visualization by comparing topic keywords against the submitted keywords from the InfoVis 2004 Contest, and we found that the additional terms provided by LDA-based keyword sets result in improved similarity between a topic keyword set and the documents in a corpus. We extended the InfoVis dataset from 8 to 20 years and collected publication metadata from our lab over a period of 21 years, and created interactive visualizations for exploring these larger datasets.
Francine Chen 0001, Patrick Chiu, Seongtaek Lim
IUI1
2016 Corpus for Customer Purchase Behavior Prediction in Social Media
Shigeyuki Sakaki, Francine Chen 0001, Mandy Korpusik, Yan-Ying Chen
LREC2
2016 Tweetviz: Visualizing Tweets for Business Intelligence
abstract
Social media offers potential opportunities for businesses to extract business intelligence. This paper presents Tweetviz, an interactive tool to help businesses extract actionable information from a large set of noisy Twitter messages. Tweetviz visualizes the tweet sentiment of business locations, identifies other business venues that Twitter users visit, and estimates some simple demographics of the Twitter users frequenting a business. A user study to evaluate the system's ability indicates that Tweetviz can provide an overview of a business's issues and sentiment as well as information aiding users in creating customer profiles.
Bas Sijtsma, Pernilla Qvarfordt, Francine Chen 0001
SIGIR3
2015 Inferring crowd-sourced venues for tweets
abstract
Knowing the geo-located venue of a tweet can facilitate better understanding of a user's geographic context, allowing apps to more precisely present information, recommend services, and target advertisements. However, due to privacy concerns, few users choose to enable geotagging of their tweets, resulting in a small percentage of tweets being geotagged; furthermore, even if the geo-coordinates are available, the closest venue to the geolocation may be incorrect. In this paper, we present a method for providing a ranked list of geo-located venues for a non-geotagged tweet, which simultaneously indicates the venue name and the geo-location at a very fine-grained granularity. In our proposed method for Venue Inference for Tweets (VIT), we construct a heterogeneous social network in order to analyze the embedded social relations, and leverage available but limited geographic data to estimate the geo-located venue of tweets. A single classifier is trained to estimate the probability of a tweet and a geo-located venue being linked, rather than training a separate model for each venue. We examine the performance of four types of social relation features and three types of geographic features embedded in a social network when inferring whether a tweet and a venue are linked, with a best accuracy of over 88%. We use the classifier probability estimates to rank the candidate geo-located venues of a non-geotagged tweet from over 19k possibilities, and observed an average top-5 accuracy of 29%.
Bokai Cao, Francine Chen 0001, Dhiraj Joshi, Philip S. Yu
IEEE BigData2
2014 Do Topic-Dependent Models Improve Microblog Sentiment Estimation?
Francine Chen 0001, Seyed Hamid Mirisaee
ICWSM1
2014 Scalable Image Search with Multiple Index Tables
abstract
Motivated by scalable partial-duplicate visual search, there has been growing interest in a wealth of compact and efficient binary feature descriptors (e.g. ORB, FREAK, BRISK). Typically, binary descriptors are clustered into codewords and quantized with Hamming distance, following the conventional bag-of-words strategy. However, such codewords formulated in Hamming space do not present obvious indexing and search performance improvement as compared to the Euclidean codewords. In this paper, without explicit codeword construction, we explore the use of partial binary descriptors as direct codebook indices (addresses). We propose a novel approach to build multiple index tables which concurrently check for collision of the same hash values. The evaluation is performed on two public image datasets: DupImage and Holidays. The experimental results demonstrate the indexing efficiency and retrieval accuracy of our approach.
Qiong Liu 0003, Francine Chen 0001, Dhiraj Joshi, Qi Tian 0001
ICMR3
2013 SmartDCap: semi-automatic capture of higher quality document images from a smartphone
abstract
People frequently capture photos with their smartphones, and some are starting to capture images of documents. However, the quality of captured document images is often lower than expected, even when an application that performs post-processing to improve the image is used. To improve the quality of captured images before post-processing, we developed the Smart Document Capture (SmartDCap) application that provides real-time feedback to users about the likely quality of a captured image. The quality measures capture the sharpness and framing of a page or regions on a page, such as a set of one or more columns, a part of a column, a figure, or a table. Using our approach, while users adjust the camera position, the application automatically determines when to take a picture of a document to produce a good quality result. We performed a subjective evaluation comparing SmartDCap and the Android Ice Cream Sandwich (ICS) camera application; we also used raters to evaluate the quality of the captured images. Our results indicate that users find SmartDCap to be as easy to use as the standard ICS camera application. Also, images captured using SmartDCap are sharper and better framed on average than images using the ICS camera application.
Francine Chen 0001, Scott A. Carter, Laurent Denoue, Jayant Kumar
IUI1
2012 Sharpness estimation for document and scene images
Jayant Kumar, Francine Chen 0001, David S. Doermann
ICPR2
2012 Genre identification for office document search and browsing
Francine Chen 0001, Andreas Girgensohn, Matthew Cooper 0002, Yijuan Lu, Gerry Filby
Int. J. Document Anal. Recognit.1
2011 Multi-touch document folding: gesture models, fold directions and symmetries
abstract
For document visualization, folding techniques provide a focus-plus-context approach with fairly high legibility on flat sections. To enable richer interaction, we explore the design space of multi-touch document folding. We discuss several design considerations for simple modeless gesturing and compatibility with standard Drag and Pinch gestures. We categorize gesture models along the characteristics of Symmetric/Asymmetric and Serial/Parallel, which yields three gesture models. We built a prototype document workspace application that integrates folding and standard gestures, and a system for testing the gesture models. A user study was conducted to compare the three models and to analyze the factors of fold direction, target symmetry, and target tolerance in user performance when folding a document to a specific shape. Our results indicate that all three factors were significant for task times, and parallelism was greater for symmetric targets.
Patrick Chiu, Chunyuan Liao, Francine Chen 0001
CHI3
2011 DiG: a task-based approach to product search
abstract
While there are many commercial systems to help people browse and compare products, these interfaces are typically product centric. To help users identify products that match their needs more efficiently, we instead focus on building a task centric interface and system. Based on answers to initial questions about the situations in which they expect to use the product, the interface identifies products that match their needs, and exposes high-level product features related to their tasks, as well as low-level information including customer reviews and product specifications. We developed semi-automatic methods to extract the high-level features used by the system from online product data. These methods identify and group product features, mine and summarize opinions about those features, and identify product uses. User studies verified our focus on high-level features for browsing products and low-level features and specifications for comparing products.
Scott A. Carter, Francine Chen 0001, Aditi S. Muralidharan, Jeremy Pickens
IUI2
2010 Picture detection in document page images
abstract
We present a method for picture detection in document page images, which can come from scanned or camera images, or rendered from electronic file formats. Our method uses OCR to separate out the text and applies the Normalized Cuts algorithm to cluster the non-text pixels into picture regions. A refinement step uses the captions found in the OCR text to deduce how many pictures are in a picture region, thereby correcting for under- and over-segmentation. A performance evaluation scheme is applied which takes into account the detection quality and fragmentation quality. We benchmark our method against the ABBYY application on page images from conference papers.
Patrick Chiu, Francine Chen 0001, Laurent Denoue
ACM Symposium on Document Engineering2
2010 FormCracker: interactive web-based form filling
abstract
Filling out document forms distributed by email or hosted on the Web is still problematic and usually requires a printer and scanner. Users commonly download and print forms, fill them out by hand, scan and email them. Even if the document is form-enabled (PDFs with FDF information), to read the file users still have to launch a separate application which may not be available, especially on mobile devices.
Laurent Denoue, John Adcock, Scott A. Carter, Patrick Chiu, Francine Chen 0001
ACM Symposium on Document Engineering5
2010 DocuBrowse: faceted searching, browsing, and recommendations in an enterprise context
abstract
Browsing and searching for documents in large, online enterprise document repositories are common activities. While internet search produces satisfying results for most user queries, enterprise search has not been as successful because of differences in document types and user requirements. To support users in finding the information they need in their online enterprise repository, we created DocuBrowse, a faceted document browsing and search system. Search results are presented within the user-created document hierarchy, showing only directories and documents matching selected facets and containing text query terms. In addition to file properties such as date and file size, automatically detected document types, or genres, serve as one of the search facets. Highlighting draws the user's attention to the most promising directories and documents while thumbnail images and automatically identified keyphrases help select appropriate documents. DocuBrowse utilizes document similarities, browsing histories, and recommender system techniques to suggest additional promising documents for the current facet and content filters.
Andreas Girgensohn, Frank M. Shipman III, Francine Chen 0001, Lynn Wilcox
IUI3
2010 FACT: fine-grained cross-media interaction with documents via a portable hybrid paper-laptop interface
abstract
FACT is an interactive paper system for fine-grained interaction with documents across the boundary between paper and computers. It consists of a small camera-projector unit, a laptop, and ordinary paper documents. With the camera-projector unit pointing to a paper document, the system allows a user to issue pen gestures on the paper document for selecting fine-grained content and applying various digital functions. For example, the user can choose individual words, symbols, figures, and arbitrary regions for keyword search, copy and paste, web search, and remote sharing. FACT thus enables a computer-like user experience on paper. This paper interaction can be integrated with laptop interaction for cross-media manipulations on multiple documents and views. We present the infrastructure, supporting techniques and interaction design, and demonstrate the feasibility via a quantitative experiment. We also propose applications such as document manipulation, map navigation and remote collaboration.
Chunyuan Liao, Qiong Liu 0003, Patrick Chiu, Francine Chen 0001
ACM Multimedia5
2008 Context and observation driven latent variable model for human pose estimation
abstract
Current approaches to pose estimation and tracking can be classified into two categories: generative and discriminative. While generative approaches can accurately determine human pose from image observations, they are computationally expensive due to search in the high dimensional human pose space. On the other hand, discriminative approaches do not generalize well, but are computationally efficient. We present a hybrid model that combines the strengths of the two in an integrated learning and inference framework. We extend the Gaussian process latent variable model (GPLVM) to include an embedding from observation space (the space of image features) to the latent space. GPLVM is a generative model, but the inclusion of this mapping provides a discriminative component, making the model observation driven. Observation Driven GPLVM (OD-GPLVM) not only provides a faster inference approach, but also more accurate estimates (compared to GPLVM) in cases where dynamics are not sufficient for the initialization of search in the latent space. We also extend OD-GPLVM to learn and estimate poses from parameterized actions/gestures. Parameterized gestures are actions which exhibit large systematic variation in joint angle space for different instances due to difference in contextual variables. For example, the joint angles in a forehand tennis shot are function of the height of the ball (Figure 2). We learn these systematic variations as a function of the contextual variables. We then present an approach to use information from scene/objects to provide context for human pose estimation for such parameterized actions.
Abhinav Gupta 0001, Trista Pei-Chun Chen, Francine Chen 0001, Don Kimber, Larry Davis 0001
CVPR3
2008 Audio privacy: reducing speech intelligibility while preserving environmental sounds
abstract
Audio monitoring has many applications but also raises privacy concerns. In an attempt to help alleviate these concerns, we have developed a method for reducing the intelligibility of speech while preserving intonation and the ability to recognize most environmental sounds. The method is based on identifying vocalic regions and replacing the vocal tract transfer function of these regions with the transfer function from prerecorded vowels, where the identity of the replacement vowel is independent of the identity of the spoken syllable. The audio signal is then re-synthesized using the original pitch and energy, but with the modified vocal tract transfer function. We performed an intelligibility study which showed that environmental sounds remained recognizable but speech intelligibility can be dramatically reduced to a 7% word recognition rate.
Francine Chen 0001, John Adcock, Shruti Krishnagiri
ACM Multimedia1
2007 Robust People Detection and Tracking in a Multi-Camera Indoor Visual Surveillance System
abstract
In this paper we describe the analysis component of an indoor, real-time, multi-camera surveillance system. The analysis includes: (1) a novel feature-level foreground segmentation method which achieves efficient and reliable segmentation results even under complex conditions, (2) an efficient greedy search based approach for tracking multiple people through occlusion, and (3) a method for multi-camera handoff that associates individual trajectories in adjacent cameras. The analysis is used for an 18 camera surveillance system that has been running continuously in an indoor business over the past several months. Our experiments demonstrate that the processing method for people detection and tracking across multiple cameras is fast and robust.
Tao Yang 0006, Francine Chen 0001, Don Kimber, Jim Vaughan
ICME2
2007 DOTS: support for effective video surveillance
abstract
DOTS (Dynamic Object Tracking System) is an indoor, real-time, multi-camera surveillance system, deployed in a real office setting. DOTS combines video analysis and user interface components to enable security personnel to effectively monitor views of interest and to perform tasks such as tracking a person. The video analysis component performs feature-level foreground segmentation with reliable results even under complex conditions. It incorporates an efficient greedy-search approach for tracking multiple people through occlusion and combines results from individual cameras into multi-camera trajectories. The user interface draws the users. attention to important events that are indexed for easy reference at a later time. Different views within the user interface provide spatial information for easier navigation. Our system, with over twenty video cameras installed in hallways and other public spaces in our office building, has been in constant use for almost a year.
Andreas Girgensohn, Don Kimber, Jim Vaughan, Tao Yang 0006, Frank M. Shipman III, Thea Turner, Eleanor Gilbert Rieffel, Lynn Wilcox, Francine Chen 0001, Anthony Dunnigan
ACM Multimedia9
2006 Improving Probabilistic Latent Semantic Analysis with Principal Component Analysis
Ayman Farahat, Francine Chen 0001
EACL2
2004 Multiple Similarity Measures and Source-Pair Information in Story Link Detection
Francine Chen 0001, Ayman Farahat, Thorsten Brants
HLT-NAACL1
2003 Optimizing Story Link Detection is not Equivalent to Optimizing New Event Detection
abstract
Link detection has been regarded as a core technology for the Topic Detection and Tracking tasks of new event detection. In this paper we formulate story link detection and new event detection as information retrieval task and hypothesize on the impact of precision and recall on both systems. Motivated by these arguments, we introduce a number of new performance enhancing techniques including part of speech tagging, new similarity measures and expanded stop lists. Experimental results validate our hypothesis.
Ayman Farahat, Francine Chen 0001, Thorsten Brants
ACL2
2003 Story Link Detection and New Event Detection are Asymmetric
Francine Chen 0001, Ayman Farahat, Thorsten Brants
HLT-NAACL1
2003 A System for new event detection
abstract
We present a new method and system for performing the New Event Detection task, i.e., in one or multiple streams of news stories, all stories on a previously unseen (new) event are marked. The method is based on an incremental TF-IDF model. Our extensions include: generation of source-specific models, similarity score normalization based on document-specific averages, similarity score normalization based on source-pair specific averages, term reweighting based on inverse event frequencies, and segmentation of the documents. We also report on extensions that did not improve results. The system performs very well on TDT3 and TDT4 test data and scored second in the TDT-2002 evaluation.
Thorsten Brants, Francine Chen 0001
SIGIR2
2002 Topic-based document segmentation with probabilistic latent semantic analysis
abstract
This paper presents a new method for topic-based document segmentation, i.e., the identification of boundaries between parts of a document that bear on different topics. The method combines the use of the Probabilistic Latent Semantic Analysis (PLSA) model with the method of selecting segmentation points based on the similarity values between pairs of adjacent blocks. The use of PLSA allows for a better representation of sparse information in a text block, such as a sentence or a sequence of sentences. Furthermore, segmentation performance is improved by combining different instantiations of the same model, either using different random initializations or different numbers of latent classes. Results on commonly available data sets are significantly better than those of other state-of-the-art systems.
Thorsten Brants, Francine Chen 0001, Ioannis Tsochantaridis
CIKM2
2002 AuGEAS: authoritativeness grading, estimation, and sorting
abstract
When searching for content in in a large heterogeneous document collections like the World Wide Web it is not easy to know which documents provide reliable authoritative information about a subject. The problem is particularly pointed as it concerns content search for "high-value" informational needs such as retrieving medical information, where the cost of error may be high. In this paper, a method is described for estimating the authoritativeness of a document based on textual, non-topical cues. This method is complementary to estimates of authoritativeness based on link structure, such as the PageRank and HITS algorithms. This method is particularly suited to "high-value" content search where the user is interested in searching for information about a specific topic. A method for combining textual estimates of authoritativeness with link analysis is also presented. The types of textual cues to authoritativeness that are easily computed and utilized by our method are described, as well as the method used to select a subset of cues to increase the computation speed. Methods for applying authoritativeness estimates to re-ranking documents returned from search engines, combining textual authoritativeness with social authority, and use in query expansion are also presented. By combining textual authority with link analysis, a more complete and robust estimate can be made of a document's authoritativeness.
Ayman Farahat, Geoffrey Nunberg, Francine Chen 0001
CIKM3
2002 A Hierarchical Model for Clustering and Categorising Documents
Éric Gaussier, Cyril Goutte, Kris Popat, Francine Chen 0001
ECIR4
2001 Text Classification in a Hierarchical Mixture Model for Small Training Sets
abstract
Documents are commonly categorized into hierarchies of topics, such as the ones maintained by Yahoo! and the Open Directory project, in order to facilitate browsing and other interactive forms of information retrieval. In addition, topic hierarchies can be utilized to overcome the sparseness problem in text categorization with a large number of categories, which is the main focus of this paper. This paper presents a hierarchical mixture model which extends the standard naive Bayes classifier and previous hierarchical approaches. Improved estimates of the term distributions are made by differentiation of words in the hierarchy according to their level of generality/specificity. Experiments on the Newsgroups and the Reuters-21578 dataset indicate improved performance of the proposed classifier in comparison to other state-of-the-art methods on datasets with a small number of positive examples.
Kristina Toutanova, Francine Chen 0001, Kris Popat, Thomas Hofmann 0001
CIKM2
2000 Introduction
Henry S. Baird, Francine Chen 0001
Inf. Retr.2
1998 Summarization of Imaged Documents without OCR
Francine Chen 0001, Dan S. Bloomberg
Comput. Vis. Image Underst.1
1997 Extraction of Indicative Summary Sentences from Imaged Documents
abstract
A system for selecting sentences from an imaged document for presentation as part of a document summary is presented. The extracts are identified without the use of optical character recognition. The sentences are selected based on a set of discrete features characterizing the words within a sentence and the location of the sentence within the imaged document. Each sentence is scored based on the values of the discrete features using a statistically based classifier. The imaged document is processed to identify the word locations, the reading order of words, and the location of sentence and paragraph boundaries in the text. The words are grouped into equivalence classes to mimic the terms in a text document. A sample extract for a technical document is shown, and evaluation against a set of abstracts created by a professional abstracting company is given. These results are compared with text-based abstracts.
Francine Chen 0001, Dan S. Bloomberg
ICDAR1
1996 Document image summarization without OCR
abstract
A system for selecting excerpts directly from imaged text without performing optical character recognition is described. The images are segmented to find text regions, text lines and words, and sentence and paragraph boundaries are identified. A set of word equivalence classes is computed based on the rank blur hit-miss transform. This information is used to identify stop words and keywords. Sentences for presentation as part of a summary are then selected based on keywords and on the location of the sentences.
Dan S. Bloomberg, Francine Chen 0001
ICIP (2)2
1995 A comparison of discrete and continuous hidden Markov models for phrase spotting in text images
abstract
In spotting for phrases in text images, speed and accuracy are important considerations. In a hidden Markov model (HMM) based spotter recognition time is dominated by the time required to compute the state conditional observation probabilities. These probabilities are a measure of how well the data match each state in the model. In this paper discrete and continuous hidden Markov models are compared based on speed and accuracy in spotting for phrases in text images. For the discrete HMM, vector quantization is used to associate each continuous feature vector with a discrete value. For the continuous HMMs, the observation distributions for the feature vectors are modeled by either a single Gaussian, or a mixture of two Gaussians. Comparisons were made on a subset of the UW English Document Image Database I. The best accuracy was observed when a mixture of two Gaussians was used in the continuous HMM. The discrete HMM provides for faster spotting particularly when long phrases are used.
Francine Chen 0001, Lynn Wilcox, Dan S. Bloomberg
ICDAR1
1995 A Trainable Document Summarizer
abstract
To summarize is to reduce in complexity, and hence in length, while retaining some of the essential qualities of the original.This paper focusses on document extracts, a particular kind of computed document summary.Document extracts consisting of roughly 20% of the original cart be as informative as the full text of a document, which suggests that even shorter extracts may be useful indicative summmies.The trends in our results are in agreement with those of Ed- mundson who used a subjectively weighted combination of features as opposed to training the feature weights using a corpus.We have developed a trainable summarization program that is grounded in a sound statistical framework.
Julian Kupiec, Jan O. Pedersen 0001, Francine Chen 0001
SIGIR3
1994 Segmentation of speech using speaker identification
abstract
This paper describes techniques for segmentation of conversational speech based on speaker identity. Speaker segmentation is performed using Viterbi decoding on a hidden Markov model network consisting of interconnected speaker sub-networks. Speaker sub-networks are initialized using Baum-Welch training on data labeled by speaker, and are iteratively retrained based on the previous segmentation. If data labeled by speaker is not available, agglomerative clustering is used to approximately segment the conversational speech according to speaker prior to Baum-Welch training. The distance measure for the clustering is a likelihood ratio in which speakers are modeled by Gaussian distributions. The distance between merged segments is recomputed at each stage of the clustering, and a duration model is used to bias the likelihood ratio. Segmentation accuracy using agglomerative clustering initialization matches accuracy using initialization with speaker labeled data.>
Lynn Wilcox, Francine Chen 0001, Don Kimber, Vijay Balasubramanian
ICASSP (1)2
1993 Word spotting in scanned images using hidden Markov models
Francine Chen 0001, Lynn Wilcox, Dan S. Bloomberg
ICASSP (5)1
1993 Detecting and locating partially specified keywords in scanned images using hidden Markov models
abstract
A hidden Markov model (HMM) based system for detecting locating, or spotting, user-specified keywords in scanned images is described. The system is font-independent, and no pre-segmentation of text and graphics is required. The bounding boxes of potential lines of text are extracted from the image using morphology. Feature vectors based on the external shape and internal structure of characters are computed for each bounding box. A keyword HMM is created by concatenating appropriate context-dependent character HMMs. The non-keyword HMM is based on context-dependent sub-character models. Keywords are spotted using Viterbi decoding on an HMM network created from the keyword and non-keyword HMMs. This model allows detection of keywords embedded in a line without pre-segmentation of the line into words or characters. Thus keywords may be specified by a baseform and variants of the keyword can be detected.>
Francine Chen 0001, Lynn Wilcox, Dan S. Bloomberg
ICDAR1
1992 The use of emphasis to automatically summarize a spoken discourse
abstract
The authors describe a method for exploiting prosodic information in natural, conventional speed for the purpose of automatically creating an audio summary. The method is based on identifying emphasized speech and then using proximity measures on the emphasized regions to select summarizing excerpts. Emphasized speech is recognized using a hidden Markov model using only non-spectral, periodic information. Syllable-based models were created and the models trained on spontaneous speech in which words had been labeled by a panel of listeners for degree of emphasis. Emphatic speech from one speaker was automatically detected and summarizing excerpts were identified, with no noticeable difference when compared to excerpts selected by individual subjects. The extensibility of the emphasis detector to other speakers was tested on a small sample of telephone speech by ten other speakers.>
Francine Chen 0001, Margaret Withgott
ICASSP1
1990 Identification of contextual factors for pronunciation networks
abstract
A data-intensive, semiautomatic method is presented for identifying subsets of contextual factors which are useful for predicting the allophonic realizations of dictionary phonemes. The method organizes contextual descriptions of phonological variation into context trees. Context trees are computed using a combination of decision tree induction for factor selection and hierarchical clustering for forming natural groups of factor values. How the resulting context trees can be used to provide allophones in creating pronunciation networks is described. A phoneme-level representation with a flexible context description is used. It allows modeling of effects extending across syllables and word boundaries.>
Francine Chen 0001
ICASSP1
1990 Application of Markov random fields to formant extraction
abstract
The theory of Markov random fields (MRFs) is used to extract formants from linear predictive coding (LPC) pole locations. In this model, the spectral peaks are presented in a time-frequency plot, using LPC poles as noisy indications of these peaks, and the formants are represented as lines connecting the spectral peaks. The spectral peaks and formants are modeled by a site process MRF and a line process MRF, respectively. The a posteriori probability of the spectral peaks and formants, given the LPC poles, quantifies the continuity constraints on the formants, the relationship between spectral peaks and formant locations, and the noise of the LPC poles. The formants are chosen by maximizing the a posteriori probability using simulated annealing. The parameters characterizing the MRFs are estimated by maximizing a pseudolikelihood function derived from LPC pole data with correctly labeled formants.>
Lynn Wilcox, Francine Chen 0001
ICASSP2
1986 Lexical access and verification in a broad phonetic approach to continuous digit recognition
abstract
This paper describes an implementation of a robust method of lexical access and a detailed phonetic verification component for recognizing continuous digits using a broad phonetic approach. The lexical access component uses a scoring method which takes into account soft labeling errors due to input signal variability. Verification is based on the use of a small set of detailed acoustic features which characterize phone hypotheses. Evaluation of the lexical access method on a database of 74 new random length digit strings, each spoken by 5 new speakers, shows the method to be tolerant to front-end errors and variations in pronunciation. Evaluation of the verification component indicates that use of a few detailed phonetic features is adequate for verification of phones in the digit vocabulary.
Francine Chen 0001
ICASSP1
1984 Application of allophonic and lexical constraints in continuous digit recognition
abstract
This paper considers the role of allophonic and lexical constraints in the recognition of continuous digit strings. We first describe a study using both narrow and broad ideal phonetic representations of digit strings for lexical access. This study suggests that allophonic and lexical constraints are very powerful in the digit recognition task. We then describe a system which segments the speech signal into broad phonetic classes and applies allophonic and lexical constraints to produce word hypotheses from a spoken digit string. The system was evaluated on three new speakers (one male and two female). After lexical access, the correct digit was not among the set of hypotheses only 1% of the time.
Francine Chen 0001, Victor Zue
ICASSP1