Christopher T. Jordan

dblp:j/CTJordan · also Chris Jordan · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
0since 2021 · last 2010
0009-0007-1942-1752ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Face, body and person analysis · 66% Learning paradigms · 18% Vision and language · 16%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 58% Distributed systems · 36% High-performance computing · 5%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
human pose estimation
0.112010
Adaptive pose priors for pictorial structures · CVPR 2010
Computer vision › Face, body and person analysis › human pose estimation
pictorial structures
0.112010
Adaptive pose priors for pictorial structures · CVPR 2010
Distributed systems
grid computing
0.122005
Massive High-Performance Global File Systems for Grid computing · SC 2005
The Grid2003 Production Grid: Principles and Practice · HPDC 2004
Computer vision › Face, body and person analysis
face recognition
0.112009
Learning from ambiguously labeled images · CVPR 2009
Machine learning › Learning paradigms
weakly supervised learning
0.112009
Learning from ambiguously labeled images · CVPR 2009
Information retrieval
interactive information retrieval
0.112009
wikiSearch: enabling interactivity in search · SIGIR 2009
Information retrieval › search engines
search engine architecture
0.112009
wikiSearch: enabling interactivity in search · SIGIR 2009
Multimedia analysis and retrieval › cross-modal alignment
video-text alignment
0.112008
Movie/Script: Alignment and Parsing of Video and Text Transcription · ECCV (4) 2008
Information retrieval › text analysis › text preprocessing
morphological analysis
0.112006
Swordfish: an unsupervised Ngram based approach to morphological analysis · SIGIR 2006
Storage systems › file systems
distributed file system
0.112005
Massive High-Performance Global File Systems for Grid computing · SC 2005
Storage systems › distributed storage
global file system
0.112005
Massive High-Performance Global File Systems for Grid computing · SC 2005
Storage systems › file systems › distributed file system
wide-area file system
0.112005
Massive High-Performance Global File Systems for Grid computing · SC 2005
Computer vision › Face, body and person analysis › person identification
character identification in video
0.012009
Learning from ambiguously labeled images · CVPR 2009
Information retrieval
search interfaces
0.012009
wikiSearch: enabling interactivity in search · SIGIR 2009
High-performance computing
distributed computing infrastructure
0.012004
The Grid2003 Production Grid: Principles and Practice · HPDC 2004

Methods — techniques the papers use, named apart from their topics

semi-parametric model · 0.1nearest neighbor · 0.1kernel regression · 0.1fibre channel frame encoding · 0.1TCP/IP · 0.1wikipedia corpus · 0.1convex surrogate loss minimization · 0.1n-gram probability · 0.1log odds · 0.1joint probability · 0.1
YearPublicationVenuePosition
2010 Adaptive pose priors for pictorial structures
abstract
Pictorial structure (PS) models are extensively used for part-based recognition of scenes, people, animals and multi-part objects. To achieve tractability, the structure and parameterization of the model is often restricted, for example, by assuming tree dependency structure and unimodal, data-independent pairwise interactions. These expressivity restrictions fail to capture important patterns in the data. On the other hand, local methods such as nearest-neighbor classification and kernel density estimation provide non-parametric flexibility but require large amounts of data to generalize well. We propose a simple semi-parametric approach that combines the tractability of pictorial structure inference with the flexibility of non-parametric methods by expressing a subset of model parameters as kernel regression estimates from a learned sparse set of exemplars. This yields query-specific, image-dependent pose priors. We develop an effective shape-based kernel for upper-body pose similarity and propose a leave-one-out loss function for learning a sparse subset of exemplars for kernel regression. We apply our techniques to two challenging datasets of human figure parsing and advance the state-of-the-art (from 80% to 86% on the Buffy dataset), while using only 15% of the training data as exemplars.
Benjamin Sapp, Christopher T. Jordan, Ben Taskar
CVPR2
2009 Learning from ambiguously labeled images
abstract
In many image and video collections, we have access only to partially labeled data. For example, personal photo collections often contain several faces per image and a caption that only specifies who is in the picture, but not which name matches which face. Similarly, movie screenplays can tell us who is in the scene, but not when and where they are on the screen. We formulate the learning problem in this setting as partially-supervised multiclass classification where each instance is labeled ambiguously with more than one label. We show theoretically that effective learning is possible under reasonable assumptions even when all the data is weakly labeled. Motivated by the analysis, we propose a general convex learning formulation based on minimization of a surrogate loss appropriate for the ambiguous label setting. We apply our framework to identifying faces culled from Web news sources and to naming characters in TV series and movies. We experiment on a very large dataset consisting of 100 hours of video, and in particular achieve 6% error for character naming on 16 episodes of LOST.
Timothée Cour, Benjamin Sapp, Christopher T. Jordan, Ben Taskar
CVPR3
2009 wikiSearch: enabling interactivity in search
abstract
wikiSearch, is a search engine customized for the Wikipedia corpus but with design features that may be generalized to other search systems. Its features enhance basic functionality and enable more fluid interactivity while supporting both workflow in the search process and the experimental process used in lab testing.
Elaine Toms, Tayze Mackenzie, Christopher T. Jordan, Sam Hall
SIGIR3
2009 Addressing gaps in knowledge while reading
abstract
Abstract Reading is a common everyday activity for most of us. In this article, we examine the potential for using Wikipedia to fill in the gaps in one's own knowledge that may be encountered while reading. If gaps are encountered frequently while reading, then this may detract from the reader's final understanding of the given document. Our goal is to increase access to explanatory text for readers by retrieving a single Wikipedia article that is related to a text passage that has been highlighted. This approach differs from traditional search methods where the users formulate search queries and review lists of possibly relevant results. This explicit search activity can be disruptive to reading. Our approach is to minimize the user interaction involved in finding related information by removing explicit query formulation and providing a single relevant result. To evaluate the feasibility of this approach, we first examined the effectiveness of three contextual algorithms for retrieval. To evaluate the effectiveness for readers, we then developed a functional prototype that uses the text of the abstract being read as context and retrieves a single relevant Wikipedia article in response to a passage the user has highlighted. We conducted a small user study where participants were allowed to use the prototype while reading abstracts. The results from this initial study indicate that users found the prototype easy to use and that using the prototype significantly improved their stated understanding and confidence in that understanding of the academic abstracts they read.
Christopher T. Jordan, Carolyn R. Watters
J. Assoc. Inf. Sci. Technol.1
2008 Cyberinfrastructure Collaboration for Distributed Digital Preservation
abstract
The data deluge is beginning to have an effect on libraries and archives. As custodians of the scholarly record, libraries and archives are being asked to play an active role in long-term digital preservation in both science and the humanities. A report to the National Science Foundation from the fall 2006 ARL Workshop on the role of academic libraries in the digital data universe states that "the group found that research and academic libraries need to expand their portfolios to include activities related to storage, preservation and curation of digital scientific and engineering data." (To Stand, p 42) One of the major trends in this area is the notion of partnerships, of considering the full set of skills necessary to preserve data for the long term and recognizing that a single group or discipline does not have expertise in all aspects of digital preservation. Libraries and archives provide expertise in information management, organization and accessibility. Computer scientists and engineers provide expertise in the portfolio of technologies required to support digital preservation. Domain scientists and humanities scholars provide expertise in the content of the data to be preserved. In order to be effective, these groups must work together. This poster describes three such collaborations based at the Texas Advanced Computer Center, the San Diego Supercomputer Center, and Indiana University.
Christopher T. Jordan, Robert H. McDonald, David Minor, Ardys Kozbial
eScience1
2008 Movie/Script: Alignment and Parsing of Video and Text Transcription
Timothée Cour, Christopher T. Jordan, Eleni Miltsakaki, Ben Taskar
ECCV (4)2
2006 Marching Towards Nirvana: Configurations for Very High Performance Parallel File Systems
abstract
Over the past 7 years, the San Diego Supercomputer Center has worked to produce the highest possible performance file systems available to the National Science Foundation community in the USA. Most of this was done with GPFS, IBM's parallel file system, but several distinctly different configurations were designed and implemented with numerous lessons learned in the process. All of these systems provided transfer rates in the multiple GB/s range. In this paper, we detail the configurations and their intended modes of operation and, as much as possible, show the resulting performance. We attempt to describe the advantages and disadvantages of each approach with an emphasis on the implications for future systems
Phil Andrews, Christopher T. Jordan, Wayne Pfeiffer
CLUSTER2
2006 Swordfish: an unsupervised Ngram based approach to morphological analysis
abstract
Extracting morphemes from words is a nontrivial task. Rule based stemming approaches such as Porter's algorithm have encountered some success, however they are restricted by their ability to identify a limited number of affixes and are language dependent. When dealing with languages with many affixes, rule based approaches generally require many more rules to deal with all the possible word forms. Deriving these rules requires a larger effort on the part of linguists and in some instances can be simply impractical. We propose an unsupervised ngram based approach, named Swordfish. Using ngram probabilities in the corpus, possible morphemes are identified. We look at two possible methods for identifying candidate morphemes, one using joint probabilities between two ngrams, and the second based on log odds between prefix probabilities. Initial results indicate the joint probability approach to be better for English while the prefix ratio approach is better for Finnish and Turkish.
Christopher T. Jordan, John Healy, Vlado Keselj
SIGIR1
2005 Scaling a Global File System to the Greatest Possible Extent, Performance, Capacity, and Number of Users
abstract
We investigate here, both theoretically and by demonstration, scaling file storage to the very widest possible extents. We use IBM's GPFS file system, with extensions developed by the San Diego Supercomputer Center in collaboration with IBM. Geographically, the file system extends across the United States, including Pittsburgh, Illinois, and San Diego, California, with the TeraGrid 40 Gb/s backbone providing the wide area network connectivity. We show the results from two demonstrations, at each of the past two supercomputing conferences, SC03 in Phoenix, Arizona, and SC04 in Pittsburgh, Pennsylvania. The second demonstration was purposely designed to presage an intended production facility across the National Science Foundation's TeraGrid.
Phil Andrews, Bryan Banister, Patricia A. Kovatch, Christopher T. Jordan, Roger L. Haskin
MSST4
2005 Massive High-Performance Global File Systems for Grid computing
abstract
In this paper we describe the evolution of Global File Systems from the concept of a few years ago, to a first demonstration using hardware Fibre Channel frame encoding into IP packets, to a native GFS, to a full prototype demonstration, and finally to a production implementation. The surprisingly excellent performance of the Global File Systems over standard TCP/IP Wide Area Networks has made them a viable candidate for the support of Grid Supercomputing. The implementation designs and performance results are documented within this paper. We also motivate and describe the authentication extensions we made to the IBM GPFS file system, in collaboration with IBM. In several ways Global File Systems are superior to the original approach of wholesale file movement between grid sites and we speculate as to future modes of operation.
Phil Andrews, Patricia A. Kovatch, Christopher T. Jordan
SC3
2004 The Grid2003 Production Grid: Principles and Practice
Ian T. Foster, Jerry Gieraltowski, Scott Gose, Natalia Maltsev, Edward N. May, Alexis A. Rodriguez, Dinanath Sulakhe, A. Vaniachine, Jim Shank, Saul Youssef, David Adams, Richard Baker 0003, Wensheng Deng, Dantong Yu, Iosif Legrand, Conrad Steenberg, M. Anzar Afaq, Eileen Berman, James Annis, L. A. T. Bauerdick, Michael Ernst, Ian Fisk, Lisa Giacchetti, Gregory E. Graham, Anne Heavey, Joseph Kaiser, Nickolai Kuropatkin, Ruth Pordes, Vijay Sekhri, John Weigand, Yujun Wu, Keith Baker, Lawrence Sorrillo, John Huth, Matthew Allen, Leigh Grundhoefer, John Hicks, Fred Luehring, Steve Peck, Robert Quick, Stephen C. Simms, George Fekete, Jan vandenBerg, Kihyeon Cho, Kihwan Kwon, Dongchul Son, Hyoungwoo Park, Shane Canon, Keith R. Jackson, David E. Konerding, Jason Lee 0001, Doug Olson, Iwona Sakrejda, Brian Tierney, Mark Green 0001, Russ Miller, James Letts, Terrence Martin, David Bury, Catalin Dumitrescu, Daniel Engh, Robert W. Gardner, Marco Mambelli, Yuri Smirnov, Jens-S. Vöckler, Michael Wilde, Yong Zhao 0009, Paul Avery, Richard Cavanaugh, Bockjoo Kim, Craig Prescott, Jorge Rodríguez 0002, Andrew Zahn, Shawn McKee, Christopher T. Jordan, James E. Prewett, Timothy L. Thomas, Horst Severini, Ben Clifford, Ewa Deelman, Larry Flon, Carl Kesselman, Gaurang Mehta, Nosa Olomu, Karan Vahi, Kaushik De, Patrick McGuigan, Mark Sosebee, Dan Bradley, Peter Couvares, Alan DeSmet, Carey Kireyev, Erik Paulson 0001, Alain J. Roy, Scott Koranda, Brian Moe, Bobby Brown, Paul Sheldon
HPDC79