Bingjun Sun

dblp:34/4451 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
0since 2021 · last 2011
0000-0002-6036-3730ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 7 first-authorArtificial intelligence and machine learning · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Information retrieval · 76% Query processing and optimization · 10% Data mining · 8%
Artificial intelligence
1 paper
Probabilistic and Bayesian machine learning · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
indexing
0.222011
Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011
Mining, indexing, and searching for textual chemical molecule information on the web · WWW 2008
Information retrieval › indexing › index compression
index pruning
0.112011
Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011
Query processing and optimization › selection queries
partial match query
0.112011
Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011
Information retrieval
query processing
0.112011
Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011
Information retrieval
search engines
0.112011
Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
hierarchical conditional random field
0.112008
Mining, indexing, and searching for textual chemical molecule information on the web · WWW 2008
Information retrieval › document retrieval
domain-specific retrieval
0.112008
Mining, indexing, and searching for textual chemical molecule information on the web · WWW 2008
Data mining › pattern mining › sequential pattern mining
frequent sequence mining
0.112008
Mining, indexing, and searching for textual chemical molecule information on the web · WWW 2008
Information retrieval › query understanding
query modeling
0.112007
Extraction and search of chemical formulae in text documents on the web · WWW 2007
Web and social media mining
social network analysis
0.112007
Predicting Blogging Behavior Using Temporal and Social Networks · ICDM 2007
Information retrieval › text analysis › topic analysis
topic detection and tracking
0.112007
Topic segmentation with shared topic detection and alignment of multiple documents · SIGIR 2007
Information retrieval › text analysis › text segmentation
topic segmentation
0.112007
Topic segmentation with shared topic detection and alignment of multiple documents · SIGIR 2007
Data mining
behavior modeling
0.012007
Predicting Blogging Behavior Using Temporal and Social Networks · ICDM 2007

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.2conditional random field · 0.2index pruning · 0.2hierarchical text segmentation · 0.2frequent subsequence mining · 0.2hierarchical conditional random fields · 0.1hierarchical conditional random field · 0.1regression · 0.1general regression neural network · 0.1extreme learning machine · 0.1entropy-based term weighting · 0.1
YearPublicationVenuePosition
2011 Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents
abstract
End-users utilize chemical search engines to search for chemical formulae and chemical names. Chemical search engines identify and index chemical formulae and chemical names appearing in text documents to support efficient search and retrieval in the future. Identifying chemical formulae and chemical names in text automatically has been a hard problem that has met with varying degrees of success in the past. We propose algorithms for chemical formula and chemical name tagging using Conditional Random Fields (CRFs) and Support Vector Machines (SVMs) that achieve higher accuracy than existing (published) methods. After chemical entities have been identified in text documents, they must be indexed. In order to support user-provided search queries that require a partial match between the chemical name segment used as a keyword or a partial chemical formula, all possible (or a significant number of) subformulae of formulae that appear in any document and all possible subterms (e.g., “methyl”) of chemical names (e.g., “methylethyl ketone”) must be indexed. Indexing all possible subformulae and subterms results in an exponential increase in the storage and memory requirements as well as the time taken to process the indices. We propose techniques to prune the indices significantly without reducing the quality of the returned results significantly. Finally, we propose multiple query semantics to allow users to pose different types of partial search queries for chemical entities. We demonstrate empirically that our search engines improve the relevance of the returned results for search queries involving chemical entities.
Bingjun Sun, Prasenjit Mitra 0001, C. Lee Giles, Karl T. Mueller
ACM Trans. Inf. Syst.1
2010 Human-Agent Collaboration for Time-Stressed Multicontext Decision Making
abstract
Multicontext team decision making under time stress is an extremely challenging issue faced by various real-world application domains. In this paper, we employ an experience-based cognitive agent architecture (called R-CAST) to address the informational challenges associated with military command and control (C2) decision-making teams, the performance of which can be significantly affected by dynamic context switching and tasking complexities. Using context switching frequency and task complexity as two factors, we conducted an experiment to evaluate whether the use of R-CAST agents as teammates and decision aids can benefit C2decision-making teams. Members from a U.S. Army Reserve Officer Training Corps organization were randomly recruited as human participants. They were grouped into ten human-human teams, each composed of two participants, and ten human-agent teams, each composed of one participant and two R-CAST agents, as teammates and decision aids. The statistical inference of experimental results indicates that R-CAST agents can significantly improve the performance of C2teams in multicontext decision making under varying time-stressed situations.
Xiaocong Fan, Michael D. McNeese, Bingjun Sun, Tim Hanratty, Laurel Allender, John Yen
IEEE Trans. Syst. Man Cybern. Part A3
2009 Independent informative subgraph mining for graph information retrieval
abstract
In order to enable scalable querying of graph databases, intelligent selection of subgraphs to index is essential. An improved index can reduce response times for graph queries significantly. For a given subgraph query, graph candidates that may contain the subgraph are retrieved using the graph index and subgraph isomorphism tests are performed to prune out unsatisfied graphs. However, since the space of all possible subgraphs of the whole set of graphs is prohibitively large, feature selection is required to identify a good subset of subgraph features for indexing. Thus, one of the key issues is: given the set of all possible subgraphs of the graph set, which subset of features is the optimal such that the algorithm retrieves the smallest set of candidate graphs and reduces the number of subgraph isomorphism tests? We introduce a graph search method for subgraph queries based on subgraph frequencies. Then, we propose several novel feature selection criteria, Max-Precision, Max-Irredundant-Information, and Max-Information-Min-Redundancy, based on mutual information. Finally we show theoretically and empirically that our proposed methods retrieve a smaller candidate set than previous methods. For example, using the same number of features, our method improve the precision for the query candidate set by 4%-13% in comparison to previous methods. As a result the response time of subgraph queries also is improved correspondingly.
Bingjun Sun, Prasenjit Mitra 0001, C. Lee Giles
CIKM1
2009 Learning to rank graphs for online similar graph search
abstract
Many applications in structure matching require the ability to search for graphs that are similar to a query graph, i.e., similarity graph queries. Prior works, especially in chemoinformatics, have used the maximum common edge subgraph (MCEG) to compute the graph similarity. This approach is prohibitively slow for real-time queries. In this work, we propose an algorithm that extracts and indexes subgraph features from a graph dataset. It computes the similarity of graphs using a linear graph kernel based on feature weights learned offline from a training set generated using MCEG. We show empirically that our proposed algorithm of learning to rank graphs can achieve higher normalized discounted cumulative gain compared with existing optimal methods based on MCEG. The running time of our algorithm is orders of magnitude faster than these existing methods.
Bingjun Sun, Prasenjit Mitra 0001, C. Lee Giles
CIKM1
2008 Mining, indexing, and searching for textual chemical molecule information on the web
abstract
Current search engines do not support user searches for chemical entities (chemical names and formulae) beyond simple keyword searches. Usually a chemical molecule can be represented in multiple textual ways. A simple keyword search would retrieve only the exact match and not the others. We show how to build a search engine that enables searches for chemical entities and demonstrate empirically that it improves the relevance of returned documents. Our search engine first extracts chemical entities from text, performs novel indexing suitable for chemical names and formulae, and supports different query models that a scientist may require. We propose a model of hierarchical conditional random fields for chemical formula tagging that considers long-term dependencies at the sentence level. To substring searches of chemical names, a search engine must index substrings of chemical names. Indexing all possible sub-sequences is not feasible in practice. We propose an algorithm for independent frequent subsequence mining to discover sub-terms of chemical names with their probabilities. We then propose an unsupervised hierarchical text segmentation (HTS) method to represent a sequence with a tree structure based on discovered independent frequent subsequences, so that sub-terms on the HTS tree should be indexed. Query models with corresponding ranking functions are introduced for chemical name searches. Experiments show that our approaches to chemical entity tagging perform well. Furthermore, we show that index pruning can reduce the index size and query time without changing the returned ranked results significantly. Finally, experiments show that our approaches out-perform traditional methods for document search with ambiguous chemical terms.
Bingjun Sun, Prasenjit Mitra 0001, C. Lee Giles
WWW1
2007 Predicting Blogging Behavior Using Temporal and Social Networks
abstract
Modeling the behavior of bloggers is an important problem with various applications in recommender systems, targeted advertising, and event detection. In this paper, we propose three models by combining content, temporal, social dimensions: the general blogging-behavior model, the profile-based blogging-behavior model and the social- network and profile-based blogging-behavior model. The models are based on two regression techniques: Extreme Learning Machine (ELM), and Modified General Regression Neural Network (MGRNN). We choose one of the largest blogs, a political blog, DailyKos1, for our empirical evaluation. Experiments show that the social network and profile-based blogging behavior model with ELM regression techniques produce good results for the most active bloggers and can be used to predict blogging behavior.
Bi Chen, Qiankun Zhao, Bingjun Sun, Prasenjit Mitra 0001
ICDM3
2007 Topic segmentation with shared topic detection and alignment of multiple documents
abstract
Topic detection and tracking and topic segmentation play an important role in capturing the local and sequential information of documents. Previous work in this area usually focuses on single documents, although similar multiple documents are available in many domains. In this paper, we introduce a novel unsupervised method for shared topic detection and topic segmentation of multiple similar documents based on mutual information (MI) and weighted mutual information (WMI) that is a combination of MI and term weights. The basic idea is that the optimal segmentation maximizes MI (or WMI). Our approach can detect shared topics among documents. It can find the optimal boundaries in a document, and align segments among documents at the same time. It also can handle single-document segmentation as a special case of the multi-document segmentation and alignment. Our methods can identify and strengthen cue terms that can be used for segmentation and partially remove stop words by using term weights based on entropy learned from multiple documents. Our experimental results show that our algorithm works well for the tasks of single-document segmentation, shared topic detection, and multi-document segmentation. Utilizing information from multiple documents can tremendously improve the performance of topic segmentation, and using WMI is even better than using MI for the multi-document segmentation.
Bingjun Sun, Prasenjit Mitra 0001, C. Lee Giles, John Yen, Hongyuan Zha
SIGIR1
2007 Extraction and search of chemical formulae in text documents on the web
abstract
Often scientists seek to search for articles on the Web related to a particular chemical. When a scientist searches for a chemical formula using a search engine today, she gets articles where the exact keyword string expressing the chemical formula is found. Searching for the exact occurrence of keywords during searching results in two problems for this domain: a) if the author searches for CH4 and the article has H4C, the article is not returned, and b) ambiguous searches like "He" return all documents where Helium is mentioned as well as documents where the pronoun "he" occurs. To remedy these deficiencies, we propose a chemical formula search engine. To build a chemical formula search engine, we must solve the following problems: 1) extract chemical formulae from text documents, 2) index chemical formulae, and 3) designranking functions for the chemical formulae. Furthermore, query models are introduced for formula search, and for each a scoring scheme based on features of partial formulae is proposed tomeasure the relevance of chemical formulae and queries. We evaluate algorithms for identifying chemical formulae in documents using classification methods based on Support Vector Machines(SVM), and a probabilistic model based on conditional random fields (CRF). Different methods for SVM and CRF to tune the trade-off between recall and precision forim balanced data are proposed to improve the overall performance. A feature selection method based on frequency and discrimination isused to remove uninformative and redundant features. Experiments show that our approaches to chemical formula extraction work well, especially after trade-off tuning. The results also demonstrate that feature selection can reduce the index size without changing ranked query results much.
Bingjun Sun, Qingzhao Tan, Prasenjit Mitra 0001, C. Lee Giles
WWW1
2006 Multi-task text segmentation and alignment based on weighted mutual information
abstract
Text segmentation is important for text analysis, while text alignment is to determine shared sub-topics among similar documents. Multi-task text segmentation and alignment is the extension of single-task segmentation to utilize information of multi-source documents. In this paper we introduce a novel domain-independent unsupervised method for multi-task segmentation and alignment based on the idea that the optimal segmentation and alignment maximizes weighted mutual information, mutual information with term weights. The experiment results show that our approach works well.
Bingjun Sun, Hongyuan Zha, John Yen
CIKM1
2004 Energy-Efficient Scheduling Algorithms of Object Retrieval on Indexed Parallel Broadcast Channels
abstract
With the goal of providing "timely and reliable" access to information in a mobile computing environment, mobile units and the wireless medium operate under constraints on energy, bandwidth, and connectivity. Among these limitations, power limitation of mobile units is one of the key issues. In a mobile computing environment, broadcasting has proved to be an effective method to distribute public data. Efficient methods for allocating and retrieving objects on parallel indexed broadcast channels have been proposed to manage power consumption and access latency. Employment of parallel channels also brings out the notion of conflicts. To minimize the effect of conflicts on both access latency and power consumption, one has to develop schemes to schedule access to the objects that minimizes the number of passes over the parallel channels. This work extends our past efforts and proposes two new scheduling algorithms that can find the minimum number of passes and inside channel switches. The simulation results show that the proposed scheduling algorithms relative to our previous work have a great impact on energy consumption and access latency. The proposed scheduling algorithms are simulated and results are presented.
Bingjun Sun, Ali R. Hurson, John Hannan
ICPP1