EDBT 2026 Demo / reviewers in the wild / expert
Zhijun Yin
dblp:53/5414
· DBLP profile ↗
38ranked-venue papers
16as first author
10since 2021 · last 2026
0000-0002-3075-1337ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 10 first-author · 9 since 2021Databases, data management, data science and information retrieval · 17 · 10 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 4 first-authorSecurity and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Affordances as a Lens: A Systematic Review of Social Media Benefits and Risks for AdolescentsabstractAlthough social media use is ubiquitous among adolescents, research often emphasizes benefits and risks without considering the design-based affordances that may influence these experiences. Affordances enable or constrain user behavior and thus provide a valuable lens for promoting positive online engagement. To this end, we developed an affordance-centered framework and systematically coded 66 empirical studies on U.S. adolescents published in the past decade to map how platform features relate to specific user activities, the affordances they instantiate, and associated benefits and risks. For instance, Discoverability affordance, driven by algorithmic recommendations, supported Cognitive engagement by facilitating learning while heightening exposure to misinformation. Visibility affordance enabled Identity exploration through ephemeral sharing, lowering performance anxiety but raising post-disclosure regret when content reached hostile viewers. By clarifying how specific affordances facilitate adolescent experiences, our work contributes a design-grounded framework for future research that advances both theoretical understanding and the development of youth-centered interventions. Abdulmalik Alluhidan, Renkai Ma, Jinkyung Park, Zainab Agha, Zhijun Yin, Pamela J. Wisniewski |
CHI | 5 |
| 2025 | Using Large Language Model for Efficient Extraction of Treatment Discontinuation Information - A Study of Online Breast Cancer Community Posts
Qingyuan Song, Jessie Yang, Ndidiamaka Obi, Congning Ni, Jeremy L. Warner, Qingxia Chen, S. Trent Rosenbloom, Bradley A. Malin, Zhijun Yin |
AIME (2) | 10 |
| 2025 | Leveraging scHi-C Data for Integrated Single-Cell Omics AnalysisabstractThe integration of single-cell multi-omics data is essential for deciphering complex gene regulatory programs. While single-cell Hi-C (scHi-C) provides insight into 3D genome architecture, its utilization in multi-omics integration remains underexplored. Here, we use mouse brain single-cell multi-omics integration as a case study to demonstrate the dual utility of scHiC data within a knowledge graph-based integrative framework. First, we use scHi-C as a gene regulatory prior to construct a Hi-C-driven guidance graph. This approach enhances integration of scRNA-seq and scATAC-seq data, resulting in an improved alignment score (FOSCTTM$=0.0356)$. Second, we show that the framework can directly integrate scHi-C as a primary data modality with scRNA-seq. We validate this method on a paired scRNA-seq and scHi-C dataset, achieving 85.0% mapping accuracy, and demonstrate its power in a cross-modal labeltransfer application. This direct integration successfully refines a broader neuronal cluster into finer, distinct hippocampal granule and pyramidal subtypes. Our work presents an effective strategy for leveraging scHi-C data in multi-omics integration, providing a powerful tool to dissect cellular regulatory programs through combining 3D genome organization with gene expression. Weixin Liu 0001, Rui Chen 0021, Yuting Tan 0005, Xue Zhong, Bingshan Li, Zhijun Yin |
BIBM | 7 |
| 2025 | From Voice to Diagnosis: A Hybrid Approach to Multi-Label Disease Classification with Uncertainty AwarenessabstractThe human voice encodes a wealth of acoustic biomarkers linked to various health conditions, including but not limited to neurological, mood, respiratory, and laryngeal disorders. Recent advancements in artificial intelligence (AI) offer a promising solution for leveraging voice to perform noninvasive, cost-effective, and scalable screening of health-related conditions. However, classifying comorbid clinical conditions presents a significant multi-label classification challenge, which is further complicated by class imbalance, feature noise, and the “black-box” nature of modern deep learning models, a critical barrier to model interpretability, error analysis, and potential clinical translation. To address these challenges, we introduce and evaluate a two-stage framework for voice-based, multi-label disease classification on a newly curated dataset from the NIH Bridge2AI initiative. This hybrid approach fuses theory-driven, handcrafted acoustic features with data-driven deep features from a pre-trained ResNet-18. To identify the most effective model configuration, we enhance a Feed-Forward Neural Network (FFNN) with an attention mechanism and Focal Loss, systematically comparing it against traditional classifiers and end-to-end fine-tuning benchmarks. Finally, we incorporate an uncertainty quantification (UQ) layer using Monte Carlo (MC) Dropout and Deep Ensembles to assess prediction reliability. Our results show that the proposed two-stage FFNN with Focal Loss and Attention emerges as the best-performing model in our evaluation, delivering the best balance of high discriminative power (Macro AUC of$\mathbf{0. 8 1 0}$) and robust classification accuracy (Macro F1 of 0.610). While an end-to-end model achieved the highest Macro F1 score (0.638), our two-stage approach proved to be more reliable and computationally efficient. Furthermore, our UQ analysis identified MC Dropout with Predictive Entropy as a practical and effective method, revealing a statistically significant correlation between model uncertainty and prediction error (overall$\mathbf{r} \boldsymbol{=} \mathbf{0. 2 8 0}$,$\mathbf{p}<\mathbf{0. 0 0 1}$), with the strongest link for Voice Disorders ($\mathbf{r} \boldsymbol{=} \mathbf{0. 3 8 0}$). In summary, this study presents a robust technical blueprint for a voice-based AI system, establishing a strong performance benchmark to guide future model development using the Bridge2AI voice dataset. Weixin Liu 0001, Bowen Qu, Matthew E. Pontell, Maria E. Powell, Bradley A. Malin, Zhijun Yin |
BIBM | 6 |
| 2025 | Integrating Single Cell RNA Sequencing Data and Protein Embeddings to Infer Cell-Cell Communication in Alzheimer's DiseaseabstractCell-cell communication (CCC) plays a critical role in the pathogenesis of Alzheimer's disease (AD), yet most computational methods for CCC inference rely exclusively on transcriptomic data, showing low consistency across datasets or methods. In this study, we present a Protein Embedding-Infused Cell Talk (PEICTalk) inference method that integrates singlecell RNA-sequencing (scRNA-seq) data with protein embeddings from the ESM-2 language model to improve the accuracy and interpretability of CCC inference. We applied PEICTalk to two large-scale scRNA-seq datasets, after harmonizing celltype annotations using the SEA-AD taxonomy via MapMyCells. PEICTalk outperformed standard tools such as CellChat and CellPhoneDB, identifying more biologically meaningful ligandreceptor interactions with higher cross-dataset reproducibility. Ablation analysis confirmed the crucial value of protein-level information. This integrative strategy offers a robust foundation for uncovering novel intercellular mechanisms in AD and may serve as a blueprint for future CCC studies. Yuting Tan 0005, Rui Chen 0021, Anshul Tiwari, Zhexing Wen, Xue Zhong, Zhijun Yin, Bingshan Li |
BIBM | 7 |
| 2025 | Catalysts of Conversation: Examining Interaction Dynamics Between Topic Initiators and Commentors in Alzheimer's Disease Online CommunitiesabstractInformal caregivers (e.g., family members or friends) of people living with Alzheimer's Disease and Related Dementias (ADRD) face substantial challenges and often seek support through online communities. Understanding the factors driving engagement within these platforms is crucial, as it can enhance communities' long-term value to meet their needs effectively. This study investigated the user interaction dynamics within two large, popular ADRD communities, TalkingPoint and ALZConnected, focusing on topic initiator engagement, initial post content, and the linguistic patterns of comments at the thread level. Using analytical methods such as propensity score matching, topic modeling, and predictive modeling, we found that active topic initiator engagement drives a higher comment volume, and reciprocal replies from topic initiators encourage further commentor engagement at the community level. Practical caregiving topics prompt more re-engagement of topic initiators, while emotional support topics attract more comments from commentors. Additionally, the linguistic complexity and emotional tone of a comment are associated with its likelihood of receiving replies from topic initiators. These findings highlight the importance of fostering active and reciprocal engagement and providing effective strategies to enhance sustainability in ADRD caregiving and broader health-related online communities. Congning Ni, Qingxia Chen, Patricia Commiskey, Qingyuan Song, Bradley A. Malin, Zhijun Yin |
WWW | 7 |
| 2024 | Improving Genetic Perturbation Response Prediction with an Enhanced Biological Knowledge GraphabstractPerturb-seq is a technique that combines scRNA-seq and CRISPR to explore cellular system operations and disease-associated genes, providing profound insights into the mechanisms behind biological processes. Although powerful, such a method is limited by its scalability for its cost-intensive and time-consuming nature, which calls for in silico prediction of genetic perturbation responses. Among all computational methods, GEARS represents the state-of-the-art by explicitly modeling the response of each gene to the perturbed gene, exploiting gene-gene relationships derived from Gene Ontology annotations. However, our evaluation of Gene Ontology annotations indicated that they are insufficient as the sole source of prior knowledge for predicting genetic perturbation responses. Therefore, they cannot fully support predicting genetic perturbation responses. We addressed this gap by constructing an augmented gene ontology network that incorporates extensive knowledge of diseases, drugs, and genes to capture nuanced gene-gene relationships not indicated by Gene Ontology alone. By replacing only the Gene Ontology graph in GEARS, our method outperforms GEARS in both single gene and combinational perturbation predictions. These findings suggest the effectiveness and importance of incorporating finer prior knowledge in predicting genetic perturbation responses, thereby encouraging future works on improving knowledge representation for single-cell perturbation prediction. Rui Chen 0021, Yuting Tan 0005, Xue Zhong, Bingshan Li, Zhijun Yin |
BIBM | 6 |
| 2022 | "Rough Day ... Need a Hug": Learning Challenges and Experiences of the Alzheimer's Disease and Related Dementia Caregivers on Reddit
Congning Ni, Bradley A. Malin, Angela L. Jefferson, Patricia Commiskey, Zhijun Yin |
ICWSM | 6 |
| 2022 | Dynamically adjusting case reporting policy to maximize privacy and public health utility in the face of a pandemicabstractOBJECTIVE: Supporting public health research and the public's situational awareness during a pandemic requires continuous dissemination of infectious disease surveillance data. Legislation, such as the Health Insurance Portability and Accountability Act of 1996 and recent state-level regulations, permits sharing deidentified person-level data; however, current deidentification approaches are limited. Namely, they are inefficient, relying on retrospective disclosure risk assessments, and do not flex with changes in infection rates or population demographics over time. In this paper, we introduce a framework to dynamically adapt deidentification for near-real time sharing of person-level surveillance data. MATERIALS AND METHODS: The framework leverages a simulation mechanism, capable of application at any geographic level, to forecast the reidentification risk of sharing the data under a wide range of generalization policies. The estimates inform weekly, prospective policy selection to maintain the proportion of records corresponding to a group size less than 11 (PK11) at or below 0.1. Fixing the policy at the start of each week facilitates timely dataset updates and supports sharing granular date information. We use August 2020 through October 2021 case data from Johns Hopkins University and the Centers for Disease Control and Prevention to demonstrate the framework's effectiveness in maintaining the PK11 threshold of 0.01. RESULTS: When sharing COVID-19 county-level case data across all US counties, the framework's approach meets the threshold for 96.2% of daily data releases, while a policy based on current deidentification techniques meets the threshold for 32.3%. CONCLUSION: Periodically adapting the data publication policies preserves privacy while enhancing public health utility through timely updates and sharing epidemiologically critical features. J. Thomas Brown, Chao Yan 0004, Weiyi Xia, Zhijun Yin, Zhiyu Wan, Aris Gkoulalas-Divanis, Murat Kantarcioglu, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Identifying Posts from Reddit Talking About ADRD Caring Experiences Using Machine Learning
Congning Ni, Jonathan Hughes, Andrew Nam, Zhijun Yin |
AMIA | 4 |
| 2019 | Biomedical Research Cohort Membership Disclosure on Social Media
Yongtai Liu, Chao Yan 0004, Zhijun Yin, Zhiyu Wan, Weiyi Xia, Murat Kantarcioglu, Yevgeniy Vorobeychik, Ellen Wright Clayton, Bradley A. Malin |
AMIA | 3 |
| 2019 | Why Patient Portal Messages Indicate Risk of Readmission for Patients with Ischemic Heart Disease
Lina M. Sulieman, Zhijun Yin, Bradley A. Malin |
AMIA | 2 |
| 2019 | Patient Messaging Content Associated with Initiating Hormonal Therapy after a Breast Cancer Diagnosis
Zhijun Yin, Jeremy L. Warner, Qingxia Chen, Bradley A. Malin |
AMIA | 1 |
| 2019 | A systematic literature review of machine learning in online personal health dataabstractOBJECTIVE: User-generated content (UGC) in online environments provides opportunities to learn an individual's health status outside of clinical settings. However, the nature of UGC brings challenges in both data collecting and processing. The purpose of this study is to systematically review the effectiveness of applying machine learning (ML) methodologies to UGC for personal health investigations. MATERIALS AND METHODS: We searched PubMed, Web of Science, IEEE Library, ACM library, AAAI library, and the ACL anthology. We focused on research articles that were published in English and in peer-reviewed journals or conference proceedings between 2010 and 2018. Publications that applied ML to UGC with a focus on personal health were identified for further systematic review. RESULTS: We identified 103 eligible studies which we summarized with respect to 5 research categories, 3 data collection strategies, 3 gold standard dataset creation methods, and 4 types of features applied in ML models. Popular off-the-shelf ML models were logistic regression (n = 22), support vector machines (n = 18), naive Bayes (n = 17), ensemble learning (n = 12), and deep learning (n = 11). The most investigated problems were mental health (n = 39) and cancer (n = 15). Common health-related aspects extracted from UGC were treatment experience, sentiments and emotions, coping strategies, and social support. CONCLUSIONS: The systematic review indicated that ML can be effectively applied to UGC in facilitating the description and inference of personal health. Future research needs to focus on mitigating bias introduced when building study cohorts, creating features from free text, improving clinical creditability of UGC, and model interpretability. Zhijun Yin, Lina M. Sulieman, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Learning When Communications Between Healthcare Providers Indicate Hormonal Therapy Medication Discontinuation
Zhijun Yin, Jeremy L. Warner, Bradley A. Malin |
AMIA | 1 |
| 2018 | It's all in the timing: calibrating temporal penalties for biomedical data sharingabstractObjective: Biomedical science is driven by datasets that are being accumulated at an unprecedented rate, with ever-growing volume and richness. There are various initiatives to make these datasets more widely available to recipients who sign Data Use Certificate agreements, whereby penalties are levied for violations. A particularly popular penalty is the temporary revocation, often for several months, of the recipient's data usage rights. This policy is based on the assumption that the value of biomedical research data depreciates significantly over time; however, no studies have been performed to substantiate this belief. This study investigates whether this assumption holds true and the data science policy implications. Methods: This study tests the hypothesis that the value of data for scientific investigators, in terms of the impact of the publications based on the data, decreases over time. The hypothesis is tested formally through a mixed linear effects model using approximately 1200 publications between 2007 and 2013 that used datasets from the Database of Genotypes and Phenotypes, a data-sharing initiative of the National Institutes of Health. Results: The analysis shows that the impact factors for publications based on Database of Genotypes and Phenotypes datasets depreciate in a statistically significant manner. However, we further discover that the depreciation rate is slow, only ∼10% per year, on average. Conclusion: The enduring value of data for subsequent studies implies that revoking usage for short periods of time may not sufficiently deter those who would violate Data Use Certificate agreements and that alternative penalty mechanisms may need to be invoked. Weiyi Xia, Zhiyu Wan, Zhijun Yin, James Gaupp, Yongtai Liu, Ellen Wright Clayton, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 3 |
| 2018 | The therapy is making me sick: how online portal communications between breast cancer patients and physicians indicate medication discontinuationabstractObjective: Online platforms have created a variety of opportunities for breast patients to discuss their hormonal therapy, a long-term adjuvant treatment to reduce the chance of breast cancer occurrence and mortality. The goal of this investigation is to ascertain the extent to which the messages breast cancer patients communicated through an online portal can indicate their potential for discontinuing hormonal therapy. Materials and Methods: We studied the de-identified electronic medical records of 1106 breast cancer patients who were prescribed hormonal therapy at Vanderbilt University Medical Center over a 12-year period. We designed a data-driven approach to investigate patients' patterns of messaging with healthcare providers, the topics they communicated, and the extent to which these messaging behaviors associate with the likelihood that a patient will discontinue a prescribed 5-year regimen of therapy. Results: The results indicates that messaging rate over time [hazard ratio (HR) = 1.373, P = 0.002], mentions of side effects (HR = 1.214, P = 0.006), and surgery-related topics (HR = 1.170, P = 0.034) were associated with increased risk of early medication discontinuation. In contrast, seeking professional suggestions (HR = 0.766, P = 0.002), expressing gratitude to healthcare providers (HR = 0.872, P = 0.044), and mentions of drugs used to treat side effects (HR = 0.807, P = 0.013) were associated with decreased risk of medication discontinuation. Discussion and Conclusion: This investigation suggests that patient-generated content can inform the study of health-related behaviors. Given that approximately 50% of breast cancer patients do not complete a course of hormonal therapy as described, the identification of factors associated with medication discontinuation can facilitate real-time interventions to prevent early discontinuation. Zhijun Yin, Morgan Harrell, Jeremy L. Warner, Qingxia Chen, Daniel Fabbri, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 1 |
| 2017 | Talking About My Care: Detecting Mentions of Hormonal Therapy Adherence Behavior in an Online Breast Cancer Community
Zhijun Yin, Wei Xie 0002, Bradley A. Malin |
AMIA | 1 |
| 2017 | Reciprocity and its Association with Treatment Adherence in an Online Breast Cancer ForumabstractOnline health communities (OHCs) are increasingly relied upon by individuals exchanging social support for diagnoses and treatment regimens. It has been shown that social support from trusted relationships (e.g., family and friends) positively influence treatment adherence in offline environments, but much less is known about the online setting. In this study, we focus on how relationships established in an online breast cancer discussion board induce reciprocity (specifically in the form of reciprocal exchange of support) and its impact on adherence to a five-year hormonal therapy, a highly prevalent long-term treatment for breast cancers, with varying completion rates. We measure reciprocity as responses to forum posts to analyze interactions of over 6,000 patients and 100,000 responses. In doing so, we assess how reciprocity is related to time active in the OHC and the tones communicated by authors in their posts (e.g., emotions, writing styles and social tendencies). We further assess if such reciprocity is associated with treatment adherence. We find the volume of the reciprocity is positively associated with completing the five-year protocol, rather than the rate of the reciprocity or the fraction of the posts that received replies. Zhijun Yin, Bradley A. Malin |
CBMS | 1 |
| 2017 | The Power of the Patient Voice: Learning Indicators of Treatment Adherence From An Online Breast Cancer Forum
Zhijun Yin, Bradley A. Malin, Jeremy L. Warner, Pei-Yun Sabrina Hsueh, Ching-Hua Chen |
ICWSM | 1 |
| 2016 | #PrayForDad: Learning the Semantics Behind Why Social Media Users Disclose Health Information
Zhijun Yin, You Chen 0001, Daniel Fabbri, Jimeng Sun 0001, Bradley A. Malin |
ICWSM | 1 |
| 2015 | Mining Twitter as a First Step toward Assessing the Adequacy of Gender Identification Terms on Intake Forms
Amanda Hicks, William R. Hogan, Michael W. Rutherford, Bradley A. Malin, Mengjun Xie, Christiane Fellbaum, Zhijun Yin, Daniel Fabbri, Josh Hanna, Jiang Bian 0001 |
AMIA | 7 |
| 2012 | RankCompete: Simultaneous ranking and clustering of information networks
Liangliang Cao, Xin Jin 0001, Zhijun Yin, Andrey Del Pozo, Jiebo Luo 0001, Jiawei Han 0001, Thomas S. Huang |
Neurocomputing | 3 |
| 2012 | Latent Community Topic Analysis: Integration of Community Discovery with Topic ModelingabstractThis article studies the problem of latent community topic analysis in text-associated graphs. With the development of social media, a lot of user-generated content is available with user networks. Along with rich information in networks, user graphs can be extended with text information associated with nodes. Topic modeling is a classic problem in text mining and it is interesting to discover the latent topics in text-associated graphs. Different from traditional topic modeling methods considering links, we incorporate community discovery into topic analysis in text-associated graphs to guarantee the topical coherence in the communities so that users in the same community are closely linked to each other and share common latent topics. We handle topic modeling and community discovery in the same framework. In our model we separate the concepts of community and topic, so one community can correspond to multiple topics and multiple communities can share the same topic. We compare different methods and perform extensive experiments on two real datasets. The results confirm our hypothesis that topics could help understand community structure, while community structure could help model topics. Zhijun Yin, Liangliang Cao, Quanquan Gu, Jiawei Han 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2011 | LPTA: A Probabilistic Model for Latent Periodic Topic AnalysisabstractThis paper studies the problem of latent periodic topic analysis from time stamped documents. The examples of time stamped documents include news articles, sales records, financial reports, TV programs, and more recently, posts from social media websites such as Flickr, Twitter, and Face book. Different from detecting periodic patterns in traditional time series database, we discover the topics of coherent semantics and periodic characteristics where a topic is represented by a distribution of words. We propose a model called LPTA (Latent Periodic Topic Analysis) that exploits the periodicity of the terms as well as term co-occurrences. To show the effectiveness of our model, we collect several representative datasets including Seminar, DBLP and Flickr. The results show that our model can discover the latent periodic topics effectively and leverage the information from both text and time well. Zhijun Yin, Liangliang Cao, Jiawei Han 0001, ChengXiang Zhai, Thomas S. Huang |
ICDM | 1 |
| 2011 | Kipnis-Shamir Attack on Unbalanced Oil-Vinegar Scheme
Lei Hu 0003, Jintai Ding, Zhijun Yin |
ISPEC | 4 |
| 2011 | Diversified Trajectory Pattern Ranking in Geo-tagged Social MediaabstractSocial media such as those residing in the popular photo sharing websites is attracting increasing attention in recent years. As a type of user-generated data, wisdom of the crowd is embedded inside such social media. In particular, millions of users upload to Flickr their photos, many associated with temporal and geographical information. In this paper, we investigate how to rank the trajectory patterns mined from the uploaded photos with geotags and timestamps. The main objective is to reveal the collective wisdom recorded in the seemingly isolated photos and the individual travel sequences reflected by the geo-tagged photos. Instead of focusing on mining frequent trajectory patterns from geo-tagged social media, we put more effort into ranking the mined trajectory patterns and diversifying the ranking results. Through leveraging the relationships among users, locations and trajectories, we rank the trajectory patterns. We then use an exemplar-based algorithm to diversify the results in order to discover the representative trajectory patterns. We have evaluated the proposed framework on 12 different cities using a Flickr dataset and demonstrated its effectiveness. Zhijun Yin, Liangliang Cao, Jiawei Han 0001, Jiebo Luo 0001, Thomas S. Huang |
SDM | 1 |
| 2011 | Geographical topic discovery and comparisonabstractThis paper studies the problem of discovering and comparing geographical topics from GPS-associated documents. GPS-associated documents become popular with the pervasiveness of location-acquisition technologies. For example, in Flickr, the geo-tagged photos are associated with tags and GPS locations. In Twitter, the locations of the tweets can be identified by the GPS locations from smart phones. Many interesting concepts, including cultures, scenes, and product sales, correspond to specialized geographical distributions. In this paper, we are interested in two questions: (1) how to discover different topics of interests that are coherent in geographical regions? (2) how to compare several topics across different geographical locations? To answer these questions, this paper proposes and compares three ways of modeling geographical topics: location-driven model, text-driven model, and a novel joint model called LGTA (Latent Geographical Topic Analysis) that combines location and text. To make a fair comparison, we collect several representative datasets from Flickr website including Landscape, Activity, Manhattan, National park, Festival, Car, and Food. The results show that the first two methods work in some datasets but fail in others. LGTA works well in all these datasets at not only finding regions of interests but also providing effective comparisons of the topics across different locations. The results confirm our hypothesis that the geographical distributions can help modeling topics, while topics provide important cues to group different geographical regions. Zhijun Yin, Liangliang Cao, Jiawei Han 0001, ChengXiang Zhai, Thomas S. Huang |
WWW | 1 |
| 2010 | A Unified Framework for Link Recommendation Using Random WalksabstractThe phenomenal success of social networking sites, such as Facebook, Twitter and LinkedIn, has revolutionized the way people communicate. This paradigm has attracted the attention of researchers that wish to study the corresponding social and technological problems. Link recommendation is a critical task that not only helps increase the linkage inside the network and also improves the user experience. In an effective link recommendation algorithm it is essential to identify the factors that influence link creation. This paper enumerates several of these intuitive criteria and proposes an approach which satisfies these factors. This approach estimates link relevance by using random walk algorithm on an augmented social graph with both attribute and structure information. The global and local influences of the attributes are leveraged in the framework as well. Other than link recommendation, our framework can also rank the attributes in the network. Experiments on DBLP and IMDB data sets demonstrate that our method outperforms state-of-the-art methods for link recommendation. Zhijun Yin, Manish Gupta 0001, Tim Weninger, Jiawei Han 0001 |
ASONAM | 1 |
| 2010 | LINKREC: a unified framework for link recommendation with user attributes and graph structureabstractWith the phenomenal success of networking sites (e.g., Facebook, Twitter and LinkedIn), social networks have drawn substantial attention. On online social networking sites, link recommendation is a critical task that not only helps improve user experience but also plays an essential role in network growth. In this paper we propose several link recommendation criteria, based on both user attributes and graph structure. To discover the candidates that satisfy these criteria, link relevance is estimated using a random walk algorithm on an augmented social graph with both attribute and structure information. The global and local influence of the attributes is leveraged in the framework as well. Besides link recommendation, our framework can also rank attributes in a social network. Experiments on DBLP and IMDB data sets demonstrate that our method outperforms state-of-the-art methods based on network structure and node attribute information for link recommendation. Zhijun Yin, Manish Gupta 0001, Tim Weninger, Jiawei Han 0001 |
WWW | 1 |
| 2009 | Query Expansion Using External Evidence
Zhijun Yin, Milad Shokouhi, Nick Craswell |
ECIR | 1 |
| 2009 | RankClus: integrating clustering with ranking for heterogeneous information network analysisabstractAs information networks become ubiquitous, extracting knowledge from information networks has become an important task. Both ranking and clustering can provide overall views on information network data, and each has been a hot topic by itself. However, ranking objects globally without considering which clusters they belong to often leads to dumb results, e.g., ranking database and computer architecture conferences together may not make much sense. Similarly, clustering a huge number of objects (e.g., thousands of authors) in one huge cluster without distinction is dull as well. Yizhou Sun, Jiawei Han 0001, Peixiang Zhao 0001, Zhijun Yin, Hong Cheng 0001 |
EDBT | 4 |
| 2009 | Exploring social tagging graph for web object classificationabstractThis paper studies web object classification problem with the novel exploration of social tags. Automatically classifying web objects into manageable semantic categories has long been a fundamental preprocess for indexing, browsing, searching, and mining these objects. The explosive growth of heterogeneous web objects, especially non-textual objects such as products, pictures, and videos, has made the problem of web classification increasingly challenging. Such objects often suffer from a lack of easy-extractable features with semantic information, interconnections between each other, as well as training examples with category labels. Zhijun Yin, Rui Li 0049, Qiaozhu Mei, Jiawei Han 0001 |
KDD | 1 |
| 2009 | CLHQS: Hierarchical Query Suggestion by Mining Clickthrough Log
Depin Chen, Ning Liu 0001, Zhijun Yin, Yang Tong, Jun Yan 0001, Zheng Chen 0001 |
PAKDD | 3 |
| 2009 | GAD: General Activity Detection for Fast Clustering on Large DataabstractIn this paper, we propose GAD (General Activity Detection) for fast clustering on large scale data. Within this framework we design a set of algorithms for different scenarios: (1) Exact GAD algorithm E-GAD, which is much faster than K-Means and gets the same clustering result. (2) Approximate GAD algorithms with different assumptions, which are faster than E-GAD while achieving different degrees of approximation. (3) GAD based algorithms to handle the “large clusters” problem which appears in many large scale clustering applications. Two existing activity detection algorithms GT and CGAUTC are special cases under the framework. The most important contribution of our work is that the framework is the general solution to exploit activity detection for fast clustering in both exact and approximate senarios, and our proposed algorithms within the framework can achieve very high speed. Extensive experiments have been conducted on several large datasets from various real world applications; the results show that our proposed algorithms are effective and efficient. Xin Jin 0001, Sangkyum Kim, Jiawei Han 0001, Liangliang Cao, Zhijun Yin |
SDM | 5 |
| 2008 | Sampling cube: a framework for statistical olap over sampling dataabstractSampling is a popular method of data collection when it is impossible or too costly to reach the entire population. For example, television show ratings in the United States are gathered from a sample of roughly 5,000 households. To use the results effectively, the samples are further partitioned in a multidimensional space based on multiple attribute values. This naturally leads to the desirability of OLAP (Online Analytical Processing) over sampling data. However, unlike traditional data, sampling data is inherently uncertain, i.e., not representing the full data in the population. Thus, it is desirable to return not only query results but also the confidence intervals indicating the reliability of the results. Moreover, a certain segment in a multidimensional space may contain none or too few samples. This requires some additional analysis to return trustable results.In this paper we propose a Sampling Cube framework, which efficiently calculates confidence intervals for any multidimensional query and uses the OLAP structure to group similar segments to increase sampling size when needed. Further, to handle high dimensional data, a Sampling Cube Shell method is proposed to effectively reduce the storage requirement while still preserving query result quality. Xiaolei Li 0001, Jiawei Han 0001, Zhijun Yin, Jae-Gil Lee 0001, Yizhou Sun |
SIGMOD Conference | 3 |
| 2008 | BibNetMiner: mining bibliographic information networksabstractOnline bibliographic databases, such as DBLP in computer science and PubMed in medical sciences, contain abundant information about research publications in different fields. Each such database forms a gigantic information network (hence called BibNet), connecting in complex ways research papers, authors, conferences/journals, and possibly citation information as well, and provides a fertile land for information network analysis. Our BibNetMiner is designed for sophisticated information network mining on such bibliographic databases. In this demo, we will take the DBLP database as an example, demonstrate several attractive functions of BibNetMiner, including clustering, ranking and profiling of conferences and authors based on the research subfields. A user-friendly, visualization-enhanced interface will be provided to facilitate interactive exploration of a bibliographic database. This project will serve as an example to demonstrate the power of links in information network mining. Since the dataset is large and the network is heterogeneous, such a study will benefit the research on the analysis of massive heterogeneous information networks. Yizhou Sun, Zhijun Yin, Hong Cheng 0001, Jiawei Han 0001, Xiaoxin Yin, Peixiang Zhao 0001 |
SIGMOD Conference | 3 |
| 2005 | Complexity Estimates for the F4 Attack on the Perturbed Matsumoto-Imai Cryptosystem
Jintai Ding, Jason E. Gower, Dieter Schmidt, Christopher Wolf, Zhijun Yin |
IMACC | 5 |