Zi Yin

dblp:123/3786 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-authorComputer networks · 2Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 51% Learning theory · 17% Transfer learning and domain adaptation · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 75% Query processing and optimization · 25%
Computer networks
2 papers
Internet of things and sensor networks · 61% Network measurement and analytics · 21% Wireless networking · 18%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.722018
The Global Anchor Method for Quantifying Linguistic Shifts and Domain Adaptation · NeurIPS 2018
On the Dimensionality of Word Embedding · NeurIPS 2018
Machine learning › Learning theory › statistical learning theory
bias-variance tradeoff
0.312018
On the Dimensionality of Word Embedding · NeurIPS 2018
Machine learning › Transfer learning and domain adaptation › domain shift
domain shift detection
0.312018
The Global Anchor Method for Quantifying Linguistic Shifts and Domain Adaptation · NeurIPS 2018
Machine learning › Representation and self-supervised learning › representation matching › feature alignment
embedding alignment
0.312018
The Global Anchor Method for Quantifying Linguistic Shifts and Domain Adaptation · NeurIPS 2018
Internet of things and sensor networks
time synchronization
0.312018
Exploiting a Natural Network Effect for Scalable, Fine-grained Clock Synchronization · NSDI 2018
Natural language and speech › Question answering and dialogue systems › conversational agents
chatbot
0.312017
DeepProbe: Information Directed Sequence Understanding and Chatbot Design via Recurrent Neural Networks · KDD 2017
Query processing and optimization
query rewriting
0.312017
DeepProbe: Information Directed Sequence Understanding and Chatbot Design via Recurrent Neural Networks · KDD 2017
Information retrieval
query understanding
0.312017
DeepProbe: Information Directed Sequence Understanding and Chatbot Design via Recurrent Neural Networks · KDD 2017
Information retrieval › ranking › search ranking
relevance ranking
0.312017
DeepProbe: Information Directed Sequence Understanding and Chatbot Design via Recurrent Neural Networks · KDD 2017
Information retrieval › ranking › relevance estimation
relevance scoring
0.312017
DeepProbe: Information Directed Sequence Understanding and Chatbot Design via Recurrent Neural Networks · KDD 2017
Computational social science and digital humanities
language evolution
0.112018
The Global Anchor Method for Quantifying Linguistic Shifts and Domain Adaptation · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

graph laplacian · 0.7global anchor method · 0.7sequence-to-sequence · 0.6recurrent neural network · 0.6entropy · 0.6attention · 0.6matrix perturbation theory · 0.3
YearPublicationVenuePosition
2019 SIMON: A Simple and Scalable Method for Sensing, Inference and Measurement in Data Center Networks
Yilong Geng, Zi Yin, Ashish Naik, Balaji Prabhakar, Mendel Rosenblum, Amin Vahdat
NSDI3
2018 On the Dimensionality of Word Embedding
abstract
In this paper, we provide a theoretical understanding of word embedding and its dimensionality. Motivated by the unitary-invariance of word embedding, we propose the Pairwise Inner Product (PIP) loss, a novel metric on the dissimilarity between word embeddings. Using techniques from matrix perturbation theory, we reveal a fundamental bias-variance trade-off in dimensionality selection for word embeddings. This bias-variance trade-off sheds light on many empirical observations which were previously unexplained, for example the existence of an optimal dimensionality. Moreover, new insights and discoveries, like when and how word embeddings are robust to over-fitting, are revealed. By optimizing over the bias-variance trade-off of the PIP loss, we can explicitly answer the open question of dimensionality selection for word embedding.
Zi Yin
NeurIPS1
2018 The Global Anchor Method for Quantifying Linguistic Shifts and Domain Adaptation
abstract
Language is dynamic, constantly evolving and adapting with respect to time, domain or topic. The adaptability of language is an active research area, where researchers discover social, cultural and domain-specific changes in language using distributional tools such as word embeddings. In this paper, we introduce the global anchor method for detecting corpus-level language shifts. We show both theoretically and empirically that the global anchor method is equivalent to the alignment method, a widely-used method for comparing word embeddings, in terms of detecting corpus-level language shifts. Despite their equivalence in terms of detection abilities, we demonstrate that the global anchor method is superior in terms of applicability as it can compare embeddings of different dimensionalities. Furthermore, the global anchor method has implementation and parallelization advantages. We show that the global anchor method reveals fine structures in the evolution of language and domain adaptation. When combined with the graph Laplacian technique, the global anchor method recovers the evolution trajectory and domain clustering of disparate text corpora.
Zi Yin, Vin Sachidananda, Balaji Prabhakar
NeurIPS1
2018 Exploiting a Natural Network Effect for Scalable, Fine-grained Clock Synchronization
Yilong Geng, Zi Yin, Ashish Naik, Balaji Prabhakar, Mendel Rosenblum, Amin Vahdat
NSDI3
2017 DeepProbe: Information Directed Sequence Understanding and Chatbot Design via Recurrent Neural Networks
abstract
Information extraction and user intention identification is a central topic in modern query understanding and recommendation systems. In this paper, we propose DeepProbe, a generic information-directed interaction framework which is built around an attention-based sequence to sequence (seq2seq) recurrent neural network. DeepProbe can rephrase, evaluate, and even actively ask questions, leveraging the generative ability and likelihood estimation made possible by seq2seq models. DeepProbe makes decisions based on a derived uncertainty (entropy) measure conditioned on user inputs, possibly with multiple rounds of interactions. Three applications, namely a rewritter, a relevance scorer and a chatbot for ad recommendation, were built around DeepProbe, with the first two serving as precursory building blocks for the third. We first use the seq2seq model in DeepProbe to rewrite a user query into one of standard query form, which is submitted to an ordinary recommendation system. Secondly, we evaluate DeepProbe's seq2seq model-based relevance scoring. Finally, we build a chatbot prototype capable of making active user interactions, which can ask questions that maximize information gain, allowing for a more efficient user intention idenfication process. We evaluate first two applications by 1) comparing with baselines by BLEU and AUC, and 2) human judge evaluation. Both demonstrate significant improvements compared with current state-of-the-art systems, proving their values as useful tools on their own, and at the same time laying a good foundation for the ongoing chatbot application.
Zi Yin, Keng-hao Chang, Ruofei Zhang
KDD1