Yingtao Tian

dblp:180/5335 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
8since 2021 · last 2024
0000-0002-3602-259XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Optimization for machine learning · 23% Reinforcement learning · 22% Language models and text generation · 18%
Databases, data mining, and information retrieval
5 papers
Knowledge graphs · 50% Data mining · 28% Data models and query languages · 14%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 50% Environmental and earth informatics · 50%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
black-box optimization
0.812024
Position: Leverage Foundational Models for Black-Box Optimization · ICML 2024
Natural language and speech › Language models and text generation
large language model
0.812024
Position: Leverage Foundational Models for Black-Box Optimization · ICML 2024
Machine learning › Optimization for machine learning
evolutionary computation
0.712023
NeuroEvoBench: Benchmarking Evolutionary Optimizers for Deep Learning Applications · NeurIPS 2023
Machine learning › Reinforcement learning
exploration
0.712023
DEIR: Efficient and Robust Exploration through Discriminative-Model-Based Episodic Intrinsic Rewards · IJCAI 2023
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.712023
DEIR: Efficient and Robust Exploration through Discriminative-Model-Based Episodic Intrinsic Rewards · IJCAI 2023
Computer vision › Video understanding and tracking
spatio-temporal modeling
0.712023
Digital Typhoon: Long-term Satellite Image Dataset for the Spatio-Temporal Modeling of Tropical Cyclones · NeurIPS 2023
Computational social science and digital humanities
digital history
0.712023
MingOfficial: A Ming Official Career Dataset and a Historical Context-Aware Representation Learning Framework · EMNLP 2023
Environmental and earth informatics › meteorology
meteorological analysis
0.712023
Digital Typhoon: Long-term Satellite Image Dataset for the Spatio-Temporal Modeling of Tropical Cyclones · NeurIPS 2023
Data mining › representation learning
graph representation learning
0.712023
MingOfficial: A Ming Official Career Dataset and a Historical Context-Aware Representation Learning Framework · EMNLP 2023
Knowledge graphs
knowledge graph embedding
0.622018
Co-training Embeddings of Knowledge Graphs and Entity Descriptions for Cross-lingual Entity Alignment · IJCAI 2018
Multilingual Knowledge Graph Embeddings for Cross-lingual Knowledge Alignment · IJCAI 2017
Natural language and speech › Question answering and dialogue systems
question understanding
0.412020
A Natural Language Interface for Database: Achieving Transfer-learnability Using Adversarial Method for Question Understanding · ICDE 2020
Data models and query languages › natural language interface
natural language interface to database
0.412020
A Natural Language Interface for Database: Achieving Transfer-learnability Using Adversarial Method for Question Understanding · ICDE 2020
Machine learning › Graph learning › limited supervision › multi-view semi-supervised learning
co-training
0.312018
Co-training Embeddings of Knowledge Graphs and Entity Descriptions for Cross-lingual Entity Alignment · IJCAI 2018
Natural language and speech › Language models and text generation › controllable text generation
grammar-guided generation
0.312018
Syntax-Directed Variational Autoencoder for Structured Data · ICLR (Poster) 2018
Machine learning › Representation and self-supervised learning
structured representation
0.312018
Syntax-Directed Variational Autoencoder for Structured Data · ICLR (Poster) 2018
Machine learning › Generative modeling
variational autoencoder
0.312018
Syntax-Directed Variational Autoencoder for Structured Data · ICLR (Poster) 2018
Knowledge graphs › knowledge graph alignment › entity alignment
cross-lingual entity alignment
0.312018
Co-training Embeddings of Knowledge Graphs and Entity Descriptions for Cross-lingual Entity Alignment · IJCAI 2018
Knowledge graphs › knowledge graph alignment
cross-lingual knowledge linking
0.312017
Multilingual Knowledge Graph Embeddings for Cross-lingual Knowledge Alignment · IJCAI 2017
Knowledge graphs
link prediction
0.312017
Multilingual Knowledge Graph Embeddings for Cross-lingual Knowledge Alignment · IJCAI 2017
Information retrieval › evaluation
benchmark dataset
0.212023
Digital Typhoon: Long-term Satellite Image Dataset for the Spatio-Temporal Modeling of Tropical Cyclones · NeurIPS 2023
Data mining
clustering
0.212023
MingOfficial: A Ming Official Career Dataset and a Historical Context-Aware Representation Learning Framework · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

satellite image processing · 2.0inter-satellite calibration · 2.0graph neural network · 1.3transformer · 0.8sequence model · 0.8performance prediction · 0.8mutual information · 0.7meta-learning · 0.7evolutionary optimization · 0.7discriminative forward model · 0.7neural network classifier · 0.4deep sequence model · 0.4adversarial methods · 0.4adversarial method · 0.4entity description embedding · 0.3translation-based embedding · 0.3linear transformation · 0.3axis calibration · 0.3
YearPublicationVenuePosition
2024 DiffCJK: Conditional Diffusion Model for High-Quality and Wide-coverage CJK Character Generation
Yingtao Tian
ICCC1
2024 Position: Leverage Foundational Models for Black-Box Optimization
abstract
Undeniably, Large Language Models (LLMs) have stirred an extraordinary wave of innovation in the machine learning research domain, resulting in substantial impact across diverse fields such as reinforcement learning, robotics, and computer vision. Their incorporation has been rapid and transformative, marking a significant paradigm shift in the field of machine learning research. However, the field of experimental design, grounded on black-box optimization, has been much less affected by such a paradigm shift, even though integrating LLMs with optimization presents a unique landscape ripe for exploration. In this position paper, we frame the field of black-box optimization around sequence-based foundation models and organize their relationship with previous literature. We discuss the most promising ways foundational language models can revolutionize optimization, which include harnessing the vast wealth of information encapsulated in free-form text to enrich task comprehension, utilizing highly flexible sequence models such as Transformers to engineer superior optimization strategies, and enhancing performance prediction over previously unseen search spaces.
Xingyou Song, Yingtao Tian, Robert T. Lange, Chansoo Lee, Yujin Tang, Yutian Chen 0001
ICML2
2023 MingOfficial: A Ming Official Career Dataset and a Historical Context-Aware Representation Learning Framework
abstract
In Chinese studies, understanding the nuanced traits of historical figures, often not explicitly evident in biographical data, has been a key interest.However, identifying these traits can be challenging due to the need for domain expertise, specialist knowledge, and context-specific insights, making the eprocess time-consuming and difficult to scale.Our focus on studying officials from China's Ming Dynasty is no exception.To tackle this challenge, we propose MingOfficial, a large-scale multi-modal dataset consisting of both structured (career records, annotated personnel types) and text (historical texts) data for 13, 031 officials.We further couple the dataset with a graph neural network (GNN) to combine both modalities in order to allow investigation of social structures and provide features to boost down-stream tasks.Experiments show that our proposed MingOfficial could enable exploratory analysis of official identities, and also significantly boost performance in tasks such as identifying nuance identities (e.g.civil officials holding military power) from 24.6% to 98.2% F 1 score in holdout test set.By making MingOfficial publicly available at https://data.depositar.io/ en/dataset/ming_official as both a dataset and an interactive tool, we aim to stimulate further research into the role of social context and representation learning in identifying individual characteristics, and hope to provide inspiration for computational approaches in other fields beyond Chinese studies.
You-Jun Chen, Hsin-Yi Hsieh, Yingtao Tian, Bert Chan, Yu-Sin Liu, Richard Tzong-Han Tsai
EMNLP4
2023 DEIR: Efficient and Robust Exploration through Discriminative-Model-Based Episodic Intrinsic Rewards
abstract
Exploration is a fundamental aspect of reinforcement learning (RL), and its effectiveness is a deciding factor in the performance of RL algorithms, especially when facing sparse extrinsic rewards. Recent studies have shown the effectiveness of encouraging exploration with intrinsic rewards estimated from novelties in observations. However, there is a gap between the novelty of an observation and an exploration, as both the stochasticity in the environment and the agent's behavior may affect the observation. To evaluate exploratory behaviors accurately, we propose DEIR, a novel method in which we theoretically derive an intrinsic reward with a conditional mutual information term that principally scales with the novelty contributed by agent explorations, and then implement the reward with a discriminative forward model. Extensive experiments on both standard and advanced exploration tasks in MiniGrid show that DEIR quickly learns a better policy than the baselines. Our evaluations on ProcGen demonstrate both the generalization capability and the general applicability of our intrinsic reward.
Shanchuan Wan, Yujin Tang, Yingtao Tian, Tomoyuki Kaneko
IJCAI3
2023 Digital Typhoon: Long-term Satellite Image Dataset for the Spatio-Temporal Modeling of Tropical Cyclones
abstract
This paper presents the official release of the Digital Typhoon dataset, the longest typhoon satellite image dataset for 40+ years aimed at benchmarking machine learning models for long-term spatio-temporal data. To build the dataset, we developed a workflow to create an infrared typhoon-centered image for cropping using Lambert azimuthal equal-area projection referring to the best track data. We also address data quality issues such as inter-satellite calibration to create a homogeneous dataset. To take advantage of the dataset, we organized machine learning tasks by the types and targets of inference, with other tasks for meteorological analysis, societal impact, and climate change. The benchmarking results on the analysis, forecasting, and reanalysis for the intensity suggest that the dataset is challenging for recent deep learning models, due to many choices that affect the performance of various models. This dataset reduces the barrier for machine learning researchers to meet large-scale real-world events called tropical cyclones and develop machine learning models that may contribute to advancing scientific knowledge on tropical cyclones as well as solving societal and sustainability issues such as disaster reduction and climate change. The dataset is publicly available at http://agora.ex.nii.ac.jp/digital-typhoon/dataset/ and https://github.com/kitamoto-lab/digital-typhoon/.
Asanobu Kitamoto, Jared Hwang, Bastien Vuillod, Lucas Gautier, Yingtao Tian, Tarin Clanuwat
NeurIPS5
2023 NeuroEvoBench: Benchmarking Evolutionary Optimizers for Deep Learning Applications
abstract
Recently, the Deep Learning community has become interested in evolutionary optimization (EO) as a means to address hard optimization problems, e.g. meta-learning through long inner loop unrolls or optimizing non-differentiable operators. One core reason for this trend has been the recent innovation in hardware acceleration and compatible software -- making distributed population evaluations much easier than before. Unlike for gradient descent-based methods though, there is a lack of hyperparameter understanding and best practices for EO – arguably due to severely less `graduate student descent' and benchmarking being performed for EO methods. Additionally, classical benchmarks from the evolutionary community provide few practical insights for Deep Learning applications. This poses challenges for newcomers to hardware-accelerated EO and hinders significant adoption. Hence, we establish a new benchmark of EO methods (NEB) tailored toward Deep Learning applications and exhaustively evaluate traditional and meta-learned EO. We investigate core scientific questions including resource allocation, fitness shaping, normalization, regularization & scalability of EO. The benchmark is open-sourced at https://github.com/neuroevobench/neuroevobench under Apache-2.0 license.
Robert T. Lange, Yujin Tang, Yingtao Tian
NeurIPS3
2022 Simultaneous Multiple-Prompt Guided Generation Using Differentiable Optimal Transport
Yingtao Tian, Marco Cuturi, David Ha
ICCC1
2021 Ukiyo-e Analysis and Creativity with Attribute and Geometry Annotation
Yingtao Tian, Tarin Clanuwat, Chikahiko Suzuki, Asanobu Kitamoto
ICCC1
2020 KaoKore: A Pre-modern Japanese Art Facial Expression Dataset
Yingtao Tian, Chikahiko Suzuki, Tarin Clanuwat, Mikel Bober-Irizar, Alex Lamb, Asanobu Kitamoto
ICCC1
2020 A Natural Language Interface for Database: Achieving Transfer-learnability Using Adversarial Method for Question Understanding
abstract
Relational database management systems (RDBMSs) are powerful because they are able to optimize and execute queries against relational databases. However, when it comes to NLIDB (natural language interface for databases), the entire system is often custom-made for a particular database. Overcoming the complexity and expressiveness of natural languages so that a single NLI can support a variety of databases is an unsolved problem. In this work, we show that it is possible to separate data specific components from latent semantic structures in expressing relational queries in a natural language. With the separation, transferring an NLI from one database to another becomes possible. We develop a neural network classifier to detect data specific components and an adversarial mechanism to locate them in a natural language question. We then introduce a general purpose transfer-learnable NLI that focuses on the latent semantic structure. We devise a deep sequence model that translates the latent semantic structure to an SQL query. Experiments show that our approach outperforms previous NLI methods on the WikiSQL [49] dataset, and the model we learned can be applied to other benchmark datasets without retraining.
Wenlu Wang, Yingtao Tian, Haixun Wang, Wei-Shinn Ku
ICDE2
2019 Fast and Accurate Network Embeddings via Very Sparse Random Projection
abstract
We present FastRP, a scalable and performant algorithm for learning distributed node representations in a graph. FastRP is over 4,000 times faster than state-of-the-art methods such as DeepWalk and node2vec, while achieving comparable or even better performance as evaluated on several real-world networks on various downstream tasks. We observe that most network embedding methods consist of two components: construct a node similarity matrix and then apply dimension reduction techniques to this matrix. We show that the success of these methods should be attributed to the proper construction of this similarity matrix, rather than the dimension reduction method employed. FastRP is proposed as a scalable algorithm for network embeddings. Two key features of FastRP are: 1) it explicitly constructs a node similarity matrix that captures transitive relationships in a graph and normalizes matrix entries based on node degrees; 2) it utilizes very sparse random projection, which is a scalable optimization-free method for dimension reduction. An extra benefit from combining these two design choices is that it allows the iterative computation of node embeddings so that the similarity matrix need not be explicitly constructed, which further speeds up FastRP. FastRP is also advantageous for its ease of implementation, parallelization and hyperparameter tuning. The source code is available at https://github.com/GTmac/FastRP.
Haochen Chen, Syed Fahad Sultan, Yingtao Tian, Muhao Chen 0001, Steven Skiena
CIKM3
2019 Learning to Represent Bilingual Dictionaries
abstract
Bilingual word embeddings have been widely used to capture the correspondence of lexical semantics in different human languages.However, the cross-lingual correspondence between sentences and words is less studied, despite that this correspondence can significantly benefit many applications such as crosslingual semantic search and textual inference.To bridge this gap, we propose a neural embedding model that leverages bilingual dictionaries 1 .The proposed model is trained to map the lexical definitions to the cross-lingual target words, for which we explore with different sentence encoding techniques.To enhance the learning process on limited resources, our model adopts several critical learning strategies, including multi-task learning on different bridges of languages, and joint learning of the dictionary model with a bilingual word embedding model.We conduct experiments on two new tasks.In the cross-lingual reverse dictionary retrieval task, we demonstrate that our model is capable of comprehending bilingual concepts based on descriptions, and the proposed learning strategies are effective.In the bilingual paraphrase identification task, we show that our model effectively associates sentences in different languages via a shared embedding space, and outperforms existing approaches in identifying bilingual paraphrases.
Muhao Chen 0001, Yingtao Tian, Haochen Chen, Kai-Wei Chang 0001, Steven Skiena, Carlo Zaniolo
CoNLL2
2019 Social Relation Inference via Label Propagation
Yingtao Tian, Haochen Chen, Bryan Perozzi, Muhao Chen 0001, Steven Skiena
ECIR (1)1
2019 SpatialNLI: A Spatial Domain Natural Language Interface to Databases Using Spatial Comprehension
abstract
A natural language interface (NLI) to databases is an interface that translates a natural language question to a structured query that is executable by database management systems (DBMS). However, an NLI that is trained in the general domain is hard to apply in the spatial domain due to the idiosyncrasy and expressiveness of the spatial questions. Inspired by the machine comprehension model, we propose a spatial comprehension model that is able to recognize the meaning of spatial entities based on the semantics of the context. The spatial semantics learned from the spatial comprehension model is then injected to the natural language question to ease the burden of capturing the spatial-specific semantics. With our spatial comprehension model and information injection, our NLI for the spatial domain, named SpatialNLI, is able to capture the semantic structure of the question and translate it to the corresponding syntax of an executable query accurately. We also experimentally ascertain that SpatialNLI outperforms state-of-the-art methods.
Wenlu Wang, Wei-Shinn Ku, Yingtao Tian, Haixun Wang
SIGSPATIAL/GIS4
2018 Enhanced Network Embeddings via Exploiting Edge Labels
abstract
Network embedding methods aim at learning low-dimensional latent representation of nodes in a network. While achieving competitive performance on a variety of network inference tasks such as node classification and link prediction, these methods treat the relations between nodes as a binary variable and ignore the rich semantics of edges. In this work, we attempt to learn network embeddings which simultaneously preserve network structure and relations between nodes. Experiments on several real-world networks illustrate that by considering different relations between different node pairs, our method is capable of producing node embeddings of higher quality than a number of state-of-the-art network embedding methods, as evaluated on a challenging multi-label node classification task.
Haochen Chen, Yingtao Tian, Bryan Perozzi, Muhao Chen 0001, Steven Skiena
CIKM3
2018 Simple Neologism Based Domain Independent Models to Predict Year of Authorship
abstract
We present domain independent models to date documents based only on neologism usage patterns. Our models capture patterns of neologism usage over time to date texts, provide insights into temporal locality of word usage over a span of 150 years, and generalize to various domains like News, Fiction, and Non-Fiction with competitive performance. Quite intriguingly, we show that by modeling only the distribution of usage counts over neologisms (the model being agnostic of the particular words themselves), we achieve competitive performance using several orders of magnitude fewer features (only 200 input features) compared to state of the art models some of which use 200K features.
Vivek Kulkarni, Yingtao Tian, Parth Dandiwala, Steven Skiena
COLING2
2018 Syntax-Directed Variational Autoencoder for Structured Data
Hanjun Dai, Yingtao Tian, Bo Dai 0001, Steven Skiena
ICLR (Poster)2
2018 Co-training Embeddings of Knowledge Graphs and Entity Descriptions for Cross-lingual Entity Alignment
abstract
Multilingual knowledge graph (KG) embeddings provide latent semantic representations of entities and structured knowledge with cross-lingual inferences, which benefit various knowledge-driven cross-lingual NLP tasks. However, precisely learning such cross-lingual inferences is usually hindered by the low coverage of entity alignment in many KGs. Since many multilingual KGs also provide literal descriptions of entities, in this paper, we introduce an embedding-based approach which leverages a weakly aligned multilingual KG for semi-supervised cross-lingual learning using entity descriptions. Our approach performs co-training of two embedding models, i.e. a multilingual KG embedding model and a multilingual literal description embedding model. The models are trained on a large Wikipedia-based trilingual dataset where most entity alignment is unknown to training. Experimental results show that the performance of the proposed approach on the entity alignment task improves at each iteration of co-training, and eventually reaches a stage at which it significantly surpasses previous approaches. We also show that our approach has promising abilities for zero-shot entity alignment, and cross-lingual KG completion.
Muhao Chen 0001, Yingtao Tian, Kai-Wei Chang 0001, Steven Skiena, Carlo Zaniolo
IJCAI2
2018 On2Vec: Embedding-based Relation Prediction for Ontology Population
abstract
Populating ontology graphs represents a long-standing problem for the Semantic Web community. Recent advances in translation-based graph embedding methods for populating instance-level knowledge graphs lead to promising new approaching for the ontology population problem. However, unlike instance-level graphs, the majority of relation facts in ontology graphs come with comprehensive semantic relations, which often include the properties of transitivity and symmetry, as well as hierarchical relations. These comprehensive relations are often too complex for existing graph embedding methods, and direct application of such methods is not feasible. Hence, we propose On2Vec, a novel translation-based graph embedding method for ontology population. On2Vec integrates two model components that effectively characterize comprehensive relation facts in ontology graphs. The first is the Component-specific Model that encodes concepts and relations into low-dimensional embedding spaces without a loss of relational properties; the second is the Hierarchy Model that performs focused learning of hierarchical relation facts. Experiments on several well-known ontology graphs demonstrate the promising capabilities of On2Vec in predicting and verifying new relation facts. These promising results also make possible significant improvements in related methods.
Muhao Chen 0001, Yingtao Tian, Xuelu Chen, Zijun Xue, Carlo Zaniolo
SDM2
2017 Multilingual Knowledge Graph Embeddings for Cross-lingual Knowledge Alignment
abstract
Many recent works have demonstrated the benefits of knowledge graph embeddings in completing monolingual knowledge graphs. Inasmuch as related knowledge bases are built in several different languages, achieving cross-lingual knowledge alignment will help people in constructing a coherent knowledge base, and assist machines in dealing with different expressions of entity relationships across diverse human languages. Unfortunately, achieving this highly desirable cross-lingual alignment by human labor is very costly and error-prone. Thus, we propose MTransE, a translation-based model for multilingual knowledge graph embeddings, to provide a simple and automated solution. By encoding entities and relations of each language in a separated embedding space, MTransE provides transitions for each embedding vector to its cross-lingual counterparts in other spaces, while preserving the functionalities of monolingual embeddings. We deploy three different techniques to represent cross-lingual transitions, namely axis calibration, translation vectors, and linear transformations, and derive five variants for MTransE using different loss functions. Our models can be trained on partially aligned graphs, where just a small portion of triples are aligned with their cross-lingual counterparts. The experiments on cross-lingual entity matching and triple-wise alignment verification show promising results, with some variants consistently outperforming others on different tasks. We also explore how MTransE preserves the key properties of its monolingual counterpart.
Muhao Chen 0001, Yingtao Tian, Mohan Yang, Carlo Zaniolo
IJCAI2