Tianyi Jiang

dblp:83/2438 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 77% Bioinformatics and computational biology · 23%
Artificial intelligence
1 paper
Speech recognition and synthesis · 100%
Databases, data mining, and information retrieval
5 papers
Data mining · 92% Information retrieval · 8%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Medical and health informatics
brain-computer interface
1.012026
Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired Benchmark · ACL (1) 2026
Data mining › clustering › user clustering
customer segmentation
0.352009
Improving Personalization Solutions through Optimal Segmentation of Customer Bases · IEEE Trans. Knowl. Data Eng. 2009
Dynamic Micro Targeting: Fitness-Based Approach to Predicting Individual Preferences · ICDM 2007
Segmenting Customers from Population to Individuals: Does 1-to-1 Keep Your Customers Forever? · IEEE Trans. Knowl. Data Eng. 2006
Data mining › behavior modeling
customer behavior modeling
0.232007
Dynamic Micro Targeting: Fitness-Based Approach to Predicting Individual Preferences · ICDM 2007
Improving Personalization Solutions through Optimal Segmentation of Customer Bases · ICDM 2006
Divide and Prosper: Comparing Models of Customer Behavior From Populations to Individuals · ICDM 2004
Mathematical optimization
combinatorial optimization
0.222009
Improving Personalization Solutions through Optimal Segmentation of Customer Bases · IEEE Trans. Knowl. Data Eng. 2009
Improving Personalization Solutions through Optimal Segmentation of Customer Bases · ICDM 2006
Data mining › predictive modeling
classification
0.112006
Segmenting Customers from Population to Individuals: Does 1-to-1 Keep Your Customers Forever? · IEEE Trans. Knowl. Data Eng. 2006
Information retrieval › user interaction
personalization
0.132007
Dynamic Micro Targeting: Fitness-Based Approach to Predicting Individual Preferences · ICDM 2007
Improving Personalization Solutions through Optimal Segmentation of Customer Bases · ICDM 2006
Divide and Prosper: Comparing Models of Customer Behavior From Populations to Individuals · ICDM 2004
Data mining
predictive modeling
0.012004
Divide and Prosper: Comparing Models of Customer Behavior From Populations to Individuals · ICDM 2004
Electronic design automation › logic synthesis
FPGA synthesis
0.012004
High level area, delay and power estimation for FPGAs · FPGA 2004

Methods — techniques the papers use, named apart from their topics

teacher-forcing-free evaluation · 2.0high-density EEG · 2.0direct grouping · 0.2hierarchical clustering · 0.2combinatorial optimization · 0.2affinity propagation · 0.2clustering · 0.2predictive modeling · 0.1classifier comparison · 0.1segmentation techniques · 0.0power macro-modeling · 0.0design space exploration · 0.0classifier · 0.0
YearPublicationVenuePosition
2026 Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired Benchmark
abstract
Translating brain signals into text could restore communication for people with severe paralysis, yet practically usable systems to date rely on invasive electrocorticography (ECoG).Electroencephalography (EEG) offers a non-invasive alternative, and EEG-to-text (EEG2Text) has been widely explored.Interestingly, however, EEG2Text models generally rely on teacher-forcing evaluation; without it, they fail to generate meaningful decoding.This reliance prevents EEG2Text from being applied in real-world, non-academic settings.This has fueled numerous debates about whether EEG2Text is a meaningful direction, by extension, and whether EEG truly contains decodable linguistic information.Here, using a neuropsychology-informed paradigm, we find that existing EEG2Text benchmarks have neglected EEG instability, a flaw that has confounded inference and sparked debate.Our experiments furnish key evidence for the feasibility of teacher-forcing-free EEG2Text decoding.Accordingly, we assemble the Corpus OF Eeg-To-Text (COFETT) using a 128-channel highdensity EEG cap, providing a benchmark dedicated to evaluating EEG2Text models.In comparisons with multiple existing benchmarks, COFETT achieves SOTA ability to distinguish among model performances and enables robust, teacher-forcing-free evaluation, thereby opening a path toward practical EEG2Text applications.COFETT is open sourced in https: //github.com/baoyudu/COFETT.
Tianyi Jiang, Kai Xiong 0002
ACL (1)4
2026 KScaNN: Scalable Approximate Nearest Neighbor Search on Kunpeng
abstract
Approximate Nearest Neighbor Search (ANNS) is a cornerstone algorithm for information retrieval, recommendation systems, and machine learning applications. While x86-based architectures have historically dominated this domain, the increasing adoption of ARM-based servers in industry presents a critical need for ANNS solutions optimized on ARM architectures. A naive port of existing x86 ANNS algorithms to ARM platforms results in a substantial performance deficit, failing to leverage the unique capabilities of the underlying hardware. To address this challenge, we introduce KScaNN, a novel ANNS algorithm co-designed for the Kunpeng 920 ARM architecture. KScaNN embodies a holistic approach that synergizes sophisticated, data aware algorithmic refinements with carefully-designed hardware specific optimizations. Its core contributions include: 1) novel algorithmic techniques, including a hybrid intra-cluster search strategy and an improved PQ residual calculation method, which optimize the search process at a higher level; 2) an ML-driven adaptive search module that provides adaptive, per-query tuning of search parameters, eliminating the inefficiencies of static configurations; and 3) highly-optimized SIMD kernels for ARM that maximize hardware utilization for the critical distance computation workloads. The experimental results demonstrate that KScaNN not only closes the performance gap but establishes a new standard, achieving up to a 1.63x speedup over the fastest x86-based solution. This work provides a definitive blueprint for achieving leadership-class performance for vector search on modern ARM architectures and underscores
Oleg Senkevich, Siyang Xu, Tianyi Jiang, Alexander Radionov, Jan Tabaszewski, Dmitriy Malyshev, Daihao Xue, Licheng Yu, Weidi Zeng, Xin Yao 0008, Siyu Huang, Gleb Neshchetkin, Qiuling Pan, Yaoyao Fu
ICDE3
2025 Knowledge-enhanced Relation Graph and Task Sampling for few-shot molecular property prediction
Zeyu Wang 0011, Tianyi Jiang, Yao Lu 0041, Xiaoze Bao, Shanqing Yu, Qi Xuan 0001
Inf. Sci.2
2025 Molecular Connectivity Index-Based Data Augmentation for Molecular Property Prediction
abstract
Recent years have seen a rapid growth of machine learning in cheminformatics problems. In order to tackle the problem of insufficient training data in reality, more and more researchers pay attention to data augmentation technology. However, few researchers pay attention to the problem of construction rules and domain information of data, which will directly impact the quality of augmented data and the augmentation performance. While in graph-based molecular research, the molecular connectivity index, as a critical topological index, can directly or indirectly reflect the topology-based physicochemical properties and biological activities. In this paper, we propose a novel data augmentation technique that modifies the topology of the molecular graph to generate augmented data with the same molecular connectivity index as the original data. The molecular connectivity index combined with data augmentation technology helps to retain more topology-based molecular properties information and generate more reliable data. Furthermore, we adopt five benchmark datasets to test our proposed models, and the results indicate that the augmented data generated based on important molecular topology features can effectively improve the prediction accuracy of molecular properties, which also provides a new perspective on data augmentation in cheminformatics studies.
Zeyu Wang 0011, Tianyi Jiang, Jinhuan Wang, Jiafei Shao, Qi Xuan 0001
IEEE Trans. Comput. Biol. Bioinform.2
2024 Mix-Key: graph mixup with key structures for molecular property prediction
abstract
Molecular property prediction faces the challenge of limited labeled data as it necessitates a series of specialized experiments to annotate target molecules. Data augmentation techniques can effectively address the issue of data scarcity. In recent years, Mixup has achieved significant success in traditional domains such as image processing. However, its application in molecular property prediction is relatively limited due to the irregular, non-Euclidean nature of graphs and the fact that minor variations in molecular structures can lead to alterations in their properties. To address these challenges, we propose a novel data augmentation method called Mix-Key tailored for molecular property prediction. Mix-Key aims to capture crucial features of molecular graphs, focusing separately on the molecular scaffolds and functional groups. By generating isomers that are relatively invariant to the scaffolds or functional groups, we effectively preserve the core information of molecules. Additionally, to capture interactive information between the scaffolds and functional groups while ensuring correlation between the original and augmented graphs, we introduce molecular fingerprint similarity and node similarity. Through these steps, Mix-Key determines the mixup ratio between the original graph and two isomers, thus generating more informative augmented molecular graphs. We extensively validate our approach on molecular datasets of different scales with several Graph Neural Network architectures. The results demonstrate that Mix-Key consistently outperforms other data augmentation methods in enhancing molecular property prediction on several datasets.
Tianyi Jiang, Zeyu Wang 0011, Wenchao Yu, Jinhuan Wang, Shanqing Yu, Xiaoze Bao, Qi Xuan 0001
Briefings Bioinform.1
2017 Increasing physics realism when evolving micro behaviors for 3D RTS games
abstract
We attack the problem of evolving high performance micro behaviors in 3D RTS-game-like simulations. Prior work had shown the potential for the Meta-Search approach to evolve high performance micro for RTS games like StarCraft. We extend this work by moving to 3D and by moving to more realistic physics for simulating the movement of entities in our RTS-game-like simulation. We compare the evolved micro performance of our entities with different physics models of motion on the same scenarios against identical opponent units in a 3D RTS simulation. Results show that our genetic algorithm approach works to reliably evolve high quality 3D micro behaviors for entities independent of the physics model used. Furthermore, experiments show that the entity's acceleration has more of an effect on performance than rotation speed. Our work provides evidence for the generalizability of an evolutionary approach to generating complex behavior for 3D RTS games, training simulations, and real-world unmanned vehicles.
Siming Liu 0001, Sushil J. Louis, Tianyi Jiang, Rui Wu 0003
CEC3
2009 Dynamic micro-targeting: fitness-based approach to predicting individual preferences
Tianyi Jiang, Alexander Tuzhilin
Knowl. Inf. Syst.1
2009 Improving Personalization Solutions through Optimal Segmentation of Customer Bases
abstract
On the Web, where the search costs are low and the competition is just a mouse click away, it is crucial to segment the customers intelligently in order to offer more targeted and personalized products and services to them. Traditionally, customer segmentation is achieved using statistics-based methods that compute a set of statistics from the customer data and group customers into segments by applying distance-based clustering algorithms in the space of these statistics. In this paper, we present a direct grouping-based approach to computing customer segments that groups customers not based on computed statistics, but in terms of optimally combining transactional data of several customers to build a data mining model of customer behavior for each group. Then, building customer segments becomes a combinatorial optimization problem of finding the best partitioning of the customer base into disjoint groups. This paper shows that finding an optimal customer partition is NP-hard, proposes several suboptimal direct grouping segmentation methods, and empirically compares them among themselves, traditional statistics-based hierarchical and affinity propagation-based segmentation, and one-to-one methods across multiple experimental conditions. It is shown that the best direct grouping method significantly dominates the statistics-based and one-to-one approaches across most of the experimental conditions, while still being computationally tractable. It is also shown that the distribution of the sizes of customer segments generated by the best direct grouping method follows a power law distribution and that microsegmentation provides the best approach to personalization.
Tianyi Jiang, Alexander Tuzhilin
IEEE Trans. Knowl. Data Eng.1
2007 Dynamic Micro Targeting: Fitness-Based Approach to Predicting Individual Preferences
abstract
It is crucial to segment customers intelligently in order to offer more targeted and personalized products and services. Traditionally, customer segmentation is achieved using statistics-based methods that compute a set of statistics from the customer data and group customers into segments by applying clustering algorithms. Recent research proposed a direct grouping-based approach that combines customers into segments by optimally combining transactional data of several customers and building a data mining model of customer behavior for each group. This paper proposes a new micro targeting method that builds predictive models of customer behavior not on the segments of customers but rather on the customer-product groups. This micro-targeting method is more general than the previously considered direct grouping method. We empirically show that it significantly outperforms the direct grouping and statistics-based segmentation methods across multiple experimental conditions and that it generates predominately small-sized segments, thus providing additional support for the micro-targeting approach to personalization.
Tianyi Jiang, Alexander Tuzhilin
ICDM1
2006 Improving Personalization Solutions through Optimal Segmentation of Customer Bases
abstract
On the Web, where the search costs are low and the competition is just a mouse click away, it is crucial to segment the customers intelligently in order to offer more targeted and personalized products and services to them. Traditionally, customer segmentation is achieved using statistics-based methods that compute a set of statistics from the customer data and group customers into segments by applying distance-based clustering algorithms in the space of these statistics. In this paper, we present a direct grouping based approach to computing customer segments that groups customers not based on computed statistics, but in terms of optimally combining transactional data of several customers to build a data mining model of customer behavior for each group. Then building customer segments becomes a combinatorial optimization problem of finding the best partitioning of the customer base into disjoint groups. The paper shows that finding an optimal customer partition is NP-hard, proposes a suboptimal direct grouping segmentation method and empirically compares it against traditional statistics-based segmentation and 1-to-1 methods across multiple experimental conditions. We show that the direct grouping method significantly dominates the statistics-based and 1-to-1 approaches across all the experimental conditions, while still being computationally tractable. We also show that there are very few size-one customer segments generated by the best direct grouping method and that micro-segmentation provides the best approach to personalization.
Tianyi Jiang, Alexander Tuzhilin
ICDM1
2006 Segmenting Customers from Population to Individuals: Does 1-to-1 Keep Your Customers Forever?
abstract
There have been various claims made in the marketing community about the benefits of 1-to-1 marketing versus traditional customer segmentation approaches and how much they can improve understanding of customer behavior. However, few rigorous studies exist that systematically compare these approaches. In this paper, we conducted such a study and compared the predictive performance of aggregate, segmentation, and 1-to-1 marketing approaches across a broad range of experimental settings, such as multiple segmentation levels, multiple real-world marketing data sets, multiple dependent variables, different types of classifiers, different segmentation techniques, and different predictive measures. Our experiments show that both 1-to-1 and segmentation approaches significantly outperform aggregate modeling. Reaffirming anecdotal evidence of the benefits of 1-to-1 marketing, our experiments show that the 1-to-1 approach also dominates the segmentation approach for the frequently transacting customers. However, our experiments also show that segmentation models taken at the best granularity levels dominate 1-to-1 models when modeling customers with little transactional data using effective clustering methods. In addition, the peak performance of segmentation models are reached at the finest granularity levels, skewed towards the 1-to-1 case. This finding adds support for the microsegmentation approach and suggests that 1-to-1 marketing may not always be the best solution
Tianyi Jiang, Alexander Tuzhilin
IEEE Trans. Knowl. Data Eng.1
2004 High level area, delay and power estimation for FPGAs
abstract
This paper describes an approach for high-level estimation of area, delay and power for FPGA synthesis. This approach has been integrated within the PACT compiler framework which has an automated design space exploration pass that determines the effects of various compiler optimizations on the synthesized hardware. Such a pass needs early estimation of area, delay and power. Towards this end, we have developed area and delay models for various RTL level operators such as adders, multipliers, and logical operators, which are parameterized with the bit widths of the devices. We have also derived high-level equation based power macro-models which take into account input switching activities, input spatial correlation and input bit width. These models are derived by actual synthesis of the RTL operators using back-end logic synthesis and place-and-route tools. Experimental results show that these area, delay and power models are accurate and efficient.
Tianyi Jiang, Xiaoyong Tang, Prithviraj Banerjee
FPGA1
2004 Macro-models for high level area and power estimation on FPGAs
abstract
As more and more complex applications are implemented on FPGAs, high-level design tools are needed to reduce the design time. A good high-level synthesis tool usually has an automated design space exploration pass to determine the effects of various compiler optimizations on the area and power of the synthesized hardware. Such a pass needs early estimation of area and power. Towards this end, we have developed high-level equation based area and power macro-models for various RTL level operators such as adders, multipliers, and logical operators. The area model is parameterized with the bit width of the device and the power model takes into account input switching activity and input spatial correlation as well as input bit width. These models are derived by actual synthesis of these RTL operators using back-end logic synthesis and place-and-route tools. Compared with the other approaches, our method generated a uniform macro-model for each operator with fewer coefficients and sometimes lower degrees. It is also easier to analyze the power sensitivity to different parameters. Experimental results show that these area and power models are accurate and efficient.
Tianyi Jiang, Xiaoyong Tang, Prithviraj Banerjee
ACM Great Lakes Symposium on VLSI1
2004 Divide and Prosper: Comparing Models of Customer Behavior From Populations to Individuals
abstract
This paper compares customer segmentation, 1-to-1, and aggregate marketing approaches across a broad range of experimental settings, including multiple segmentation levels, marketing datasets, dependent variables, and different types of classifiers, segmentation techniques, and predictive measures. Our experimental results show that, overall, 1-to-1 modeling significantly outperforms the aggregate approach among high-volume customers and is never worse than aggregate approach among low-volume customers. Moreover, the best segmentation techniques tend to outperform 1-to-l modeling among low-volume customers.
Tianyi Jiang, Alexander Tuzhilin
ICDM1