Ting-Han Fan

dblp:213/0948 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 40% Language models and text generation · 39% Optimization for machine learning · 14%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › compositional generalization › length generalization
length extrapolation
1.222023
Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis · ACL (1) 2023
KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation · NeurIPS 2022
Machine learning › Deep learning architectures and training
positional encoding
1.222023
Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis · ACL (1) 2023
KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation · NeurIPS 2022
Machine learning › Deep learning architectures and training
transformer
1.222023
Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis · ACL (1) 2023
KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation · NeurIPS 2022
Natural language and speech › Language models and text generation
in-context learning
0.912025
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning? · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning? · NeurIPS 2025
Machine learning › Deep learning architectures and training › positional encoding
relative positional encoding
0.712023
Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis · ACL (1) 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable model
0.612022
Training Discrete Deep Generative Models via Gapped Straight-Through Estimator · ICML 2022
Machine learning › Optimization for machine learning
gradient estimation
0.612022
Training Discrete Deep Generative Models via Gapped Straight-Through Estimator · ICML 2022
Machine learning › Optimization for machine learning › gradient estimation
straight-through estimator
0.612022
Training Discrete Deep Generative Models via Gapped Straight-Through Estimator · ICML 2022
Natural language and speech › Language models and text generation › large language model reasoning
long-context reasoning
0.312025
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning? · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 0.9many-shot ICL · 0.9receptive field analysis · 0.7cumulative normalized gradient · 0.7variance reduction · 0.6kernelization · 0.6gumbel-softmax · 0.6conditionally positive definite kernels · 0.6
YearPublicationVenuePosition
2025 MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
abstract
The ability to recognize patterns from examples and apply them to new ones is a primal ability for general intelligence, and is widely studied by psychology and AI researchers. Many benchmarks have been proposed to measure such ability for Large Language Models (LLMs); however, they focus on few-shot (usually <10) setting and lack evaluation for aggregating many pieces of information from long contexts. On the other hand, the ever-growing context length of LLMs have brought forth the novel paradigm of many-shot In-Context Learning (ICL), which addresses new tasks with hundreds to thousands of examples without expensive and inefficient fine-tuning. However, many-shot evaluations often focus on classification, and popular long-context LLM tasks such as Needle-In-A-Haystack (NIAH) seldom require complicated intelligence for integrating many pieces of information. To fix the issues from both worlds, we propose MIR-Bench, the first many-shot in-context reasoning benchmark for pattern recognition that asks LLM to predict output via input-output examples from underlying functions with diverse data format. Based on MIR-Bench, we study many novel problems for many-shot in-context reasoning, and acquired many insightful findings including scaling effect, robustness, inductive vs. transductive reasoning, retrieval Augmented Generation (RAG), coding for inductive reasoning, cross-domain generalizability, etc. Our dataset is available at https://huggingface.co/datasets/kaiyan289/MIR-Bench.
Zhan Ling, Ting-Han Fan, Lingfeng Shen, Zhengyin Du, Jiecao Chen
NeurIPS5
2023 Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis
abstract
Length extrapolation permits training a transformer language model on short sequences that preserves perplexities when tested on substantially longer sequences.A relative positional embedding design, ALiBi, has had the widest usage to date.We dissect ALiBi via the lens of receptive field analysis empowered by a novel cumulative normalized gradient tool.The concept of receptive field further allows us to modify the vanilla Sinusoidal positional embedding to create Sandwich, the first parameter-free relative positional embedding design that truly length information uses longer than the training sequence.Sandwich shares with KERPLE and T5 the same logarithmic decaying temporal bias pattern with learnable relative positional embeddings; these elucidate future extrapolatable positional embedding design.
Ta-Chung Chi, Ting-Han Fan, Alexander I. Rudnicky, Peter J. Ramadge
ACL (1)2
2022 Training Discrete Deep Generative Models via Gapped Straight-Through Estimator
abstract
While deep generative models have succeeded in image processing, natural language processing, and reinforcement learning, training that involves discrete random variables remains challenging due to the high variance of its gradient estimation process. Monte Carlo is a common solution used in most variance reduction approaches. However, this involves time-consuming resampling and multiple function evaluations. We propose a Gapped Straight-Through (GST) estimator to reduce the variance without incurring resampling overhead. This estimator is inspired by the essential properties of Straight-Through Gumbel-Softmax. We determine these properties and show via an ablation study that they are essential. Experiments demonstrate that the proposed GST estimator enjoys better performance compared to strong baselines on two discrete deep generative modeling tasks, MNIST-VAE and ListOps.
Ting-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. Ramadge
ICML1
2022 KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation
abstract
Relative positional embeddings (RPE) have received considerable attention since RPEs effectively model the relative distance among tokens and enable length extrapolation. We propose KERPLE, a framework that generalizes relative position embedding for extrapolation by kernelizing positional differences. We achieve this goal using conditionally positive definite (CPD) kernels, a class of functions known for generalizing distance metrics. To maintain the inner product interpretation of self-attention, we show that a CPD kernel can be transformed into a PD kernel by adding a constant offset. This offset is implicitly absorbed in the Softmax normalization during self-attention. The diversity of CPD kernels allows us to derive various RPEs that enable length extrapolation in a principled way. Experiments demonstrate that the logarithmic variant achieves excellent extrapolation performance on three large language modeling datasets. Our implementation and pretrained checkpoints are released at~\url{https://github.com/chijames/KERPLE.git}.
Ta-Chung Chi, Ting-Han Fan, Peter J. Ramadge, Alexander I. Rudnicky
NeurIPS2
2021 A Contraction Approach to Model-based Reinforcement Learning
abstract
Despite its experimental success, Model-based Reinforcement Learning still lacks a complete theoretical understanding. To this end, we analyze the error in the cumulative reward using a contraction approach. We consider both stochastic and deterministic state transitions for continuous (non-discrete) state and action spaces. This approach doesn’t require strong assumptions and can recover the typical quadratic error to the horizon. We prove that branched rollouts can reduce this error and are essential for deterministic transitions to have a Bellman contraction. Our analysis of policy mismatch error also applies to Imitation Learning. In this case, we show that GAN-type learning has an advantage over Behavioral Cloning when its discriminator is well-trained.
Ting-Han Fan, Peter J. Ramadge
AISTATS1
2018 Rumor Source Detection: A Probabilistic Perspective
abstract
In this paper we consider the problem of rumor source detection in a network. Our main contribution is an efficient Belief-Propagation-based (BP) algorithm to compute the joint likelihood function of the source location and the spreading time for the general continuous-time Susceptible-Infected epidemic model on trees. As a result, many probabilistic detection algorithms, including the joint maximum likelihood estimator, can be implemented with time complexity being nearly linear in the product of the size of the graph and the effective range of the spreading time. This is in sharp contrast to the widely employed discrete-time epidemic models where the complexity in computing the likelihood function of the source location is exponential. To extend the BP algorithm to general graphs, we propose a “Gamma Generated Tree” heuristic to convert the original graph to a tree with heterogeneous infection rates over edges. Compared to state-of-the-art methods, simulation results show that our algorithm provides better estimates of the source when the graph topology is similar to trees. As a byproduct, the spreading time can also be estimated, which is useful in some applications.
Ting-Han Fan, I-Hsiang Wang
ICASSP1
2017 A New Social Network Model of Online Forums
abstract
Modeling online opinion dynamics plays an important role toward in-depth comprehension of collective interactive behavior in modern human society and future digital society. Some statistical models based on network science or social networks have been known for years. In this paper, we propose a novel network model of online forums. Different from the existing models in which model selection is obscure, we accommodate reasoning in both mathematical and psychological contexts such that the basic parameters of the network model can be identified in an intuitive way. The proposed first-order-aging-with-fitness (FOAF) model extends from the well-known BA model, by adjusting weights for younger nodes, to better reflect the truth of Internet forums. We successfully verify the FOAF model and consequent entire social network model of wider applicability and better alignment with real online opinion data.
Ting-Han Fan, Kwang-Cheng Chen
GLOBECOM1