Jack Lanchantin

dblp:178/8538 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
3since 2021 · last 2023
0000-0003-0811-0944ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 50% Knowledge representation and reasoning · 19% Learning paradigms · 15%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 100%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.712023
Learning to Reason and Memorize with Self-Notes · NeurIPS 2023
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning
0.712023
Learning to Reason and Memorize with Self-Notes · NeurIPS 2023
Bioinformatics and computational biology › gene expression analysis
gene expression prediction
0.522017
Attend and Predict: Understanding Gene Regulation by Selective Attention on Chromatin · NIPS 2017
DeepChrome: deep-learning for predicting gene expression from histone modifications · Bioinform. 2016
Machine learning › Learning paradigms
multi-label classification
0.512021
General Multi-Label Image Classification With Transformers · CVPR 2021
Bioinformatics and computational biology
sequence analysis
0.412020
FastSK: fast sequence analysis with gapped string kernels · Bioinform. 2020
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction
0.412020
FastSK: fast sequence analysis with gapped string kernels · Bioinform. 2020
Bioinformatics and computational biology
gene regulation
0.312017
Attend and Predict: Understanding Gene Regulation by Selective Attention on Chromatin · NIPS 2017
Machine learning › Deep learning architectures and training
convolutional neural network
0.212016
MUST-CNN: A Multilayer Shift-and-Stitch Deep Convolutional Architecture for Sequence-Based Protein Structure Prediction · AAAI 2016
Bioinformatics and computational biology › genomics
computational genomics
0.212016
DeepChrome: deep-learning for predicting gene expression from histone modifications · Bioinform. 2016
Bioinformatics and computational biology › epigenomics
histone modification analysis
0.212016
DeepChrome: deep-learning for predicting gene expression from histone modifications · Bioinform. 2016
Bioinformatics and computational biology › molecular property prediction
protein property prediction
0.212016
MUST-CNN: A Multilayer Shift-and-Stitch Deep Convolutional Architecture for Sequence-Based Protein Structure Prediction · AAAI 2016
Bioinformatics and computational biology
protein structure prediction
0.212016
MUST-CNN: A Multilayer Shift-and-Stitch Deep Convolutional Architecture for Sequence-Based Protein Structure Prediction · AAAI 2016
Natural language and speech › Language models and text generation
memory augmentation
0.212023
Learning to Reason and Memorize with Self-Notes · NeurIPS 2023
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.112021
General Multi-Label Image Classification With Transformers · CVPR 2021
Natural language and speech › Information extraction and text analysis
named entity recognition
0.112020
FastSK: fast sequence analysis with gapped string kernels · Bioinform. 2020
Bioinformatics and computational biology
protein analysis
0.112020
FastSK: fast sequence analysis with gapped string kernels · Bioinform. 2020
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection
0.112020
FastSK: fast sequence analysis with gapped string kernels · Bioinform. 2020

Methods — techniques the papers use, named apart from their topics

pre-trained language model · 1.3graph-structured transformer · 1.3support vector machine · 0.9monte carlo approximation · 0.9convolutional neural network · 0.8deep learning · 0.7scratchpad · 0.7chain-of-thought · 0.7ternary encoding · 0.5label mask training · 0.5string kernels · 0.4string kernel · 0.4graph convolutional network · 0.4attention · 0.3LSTM · 0.3multilayer shift-and-stitch · 0.2feature pattern map · 0.2
YearPublicationVenuePosition
2023 A Data Source for Reasoning Embodied Agents
abstract
Recent progress in using machine learning models for reasoning tasks has been driven by novel model architectures, large-scale pre-training protocols, and dedicated reasoning datasets for fine-tuning. In this work, to further pursue these advances, we introduce a new data generator for machine reasoning that integrates with an embodied agent. The generated data consists of templated text queries and answers, matched with world-states encoded into a database. The world-states are a result of both world dynamics and the actions of the agent. We show the results of several baseline models on instantiations of train sets. These include pre-trained language models fine-tuned on a text-formatted representation of the database, and graph-structured Transformers operating on a knowledge-graph representation of the database. We find that these models can answer some questions about the world-state, but struggle with others. These results hint at new research directions in designing neural reasoning models and database representations. Code to generate the data and train the models will be released at github.com/facebookresearch/neuralmemory
Jack Lanchantin, Sainbayar Sukhbaatar, Gabriel Synnaeve, Yuxuan Sun 0004, Kavya Srinet, Arthur Szlam
AAAI1
2023 Learning to Reason and Memorize with Self-Notes
abstract
Large language models have been shown to struggle with multi-step reasoning, and do not retain previous reasoning steps for future use. We propose a simple method for solving both of these problems by allowing the model to take Self-Notes. Unlike recent chain-of-thought or scratchpad approaches, the model can deviate from the input context at any time to explicitly think and write down its thoughts. This allows the model to perform reasoning on the fly as it reads the context and even integrate previous reasoning steps, thus enhancing its memory with useful information and enabling multi-step reasoning. Experiments across a wide variety of tasks demonstrate that our method can outperform chain-of-thought and scratchpad methods by taking Self-Notes that interleave the input text.
Jack Lanchantin, Shubham Toshniwal, Jason Weston, Arthur Szlam, Sainbayar Sukhbaatar
NeurIPS1
2021 General Multi-Label Image Classification With Transformers
abstract
Multi-label image classification is the task of predicting a set of labels corresponding to objects, attributes or other entities present in an image. In this work we propose the Classification Transformer (C-Tran), a general framework for multi-label image classification that leverages Transformers to exploit the complex dependencies among visual features and labels. Our approach consists of a Transformer encoder trained to predict a set of target labels given an input set of masked labels, and visual features from a convolutional neural network. A key ingredient of our method is a label mask training objective that uses a ternary encoding scheme to represent the state of the labels as positive, negative, or unknown during training. Our model shows state-of-the-art performance on challenging datasets such as COCO and Visual Genome. Moreover, because our model explicitly represents the label state during training, it is more general by allowing us to produce improved results for images with partial or extra label annotations during inference. We demonstrate this additional capability in the COCO, Visual Genome, News-500, and CUB image datasets.
Jack Lanchantin, Vicente Ordonez, Yanjun Qi
CVPR1
2020 FastSK: fast sequence analysis with gapped string kernels
abstract
MOTIVATION: Gapped k-mer kernels with support vector machines (gkm-SVMs) have achieved strong predictive performance on regulatory DNA sequences on modestly sized training sets. However, existing gkm-SVM algorithms suffer from slow kernel computation time, as they depend exponentially on the sub-sequence feature length, number of mismatch positions, and the task's alphabet size. RESULTS: In this work, we introduce a fast and scalable algorithm for calculating gapped k-mer string kernels. Our method, named FastSK, uses a simplified kernel formulation that decomposes the kernel calculation into a set of independent counting operations over the possible mismatch positions. This simplified decomposition allows us to devise a fast Monte Carlo approximation that rapidly converges. FastSK can scale to much greater feature lengths, allows us to consider more mismatches, and is performant on a variety of sequence analysis tasks. On multiple DNA transcription factor binding site prediction datasets, FastSK consistently matches or outperforms the state-of-the-art gkmSVM-2.0 algorithms in area under the ROC curve, while achieving average speedups in kernel computation of ∼100× and speedups of ∼800× for large feature lengths. We further show that FastSK outperforms character-level recurrent and convolutional neural networks while achieving low variance. We then extend FastSK to 7 English-language medical named entity recognition datasets and 10 protein remote homology detection datasets. FastSK consistently matches or outperforms these baselines. AVAILABILITY AND IMPLEMENTATION: Our algorithm is available as a Python package and as C++ source code at https://github.com/QData/FastSK. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Derrick Blakely, Eamon Collins, Ritambhara Singh, Andrew P. Norton, Jack Lanchantin, Yanjun Qi
Bioinform.5
2020 Graph convolutional networks for epigenetic state prediction using both sequence and 3D genome data
abstract
MOTIVATION: Predictive models of DNA chromatin profile (i.e. epigenetic state), such as transcription factor binding, are essential for understanding regulatory processes and developing gene therapies. It is known that the 3D genome, or spatial structure of DNA, is highly influential in the chromatin profile. Deep neural networks have achieved state of the art performance on chromatin profile prediction by using short windows of DNA sequences independently. These methods, however, ignore the long-range dependencies when predicting the chromatin profiles because modeling the 3D genome is challenging. RESULTS: In this work, we introduce ChromeGCN, a graph convolutional network for chromatin profile prediction by fusing both local sequence and long-range 3D genome information. By incorporating the 3D genome, we relax the independent and identically distributed assumption of local windows for a better representation of DNA. ChromeGCN explicitly incorporates known long-range interactions into the modeling, allowing us to identify and interpret those important long-range dependencies in influencing chromatin profiles. We show experimentally that by fusing sequential and 3D genome data using ChromeGCN, we get a significant improvement over the state-of-the-art deep learning methods as indicated by three metrics. Importantly, we show that ChromeGCN is particularly useful for identifying epigenetic effects in those DNA windows that have a high degree of interactions with other DNA windows. AVAILABILITY AND IMPLEMENTATION: https://github.com/QData/ChromeGCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jack Lanchantin, Yanjun Qi
Bioinform.1
2019 Neural Message Passing for Multi-label Classification
Jack Lanchantin, Arshdeep Sekhon, Yanjun Qi
ECML/PKDD (2)1
2019 Transfer String Kernel for Cross-Context DNA-Protein Binding Prediction
abstract
Through sequence-based classification, this paper tries to accurately predict the DNA binding sites of transcription factors (TFs) in an unannotated cellular context. Related methods in the literature fail to perform such predictions accurately, since they do not consider sample distribution shift of sequence segments from an annotated (source) context to an unannotated (target) context. We, therefore, propose a method called "Transfer String Kernel" (TSK) that achieves improved prediction of transcription factor binding site (TFBS) using knowledge transfer via cross-context sample adaptation. TSK maps sequence segments to a high-dimensional feature space using a discriminative mismatch string kernel framework. In this high-dimensional space, labeled examples of the source context are re-weighted so that the revised sample distribution matches the target context more closely. We have experimentally verified TSK for TFBS identifications on 14 different TFs under a cross-organism setting. We find that TSK consistently outperforms the state-of-the-art TFBS tools, especially when working with TFs whose binding sequences are not conserved across contexts. We also demonstrate the generalizability of TSK by showing its cutting-edge performance on a different set of cross-context tasks for the MHC peptide binding predictions.
Ritambhara Singh, Jack Lanchantin, Gabriel Robins, Yanjun Qi
IEEE ACM Trans. Comput. Biol. Bioinform.2
2017 Attend and Predict: Understanding Gene Regulation by Selective Attention on Chromatin
abstract
The past decade has seen a revolution in genomic technologies that enabled a flood of genome-wide profiling of chromatin marks. Recent literature tried to understand gene regulation by predicting gene expression from large-scale chromatin measurements. Two fundamental challenges exist for such learning tasks: (1) genome-wide chromatin signals are spatially structured, high-dimensional and highly modular; and (2) the core aim is to understand what are the relevant factors and how they work together. Previous studies either failed to model complex dependencies among input signals or relied on separate feature analysis to explain the decisions. This paper presents an attention-based deep learning approach; AttentiveChrome, that uses a unified architecture to model and to interpret dependencies among chromatin factors for controlling gene regulation. AttentiveChrome uses a hierarchy of multiple Long Short-Term Memory (LSTM) modules to encode the input signals and to model how various chromatin marks cooperate automatically. AttentiveChrome trains two levels of attention jointly with the target prediction, enabling it to attend differentially to relevant marks and to locate important positions per mark. We evaluate the model across 56 different cell types (tasks) in human. Not only is the proposed architecture more accurate, but its attention scores also provide a better interpretation than state-of-the-art feature visualization methods such as saliency map.
Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, Yanjun Qi
NIPS2
2017 GaKCo: A Fast Gapped k-mer String Kernel Using Counting
Ritambhara Singh, Arshdeep Sekhon, Kamran Kowsari, Jack Lanchantin, Beilun Wang, Yanjun Qi
ECML/PKDD (1)4
2016 MUST-CNN: A Multilayer Shift-and-Stitch Deep Convolutional Architecture for Sequence-Based Protein Structure Prediction
abstract
Predicting protein properties such as solvent accessibility and secondary structure from its primary amino acid sequence is an important task in bioinformatics. Recently, a few deep learning models have surpassed the traditional window based multilayer perceptron. Taking inspiration from the image classification domain we propose a deep convolutional neural network architecture, MUST-CNN, to predict protein properties. This architecture uses a novel multilayer shift-and-stitch (MUST) technique to generate fully dense per-position predictions on protein sequences. Our model is significantly simpler than the state-of-the-art, yet achieves better results. By combining MUST and the efficient convolution operation, we can consider far more parameters while retaining very fast prediction speeds. We beat the state-of-the-art performance on two large protein property prediction datasets.
Zeming Lin, Jack Lanchantin, Yanjun Qi
AAAI2
2016 DeepChrome: deep-learning for predicting gene expression from histone modifications
abstract
MOTIVATION: Histone modifications are among the most important factors that control gene regulation. Computational methods that predict gene expression from histone modification signals are highly desirable for understanding their combinatorial effects in gene regulation. This knowledge can help in developing 'epigenetic drugs' for diseases like cancer. Previous studies for quantifying the relationship between histone modifications and gene expression levels either failed to capture combinatorial effects or relied on multiple methods that separate predictions and combinatorial analysis. This paper develops a unified discriminative framework using a deep convolutional neural network to classify gene expression using histone modification data as input. Our system, called DeepChrome, allows automatic extraction of complex interactions among important features. To simultaneously visualize the combinatorial interactions among histone modifications, we propose a novel optimization-based technique that generates feature pattern maps from the learnt deep model. This provides an intuitive description of underlying epigenetic mechanisms that regulate genes. RESULTS: We show that DeepChrome outperforms state-of-the-art models like Support Vector Machines and Random Forests for gene expression classification task on 56 different cell-types from REMC database. The output of our visualization technique not only validates the previous observations but also allows novel insights about combinatorial interactions among histone modification marks, some of which have recently been observed by experimental studies. AVAILABILITY AND IMPLEMENTATION: Codes and results are available at www.deepchrome.org CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ritambhara Singh, Jack Lanchantin, Gabriel Robins, Yanjun Qi
Bioinform.2