Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Ofir Press, Noah A. Smith, Mike Lewis

2021-08-27ICLR 2022 4Word Embeddings Playing the Game of 2048

Paper PDF Code Code Code Code Code Code Code(official)Code Code Code

Abstract

Since the introduction of the transformer model by Vaswani et al. (2017), a fundamental question has yet to be answered: how does a model achieve extrapolation at inference time for sequences that are longer than it saw during training? We first show that extrapolation can be enabled by simply changing the position representation method, though we find that current methods do not allow for efficient extrapolation. We therefore introduce a simpler and more efficient position method, Attention with Linear Biases (ALiBi). ALiBi does not add positional embeddings to word embeddings; instead, it biases query-key attention scores with a penalty that is proportional to their distance. We show that this method trains a 1.3 billion parameter model on input sequences of length 1024 that extrapolates to input sequences of length 2048, achieving the same perplexity as a sinusoidal position embedding model trained on inputs of length 2048 but training 11% faster and using 11% less memory. ALiBi's inductive bias towards recency also leads it to outperform multiple strong position methods on the WikiText-103 benchmark.

Related Papers

Speak2Sign3D: A Multi-modal Pipeline for English Speech to American Sign Language Animation2025-07-09 Computational Detection of Intertextual Parallels in Biblical Hebrew: A Benchmark Study Using Transformer-Based Language Models2025-06-30 Including Semantic Information via Word Embeddings for Skeleton-based Action Recognition2025-06-23 Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings2025-06-21 Characterizing Linguistic Shifts in Croatian News via Diachronic Word Embeddings2025-06-16 Learning Obfuscations Of LLM Embedding Sequences: Stained Glass Transform2025-06-11 Recommender systems, stigmergy, and the tyranny of popularity2025-06-06 Static Word Embeddings for Sentence Semantic Representation2025-06-05