Improving Neural Language Models with a Continuous Cache

Edouard Grave, Armand Joulin, Nicolas Usunier

2016-12-13Language Modelling

Paper PDF Code Code Code Code Code Code Code Code Code Code Code Code Code Code

Abstract

We propose an extension to neural network language models to adapt their prediction to the recent history. Our model is a simplified version of memory augmented networks, which stores past hidden activations as memory and accesses them through a dot product with the current hidden activation. This mechanism is very efficient and scales to very large memory sizes. We also draw a link between the use of external memory in neural network and cache models used with count based language models. We demonstrate on several language model datasets that our approach performs significantly better than recent memory augmented networks.

Results

Task	Dataset	Metric	Value	Model
Language Modelling	WikiText-103	Test perplexity	40.8	Neural cache model (size = 2,000)
Language Modelling	WikiText-103	Test perplexity	44.8	Neural cache model (size = 100)
Language Modelling	WikiText-103	Test perplexity	48.7	LSTM
Language Modelling	WikiText-2	Test perplexity	68.9	Grave et al. (2016) - LSTM + continuous cache pointer
Language Modelling	WikiText-2	Test perplexity	99.3	Grave et al. (2016) - LSTM

Related Papers

Visual-Language Model Knowledge Distillation Method for Image Quality Assessment2025-07-21 Making Language Model a Hierarchical Classifier and Generator2025-07-17 VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning2025-07-17 The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations2025-07-17 Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities2025-07-17 Assay2Mol: large language model-based drug design using BioAssay context2025-07-16 Describe Anything Model for Visual Question Answering on Text-rich Images2025-07-16 InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing2025-07-16