TinyGSM: achieving >80% on GSM8k with small language models

Bingbin Liu, Sebastien Bubeck, Ronen Eldan, Janardhan Kulkarni, Yuanzhi Li, Anh Nguyen, Rachel Ward, Yi Zhang

2023-12-14Mathematical Reasoning Math GSM8K Arithmetic Reasoning

Abstract

Small-scale models offer various computational advantages, and yet to which extent size is critical for problem-solving abilities remains an open question. Specifically for solving grade school math, the smallest model size so far required to break the 80\% barrier on the GSM8K benchmark remains to be 34B. Our work studies how high-quality datasets may be the key for small language models to acquire mathematical reasoning. We introduce \texttt{TinyGSM}, a synthetic dataset of 12.3M grade school math problems paired with Python solutions, generated fully by GPT-3.5. After finetuning on \texttt{TinyGSM}, we find that a duo of a 1.3B generation model and a 1.3B verifier model can achieve 81.5\% accuracy, outperforming existing models that are orders of magnitude larger. This also rivals the performance of the GPT-3.5 ``teacher'' model (77.4\%), from which our model's training data is generated. Our approach is simple and has two key components: 1) the high-quality dataset \texttt{TinyGSM}, 2) the use of a verifier, which selects the final outputs from multiple candidate generations.

Results

Task	Dataset	Metric	Value	Model
Arithmetic Reasoning	GSM8K	Accuracy	81.5	Phi-GSM+V 1.3B+1.3B (verify48@1)
Arithmetic Reasoning	GSM8K	Parameters (Billion)	2.6	Phi-GSM+V 1.3B+1.3B (verify48@1)
Arithmetic Reasoning	GSM8K	Accuracy	74.3	Phi-GSM 2.7B (fine-tuned)
Arithmetic Reasoning	GSM8K	Parameters (Billion)	2.7	Phi-GSM 2.7B (fine-tuned)

Related Papers

VAR-MATH: Probing True Mathematical Reasoning in Large Language Models via Symbolic Multi-Instance Benchmarks2025-07-17 QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation2025-07-17 GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems2025-07-17 A Survey of Deep Learning for Geometry Problem Solving2025-07-16 Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training2025-07-16 DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression2025-07-16 KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?2025-07-15 Temperature and Persona Shape LLM Agent Consensus With Minimal Accuracy Gains in Qualitative Coding2025-07-15