Fine-tuning Large Language Models for Entity Matching

Aaron Steiner, Ralph Peeters, Christian Bizer

2024-09-12Entity Resolution Data Integration Prompt Engineering

Abstract

Generative large language models (LLMs) are a promising alternative to pre-trained language models for entity matching due to their high zero-shot performance and ability to generalize to unseen entities. Existing research on using LLMs for entity matching has focused on prompt engineering and in-context learning. This paper explores the potential of fine-tuning LLMs for entity matching. We analyze fine-tuning along two dimensions: 1) the representation of training examples, where we experiment with adding different types of LLM-generated explanations to the training set, and 2) the selection and generation of training examples using LLMs. In addition to the matching performance on the source dataset, we investigate how fine-tuning affects the models ability to generalize to other in-domain datasets as well as across topical domains. Our experiments show that fine-tuning significantly improves the performance of the smaller models while the results for the larger models are mixed. Fine-tuning also improves the generalization to in-domain datasets while hurting cross-domain transfer. We show that adding structured explanations to the training set has a positive impact on the performance of three out of four LLMs, while the proposed example selection and generation methods, only improve the performance of Llama 3.1 8B while decreasing the performance of GPT-4o-mini.

Results

Task	Dataset	Metric	Value	Model
Data Integration	Abt-Buy	F1 (%)	94.09	gpt-4o-mini-2024-07-18_fine_tuned
Data Integration	Abt-Buy	F1 (%)	92.2	gpt-4o-2024-08-06
Data Integration	Abt-Buy	F1 (%)	87.68	gpt-4o-mini-2024-07-18
Data Integration	Abt-Buy	F1 (%)	87.34	Meta-Llama-3.1-8B-Instruct_fine_tuned
Data Integration	Abt-Buy	F1 (%)	79.12	Meta-Llama-3.1-70B-Instruct
Data Integration	Abt-Buy	F1 (%)	56.57	Meta-Llama-3.1-8B-Instruct
Data Integration	WDC Products-80%cc-seen-medium	F1 (%)	87.1	gpt-4o-2024-08-06_fine_tuned_wdc_small
Data Integration	WDC Products-80%cc-seen-medium	F1 (%)	84.38	gpt-4o-mini-2024-07-18_structured_explanations
Data Integration	WDC Products-80%cc-seen-medium	F1 (%)	81.61	gpt-4o-mini-2024-07-18
Data Integration	WDC Products-80%cc-seen-medium	F1 (%)	76.7	Llama3.1_70B_structured_explanations
Data Integration	WDC Products-80%cc-seen-medium	F1 (%)	75.2	Llama3.1_70B
Data Integration	WDC Products-80%cc-seen-medium	F1 (%)	74.37	Llama3.1_8B_error-based_example_selection
Data Integration	WDC Products-80%cc-seen-medium	F1 (%)	74.13	Llama3.1_8B_structured_explanations
Data Integration	WDC Products-80%cc-seen-medium	F1 (%)	53.36	Llama3.1_8B
Data Integration	Amazon-Google	F1 (%)	80.25	gpt-4o-mini-2024-07-18_fine_tuned
Data Integration	Amazon-Google	F1 (%)	63.45	gpt-4o-2024-08-06
Data Integration	Amazon-Google	F1 (%)	59.2	gpt-4o-mini-2024-07-18
Data Integration	Amazon-Google	F1 (%)	51.44	Meta-Llama-3.1-70B-Instruct
Data Integration	Amazon-Google	F1 (%)	50	Meta-Llama-3.1-8B-Instruct_fine_tuned
Data Integration	Amazon-Google	F1 (%)	49.16	Meta-Llama-3.1-8B-Instruct
Data Integration	WDC Products	F1 (%)	87.07	gpt-4o-2024-08-06_fine_tuned_wdc_small
Entity Resolution	Abt-Buy	F1 (%)	94.09	gpt-4o-mini-2024-07-18_fine_tuned
Entity Resolution	Abt-Buy	F1 (%)	92.2	gpt-4o-2024-08-06
Entity Resolution	Abt-Buy	F1 (%)	87.68	gpt-4o-mini-2024-07-18
Entity Resolution	Abt-Buy	F1 (%)	87.34	Meta-Llama-3.1-8B-Instruct_fine_tuned
Entity Resolution	Abt-Buy	F1 (%)	79.12	Meta-Llama-3.1-70B-Instruct
Entity Resolution	Abt-Buy	F1 (%)	56.57	Meta-Llama-3.1-8B-Instruct
Entity Resolution	WDC Products-80%cc-seen-medium	F1 (%)	87.1	gpt-4o-2024-08-06_fine_tuned_wdc_small
Entity Resolution	WDC Products-80%cc-seen-medium	F1 (%)	84.38	gpt-4o-mini-2024-07-18_structured_explanations
Entity Resolution	WDC Products-80%cc-seen-medium	F1 (%)	81.61	gpt-4o-mini-2024-07-18
Entity Resolution	WDC Products-80%cc-seen-medium	F1 (%)	76.7	Llama3.1_70B_structured_explanations
Entity Resolution	WDC Products-80%cc-seen-medium	F1 (%)	75.2	Llama3.1_70B
Entity Resolution	WDC Products-80%cc-seen-medium	F1 (%)	74.37	Llama3.1_8B_error-based_example_selection
Entity Resolution	WDC Products-80%cc-seen-medium	F1 (%)	74.13	Llama3.1_8B_structured_explanations
Entity Resolution	WDC Products-80%cc-seen-medium	F1 (%)	53.36	Llama3.1_8B
Entity Resolution	Amazon-Google	F1 (%)	80.25	gpt-4o-mini-2024-07-18_fine_tuned
Entity Resolution	Amazon-Google	F1 (%)	63.45	gpt-4o-2024-08-06
Entity Resolution	Amazon-Google	F1 (%)	59.2	gpt-4o-mini-2024-07-18
Entity Resolution	Amazon-Google	F1 (%)	51.44	Meta-Llama-3.1-70B-Instruct
Entity Resolution	Amazon-Google	F1 (%)	50	Meta-Llama-3.1-8B-Instruct_fine_tuned
Entity Resolution	Amazon-Google	F1 (%)	49.16	Meta-Llama-3.1-8B-Instruct
Entity Resolution	WDC Products	F1 (%)	87.07	gpt-4o-2024-08-06_fine_tuned_wdc_small

Fine-tuning Large Language Models for Entity Matching

Abstract

Results

Related Papers

Fine-tuning Large Language Models for Entity Matching

Abstract

Results

Related Papers