Training Graph Neural Networks with 1000 Layers

Guohao Li, Matthias Müller, Bernard Ghanem, Vladlen Koltun

2021-06-14Graph Sampling Node Property Prediction

Abstract

Deep graph neural networks (GNNs) have achieved excellent results on various tasks on increasingly large graph datasets with millions of nodes and edges. However, memory complexity has become a major obstacle when training deep GNNs for practical applications due to the immense number of nodes, edges, and intermediate activations. To improve the scalability of GNNs, prior works propose smart graph sampling or partitioning strategies to train GNNs with a smaller set of nodes or sub-graphs. In this work, we study reversible connections, group convolutions, weight tying, and equilibrium models to advance the memory and parameter efficiency of GNNs. We find that reversible connections in combination with deep network architectures enable the training of overparameterized GNNs that significantly outperform existing methods on multiple datasets. Our models RevGNN-Deep (1001 layers with 80 channels each) and RevGNN-Wide (448 layers with 224 channels each) were both trained on a single commodity GPU and achieve an ROC-AUC of $87.74 \pm 0.13$ and $88.24 \pm 0.15$ on the ogbn-proteins dataset. To the best of our knowledge, RevGNN-Deep is the deepest GNN in the literature by one order of magnitude. Please visit our project website https://www.deepgcns.org/arch/gnn1000 for more information.

Results

Task	Dataset	Metric	Value	Model
Node Property Prediction	ogbn-arxiv	Number of params	2098256	RevGAT+N.Adj+LabelReuse+SelfKD
Node Property Prediction	ogbn-arxiv	Number of params	2098256	RevGAT+NormAdj+LabelReuse
Node Property Prediction	ogbn-products	Number of params	2945007	RevGNN-112
Node Property Prediction	ogbn-proteins	Number of params	68471608	RevGNN-Wide
Node Property Prediction	ogbn-proteins	Number of params	20031384	RevGNN-Deep

Training Graph Neural Networks with 1000 Layers

Abstract

Results

Related Papers

Training Graph Neural Networks with 1000 Layers

Abstract

Results

Related Papers