Learning local discrete features in explainable-by-design convolutional neural networks

Pantelis I. Kaplanoglou, Konstantinos Diamantaras

2024-10-31Image Classification Explainable Artificial Intelligence (XAI)Interpretable Machine Learning

Abstract

Our proposed framework attempts to break the trade-off between performance and explainability by introducing an explainable-by-design convolutional neural network (CNN) based on the lateral inhibition mechanism. The ExplaiNet model consists of the predictor, that is a high-accuracy CNN with residual or dense skip connections, and the explainer probabilistic graph that expresses the spatial interactions of the network neurons. The value on each graph node is a local discrete feature (LDF) vector, a patch descriptor that represents the indices of antagonistic neurons ordered by the strength of their activations, which are learned with gradient descent. Using LDFs as sequences we can increase the conciseness of explanations by repurposing EXTREME, an EM-based sequence motif discovery method that is typically used in molecular biology. Having a discrete feature motif matrix for each one of intermediate image representations, instead of a continuous activation tensor, allows us to leverage the inherent explainability of Bayesian networks. By collecting observations and directly calculating probabilities, we can explain causal relationships between motifs of adjacent levels and attribute the model's output to global motifs. Moreover, experiments on various tiny image benchmark datasets confirm that our predictor ensures the same level of performance as the baseline architecture for a given count of parameters and/or layers. Our novel method shows promise to exceed this performance while providing an additional stream of explanations. In the solved MNIST classification task, it reaches a comparable to the state-of-the-art performance for single models, using standard training setup and 0.75 million parameters.

Results

Task	Dataset	Metric	Value	Model
Image Classification	Fashion-MNIST	Accuracy	93.45	R-ExplaiNet-26
Image Classification	Fashion-MNIST	Percentage error	6.55	R-ExplaiNet-26
Image Classification	Fashion-MNIST	Trainable Parameters	892362	R-ExplaiNet-26
Image Classification	Oracle-MNIST	Accuracy	96.93	R-ExplaiNet-26
Image Classification	Oracle-MNIST	Trainable Parameters	892362	R-ExplaiNet-26
Image Classification	CIFAR-10	Percentage correct	94.15	R-ExplaiNet-26
Image Classification	Kuzushiji-MNIST	Accuracy	98.78	R-ExplaiNet-26
Image Classification	Kuzushiji-MNIST	Error	1.22	R-ExplaiNet-26
Image Classification	Kuzushiji-MNIST	Trainable Parameters	892362	R-ExplaiNet-26
Image Classification	MNIST	Accuracy	99.8	R-ExplaiNet-22 (single model)
Image Classification	MNIST	Percentage error	0.2	R-ExplaiNet-22 (single model)
Image Classification	MNIST	Trainable Parameters	743882	R-ExplaiNet-22 (single model)

Learning local discrete features in explainable-by-design convolutional neural networks

Abstract

Results

Related Papers

Learning local discrete features in explainable-by-design convolutional neural networks

Abstract

Results

Related Papers