Explainable Artificial Intelligence (VTR UGE 21), day 2 of 5

Model-Agnostic Explanation Methods

Prof. Dr. Utku Kose, Süleyman Demirel University

Explanation lab

Part A explains a case of a model with two features in three ways at once. Part B matches questions with the methods that answer them. Part C computes Shapley values by hand for a model with three features.

Part A: Three explanations of one case

The heat map shows a model of two features, a and b, with an interaction: The effect of one feature depends on the value of the other. Click a point on the map to choose the case to explain. The panel then explains it in three ways: Exact Shapley values against the data mean, a local linear surrogate in the manner of LIME (local interpretable model-agnostic explanations) whose kernel width and seed you set, and the nearest counterfactual, drawn as a path to the decision boundary, the line on which the model output is 0.5.

The two Shapley values add up to the model output minus 0.5, the output at the data mean. A bar named LIME shows the slope of the surrogate times the distance of the case from the data mean, so that it can be compared with the Shapley value next to it. The LIME fidelity is the share of the variation of the model near the case that the surrogate reproduces, a number between 0 and 1. The formulas are on the lecture page, under the headings on LIME, on Shapley values and on choosing a method.

Open in ColabContinue in Colab, section 4: LIME and SHAP (Shapley additive explanations) for a real patient, side by side.

Part B: Which method answers which question?

Each question below is answered best by one family of methods. Choose the method for every question and press Check the matches. In the list, ICE stands for individual conditional expectation, and a local attribution states how much each feature contributed to one prediction.

QuestionMethodFeedback

Open in ColabContinue in Colab, section 3: partial dependence and ICE curves for any measurement.

Part C: Shapley values by hand

A model with three features, A, B and C, has the outputs below when only some features take the patient's values and the others take their average. Compute the Shapley value of each feature: the average, over the six orders in which the features can join, of the change of the output at the moment the feature joins.

Features presentnoneABCA, B A, CB, CA, B, C
Model output0.300.500.40 0.300.700.550.450.80

Open in ColabContinue in Colab, section 8: exact Shapley values next to KernelSHAP and TreeSHAP, two algorithms of the SHAP family.