1.3 Explainability with multimodal models
Attention, LIME and SHAP, grounding in RAG, model cards and proxy models.
Key points
Explainable AI (XAI) is a set of tools and techniques that help people understand why a model decided something. For images, audio and text, attention can be visualized to show which parts of the input mattered.
What NVIDIA says (2)
“is a set of tools and techniques used by organizations to help people better understand why a model makes certain decisions and how it works.”
“For some data — images, audio and text — similar results can be visualized through the use of”
LIME and SHAP assign credit to each input feature for one prediction. NVIDIA notes they give literal mathematical answers that can be shown to many audiences.
What NVIDIA says (1)
“Techniques with names like LIME and SHAP offer very literal mathematical answers to this question”
Retrieval-augmented generation (RAG) fetches relevant data and gives it to the model as context. Grounding means converting other modalities into one primary modality, here text. The text descriptions also make the retrieved evidence readable by people.
What NVIDIA says (2)
“The key benefit here is that the metadata generated from the information-rich image is extremely helpful in answering objective questions.”
“The key disadvantages are preprocessing costs and losing some nuance from the image.”
RAG retrieves documents and passes them to the model. The model can then cite them, so users can verify each claim. A multimodal RAG app can cite retrieved images and charts the same way.
What NVIDIA says (1)
“Retrieval-augmented generation gives models sources they can cite, like footnotes in a research paper, so users can check any claims.”
A model card is a short document that describes a model, its uses and its limits. Model Card++ adds four subsections on trust topics.
What NVIDIA says (1)
“Four subsections detailing model-specific information concerning Bias, Explainability, Privacy, and Safety and Security.”
Proxy modeling uses a simple model to approximate a complex one. It gives a sense of the whole model, but it is only an approximation.
What NVIDIA says (1)
“simpler, more easily comprehended models like decision trees can be used to approximately describe the more detailed AI model.”
Pedigree is the history of how a model was made. Knowing it helps people judge when the model's outputs make sense, including for multimodal models trained on images and text.
What NVIDIA says (1)
“Explaining the pedigree of the model: How was the model trained? What data was used? How was the impact of any bias in the training data measured and mitigated?”
Key terms
- Attention map: A picture of attention weights showing which parts of the input a model focused on.
- Explainable AI: Tools and techniques that help people understand why a model made a decision and how it works.
- Model card: A short document that describes a model, its intended use and its limits for developers and users.
- Retrieval-augmented generation: A method that retrieves relevant data at query time and gives it to the model as context.
Sample question
How can a model with attention help explain a decision about an image, audio or text input?
Show the answer
Answer: Visualize where the model's attention fell, so the model shows its work
Explainable AI (XAI) is a set of tools and techniques that help people understand why a model decided something. For images, audio and text, attention can be visualized to show which parts of the input mattered.
What NVIDIA says (2)
“is a set of tools and techniques used by organizations to help people better understand why a model makes certain decisions and how it works.”
“For some data — images, audio and text — similar results can be visualized through the use of”
Practice 1.3 (7 questions) Full Experimentation guide
← 1.2 Managing and preprocessing multimodal data · 1.4 Testing multimodal data quality →