<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Explainable AI |</title><link>https://sabrina-henry.github.io/tags/explainable-ai/</link><atom:link href="https://sabrina-henry.github.io/tags/explainable-ai/index.xml" rel="self" type="application/rss+xml"/><description>Explainable AI</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><image><url>https://sabrina-henry.github.io/media/icon_hu_da05098ef60dc2e7.png</url><title>Explainable AI</title><link>https://sabrina-henry.github.io/tags/explainable-ai/</link></image><item><title>SIGMA: Self-Guided Integrated Gradient Method for Attribution</title><link>https://sabrina-henry.github.io/projects/sigma/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://sabrina-henry.github.io/projects/sigma/</guid><description>&lt;p&gt;Modern neural networks can have very high confidence — but confidence alone does not tell us &lt;em&gt;why&lt;/em&gt; a prediction was made.&lt;/p&gt;
&lt;p&gt;Feature attribution methods aim to answer this &lt;em&gt;why&lt;/em&gt; by highlighting which parts of an input most influenced a model&amp;rsquo;s decision. In practice, however, the way we choose to explain a prediction can actually shape the explanation itself.&lt;/p&gt;
&lt;h2 id="what-does-it-mean-to-explain-a-prediction"&gt;What does it mean to explain a prediction?&lt;/h2&gt;
&lt;p&gt;Given an input and a trained model, an attribution method assigns importance scores to individual input features — pixels in an image, time points on a graph, or tokens in text — indicating how strongly each contributed to the final prediction.&lt;/p&gt;
&lt;p&gt;Among the most widely used approaches are &lt;em&gt;path-based&lt;/em&gt; methods, which accumulate model gradients as the input is gradually transformed, telling us how sensitive each input feature is to the output prediction.&lt;/p&gt;
&lt;div class="interactive-figure"&gt;
&lt;div class="figure-frame"&gt;
&lt;img src="Interactive_figs/ladybug.png" alt="Input image" class="figure-image base-image"&gt;
&lt;img src="Interactive_figs/SIGMA_ladybug.png" alt="Integrated Gradients attribution" class="figure-image attribution-image"&gt;
&lt;/div&gt;
&lt;button class="toggle-btn" onclick="toggleAttribution(this)"&gt;Click to Show Attribution&lt;/button&gt;
&lt;figcaption&gt;The model predicted the class 'Ladybug' for this input image, click to highlight which pixels were most important in this prediction.&lt;/figcaption&gt;
&lt;/div&gt;
&lt;h2 id="the-baseline-assumption"&gt;The baseline assumption&lt;/h2&gt;
&lt;p&gt;Integrated Gradients explains a prediction by integrating gradients along a straight path from a &lt;em&gt;baseline&lt;/em&gt; input to the image being explained. The baseline is intended to represent the absence of information — often a black or blurred image. However, different baselines can lead to noticeably different explanations, even when the model and input remain unchanged.&lt;/p&gt;
&lt;figure class="interactive-figure" data-figure="baseline-1"&gt;
&lt;div class="figure-frame-overlay"&gt;
&lt;img src="Interactive_figs/ladybug.png" alt="Input image" class="figure-image base-image"&gt;
&lt;img src="" alt="Attribution overlay" class="figure-image attribution-image" data-overlay&gt;
&lt;/div&gt;
&lt;div class="toggle-row" role="group" aria-label="Choose baseline"&gt;
&lt;button class="toggle-btn baseline-btn" aria-pressed="true" data-src="Interactive_figs/IG_black_baseline.png" onclick="setBaseline(this)"&gt;Black&lt;/button&gt;
&lt;button class="toggle-btn baseline-btn" aria-pressed="false" data-src="Interactive_figs/IG_noise_baseline.png" onclick="setBaseline(this)"&gt;Noise&lt;/button&gt;
&lt;button class="toggle-btn baseline-btn" aria-pressed="false" data-src="Interactive_figs/IG_blur_baseline.png" onclick="setBaseline(this)"&gt;Blurred&lt;/button&gt;
&lt;/div&gt;
&lt;figcaption&gt;Integrated Gradients attribution for the same input using different baselines.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="saturation-along-attribution-paths"&gt;Saturation along attribution paths&lt;/h2&gt;
&lt;p&gt;As the input moves away from the baseline, the model&amp;rsquo;s confidence often increases rapidly before entering a saturated region where further changes have little effect on the output. Gradients accumulated in these flat regions can dominate the final attribution, despite contributing little to the actual decision.&lt;/p&gt;
&lt;figure class="interactive-figure figure-large"&gt;
&lt;div class="figure-frame"&gt;
&lt;img src="Interactive_figs/saturation-01.png" alt="Saturation along attribution paths" class="figure-image"&gt;
&lt;/div&gt;
&lt;/figure&gt;
&lt;h2 id="thinking-in-terms-of-confidence-landscapes"&gt;Thinking in terms of confidence landscapes&lt;/h2&gt;
&lt;p&gt;Rather than viewing attribution as interpolation between two images, it can be helpful to think of the model&amp;rsquo;s output as a &lt;em&gt;confidence landscape&lt;/em&gt; over input space. An informative explanation should focus on the regions where the model&amp;rsquo;s belief actually changes — not where it is already certain.&lt;/p&gt;
&lt;figure class="interactive-figure figure-large"&gt;
&lt;div class="figure-frame"&gt;
&lt;img src="Interactive_figs/conflandscape.png" alt="Confidence landscape" class="figure-image"&gt;
&lt;/div&gt;
&lt;/figure&gt;
&lt;h2 id="a-self-guided-path"&gt;A self-guided path&lt;/h2&gt;
&lt;p&gt;SIGMA removes the need for a baseline entirely. Instead of following a predefined path, SIGMA constructs its own trajectory by iteratively perturbing the input in directions that reduce the model&amp;rsquo;s confidence in the predicted class. Gradients are accumulated along this path and weighted by the corresponding drop in confidence.&lt;/p&gt;
&lt;figure class="interactive-figure figure-video"&gt;
&lt;div class="figure-frame"&gt;
&lt;video class="article-video" autoplay muted loop playsinline controls preload="auto"&gt;
&lt;source src="Interactive_figs/sigma_side_by_side.mp4" type="video/mp4"&gt;
Sorry — your browser can't play this video.
&lt;/video&gt;
&lt;/div&gt;
&lt;figcaption&gt;SIGMA trajectory through a prediction landscape with corresponding perturbed input and attribution evolution.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure class="interactive-figure figure-method"&gt;
&lt;div class="figure-frame"&gt;
&lt;img src="Interactive_figs/SIGMA.png" alt="SIGMA attribution method" class="figure-image"&gt;
&lt;/div&gt;
&lt;/figure&gt;
&lt;h2 id="what-changes-in-practice"&gt;What changes in practice?&lt;/h2&gt;
&lt;p&gt;Because SIGMA follows the model&amp;rsquo;s confidence downhill, it avoids early saturation and continues to collect informative gradients throughout the path. The resulting attribution maps tend to be more spatially coherent and less influenced by regions that have little effect on the prediction.&lt;/p&gt;
&lt;figure class="interactive-figure figure-large"&gt;
&lt;div class="figure-frame"&gt;
&lt;img src="Interactive_figs/further_results_qualitative.png" alt="Further results" class="figure-image"&gt;
&lt;/div&gt;
&lt;/figure&gt;
&lt;h2 id="faithfulness-to-model-behaviour"&gt;Faithfulness to model behaviour&lt;/h2&gt;
&lt;p&gt;By accumulating gradients in proportion to confidence change, SIGMA aligns attribution strength with the model&amp;rsquo;s sensitivity. Revealing input features in order of attribution importance reconstructs the model&amp;rsquo;s confidence more efficiently than random or saturated-gradient explanations.&lt;/p&gt;
&lt;h2 id="beyond-explanation"&gt;Beyond explanation&lt;/h2&gt;
&lt;p&gt;Following confidence collapse naturally produces perturbed inputs that remain recognisable to humans but receive near-zero confidence from the model. These low-confidence variants expose decision boundaries and can be reused as training augmentations, linking interpretability and model reliability.&lt;/p&gt;
&lt;div class="perturb-widget"&gt;
&lt;div class="perturb-buttons" role="tablist" aria-label="Choose perturbation"&gt;
&lt;button class="perturb-btn is-active" data-mode="clean" aria-selected="true"&gt;
&lt;img src="Interactive_figs/clean_button.png" alt="Clean example"&gt;
&lt;/button&gt;
&lt;button class="perturb-btn" data-mode="gaussian" aria-selected="false"&gt;
&lt;img src="Interactive_figs/gaussian_button.png" alt="Gaussian example"&gt;
&lt;/button&gt;
&lt;button class="perturb-btn" data-mode="fgsm" aria-selected="false"&gt;
&lt;img src="Interactive_figs/fgsm_button.png" alt="FGSM example"&gt;
&lt;/button&gt;
&lt;button class="perturb-btn" data-mode="sigma" aria-selected="false"&gt;
&lt;img src="Interactive_figs/sigma_button.png" alt="SIGMA example"&gt;
&lt;/button&gt;
&lt;/div&gt;
&lt;figure class="chart-figure"&gt;
&lt;img id="chartImg" src="Interactive_figs/web_bar_original.png" alt="Bar chart for Clean perturbation"&gt;
&lt;figcaption id="chartCap"&gt;Original network performance (Clean)&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;/div&gt;
&lt;script&gt;
(function () {
const CHARTS = {
clean: { src: "Interactive_figs/web_bar_original.png", cap: "Original network performance (Clean)" },
gaussian: { src: "Interactive_figs/web_bar_gaussian.png", cap: "Network performance under Gaussian noise augmentations" },
fgsm: { src: "Interactive_figs/web_bar_FGSM.png", cap: "Network performance under FGSM augmentations" },
sigma: { src: "Interactive_figs/web_bar_SIGMA.png", cap: "Network performance under SIGMA augmentations" }
};
const chartImg = document.getElementById("chartImg");
const chartCap = document.getElementById("chartCap");
function setActive(btn) {
document.querySelectorAll(".perturb-btn").forEach(b =&gt; {
const active = (b === btn);
b.classList.toggle("is-active", active);
b.setAttribute("aria-selected", active ? "true" : "false");
});
}
function swapChart(mode) {
const d = CHARTS[mode];
chartImg.src = d.src;
chartImg.alt = "Bar chart for " + mode;
chartCap.textContent = d.cap;
}
document.querySelectorAll(".perturb-btn").forEach(btn =&gt; {
btn.addEventListener("click", () =&gt; {
setActive(btn);
swapChart(btn.dataset.mode);
});
});
})();
&lt;/script&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;SIGMA offers a shift in perspective for attribution based interpretability — from explaining predictions relative to arbitrary references, to explaining them by following the model&amp;rsquo;s own confidence. By treating explanations as paths shaped by the model itself, we gain both clearer attributions and a deeper understanding of how confidence emerges and collapses.&lt;/p&gt;
&lt;script&gt;
function toggleAttribution(button) {
const figure = button.closest(".interactive-figure");
if (!figure) return;
const attribution = figure.querySelector(".attribution-image");
if (!attribution) return;
const isVisible = attribution.classList.toggle("visible");
button.textContent = isVisible ? "Hide attribution" : "Show attribution";
button.setAttribute("aria-pressed", isVisible ? "true" : "false");
}
function setBaseline(button) {
const figure = button.closest(".interactive-figure");
if (!figure) return;
const overlay = figure.querySelector("img[data-overlay]");
if (!overlay) return;
const buttons = figure.querySelectorAll(".baseline-btn");
buttons.forEach(btn =&gt; {
const isActive = btn === button;
btn.classList.toggle("active", isActive);
btn.setAttribute("aria-pressed", isActive ? "true" : "false");
});
const src = button.getAttribute("data-src").trim();
overlay.src = src;
overlay.classList.add("visible");
}
document.addEventListener("DOMContentLoaded", () =&gt; {
document.querySelectorAll(".interactive-figure").forEach(fig =&gt; {
const overlay = fig.querySelector("img[data-overlay]");
const active = fig.querySelector(".baseline-btn.active");
if (overlay &amp;&amp; active) setBaseline(active);
});
});
&lt;/script&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Paper:&lt;/strong&gt;
&lt;br&gt;
&lt;strong&gt;Code:&lt;/strong&gt;
&lt;br&gt;
&lt;strong&gt;Contact:&lt;/strong&gt;
&lt;/p&gt;</description></item></channel></rss>