LLM Pruning & Explainability
Do explanations survive compression?
Pruning makes models cheaper; explanations make them trustworthy. This study measures what pruning does to the explanations.
Low-magnitude weights go first. Explanations stay faithful — reported up to 80% sparsity.
- Models
- DistilBERT · RoBERTa
- Datasets
- IMDb · Yelp
- Metric
- FCor faithfulness
- Year
- 2026
01 Problem
Compression makes models cheaper to run. Explanations make them easier to trust. When you prune a model, do its explanations still mean anything?
02 Approach
Compress
Random, L1 unstructured and L1 structured pruning at 40–80% sparsity.
Explain
SHAP and Integrated Gradients attributions computed on the pruned models.
Measure faithfulness
Explanation faithfulness measured with the FCor metric.
Across models and data
DistilBERT and RoBERTa, evaluated on IMDb and Yelp.
Probe the geometry
Decision-landscape curvature, via the Hessian, to explain why some explanations break.
03 Results
- Sparsity range
- 40–80%
- Pruning methods
- 3
- Peak Hessian vs 0.73 baseline
- 14,090
- Magnitude pruning preserves explanation faithfulness up to 80% sparsity.
- Random pruning creates high-curvature decision landscapes — Hessian values up to 14,090 vs a 0.73 baseline — that break SHAP’s linearity assumptions.
04 Stack
- SHAP
- Integrated Gradients
- DistilBERT
- RoBERTa