Counterfactual Latent Representations: A Neurosymbolic Approach
MCML Authors
Abstract
Abstract
Counterfactual fairness, i.e., asking how a prediction would change under interventions on sensitive attributes, is hard to operationalize for unstructured inputs such as text. We propose a modular neurosymbolic framework that decouples representation learning from causal intervention: a pretrained encoder maps text to a latent representation, a semantic decoder grounds part of this space in a structured causal subgraph, and a neural manipulator learns to transform the latent representation analogously to a symbolic counterfactual operation on the decoded features, trained under a consistency constraint between the two. The approach localizes causal assumptions in an explicit subgraph, trading end-to-end optimization for auditability.
inproceedings KK26b
ECAF 2026
5th European Conference on Algorithmic Fairness. Ghent, Belgium, Sep 02-04, 2026. Spotlight Presentation. To be published.Authors
L. Kestel • C. KernResearch Area
BibTeXKey: KK26b