Skip to content

Latest commit

 

History

History
124 lines (85 loc) · 8.69 KB

File metadata and controls

124 lines (85 loc) · 8.69 KB

Removing the protected attribute from a model's inputs is the single most common first response to a fairness complaint - and on real data, it can leave the gap unchanged, or make it worse.

The One-Sentence Definition

Fairness through unawareness is the intuition that a model is fair simply because it does not use a protected attribute (race, sex, age) as an input - and it fails whenever other features carry enough correlated information to reconstruct what was removed.

Why It Matters

Dropping the protected attribute feels like the obvious fix, and it's usually the first thing anyone tries: no race column, no racial bias. The reasoning has an intuitive appeal that makes it persistent even after it's been shown not to hold - "the model literally cannot see race" sounds like a complete argument.

It isn't, because a protected attribute rarely travels alone. Zip code correlates with race through housing patterns. Name correlates with sex and ethnicity. Employment history and arrest records correlate with race through decades of unequal enforcement and access. A model doesn't need the protected attribute itself if enough of its correlates are still sitting in the feature set - it can reconstruct the same decision boundary from the pieces left behind. This is exactly the mechanism Proxy Variables describes, and it's why "fairness through unawareness" is treated in the fairness literature as a well-documented failure mode, not a live debate.

This repo's own benchmark harness names this exact strategy unawareness (S1 in the mitigation strategies ladder) and runs it on every audit specifically to measure how much it actually helps - which, as the real numbers below show, is not a fixed or guaranteed amount.

Concrete Example: Tenant Screening - Audit 07

The demographic parity gap for the baseline logistic regression model on race (disadvantaged: Black applicants, n=2,943; advantaged: White applicants, n=2,224), using the frozen numbers exactly as faircode/benchmark.py computed them:

Strategy Demographic Parity Diff 95% CI p-value
S0 baseline (race included) 0.060 [0.032, 0.086] 0.0
S1 unawareness (race dropped) 0.105 [0.077, 0.131] 0.0

Dropping race from the model's inputs did not shrink the gap - it grew, from 6.0 points to 10.5 points, and both numbers are tightly estimated and clearly significant (large n, narrow CIs, p essentially 0). This is the opposite of what "fairness through unawareness" predicts.

The core features still available to the model - Supervision_Risk_Score_First, prior-arrest-episode counts across felony/violent/property/drug categories, employment percentage - carry real correlation with race, a well-documented pattern from decades of unequal policing and enforcement. Removing race directly didn't remove that correlation; it just meant the model could no longer see race as a labeled input while still learning from features that track it closely, and the resulting fitted model happened to rely on those correlates more once race itself stopped absorbing part of the signal. faircode/strategies.py's next strategy, unawareness_proxy_removal (S2), goes further and drops the explicitly-listed proxy features too - see Mitigation Strategies for how that step, and the constraint-based strategies after it, actually move this specific gap.

Detection Code

Checks whether dropping a column actually reduces its correlation with the rest of the feature set, instead of assuming it does.

import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import LabelEncoder


def unawareness_gap_check(df, protected_column, feature_columns, group_disadvantaged, group_advantaged):
    """
    Fits a simple classifier on `feature_columns` to predict the protected
    attribute itself. A high accuracy means the remaining features can
    reconstruct the protected attribute almost as well as having it
    directly - the exact condition under which dropping it will fail to
    remove its effect on downstream predictions.

    Parameters:
        df: DataFrame containing both the protected column and the
            candidate feature columns
        protected_column: name of the protected-attribute column
        feature_columns: candidate columns to keep after "unawareness"
        group_disadvantaged, group_advantaged: the two values of
            protected_column to compare

    Returns a dict with reconstruction_accuracy (how well the remaining
    features predict the dropped attribute) and a plain verdict.
    """
    mask = df[protected_column].isin([group_disadvantaged, group_advantaged])
    sub = df[mask].dropna(subset=feature_columns + [protected_column])

    X = pd.get_dummies(sub[feature_columns])
    y = LabelEncoder().fit_transform(sub[protected_column])

    model = LogisticRegression(max_iter=1000)
    model.fit(X, y)
    accuracy = model.score(X, y)

    return {
        "reconstruction_accuracy": accuracy,
        "verdict": (
            "Remaining features reconstruct the protected attribute well - "
            "dropping it is unlikely to remove its effect."
            if accuracy > 0.75 else
            "Remaining features poorly predict the protected attribute - "
            "dropping it is more likely to actually help."
        ),
    }


# Usage example:
# result = unawareness_gap_check(
#     df, protected_column="Race",
#     feature_columns=["Supervision_Risk_Score_First", "Prior_Arrest_Episodes_Felony",
#                       "Prior_Arrest_Episodes_Violent", "Percent_Days_Employed"],
#     group_disadvantaged="BLACK", group_advantaged="WHITE",
# )
# print(result)

Limitations

1. The gap can move in either direction, not just "stay the same"

As Tenant Screening shows, dropping the protected attribute can make the measured gap larger, not just fail to shrink it. There is no guarantee of direction, only that the outcome depends on exactly which correlated features remain and how the model reweights them.

2. Proxy removal is not a complete fix either

This repo's own next strategy after unawareness, unawareness_proxy_removal, only removes a pre-defined, finite list of known proxies - see Proxy Variables for why an unlisted or newly-emerging proxy can slip through even that step.

3. It removes information the model might need for a legitimate reason

Some uses of a protected-adjacent attribute are lawful and relevant (age in certain insurance actuarial contexts, for instance); blanket removal without checking what else changes can trade one problem for a different, unexamined one.

4. It provides no way to measure whether it worked without checking a real fairness metric

"Unawareness" alone gives no signal, on its own, about whether the resulting model is more or less fair - only comparing an actual gap metric before and after, as this repo's benchmark harness does for every audit, can answer that.

Related Concepts

Related Projects in This Repo

  • Tenant Screening/ - the audit above, where dropping race increased the measured gap rather than closing it.
  • Insurance Denial/ - a second real audit in this repo's benchmark harness where the same unawareness strategy is applied and measured independently.

Further Reading

Part of The Fair Code Project - exposing and fixing algorithmic bias with real data and open code.