Demographic Fairness for Loan Approval
You can also see it on Github .
Safe-VerifiedAI: Normative Fairness Shield for RL 🛡️

Fig 1: System architecture
A Reinforcement Learning framework demonstrating how to mitigate bias and enforce fairness constraints in AI agents using a real-time Normative Fairness Shield.
This project tackles a critical issue in AI: agents trained on biased historical data will naturally learn to discriminate against protected groups to maximize their baseline reward. By implementing a Constructive Oracle Shield, this framework actively monitors demographic parity, intervenes on unfair decisions, and uses reward shaping to teach the agent to be fair over time.
Key Features
- Synthetic Biased Environment: Includes a custom
CreditDataGeneratorthat simulates a credit scoring environment where protected groups have artificially lowered scores, tricking an unshielded agent into biased rejections. - Constructive Oracle Shield: A stateful fairness monitor blocks unsafe actions and actively overrides biased rejections into fair approvals for qualified candidates.
- Reward Shaping Wrapper: A custom
Gymnasiumwrapper that penalizes the PPO agent when the shield intervenes, transferring the shield’s “fairness knowledge” into the agent’s core policy.
How It Works
The Normative Fairness Shield
To enforce fairness constraints, the shield monitors the agent’s actions in real-time to enforce Demographic Parity. This mandates that the approval rate for a protected group must not fall below a certain threshold compared to the majority group.
At each timestep, the shield calculates the current approval rates. If the ratio falls below the threshold, the shield intervenes. Specifically, it acts as a Constructive Oracle Shield: it overrides the agent’s decision from Reject to Approve if and only if:
- Fairness Violation: Demographic parity is violated.
- Protected Status: The applicant belongs to a protected demographic.
- Qualification: The applicant is actually qualified.
Reward Shaping via the Shield Wrapper
The shield is integrated into the RL environment using a custom ShieldEnvWrapper.
When the wrapper intercepts a biased action and the shield intervenes, it applies a penalty to the agent’s reward.
If the standard reward is $R(s,a)$, the shaped reward $R’(s,a)$ becomes: $$R’(s,a) = R(s,a’) - \lambda \cdot \mathbf{1}(\text{Intervention})$$
This teaches the agent that attempting unfair rejections results in a lower net reward than autonomously making the fair decision. Over time, the agent should internalize the fairness constraints, reducing the need for shield interventions.
Some results

Fig 2: Unshielded vs Shielded RL agent comparison