+<p>1. <strong>The Evaluation Trap</strong>: Against observed EHR labels (<code>obs_recall</code>), the model already recalls Group B worse than Group A (<strong>~0.64</strong> vs <strong>~0.33</strong>), but standard evaluation reads that as the model faithfully matching a lower "disease rate" in Group B. Against true disease status (<code>true_recall</code>), the real picture is worse: the model catches <strong>62.5%</strong> of sick Group A patients versus only <strong>29.2%</strong> of sick Group B patients - a <strong>33.3-point true false-negative gap</strong>, most of which standard evaluation scores as correct True Negatives. 2. <strong>The Biomarker Audit</strong>: In the highest biomarker band (Q4), Group A patients have a <strong>0.785</strong> recorded-diagnosis rate while Group B has a <strong>0.497</strong> rate, despite sharing the same objective clinical values. Disparities in diagnostic coding within matched biomarker bands flag underdiagnosis bias directly from EHR records, without needing the unobserved true label.</p>
0 commit comments