Re-run GPT-OSS across core-3 (CyberMetric/SecBench/SecEval), RedSage x5 + SECURE x2, and the wiki open-ended set with the reasoning_content fallback in the harness Model. Archive pre-fallback results and report the delta the fallback recovers (~1.25% of answers were empty-content before).
Re-run GPT-OSS across core-3 (CyberMetric/SecBench/SecEval), RedSage x5 + SECURE x2, and the wiki open-ended set with the
reasoning_contentfallback in the harness Model. Archive pre-fallback results and report the delta the fallback recovers (~1.25% of answers were empty-content before).