Point 2 is the one I’d push further. Once a probability has done its real job, exposing that two people in the room are 40pp apart, there’s a second lever most skip: checking whether the reasoning each side exposed was actually licensed by the evidence, or just confidently asserted. The interesting part of a 40pp gap usually isn’t the number. It’s that one person treated as established what the evidence only suggested.
That also reframes the accountability worry. Brier scores need a tournament’s worth of resolved questions to mean much, which slow or one-shot decisions never supply. But you can benchmark a different way: name the load-bearing reasons up front, seal them before the outcome, then later check whether the ones you flagged as weak were the ones that actually drove the result. That works on a single decision, and it survives the clearance filter you describe, because what reaches the minister isn’t a percentage, it’s “here is why this was a reasonable call on what we knew at the time.”
I work on exactly this in pharma R&D, where go/no-go calls have dated readouts a year or two out, so the reasoning can be scored against reality without waiting for a tournament’s worth of questions. The decision-makers there want the defensible record, not the probability, which matches your experience.
Point 2 is the one I’d push further. Once a probability has done its real job, exposing that two people in the room are 40pp apart, there’s a second lever most skip: checking whether the reasoning each side exposed was actually licensed by the evidence, or just confidently asserted. The interesting part of a 40pp gap usually isn’t the number. It’s that one person treated as established what the evidence only suggested.
That also reframes the accountability worry. Brier scores need a tournament’s worth of resolved questions to mean much, which slow or one-shot decisions never supply. But you can benchmark a different way: name the load-bearing reasons up front, seal them before the outcome, then later check whether the ones you flagged as weak were the ones that actually drove the result. That works on a single decision, and it survives the clearance filter you describe, because what reaches the minister isn’t a percentage, it’s “here is why this was a reasonable call on what we knew at the time.”
I work on exactly this in pharma R&D, where go/no-go calls have dated readouts a year or two out, so the reasoning can be scored against reality without waiting for a tournament’s worth of questions. The decision-makers there want the defensible record, not the probability, which matches your experience.