The trust point above, that demand for rationales is partly distrust of the numbers, is the one I’d build on. A lot of that demand isn’t really a request for more words. It’s a way of asking “should I believe this wasn’t reverse-engineered from what you already expected?” That’s a question about a rationale’s integrity, not its length, and it’s also why “now it’s just LLM hallucinations” feels like a fair worry.
Which points at a job distinct from producing rationales: evaluating them. For each load-bearing claim, was it actually established by the evidence, or only suggested by it and then written up as established? Two rationales of equal length can differ enormously on that, and that’s where most of the persuasive weight should sit.
The cleanest way I’ve found to make it checkable is to seal the assessment before the outcome and judge it only on what was knowable at the time. Then hindsight can’t quietly relabel a lucky call as sound reasoning, and a reader can verify that for themselves.
I work on this in pharma R&D, where unlike AGI the outcomes are dated and land in a year or two, so the discipline is testable against reality on a real clock. Different domain, same hole. Glad to share a worked example if it’s useful.
The trust point above, that demand for rationales is partly distrust of the numbers, is the one I’d build on. A lot of that demand isn’t really a request for more words. It’s a way of asking “should I believe this wasn’t reverse-engineered from what you already expected?” That’s a question about a rationale’s integrity, not its length, and it’s also why “now it’s just LLM hallucinations” feels like a fair worry.
Which points at a job distinct from producing rationales: evaluating them. For each load-bearing claim, was it actually established by the evidence, or only suggested by it and then written up as established? Two rationales of equal length can differ enormously on that, and that’s where most of the persuasive weight should sit.
The cleanest way I’ve found to make it checkable is to seal the assessment before the outcome and judge it only on what was knowable at the time. Then hindsight can’t quietly relabel a lucky call as sound reasoning, and a reader can verify that for themselves.
I work on this in pharma R&D, where unlike AGI the outcomes are dated and land in a year or two, so the discipline is testable against reality on a real clock. Different domain, same hole. Glad to share a worked example if it’s useful.