When models disagree, measuring ensemble variance helps spot risky inputs that...
https://front-wiki.win/index.php/How_to_Instrument_Per-Instance_Outputs_for_Disagreement_Monitoring
When models disagree, measuring ensemble variance helps spot risky inputs that need extra attention. For example, cases in the top 1-2% of variance can be routed to human review to prevent errors