The Short Version
The available evidence indicates these win-probability models are broadly calibrated over many games and usually align reasonably well with actual outcomes. They are not perfect: some studies find underdog bias and phase-specific quirks, and they are not clearly superior to simple baseline models. The comparison to human intuition is supported more indirectly than directly, but the overall claim is largely accurate.