There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.
If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.
If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.