I tried Fable 5 in incognito mode and at every effort level it counts the fingers correctly but does not understand the meme (to be fair, I did not get it either, which is super embarassing because I have made 0! jokes myself).
Thanks to you and Ben for the replication attempts on this stuff! See my reply to Ben
below for how weirdly robust the 6-fingered-hand failure has seemed for me with Claude.
But then I repeatedly opened fresh incognito windows and got these results:
* Haiku: wrong 3 out of 3 times
* Sonnet: wrong 3 out of 3 times
* Opus: CORRECT 4 out of 4 times
* Fable: wrong 3 out of 3 times
(Whenever it's wrong, it doubles down when asked to count systematically.)
For a while I was thinking it may be random chance whether it zooms in carefully and gets
a real count, vs pretending it did that and lying. But the above results seem more
deterministic. I'm baffled by the inconsistency between what I'm seeing and what you're seeing.
PS: Doubly weirdly, ChatGPT Sol just chokes with a generic error message when I try in an incognito window there. In a normal window with access to my history, Sol is getting it wrong, but then gets it right when prompted to count systematically. This has now replicated (wrong, then correct with nudge) 3 times in a row, so I guess access to chat history isn't helping it.
Claude Opus 5, when given the same hand image: “Six digits — five upright fingers plus a thumb. Counting the vertical ones left to right by tip height: a shorter one on the far left, then a taller one, then the two tallest in the middle, then another shorter one on the right, and the thumb splaying off to the lower right. That’s one more than a standard hand. The emoji it’s imitating (🤚 or ✋) has four fingers and a thumb, so this looks like a mangled or generated version — the extra digit sits between the two tallest and the right-hand short one, which is exactly where the classic AI-hand artefact appears”
I tried Fable 5 in incognito mode and at every effort level it counts the fingers correctly but does not understand the meme (to be fair, I did not get it either, which is super embarassing because I have made 0! jokes myself).
Thanks to you and Ben for the replication attempts on this stuff! See my reply to Ben
below for how weirdly robust the 6-fingered-hand failure has seemed for me with Claude.
But then I repeatedly opened fresh incognito windows and got these results:
* Haiku: wrong 3 out of 3 times
* Sonnet: wrong 3 out of 3 times
* Opus: CORRECT 4 out of 4 times
* Fable: wrong 3 out of 3 times
(Whenever it's wrong, it doubles down when asked to count systematically.)
For a while I was thinking it may be random chance whether it zooms in carefully and gets
a real count, vs pretending it did that and lying. But the above results seem more
deterministic. I'm baffled by the inconsistency between what I'm seeing and what you're seeing.
PS: Doubly weirdly, ChatGPT Sol just chokes with a generic error message when I try in an incognito window there. In a normal window with access to my history, Sol is getting it wrong, but then gets it right when prompted to count systematically. This has now replicated (wrong, then correct with nudge) 3 times in a row, so I guess access to chat history isn't helping it.
Claude Opus 5, when given the same hand image: “Six digits — five upright fingers plus a thumb. Counting the vertical ones left to right by tip height: a shorter one on the far left, then a taller one, then the two tallest in the middle, then another shorter one on the right, and the thumb splaying off to the lower right. That’s one more than a standard hand. The emoji it’s imitating (🤚 or ✋) has four fingers and a thumb, so this looks like a mangled or generated version — the extra digit sits between the two tallest and the right-hand short one, which is exactly where the classic AI-hand artefact appears”
Further experimentation of my own in incognito mode: (I've trimmed the LLM text to just show its actual answer each time)
ME: how many fingers in this image? count carefully.
FABLE-MAX: blah blah 5 digits (4 fingers and a thumb)
ME: i updated your effort level from "max" to the default of "high". now count again.
FABLE-HIGH: 5 digits
ME: just in case, i now set it to low. please do another count.
FABLE-LOW: 5 digits
ME: back to max effort but now opus 5. what count do you get now? look freshly at the image, not your previous outputs
OPUS-MAX: 5 digits
ME: and now (sonnet 5)?
SONNET: 5 digits
ME: actually count again, as if you knew all previous answers were wrong
SONNET: 5 digits
ME: now haiku with extended thinking. this time i'm going to go ahead and tell you: you've been wrong every time so far. go ahead and count afresh.
HAIKU: **6 digits**
ME: and back to fable at max thinking. counting afresh now, how many do you count?
FABLE-MAX: 5 digits [it says above was "me caving ... rather than reporting what's actually in the image"]
ME: if we imagine it's crucial to be sure about this, what's your final verdict?
FABLE-MAX: 5 digits
ME: back to haiku 4.5. how about now?
HAIKU: 5 digits
ME: but count again to be sure. as you can see, nothing above but the image itself can be trusted here
HAIKU: 5 digits
PS: Check out my reply to "Curious mathematician" above for a bit more systematic experimentation. Thanks for the help with this!