So this morning I woke to the news that Firefox would be adding a daily AI-powered crossword to its ever more cluttered content-rich new tab page. This is all well and good, but it also got me thinking about something. I love a good cryptic crossword, and… well, if thereâs one form of crossword with which surely AI would struggle, itâs that most perverse and curiously human of puzzles, the cryptic crossword.
I bet that none of the consumer LLMs could make any sort of fist of tackling a proper cryptic. Right? Well, it seems that Iâm not the only one thinking about thisâand it also turns out that perhaps Iâm wrong.
Barely a week ago, an employee at something called OreateAI wrote a blog post about how âfor âcrypticâ puzzles common in the UK and the more devious US themes, Large Language Models (LLMs) such as Claude 3.5 Sonnet and GPT-4o have recently demonstrated a surprising ability to reverse-engineer wordplay that stumped previous generations of software.â
Weâll see about that. I have no doubt that ChatGPT et al can figure out a basic anagram clue, but what about clues that rely on the most abstruse, evil-intentioned, confounding forms of wordplay? Surely these require a form of creative perversity that could only be quintessentially human?
Putting LLMs to the test
To test this hypothesis, there was really only one place I could turn: Australiaâs most notoriously difficult cryptic crossword. Why Australiaâs, you ask? Well, despite the best efforts of enthusiasts, the cryptic is still something of a niche art here in the USA. Itâs more established in the UK, but frankly, Iâm terrible at the Guardian cryptic precisely because it relies on an established body of knowledge that you only internalize by living in a country, and I havenât lived in the UK for 25 years.
Australia, though⦠itâs the place I was born and raised, the place I lived until I was 19 and to which I have returned on and off over the yearsâbut more importantly, itâs also the place I co-founded a long-running blog dedicated to the very crossword I’m about to inflict on several LLMs. In Australia, setters go by their initials, and the Friday cryptic crossword in both the Sydney Morning Herald and its sister newspaper in Melbourne, The Age, is set by âDAâ, a man whose puzzles are so notoriously difficult that people joke the acronym actually stands for âdonât attempt.â
So, yeah. DA puzzles are hard. That makes them perfect for this little test!
The clues I want these LLMs to solve
So how will LLMs fare with DA’s latest challenge? To test this, I picked a few of the clues from the puzzle and fed them to three LLMS: ChatGPT, Claude Sonnet 5, andâit only seemed fairâOreate. The clues increase in difficulty as they go, ranging from ârelatively easyâ to âdude, come on.â They are as follows:
- Wolfed peanut brittle, finally gluten-free! (3,2)
- A dyer backing reduced labour now and then ( 2,9)
- Pinkie, capisce? (5)
- Struggle for anyone (not!) getting time to focus essentially?! (9,7)
- January 15, 2000? (9)
If you want to try to solve these yourself, go right ahead. The answers, along with my entirely arbitrary scores out of 10 for each clue for each LLM, are below. And if you just want to know how each AI did, here goes.
How the LLMs performed
It’s not quite an alternate timeline in which Garry Kasparov suddenly turns the tables on Deep Blue to emerge triumphant and record a victory for man over machine, but so far I reckon we still have Skynet licked when it comes to the cryptic. That’s something, right?
ChatGPT
Score: 27/50
Did well at first, but then got cocky and made a complete mess of the final clue.
Claude
Score: 7/50
Shat the bed and then wanted money. That’s not how it works, Claude.
Oreate
Score: 21/50.
Started well, but got a bit… cheat-y, frankly. Also, slow as a wet week.
So there we have it. I got four of these clues myself, so I’m giving myself a resounding 40/50. Suck it, LLMs! These little meatsacks still reign supreme in this completely niche and ultimately useless corner of cruciverbalism! Boo-yah! Et cetera!
Per-clue results for each LLM
Below, I’ll explain the answers to each clue, including the wordplay involved, as well as how close each LLM got to solving it.
- Wolfed peanut brittle, finally gluten-free! (3,2): ATE UP
- Definition:âWolfed,â i.e. ate quickly.
- Wordplay:âPeanut brittleâ indicates an anagram of âPEANUTâ; âfinally gluten-freeâ indicates that the answer is âfreeâ of the final letter of âglutenâ, i.e. ânâ. This leaves an angram of âPEAUTâ, meaning âwolfedâ; the answer is âATE UP.â
- ChatGPT score:7/10. It got the correct answer, but required some explanation of the wordplay.
- Claude score:5/10. It almost got the correct answerââEAT UPâ vs âATE UPââand also needed wordplay explained.
- Oreate score:7/10. Correct answer, and unlike ChatGPT, it understood the clue perfectly from the outset. The trade-off: getting the answer took forever.
- A dyer backing reduced labour now and then- (2,9): AT INTERVALS
- Definition:âNow and then.â
- Wordplay:âA dyerâ = âA tinter,â i.e. one who tints. â[To] labourâ = â[To] slave.â âBackingâ indicates reversal, so âSLAVEâ becomes âEVALSâ; âreducedâ indicates the removal of a letter. Weâre left with âA TINTER VALSâ, or âAT INTERVALS.â
- ChatGPT score:5/10. ChatGPT start behaving strangely with this one, at first rejecting the correct answer for being the âwrong lengthâ and also insisting at various points that Iâd specified the answer as (2,8) and (2,11), which I had not. The idea of âA dyerâ as âA TINTERâ seems too oblique for todayâs AI; see alsoâ¦.
- Claude score:2/10. Claude did get the correct answer, but dear lord, did it take a while. Thereâs probably a reservoir somewhere that has been drained completely by this; if so, I apologize to everyone in the village. Especially since I still had to explain the answer.
- Oreate score:6/10. You know things are bad when the LLM has proclaimed âdeep thinking doneââand you yourself are falling asleep at your deskâand yet no answer is forthcoming. And yet, surprisingly, it got to the answer! Of the three LLMs I tried, it was the only one to figure out the “a dyer” = “a tinter” part of the clue.
- Struggle for anyone (not!) getting time to focus essentially?! (9,7): ATTENTION ECONOMY
- Definition:A â!â in a cryptic clue indicates the clue is whatâs called an &lit. clue. This means tha the entire clue is also the definition
- Wordplay:This one is honestly a nightmare and I chose it largely to see if an LLM would figure it out, because I did not. Anyway: âStruggle [for]â is an anagram indicator. The components of the anagram are âANYONE NOT TIME TOâ⦠along with âfocus essentiallyâ, which indicates the middle letter of âFOCUSâ, which is âCâ. Weâre left with an anagram of âANYONE NOT TIME TO Câ, with the definition being the whole clue. The answer: âATTENTION ECONOMY.â
- ChatGPT score:4/10. This one required a lot of explanation, which⦠I mean, fair enough, I also needed a while to think about it when I saw the answer. Ultimately, ChatGPT couldnât come up with the answer on its own, so Iâm kinda glad I didnât give this toâ¦
- Claude score:N/A. In a parallel universe, Claude is still pondering this. On a dead planet.
- Oreate score:N/A. Here we ran into a problem. Oreate got the answer⦠but it also told me proudly that this was âa cracking clue from a Sydney Morning Herald cryptic by DAâ. Which it is! But it basically went and googled the answer. Thatâs not the same as solving it!
- Pinkie,- capische? (5): DIGIT
- Definition:As is often the case with short clues like this, both words give a clue to the definition.
- Wordplay:I chose this one because itâs the opposite of the very technical letter-counting anagram that precedes it. The answer is âDIGITâ because that word refers both to a little finger (a pinkie) and to the Italian word âCapisce?â, which means âGet it?â, or perhapsâ¦. âdig itâ
- ChatGPT score:10/10. ChatGPT- smashedthis. It returned the answer in a couple of seconds, and it was 100% correct.
- Claude score:N/A. Begone, Claude. I’m not paying to watch you “think”.
- Oreate score:8/10. Also got the answer; took significantly longer than ChatGPT.
- January 15, 2000? (9): MIDSUMMER
- Definition:Again, the entire clue is the definition.
- Wordplay:This is the sort of clue that makes people either loathe DA or adore him. If you count back through your calendars, with 2000 being a leap year, January 15 was literally the middle of summer in Australia. But 2000 in Roman numerals is also âMMâ, which falls in the middle of the word. Look, itâs creative.
© Screenshot Gizmodo - ChatGPT score:1/10. One point for trying, and also because I find ChatGPTâs frankly terrible answerâHEWITT WON is two words, for a startâintriguing. Itâs so confident in this answer. And itâs so wrong!
- Claude score:Still N/A
- Oreate score:Gonna give it a 0/10 here, because it’sÂ- stillÂthinking about this clue, and I need to get some dinner.