
TLDR
I trained a 59M-parameter transformer from scratch to play Wheel of Fortune, and named it transformer-solve. It plays head to head against smart-solve, a conventional solver built on word lists. transformer-solve first learned English from a billion characters of text, then trained on 5.8 million game lines labelled with smart-solve’s moves. Training took 32 minutes on one rented H100.
On 3,349 puzzles neither had seen, transformer-solve solved 98.7% of MAIN games and smart-solve 98.3%. smart-solve needed fewer letters, 9.04 on average against 12.33, and won 67% of the head-to-head games. In BONUS, smart-solve solved 79.8% and transformer-solve 77.2%.
transformer-solve learned to imitate its teacher, not to beat it. Most of the work was data, not model code, and the biggest gain came from tuning when it may solve. smart-solve fails on words missing from its lists, transformer-solve on spelling.
Introduction
I wanted to know whether a small language model could learn to play Wheel of Fortune. Not just read the board, but play: choose letters, decide when it knows enough to solve, and do it without a strategy written into its code. The test was to give it a teacher and see how closely it could imitate it.
So I built two solvers.
smart-solve is deliberately conventional. It has a word list, and it calls the letter expected to narrow the remaining candidates the most. It plays well with no machine learning at all.
transformer-solve is a small character-level language model trained from scratch. It first learns the patterns of written English, then trains on games labelled with smart-solve’s moves. A language model can produce plausible text. The question is whether that is enough to learn a game from examples.
Both played the same few thousand puzzles, which neither had seen before.
The result was closer than I expected. The two solved at nearly the same rate, but the transformer needed more letters. It had learned to play mostly by imitating its teacher. The interesting part is why, and whether I can close the gap.
This is the story of the two solvers, the data behind them, the mistakes they made, and what the experiment taught me about training a small model to do something useful.
Head-to-head solvers
The demo puts the two solvers head to head. Both get the same puzzle under the same rules, and each plays its own board beside the other’s, one move at a time.
A scoreboard compares the two games:
- A solver that solves beats one that doesn’t.
- If both solve, fewer misses wins, then fewer letters.
- If those are equal, it’s a tie. If neither solves, both lose.
Three modes pick the puzzles:
- Random plays held-out puzzles one after another, until you stop it.
- Demo plays seven preset examples, then stops.
- Your puzzle plays one puzzle you type in, in a category you pick.
MAIN and BONUS switch between the main game and the bonus round.
The game as played here
Wheel of Fortune puzzles are short phrases with a category, laid out on a board of four rows. The rows are 12, 14, 14 and 12 squares wide, 52 squares in all. That layout has been on the show since 1997. Letters start hidden. Spaces and punctuation always show.
The code writes a board as one line, with an underscore for each hidden letter. A _A__L_ __ ___ ______ is A CANDLE IN THE WINDOW after A and L are called. Category names come from the show: Thing, Phrase, Food And Drink, Before And After, Rhyme Time and 44 others.
The main game (MAIN)
From any board, a solver makes one of three moves:
- CALL a consonant:
CALL T. Every T on the board turns over. - BUY a vowel:
BUY E. - SOLVE:
SOLVE A CANDLE IN THE WINDOW.
The rules:
- Invalid moves end the game. Calling a letter twice, buying a consonant, or calling a vowel are all invalid.
- The money rule. On the show a vowel costs $250, and money only comes from consonants that hit. The games here keep no money, so the rule is read off the board: a vowel can be bought once a called consonant shows on it.
- One SOLVE. The first SOLVE ends the game, right or wrong.
A MAIN game ends one of three ways: solved, wrong, or invalid. A solved game is also measured by how many letters were called before the SOLVE, and how many of those were misses, letters not in the puzzle. On the show a miss loses the turn. Fewer is better for both. A solver that waits until the board is full always solves, so the letter count is the real measure.
The bonus round (BONUS)
The bonus round is one shot:
- R S T L N E are turned over.
- The solver picks 3 more consonants and 1 vowel:
PICK DHM A. - Those letters turn over, and the solver gets one SOLVE.
BONUS is scored on the solve rate alone. It is harder than MAIN: the solver sees ten letters at most, picks four of them blind, and gets one guess.
What is left out
There is no wheel, no money, no turns and no opponent on the board. Each solver plays its own copy of the puzzle from the first letter to the solve.
Fair comparison
The two solvers are held to the same terms:
- Same puzzles. Both play the same 3,349 held-out puzzles, with the same move rules.
- No peeking. Neither solver sees a held-out answer. Both know only the training puzzles, the word list, and the category lists. transformer-solve also knows the pretraining text.
- The caller holds the answer. The evaluation and the server send a solver only the category, the board and the called letters, never the answer.
Building & Training the Model
This section walks through the build one phase at a time, from the scraped puzzles to the C server behind the demo. Every phase has its script, its inputs and outputs, a sample of the data, and the commands to run it. The setup section covers what you need. Every generated file has a published SHA-256 checksum, so you can check that your rebuild matches mine byte for byte.
The project is in its own repository, dannyheskett/wof-solver. All links on this page point at the tag v1, so the code you read here is the code that produced the results.
| Directory | What is in it |
|---|---|
scripts/ | One Python script per phase, named by phase (01_fetch_puzzles.py to 13_build_demo.py), plus three shared modules |
data/ | The puzzles, the splits, the category lists, the results. DATA.md documents every file |
server/ | wof-server, the C program that runs both solvers and serves the demo |
client/ | The demo game, in C with raylib, compiled to WebAssembly |
NOTES.md | The project notes: rules, results, hosting, the cost log |
Setup
Everything except training (phase 9) runs on an ordinary Linux machine.
What you need:
| Python | 3.12, with the packages in requirements.txt: numpy 1.26.4, pyarrow 25.0.1, torch 2.14.0 |
| Disk | about 10 GB: 6.3 GB of downloaded sources, 2.9 GB of generated training files, 0.6 GB of checkpoints |
| A C compiler | for the C server. Any C99 compiler with x86-64-v3 support (AVX2, F16C) |
| A GPU | only to retrain (phase 9). The trained checkpoints are published, so you can skip training |
Steps:
- Clone the repository and enter the project:
git clone https://github.com/dannyheskett/wof-solver.git cd wof-solver git checkout v1 - Make a virtual environment and install the pinned packages. The extra index gives the CPU build of torch, which is all phases 10 to 13 need:
uv venv -p 3.12 .venv uv pip install -p .venv/bin/python -r requirements.txt \ --extra-index-url https://download.pytorch.org/whl/cpu --index-strategy unsafe-best-match - Every command on this page runs from the root of the clone with
.venv/bin/python.
Published artifacts. Two things in the pipeline can’t be repeated exactly: the scrape in phase 1 (the site changes) and training on a GPU (not bit-for-bit repeatable). Their outputs are published, so every other phase rebuilds from the same inputs I used:
data/puzzles.csv, the scrape, is in the repository.The checkpoints and the raw Wikidata results are attached to the
v1release:Asset Size What it is pre-ckpt.pt237 MB the pretrained transformer, the start of finetuning ft2b-ckpt.pt237 MB the finetuned transformer behind every result wof.bin118 MB the same weights as float16, for the C server wikidata-raw.tar.gz2.7 MB the Wikidata query results phase 4 reads mkdir -p runs/pre runs/ft2b base=https://github.com/dannyheskett/wof-solver/releases/download/v1 curl -L -o runs/pre/ckpt.pt $base/pre-ckpt.pt curl -L -o runs/ft2b/ckpt.pt $base/ft2b-ckpt.pt
Checking a rebuild. data/SHA256SUMS lists a checksum for every file the pipeline writes, including the large ones that are not in git. After any phase:
cd data && sha256sum -c --ignore-missing SHA256SUMS
Every line should say OK. --ignore-missing skips the files you haven’t built yet.
Phase 1: Fetching the puzzles
Script: 01_fetch_puzzles.py. Reads the wofanswers.com API, writes data/puzzles.csv.
wofanswers.com is a fan site that lists every puzzle from the show and from the mobile game, with a solver page for each category. The site’s pages call a JSON API at /api/collections/wof-answers/items, which takes filters as query parameters and returns rows like:
{"meta": {"total": 13159},
"data": [{"answer": "A CANDLE IN THE WINDOW", "category": "Around The House"}, ...]}
The API has two problems. It returns at most 1,000 rows per call, and its offset paging overlaps: asking for rows 1,000 to 2,000 repeats some rows from the first page and skips others. Paging through it gives a list that looks complete and isn’t.
So the script never pages. It splits the data with filters until every slice fits in one call. It starts with the category, then splits big categories by word count, then by letter count. At each split it checks that the children’s totals add up to the parent’s, so a slice can’t go missing silently:
def collect(filters, total, depth, rows, seen):
"""Collect every row matching filters, splitting until slices fit in one call."""
resp = query(filters)
if resp["meta"]["total"] != total:
raise SystemExit(f"{filters}: total {resp['meta']['total']} != expected {total}")
if total <= LIMIT:
if len(resp["data"]) != total:
raise SystemExit(f"{filters}: got {len(resp['data'])} rows, total says {total}")
rows += resp["data"]
return
if depth == len(SPLITS):
raise SystemExit(f"{filters}: {total} rows and no splits left")
field = SPLITS[depth]
...
for values in (guess, RANGES[field]):
children = []
for v in values:
child = dict(filters, **{field: v})
n = query(child)["meta"]["total"]
if n:
children.append((child, n))
if sum(n for _, n in children) == total:
break
...
for child, n in children:
collect(child, n, depth + 1, rows, seen)
SPLITSis["word_count", "letter_count"], andLIMITis 1,000.guessis the set of split values seen in the slice’s first 1,000 rows. When those don’t cover the total, the script tries every value inRANGES(1 to 20 words, 1 to 80 letters).query()caches each response indata/.cache/fetch/, keyed by a hash of the URL, so a rerun resumes without fetching again. There is a one-second pause between calls.
At the end, the script checks that the rows it collected match the API’s grand total. The API lists 24 puzzles twice, so the script writes only the distinct (answer, category) pairs.
Steps:
- Run the fetch:It prints each category’s row count as it goes, and a final line with the totals.
.venv/bin/python scripts/01_fetch_puzzles.py - You don’t need to run this to rebuild the project. The site changes as new episodes air, so a fresh fetch gives a different file. The
puzzles.csvin the repository is the fetch every later phase used.
Output: data/puzzles.csv, 68,197 rows, 2.1 MB.
answer,category
A CANDLE IN THE WINDOW,Around The House
A CANVAS BACKPACK,Around The House
A JAR OF PENNIES,Around The House
...
FAME AND FORTUNE,Phrase
FAMILIARITY BREEDS CONTEMPT,Phrase
FAMILY DINING EXPERIENCE,Phrase
Phase 2: Cleaning and the control split
Script: 02_clean_puzzles.py. Reads puzzles.csv, writes data/train.csv and data/control.csv.
02_clean_puzzles.py applies five rules, in order:
- Merge category names. Three categories appear under two names.
STAR & ROLEbecomesStar And Role,CharacterbecomesFictional Character, andFictional PlacebecomesPlace. - Drop answers with characters the board can’t show. The board has the 26 letters, space, and
- & ' . ! ? ,. Answers with$,%, digits and the like go. - Drop answers that don’t fit the TV board. The source includes mobile-game puzzles, and some are longer than the TV board allows.
- Drop exact repeats of (answer, category).
- Split off control. An answer goes to control if the SHA-1 of its text falls in the bottom 5%.
The board check is shared with the rest of the project, in wof_common.py:
# The TV board: rows of 12/14/14/12 squares, in use since 1997
BOARD_ROWS = (12, 14, 14, 12)
BOARD_CHARS = set("ABCDEFGHIJKLMNOPQRSTUVWXYZ -&'.!?,")
def pieces(answer):
"""Split into (text, joined) pieces. A row can break between words or
right after a hyphen. joined=True means the piece continues the previous
one with no space, e.g. BLACK-AND-WHITE -> BLACK- AND- WHITE."""
out = []
for word in answer.split():
parts = word.split("-")
for k, part in enumerate(parts):
text = part + ("-" if k < len(parts) - 1 else "")
if text:
out.append((text, k > 0))
return out
def fits_board(answer):
"""Greedy wrap. Filling each row as far as it goes is optimal when the
rows are fixed and in order, so a False here means no layout fits."""
ps = pieces(answer)
i = 0
for width in BOARD_ROWS:
used = 0
while i < len(ps):
text, joined = ps[i]
need = len(text) + (1 if used and not joined else 0)
if used + need > width:
break
used += need
i += 1
return i == len(ps)
The greedy wrap is enough. The rows have fixed widths and come in a fixed order, so pushing as many words as fit onto each row never makes a later row worse. If the greedy wrap runs out of rows, no layout fits.
The control split is one function:
CONTROL_PCT = 5
def in_control(answer):
h = int(hashlib.sha1(answer.encode()).hexdigest(), 16)
return h % 100 < CONTROL_PCT
The split uses only the answer text. That has two consequences:
- It is the same on every run and on every machine. There is no random seed to record.
- An answer filed under two categories lands on one side. GULLIVER’S TRAVELS is listed under both Book Title and Title. If the two copies were split separately, one could train a solver while the other tests it. The script ends by asserting that no answer is on both sides.
Steps:
- Run it:
.venv/bin/python scripts/02_clean_puzzles.py - It prints a report of every drop:
read 68197 dropped 133 characters dropped 26 too long for board bad characters: "$%()*/169:; train 64689 -> data/train.csv control 3349 -> data/control.csv (3349 distinct answers)
Output: data/control.csv, 3,349 puzzles, 108 KB, and data/train.csv with 64,689 puzzles until phase 3 runs.
answer,category
A ROLL OF FILM,Around The House
ACCENT FURNITURE,Around The House
...
Phase 3: Dev split
Script: 03_split_dev.py. Reads train.csv, writes data/train.csv and data/dev.csv.
transformer-solve has one setting to tune: how sure it must be before it solves. Tuning that on the control puzzles would leak them into the result, so phase 3 holds 300 more answers out of training, as the dev set.
The dev answers are the 300 with the lowest SHA-1 of dev|<answer>. The dev| prefix makes this hash independent of the control hash, so the two splits don’t line up.
DEV_SIZE = 300
def key(answer):
return hashlib.sha1(f"dev|{answer}".encode()).hexdigest()
def main():
pool = wc.read_puzzles("train.csv")
if (wc.DATA / "dev.csv").exists():
pool += wc.read_puzzles("dev.csv")
pool = sorted(set(pool), key=lambda p: (p[1], p[0])) # dev rows may be in both
dev_answers = set(sorted({a for a, _ in pool}, key=key)[:DEV_SIZE])
sides = {"train.csv": [], "dev.csv": []}
for answer, category in pool:
sides["dev.csv" if answer in dev_answers else "train.csv"].append((answer, category))
The script reads train.csv and any existing dev.csv back together before splitting. That makes it safe to run twice: a second run gives the same two files.
Steps:
- Run it:It prints:
.venv/bin/python scripts/03_split_dev.pytrain.csv 64389 puzzles dev.csv 300 puzzles
Output:
| File | Rows | Size |
|---|---|---|
data/train.csv | 64,389 | 2.0 MB |
data/dev.csv | 300 | 12 KB |
data/control.csv | 3,349 | 108 KB |
answer,category
ANTIQUE POLAROID CAMERA,Around The House
BAMBOO PLACE MATS,Around The House
...
From here on, every phase reads these three files and never changes them.
The three puzzle files
The end state:
control.csvis held out from everything. Neither solver learns from it, and every result on this page is measured on it.dev.csvis held out from the transformer’s training. It is used only to tune one setting, the confidence transformer-solve needs before it solves.train.csvis what both solvers learn from. smart-solve also addsdev.csvto its word list, since it has nothing to tune.
puzzles.csv and the three split files have the same layout: CSV with a header row and two columns, answer and category, sorted by category and then answer.
Four categories make up 48.6% of the training puzzles: Thing, Food And Drink, Phrase and What Are You Doing. Results on this page are over all 3,349 control puzzles, so those four carry the most weight.
Phase 4: Category lists
The training puzzles cover 64,389 answers. That is a lot of Wheel of Fortune, but not much of the world. A held-out puzzle like a film title or a city the model has never seen can still be solved if the model knows that the name exists. Phase 4 builds lists of real names and phrases for the show’s categories from four open sources. They feed two things later: the transformer’s LIST training lines in phase 8, and the wordplay categories.
| Script | Reads | Writes | Lines |
|---|---|---|---|
fetch_wikidata.py | the QLever Wikidata endpoint | data/sources/raw/wikidata/*.json (42 files) | |
build_wikidata.py | the 42 JSON files | data/sources/wikidata.tsv | 116,403 |
build_wiktionary.py | the kaikki.org Wiktionary extract | data/sources/wiktionary.tsv | 21,487 |
build_wordnet.py | Open English WordNet 2025 | data/sources/wordnet.tsv | 38,256 |
build_wordplay.py | the three lists above, the puzzles, the CMU Pronouncing Dictionary | data/sources/wordplay.tsv | 47,032 |
All four outputs have the same layout: tab-separated category and phrase, no header, sorted. The four TSVs are in the repository. The raw inputs, with their URLs and checksums, are listed in data/sources/SOURCE.md.
One writer for every list
Every builder produces (category, phrase) pairs and hands them to write() in src_common.py. That one function applies the board rules to all four lists:
def normalize(text):
"""Board form of a phrase, or None if it holds characters a board cannot
show (digits, symbols, other scripts). Parenthesised notes are dropped:
"Titanic (1997 film)" becomes TITANIC."""
text = PARENS.sub("", text.translate(TRANSLATE))
text = unicodedata.normalize("NFKD", text)
text = "".join(ch for ch in text if not unicodedata.combining(ch)).upper()
text = SPACES.sub(" ", text).strip(" ,.")
if not text or set(text) - BOARD_CHARS or not any(ch.isalpha() for ch in text):
return None
return text
def write(name, pairs):
control = control_solutions()
seen = set()
counts = Counter()
dropped = Counter()
for category, phrase in pairs:
p = normalize(phrase)
if p is None:
dropped["characters"] += 1
elif not fits_board(p):
dropped["too long for board"] += 1
elif p in control:
dropped["control solution"] += 1
elif (category, p) in seen:
dropped["repeat"] += 1
else:
seen.add((category, p))
counts[category] += 1
normalize() does the following:
- strips accents, so É becomes E.
- maps the letters that don’t decompose (ß to SS, Ł to L), and curly quotes to straight ones.
- drops parenthesised notes.
- uppercases the text.
A phrase with anything the board can’t show (digits, symbols, other scripts) is dropped. So is one that doesn’t fit the 12/14/14/12 board.
The control check is the important line. Any phrase that is a control answer is dropped, from every list. Wikidata knows thousands of film titles, and some of them are control puzzles. Without this line, the transformer would study the test.
Wikidata: names, titles and places
Wikidata is the structured database behind Wikipedia, and it is CC0. fetch_wikidata.py runs 42 SPARQL queries against QLever, a fast public mirror of Wikidata. Each query returns the English label of every item of a class, with its sitelink count: the number of Wikipedia language editions with an article on it. The script uses that count as a fame filter, so obscure items are cut.
QUERIES = {
# titles
"films": simple(["Q11424"], 10),
"tv_series": simple(["Q5398426"], 8),
"novels": simple(["Q8261", "Q7725634"], 10),
"songs": simple(["Q7366", "Q134556"], 8),
# people and groups
"actors": people("Q33999", 30),
"singers": people("Q177220", 30),
...
# joins on the lists above, so they run last
"books_authors": label_of("novels", "wdt:P50", "author", min_links=15),
"songs_artists": label_of("songs", "wdt:P175", "artist"),
...
}
simple(["Q11424"], 10) means: every item that is an instance of film (Q11424) with an English label and at least 10 sitelinks. people("Q33999", 30) is every human whose occupation is actor, with at least 30. A few lists are joins on earlier ones. books_authors takes the novels already fetched and asks for each one’s author.
QLever stops any query at 30 seconds. A join over a whole class (every novel with its author) runs longer than that, so a join is fetched in two steps. First comes the base list, then the linked values for 500 items at a time in a VALUES block. Each list is saved as raw JSON:
[{"item": "http://www.wikidata.org/entity/Q1000394", "label": "This Modern Age", "links": "10"},
{"item": "http://www.wikidata.org/entity/Q1000826", "label": "Guns of the Magnificent Seven", "links": "14"},
...]
build_wikidata.py maps each raw file to show categories and formats the joins the way the show writes them:
| Category | Built from | Example line |
|---|---|---|
| Movie Title | films | THE GENERAL'S DAUGHTER |
| On The Map | countries, states, cities, rivers, lakes, islands… and US cities as CITY STATE | ELKHART KANSAS |
| Star And Role | film cast with character names, only for actors in the people lists | HEATH LEDGER AS NED KELLY |
| Title/Author | novels joined to their authors | TITLE BY AUTHOR |
| Song Artist | songs joined to their performers | SONG BY ARTIST |
| Husband And Wife | spouses, both in the people lists | shared last names go once: BETTINA & LUDWIG ACHIM VON ARNIM |
| Proper Name, Show Biz | actors, singers, writers, athletes, brands, companies, bands, teams |
Steps:
- The Wikidata endpoint changes daily, so a fresh fetch gives different results. Use the published raw files to rebuild exactly:
mkdir -p data/sources/raw curl -L https://github.com/dannyheskett/wof-solver/releases/download/v1/wikidata-raw.tar.gz \ | tar -xz -C data/sources/raw - Or fetch your own, with the result differing from mine. The script pauses 20 seconds between queries, and runs the joins in batches of 500 items:
.venv/bin/python scripts/04_category_lists/fetch_wikidata.py - Build the list:
.venv/bin/python scripts/04_category_lists/build_wikidata.pywikidata: 116,403 lines -> data/sources/wikidata.tsv 21,697 On The Map 21,294 Movie Title 20,412 Proper Name 9,571 Show Biz 7,709 Star And Role ... dropped: too long for board 1,395, characters 4,354, repeat 19,796, control solution 407
The 407 dropped control solutions are puzzles from the held-out set that Wikidata also knows.
Wiktionary: phrases and “-ing” activities
Two of the show’s biggest categories are Phrase and What Are You Doing (TAKING A WALK, BAKING COOKIES). Wiktionary has both, as multiword entries. kaikki.org publishes the English Wiktionary as one JSON object per entry, 3.3 GB in all. build_wiktionary.py reads it one line at a time:
PHRASE_POS = {"phrase", "proverb", "prep_phrase", "intj"}
IDIOM_POS = {"noun", "adj", "adv"}
DEAD = {"obsolete", "archaic", "form-of", "alt-of", "misspelling"}
BAD = {"vulgar", "offensive", "derogatory", "slur", "ethnic"}
...
if " " not in word or not senses:
continue
if all(t & DEAD for t in tags) or any(t & BAD for t in tags):
continue
if pos in PHRASE_POS or (pos in IDIOM_POS and any("idiomatic" in t for t in tags)):
if is_common(word):
phrases.append(word)
elif pos == "verb":
verbs.append(word)
- Phrase: multiword phrases, proverbs and interjections, plus idiomatic nouns, adjectives and adverbs.
- What Are You Doing: multiword verbs (“take a walk”), with the first word turned into its present participle. The participle comes from the verb’s own Wiktionary entry when it has one (
taketotaking), and from spelling rules otherwise. - Dropped:
- entries where every sense is obsolete or archaic, and any entry with a sense tagged vulgar or offensive.
- entries with a word missing from the SCOWL word list, which drops Latin and other foreign phrases.
- verb phrases with “someone” or “something” in them.
- Rewritten: “one’s” becomes MY and “oneself” becomes MYSELF, the way the show writes them: GUARDING MY TONGUE.
Steps:
- Download the extract (3.3 GB). kaikki.org rebuilds it from new Wiktionary dumps, so today’s file will differ from mine. The
wiktionary.tsvin the repository is the one every later phase used.curl -L -o data/sources/raw/kaikki-english.jsonl \ https://kaikki.org/dictionary/English/kaikki.org-dictionary-English.jsonl - Build it:
.venv/bin/python scripts/04_category_lists/build_wiktionary.pyparticiples 41,187, phrases 9,518, verb phrases 12,282 wiktionary: 21,487 lines -> data/sources/wiktionary.tsv 12,224 What Are You Doing 9,263 Phrase dropped: repeat 157, control solution 63, too long for board 93
Sample lines:
What Are You Doing BARING MY BREAST
What Are You Doing COMING BACK FROM THE DEAD
What Are You Doing GUARDING MY TONGUE
WordNet: everyday nouns
The show’s Thing, Around The House and In The Kitchen puzzles are everyday nouns. Open English WordNet organises nouns into a tree of synsets (groups of synonyms), each filed in a lexicographer file such as noun.food or noun.artifact. build_wordnet.py takes categories from whole files, or from everything below a root synset:
LEXFILES = {
"noun.food": ["Food And Drink"],
"noun.animal": ["Living Thing"],
"noun.plant": ["Living Thing"],
"noun.artifact": ["Thing"],
"noun.object": ["Thing"],
"noun.person": ["Person"],
}
# (category, lexfile, lemma of the root synset[, word in its definition])
ROOTS = [
("Occupation", "noun.person", "professional"),
...
("Around The House", "noun.artifact", "furniture"),
("Around The House", "noun.artifact", "home appliance"),
...
("In The Kitchen", "noun.artifact", "kitchen utensil"),
("In The Kitchen", "noun.artifact", "cutlery", "eating"),
...
]
A root can carry a word from its definition, to pick one sense of an ambiguous lemma (a dictionary headword): “cutlery” with “eating” is the knives and forks, not the cutting tools. Capitalised lemmas (proper names, Latin genus names) are skipped, as are lemmas with any word outside the SCOWL list. People is built from the Person lemmas with simple plural rules.
Steps:
- Download WordNet 2025 (11 MB):
curl -L -o data/sources/raw/english-wordnet-2025.xml.gz \ https://github.com/globalwordnet/english-wordnet/releases/download/2025-edition/english-wordnet-2025.xml.gz - Build it:It prints each root with its synset and lemma counts, then the totals:
.venv/bin/python scripts/04_category_lists/build_wordnet.pywordnet: 38,256 lines -> data/sources/wordnet.tsv 12,472 Thing 7,627 Living Thing 5,978 Person 5,944 People 2,430 Food And Drink ...
Wordplay: the joined categories
Four of the show’s categories are wordplay on other phrases. build_wordplay.py makes them from the three lists above plus the training puzzles, so it runs last:
| Category | Rule | Example |
|---|---|---|
| Same Letter | every word starts with the same letter | CREEPY CRAWLY CREATURES |
| Rhyme Time | two different words of 3+ letters rhyme | HIKING AND BIKING |
| Before And After | phrase A’s last word starts phrase B, joined on it | RUBBER BAND + BAND OF BROTHERS, so RUBBER BAND OF BROTHERS |
| Same Name | two phrases ending in the same word | BOWLING BALL + DEBUTANTE BALL, so BOWLING & DEBUTANTE BALL |
Rhymes come from the CMU Pronouncing Dictionary. Two words rhyme when their sounds match from the last stressed vowel on:
def load_rhymes():
rhyme = {}
for line in open(RAW / "cmudict.dict", encoding="latin-1"):
word, *phones = line.split("#")[0].split()
word = re.sub(r"\(\d+\)$", "", word).upper()
stressed = [i for i, p in enumerate(phones) if p[-1] in "12"]
if stressed and word not in rhyme:
rhyme[word] = " ".join(p.rstrip("012") for p in phones[stressed[-1]:])
return rhyme
HIKING is HH AY1 K IH0 NG, so its rhyme key is AY K IH NG. BIKING has the same key.
The two joins give millions of candidates (2,847,612 for Before And After). The script shuffles them with a fixed seed, random.Random("wof-wordplay-v1"), and keeps 20,000 of each. The fixed seed makes every run give the same file. The joins are mechanical, so many are odd (CUTTING RED TAPE RECORDING, A BRISK & DAYTIME JOG). They teach the format of the category more than real phrases.
Steps:
- Download the CMU dictionary, pinned to the commit used:
curl -L -o data/sources/raw/cmudict.dict \ https://raw.githubusercontent.com/cmusphinx/cmudict/0f8072f814306c5ee4fbf992ed853601b12c01f9/cmudict.dict - Build it:
.venv/bin/python scripts/04_category_lists/build_wordplay.py194,826 phrases, 126,020 pronunciations candidates: Before And After 2,847,612, Same Name 44,342 wordplay: 47,032 lines -> data/sources/wordplay.tsv 19,915 Before And After 18,357 Same Name 7,297 Same Letter 1,463 Rhyme Time
smart-solve
smart-solve is the opponent, and the transformer’s teacher: its moves label the training data in phases 6 and 7. It has no training step. It is code in two shared modules, wof_common.py (the word list and the letter choice) and solvers.py (the game moves).
The word list
smart-solve’s vocabulary is two things together:
- The SCOWL word list, size 50, from Debian’s
wamericanpackage. It has 104,334 words, indata/wordlist/american-english.txt. - Every word of
train.csvanddev.csv.
Words are grouped by shape: their length, with apostrophes in place. CAN'T has shape ___'_. For each shape the code builds a matrix of position bitmasks. P[i, c] has bit k set when word i has letter c at position k:
def posmasks(word):
row = [0] * 26
for k, ch in enumerate(word):
if ch in IDX:
row[IDX[ch]] |= 1 << k
return row
For FILM, the row has F = 0001, I = 0010, L = 0100, M = 1000, and 0 for every other letter. That one integer per letter answers the question the solver asks over and over: “if this letter is called, which squares light up?”
Candidates
A board word’s candidates are the words of its shape that match every called letter exactly: the letter shows where it shows, and nowhere else. A called letter that missed must not appear at all. With the bitmasks, that is one comparison per called letter:
if cols:
want = np.array(wc.posmasks(key), dtype=np.int32)[cols]
idx = np.nonzero((P[:, cols] == want).all(axis=1))[0]
key is the board word as shown, such as F_LM, and cols are the called letters. For F_LM with A F L M O S called, a candidate must have F at position 0, L at 2 and M at 3, no other F, L or M, and no A, O or S, since those were called and don’t show in this word.
Weights
Each candidate weighs 1 plus how often it appears in training puzzles of the board’s category. This is how the category steers the choice. In Around The House, LAMP appears in 18 training puzzles and LIMP in none, so LAMP weighs 19 and LIMP weighs 1.
def weights(self, category, sh):
key = (category, sh)
if key not in self._weights:
counts = self.counts[category]
self._weights[key] = np.array(
[1.0 + counts.get(w, 0) for w in self.words[sh]])
return self._weights[key]
Choosing a letter
The core of smart-solve is one score per letter: the expected log2 of the number of candidates left after calling it, summed over the board’s words. A letter that splits the candidates into many small groups scores low. A letter that leaves them in one big group scores high. smart-solve calls the letter with the lowest score.
P = self.v.P[it["shape"]][idx]
w = self.v.weights(self.category, it["shape"])[idx]
total = w.sum()
out = {}
for letter in letters:
_, inv, n = np.unique(P[:, IDX[letter]], return_inverse=True,
return_counts=True)
W = np.bincount(inv, weights=w)
out[letter] = float((W / total * np.log2(n)).sum())
For each letter, P[:, IDX[letter]] is the bitmask of where that letter sits in each candidate. Candidates with the same bitmask would look the same after the call, so they form a group. n is each group’s size and W its total weight. The score is the weighted average of log2(n): the bits of uncertainty left. Ties go to the more common letter.
Here are the scores on the opening board of A ROLL OF FILM, _ ____ __ ____ in Around The House. Before any call, the four words hold 4.70, 11.76, 8.27 and 11.76 bits of uncertainty, 36.49 in all. Each column is the expected bits left in one word after the call:
| Letter | _ | ____ | __ | ____ | Total |
|---|---|---|---|---|---|
| S | 4.54 | 10.30 | 7.73 | 10.30 | 32.85 |
| L | 4.54 | 10.49 | 7.84 | 10.49 | 33.35 |
| R | 4.54 | 10.57 | 7.73 | 10.57 | 33.39 |
| T | 4.54 | 10.62 | 7.67 | 10.62 | 33.46 |
| N | 4.54 | 10.72 | 7.53 | 10.72 | 33.51 |
| D | 4.54 | 10.83 | 7.75 | 10.83 | 33.95 |
| Z | 4.54 | 11.55 | 8.09 | 11.55 | 35.72 |
S wins because it splits the four-letter words best. It appears in many of them, in many different positions, so seeing where it lands (or that it doesn’t) tells the most. N would do better on the two-letter word, but the two four-letter words count twice. Z barely moves anything. Vowels aren’t in the table because none can be bought yet. The one-letter word scores the same for every consonant: its 26 candidates are the single letters, and each consonant matches exactly one of them.
Two rules shape the choice:
- The money rule. Until a called consonant shows on the board, vowels are left out.
- When to solve. smart-solve solves when every hidden word has exactly one candidate, and fills each word with it.
A game
Here is smart-solve playing the control puzzle A ROLL OF FILM, in Around The House. The candidate counts are per hidden word, read from SmartSolver.view() at each turn. A word drops out of the list once it is fully shown:
board called candidates per hidden word move
_ ____ __ ____ [26, 3473, 308, 3473] CALL S
_ ____ __ ____ S [25, 2437, 275, 2437] CALL L
_ __LL __ __L_ LS [24, 47, 252, 159] BUY A
A __LL __ __L_ ALS [38, 214, 105] BUY O
A _OLL O_ __L_ ALOS [6, 14, 63] CALL F
A _OLL OF F_L_ AFLOS [6, 3] CALL M
A _OLL OF F_LM AFLMOS [5, 1] CALL R
A ROLL OF F_LM AFLMORS [1] SOLVE A ROLL OF FILM
S misses, but it still helps: 1,036 four-letter words with an S drop out of each four-letter slot. L hits in two words, which cuts ____ from 2,437 candidates to 47. Once L shows, vowels can be bought. After seven letters, the last word is F_LM with one candidate, and smart-solve solves. In data/eval/smart-solve.tsv this game is:
main Around The House A ROLL OF FILM solved letters=7 buys=2 misses=1 order=SLAOFMR
The bonus round
In the bonus round the solver names four letters at once, with no feedback between them. best_set tries every combination of 3 consonants and 1 vowel that are not R S T L N E. For each combination it computes the chance that filling every word with its heaviest candidate will be right once those letters show. It keeps the best:
for cs in itertools.combinations(consonants, n_cons):
for vs in itertools.combinations(vowels, n_vow):
chosen = cs + vs
tie = sorted(rank[c] for c in chosen)
p = 1.0
for codes, w, total in live:
key = np.zeros(len(w), dtype=np.int64)
for c in chosen:
inv, n = codes[c]
key = key * n + inv
first = np.unique(key, return_index=True)[1]
p *= w[first].sum() / total
That is 16 consonants choose 3 (Y counts as a consonant) times 4 vowels: 2,240 combinations per puzzle. For each word, key gives every candidate a code for how the four letters would light it up. Candidates with the same code can’t be told apart. Within a group the solver will guess the heaviest, so the chance of being right is the weight of each group’s heaviest member over the total. Words are taken as independent, so the chances multiply.
On A ROLL OF FILM:
_ R_LL __ __L_ PICK CKM O -> _ ROLL O_ __LM SOLVE A ROLL OF FILM
What smart-solve can’t do
smart-solve only knows words in its list. 163 of the 10,649 words in the control puzzles (1.5%) are not in it, in 159 puzzles. On those, it either never pins the word, or pins a wrong word that happens to fit. All 57 of its wrong MAIN solves come from those puzzles. Some missing words are typos in the source data (HADRIN, LASANGA, CHEDDER). It also has no idea which phrases are real. To smart-solve, a board is a set of independent words.
Phase 5: Pretraining corpus
A model that only ever sees Wheel of Fortune lines has to learn English from 64,000 short phrases. Pretraining gives it English first: spelling, common words, which words follow which. Phase 5 builds that text, one billion characters, from two open datasets on Hugging Face.
| Script | Reads | Writes |
|---|---|---|
05_make_corpus.py | 5 Parquet files in data/corpus/raw/ | data/corpus/train.txt, data/corpus/val.txt |
- wikimedia/wikipedia, the
20231101.endump, files 0, 10, 20 and 30 of 41. CC BY-SA 4.0. - HuggingFaceFW/fineweb-edu, the
sample/10BTfile000_00000: web pages filtered for educational content. ODC-By 1.0.
Half comes from each. Wikipedia is clean, but it reads like an encyclopedia. FineWeb-Edu adds plainer, more everyday writing.
Cleaning
The model will only ever see board characters, so the corpus is cut down to them:
def clean_line(line):
line = unicodedata.normalize("NFKD", line.translate(TRANSLATE))
line = "".join(ch for ch in line if not unicodedata.combining(ch)).upper()
line = SPACES.sub(" ", NOT_BOARD.sub(" ", line))
line = REPEAT.sub(r"\1", STRAY.sub(r"\1", line)).strip(" ,.")
letters = sum(ch.isalpha() for ch in line)
return line if letters >= MIN_LETTERS else None
- Accents are stripped and the text is uppercased, as in phase 4.
- Every character outside
A-Z, space and- & ' . ! ? ,becomes a space. That removes digits, brackets and symbols. STRAYandREPEATtidy the punctuation left behind where numbers were. “IN 1997, THE” would become “IN , THE”, and this fixes it to “IN, THE”.- Paragraphs with fewer than 40 letters are dropped. That removes headings, captions and table rows.
- Wikipedia articles stop at their References, See also, External links or Notes section.
Interleaving and the validation split
The script reads both sources at once and always takes the next document from whichever source has contributed fewer characters so far, so they stay at half each. Each source stops at 500 million characters. About 1% of documents go to val.txt instead, chosen by the SHA-1 of the text, so the split is the same on every run:
h = int(hashlib.sha1(doc.encode()).hexdigest(), 16) % 100
if h < VAL_PCT:
val.write(doc)
Steps:
- Download the five Parquet files (3.1 GB). The URLs are pinned to the dataset commits I used:Their checksums are in
mkdir -p data/corpus/raw && cd data/corpus/raw W=https://huggingface.co/datasets/wikimedia/wikipedia/resolve/b04c8d1ceb2f5cd4588862100d08de323dccfbaa/20231101.en for n in 00000 00010 00020 00030; do curl -L -o wiki-$n.parquet "$W/train-$n-of-00041.parquet" done curl -L -o fineweb-000_00000.parquet \ "https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/sample/10BT/000_00000.parquet" cd ../../..data/corpus/SOURCE.md. - Build the corpus:
.venv/bin/python scripts/05_make_corpus.pywiki 132,888 documents 500,005,177 characters web 110,792 documents 500,001,314 characters train 1,000,006,491 characters, val 10,141,309 characters
Output:
| File | Size |
|---|---|
data/corpus/train.txt | 954 MB |
data/corpus/val.txt | 9.7 MB |
One paragraph per line, with a blank line between documents:
ANARCHISM IS A POLITICAL PHILOSOPHY AND MOVEMENT THAT IS SKEPTICAL OF ALL JUSTIFICATIONS FOR
AUTHORITY AND SEEKS TO ABOLISH THE INSTITUTIONS IT CLAIMS MAINTAIN UNNECESSARY COERCION AND
HIERARCHY, TYPICALLY INCLUDING NATION-STATES, AND CAPITALISM. ...
Phase 6: MAIN game states
Script: 06_make_main.py. Reads train.csv and control.csv, writes data/main_train.tsv and data/main_control.tsv.
The transformer learns to play by reading game states with the right move after each one. Phases 6 to 8 write those states. smart-solve plays every training puzzle many ways, and every board along the way becomes one line, labelled with the move smart-solve would make there.
The _control files hold the same kinds of lines for the control puzzles. They are used only to measure validation loss during training, never to train.
All four TSVs share one layout, tab-separated, no header:
category board called target solution
called is the letters called so far, sorted A-Z, or - before any call. solution is kept for checking and is not shown to the model.
A solver at play will miss letters smart-solve wouldn’t call, and land on boards smart-solve’s own games never pass through. If the training data only had smart-solve’s own games, the model would know what to do on the perfect path and nothing off it. So each puzzle is walked five ways:
- three times in a random letter order.
- once in order of letter frequency.
- once following smart-solve’s own picks.
Every board a walk passes through becomes one line. The label is always the move smart-solve would make on that board, whatever the walk calls next:
def walk(bw, solution, category, policy, rng):
called = set()
order = rng.sample(wc.LETTERS, 26) if policy == "random" else ORDER
while True:
b = wc.board(solution, called)
c = "".join(sorted(called)) or "-"
if bw.solved():
yield f"{category}\t{b}\t{c}\tSOLVE {solution}\t{solution}\n"
return
vowels = wc.can_buy(b, called)
best = bw.smart_pick(called, ORDER, vowels)
yield f"{category}\t{b}\t{c}\t{wc.move(best)}\t{solution}\n"
nxt = best if policy == "smart" else next(
x for x in order if x not in called and (vowels or x not in wc.VOWELS))
called.add(nxt)
bw.call(nxt)
bestis smart-solve’s pick on this board, and it becomes the line’s target.nxtis what the walk actually calls. That isbeston the smart walk, and the next letter in its own order on the others.- A walk ends at its first SOLVE line. Lines repeated within a puzzle are dropped.
- Each puzzle’s random orders are seeded from
SEED = "wof-main-v1"and the puzzle itself, so every run writes the same file.
Here are lines from one random walk through A CANDLE IN THE WINDOW. The walk calls J, K, V and U, which miss. The label stays CALL N, because N is still the best move on that board:
Around The House _ ______ __ ___ ______ - CALL S A CANDLE IN THE WINDOW
Around The House _ ____L_ __ ___ ______ L BUY A A CANDLE IN THE WINDOW
Around The House A _A__L_ __ ___ ______ AL BUY E A CANDLE IN THE WINDOW
Around The House A _A__L_ I_ ___ _I____ AIL CALL N A CANDLE IN THE WINDOW
Around The House A _A__L_ I_ ___ _I____ AIKL CALL N A CANDLE IN THE WINDOW
Around The House A _A__L_ I_ ___ _I____ AIJKL CALL N A CANDLE IN THE WINDOW
Around The House A _A__L_ I_ ___ _I____ AIJKLV CALL N A CANDLE IN THE WINDOW
Around The House A _A__L_ I_ ___ _I____ AIJKLUV CALL N A CANDLE IN THE WINDOW
Across all 5,131,187 lines in main_train.tsv, the labels are:
| Target | Lines |
|---|---|
| CALL | 2,802,502 |
| BUY | 2,007,153 |
| SOLVE | 321,532 |
The most common single labels are BUY E (970,958), BUY A (558,166), CALL S (503,917), CALL T (371,377) and CALL R (360,841). That skew comes straight from smart-solve. E is the most common letter in the puzzles, and buying it as soon as the money rule allows is smart-solve’s most frequent move. The model will inherit this habit.
The labels for the training puzzles use a word list built from train.csv alone, with weights from train.csv alone. Neither dev nor control shapes a training label.
Steps:
- Run it. It uses 11 worker processes:It prints progress every 5,000 puzzles and a summary line per file.
.venv/bin/python scripts/06_make_main.py
Output:
| File | Lines | Size |
|---|---|---|
data/main_train.tsv | 5,131,187 | 349 MB |
data/main_control.tsv | 267,216 | 19 MB |
Phase 7: BONUS picks and solves
Script: 07_make_bonus.py. Reads train.csv and control.csv, writes data/bonus_train.tsv and data/bonus_control.tsv.
The bonus round needs two kinds of lines:
- PICK lines. The board after R S T L N E, labelled with smart-solve’s
best_setpick. There is one per puzzle. - SOLVE lines. The board after R S T L N E and a set of four letters, labelled with the answer. There are up to 10 per puzzle, from different sets: smart-solve’s set, the four most common letters left, and random sets.
lines = [f"{category}\t{wc.board(solution, FREE)}\t{free}\t{pick_text(smart)}\t{solution}\n"]
sets = [frozenset(smart),
frozenset([c for c in ORDER if c in CONS][:3] + [c for c in ORDER if c in VOWS][:1])]
tries = 0
while len(set(sets)) < PICK_SETS and tries < 100:
sets.append(frozenset(rng.sample(CONS, 3) + rng.sample(VOWS, 1)))
tries += 1
Training on boards from many different sets, not just smart-solve’s set, teaches the model to read a board with any letters showing. In play, transformer-solve makes its own pick, and that may not be smart-solve’s.
Around The House _ __N_LE _N T_E __N___ ELNRST PICK CDM I A CANDLE IN THE WINDOW
Around The House _ C_NDLE IN T_E _IND__ CDEILMNRST SOLVE A CANDLE IN THE WINDOW A CANDLE IN THE WINDOW
Around The House A CANDLE _N THE __ND__ ACDEHLNRST SOLVE A CANDLE IN THE WINDOW A CANDLE IN THE WINDOW
Around The House _ __N_LE IN T_E WIN__W BEIJLNRSTW SOLVE A CANDLE IN THE WINDOW A CANDLE IN THE WINDOW
Steps:
- Run it:
.venv/bin/python scripts/07_make_bonus.py
Output:
| File | Lines | Size |
|---|---|---|
data/bonus_train.tsv | 708,279 | 59 MB |
data/bonus_control.tsv | 36,839 | 3.1 MB |
Phase 8: Token files
Script: 08_make_train.py. Reads the four TSVs, the category lists and the corpus, writes data/train/.
The model reads characters, not words. Its alphabet has 48 characters, and a character’s token id is its position in this string:
ALPHABET = "\n !&',-./0123456789?ABCDEFGHIJKLMNOPQRSTUVWXYZ_|"
That is newline, space, the board’s punctuation, digits, ?, the 26 letters, _ for a hidden square and | as a field separator. The digits and / are never used after phase 5’s cleaning, but they keep the alphabet general.
Every game line becomes a prompt and a target, ending in a newline:
MAIN|AROUND THE HOUSE|A _A__L_ __ ___ ______|AL|BUY E
BONUS|AROUND THE HOUSE|_ __N_LE _N T_E __N___|ELNRST|PICK CDM I
BONUS|AROUND THE HOUSE|A CANDLE _N THE __ND__|ACDEHLNRST|SOLVE A CANDLE IN THE WINDOW
LIST|ON THE MAP|ELKHART KANSAS
The prompt is everything up to the last |. Training puts loss only on the target and the newline. The model is never trained to predict the board it is given, only the move.
Phase 8 writes five training streams and two validation streams. Each one is a token file (one byte per character) plus an index of (start, prompt length, line length):
def add(self, prompt, target):
line = prompt + target + "\n"
self.index.append((self.pos, len(prompt), len(line)))
self.pos += len(line)
self.buf.append(line)
| Stream | Lines | Tokens | From |
|---|---|---|---|
main | 5,131,187 | 283M | main_train.tsv |
bonus | 708,279 | 52M | bonus_train.tsv |
list | 281,171 | 10M | every training puzzle and every category-list phrase, as `LIST |
list_bonus | 280,147 | 20M | the same phrases as BONUS SOLVE lines, with R S T L N E plus a seeded random 3 consonants and 1 vowel shown |
corpus | 1,000M | corpus/train.txt, the text as token ids, loss on everything | |
main_val, bonus_val | 304,055 | 17M | the control TSVs, for validation loss only |
corpus_val | 10M | corpus/val.txt |
The two list streams carry phase 4 into the model. list teaches which phrases exist in each category. list_bonus is bonus-round practice on 280,147 phrases, most of which are not puzzles. Any phrase that is a dev answer is left out of both, so the dev set stays clean for tuning.
Steps:
- Run it:It prints each stream’s lines and tokens.
.venv/bin/python scripts/08_make_train.py
Output: data/train/, 1.5 GB, 15 files:
alphabet.txt
main.bin main.idx.npy main_val.bin main_val.idx.npy
bonus.bin bonus.idx.npy bonus_val.bin bonus_val.idx.npy
list.bin list.idx.npy
list_bonus.bin list_bonus.idx.npy
corpus.bin corpus_val.bin
The transformer
model.py is a small GPT in the style of Andrej Karpathy’s nanoGPT: a decoder-only transformer that predicts the next character from the ones before it.
How a transformer reads a game line
For anyone new to transformers, here is what happens inside this one when it reads a prompt like MAIN|AROUND THE HOUSE|A _A__L_ __ ___ ______|AL|.
- Characters become tokens. Each character is replaced by its index in the 48-character alphabet. M is 32, A is 20,
|is 47, and so on. The prompt becomes a list of 48 numbers, one per character. - Tokens become vectors. The token embedding is a table with one row of 640 numbers per character. Each token is replaced by its row. A second table, the position embedding, has one row per position from 0 to 255. The row for that position is added in, so the model knows where each character sits. The prompt is now 48 vectors of 640 numbers.
- Attention mixes information between positions. In each block, every position builds three vectors from its own: a query, a key and a value. A position compares its query with the key of every earlier position, and the closer the match, the more of that position’s value it takes in. This is how the
_squares of the board can pick up information from the category at the start of the line and the called letters at the end. The causal mask means a position can look back but never forward. Ten heads do this in parallel, each with its own 64-number slice, so different heads can track different things. - The MLP works on each position alone. After attention, each vector goes through a two-layer network that widens it to 2,560 numbers, applies GELU, and narrows it back to 640. Attention moves information between positions. The MLP transforms it.
- The residual stream carries everything forward. Attention and the MLP don’t replace a position’s vector. They add to it. After 12 blocks, each vector is the sum of its starting embedding and 24 updates.
- The last vector predicts the next character. The vector at the last position, after a final LayerNorm, is compared with each of the 48 token embeddings by dot product. That gives 48 scores, the logits. Softmax turns them into probabilities.
After the prompt above, the model writes B, U, Y, a space and E, then a newline: BUY E, the same move smart-solve makes there. Each written character is fed back in as the next position.
Training adjusts the 59 million numbers so the probability of the right next character goes up. The loss is the negative log of that probability, averaged over the characters that count. A loss of 0.14 nats on main_val means the model gives the right target character a probability of about 87% as a geometric mean: e^−0.14 = 0.87.
The shape
| Setting | Value |
|---|---|
| Layers | 12 |
| Attention heads | 10, of 64 dimensions each |
Width (n_embd) | 640 |
Context (block) | 256 characters |
| Vocabulary | 48 characters |
| Parameters | 59,045,120 (59.05M), not counting position embeddings |
A 256-character context is more than any game line needs. The longest line in the training streams is 158 characters, a MAIN line with its prompt and target.
The block
class Block(nn.Module):
def __init__(self, c):
super().__init__()
self.ln1, self.ln2 = nn.LayerNorm(c.n_embd), nn.LayerNorm(c.n_embd)
self.attn = Attention(c)
self.mlp = nn.Sequential(nn.Linear(c.n_embd, 4 * c.n_embd, bias=False), nn.GELU(),
nn.Linear(4 * c.n_embd, c.n_embd, bias=False),
nn.Dropout(c.dropout))
def forward(self, x, cache=None):
x = x + self.attn(self.ln1(x), cache)
return x + self.mlp(self.ln2(x))
- Pre-norm. LayerNorm comes before attention and before the MLP, and each adds its result back to the input (the residual stream). This is the GPT-2 layout.
- Attention is PyTorch’s
scaled_dot_product_attentionwith a causal mask, so each position sees only the characters before it. - No biases in the linear layers.
- Tied output. The output layer reuses the token embedding matrix, so the score for each next character is a dot product with that character’s embedding.
Generation
To play a move, the model reads the prompt and writes characters until it writes a newline. complete() does this greedily, always taking the most likely character, with a key/value cache so each new character costs one position, not a rerun of the whole prompt:
caches = [{} for _ in self.blocks]
logits = self(torch.tensor([p], device=device), caches=caches)[0, -1]
pos = len(p)
for _ in range(max_new):
row = logits.float()
first = int(row.argmax())
banned = ban(i, out[i]) if ban else ()
if banned:
row = row.clone()
row[list(banned)] = float("-inf")
best = int(row.argmax())
...
self.logp[i].append(float(torch.log_softmax(row, dim=0)[best]))
if best == stop:
break
out[i].append(best)
logits = self(torch.tensor([[best]], device=device), caches=caches, start=pos)[0, -1]
pos += 1
Two hooks in it matter for play:
banlets the caller rule out characters at each step. transformer-solve uses it to keep moves legal, as described in transformer-solve: the model as a player.logprecords the log probability of each character written. Summed over a SOLVE, it gives the model’s confidence in its answer, and that is what decides whether transformer-solve solves.
Phase 9: Training
Training has two stages:
- Pretraining teaches the model English. It reads the corpus and learns to predict the next character, with loss on every character.
- Finetuning teaches it the game. It starts from the pretrained weights and trains on a mix of the game streams, with some corpus kept in so the English isn’t lost.
| Pretrain | Finetune | |
|---|---|---|
| Starts from | random weights | runs/pre/ckpt.pt |
| Data | corpus only | main 35%, bonus 30%, list_bonus 20%, list 5%, corpus 10% |
| Steps | 30,500 | 15,000 |
| Batch | 128 rows of 256 tokens (32,768 tokens) | the same |
| Tokens seen | 1.0B | 0.49B |
| Learning rate | 6e-4, 500 warmup steps | 3e-4, 200 warmup steps |
| Writes | runs/pre/ | runs/ft2b/ |
ft2b is the finetuned model behind every result on this page.
The training script
09_train.py builds each batch row from one stream, picked at random by the mix weights. A corpus row is 257 consecutive characters from a random point in corpus.bin. A game row packs whole lines, picked at random, one after another until the row is full. Its loss mask is on only for target characters:
def row(self, rng, n):
"""n + 1 tokens of whole random lines, and the loss mask for each
predicted position (True where the next token is target)."""
x, m = [], []
while len(x) < n + 1:
start, plen, length = self.idx[rng.integers(len(self.idx))]
x.extend(self.tok[start:start + length])
m.extend([False] * plen + [True] * (length - plen))
return np.array(x[:n + 1], dtype=np.int64), np.array(m[1:n + 1])
Positions where the mask is off get target -1, which cross_entropy(..., ignore_index=-1) skips. The model still reads the prompt. It just isn’t scored on predicting it.
The rest is a standard nanoGPT-style loop:
- AdamW with betas 0.9 and 0.95, and weight decay 0.1 on the matrices only.
- The learning rate warms up linearly, then follows a cosine decay to a tenth of its peak.
- bfloat16 autocast on the GPU, gradient clipping at 1.0, and
torch.compile. - Checkpoints. Every
--eval_everysteps it measures validation loss on each validation stream and savesruns/<name>/ckpt.pt, which holds the weights, the config, the alphabet and the arguments.log.tsvrecords the losses and throughput.
09_pod/run.sh has the exact commands used:
python3 scripts/09_train.py --name pre --phase pretrain --steps 30500 \
--eval_every 2000 --compile 2>&1 | tee pre.log
python3 scripts/09_train.py --name ft2b --phase finetune --init runs/pre/ckpt.pt \
--steps 15000 --eval_every 1000 --lr 3e-4 --warmup 200 --compile 2>&1 | tee ft2b.log
Running it on a rented GPU
I trained on one NVIDIA H100 SXM 80 GB, rented from Runpod in its Secure Cloud. Any machine with a CUDA GPU and PyTorch will do. Smaller GPUs will take longer.
Steps:
- Bundle what the GPU needs: the model, the training script,
run.shand the token files from phase 8. This writes/tmp/wof-train.tar, about 1.6 GB:bash scripts/09_pod/pack.sh - Copy it to the GPU machine and unpack it. On Runpod, with SSH set up for the pod:
rsync -P /tmp/wof-train.tar root@<pod-address>:/workspace/ ssh root@<pod-address> 'cd /workspace && tar -xf wof-train.tar' - Run both stages, about 32 minutes of training on an H100:
ssh root@<pod-address> 'cd /workspace && bash scripts/09_pod/run.sh'run.sh finetuneruns only the second stage, from an existingruns/pre/ckpt.pt. - Copy the checkpoints back, then stop the pod:
rsync -P root@<pod-address>:/workspace/runs/pre/ckpt.pt runs/pre/ rsync -P root@<pod-address>:/workspace/runs/ft2b/ckpt.pt runs/ft2b/ - Before renting anything, check the script on a CPU with a toy model, in a few minutes:
.venv/bin/python scripts/09_train.py --name smoke --phase pretrain --n_layer 2 --n_head 2 \ --n_embd 64 --batch 16 --steps 200 --eval_every 100
About exact reproduction. Training on a GPU is not bit-for-bit repeatable. torch.compile and the GPU’s parallel sums can round slightly differently from run to run, even with the same seed. A rerun will give a model with very close losses, but not identical weights. To reproduce the later phases exactly, use the published checkpoints from the release.
What the logs show
Pretraining, from runs/pre/log.tsv. Loss is cross-entropy in nats per character on the held-out corpus_val. The full log is in Appendix A.
A loss of 4.02 at step 0 is about log(48) = 3.87: random guessing over 48 characters. It ends at 0.928 nats, which is 1.34 bits per character. The H100 ran at about 820,000 tokens a second, so pretraining took 20 minutes.
Finetuning, from runs/ft2b/log.tsv, where main_val and bonus_val are measured on the control puzzles’ lines, on the target characters only.
- The game is learned fast. Most of the drop in
main_valhappens in the first 1,000 steps. - English dips, then recovers.
corpus_valrises from 0.928 to 1.03 as the game data comes in, then comes back down to 0.971 as the learning rate decays. - The bonus round levels off.
bonus_valbottoms out around step 10,000 to 12,000, at 0.083, and ends at 0.085.
Finetuning took 12 minutes at about 686,000 tokens a second.
Output:
| File | Size |
|---|---|
runs/pre/ckpt.pt | 237 MB |
runs/ft2b/ckpt.pt | 237 MB |
Both are in the release as pre-ckpt.pt and ft2b-ckpt.pt.
transformer-solve: the model as a player
The model writes text. solvers.py turns it into a player, TransformerSolver. For each move it builds the prompt, exactly as in training, and lets the model complete it:
MAIN|AROUND THE HOUSE|A _OLL O_ __L_|ALOS| -> CALL F
A free-running model can write anything. Masks hold it to the rules and the board, through the ban hook in complete():
def ban(_, new):
if not new:
return no_buy
if new in verbs:
return repeat
if len(new) >= len(solve) and new[:len(solve)] == solve:
k = len(new) - len(solve)
return list(everything - allowed[k]) if k < len(allowed) else []
return ()
- First character: B is banned until a vowel can be bought, the money rule.
- After
CALLorBUY: letters already called are banned. - After
SOLVE: each character must fit its square. Shown letters, spaces and punctuation are copied. A blank takes only an uncalled letter. The answer must end where the board ends.
The SOLVE mask means every word of an answer fills its slot exactly. The model can still pick the wrong word, but it can’t write one of the wrong length. In the evaluation, the masks changed 9,800 characters of letter moves and 3,420 characters of SOLVEs, across 44,694 MAIN moves.
Phase 10: SOLVE cutoff
The model’s first instinct is to solve too early. On the dev puzzles, a transformer-solve that accepts every SOLVE it writes solves only 64% of them. It has a confident-looking guess well before the board supports one.
So a MAIN SOLVE is accepted only when the model’s probability for that answer reaches a set cutoff. The probability is the product of the probabilities of each character after SOLVE , from logp. Below the cutoff, transformer-solve turns the SOLVE down and makes its best letter move instead.
10_tune_solve.py picks the cutoff on the 300 dev puzzles, which the model never trained on. It plays every game with every early SOLVE turned down, so each game runs until the board is full, and it logs each turned-down SOLVE with its probability and whether it was right. One run scores every cutoff:
def score(games, cut):
solved, letters = 0, []
for _, _, _, tried, final, _ in games:
hit = next(((ok, k) for p, ok, k, _ in tried if p >= cut), final)
if hit[0]:
solved += 1
letters.append(hit[1])
return solved / len(games), letters
This works because a turned-down SOLVE always leads to the same letter move. So a game with cutoff p plays exactly like the logged game up to the first SOLVE with probability at least p. That SOLVE decides the game.
Steps:
- Run it on the finetuned checkpoint. It uses 11 processes, one torch thread each:
.venv/bin/python scripts/10_tune_solve.py runs/ft2b/ckpt.pt
Result (data/eval/tune.tsv, 2,819 logged SOLVEs from 300 dev games):
| Cutoff | Solved |
|---|---|
| 0 (accept every SOLVE) | 64.0% |
| 0.50 | 80.3% |
| 0.90 | 94.0% |
| 0.95 | 95.3% |
| 0.99 | 98.3% |
Three dev games from tune.tsv show how the confidence moves. Each row is a SOLVE the model wrote and the tuning run turned down:
category solution letters p right answer
Phrase OUT WITH THE OLD AND IN WITH THE NEW 11 0.9841 1 OUT WITH THE OLD AND IN WITH THE NEW
Phrase OUT WITH THE OLD AND IN WITH THE NEW 12 0.9774 1 OUT WITH THE OLD AND IN WITH THE NEW
Phrase OUT WITH THE OLD AND IN WITH THE NEW 13 0.9925 1 OUT WITH THE OLD AND IN WITH THE NEW
Phrase THROWN FOR A LOOP 13 0.9035 1 THROWN FOR A LOOP
Phrase THROWN FOR A LOOP 14 0.7507 1 THROWN FOR A LOOP
...
Phrase THROWN FOR A LOOP 21 0.6956 1 THROWN FOR A LOOP
People TRIPLETS 4 0.9840 0 TRUMPETS
People TRIPLETS 5 0.9966 0 TRUMPETS
- OUT WITH THE OLD AND IN WITH THE NEW is right from 11 letters on. With a 0.99 cutoff, transformer-solve solves it at 13.
- THROWN FOR A LOOP is right from 13 letters on, but the model’s confidence falls as more letters show, and it never reaches 0.99. With the cutoff, it calls letters until the board is full. It still solves, but late.
- TRIPLETS goes wrong. After four letters the model writes TRUMPETS at 0.984, and after five at 0.997. That clears the cutoff, so this dev game is lost. A high probability is a strong signal, not a guarantee.
Of the 2,819 logged SOLVEs, 372 were wrong, and 215 of those had a probability under 0.5. Most wrong guesses come with low confidence, so one threshold filters out most of them.
0.99 is the cutoff used everywhere after this, in the evaluation and on the server. At 0.99, transformer-solve solved 98.3% of dev puzzles with 12.08 letters on average.
Phase 11: Evaluation
11_evaluate.py plays a solver on all 3,349 control puzzles, in both modes. Every move is checked against the rules:
def play_main(n, solution, category):
S.new_puzzle(wc.seed_for(SEED, str(n)))
called = []
while True:
board = wc.board(solution, set(called))
mv = S.main(category, board, "".join(sorted(called)))
if mv.startswith("SOLVE "):
return "solved" if mv[6:] == solution else "wrong", called
letter = check_move(mv, board, called)
if letter is None:
return "invalid", called
called.append(letter)
Steps:
- smart-solve:
.venv/bin/python scripts/11_evaluate.py smart-solve - transformer-solve. The model runs on the CPU, one game per worker process:
.venv/bin/python scripts/11_evaluate.py transformer-solve:runs/ft2b/ckpt.pt - A quick check on the first 200 puzzles writes to
/tmpinstead:.venv/bin/python scripts/11_evaluate.py transformer-solve:runs/ft2b/ckpt.pt main 200
Output: one row per puzzle and mode in data/eval/smart-solve.tsv and data/eval/transformer-solve.tsv, and one summary row per solver and mode in data/eval/summary.tsv:
mode category solution result detail
main Around The House A ROLL OF FILM solved letters=7 buys=2 misses=1 order=SLAOFMR
main Around The House A ROLL OF FILM solved letters=10 buys=4 misses=4 order=SLAOFREIKT masked=2 fitted=0 rejected=2
The first row is smart-solve and the second is transformer-solve, on the same puzzle. Both open S, L, A, O, F. smart-solve calls M and R and solves. transformer-solve turns down two SOLVEs under the cutoff and calls R, E, I, K and T before it is sure.
Phase 12: Exporting the weights
12_export_weights.py writes the checkpoint as one flat binary file that C can read without a library:
magic "WOFM" and a uint32 version (2)
uint32 vocab, block, n_layer, n_head, n_embd
char the alphabet, vocab bytes; a token's id is its index
float16 tensors (IEEE half), in order:
tok [vocab][n_embd] (also the output layer, tied)
pos [block][n_embd]
per layer: ln1.w, ln1.b, qkv, proj, ln2.w, ln2.b, fc, out
ln.w, ln.b [n_embd]
The weights are stored as float16, half the size of float32: 118 MB instead of 237 MB. The server widens each weight row to float32 as it uses it, with the F16C instruction _mm256_cvtph_ps, and does all its arithmetic in float32:
// y[t][o] = x[t] . w[o] for t < T: each weight row is widened once for all T rows.
static void linear(float *y, const float *x, const half *w, int T, int in, int out) {
static float row[MAX_IN];
for (int o = 0; o < out; o++) {
widen(row, w + (size_t)o * in, in);
for (int t = 0; t < T; t++) y[(size_t)t * out + o] = dot(x + (size_t)t * in, row, in);
}
}
Steps:
.venv/bin/python scripts/12_export_weights.py runs/ft2b/ckpt.pt runs/ft2b/wof.bin
Output: runs/ft2b/wof.bin, 118,417,996 bytes, also in the release.
Phase 13: Demo data
13_build_demo.py writes data/demo.json, everything behind the demo’s puzzle endpoints:
categories: all 49 training categories, most common first, each marked with whether control has a puzzle in it.puzzles: the 3,349 control puzzles, which Random mode draws from.presets: the seven Demo examples.
The presets are hand-picked control puzzles where transformer-solve beats smart-solve. They are the transformer’s wins, not typical games. Over the control set, smart-solve wins most head-to-head matches. Each preset has a note, such as: “smart-solve’s word list has no HEADSHOT. It solves CELEBRITY HELMSMAN.”
.venv/bin/python scripts/13_build_demo.py
The C server
The demo needs both solvers answering moves over HTTP, fast, on a small cheap machine. In Python, transformer-solve needs PyTorch, and the PyTorch install alone (727 MB) is three times the size of the model. So the server is one C program, wof-server, that runs both solvers and serves the game client:
| File | Lines | What it does |
|---|---|---|
model.c | 207 | the transformer’s forward pass, with a key/value cache |
transformer.c | 200 | transformer-solve: decoding under the same masks, and the 0.99 cutoff |
smart.c | 776 | smart-solve, ported from wof_common.py and solvers.py |
main.c | 727 | HTTP, the API, the static files |
HTTP parsing is picohttpparser (MIT) and JSON is cJSON (MIT), both vendored.
Parity: the C server plays like the Python
A port is only useful if it plays the same moves. Three tests in server/tests/ compare the C code with the Python it replaces:
| Test | What it compares | Result |
|---|---|---|
model_parity.py | log probabilities and greedy completions on 200 game prompts | largest log-probability difference 0.0096, all 200 completions the same |
tf_parity.py | transformer-solve’s moves in 100 MAIN and 100 BONUS games | 1,329 of 1,331 MAIN moves the same, and all 200 BONUS picks and solves |
smart_parity.py | smart-solve’s move at every state of all 3,349 control puzzles | all 40,223 moves the same |
smart-solve matches move for move because the port copies two details of the Python:
- The summing order. It adds scores in numpy’s pairwise order, not left to right.
- The rounding. It rounds the way Python rounds before comparing scores.
Without those, near-ties break differently.
The two transformer moves that differ are float16 effects. Both are close calls between two letters, which float16 rounding tips the other way: CALL S instead of CALL R on one empty board, and BUY O instead of BUY A on another.
Steps:
make -C server
python3 server/tests/model_parity.py runs/ft2b/ckpt.pt runs/ft2b/wof.bin
python3 server/tests/tf_parity.py runs/ft2b/ckpt.pt runs/ft2b/wof.bin
python3 server/tests/smart_parity.py 3349
The tests need torch, so run them with .venv/bin/python if that’s where you installed it.
Running it on Fly.io
The server runs on Fly.io on the smallest machine there is: one shared CPU and 256 MB of RAM.
The weights alone are 118 MB, so fitting in 256 MB takes one more step. The first float16 build read the weights into the process’s own memory, 160 MB in all. The kernel killed it at startup for running out of memory. Now the server maps the weights file into memory with mmap instead of copying it. The weights live in the kernel’s page cache, which the kernel can drop and reread from disk, and the server’s own memory is about 30 MB.
Idle machines suspend after a few minutes and wake on the next request. Measured on the deployed server:
| A transformer-solve move | about 240 ms |
| A smart-solve move | about 5 ms |
| Waking a suspended machine and making a move | about 1.1 s |
| The first move after a fresh start (weights read from disk) | about 5.6 s |
Steps:
- Build the game client first, since the server image includes it. See The game client.
- Run the server locally:
make -C server WEIGHTS=runs/ft2b/wof.bin WEB_ROOT=client/build/web PORT=8080 ./server/build/wof-server - Ask it for a move:
curl -s -X POST localhost:8080/move \ -d '{"mode":"main","category":"Phrase","board":"____ __ ___ ____","called":""}'{"smart":{"move":"CALL T","ms":7},"transformer":{"move":"CALL T","ms":194}} - Deploy to Fly, from the root of the clone, with your own app name in
server/fly.toml:fly deploy . --config server/fly.toml --dockerfile server/Dockerfile --remote-only --ha=false
The full API is documented at the top of server/src/main.c.
The game client
The demo at wof-solver.fly.dev is a small game in C, written with raylib and compiled to WebAssembly with Emscripten. It is about 1,500 lines in client/src/. It draws two boards side by side, one per solver, and plays them move by move, fetching each move from the server.
Each side shows its board, the letter strip, letters and misses, a stopwatch for the server’s move time, and the move history. For transformer-solve, the history includes the answer it was considering and its confidence, and a SOLVE under 99% shows as held back. On a tall phone screen the two boards stack.
Steps:
- Build raylib 6.0 for the web, once. This needs the Emscripten SDK on your path:
cd client bash scripts/build_raylib_web.sh - Build the client and serve it locally:The page talks to
make web make web-serve # http://localhost:8000/wof-solver.htmlhttps://wof-solver.fly.dev. Add?api=http://localhost:8080to use a local server started withCORS_ORIGINS=http://localhost:8000. - Run the native tests of the game logic:
make test
Results & Conclusions
Overall results
| Solver | MAIN solved | Wrong | Letters (median / mean) | Misses (mean) | BONUS solved |
|---|---|---|---|---|---|
| smart-solve | 98.3% | 57 | 9 / 9.04 | 1.96 | 79.8% |
| transformer-solve | 98.7% | 42 | 12 / 12.33 | 3.77 | 77.2% |
Neither solver made an invalid move.
- The solve rates are close, but smart-solve needs fewer letters. Scored head to head on each MAIN puzzle, fewer misses winning and an unsolved puzzle losing, smart-solve wins 2,245 (67%), transformer-solve wins 439 (13%), and 665 are ties.
- Missing words cause smart-solve’s wrong answers. All 57 of them are in the 159 puzzles with a word missing from its list: a list word fits the board, so it solves early with the wrong word (FILM JOUR FESTIVAL, THE COUPONS for THE GORGONS). Only 10 of transformer-solve’s 42 wrong solves are in those puzzles.
- In BONUS, they fail on different puzzles. On the 1,348 puzzles where both pick the same letters, smart-solve solves 1,135 and transformer-solve 1,096. smart-solve solves 414 puzzles that transformer-solve misses, and transformer-solve solves 326 that smart-solve misses. Accepting a solve from either would reach 89.5%.
- transformer-solve has habits. It opens the same way in most games: CALL S, then BUY E and BUY A once a consonant shows. GOOEY COOKIE DOUGH costs it five misses (S R L N T) before its first hit.
By category
Appendix B has the results for the 16 largest categories in the control set.
- Names are where the word list fails. In MAIN, smart-solve’s worst categories are Proper Name (91.4%) and Fictional Character (93.5%), the ones full of names its list doesn’t have. transformer-solve solves 98.4% and 98.7% of them.
- transformer-solve uses more letters everywhere. It calls 2.4 to 4.7 more letters than smart-solve in every one of these categories.
- In BONUS, transformer-solve leads on the wordplay categories: Before And After (77.1% against 70.1%), Same Name (76.8% against 68.1%) and Show Biz (80.6% against 66.1%). All three are well covered by the category lists from phase 4. smart-solve leads in most of the others, with its widest margin in Proper Name (71.1% against 60.2%).
How each solver gets it wrong
The evaluation files record each game’s letters, not the wrong answer itself. To see the answers, I replayed every wrong MAIN game through the C server. Every replayed game called the same letters in the same order as the evaluation.
smart-solve is wrong when a word is missing from its list and another word of the same shape fits what is showing:
| Puzzle | smart-solve’s answer | Board when it solved |
|---|---|---|
| FILM NOIR FESTIVAL | FILM JOUR FESTIVAL | F_L_ _O_R FEST__AL |
| THE GORGONS | THE COUPONS | THE _O__ON_ |
| CELEBRITY HEADSHOT | CELEBRITY HELMSMAN | _E_E_RI__ _E__S___ |
| BEEF BOURGUIGNON | BEEF APPROVINGLY | _EEF ___R__I____ |
| MEPHISTO | LEBOWSKI | _E___S__ |
| BRIGITTE BARDOT | BRIGITTE HARLOT | _RI_ITTE _AR__T |
It solves once every word has one candidate. When the true word isn’t a candidate, the one that is left looks just as certain. BEEF APPROVINGLY shows that it has no sense of which words go together. Some of its wrong answers are the source’s typos, not its own: the control puzzles include HOMEMADE THREE-CHEESE LASANGA and NATURALLY AGED CHEDDER CHEESE.
transformer-solve goes wrong differently. In 41 of its 42 wrong answers, it is off by one or two letters, often writing a word that doesn’t exist:
| Puzzle | transformer-solve’s answer |
|---|---|
| BACHELOR AUCTION | BACHELOR APCTION |
| ZOOKEEPER | BOOKEEPER |
| WHEAT GERM | WHEAT PERM |
| HIGHLY PROFITABLE | HIGHLY PROMITABLE |
| INTERNATIONAL SYSTEM OF UNITS | INTERNATIONAL SYSTEM OF KNITS |
| PINCHING YOUR CHEEKS | PUNCHING YOUR CHEEKS |
| CYNDI LAUPER | CONDI LAUPER |
| WATCHING GREMLINS | WATCHING FREMZINS |
The model writes the answer one character at a time. The masks keep each character legal for its square, but nothing checks that the word is real. With a few squares left blank, its most likely fill can be a non-word, at over 99% confidence. One of its “wrong” answers is a correction: for the control puzzle HOMEMADE THREE-CHEESE LASANGA it writes LASAGNA.
Learnings and conclusions
The data was most of the work
There are thirteen phases, and only one of them trains a model. The first eight are all data work. They scrape and clean the puzzles, split them, build the category lists and the corpus, and write 5.8 million labelled game lines. Training took 32 minutes on one GPU, and the model code is only 130 lines. The results depend much more on the data than on the model.
A small model can learn a game from a teacher
The 59M-parameter transformer started knowing nothing, not even English. After a billion characters of pretraining and half a billion of game lines, it solves 98.7% of puzzles it has never seen, against smart-solve’s 98.3%. It never made an illegal move in 3,349 games, though the masks did correct some characters along the way.
It learned the teacher’s moves, not a better strategy
Every MAIN training label is smart-solve’s pick, so the model learned to imitate it. A model trained to copy a strategy can approach that strategy. Getting past it needs a different signal, such as training on the outcome of whole games.
Knowing when to stop was the biggest lever
Accepting every SOLVE the model wrote solved 64% of dev puzzles. Requiring 99% confidence solved 98.3%, with no retraining at all. The model’s own probability for its answer turned out to be a usable measure of whether it was right. Tuning that one number on held-out puzzles mattered more than any training setting.
Rules belong in the decoder
A free-running model can call a letter twice or write an answer that doesn’t fit the board. Masking those characters out while it generates costs almost nothing and removes a whole class of failure. The model only has to learn which legal move is best.
The two solvers fail differently
smart-solve fails where its list has gaps. transformer-solve fails by misspelling. A word list and a language model cover each other’s gaps.
Held-out data has to be guarded everywhere
The control split is by answer text, so a puzzle filed under two categories can’t sit on both sides. The category lists drop every control answer before the model sees them: 680 phrases across Wikidata, Wiktionary and WordNet. Without those checks, the model would have trained on part of its own test.
Small and cheap is enough
A 207-line C forward pass with float16 weights serves the model on a 256 MB machine. A general-purpose model server was the wrong fit. An earlier version ran the model in llama.cpp’s llama-server on a 1 GB machine, and its default prompt cache pushed the weights out of memory.
Reproducible by construction
Every split is a hash of the answer text, and every random choice is seeded from the puzzle. Rerunning phases 2 through 8 from the published inputs gives byte-identical files. The one exception is GPU training, which is why the checkpoints are published.
Appendices
Appendix A: Training logs
Pretraining, from runs/pre/log.tsv:
| Step | Tokens | corpus_val loss | Elapsed |
|---|---|---|---|
| 0 | 0 | 4.021 | 2 s |
| 2,000 | 66M | 1.192 | 82 s |
| 10,000 | 328M | 1.023 | 398 s |
| 20,000 | 655M | 0.962 | 799 s |
| 30,500 | 1.0B | 0.928 | 1,220 s |
Finetuning, from runs/ft2b/log.tsv:
| Step | main_val | bonus_val | corpus_val | Elapsed |
|---|---|---|---|---|
| 0 | 2.737 | 1.642 | 0.928 | 2 s |
| 1,000 | 0.207 | 0.159 | 1.018 | 51 s |
| 5,000 | 0.168 | 0.096 | 1.025 | 243 s |
| 10,000 | 0.151 | 0.083 | 0.993 | 482 s |
| 15,000 | 0.140 | 0.085 | 0.971 | 722 s |
Appendix B: Results by category
The 16 largest categories in the control set. Letters is the mean number of letters called in solved MAIN games.
| Category | Puzzles | MAIN smart-solve | Letters | MAIN transformer-solve | Letters | BONUS smart-solve | BONUS transformer-solve |
|---|---|---|---|---|---|---|---|
| Thing | 662 | 98.9% | 8.18 | 98.9% | 11.94 | 84.3% | 78.9% |
| Food And Drink | 374 | 97.9% | 8.60 | 97.6% | 11.99 | 89.8% | 84.0% |
| Phrase | 306 | 100.0% | 10.00 | 99.0% | 12.42 | 78.1% | 78.1% |
| What Are You Doing | 272 | 100.0% | 10.01 | 97.4% | 12.77 | 73.9% | 74.6% |
| Place | 186 | 99.5% | 7.85 | 99.5% | 11.34 | 88.2% | 82.8% |
| Event | 170 | 98.8% | 8.98 | 99.4% | 12.08 | 85.3% | 82.9% |
| Before And After | 144 | 100.0% | 10.25 | 100.0% | 13.24 | 70.1% | 77.1% |
| Proper Name | 128 | 91.4% | 9.18 | 98.4% | 13.60 | 71.1% | 60.2% |
| People | 124 | 99.2% | 7.68 | 97.6% | 11.40 | 87.1% | 78.2% |
| Living Thing | 101 | 98.0% | 8.43 | 98.0% | 12.97 | 80.2% | 74.3% |
| Person | 88 | 95.5% | 6.95 | 100.0% | 11.61 | 86.4% | 84.1% |
| Fictional Character | 77 | 93.5% | 9.76 | 98.7% | 13.61 | 72.7% | 67.5% |
| On The Map | 71 | 97.2% | 8.39 | 98.6% | 12.99 | 85.9% | 74.6% |
| Same Name | 69 | 100.0% | 10.20 | 100.0% | 12.71 | 68.1% | 76.8% |
| Around The House | 68 | 100.0% | 9.13 | 100.0% | 11.99 | 82.4% | 73.5% |
| Show Biz | 62 | 98.4% | 10.15 | 98.4% | 12.74 | 66.1% | 80.6% |
Appendix C: What it cost
| Item | Cost |
|---|---|
| GPU: two Runpod sessions on an H100 SXM, about 88 minutes in all, at $3.49 an hour | about $5.15 |
| Hosting: Fly.io shared-cpu-1x, 256 MB | $2.19 a month if always on, less when suspended |
Appendix D: Licenses of the data
| Source | License | Used for |
|---|---|---|
| wofanswers.com | no license stated | puzzles.csv |
SCOWL / Debian wamerican | see data/wordlist/COPYRIGHT | smart-solve’s word list |
| Wikidata | CC0 | wikidata.tsv |
| Wiktionary, via kaikki.org | CC BY-SA | wiktionary.tsv |
| Open English WordNet 2025 | CC BY 4.0 | wordnet.tsv |
| CMU Pronouncing Dictionary | BSD-style | the rhymes in wordplay.tsv |
| Wikipedia | CC BY-SA 4.0 | pretraining text |
| FineWeb-Edu | ODC-By 1.0 | pretraining text |