The WOF Solver game: smart-solve and transformer-solve play the same Food And Drink puzzle side by side. smart-solve has solved LONG ISLAND ICED TEA after 10 letters, and transformer-solve is still calling letters.

TLDR

I trained a 59M-parameter transformer from scratch to play Wheel of Fortune, and named it transformer-solve. It plays head to head against smart-solve, a conventional solver built on word lists. transformer-solve first learned English from a billion characters of text, then trained on 5.8 million game lines labelled with smart-solve’s moves. Training took 32 minutes on one rented H100.

On 3,349 puzzles neither had seen, transformer-solve solved 98.7% of MAIN games and smart-solve 98.3%. smart-solve needed fewer letters, 9.04 on average against 12.33, and won 67% of the head-to-head games. In BONUS, smart-solve solved 79.8% and transformer-solve 77.2%.

transformer-solve learned to imitate its teacher, not to beat it. Most of the work was data, not model code, and the biggest gain came from tuning when it may solve. smart-solve fails on words missing from its lists, transformer-solve on spelling.

Introduction

I wanted to know whether a small language model could learn to play Wheel of Fortune. Not just read the board, but play: choose letters, decide when it knows enough to solve, and do it without a strategy written into its code. The test was to give it a teacher and see how closely it could imitate it.

So I built two solvers.

smart-solve is deliberately conventional. It has a word list, and it calls the letter expected to narrow the remaining candidates the most. It plays well with no machine learning at all.

transformer-solve is a small character-level language model trained from scratch. It first learns the patterns of written English, then trains on games labelled with smart-solve’s moves. A language model can produce plausible text. The question is whether that is enough to learn a game from examples.

Both played the same few thousand puzzles, which neither had seen before.

The result was closer than I expected. The two solved at nearly the same rate, but the transformer needed more letters. It had learned to play mostly by imitating its teacher. The interesting part is why, and whether I can close the gap.

This is the story of the two solvers, the data behind them, the mistakes they made, and what the experiment taught me about training a small model to do something useful.

Head-to-head solvers

The demo puts the two solvers head to head. Both get the same puzzle under the same rules, and each plays its own board beside the other’s, one move at a time.

A scoreboard compares the two games:

  • A solver that solves beats one that doesn’t.
  • If both solve, fewer misses wins, then fewer letters.
  • If those are equal, it’s a tie. If neither solves, both lose.

Three modes pick the puzzles:

  • Random plays held-out puzzles one after another, until you stop it.
  • Demo plays seven preset examples, then stops.
  • Your puzzle plays one puzzle you type in, in a category you pick.

MAIN and BONUS switch between the main game and the bonus round.

The game as played here

Wheel of Fortune puzzles are short phrases with a category, laid out on a board of four rows. The rows are 12, 14, 14 and 12 squares wide, 52 squares in all. That layout has been on the show since 1997. Letters start hidden. Spaces and punctuation always show.

The puzzle board: four rows of 12, 14, 14 and 12 squares, the short rows centred.

The code writes a board as one line, with an underscore for each hidden letter. A _A__L_ __ ___ ______ is A CANDLE IN THE WINDOW after A and L are called. Category names come from the show: Thing, Phrase, Food And Drink, Before And After, Rhyme Time and 44 others.

The main game (MAIN)

From any board, a solver makes one of three moves:

  • CALL a consonant: CALL T. Every T on the board turns over.
  • BUY a vowel: BUY E.
  • SOLVE: SOLVE A CANDLE IN THE WINDOW.

The rules:

  • Invalid moves end the game. Calling a letter twice, buying a consonant, or calling a vowel are all invalid.
  • The money rule. On the show a vowel costs $250, and money only comes from consonants that hit. The games here keep no money, so the rule is read off the board: a vowel can be bought once a called consonant shows on it.
  • One SOLVE. The first SOLVE ends the game, right or wrong.

A MAIN game ends one of three ways: solved, wrong, or invalid. A solved game is also measured by how many letters were called before the SOLVE, and how many of those were misses, letters not in the puzzle. On the show a miss loses the turn. Fewer is better for both. A solver that waits until the board is full always solves, so the letter count is the real measure.

The bonus round (BONUS)

The bonus round is one shot:

  1. R S T L N E are turned over.
  2. The solver picks 3 more consonants and 1 vowel: PICK DHM A.
  3. Those letters turn over, and the solver gets one SOLVE.

BONUS is scored on the solve rate alone. It is harder than MAIN: the solver sees ten letters at most, picks four of them blind, and gets one guess.

What is left out

There is no wheel, no money, no turns and no opponent on the board. Each solver plays its own copy of the puzzle from the first letter to the solve.

Fair comparison

The two solvers are held to the same terms:

  • Same puzzles. Both play the same 3,349 held-out puzzles, with the same move rules.
  • No peeking. Neither solver sees a held-out answer. Both know only the training puzzles, the word list, and the category lists. transformer-solve also knows the pretraining text.
  • The caller holds the answer. The evaluation and the server send a solver only the category, the board and the called letters, never the answer.

Building & Training the Model

This section walks through the build one phase at a time, from the scraped puzzles to the C server behind the demo. Every phase has its script, its inputs and outputs, a sample of the data, and the commands to run it. The setup section covers what you need. Every generated file has a published SHA-256 checksum, so you can check that your rebuild matches mine byte for byte.

The project is in its own repository, dannyheskett/wof-solver. All links on this page point at the tag v1, so the code you read here is the code that produced the results.

DirectoryWhat is in it
scripts/One Python script per phase, named by phase (01_fetch_puzzles.py to 13_build_demo.py), plus three shared modules
data/The puzzles, the splits, the category lists, the results. DATA.md documents every file
server/wof-server, the C program that runs both solvers and serves the demo
client/The demo game, in C with raylib, compiled to WebAssembly
NOTES.mdThe project notes: rules, results, hosting, the cost log

Setup

Everything except training (phase 9) runs on an ordinary Linux machine.

What you need:

Python3.12, with the packages in requirements.txt: numpy 1.26.4, pyarrow 25.0.1, torch 2.14.0
Diskabout 10 GB: 6.3 GB of downloaded sources, 2.9 GB of generated training files, 0.6 GB of checkpoints
A C compilerfor the C server. Any C99 compiler with x86-64-v3 support (AVX2, F16C)
A GPUonly to retrain (phase 9). The trained checkpoints are published, so you can skip training

Steps:

  • Clone the repository and enter the project:
    git clone https://github.com/dannyheskett/wof-solver.git
    cd wof-solver
    git checkout v1
    
  • Make a virtual environment and install the pinned packages. The extra index gives the CPU build of torch, which is all phases 10 to 13 need:
    uv venv -p 3.12 .venv
    uv pip install -p .venv/bin/python -r requirements.txt \
        --extra-index-url https://download.pytorch.org/whl/cpu --index-strategy unsafe-best-match
    
  • Every command on this page runs from the root of the clone with .venv/bin/python.

Published artifacts. Two things in the pipeline can’t be repeated exactly: the scrape in phase 1 (the site changes) and training on a GPU (not bit-for-bit repeatable). Their outputs are published, so every other phase rebuilds from the same inputs I used:

  • data/puzzles.csv, the scrape, is in the repository.

  • The checkpoints and the raw Wikidata results are attached to the v1 release:

    AssetSizeWhat it is
    pre-ckpt.pt237 MBthe pretrained transformer, the start of finetuning
    ft2b-ckpt.pt237 MBthe finetuned transformer behind every result
    wof.bin118 MBthe same weights as float16, for the C server
    wikidata-raw.tar.gz2.7 MBthe Wikidata query results phase 4 reads
    mkdir -p runs/pre runs/ft2b
    base=https://github.com/dannyheskett/wof-solver/releases/download/v1
    curl -L -o runs/pre/ckpt.pt  $base/pre-ckpt.pt
    curl -L -o runs/ft2b/ckpt.pt $base/ft2b-ckpt.pt
    

Checking a rebuild. data/SHA256SUMS lists a checksum for every file the pipeline writes, including the large ones that are not in git. After any phase:

cd data && sha256sum -c --ignore-missing SHA256SUMS

Every line should say OK. --ignore-missing skips the files you haven’t built yet.

Phase 1: Fetching the puzzles

Script: 01_fetch_puzzles.py. Reads the wofanswers.com API, writes data/puzzles.csv.

wofanswers.com is a fan site that lists every puzzle from the show and from the mobile game, with a solver page for each category. The site’s pages call a JSON API at /api/collections/wof-answers/items, which takes filters as query parameters and returns rows like:

{"meta": {"total": 13159},
 "data": [{"answer": "A CANDLE IN THE WINDOW", "category": "Around The House"}, ...]}

The API has two problems. It returns at most 1,000 rows per call, and its offset paging overlaps: asking for rows 1,000 to 2,000 repeats some rows from the first page and skips others. Paging through it gives a list that looks complete and isn’t.

So the script never pages. It splits the data with filters until every slice fits in one call. It starts with the category, then splits big categories by word count, then by letter count. At each split it checks that the children’s totals add up to the parent’s, so a slice can’t go missing silently:

def collect(filters, total, depth, rows, seen):
    """Collect every row matching filters, splitting until slices fit in one call."""
    resp = query(filters)
    if resp["meta"]["total"] != total:
        raise SystemExit(f"{filters}: total {resp['meta']['total']} != expected {total}")
    if total <= LIMIT:
        if len(resp["data"]) != total:
            raise SystemExit(f"{filters}: got {len(resp['data'])} rows, total says {total}")
        rows += resp["data"]
        return
    if depth == len(SPLITS):
        raise SystemExit(f"{filters}: {total} rows and no splits left")
    field = SPLITS[depth]
    ...
    for values in (guess, RANGES[field]):
        children = []
        for v in values:
            child = dict(filters, **{field: v})
            n = query(child)["meta"]["total"]
            if n:
                children.append((child, n))
        if sum(n for _, n in children) == total:
            break
    ...
    for child, n in children:
        collect(child, n, depth + 1, rows, seen)
  • SPLITS is ["word_count", "letter_count"], and LIMIT is 1,000.
  • guess is the set of split values seen in the slice’s first 1,000 rows. When those don’t cover the total, the script tries every value in RANGES (1 to 20 words, 1 to 80 letters).
  • query() caches each response in data/.cache/fetch/, keyed by a hash of the URL, so a rerun resumes without fetching again. There is a one-second pause between calls.

At the end, the script checks that the rows it collected match the API’s grand total. The API lists 24 puzzles twice, so the script writes only the distinct (answer, category) pairs.

Steps:

  • Run the fetch:
    .venv/bin/python scripts/01_fetch_puzzles.py
    
    It prints each category’s row count as it goes, and a final line with the totals.
  • You don’t need to run this to rebuild the project. The site changes as new episodes air, so a fresh fetch gives a different file. The puzzles.csv in the repository is the fetch every later phase used.

Output: data/puzzles.csv, 68,197 rows, 2.1 MB.

answer,category
A CANDLE IN THE WINDOW,Around The House
A CANVAS BACKPACK,Around The House
A JAR OF PENNIES,Around The House
...
FAME AND FORTUNE,Phrase
FAMILIARITY BREEDS CONTEMPT,Phrase
FAMILY DINING EXPERIENCE,Phrase

Phase 2: Cleaning and the control split

Script: 02_clean_puzzles.py. Reads puzzles.csv, writes data/train.csv and data/control.csv.

02_clean_puzzles.py applies five rules, in order:

  1. Merge category names. Three categories appear under two names. STAR & ROLE becomes Star And Role, Character becomes Fictional Character, and Fictional Place becomes Place.
  2. Drop answers with characters the board can’t show. The board has the 26 letters, space, and - & ' . ! ? ,. Answers with $, %, digits and the like go.
  3. Drop answers that don’t fit the TV board. The source includes mobile-game puzzles, and some are longer than the TV board allows.
  4. Drop exact repeats of (answer, category).
  5. Split off control. An answer goes to control if the SHA-1 of its text falls in the bottom 5%.

The board check is shared with the rest of the project, in wof_common.py:

# The TV board: rows of 12/14/14/12 squares, in use since 1997
BOARD_ROWS = (12, 14, 14, 12)
BOARD_CHARS = set("ABCDEFGHIJKLMNOPQRSTUVWXYZ -&'.!?,")

def pieces(answer):
    """Split into (text, joined) pieces. A row can break between words or
    right after a hyphen. joined=True means the piece continues the previous
    one with no space, e.g. BLACK-AND-WHITE -> BLACK- AND- WHITE."""
    out = []
    for word in answer.split():
        parts = word.split("-")
        for k, part in enumerate(parts):
            text = part + ("-" if k < len(parts) - 1 else "")
            if text:
                out.append((text, k > 0))
    return out

def fits_board(answer):
    """Greedy wrap. Filling each row as far as it goes is optimal when the
    rows are fixed and in order, so a False here means no layout fits."""
    ps = pieces(answer)
    i = 0
    for width in BOARD_ROWS:
        used = 0
        while i < len(ps):
            text, joined = ps[i]
            need = len(text) + (1 if used and not joined else 0)
            if used + need > width:
                break
            used += need
            i += 1
    return i == len(ps)

The greedy wrap is enough. The rows have fixed widths and come in a fixed order, so pushing as many words as fit onto each row never makes a later row worse. If the greedy wrap runs out of rows, no layout fits.

The control split is one function:

CONTROL_PCT = 5

def in_control(answer):
    h = int(hashlib.sha1(answer.encode()).hexdigest(), 16)
    return h % 100 < CONTROL_PCT

The split uses only the answer text. That has two consequences:

  • It is the same on every run and on every machine. There is no random seed to record.
  • An answer filed under two categories lands on one side. GULLIVER’S TRAVELS is listed under both Book Title and Title. If the two copies were split separately, one could train a solver while the other tests it. The script ends by asserting that no answer is on both sides.

Steps:

  • Run it:
    .venv/bin/python scripts/02_clean_puzzles.py
    
  • It prints a report of every drop:
    read      68197
    dropped     133  characters
    dropped      26  too long for board
             bad characters: "$%()*/169:;
    train     64689  -> data/train.csv
    control    3349  -> data/control.csv  (3349 distinct answers)
    

Output: data/control.csv, 3,349 puzzles, 108 KB, and data/train.csv with 64,689 puzzles until phase 3 runs.

answer,category
A ROLL OF FILM,Around The House
ACCENT FURNITURE,Around The House
...

Phase 3: Dev split

Script: 03_split_dev.py. Reads train.csv, writes data/train.csv and data/dev.csv.

transformer-solve has one setting to tune: how sure it must be before it solves. Tuning that on the control puzzles would leak them into the result, so phase 3 holds 300 more answers out of training, as the dev set.

The dev answers are the 300 with the lowest SHA-1 of dev|<answer>. The dev| prefix makes this hash independent of the control hash, so the two splits don’t line up.

DEV_SIZE = 300

def key(answer):
    return hashlib.sha1(f"dev|{answer}".encode()).hexdigest()

def main():
    pool = wc.read_puzzles("train.csv")
    if (wc.DATA / "dev.csv").exists():
        pool += wc.read_puzzles("dev.csv")
    pool = sorted(set(pool), key=lambda p: (p[1], p[0]))   # dev rows may be in both
    dev_answers = set(sorted({a for a, _ in pool}, key=key)[:DEV_SIZE])
    sides = {"train.csv": [], "dev.csv": []}
    for answer, category in pool:
        sides["dev.csv" if answer in dev_answers else "train.csv"].append((answer, category))

The script reads train.csv and any existing dev.csv back together before splitting. That makes it safe to run twice: a second run gives the same two files.

Steps:

  • Run it:
    .venv/bin/python scripts/03_split_dev.py
    
    It prints:
    train.csv   64389 puzzles
    dev.csv       300 puzzles
    

Output:

FileRowsSize
data/train.csv64,3892.0 MB
data/dev.csv30012 KB
data/control.csv3,349108 KB
answer,category
ANTIQUE POLAROID CAMERA,Around The House
BAMBOO PLACE MATS,Around The House
...

From here on, every phase reads these three files and never changes them.

The three puzzle files

The end state:

How the puzzles are split. Phase 1 fetches 68,197 puzzles from wofanswers.com into puzzles.csv. Phase 2 drops puzzles that don't fit the TV board, merges categories, and holds out 5% of answers by hash as control.csv (3,349 puzzles). Phase 3 holds 300 more answers out of train.csv as dev.csv, leaving 64,389 puzzles in train.csv.
  • control.csv is held out from everything. Neither solver learns from it, and every result on this page is measured on it.
  • dev.csv is held out from the transformer’s training. It is used only to tune one setting, the confidence transformer-solve needs before it solves.
  • train.csv is what both solvers learn from. smart-solve also adds dev.csv to its word list, since it has nothing to tune.

puzzles.csv and the three split files have the same layout: CSV with a header row and two columns, answer and category, sorted by category and then answer.

Four categories make up 48.6% of the training puzzles: Thing, Food And Drink, Phrase and What Are You Doing. Results on this page are over all 3,349 control puzzles, so those four carry the most weight.

Phase 4: Category lists

The training puzzles cover 64,389 answers. That is a lot of Wheel of Fortune, but not much of the world. A held-out puzzle like a film title or a city the model has never seen can still be solved if the model knows that the name exists. Phase 4 builds lists of real names and phrases for the show’s categories from four open sources. They feed two things later: the transformer’s LIST training lines in phase 8, and the wordplay categories.

ScriptReadsWritesLines
fetch_wikidata.pythe QLever Wikidata endpointdata/sources/raw/wikidata/*.json (42 files)
build_wikidata.pythe 42 JSON filesdata/sources/wikidata.tsv116,403
build_wiktionary.pythe kaikki.org Wiktionary extractdata/sources/wiktionary.tsv21,487
build_wordnet.pyOpen English WordNet 2025data/sources/wordnet.tsv38,256
build_wordplay.pythe three lists above, the puzzles, the CMU Pronouncing Dictionarydata/sources/wordplay.tsv47,032

All four outputs have the same layout: tab-separated category and phrase, no header, sorted. The four TSVs are in the repository. The raw inputs, with their URLs and checksums, are listed in data/sources/SOURCE.md.

One writer for every list

Every builder produces (category, phrase) pairs and hands them to write() in src_common.py. That one function applies the board rules to all four lists:

def normalize(text):
    """Board form of a phrase, or None if it holds characters a board cannot
    show (digits, symbols, other scripts). Parenthesised notes are dropped:
    "Titanic (1997 film)" becomes TITANIC."""
    text = PARENS.sub("", text.translate(TRANSLATE))
    text = unicodedata.normalize("NFKD", text)
    text = "".join(ch for ch in text if not unicodedata.combining(ch)).upper()
    text = SPACES.sub(" ", text).strip(" ,.")
    if not text or set(text) - BOARD_CHARS or not any(ch.isalpha() for ch in text):
        return None
    return text

def write(name, pairs):
    control = control_solutions()
    seen = set()
    counts = Counter()
    dropped = Counter()
    for category, phrase in pairs:
        p = normalize(phrase)
        if p is None:
            dropped["characters"] += 1
        elif not fits_board(p):
            dropped["too long for board"] += 1
        elif p in control:
            dropped["control solution"] += 1
        elif (category, p) in seen:
            dropped["repeat"] += 1
        else:
            seen.add((category, p))
            counts[category] += 1

normalize() does the following:

  • strips accents, so É becomes E.
  • maps the letters that don’t decompose (ß to SS, Ł to L), and curly quotes to straight ones.
  • drops parenthesised notes.
  • uppercases the text.

A phrase with anything the board can’t show (digits, symbols, other scripts) is dropped. So is one that doesn’t fit the 12/14/14/12 board.

The control check is the important line. Any phrase that is a control answer is dropped, from every list. Wikidata knows thousands of film titles, and some of them are control puzzles. Without this line, the transformer would study the test.

Wikidata: names, titles and places

Wikidata is the structured database behind Wikipedia, and it is CC0. fetch_wikidata.py runs 42 SPARQL queries against QLever, a fast public mirror of Wikidata. Each query returns the English label of every item of a class, with its sitelink count: the number of Wikipedia language editions with an article on it. The script uses that count as a fame filter, so obscure items are cut.

QUERIES = {
    # titles
    "films": simple(["Q11424"], 10),
    "tv_series": simple(["Q5398426"], 8),
    "novels": simple(["Q8261", "Q7725634"], 10),
    "songs": simple(["Q7366", "Q134556"], 8),
    # people and groups
    "actors": people("Q33999", 30),
    "singers": people("Q177220", 30),
    ...
    # joins on the lists above, so they run last
    "books_authors": label_of("novels", "wdt:P50", "author", min_links=15),
    "songs_artists": label_of("songs", "wdt:P175", "artist"),
    ...
}

simple(["Q11424"], 10) means: every item that is an instance of film (Q11424) with an English label and at least 10 sitelinks. people("Q33999", 30) is every human whose occupation is actor, with at least 30. A few lists are joins on earlier ones. books_authors takes the novels already fetched and asks for each one’s author.

QLever stops any query at 30 seconds. A join over a whole class (every novel with its author) runs longer than that, so a join is fetched in two steps. First comes the base list, then the linked values for 500 items at a time in a VALUES block. Each list is saved as raw JSON:

[{"item": "http://www.wikidata.org/entity/Q1000394", "label": "This Modern Age", "links": "10"},
 {"item": "http://www.wikidata.org/entity/Q1000826", "label": "Guns of the Magnificent Seven", "links": "14"},
 ...]

build_wikidata.py maps each raw file to show categories and formats the joins the way the show writes them:

CategoryBuilt fromExample line
Movie TitlefilmsTHE GENERAL'S DAUGHTER
On The Mapcountries, states, cities, rivers, lakes, islands… and US cities as CITY STATEELKHART KANSAS
Star And Rolefilm cast with character names, only for actors in the people listsHEATH LEDGER AS NED KELLY
Title/Authornovels joined to their authorsTITLE BY AUTHOR
Song Artistsongs joined to their performersSONG BY ARTIST
Husband And Wifespouses, both in the people listsshared last names go once: BETTINA & LUDWIG ACHIM VON ARNIM
Proper Name, Show Bizactors, singers, writers, athletes, brands, companies, bands, teams

Steps:

  • The Wikidata endpoint changes daily, so a fresh fetch gives different results. Use the published raw files to rebuild exactly:
    mkdir -p data/sources/raw
    curl -L https://github.com/dannyheskett/wof-solver/releases/download/v1/wikidata-raw.tar.gz \
        | tar -xz -C data/sources/raw
    
  • Or fetch your own, with the result differing from mine. The script pauses 20 seconds between queries, and runs the joins in batches of 500 items:
    .venv/bin/python scripts/04_category_lists/fetch_wikidata.py
    
  • Build the list:
    .venv/bin/python scripts/04_category_lists/build_wikidata.py
    
    wikidata: 116,403 lines -> data/sources/wikidata.tsv
          21,697  On The Map
          21,294  Movie Title
          20,412  Proper Name
           9,571  Show Biz
           7,709  Star And Role
           ...
        dropped: too long for board 1,395, characters 4,354, repeat 19,796, control solution 407
    

The 407 dropped control solutions are puzzles from the held-out set that Wikidata also knows.

Wiktionary: phrases and “-ing” activities

Two of the show’s biggest categories are Phrase and What Are You Doing (TAKING A WALK, BAKING COOKIES). Wiktionary has both, as multiword entries. kaikki.org publishes the English Wiktionary as one JSON object per entry, 3.3 GB in all. build_wiktionary.py reads it one line at a time:

PHRASE_POS = {"phrase", "proverb", "prep_phrase", "intj"}
IDIOM_POS = {"noun", "adj", "adv"}
DEAD = {"obsolete", "archaic", "form-of", "alt-of", "misspelling"}
BAD = {"vulgar", "offensive", "derogatory", "slur", "ethnic"}
...
            if " " not in word or not senses:
                continue
            if all(t & DEAD for t in tags) or any(t & BAD for t in tags):
                continue
            if pos in PHRASE_POS or (pos in IDIOM_POS and any("idiomatic" in t for t in tags)):
                if is_common(word):
                    phrases.append(word)
            elif pos == "verb":
                verbs.append(word)
  • Phrase: multiword phrases, proverbs and interjections, plus idiomatic nouns, adjectives and adverbs.
  • What Are You Doing: multiword verbs (“take a walk”), with the first word turned into its present participle. The participle comes from the verb’s own Wiktionary entry when it has one (take to taking), and from spelling rules otherwise.
  • Dropped:
    • entries where every sense is obsolete or archaic, and any entry with a sense tagged vulgar or offensive.
    • entries with a word missing from the SCOWL word list, which drops Latin and other foreign phrases.
    • verb phrases with “someone” or “something” in them.
  • Rewritten: “one’s” becomes MY and “oneself” becomes MYSELF, the way the show writes them: GUARDING MY TONGUE.

Steps:

  • Download the extract (3.3 GB). kaikki.org rebuilds it from new Wiktionary dumps, so today’s file will differ from mine. The wiktionary.tsv in the repository is the one every later phase used.
    curl -L -o data/sources/raw/kaikki-english.jsonl \
        https://kaikki.org/dictionary/English/kaikki.org-dictionary-English.jsonl
    
  • Build it:
    .venv/bin/python scripts/04_category_lists/build_wiktionary.py
    
    participles 41,187, phrases 9,518, verb phrases 12,282
    wiktionary: 21,487 lines -> data/sources/wiktionary.tsv
          12,224  What Are You Doing
           9,263  Phrase
        dropped: repeat 157, control solution 63, too long for board 93
    

Sample lines:

What Are You Doing	BARING MY BREAST
What Are You Doing	COMING BACK FROM THE DEAD
What Are You Doing	GUARDING MY TONGUE

WordNet: everyday nouns

The show’s Thing, Around The House and In The Kitchen puzzles are everyday nouns. Open English WordNet organises nouns into a tree of synsets (groups of synonyms), each filed in a lexicographer file such as noun.food or noun.artifact. build_wordnet.py takes categories from whole files, or from everything below a root synset:

LEXFILES = {
    "noun.food": ["Food And Drink"],
    "noun.animal": ["Living Thing"],
    "noun.plant": ["Living Thing"],
    "noun.artifact": ["Thing"],
    "noun.object": ["Thing"],
    "noun.person": ["Person"],
}
# (category, lexfile, lemma of the root synset[, word in its definition])
ROOTS = [
    ("Occupation", "noun.person", "professional"),
    ...
    ("Around The House", "noun.artifact", "furniture"),
    ("Around The House", "noun.artifact", "home appliance"),
    ...
    ("In The Kitchen", "noun.artifact", "kitchen utensil"),
    ("In The Kitchen", "noun.artifact", "cutlery", "eating"),
    ...
]

A root can carry a word from its definition, to pick one sense of an ambiguous lemma (a dictionary headword): “cutlery” with “eating” is the knives and forks, not the cutting tools. Capitalised lemmas (proper names, Latin genus names) are skipped, as are lemmas with any word outside the SCOWL list. People is built from the Person lemmas with simple plural rules.

Steps:

  • Download WordNet 2025 (11 MB):
    curl -L -o data/sources/raw/english-wordnet-2025.xml.gz \
        https://github.com/globalwordnet/english-wordnet/releases/download/2025-edition/english-wordnet-2025.xml.gz
    
  • Build it:
    .venv/bin/python scripts/04_category_lists/build_wordnet.py
    
    It prints each root with its synset and lemma counts, then the totals:
    wordnet: 38,256 lines -> data/sources/wordnet.tsv
          12,472  Thing
           7,627  Living Thing
           5,978  Person
           5,944  People
           2,430  Food And Drink
           ...
    

Wordplay: the joined categories

Four of the show’s categories are wordplay on other phrases. build_wordplay.py makes them from the three lists above plus the training puzzles, so it runs last:

CategoryRuleExample
Same Letterevery word starts with the same letterCREEPY CRAWLY CREATURES
Rhyme Timetwo different words of 3+ letters rhymeHIKING AND BIKING
Before And Afterphrase A’s last word starts phrase B, joined on itRUBBER BAND + BAND OF BROTHERS, so RUBBER BAND OF BROTHERS
Same Nametwo phrases ending in the same wordBOWLING BALL + DEBUTANTE BALL, so BOWLING & DEBUTANTE BALL

Rhymes come from the CMU Pronouncing Dictionary. Two words rhyme when their sounds match from the last stressed vowel on:

def load_rhymes():
    rhyme = {}
    for line in open(RAW / "cmudict.dict", encoding="latin-1"):
        word, *phones = line.split("#")[0].split()
        word = re.sub(r"\(\d+\)$", "", word).upper()
        stressed = [i for i, p in enumerate(phones) if p[-1] in "12"]
        if stressed and word not in rhyme:
            rhyme[word] = " ".join(p.rstrip("012") for p in phones[stressed[-1]:])
    return rhyme

HIKING is HH AY1 K IH0 NG, so its rhyme key is AY K IH NG. BIKING has the same key.

The two joins give millions of candidates (2,847,612 for Before And After). The script shuffles them with a fixed seed, random.Random("wof-wordplay-v1"), and keeps 20,000 of each. The fixed seed makes every run give the same file. The joins are mechanical, so many are odd (CUTTING RED TAPE RECORDING, A BRISK & DAYTIME JOG). They teach the format of the category more than real phrases.

Steps:

  • Download the CMU dictionary, pinned to the commit used:
    curl -L -o data/sources/raw/cmudict.dict \
        https://raw.githubusercontent.com/cmusphinx/cmudict/0f8072f814306c5ee4fbf992ed853601b12c01f9/cmudict.dict
    
  • Build it:
    .venv/bin/python scripts/04_category_lists/build_wordplay.py
    
    194,826 phrases, 126,020 pronunciations
      candidates: Before And After 2,847,612, Same Name 44,342
    wordplay: 47,032 lines -> data/sources/wordplay.tsv
          19,915  Before And After
          18,357  Same Name
           7,297  Same Letter
           1,463  Rhyme Time
    

smart-solve

smart-solve is the opponent, and the transformer’s teacher: its moves label the training data in phases 6 and 7. It has no training step. It is code in two shared modules, wof_common.py (the word list and the letter choice) and solvers.py (the game moves).

The word list

smart-solve’s vocabulary is two things together:

Words are grouped by shape: their length, with apostrophes in place. CAN'T has shape ___'_. For each shape the code builds a matrix of position bitmasks. P[i, c] has bit k set when word i has letter c at position k:

def posmasks(word):
    row = [0] * 26
    for k, ch in enumerate(word):
        if ch in IDX:
            row[IDX[ch]] |= 1 << k
    return row

For FILM, the row has F = 0001, I = 0010, L = 0100, M = 1000, and 0 for every other letter. That one integer per letter answers the question the solver asks over and over: “if this letter is called, which squares light up?”

Candidates

A board word’s candidates are the words of its shape that match every called letter exactly: the letter shows where it shows, and nowhere else. A called letter that missed must not appear at all. With the bitmasks, that is one comparison per called letter:

            if cols:
                want = np.array(wc.posmasks(key), dtype=np.int32)[cols]
                idx = np.nonzero((P[:, cols] == want).all(axis=1))[0]

key is the board word as shown, such as F_LM, and cols are the called letters. For F_LM with A F L M O S called, a candidate must have F at position 0, L at 2 and M at 3, no other F, L or M, and no A, O or S, since those were called and don’t show in this word.

Weights

Each candidate weighs 1 plus how often it appears in training puzzles of the board’s category. This is how the category steers the choice. In Around The House, LAMP appears in 18 training puzzles and LIMP in none, so LAMP weighs 19 and LIMP weighs 1.

    def weights(self, category, sh):
        key = (category, sh)
        if key not in self._weights:
            counts = self.counts[category]
            self._weights[key] = np.array(
                [1.0 + counts.get(w, 0) for w in self.words[sh]])
        return self._weights[key]

Choosing a letter

The core of smart-solve is one score per letter: the expected log2 of the number of candidates left after calling it, summed over the board’s words. A letter that splits the candidates into many small groups scores low. A letter that leaves them in one big group scores high. smart-solve calls the letter with the lowest score.

        P = self.v.P[it["shape"]][idx]
        w = self.v.weights(self.category, it["shape"])[idx]
        total = w.sum()
        out = {}
        for letter in letters:
            _, inv, n = np.unique(P[:, IDX[letter]], return_inverse=True,
                                  return_counts=True)
            W = np.bincount(inv, weights=w)
            out[letter] = float((W / total * np.log2(n)).sum())

For each letter, P[:, IDX[letter]] is the bitmask of where that letter sits in each candidate. Candidates with the same bitmask would look the same after the call, so they form a group. n is each group’s size and W its total weight. The score is the weighted average of log2(n): the bits of uncertainty left. Ties go to the more common letter.

Here are the scores on the opening board of A ROLL OF FILM, _ ____ __ ____ in Around The House. Before any call, the four words hold 4.70, 11.76, 8.27 and 11.76 bits of uncertainty, 36.49 in all. Each column is the expected bits left in one word after the call:

Letter___________Total
S4.5410.307.7310.3032.85
L4.5410.497.8410.4933.35
R4.5410.577.7310.5733.39
T4.5410.627.6710.6233.46
N4.5410.727.5310.7233.51
D4.5410.837.7510.8333.95
Z4.5411.558.0911.5535.72

S wins because it splits the four-letter words best. It appears in many of them, in many different positions, so seeing where it lands (or that it doesn’t) tells the most. N would do better on the two-letter word, but the two four-letter words count twice. Z barely moves anything. Vowels aren’t in the table because none can be bought yet. The one-letter word scores the same for every consonant: its 26 candidates are the single letters, and each consonant matches exactly one of them.

Two rules shape the choice:

  • The money rule. Until a called consonant shows on the board, vowels are left out.
  • When to solve. smart-solve solves when every hidden word has exactly one candidate, and fills each word with it.

A game

Here is smart-solve playing the control puzzle A ROLL OF FILM, in Around The House. The candidate counts are per hidden word, read from SmartSolver.view() at each turn. A word drops out of the list once it is fully shown:

board            called     candidates per hidden word    move
_ ____ __ ____              [26, 3473, 308, 3473]         CALL S
_ ____ __ ____   S          [25, 2437, 275, 2437]         CALL L
_ __LL __ __L_   LS         [24, 47, 252, 159]            BUY A
A __LL __ __L_   ALS        [38, 214, 105]                BUY O
A _OLL O_ __L_   ALOS       [6, 14, 63]                   CALL F
A _OLL OF F_L_   AFLOS      [6, 3]                        CALL M
A _OLL OF F_LM   AFLMOS     [5, 1]                        CALL R
A ROLL OF F_LM   AFLMORS    [1]                           SOLVE A ROLL OF FILM

S misses, but it still helps: 1,036 four-letter words with an S drop out of each four-letter slot. L hits in two words, which cuts ____ from 2,437 candidates to 47. Once L shows, vowels can be bought. After seven letters, the last word is F_LM with one candidate, and smart-solve solves. In data/eval/smart-solve.tsv this game is:

main	Around The House	A ROLL OF FILM	solved	letters=7 buys=2 misses=1 order=SLAOFMR

The bonus round

In the bonus round the solver names four letters at once, with no feedback between them. best_set tries every combination of 3 consonants and 1 vowel that are not R S T L N E. For each combination it computes the chance that filling every word with its heaviest candidate will be right once those letters show. It keeps the best:

        for cs in itertools.combinations(consonants, n_cons):
            for vs in itertools.combinations(vowels, n_vow):
                chosen = cs + vs
                tie = sorted(rank[c] for c in chosen)
                p = 1.0
                for codes, w, total in live:
                    key = np.zeros(len(w), dtype=np.int64)
                    for c in chosen:
                        inv, n = codes[c]
                        key = key * n + inv
                    first = np.unique(key, return_index=True)[1]
                    p *= w[first].sum() / total

That is 16 consonants choose 3 (Y counts as a consonant) times 4 vowels: 2,240 combinations per puzzle. For each word, key gives every candidate a code for how the four letters would light it up. Candidates with the same code can’t be told apart. Within a group the solver will guess the heaviest, so the chance of being right is the weight of each group’s heaviest member over the total. Words are taken as independent, so the chances multiply.

On A ROLL OF FILM:

_ R_LL __ __L_    PICK CKM O    ->    _ ROLL O_ __LM    SOLVE A ROLL OF FILM

What smart-solve can’t do

smart-solve only knows words in its list. 163 of the 10,649 words in the control puzzles (1.5%) are not in it, in 159 puzzles. On those, it either never pins the word, or pins a wrong word that happens to fit. All 57 of its wrong MAIN solves come from those puzzles. Some missing words are typos in the source data (HADRIN, LASANGA, CHEDDER). It also has no idea which phrases are real. To smart-solve, a board is a set of independent words.

Phase 5: Pretraining corpus

A model that only ever sees Wheel of Fortune lines has to learn English from 64,000 short phrases. Pretraining gives it English first: spelling, common words, which words follow which. Phase 5 builds that text, one billion characters, from two open datasets on Hugging Face.

ScriptReadsWrites
05_make_corpus.py5 Parquet files in data/corpus/raw/data/corpus/train.txt, data/corpus/val.txt

Half comes from each. Wikipedia is clean, but it reads like an encyclopedia. FineWeb-Edu adds plainer, more everyday writing.

Cleaning

The model will only ever see board characters, so the corpus is cut down to them:

def clean_line(line):
    line = unicodedata.normalize("NFKD", line.translate(TRANSLATE))
    line = "".join(ch for ch in line if not unicodedata.combining(ch)).upper()
    line = SPACES.sub(" ", NOT_BOARD.sub(" ", line))
    line = REPEAT.sub(r"\1", STRAY.sub(r"\1", line)).strip(" ,.")
    letters = sum(ch.isalpha() for ch in line)
    return line if letters >= MIN_LETTERS else None
  • Accents are stripped and the text is uppercased, as in phase 4.
  • Every character outside A-Z, space and - & ' . ! ? , becomes a space. That removes digits, brackets and symbols.
  • STRAY and REPEAT tidy the punctuation left behind where numbers were. “IN 1997, THE” would become “IN , THE”, and this fixes it to “IN, THE”.
  • Paragraphs with fewer than 40 letters are dropped. That removes headings, captions and table rows.
  • Wikipedia articles stop at their References, See also, External links or Notes section.

Interleaving and the validation split

The script reads both sources at once and always takes the next document from whichever source has contributed fewer characters so far, so they stay at half each. Each source stops at 500 million characters. About 1% of documents go to val.txt instead, chosen by the SHA-1 of the text, so the split is the same on every run:

            h = int(hashlib.sha1(doc.encode()).hexdigest(), 16) % 100
            if h < VAL_PCT:
                val.write(doc)

Steps:

  • Download the five Parquet files (3.1 GB). The URLs are pinned to the dataset commits I used:
    mkdir -p data/corpus/raw && cd data/corpus/raw
    W=https://huggingface.co/datasets/wikimedia/wikipedia/resolve/b04c8d1ceb2f5cd4588862100d08de323dccfbaa/20231101.en
    for n in 00000 00010 00020 00030; do
      curl -L -o wiki-$n.parquet "$W/train-$n-of-00041.parquet"
    done
    curl -L -o fineweb-000_00000.parquet \
      "https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/87f09149ef4734204d70ed1d046ddc9ca3f2b8f9/sample/10BT/000_00000.parquet"
    cd ../../..
    
    Their checksums are in data/corpus/SOURCE.md.
  • Build the corpus:
    .venv/bin/python scripts/05_make_corpus.py
    
    wiki    132,888 documents    500,005,177 characters
    web     110,792 documents    500,001,314 characters
    train 1,000,006,491 characters, val 10,141,309 characters
    

Output:

FileSize
data/corpus/train.txt954 MB
data/corpus/val.txt9.7 MB

One paragraph per line, with a blank line between documents:

ANARCHISM IS A POLITICAL PHILOSOPHY AND MOVEMENT THAT IS SKEPTICAL OF ALL JUSTIFICATIONS FOR
AUTHORITY AND SEEKS TO ABOLISH THE INSTITUTIONS IT CLAIMS MAINTAIN UNNECESSARY COERCION AND
HIERARCHY, TYPICALLY INCLUDING NATION-STATES, AND CAPITALISM. ...

Phase 6: MAIN game states

Script: 06_make_main.py. Reads train.csv and control.csv, writes data/main_train.tsv and data/main_control.tsv.

The transformer learns to play by reading game states with the right move after each one. Phases 6 to 8 write those states. smart-solve plays every training puzzle many ways, and every board along the way becomes one line, labelled with the move smart-solve would make there.

The _control files hold the same kinds of lines for the control puzzles. They are used only to measure validation loss during training, never to train.

All four TSVs share one layout, tab-separated, no header:

category  board  called  target  solution

called is the letters called so far, sorted A-Z, or - before any call. solution is kept for checking and is not shown to the model.

A solver at play will miss letters smart-solve wouldn’t call, and land on boards smart-solve’s own games never pass through. If the training data only had smart-solve’s own games, the model would know what to do on the perfect path and nothing off it. So each puzzle is walked five ways:

  • three times in a random letter order.
  • once in order of letter frequency.
  • once following smart-solve’s own picks.

Every board a walk passes through becomes one line. The label is always the move smart-solve would make on that board, whatever the walk calls next:

def walk(bw, solution, category, policy, rng):
    called = set()
    order = rng.sample(wc.LETTERS, 26) if policy == "random" else ORDER
    while True:
        b = wc.board(solution, called)
        c = "".join(sorted(called)) or "-"
        if bw.solved():
            yield f"{category}\t{b}\t{c}\tSOLVE {solution}\t{solution}\n"
            return
        vowels = wc.can_buy(b, called)
        best = bw.smart_pick(called, ORDER, vowels)
        yield f"{category}\t{b}\t{c}\t{wc.move(best)}\t{solution}\n"
        nxt = best if policy == "smart" else next(
            x for x in order if x not in called and (vowels or x not in wc.VOWELS))
        called.add(nxt)
        bw.call(nxt)
  • best is smart-solve’s pick on this board, and it becomes the line’s target.
  • nxt is what the walk actually calls. That is best on the smart walk, and the next letter in its own order on the others.
  • A walk ends at its first SOLVE line. Lines repeated within a puzzle are dropped.
  • Each puzzle’s random orders are seeded from SEED = "wof-main-v1" and the puzzle itself, so every run writes the same file.

Here are lines from one random walk through A CANDLE IN THE WINDOW. The walk calls J, K, V and U, which miss. The label stays CALL N, because N is still the best move on that board:

Around The House	_ ______ __ ___ ______	-	CALL S	A CANDLE IN THE WINDOW
Around The House	_ ____L_ __ ___ ______	L	BUY A	A CANDLE IN THE WINDOW
Around The House	A _A__L_ __ ___ ______	AL	BUY E	A CANDLE IN THE WINDOW
Around The House	A _A__L_ I_ ___ _I____	AIL	CALL N	A CANDLE IN THE WINDOW
Around The House	A _A__L_ I_ ___ _I____	AIKL	CALL N	A CANDLE IN THE WINDOW
Around The House	A _A__L_ I_ ___ _I____	AIJKL	CALL N	A CANDLE IN THE WINDOW
Around The House	A _A__L_ I_ ___ _I____	AIJKLV	CALL N	A CANDLE IN THE WINDOW
Around The House	A _A__L_ I_ ___ _I____	AIJKLUV	CALL N	A CANDLE IN THE WINDOW

Across all 5,131,187 lines in main_train.tsv, the labels are:

TargetLines
CALL2,802,502
BUY2,007,153
SOLVE321,532

The most common single labels are BUY E (970,958), BUY A (558,166), CALL S (503,917), CALL T (371,377) and CALL R (360,841). That skew comes straight from smart-solve. E is the most common letter in the puzzles, and buying it as soon as the money rule allows is smart-solve’s most frequent move. The model will inherit this habit.

The labels for the training puzzles use a word list built from train.csv alone, with weights from train.csv alone. Neither dev nor control shapes a training label.

Steps:

  • Run it. It uses 11 worker processes:
    .venv/bin/python scripts/06_make_main.py
    
    It prints progress every 5,000 puzzles and a summary line per file.

Output:

FileLinesSize
data/main_train.tsv5,131,187349 MB
data/main_control.tsv267,21619 MB

Phase 7: BONUS picks and solves

Script: 07_make_bonus.py. Reads train.csv and control.csv, writes data/bonus_train.tsv and data/bonus_control.tsv.

The bonus round needs two kinds of lines:

  • PICK lines. The board after R S T L N E, labelled with smart-solve’s best_set pick. There is one per puzzle.
  • SOLVE lines. The board after R S T L N E and a set of four letters, labelled with the answer. There are up to 10 per puzzle, from different sets: smart-solve’s set, the four most common letters left, and random sets.
    lines = [f"{category}\t{wc.board(solution, FREE)}\t{free}\t{pick_text(smart)}\t{solution}\n"]
    sets = [frozenset(smart),
            frozenset([c for c in ORDER if c in CONS][:3] + [c for c in ORDER if c in VOWS][:1])]
    tries = 0
    while len(set(sets)) < PICK_SETS and tries < 100:
        sets.append(frozenset(rng.sample(CONS, 3) + rng.sample(VOWS, 1)))
        tries += 1

Training on boards from many different sets, not just smart-solve’s set, teaches the model to read a board with any letters showing. In play, transformer-solve makes its own pick, and that may not be smart-solve’s.

Around The House	_ __N_LE _N T_E __N___	ELNRST	PICK CDM I	A CANDLE IN THE WINDOW
Around The House	_ C_NDLE IN T_E _IND__	CDEILMNRST	SOLVE A CANDLE IN THE WINDOW	A CANDLE IN THE WINDOW
Around The House	A CANDLE _N THE __ND__	ACDEHLNRST	SOLVE A CANDLE IN THE WINDOW	A CANDLE IN THE WINDOW
Around The House	_ __N_LE IN T_E WIN__W	BEIJLNRSTW	SOLVE A CANDLE IN THE WINDOW	A CANDLE IN THE WINDOW

Steps:

  • Run it:
    .venv/bin/python scripts/07_make_bonus.py
    

Output:

FileLinesSize
data/bonus_train.tsv708,27959 MB
data/bonus_control.tsv36,8393.1 MB

Phase 8: Token files

Script: 08_make_train.py. Reads the four TSVs, the category lists and the corpus, writes data/train/.

The model reads characters, not words. Its alphabet has 48 characters, and a character’s token id is its position in this string:

ALPHABET = "\n !&',-./0123456789?ABCDEFGHIJKLMNOPQRSTUVWXYZ_|"

That is newline, space, the board’s punctuation, digits, ?, the 26 letters, _ for a hidden square and | as a field separator. The digits and / are never used after phase 5’s cleaning, but they keep the alphabet general.

Every game line becomes a prompt and a target, ending in a newline:

MAIN|AROUND THE HOUSE|A _A__L_ __ ___ ______|AL|BUY E
BONUS|AROUND THE HOUSE|_ __N_LE _N T_E __N___|ELNRST|PICK CDM I
BONUS|AROUND THE HOUSE|A CANDLE _N THE __ND__|ACDEHLNRST|SOLVE A CANDLE IN THE WINDOW
LIST|ON THE MAP|ELKHART KANSAS

The prompt is everything up to the last |. Training puts loss only on the target and the newline. The model is never trained to predict the board it is given, only the move.

Phase 8 writes five training streams and two validation streams. Each one is a token file (one byte per character) plus an index of (start, prompt length, line length):

    def add(self, prompt, target):
        line = prompt + target + "\n"
        self.index.append((self.pos, len(prompt), len(line)))
        self.pos += len(line)
        self.buf.append(line)
StreamLinesTokensFrom
main5,131,187283Mmain_train.tsv
bonus708,27952Mbonus_train.tsv
list281,17110Mevery training puzzle and every category-list phrase, as `LIST
list_bonus280,14720Mthe same phrases as BONUS SOLVE lines, with R S T L N E plus a seeded random 3 consonants and 1 vowel shown
corpus1,000Mcorpus/train.txt, the text as token ids, loss on everything
main_val, bonus_val304,05517Mthe control TSVs, for validation loss only
corpus_val10Mcorpus/val.txt

The two list streams carry phase 4 into the model. list teaches which phrases exist in each category. list_bonus is bonus-round practice on 280,147 phrases, most of which are not puzzles. Any phrase that is a dev answer is left out of both, so the dev set stays clean for tuning.

Steps:

  • Run it:
    .venv/bin/python scripts/08_make_train.py
    
    It prints each stream’s lines and tokens.

Output: data/train/, 1.5 GB, 15 files:

alphabet.txt
main.bin        main.idx.npy        main_val.bin    main_val.idx.npy
bonus.bin       bonus.idx.npy       bonus_val.bin   bonus_val.idx.npy
list.bin        list.idx.npy
list_bonus.bin  list_bonus.idx.npy
corpus.bin      corpus_val.bin

The transformer

model.py is a small GPT in the style of Andrej Karpathy’s nanoGPT: a decoder-only transformer that predicts the next character from the ones before it.

How a transformer reads a game line

For anyone new to transformers, here is what happens inside this one when it reads a prompt like MAIN|AROUND THE HOUSE|A _A__L_ __ ___ ______|AL|.

  1. Characters become tokens. Each character is replaced by its index in the 48-character alphabet. M is 32, A is 20, | is 47, and so on. The prompt becomes a list of 48 numbers, one per character.
  2. Tokens become vectors. The token embedding is a table with one row of 640 numbers per character. Each token is replaced by its row. A second table, the position embedding, has one row per position from 0 to 255. The row for that position is added in, so the model knows where each character sits. The prompt is now 48 vectors of 640 numbers.
  3. Attention mixes information between positions. In each block, every position builds three vectors from its own: a query, a key and a value. A position compares its query with the key of every earlier position, and the closer the match, the more of that position’s value it takes in. This is how the _ squares of the board can pick up information from the category at the start of the line and the called letters at the end. The causal mask means a position can look back but never forward. Ten heads do this in parallel, each with its own 64-number slice, so different heads can track different things.
  4. The MLP works on each position alone. After attention, each vector goes through a two-layer network that widens it to 2,560 numbers, applies GELU, and narrows it back to 640. Attention moves information between positions. The MLP transforms it.
  5. The residual stream carries everything forward. Attention and the MLP don’t replace a position’s vector. They add to it. After 12 blocks, each vector is the sum of its starting embedding and 24 updates.
  6. The last vector predicts the next character. The vector at the last position, after a final LayerNorm, is compared with each of the 48 token embeddings by dot product. That gives 48 scores, the logits. Softmax turns them into probabilities.

After the prompt above, the model writes B, U, Y, a space and E, then a newline: BUY E, the same move smart-solve makes there. Each written character is fed back in as the next position.

Training adjusts the 59 million numbers so the probability of the right next character goes up. The loss is the negative log of that probability, averaged over the characters that count. A loss of 0.14 nats on main_val means the model gives the right target character a probability of about 87% as a geometric mean: e^−0.14 = 0.87.

The shape

SettingValue
Layers12
Attention heads10, of 64 dimensions each
Width (n_embd)640
Context (block)256 characters
Vocabulary48 characters
Parameters59,045,120 (59.05M), not counting position embeddings

A 256-character context is more than any game line needs. The longest line in the training streams is 158 characters, a MAIN line with its prompt and target.

The block

class Block(nn.Module):
    def __init__(self, c):
        super().__init__()
        self.ln1, self.ln2 = nn.LayerNorm(c.n_embd), nn.LayerNorm(c.n_embd)
        self.attn = Attention(c)
        self.mlp = nn.Sequential(nn.Linear(c.n_embd, 4 * c.n_embd, bias=False), nn.GELU(),
                                 nn.Linear(4 * c.n_embd, c.n_embd, bias=False),
                                 nn.Dropout(c.dropout))

    def forward(self, x, cache=None):
        x = x + self.attn(self.ln1(x), cache)
        return x + self.mlp(self.ln2(x))
  • Pre-norm. LayerNorm comes before attention and before the MLP, and each adds its result back to the input (the residual stream). This is the GPT-2 layout.
  • Attention is PyTorch’s scaled_dot_product_attention with a causal mask, so each position sees only the characters before it.
  • No biases in the linear layers.
  • Tied output. The output layer reuses the token embedding matrix, so the score for each next character is a dot product with that character’s embedding.

Generation

To play a move, the model reads the prompt and writes characters until it writes a newline. complete() does this greedily, always taking the most likely character, with a key/value cache so each new character costs one position, not a rerun of the whole prompt:

            caches = [{} for _ in self.blocks]
            logits = self(torch.tensor([p], device=device), caches=caches)[0, -1]
            pos = len(p)
            for _ in range(max_new):
                row = logits.float()
                first = int(row.argmax())
                banned = ban(i, out[i]) if ban else ()
                if banned:
                    row = row.clone()
                    row[list(banned)] = float("-inf")
                best = int(row.argmax())
                ...
                self.logp[i].append(float(torch.log_softmax(row, dim=0)[best]))
                if best == stop:
                    break
                out[i].append(best)
                logits = self(torch.tensor([[best]], device=device), caches=caches, start=pos)[0, -1]
                pos += 1

Two hooks in it matter for play:

  • ban lets the caller rule out characters at each step. transformer-solve uses it to keep moves legal, as described in transformer-solve: the model as a player.
  • logp records the log probability of each character written. Summed over a SOLVE, it gives the model’s confidence in its answer, and that is what decides whether transformer-solve solves.

Phase 9: Training

Training has two stages:

  1. Pretraining teaches the model English. It reads the corpus and learns to predict the next character, with loss on every character.
  2. Finetuning teaches it the game. It starts from the pretrained weights and trains on a mix of the game streams, with some corpus kept in so the English isn’t lost.
PretrainFinetune
Starts fromrandom weightsruns/pre/ckpt.pt
Datacorpus onlymain 35%, bonus 30%, list_bonus 20%, list 5%, corpus 10%
Steps30,50015,000
Batch128 rows of 256 tokens (32,768 tokens)the same
Tokens seen1.0B0.49B
Learning rate6e-4, 500 warmup steps3e-4, 200 warmup steps
Writesruns/pre/runs/ft2b/

ft2b is the finetuned model behind every result on this page.

The training script

09_train.py builds each batch row from one stream, picked at random by the mix weights. A corpus row is 257 consecutive characters from a random point in corpus.bin. A game row packs whole lines, picked at random, one after another until the row is full. Its loss mask is on only for target characters:

    def row(self, rng, n):
        """n + 1 tokens of whole random lines, and the loss mask for each
        predicted position (True where the next token is target)."""
        x, m = [], []
        while len(x) < n + 1:
            start, plen, length = self.idx[rng.integers(len(self.idx))]
            x.extend(self.tok[start:start + length])
            m.extend([False] * plen + [True] * (length - plen))
        return np.array(x[:n + 1], dtype=np.int64), np.array(m[1:n + 1])

Positions where the mask is off get target -1, which cross_entropy(..., ignore_index=-1) skips. The model still reads the prompt. It just isn’t scored on predicting it.

The rest is a standard nanoGPT-style loop:

  • AdamW with betas 0.9 and 0.95, and weight decay 0.1 on the matrices only.
  • The learning rate warms up linearly, then follows a cosine decay to a tenth of its peak.
  • bfloat16 autocast on the GPU, gradient clipping at 1.0, and torch.compile.
  • Checkpoints. Every --eval_every steps it measures validation loss on each validation stream and saves runs/<name>/ckpt.pt, which holds the weights, the config, the alphabet and the arguments. log.tsv records the losses and throughput.

09_pod/run.sh has the exact commands used:

python3 scripts/09_train.py --name pre --phase pretrain --steps 30500 \
    --eval_every 2000 --compile 2>&1 | tee pre.log
python3 scripts/09_train.py --name ft2b --phase finetune --init runs/pre/ckpt.pt \
    --steps 15000 --eval_every 1000 --lr 3e-4 --warmup 200 --compile 2>&1 | tee ft2b.log

Running it on a rented GPU

I trained on one NVIDIA H100 SXM 80 GB, rented from Runpod in its Secure Cloud. Any machine with a CUDA GPU and PyTorch will do. Smaller GPUs will take longer.

Steps:

  • Bundle what the GPU needs: the model, the training script, run.sh and the token files from phase 8. This writes /tmp/wof-train.tar, about 1.6 GB:
    bash scripts/09_pod/pack.sh
    
  • Copy it to the GPU machine and unpack it. On Runpod, with SSH set up for the pod:
    rsync -P /tmp/wof-train.tar root@<pod-address>:/workspace/
    ssh root@<pod-address> 'cd /workspace && tar -xf wof-train.tar'
    
  • Run both stages, about 32 minutes of training on an H100:
    ssh root@<pod-address> 'cd /workspace && bash scripts/09_pod/run.sh'
    
    run.sh finetune runs only the second stage, from an existing runs/pre/ckpt.pt.
  • Copy the checkpoints back, then stop the pod:
    rsync -P root@<pod-address>:/workspace/runs/pre/ckpt.pt  runs/pre/
    rsync -P root@<pod-address>:/workspace/runs/ft2b/ckpt.pt runs/ft2b/
    
  • Before renting anything, check the script on a CPU with a toy model, in a few minutes:
    .venv/bin/python scripts/09_train.py --name smoke --phase pretrain --n_layer 2 --n_head 2 \
        --n_embd 64 --batch 16 --steps 200 --eval_every 100
    

About exact reproduction. Training on a GPU is not bit-for-bit repeatable. torch.compile and the GPU’s parallel sums can round slightly differently from run to run, even with the same seed. A rerun will give a model with very close losses, but not identical weights. To reproduce the later phases exactly, use the published checkpoints from the release.

What the logs show

Pretraining, from runs/pre/log.tsv. Loss is cross-entropy in nats per character on the held-out corpus_val. The full log is in Appendix A.

A loss of 4.02 at step 0 is about log(48) = 3.87: random guessing over 48 characters. It ends at 0.928 nats, which is 1.34 bits per character. The H100 ran at about 820,000 tokens a second, so pretraining took 20 minutes.

Finetuning, from runs/ft2b/log.tsv, where main_val and bonus_val are measured on the control puzzles’ lines, on the target characters only.

  • The game is learned fast. Most of the drop in main_val happens in the first 1,000 steps.
  • English dips, then recovers. corpus_val rises from 0.928 to 1.03 as the game data comes in, then comes back down to 0.971 as the learning rate decays.
  • The bonus round levels off. bonus_val bottoms out around step 10,000 to 12,000, at 0.083, and ends at 0.085.

Finetuning took 12 minutes at about 686,000 tokens a second.

Output:

FileSize
runs/pre/ckpt.pt237 MB
runs/ft2b/ckpt.pt237 MB

Both are in the release as pre-ckpt.pt and ft2b-ckpt.pt.

transformer-solve: the model as a player

The model writes text. solvers.py turns it into a player, TransformerSolver. For each move it builds the prompt, exactly as in training, and lets the model complete it:

MAIN|AROUND THE HOUSE|A _OLL O_ __L_|ALOS|   ->   CALL F

A free-running model can write anything. Masks hold it to the rules and the board, through the ban hook in complete():

        def ban(_, new):
            if not new:
                return no_buy
            if new in verbs:
                return repeat
            if len(new) >= len(solve) and new[:len(solve)] == solve:
                k = len(new) - len(solve)
                return list(everything - allowed[k]) if k < len(allowed) else []
            return ()
  • First character: B is banned until a vowel can be bought, the money rule.
  • After CALL or BUY : letters already called are banned.
  • After SOLVE : each character must fit its square. Shown letters, spaces and punctuation are copied. A blank takes only an uncalled letter. The answer must end where the board ends.

The SOLVE mask means every word of an answer fills its slot exactly. The model can still pick the wrong word, but it can’t write one of the wrong length. In the evaluation, the masks changed 9,800 characters of letter moves and 3,420 characters of SOLVEs, across 44,694 MAIN moves.

Phase 10: SOLVE cutoff

The model’s first instinct is to solve too early. On the dev puzzles, a transformer-solve that accepts every SOLVE it writes solves only 64% of them. It has a confident-looking guess well before the board supports one.

So a MAIN SOLVE is accepted only when the model’s probability for that answer reaches a set cutoff. The probability is the product of the probabilities of each character after SOLVE , from logp. Below the cutoff, transformer-solve turns the SOLVE down and makes its best letter move instead.

10_tune_solve.py picks the cutoff on the 300 dev puzzles, which the model never trained on. It plays every game with every early SOLVE turned down, so each game runs until the board is full, and it logs each turned-down SOLVE with its probability and whether it was right. One run scores every cutoff:

def score(games, cut):
    solved, letters = 0, []
    for _, _, _, tried, final, _ in games:
        hit = next(((ok, k) for p, ok, k, _ in tried if p >= cut), final)
        if hit[0]:
            solved += 1
            letters.append(hit[1])
    return solved / len(games), letters

This works because a turned-down SOLVE always leads to the same letter move. So a game with cutoff p plays exactly like the logged game up to the first SOLVE with probability at least p. That SOLVE decides the game.

Steps:

  • Run it on the finetuned checkpoint. It uses 11 processes, one torch thread each:
    .venv/bin/python scripts/10_tune_solve.py runs/ft2b/ckpt.pt
    

Result (data/eval/tune.tsv, 2,819 logged SOLVEs from 300 dev games):

CutoffSolved
0 (accept every SOLVE)64.0%
0.5080.3%
0.9094.0%
0.9595.3%
0.9998.3%

Three dev games from tune.tsv show how the confidence moves. Each row is a SOLVE the model wrote and the tuning run turned down:

category  solution                              letters  p       right  answer
Phrase    OUT WITH THE OLD AND IN WITH THE NEW  11       0.9841  1      OUT WITH THE OLD AND IN WITH THE NEW
Phrase    OUT WITH THE OLD AND IN WITH THE NEW  12       0.9774  1      OUT WITH THE OLD AND IN WITH THE NEW
Phrase    OUT WITH THE OLD AND IN WITH THE NEW  13       0.9925  1      OUT WITH THE OLD AND IN WITH THE NEW
Phrase    THROWN FOR A LOOP                     13       0.9035  1      THROWN FOR A LOOP
Phrase    THROWN FOR A LOOP                     14       0.7507  1      THROWN FOR A LOOP
...
Phrase    THROWN FOR A LOOP                     21       0.6956  1      THROWN FOR A LOOP
People    TRIPLETS                              4        0.9840  0      TRUMPETS
People    TRIPLETS                              5        0.9966  0      TRUMPETS
  • OUT WITH THE OLD AND IN WITH THE NEW is right from 11 letters on. With a 0.99 cutoff, transformer-solve solves it at 13.
  • THROWN FOR A LOOP is right from 13 letters on, but the model’s confidence falls as more letters show, and it never reaches 0.99. With the cutoff, it calls letters until the board is full. It still solves, but late.
  • TRIPLETS goes wrong. After four letters the model writes TRUMPETS at 0.984, and after five at 0.997. That clears the cutoff, so this dev game is lost. A high probability is a strong signal, not a guarantee.

Of the 2,819 logged SOLVEs, 372 were wrong, and 215 of those had a probability under 0.5. Most wrong guesses come with low confidence, so one threshold filters out most of them.

0.99 is the cutoff used everywhere after this, in the evaluation and on the server. At 0.99, transformer-solve solved 98.3% of dev puzzles with 12.08 letters on average.

Phase 11: Evaluation

11_evaluate.py plays a solver on all 3,349 control puzzles, in both modes. Every move is checked against the rules:

def play_main(n, solution, category):
    S.new_puzzle(wc.seed_for(SEED, str(n)))
    called = []
    while True:
        board = wc.board(solution, set(called))
        mv = S.main(category, board, "".join(sorted(called)))
        if mv.startswith("SOLVE "):
            return "solved" if mv[6:] == solution else "wrong", called
        letter = check_move(mv, board, called)
        if letter is None:
            return "invalid", called
        called.append(letter)

Steps:

  • smart-solve:
    .venv/bin/python scripts/11_evaluate.py smart-solve
    
  • transformer-solve. The model runs on the CPU, one game per worker process:
    .venv/bin/python scripts/11_evaluate.py transformer-solve:runs/ft2b/ckpt.pt
    
  • A quick check on the first 200 puzzles writes to /tmp instead:
    .venv/bin/python scripts/11_evaluate.py transformer-solve:runs/ft2b/ckpt.pt main 200
    

Output: one row per puzzle and mode in data/eval/smart-solve.tsv and data/eval/transformer-solve.tsv, and one summary row per solver and mode in data/eval/summary.tsv:

mode  category          solution        result  detail
main  Around The House  A ROLL OF FILM  solved  letters=7 buys=2 misses=1 order=SLAOFMR
main  Around The House  A ROLL OF FILM  solved  letters=10 buys=4 misses=4 order=SLAOFREIKT masked=2 fitted=0 rejected=2

The first row is smart-solve and the second is transformer-solve, on the same puzzle. Both open S, L, A, O, F. smart-solve calls M and R and solves. transformer-solve turns down two SOLVEs under the cutoff and calls R, E, I, K and T before it is sure.

Phase 12: Exporting the weights

12_export_weights.py writes the checkpoint as one flat binary file that C can read without a library:

magic   "WOFM" and a uint32 version (2)
uint32  vocab, block, n_layer, n_head, n_embd
char    the alphabet, vocab bytes; a token's id is its index
float16 tensors (IEEE half), in order:
    tok        [vocab][n_embd]   (also the output layer, tied)
    pos        [block][n_embd]
    per layer: ln1.w, ln1.b, qkv, proj, ln2.w, ln2.b, fc, out
    ln.w, ln.b [n_embd]

The weights are stored as float16, half the size of float32: 118 MB instead of 237 MB. The server widens each weight row to float32 as it uses it, with the F16C instruction _mm256_cvtph_ps, and does all its arithmetic in float32:

// y[t][o] = x[t] . w[o] for t < T: each weight row is widened once for all T rows.
static void linear(float *y, const float *x, const half *w, int T, int in, int out) {
    static float row[MAX_IN];
    for (int o = 0; o < out; o++) {
        widen(row, w + (size_t)o * in, in);
        for (int t = 0; t < T; t++) y[(size_t)t * out + o] = dot(x + (size_t)t * in, row, in);
    }
}

Steps:

.venv/bin/python scripts/12_export_weights.py runs/ft2b/ckpt.pt runs/ft2b/wof.bin

Output: runs/ft2b/wof.bin, 118,417,996 bytes, also in the release.

Phase 13: Demo data

13_build_demo.py writes data/demo.json, everything behind the demo’s puzzle endpoints:

  • categories: all 49 training categories, most common first, each marked with whether control has a puzzle in it.
  • puzzles: the 3,349 control puzzles, which Random mode draws from.
  • presets: the seven Demo examples.

The presets are hand-picked control puzzles where transformer-solve beats smart-solve. They are the transformer’s wins, not typical games. Over the control set, smart-solve wins most head-to-head matches. Each preset has a note, such as: “smart-solve’s word list has no HEADSHOT. It solves CELEBRITY HELMSMAN.”

.venv/bin/python scripts/13_build_demo.py

The C server

The demo needs both solvers answering moves over HTTP, fast, on a small cheap machine. In Python, transformer-solve needs PyTorch, and the PyTorch install alone (727 MB) is three times the size of the model. So the server is one C program, wof-server, that runs both solvers and serves the game client:

FileLinesWhat it does
model.c207the transformer’s forward pass, with a key/value cache
transformer.c200transformer-solve: decoding under the same masks, and the 0.99 cutoff
smart.c776smart-solve, ported from wof_common.py and solvers.py
main.c727HTTP, the API, the static files

HTTP parsing is picohttpparser (MIT) and JSON is cJSON (MIT), both vendored.

Parity: the C server plays like the Python

A port is only useful if it plays the same moves. Three tests in server/tests/ compare the C code with the Python it replaces:

TestWhat it comparesResult
model_parity.pylog probabilities and greedy completions on 200 game promptslargest log-probability difference 0.0096, all 200 completions the same
tf_parity.pytransformer-solve’s moves in 100 MAIN and 100 BONUS games1,329 of 1,331 MAIN moves the same, and all 200 BONUS picks and solves
smart_parity.pysmart-solve’s move at every state of all 3,349 control puzzlesall 40,223 moves the same

smart-solve matches move for move because the port copies two details of the Python:

  • The summing order. It adds scores in numpy’s pairwise order, not left to right.
  • The rounding. It rounds the way Python rounds before comparing scores.

Without those, near-ties break differently.

The two transformer moves that differ are float16 effects. Both are close calls between two letters, which float16 rounding tips the other way: CALL S instead of CALL R on one empty board, and BUY O instead of BUY A on another.

Steps:

make -C server
python3 server/tests/model_parity.py runs/ft2b/ckpt.pt runs/ft2b/wof.bin
python3 server/tests/tf_parity.py runs/ft2b/ckpt.pt runs/ft2b/wof.bin
python3 server/tests/smart_parity.py 3349

The tests need torch, so run them with .venv/bin/python if that’s where you installed it.

Running it on Fly.io

The server runs on Fly.io on the smallest machine there is: one shared CPU and 256 MB of RAM.

The weights alone are 118 MB, so fitting in 256 MB takes one more step. The first float16 build read the weights into the process’s own memory, 160 MB in all. The kernel killed it at startup for running out of memory. Now the server maps the weights file into memory with mmap instead of copying it. The weights live in the kernel’s page cache, which the kernel can drop and reread from disk, and the server’s own memory is about 30 MB.

Idle machines suspend after a few minutes and wake on the next request. Measured on the deployed server:

A transformer-solve moveabout 240 ms
A smart-solve moveabout 5 ms
Waking a suspended machine and making a moveabout 1.1 s
The first move after a fresh start (weights read from disk)about 5.6 s

Steps:

  • Build the game client first, since the server image includes it. See The game client.
  • Run the server locally:
    make -C server
    WEIGHTS=runs/ft2b/wof.bin WEB_ROOT=client/build/web PORT=8080 ./server/build/wof-server
    
  • Ask it for a move:
    curl -s -X POST localhost:8080/move \
        -d '{"mode":"main","category":"Phrase","board":"____ __ ___ ____","called":""}'
    
    {"smart":{"move":"CALL T","ms":7},"transformer":{"move":"CALL T","ms":194}}
    
  • Deploy to Fly, from the root of the clone, with your own app name in server/fly.toml:
    fly deploy . --config server/fly.toml --dockerfile server/Dockerfile --remote-only --ha=false
    

The full API is documented at the top of server/src/main.c.

The game client

The demo at wof-solver.fly.dev is a small game in C, written with raylib and compiled to WebAssembly with Emscripten. It is about 1,500 lines in client/src/. It draws two boards side by side, one per solver, and plays them move by move, fetching each move from the server.

Each side shows its board, the letter strip, letters and misses, a stopwatch for the server’s move time, and the move history. For transformer-solve, the history includes the answer it was considering and its confidence, and a SOLVE under 99% shows as held back. On a tall phone screen the two boards stack.

Steps:

  • Build raylib 6.0 for the web, once. This needs the Emscripten SDK on your path:
    cd client
    bash scripts/build_raylib_web.sh
    
  • Build the client and serve it locally:
    make web
    make web-serve        # http://localhost:8000/wof-solver.html
    
    The page talks to https://wof-solver.fly.dev. Add ?api=http://localhost:8080 to use a local server started with CORS_ORIGINS=http://localhost:8000.
  • Run the native tests of the game logic:
    make test
    

Results & Conclusions

Overall results

SolverMAIN solvedWrongLetters (median / mean)Misses (mean)BONUS solved
smart-solve98.3%579 / 9.041.9679.8%
transformer-solve98.7%4212 / 12.333.7777.2%

Neither solver made an invalid move.

  • The solve rates are close, but smart-solve needs fewer letters. Scored head to head on each MAIN puzzle, fewer misses winning and an unsolved puzzle losing, smart-solve wins 2,245 (67%), transformer-solve wins 439 (13%), and 665 are ties.
  • Missing words cause smart-solve’s wrong answers. All 57 of them are in the 159 puzzles with a word missing from its list: a list word fits the board, so it solves early with the wrong word (FILM JOUR FESTIVAL, THE COUPONS for THE GORGONS). Only 10 of transformer-solve’s 42 wrong solves are in those puzzles.
  • In BONUS, they fail on different puzzles. On the 1,348 puzzles where both pick the same letters, smart-solve solves 1,135 and transformer-solve 1,096. smart-solve solves 414 puzzles that transformer-solve misses, and transformer-solve solves 326 that smart-solve misses. Accepting a solve from either would reach 89.5%.
  • transformer-solve has habits. It opens the same way in most games: CALL S, then BUY E and BUY A once a consonant shows. GOOEY COOKIE DOUGH costs it five misses (S R L N T) before its first hit.

By category

Appendix B has the results for the 16 largest categories in the control set.

  • Names are where the word list fails. In MAIN, smart-solve’s worst categories are Proper Name (91.4%) and Fictional Character (93.5%), the ones full of names its list doesn’t have. transformer-solve solves 98.4% and 98.7% of them.
  • transformer-solve uses more letters everywhere. It calls 2.4 to 4.7 more letters than smart-solve in every one of these categories.
  • In BONUS, transformer-solve leads on the wordplay categories: Before And After (77.1% against 70.1%), Same Name (76.8% against 68.1%) and Show Biz (80.6% against 66.1%). All three are well covered by the category lists from phase 4. smart-solve leads in most of the others, with its widest margin in Proper Name (71.1% against 60.2%).

How each solver gets it wrong

The evaluation files record each game’s letters, not the wrong answer itself. To see the answers, I replayed every wrong MAIN game through the C server. Every replayed game called the same letters in the same order as the evaluation.

smart-solve is wrong when a word is missing from its list and another word of the same shape fits what is showing:

Puzzlesmart-solve’s answerBoard when it solved
FILM NOIR FESTIVALFILM JOUR FESTIVALF_L_ _O_R FEST__AL
THE GORGONSTHE COUPONSTHE _O__ON_
CELEBRITY HEADSHOTCELEBRITY HELMSMAN_E_E_RI__ _E__S___
BEEF BOURGUIGNONBEEF APPROVINGLY_EEF ___R__I____
MEPHISTOLEBOWSKI_E___S__
BRIGITTE BARDOTBRIGITTE HARLOT_RI_ITTE _AR__T

It solves once every word has one candidate. When the true word isn’t a candidate, the one that is left looks just as certain. BEEF APPROVINGLY shows that it has no sense of which words go together. Some of its wrong answers are the source’s typos, not its own: the control puzzles include HOMEMADE THREE-CHEESE LASANGA and NATURALLY AGED CHEDDER CHEESE.

transformer-solve goes wrong differently. In 41 of its 42 wrong answers, it is off by one or two letters, often writing a word that doesn’t exist:

Puzzletransformer-solve’s answer
BACHELOR AUCTIONBACHELOR APCTION
ZOOKEEPERBOOKEEPER
WHEAT GERMWHEAT PERM
HIGHLY PROFITABLEHIGHLY PROMITABLE
INTERNATIONAL SYSTEM OF UNITSINTERNATIONAL SYSTEM OF KNITS
PINCHING YOUR CHEEKSPUNCHING YOUR CHEEKS
CYNDI LAUPERCONDI LAUPER
WATCHING GREMLINSWATCHING FREMZINS

The model writes the answer one character at a time. The masks keep each character legal for its square, but nothing checks that the word is real. With a few squares left blank, its most likely fill can be a non-word, at over 99% confidence. One of its “wrong” answers is a correction: for the control puzzle HOMEMADE THREE-CHEESE LASANGA it writes LASAGNA.

Learnings and conclusions

The data was most of the work

There are thirteen phases, and only one of them trains a model. The first eight are all data work. They scrape and clean the puzzles, split them, build the category lists and the corpus, and write 5.8 million labelled game lines. Training took 32 minutes on one GPU, and the model code is only 130 lines. The results depend much more on the data than on the model.

A small model can learn a game from a teacher

The 59M-parameter transformer started knowing nothing, not even English. After a billion characters of pretraining and half a billion of game lines, it solves 98.7% of puzzles it has never seen, against smart-solve’s 98.3%. It never made an illegal move in 3,349 games, though the masks did correct some characters along the way.

It learned the teacher’s moves, not a better strategy

Every MAIN training label is smart-solve’s pick, so the model learned to imitate it. A model trained to copy a strategy can approach that strategy. Getting past it needs a different signal, such as training on the outcome of whole games.

Knowing when to stop was the biggest lever

Accepting every SOLVE the model wrote solved 64% of dev puzzles. Requiring 99% confidence solved 98.3%, with no retraining at all. The model’s own probability for its answer turned out to be a usable measure of whether it was right. Tuning that one number on held-out puzzles mattered more than any training setting.

Rules belong in the decoder

A free-running model can call a letter twice or write an answer that doesn’t fit the board. Masking those characters out while it generates costs almost nothing and removes a whole class of failure. The model only has to learn which legal move is best.

The two solvers fail differently

smart-solve fails where its list has gaps. transformer-solve fails by misspelling. A word list and a language model cover each other’s gaps.

Held-out data has to be guarded everywhere

The control split is by answer text, so a puzzle filed under two categories can’t sit on both sides. The category lists drop every control answer before the model sees them: 680 phrases across Wikidata, Wiktionary and WordNet. Without those checks, the model would have trained on part of its own test.

Small and cheap is enough

A 207-line C forward pass with float16 weights serves the model on a 256 MB machine. A general-purpose model server was the wrong fit. An earlier version ran the model in llama.cpp’s llama-server on a 1 GB machine, and its default prompt cache pushed the weights out of memory.

Reproducible by construction

Every split is a hash of the answer text, and every random choice is seeded from the puzzle. Rerunning phases 2 through 8 from the published inputs gives byte-identical files. The one exception is GPU training, which is why the checkpoints are published.

Appendices

Appendix A: Training logs

Pretraining, from runs/pre/log.tsv:

StepTokenscorpus_val lossElapsed
004.0212 s
2,00066M1.19282 s
10,000328M1.023398 s
20,000655M0.962799 s
30,5001.0B0.9281,220 s

Finetuning, from runs/ft2b/log.tsv:

Stepmain_valbonus_valcorpus_valElapsed
02.7371.6420.9282 s
1,0000.2070.1591.01851 s
5,0000.1680.0961.025243 s
10,0000.1510.0830.993482 s
15,0000.1400.0850.971722 s

Appendix B: Results by category

The 16 largest categories in the control set. Letters is the mean number of letters called in solved MAIN games.

CategoryPuzzlesMAIN smart-solveLettersMAIN transformer-solveLettersBONUS smart-solveBONUS transformer-solve
Thing66298.9%8.1898.9%11.9484.3%78.9%
Food And Drink37497.9%8.6097.6%11.9989.8%84.0%
Phrase306100.0%10.0099.0%12.4278.1%78.1%
What Are You Doing272100.0%10.0197.4%12.7773.9%74.6%
Place18699.5%7.8599.5%11.3488.2%82.8%
Event17098.8%8.9899.4%12.0885.3%82.9%
Before And After144100.0%10.25100.0%13.2470.1%77.1%
Proper Name12891.4%9.1898.4%13.6071.1%60.2%
People12499.2%7.6897.6%11.4087.1%78.2%
Living Thing10198.0%8.4398.0%12.9780.2%74.3%
Person8895.5%6.95100.0%11.6186.4%84.1%
Fictional Character7793.5%9.7698.7%13.6172.7%67.5%
On The Map7197.2%8.3998.6%12.9985.9%74.6%
Same Name69100.0%10.20100.0%12.7168.1%76.8%
Around The House68100.0%9.13100.0%11.9982.4%73.5%
Show Biz6298.4%10.1598.4%12.7466.1%80.6%

Appendix C: What it cost

ItemCost
GPU: two Runpod sessions on an H100 SXM, about 88 minutes in all, at $3.49 an hourabout $5.15
Hosting: Fly.io shared-cpu-1x, 256 MB$2.19 a month if always on, less when suspended

Appendix D: Licenses of the data

SourceLicenseUsed for
wofanswers.comno license statedpuzzles.csv
SCOWL / Debian wamericansee data/wordlist/COPYRIGHTsmart-solve’s word list
WikidataCC0wikidata.tsv
Wiktionary, via kaikki.orgCC BY-SAwiktionary.tsv
Open English WordNet 2025CC BY 4.0wordnet.tsv
CMU Pronouncing DictionaryBSD-stylethe rhymes in wordplay.tsv
WikipediaCC BY-SA 4.0pretraining text
FineWeb-EduODC-By 1.0pretraining text