Methodology

How Word Sieve searches its word list

Word Sieve searches one word list, an edition of ENABLE1, and every answer depends on which words that edition contains and on the letters, pattern or clues you enter. This page explains how the list became the index your browser loads, what each search does, and where the results stop. The guides use figures computed from the same file.

Word listENABLE1 · public domain Words172,781 Indexed168,509 · 2–15 letters Anagram keys152,185 Index shards14 · 3,203,788 bytes

What is measured

What is measured. A match means an entry in this site's word list fits the letters, pattern or clues you entered. The solvers search the indexed length range and report a total even when only part of the result is shown on screen. The word maker orders matches by the standard letter-value table printed on its page. The clue filter ranks candidates by letter frequency at each position among the surviving words, counting a repeated letter only once in each candidate's rank score. The rankings set the order only: whether a word is in the result depends on what you entered.

How it is kept. The list is served as plain text, with a provenance note recording its source, hashes and publication exclusions. Before publication, an authoring tool groups words by their sorted letters and writes an index file for each word length. A search loads the length files it needs. Repository checks regenerate the index and compare it with the shipped files byte for byte. The site is static, so no index is generated at deployment.

What it cannot tell you. A match cannot tell you whether your game's dictionary accepts the word, or which move to play on its board. The solvers find only words in this edition, and only in the 2–15-letter index: the other 4,272 entries in the full file are sixteen letters or longer.

The word list

ENABLE1, the Enhanced North American Benchmark Lexicon, was researched and compiled by M. Leo Cooper and Alan Beale and released into the public domain. Word Sieve uses an ENABLE1-derived edition: upstream ENABLE1 with 39 entries excluded before publication, which leaves 172,781 entries. They are written in lowercase a–z only, with no capitals, hyphens, apostrophes or diacritics.

You can read the exact edition in wordlist/enable1.txt. The source download, both SHA-256 hashes, the exclusions, the appended newline and a historical comparison of mirrors are described in wordlist/PROVENANCE.md. That note also quotes the public-domain statements that are the basis for redistribution, and a copy of the README is served beside the list.

The official tournament lists are separate proprietary compilations: TWL and NWL in North America, and Collins elsewhere. They are not used here, and this edition is the only word list the site ships. A result shows what this edition contains; a tournament or your game's own judge decides what it accepts.

Word Sieve is an independent tool with no affiliation to, endorsement from, or connection with any word game or its publisher. Game names on this site only identify the game a page discusses.

The index

Sorting a word's letters gives it the same key as every anagram made from those letters. For example, aelpp is the key for appel, apple and pepla. An exact one-word anagram search looks up that key once. A search for shorter words within a rack uses several keys.

index = { sorted(word) → [every word with those letters] }

Why it is split by length

The index holds 152,185 keys in 3,203,788 bytes across fourteen files, shard-02.txt through shard-15.txt. Each file contains words of one length. In the unscrambler, a five-letter rack uses shards 2 through 5, about 127 KB, and a seven-letter rack uses 2 through 7, about 645 KB. Compression can cut the transfer to an estimated third of these figures. The byte count under the results is the loaded shard text before compression, and it can include data that an earlier search on the same page already loaded.

Each page stays under 100 KB uncompressed, together with the stylesheets and scripts it loads. That limit leaves out the index files a search fetches, which are separate downloads, so a solve uses more data than the page weight.

Three search paths, and two of them are scans

For a rack search without blanks, the solver enumerates distinct combinations of letters and looks each one up in the index. A seven-letter rack has at most 120 combinations of two letters or more; repeated letters can reduce that number. So a seven-letter rack costs at most 120 map lookups, and the rack is never compared with all 172,781 entries. Exact anagrams need one lookup, of the whole rack's key, and two-word anagrams look up keys for the rack's possible splits.

A blank can supply any letter. Generating substituted keys would mean 26 substitutions per combination for one blank, or 351 for two, multiplying the work on a long rack. Instead, the blank path scans the keys in the loaded shards. It counts the letters each key needs beyond those in your rack and keeps the key if the blanks can cover that deficit. The scan is bounded by the loaded keys, with 152,185 in the full index.

Crossword patterns, board clues and browse filters use a third path: a scan over words. A sorted-letter key has lost the positions needed to answer these questions. Each of these searches loads the shard for one word length and checks its words against the constraints. The page then reports how many words it searched: every word of that length.

The repository checks compare the two rack-search paths on blank-free racks. They also compare the shipped rack solver with an independent brute-force scan of all 172,781 entries on a set of probe racks, and confirm that the comparison catches a planted difference. These checks run on selected inputs.

Search limits

The edition is fixed and does not follow current usage

ENABLE1 was compiled decades ago, with editorial decisions about what counts as a word. It is not a source for proper names, punctuated spellings or many recent coinages. If an entry looks unfamiliar or a modern word is missing, check another dictionary.

Your game's dictionary makes the ruling

A game may reject a word found here or accept one missing from this edition. Check that game's own dictionary or rules when acceptance matters. Dictionary disagreement alone does not mean the solver searched its list incorrectly.

The solvers do not see your board

The solvers use the letters, patterns and clues you enter. They do not read your game board or account for open scoring squares, other played tiles or an opponent's rack. The word maker adds up fixed letter values and does not calculate the value of a move.

Rack searches stop at fifteen characters

The longest indexed word is fifteen letters. A rack of sixteen characters or more is refused with a message, and nothing is trimmed. Some searches have a tighter cap, as shown on their pages.

Rack searches accept up to two blanks

On rack pages that accept blanks, use ? for each blank tile. The limit is two, separate from the fifteen-character rack limit. A rack with a third blank is refused. Exact-anagram and jumble searches do not accept blanks.

Long results stop at 1,500 on screen

The display limit is 1,500 results. The unscrambler shows the longest matches first, with alphabetical order within each length. The word maker shows the highest letter-value totals first, and the clue solver uses its frequency ranking. A capped list reports how many matches were found and how many are not displayed. The words left off are the last in that page's order: the shortest on the unscrambler, the lowest totals on the word maker and the lowest-ranked on the clue solver.

Questions

Why does the site download a dictionary to my browser?
So your letters can stay in the browser. They are not sent to a server for solving or any other purpose: the solver fetches only this site's index files, chosen by word length, and those requests carry none of your letters. The site's Content-Security-Policy allows connections only to this origin. The cost is the index download, split by length so a solve loads only the files it needs. The size the page reports is the loaded text, before compression.
How is the index built, and can I check it?
Yes. The list and the shard files are plain text you can read, and the provenance note identifies the source edition and describes the checks. The authoring tool reads wordlist/enable1.txt, sorts each indexed word's letters into a key, and groups the words that share a key. It writes one file per word length, with keys and words in a fixed order, and the repository checks confirm that a fresh run reproduces the shipped files exactly. A separate check tests the real rack solver on probe racks against an independent brute-force scan of the whole list.
Why are 4,272 words in the file but not in the index?
They are 16 letters or longer, past this site's 15-character rack limit and its 2–15-letter index. The file still contains them: of its 172,781 entries, 168,509 are indexed. The solvers skip the longer words, which stay readable in wordlist/enable1.txt.
Does the site keep anything on my device?
No, nothing you enter. Word Sieve sets no cookie, writes nothing to local storage, session storage or IndexedDB, and registers no service worker, so no rack history, result, setting or preference is kept between page loads. During a solve, the rack exists only in the text box and page memory. Your browser can cache the site's public page and dictionary files under its HTTP caching headers; those files contain no entered racks.
What would make an answer here wrong?
A dictionary difference or a bug. A match can be right for this list and still be refused by your game: ENABLE1 was compiled decades ago, and Word Sieve's edition, with 39 entries excluded, is neither a game's official dictionary nor updated for current usage. For bugs, the repository checks compare generated files and test solver results against independent scans, on selected inputs. If a result looks wrong, check the letters or clues you entered and the dictionary your game uses.

Files and tools