About this tool
A free keyword clustering tool that groups terms by meaning and never sees your data. Here is what it does, how it does it, and where the limits are.
Why it exists
This started as a research tool, not a product. The underlying work — building graphs out of text, embedding short strings into vectors, and finding communities in the result — is the same set of methods used across a lot of computational research. Keyword lists turn out to be an almost perfect fit for it: thousands of short phrases where the interesting structure is semantic rather than lexical.
Existing clustering tools mostly ran server-side and charged per keyword, which made experimenting with the method expensive. Building a local version was cheaper than paying for one. Once it worked, publishing it cost nothing extra — the analysis runs on your machine, so a thousand users cost the same to serve as one.
The method, in full
Nothing here is a black box, so it is worth stating plainly what happens between your CSV and the map:
-
01
Column detection
The file is scanned to find the keyword column and, if present, a search-volume column. Detection scores actual cell content rather than trusting headers, so exports from Search Console, Ahrefs, Semrush and Keyword Planner all work without configuration. Comma, semicolon and tab separators are all handled.
-
02
Embedding
Each keyword is passed through
all-MiniLM-L6-v2, a compact sentence-transformer model, producing a 384-dimensional vector per keyword. That model is English-trained, so a non-English list groups better with the multilingual option in the toolbar — it is offered rather than default because the weights are 113 MB against 22 MB. Phrases that mean similar things land near each other in that space even when they share no words — which is the whole point, and the thing that word-matching approaches cannot do. -
03
Clustering
Tight seed clusters are formed from high-similarity pairs, near-duplicate seeds are merged, and every remaining keyword is assigned to its nearest cluster above a similarity floor. The result is that every keyword lands somewhere, rather than a large "unclustered" bucket that you then have to sort by hand.
-
04
Labelling and intent
Each cluster is named from its most distinctive terms, and classified as informational, commercial or transactional based on the phrasing patterns in its keywords.
-
05
The graph
Clusters become nodes, sized by keyword count or search volume; edges connect clusters whose centroids are close. Layout is force-directed, so related topics physically settle near each other and you can see the shape of a content area rather than reading it off a spreadsheet.
Two workspaces
The pipeline above runs on whatever text a file gives it, so the same engine covers two quite different jobs. Which one you get is decided by the file you drop in, not by a setting you have to find.
-
01
Keywords
Any export with a keyword column produces the topic map and the content plan. If the export also carries a position and a landing page — a Search Console performance export does, and so do Ahrefs and Semrush organic exports — two more views appear: the queries already sitting on page two, and the queries where more than one of your own URLs is taking impressions.
-
02
Pages
A crawl export from Screaming Frog is read one row per page, and the model embeds page titles instead of keywords. You get the topic map of the site you actually have, every on-page and technical problem the crawl file can prove, and pairs of pages that cover the same topic without linking to each other. Add the All Inlinks export as a second file and the pairs that already link are filtered out.
The boundary is worth stating plainly: this tool never fetches your website. A browser cannot read a response from another domain, and a server that could would be the end of the guarantee everything else here rests on. You run the crawler, and its output is read on your device exactly like a keyword file.
A cluster table tells you what groups exist. The graph tells you which groups are adjacent — which is where pillar pages, internal links and content gaps actually come from. Two clusters sitting close together with nothing between them is usually a page you have not written yet.
Privacy is architectural, not a policy
Most tools promise not to misuse your uploads. This one cannot misuse them, because it never receives them. The model is downloaded once and cached by your browser; after that, your CSV is read, embedded and clustered entirely on your own device. There is no account, no database and no upload endpoint.
That distinction matters most for agencies and in-house teams: a keyword list is a map of commercial strategy, and running it through a third-party server is often the step that requires a data-processing agreement. Here there is nothing to sign, because there is no processing to agree to. You can confirm it yourself — open your browser's network tab while an analysis runs, or disconnect from the internet after the model has loaded and watch it keep working.
What it deliberately does not do
Rank tracking, backlink indexes, PageSpeed scores and SERP scraping are all absent, and will stay absent. Every one of them needs either a server that fetches the web on your behalf or a paid third-party data provider, which would break the one guarantee this tool is built on. They are also what every competing product already sells, which makes them the worst possible place to compete from.
The crawl analysis is the one case worth distinguishing carefully. This tool does not audit your site — it reads an audit you already ran. Screaming Frog does the crawling; its export is then read here like any other file. That is the difference between a tool that needs access to your website and one that needs a file you already have.
The honest limitation that follows: this tool plans content, it does not measure rankings. It answers "how many pages should exist and what should each cover", not "where do I rank today". Use it alongside Search Console, not instead of it.
How it stays free
There is no paid tier, no advertising, no affiliate links and no data resale. Files up to 5,000 keywords work with no account at all. Larger sets are also free — email for an unlock code and you get one back, usually the same day.
The reason for the email step is not monetisation, it is feedback. A short note about what you are working on is far more useful for deciding what to build next than any analytics dashboard, and it is the only user research a project like this gets.
Who builds it
A one-person project, built and maintained by a researcher working with graph analysis and embedding models. There is no company behind it, no investors and no growth target — which is precisely why it can afford to have no paid tier and to refuse the features that would require collecting your data.
Related open work: Awesome Keyword Clustering, a curated list of clustering tools, libraries, embedding models and papers — including the ones that compete with this site.
Questions, bug reports and feature requests all go to info@graphmykeywords.com.
Try it on your own list
Drop a keyword CSV in and see the clusters. Free, no signup, and the file never leaves your browser.
Open the tool