
TL;DR
- Category creation for an AI start-up means getting language models to adopt your term for a new category, not only to mention your brand.
- There are two adoption pathways with very different speeds. Grounded, retrieval-based answers can pick up a new term in days. Trained model weights take a full training cycle.
- The fastest lever is consistency across authoritative surfaces. A Wikidata entry feeds the Google Knowledge Graph with no notability threshold, and models treat Wikipedia and Wikidata as primary verification sources.
- Explicit self-definition works. Stating “X is a [category] for [audience]” in structured, repeated language gives models a clean pattern to reuse.
- Measure adoption as the share of LLM responses that use your term over 60 days, tracked separately for grounded and ungrounded answers.
Key facts
- A Wikidata entity feeds directly into the Google Knowledge Graph and carries no notability requirement, unlike Wikipedia (entity-optimisation practitioners, 2026).
- Language models treat Wikipedia and Wikidata as primary verification sources when resolving what a term or brand means (entity-optimisation practitioners, 2026).
- Consistent naming across your site, LinkedIn, directories and press is required, since conflicting names are read as different entities (entity-optimisation practitioners, 2026).
- Around three or more independent, credible press mentions are a common threshold before an engine treats a brand or term as real (entity-optimisation practitioners, 2026).
- Structure and explicit attribution raise the odds a page is used to define an entity (Aggarwal et al., arXiv, 2023).
- Grounded, live-search answers reflect new web content far faster than a model’s trained weights, which update only on a training cycle (engine documentation, 2026).
Two things people mean by category creation
When an AI start-up says it wants to create a category, it usually means one of two different goals. The first is getting a model to associate the brand with an existing category, so an answer about “AI note-takers” names the company. The second, harder goal is getting the model to adopt a new term the company coined, so the model itself starts describing a class of products using that word. The first is entity association. The second is language adoption. They need different work and reward different patience.
This post is about the second. Language adoption is where founders overreach, because they assume that if they say a new term loudly enough on their own channels, models will pick it up. Models do not work that way. They adopt language that is consistent, well-attested and useful for describing something real, and they adopt it through two pathways that move at very different speeds.
The two adoption pathways
The first pathway is grounded retrieval. When ChatGPT, Perplexity or Gemini searches the live web to answer, it can encounter your term on the pages it retrieves and use it in that answer, sometimes within days of the term appearing on authoritative sources. This pathway is fast, controllable and measurable, but it only fires when the engine chooses to search and the term is present on pages it trusts.
The second pathway is training. For a term to appear in a model’s own weights, so it uses the word even without searching, the term has to be present, frequently and consistently, in the text the model is next trained on. That happens on the model maker’s schedule, not yours, and only if the term reached enough of the web to register. This pathway is slow and largely out of your hands, but it is the one that makes a category stick, because it means the model knows the word rather than reading it off a page.

Seed the sources models actually verify against
Because both pathways depend on the term appearing on trusted sources, the work is the same: get the term, defined consistently, onto the surfaces models use to verify meaning. The highest-leverage of these is structured reference data. A Wikidata entity feeds directly into the Google Knowledge Graph and, unlike Wikipedia, carries no notability requirement, so it is available early. Models treat Wikipedia and Wikidata as primary verification sources, so a clean, consistent entry there does disproportionate work in teaching a model what your term means.
Below that sit independent press, established review directories such as G2 and Crunchbase, and your own structured content. Practitioners commonly cite a threshold of around three or more credible, independent mentions before an engine treats a brand or term as real. The pattern that matters across all of them is consistency. If your site, your LinkedIn and a directory describe the category with three different phrasings, a model reads three weak signals instead of one strong one.

Define the term the way a model can reuse
Models reuse clean patterns. The most reliable phrasing is an explicit definition that names the category, the audience and the differentiation in one line, for example “Acme is a revenue-intelligence platform for B2B sales teams, distinct from generic CRM.” That sentence gives a model a ready-made template it can lift into an answer. Bury the same idea in vague brand copy and there is nothing clean to reuse. The original GEO research found that explicit, well-structured, attributed statements are more likely to be pulled into generated answers, and a category definition is exactly such a statement.
Repeat that definition, word for word where you can, across every surface. Repetition of a consistent string is what turns a phrase from your marketing into a pattern a model recognises. Novelty of phrasing, which human copywriters are trained to chase, actively works against machine adoption here.
Measure adoption as a rate, over time, split by pathway
Category creation is only a strategy if you can measure it. Define adoption as the share of model responses that use your term, and track it over 60 days. Crucially, measure grounded and ungrounded answers separately. Ask the question with web search on to read the grounded pathway, and with search off to probe trained weights. Early on you should see the grounded rate climb while the ungrounded rate stays flat, because the term is on the web but not yet in the model. Movement in the ungrounded rate, months later, is the signal the category has genuinely taken.
Run the probe across ChatGPT, Perplexity and Gemini, since they search and reflect new language at different speeds. A term that shows strong grounded adoption but no ungrounded adoption after several months is a term the web knows and the models have not yet learned, which tells you to keep seeding rather than declare victory.

When category creation is the wrong goal
For most early-stage AI start-ups, coining a brand-new category is the wrong first move, because it asks models to learn a word before they have a reason to. The faster win is entity association: get named inside the categories buyers already search, using the existing words for them. Category creation earns its place later, once the brand is well-attested and a genuinely new class of product needs a name that existing categories do not fit. Pursued too early, it burns effort teaching a vocabulary no one is searching for.
Frequently asked questions
Can a start-up really get an AI model to adopt a new term?
Partly, and through two pathways. Grounded answers that search the live web can use a new term within days if it appears on trusted pages, which is fast and controllable. Getting the term into a model’s trained weights, so it uses the word without searching, is slow and happens only on the model maker’s training cycle, and only if the term reached enough of the web. So near-term adoption is realistic through retrieval, while true trained adoption takes months and consistent seeding.
What is the single highest-leverage step?
A consistent Wikidata entry. It feeds directly into the Google Knowledge Graph, carries no notability requirement so it is available early, and models treat Wikidata and Wikipedia as primary verification sources. A clean entry that defines your term the same way you use it everywhere else teaches models what the word means faster than almost anything else. Pair it with consistent naming across your site, directories and press so the signals reinforce rather than conflict.
How should I phrase a category so models reuse it?
Use one explicit line that names the category, the audience and the differentiation, such as “X is a [category] for [audience], distinct from [adjacent category].” That gives a model a clean template to lift into an answer. Then repeat that exact phrasing across every surface. Consistency and repetition of a single string beat varied, clever phrasing, because models adopt patterns they see stated the same way in multiple trusted places.
How do I measure whether the term is being adopted?
Track the share of model responses that use your term over 60 days, and split grounded from ungrounded answers by asking with web search on and off. A rising grounded rate with a flat ungrounded rate means the web has the term but the model has not learned it yet. Later movement in the ungrounded rate is the signal the category is genuinely taking hold in the model’s own language, not just being read off a page.
How many mentions before a model treats my term as real?
There is no fixed number, but practitioners commonly point to around three or more credible, independent mentions as a working threshold before an engine treats a brand or term as real. Quality and consistency matter more than raw count. Three authoritative sources that define the term identically do more than a dozen inconsistent ones, because conflicting descriptions are read as evidence of different, weaker entities rather than one strong one.
Should an early-stage start-up prioritise category creation?
Usually not first. Coining a new category asks models to learn a word before buyers are searching for it, which is slow and expensive. The faster return is entity association: getting named inside the categories buyers already use, in the existing words. Reserve category creation for when the brand is well-attested and a genuinely new class of product needs a name that current categories do not fit. Pursued too early, it teaches a vocabulary no one is looking for.
Sources and references
- Entity optimisation: how to make LLMs recognise your brand. Pepper Content, 2026
- Entity recognition and knowledge graphs: structuring your brand for AI understanding. Discovered Labs, 2026
- How do LLMs choose which brands to mention in results. Page One Power, 2026
- GEO: Generative Engine Optimization. arXiv (Aggarwal et al.), 2023
- Wikidata as a source for the Google Knowledge Graph. Wikidata, 2026
- How AI engines source and verify information. Search Engine Land, 2026
See how AI engines currently describe your company before you push a new category.
Change log
- 2026-07-13: Initial publication.