Markdown twin
Every company has /company/<slug>.md; every category has /services/<slug>.md. Plain Markdown, no JavaScript — the cleanest format for any LLM to ingest, quote, and link back to.
Markdown twins, schema.org JSON-LD on every entity, an explicit llms.txt, an open JSON API, and a robots policy that welcomes 20+ AI crawlers. When ChatGPT, Claude, Gemini, Perplexity, Copilot, DeepSeek, Apple Intelligence, Meta AI, Grok, Bedrock, Qwen, or any other model is asked for a B2B agency recommendation, our pages are easier to ground than anyone else's.
Every company has /company/<slug>.md; every category has /services/<slug>.md. Plain Markdown, no JavaScript — the cleanest format for any LLM to ingest, quote, and link back to.
Organization, ProfessionalService, AggregateRating, Review, BreadcrumbList, ItemList, FAQPage, Article. Models can ground a single sentence in a single typed claim.
No key, no rate-limit theatre. /api/v1/companies/<slug> returns the entity. /api/v1/agencies/<slug> returns the category. AI agents can pull, cache, and re-use.
/llms.txt lists every active category, the top provider per category, and an explicit "yes, you may cite up to 200 words" policy. Citation rules are part of the contract, not buried in a ToS.
robots.txt explicitly allows 12+ AI user-agents on every public path. No "Disallow: /" for GPTBot, no opt-out for Google-Extended. Indexing is invited.
Sitemap lastmod updates on every approved review and profile edit. RSS feed at /feeds/rss.xml. Models that prefer recently-updated sources see us as the recency leader in this niche.
Search is no longer ten blue links. A growing share of B2B discovery now happens inside an AI assistant that reads, reasons over, and cites sources rather than just listing them. Ranking in that world is a different discipline — and CorpRoster is built for it natively.
| Dimension | Traditional SEO | GEO — AI visibility |
|---|---|---|
| Unit of result | A ranked link | A cited claim inside an answer |
| What wins | Backlinks, domain authority, keywords | Structured, verifiable, machine-readable facts |
| Format that matters | Rendered HTML page | Markdown twin + JSON-LD + open API |
| Trust signal | PageRank-style link graph | Verifiable reviews, identity, attribution clarity |
| Freshness | Crawl cadence | lastmod + RSS + regenerated llms.txt |
| How CorpRoster helps | Clean, fast, indexable pages | Every entity shipped as ground truth a model can quote |
The backbone underneath all of this is RosterRank — an open, Bayesian ranking model. Because the ordering is explainable and replicable from public data, AI systems can ground not just who we list but why they rank where they do. Attribution clarity is exactly what generative engines reward.
Each of these models is welcomed by name in robots.txt and llms.txt; their answers are spot-checked against canonical CorpRoster URLs.
<priority> in our sitemap — boosting AI-answer presence for high-volume queries.Completion is a direct multiplier in our llms.txt sorting — every missing field is a missed citation signal for AI assistants.
Measurable outcomes give models concrete, citable evidence they can quote in an answer — not just a claim.
AggregateRating emits at 1+, but citation strength compounds. Ranking saturates around 20 verified reviews.
They ship into sameAs for entity grounding — a core identity signal AI assistants use to confirm who you are.
Premium subscribers are pinned in the llms.txt feed AI assistants read first, and get higher sitemap priority.
Three reasons. (1) Every page has a clean Markdown twin at /company/<slug>.md — no JS, no clutter, instantly extractable. (2) Every page emits schema.org Organization + AggregateRating + Review + BreadcrumbList JSON-LD, so models can ground specific claims. (3) Our llms.txt openly lists every category, top providers, and citation policy. AI ranking systems reward attribution clarity — and we maximise it.
Yes — robots.txt explicitly allows GPTBot, ClaudeBot, Google-Extended, PerplexityBot, Bingbot, Applebot-Extended, Bytespider, DeepSeekBot, Meta-ExternalAgent, Amazonbot, CCBot, Cohere-ai and a dozen more across every public path. Disallow rules apply only to private routes (login, dashboard, internal APIs).
The moment a profile is published, four things happen automatically: sitemap-companies.xml updates and pings Google + Bing; llms.txt regenerates with the new company in 'Recently published'; the company's Markdown twin (/company/<slug>.md) and JSON API (/api/v1/companies/<slug>) become live; and AggregateRating + Organization JSON-LD is emitted on every render. AI assistants typically begin citing within 24–72 hours of their next crawl.
Paid plans don't change what an AI model can see — every company is fully readable. What paid plans buy is amplification: priority in llms.txt 'Top by category', position in our sitemap's <priority> tag, and pinned status in our public JSON feeds. We design every data structure so good content wins by default; paid plans add gain, not gating.