No black boxes. Below: the peer-reviewed paper, the official guidelines, and the primary docs behind every pillar — plus an honest label for what's measured, what's documented, and what's our judgment.
Princeton researchers coined “Generative Engine Optimization” and benchmarked 9 methods on Perplexity.ai and GPT-4 with search across 10,000 queries. Statistics, quotations, and cited sources won decisively (up to ~40% gains); style tweaks did almost nothing. Our Citability pillar and weighting philosophy come straight from these results.
arxiv.org/abs/2311.09735Statistics Addition, Quotation Addition and Cite Sources were the top-performing GEO methods — up to ~30–40% visibility lift. Fluency/style rewrites barely moved results.
GEO: Generative Engine Optimization — Aggarwal et al., Princeton, Nov 2023Retrieval systems prefer self-contained, directly-answering chunks; the GEO paper's best methods all reward directly quotable blocks. Our 25% weight reflects this combined evidence.
GEO paper (ibid.) + RAG retrieval literatureGoogle's 170-page rater guidelines define Experience, Expertise, Authoritativeness, Trust — authorship, dates, sourcing. AI Overviews/Gemini inherit these signals.
Google Search Quality Rater GuidelinesBot operators publish exactly what they need: allowed user-agents (GPTBot, PerplexityBot, ClaudeBot), renderable HTML, fast responses. Observable requirements, not theory.
OpenAI GPTBot docs · PerplexityBot & ClaudeBot crawling policies · robots.txt conventionsschema.org vocabulary + Google's structured-data docs define machine-readable facts; citation engines observably extract schema-marked content cleanly (esp. FAQPage).
schema.org · Google Search Central structured data docsRecency signals matter to RAG retrievers and training-cutoff-sensitive models; third-party mentions (reviews, forums) dominate LLM training corpora. Smallest weight = weakest direct evidence.
RAG recency literature + corpus composition studies