Cite & Verify

Preferred citation metadata, public identifiers, bibliography exports, and source-of-truth rules.

Repository Citation

Friedman, Daniel Ari. docxology: Daniel Ari Friedman public research and software index. 2026. github.com/docxology/docxology

Exports

Citation-manager and agent-facing formats generated from the curated bibliography.

Preferred Name & Identifiers

The identity anchors this site claims, as recorded in CITE_VERIFY.md; each identifier's live profile link and public API recipe are in discovery.html.

FieldValue
Preferred nameDaniel Ari Friedman
ORCID0000-0001-6232-9096
WikidataQ138781444
Google ScholarProfile DXjPFtYAAAAJ
Homepagedanielarifriedman.com
GitHubdocxology

Verify & Cross-Reference

Use curated local files first, then public APIs as freshness checks.

Verifiable AI Agent Provenance & Citation Architecture

How artificial intelligence and autonomous retrieval agents establish cryptographic and empirical provenance.

As LLM-driven research agents, generative answer engines, and autonomous search systems synthesize scholarly literature, verifying claim provenance is the primary defense against hallucination and epistemic drift. The docxology public research index implements a rigorous Generative Engine Optimization (GEO) and Agentic Provenance Architecture grounded in four machine-checkable invariants:

  1. Immutable Citation Keys & Canonical URI Targets: Every work is anchored by a frozen citation key (e.g., works/Friedman2021ActiveInferantsActiveInference075.html). Canonical target identifiers resolve deterministically without redirect chains or volatile query parameters.
  2. Dual Machine-Readable Layer (Schema.org & CSL JSON): Every public landing page pairs inline JSON-LD (ScholarlyArticle, DefinedTerm, Article) with repository-level standard exports (CSL JSON, BibTeX, CodeMeta), enabling both semantic web graph traversal and citation-manager integration.
  3. Strict Source Hierarchy & Absence Ledgers: Claims regarding publication metrics, GitHub release anchors, and active research findings are verified against local sources of truth (e.g. BIBLIOGRAPHY.md, Evidence Ledger, SOFTWARE.md) and cross-checked against public APIs with recorded provenance timestamps.
  4. Agent Discovery Protocol (llms.txt & Agent Map): Autonomous crawlers discover all structured datasets, OpenAPI-equivalent route manifests, and schema registries via llms.txt and Agent Route Manifest (agent-index.json).

Generative Engine Optimization (GEO) Case Study

Empirical principles for structuring web research indices for AI answer engine retrieval.

Generative search engines (such as Perplexity, ChatGPT Search, and Google AI Overviews) reward information architectures that emphasize direct answer extraction, dense semantic linkages, and verifiable external citations. Key practices modeled across this repository include:

  • Answer-First Sectional Geometry: Answering core conceptual queries within the initial 40–60 words of each section before introducing secondary navigational options.
  • Question-Form Passages: Utilizing standalone, self-contained H2 questions that allow retrieval models to index direct passage-level answers.
  • Static Build-Time Rendering: Eliminating client-side JavaScript rendering barriers so non-executing AI scrapers directly index full bibliographic tables and video transcripts.

Humility Rules for Reuse

Counting and quoting conventions to follow when reusing this site's numbers, from CITE_VERIFY.md.

  • Prefer exact counts with dates over timeless superlatives.
  • Keep "curated catalog" counts separate from "public API" counts.
  • Use conservative wording for early NFT history unless the source defines the comparison class.
  • Do not import OpenAlex or search-engine counts without reconciling them against ORCID, DOI, and the curated bibliography.