Computational · Paper · 2026

Editorial Quality at Scale: A Reproducible Prose-Review Pipeline

Daniel Ari Friedman

Zenodo

Download PDF Publication source Paper folder on GitHub Read extracted text

Overview

This paper documents template_prose_project, the prose-focused exemplar of the Research Project Template (https://github.com/docxology/template). It pairs the template's two-layer architecture with the prose analysis infrastructure (https://github.com/docxology/template/tree/main/infrastructure/prose) (readability metrics, structural outline, editorial quality flags) and the reference validation infrastructure (https://github.com/docxology/template/tree/main/infrastructure/reference) (BibTeX validation), demonstrating that rigorous editorial review can be expressed as a configurable, deterministic pipeline with no novel domain algorithm of its own.

A single manuscript/config.yaml defines target grade-level bands, citation-density floors, structural rules (every section has an H1, no heading levels skipped), and bibliography-consistency policy. The pipeline reads the manuscript, runs the prose analysers, cross-checks every citation against manuscript/references.bib, evaluates the configured checks, and writes a deterministic markdown review report alongside three figures (per-file word counts, readability metrics, citation density) and a JSON manuscript_report.json suitable for CI artefacts.

Run snapshot. The current configuration analyses 8 file(s) totalling 1742 words across 86 sentence(s) and 64 paragraph(s). Average Flesch-Kincaid grade level is 15.93; average Gunning Fog index is 16.69; the manuscript references 6 unique citation key(s); the longest section is 413 words and the shortest is 17. These numbers are auto-substituted by scripts/z_generate_manuscript_variables.py after every run, so the abstract tracks the JSON outputs in output/.

The contribution is methodological and architectural: a generic, reusable prose-quality module (infrastructure/prose/) that any project in the template can opt into, plus a minimal, configurable exemplar (projects/templates/template_prose_project/) that wires it to the bibliography and the manuscript pipeline.

Keywords: prose analysis, readability, editorial review, reproducible manuscript review, scientific infrastructure

---
Associated artifacts
GitHub release: v0.4.2 (https://github.com/docxology/template_prose_project/releases/tag/v0.4.2)
DOI: https://doi.org/10.5281/zenodo.20417104
Zenodo: https://zenodo.org/records/20417104
PDF SHA-256: 290d21b10bd588b978d6a3200cdf0e3c2441ca86fcdc777ab41975fa910a260e

prose analysisreadabilityeditorial reviewreproducible researchmanuscript quality

Overview source: Curated paper metadata.

Methods and contributions

Read the source for the full argument, qualifications, and evidence.

Findings and contributions

  • The paper reports that editorial review can be expressed as a configurable, deterministic pipeline with no novel domain algorithm of its own.
  • On the bundled manuscript, the run analysed 8 files totalling 1731 words, with average Flesch-Kincaid grade 15.87 and Gunning Fog 16.67.
  • Because no external service is consulted, a second run on the same inputs produces byte-identical JSON (modulo timestamp metadata).
  • The stated contribution is architectural: a reusable prose-quality module any template project can opt into, plus a minimal configurable exemplar.

Methods

  • Five-stage pure-function pipeline: read, analyse, cross-check, evaluate, render — template_prose_project implements editorial review as five pure-function stages in src/, with scripts limited to argument parsing and I/O.
  • Readability metrics: Flesch Reading Ease, Flesch-Kincaid grade, Gunning Fog — Per-file metrics are computed with textbook readability formulae over a vowel-group syllable heuristic.
  • Heuristic quality flags: passive voice, hedge density, citation density — A quality analyser flags passive-voice candidates, hedge words, Pandoc [@key] citation density, and long sentences.
  • Citation-key cross-check against references.bib with configurable policy — Every cited key is matched against the BibTeX file, with fail_on_missing / fail_on_unused settings in config.yaml.
  • Run-twice diff test for byte-identical JSON output — Reproducibility is checked locally by running the pipeline twice and diffing manuscript_report.json.

Summary sources: Paper metadata and evidence · Extracted source text.

PDF downloads

Archived files available directly from this site.

These files may be different versions or companion documents. The archive does not identify a latest edition.

Citation

Citation metadata follows the unified bibliography.

Friedman, Daniel Ari. 2026. Editorial Quality at Scale: A Reproducible Prose-Review Pipeline. Zenodo. DOI: 10.5281/zenodo.20417104. URL: https://doi.org/10.5281/zenodo.20417104.
Download bibliography

Catalog details and resources

Catalog row122
Citation keyFriedman2026EditorialQualityAtScale122
Platform availability

Related in Computational

Other catalogued works in the same domain.

View all Computational works, software & media →