Overview
This exemplar demonstrates a data descriptor workflow in which the schema, file inventory, provenance chain, license boundary, and validation gate are treated as first-class research artifacts rather than afterthoughts. It ships a small, public, synthetic demonstration dataset (two CSV files under data/fixtures/) and a machine-readable descriptor (data/example_descriptor.json) that declares each file's media type, sha256 checksum, and row count alongside a six-field data dictionary with typed constraints. A tested validation library (src/data_descriptor/) checks the descriptor's shape, safety, and completeness; recomputes each declared checksum and row count against the bytes on disk; and emits a deterministic, metadata-only release manifest suitable for pre-publication review. Every figure and quantitative claim in this manuscript is produced by that library and regenerated on demand, so the prose describes structure and provenance rather than transcribing values that would drift. This is a template with a demonstration dataset: it makes no scientific claim about the data, only about how to describe and release a dataset responsibly.
Overview source: Curated paper metadata.
Methods and contributions
Read the source for the full argument, qualifications, and evidence.
Findings and contributions
- The clean fixture descriptor produces zero validation findings, while the deliberately perturbed demo produces several errors and warnings.
- For the shipped fixture, both files verify: declared and actual row counts agree and each recomputed checksum matches, leaving the readiness score unpenalised.
- The paper reports that its zero-mock test suite exceeds the 90% project coverage gate.
- The work explicitly makes no scientific claim about the synthetic data; claims are limited to how to describe and release a dataset.
Methods
- Synthetic demo dataset: two CSV fixtures plus JSON descriptor — Ships two small deterministic synthetic CSV files (measurements, subjects) described by a machine-readable descriptor with checksums and row counts.
- Six-field data dictionary with typed constraints — Declares six fields with types, nullability, units, regex patterns, closed enumerations and numeric bounds for the measurement table.
- Order-independent sha256 schema fingerprint — descriptor_fingerprint() hashes (name, type, nullable) triples so reordering fields does not change the schema fingerprint.
- Validation gate with readiness score and perturbed negative control — validate_descriptor() emits severity-tagged findings folded into a readiness score; a deliberately broken copy tests that the gate reacts.
- Byte-level descriptor-to-file verification — verify_descriptor_files() recomputes sha256 and row counts of present files and reports verified, mismatch, or absent status.
Summary sources: Paper metadata and evidence · Extracted source text.
PDF downloads
Archived files available directly from this site.
- Friedman_2026_Data_a542f49d.pdf PDF · 0.19 MiB · extracted-text source
Citation
Citation metadata follows the unified bibliography.
Catalog details and resources
Related in Computational
Other catalogued works in the same domain.