Overview
Extracted from the local paper documentation when available.
This paper formulates on-policy distillation as active inference in finite variational models, with exact claims only for declared objects and interpretive claims explicitly bounded outside them. In the construction, the intractable teacher policy plays the role of the generative model $p(o,s)$, the tractable student policy is the approximate posterior $q(s)$, and the per-token reverse-KL...
Use Notes
Concise findings and methods pulled from README/SKILL documentation.
Citation
Plain-text citation for quick reuse.
Primary source Documentation Full Text Image Gallery Source repository BibTeX
Related in Active Inference
Other catalogued works in the same domain.