Skip to main content
Every published number names its task, data snapshot, baseline, controls, sample size, and uncertainty. A model card is not a substitute for an evaluation.

The beta standard

Results report n, intervals clustered by person, and cost per thousand scored people. Stated answers and revealed actions are separate task families. Novel-domain transfer is reported separately from in-domain fit.

Public record

Our own evaluation code lives in the HFM research repository. External suites are tracked with their version, command, date, arm, metric, baseline, and full result, including runs we do not feature. The public record makes emphasis distinguishable from existence. The current research record includes next-item transfer on Douban, Steam, Amazon, and MovieLens, Semantic ID studies, and survey probes. Those are research results, not guarantees for every client deployment. A client model card carries its own data snapshot, temporal split, controls, and measured result. See the API model cards and the research repository.