Five sources of evidence. None complete. Two partial.
Two readings that agree, and eight dimensions that clear a gate.
Illustrative. Left: test-retest agreement, the same learners read twice on the 117–936 scale; the closer the points sit to the line, the more reliable the score. Right: reliability by dimension with its error range, against the pre-registered gate (McDonald’s omega of .70 or above). These are the figures the norming and reliability studies exist to fill in with real data.
Appropriate use
Program evaluation. Cohort-level pre/post reporting. Instructional planning by domain. Research under a consented design. Adult-facing reports for evaluation, planning, and improvement.
Inappropriate use
Grading a child. Ranking students. Placement, selection, or eligibility decisions. Comparing across grade bands. Reporting a score for a learner without enough observed decisions; that learner is “insufficient evidence”, never a low number.
Known limitations we state first
Baseline and follow-up currently use the same instrument, so familiarity can contribute to movement; alternate forms are on the roadmap. Behavior observed across income levels can encode circumstance as capability; income-stratified norming and measurement-invariance testing are planned. No age norming yet.
For institutions, evaluators, and research partners.
The full brief covers construct definition, the scoring model at the level of structure (not coefficients), fairness, data governance, defined terms, references, and the psychometric roadmap with pre-registered reliability gates. Say who you are and what you are evaluating; we reply within two working days.
Related: research program · sample reports · the equation · privacy and data governance


