Skip to main content
A scientist at a microscope, graded in navy and gold
FLIQ Score™
Validation · FLIQ Validation Brief v2.0

Where the FLIQ Score™ stands on validity.

The FLIQ Score™ is an emerging measure. Its scoring method, the FLIQ Score Model (FSM), is designed to observe financial decision-making across eight dimensions; it has not yet been validated in the psychometric sense. This page states the current status against the five sources of validity evidence in the AERA/APA/NCME Standards, and the limits that follow. The full brief is available on request.

Current status

Five sources of evidence. None complete. Two partial.

Test content

The eight domains mirror the Financial Foundations™ curriculum one-to-one and the 2021 national financial-literacy standards. Item-level content review by external educators has not been run.

Partial

Response processes

Scores are read from recorded decisions and stated reasons inside structured scenarios rather than recalled answers. Think-aloud or cognitive-interview studies have not been run.

Partial

Internal structure

The scoring model is specified and reproducible. Reliability (internal consistency, test-retest) and factor structure have not been estimated; a norming sample is required.

Not established

Relations to other variables

No convergent, discriminant, or predictive evidence yet. The first cohorts will be the first data.

Not established

Consequences of use

Guardrails are in force: no grade, no ranking, no placement or selection use, no cross-grade comparison. Evidence of consequences in practice does not yet exist.

Partial
What the studies will produce

Two readings that agree, and eight dimensions that clear a gate.

Illustrative: a scatter of first readings against second readings along the agreement line, and eight dimension reliability bars against a pre-registered gate

Illustrative. Left: test-retest agreement, the same learners read twice on the 117–936 scale; the closer the points sit to the line, the more reliable the score. Right: reliability by dimension with its error range, against the pre-registered gate (McDonald’s omega of .70 or above). These are the figures the norming and reliability studies exist to fill in with real data.

Appropriate use

Program evaluation. Cohort-level pre/post reporting. Instructional planning by domain. Research under a consented design. Adult-facing reports for evaluation, planning, and improvement.

Inappropriate use

Grading a child. Ranking students. Placement, selection, or eligibility decisions. Comparing across grade bands. Reporting a score for a learner without enough observed decisions; that learner is “insufficient evidence”, never a low number.

Known limitations we state first

Baseline and follow-up currently use the same instrument, so familiarity can contribute to movement; alternate forms are on the roadmap. Behavior observed across income levels can encode circumstance as capability; income-stratified norming and measurement-invariance testing are planned. No age norming yet.

Request the full brief

For institutions, evaluators, and research partners.

The full brief covers construct definition, the scoring model at the level of structure (not coefficients), fairness, data governance, defined terms, references, and the psychometric roadmap with pre-registered reliability gates. Say who you are and what you are evaluating; we reply within two working days.

Related: research program · sample reports · the equation · privacy and data governance

We respond within 1–2 business days  ·  Your information is never shared