The sector's problem is the distance between knowing and doing.
Fernandes, Lynch & Netemeyer (2014), across 201 studies: financial-literacy interventions explained about 0.1% of the variance in financial behavior, with effects decaying to negligible at 20 months or more.
Kaiser, Lusardi, Menkhoff & Urban (2022), across 76 randomized experiments in 33 countries: positive causal effects on knowledge and behavior, economically meaningful, three to five times larger than earlier reviews once study quality and publication bias are handled.
Bruhn et al. (2016), the largest youth RCT (868 Brazilian schools, ~20,000 students): knowledge rose a quarter of a standard deviation, saving and planning improved, and use of expensive credit and late repayment also rose. Teaching about a product can increase its use.
Read together: financial education moves behavior, modestly; the effects decay; and nearly every study measures behavior by self-report after the fact. An instrument that observes decisions directly, before and after instruction and again at 3, 6 and 12 months, sits in exactly that gap. That is the FLIQ thesis. It is a thesis until the studies below make it a finding.
Ten instruments whose problem resembles ours.
| Instrument | What it measures | How it earned its standing | What FLIQ borrows |
|---|---|---|---|
| PISA Financial Literacy (OECD) | 15-year-olds' financial literacy from performance tasks; ~100,000 students in 20 economies (2022) | IRT scaling on a 0–1,000 metric; empirically set proficiency levels; trend items across cycles | Framework discipline (content × process × context), IRT scaling, empirically defined bands |
| CFPB Financial Well-Being Scale | Adult financial well-being; 10 items, one latent variable | IRT development; common 0–100 metric across short and long forms; national norms | A common metric across forms; the norming report as a public artifact |
| Balloon Analogue Risk Task (clinical psychology) | Risk-taking propensity from a behavioral task | Task performance predicts self-reported real-world risk behavior in adults and adolescents, beyond demographics (Lejuez et al., 2002, 2003) | The validation paradigm: a simulated behavior earns standing by predicting real behavior |
| Delay-discounting tasks (behavioral economics) | Preference for smaller-sooner over larger-later rewards | Present bias predicts credit-card debt; discount factors predict creditworthiness (Meier & Sprenger, 2010, 2012) | A laboratory choice measure that predicts financial outcomes; a convergent measure for Saving and Borrowing |
| Stealth assessment / evidence-centered design (game-based measurement) | Competencies inferred from gameplay telemetry | Mislevy's competency–evidence–task models with Bayesian scoring; Shute's program showed game-based assessment can be valid and reliable | The scoring architecture: FLIQ's domain × trait grid and confidence term are an ECD structure |
| SELweb (school-based social-emotional assessment) | Children's social-emotional comprehension by direct assessment | Browser-delivered; N ≈ 4,400; internal consistency .79–.88; temporal stability .55–.79; confirmed factor structure | The closest procedural sibling and a template validation sequence |
| NIH Toolbox (cognition) | Executive function and other domains, ages 3–85 | Population norming (N ≈ 3,800); published scoring-algorithm comparisons | Norming plan structure; age-band handling |
| PROMIS (clinical outcomes) | Patient-reported outcomes across domains | IRT-calibrated item banks; interchangeable short forms; adaptive testing | Calibrate once, deliver many forms: how the same-instrument confound is removed |
| Lexile Framework (reading) | Reader ability and text complexity on one scale | Rasch-based scale licensed across publishers and assessments | One scale, many licensees, reported as a unit (1050L; 687 FLIQ) |
| FICO Score (consumer credit) | Likelihood of credit default | Predictive validity against realized default; continuous recalibration | Validate against later real outcomes. FLIQ differs in purpose: it is engineered to move with learning and is never used for selection |
Deliberately absent: self-report character scales, because self-report is the weakness FLIQ exists to avoid; and the Iowa Gambling Task, a strong clinical analogue whose scoring reliability in children is a caution rather than a model.
Eight practices, in the order they matter.
Competency model (eight domains × four traits), evidence model (which observables bear on which competencies, with weights), task model (the scenarios). The one document that lets a measurement group evaluate the instrument without reading code.
PROMIS's logic. Removes the baseline/follow-up familiarity confound we currently disclose.
400 to 600 consented learners by age band and household income, as CFPB, NIH Toolbox and SELweb did.
Omega ≥ .70; test-retest ≥ .65 at two to four weeks; standard errors by band.
DIF across income, gender, race and ethnicity, language. The specific threat: an instrument observing money behavior across incomes can encode circumstance as capability (Watts, Duncan & Quan, 2018).
Does a decision in simulation predict a decision in an applied setting, and persist at 3, 6, 12 months? This is the instrument's central claim.
Fernandes et al. found effects gone by 20 months; Kaiser et al. found them durable in better studies. A matched 3/6/12-month schedule is how FLIQ contributes to that argument.
What Works Clearinghouse design standards; ESSA evidence tiers; CONSORT-style attrition flow; pre-registration on OSF.
Ten phases, each supplying what the next requires.
Sample sizes are planning figures; a partner's power analysis governs. Phase 0 is the recommended first collaboration: it needs no data agreement and produces a paper.
| Phase | Study | Design and output | Planning n |
|---|---|---|---|
| 0 | Sector meta-analysis | Systematic review of K–12 financial education restricted to behavioral outcomes; PRISMA; effect-size and decay estimates. A desk study: no data agreement, a publishable paper, and the priors for every later power calculation. | — |
| 1 | ECD documentation and expert review | Competency, evidence and task models written; blind review of scenario–domain mapping. | 5–7 reviewers |
| 2 | Cognitive interviews | Think-aloud across age bands during The $100 Week™. | 24–36 children |
| 3 | Norming and internal structure | Stratified consented sample; CFA; general-factor test; IRT calibration. | 400–600 |
| 4 | Reliability and alternate forms | Test-retest; parallel forms from the calibrated bank. | 150–250 |
| 5 | Measurement invariance | DIF and invariance on the norming sample and first cohorts. | from 3 |
| 6 | Convergent and discriminant validity | Concurrent delay-discounting, risk, executive-function and knowledge measures. | 200–300 |
| 7 | Program effect with comparison | Cluster-randomized or matched comparison across classrooms; alternate forms; fidelity measured. | 20–30 classrooms |
| 8 | Transfer and persistence | Learners with an applied or real-world node followed at 3, 6, 12 months. | from 7 |
| 9 | Consequential validity | How reports are used by educators, districts and families. | 15–25 sites |
- Kaiser, Lusardi, Menkhoff & Urban (2022). Financial education affects financial knowledge and downstream behaviors. Journal of Financial Economics, 145(2). www.sciencedirect.com/science/article/abs/pii/S0304405X21004281
- Fernandes, Lynch & Netemeyer (2014). Financial literacy, financial education, and downstream financial behaviors. Management Science, 60(8). ideas.repec.org/a/inm/ormnsc/v60y2014i8p1861-1883.html
- Bruhn et al. (2016). The impact of high school financial education: Evidence from a large-scale evaluation in Brazil. World Bank. documents1.worldbank.org/curated/en/753501468015879809/pdf/WPS6723.pdf
- OECD (2024). PISA 2022 Results, Volume IV: How Financially Smart Are Students? www.oecd.org/en/publications/pisa-2022-results-volume-iv_5a849c2a-en.html
- CFPB (2017). Financial Well-Being Scale: Scale development technical report. www.consumerfinance.gov/data-research/research-reports/financial-well-being-technical-report/
- Lejuez et al. (2003). Evaluation of the BART as a predictor of adolescent real-world risk-taking behaviours. Journal of Adolescence. pubmed.ncbi.nlm.nih.gov/12887935/
- Meier & Sprenger (2012). Time discounting predicts creditworthiness. Psychological Science. www.researchgate.net/publication/51867630_Time_Discounting_Predicts_Creditworthiness
- Shute, V. J. Stealth assessment primer; lessons learned and best practices. myweb.fsu.edu/vshute/pdf/SA_Primer.pdf
- McKown et al. (2016). Web-based assessment of children's social-emotional comprehension (SELweb). xsel-labs.com/wp-content/uploads/2023/03/McKown-et-al.-2016-SELweb-FINAL.pdf
- Fries et al. (2014). Item response theory, computerized adaptive testing, and PROMIS. Journal of Rheumatology, 41(1). www.jrheum.org/content/41/1/153
- NIH Toolbox executive function tasks in a U.S. norming sample (2024). pmc.ncbi.nlm.nih.gov/articles/PMC11841212/
- AERA, APA & NCME (2014). Standards for Educational and Psychological Testing. www.testingstandards.net/
Companion documents: the FLIQ Validation Brief v2.0 and the technical specification of the scoring model (under agreement). To take on a study, start at the Research Program.

