Skip to main content
A physicist working through equations on a chalkboard, graded in navy and gold
FLIQ Score™
Research Foundations · v1.0, September 2026

What the FLIQ Score™ is built on, and what it learns from.

Four questions a research office asks before a first call: what construct is being measured and what is its sector's evidence baseline; which instruments in other fields measure something analogous and how they were validated; which of those methods carry over to the FLIQ Score Model (FSM); and what the study program is from here. Every citation below is to a primary source.

The construct and its baseline

The sector's problem is the distance between knowing and doing.

0.1%

Fernandes, Lynch & Netemeyer (2014), across 201 studies: financial-literacy interventions explained about 0.1% of the variance in financial behavior, with effects decaying to negligible at 20 months or more.

76 RCTs

Kaiser, Lusardi, Menkhoff & Urban (2022), across 76 randomized experiments in 33 countries: positive causal effects on knowledge and behavior, economically meaningful, three to five times larger than earlier reviews once study quality and publication bias are handled.

+0.25 SD, and more credit

Bruhn et al. (2016), the largest youth RCT (868 Brazilian schools, ~20,000 students): knowledge rose a quarter of a standard deviation, saving and planning improved, and use of expensive credit and late repayment also rose. Teaching about a product can increase its use.

Read together: financial education moves behavior, modestly; the effects decay; and nearly every study measures behavior by self-report after the fact. An instrument that observes decisions directly, before and after instruction and again at 3, 6 and 12 months, sits in exactly that gap. That is the FLIQ thesis. It is a thesis until the studies below make it a finding.

Analogues in other sectors

Ten instruments whose problem resembles ours.

InstrumentWhat it measuresHow it earned its standingWhat FLIQ borrows
PISA Financial Literacy (OECD)15-year-olds' financial literacy from performance tasks; ~100,000 students in 20 economies (2022)IRT scaling on a 0–1,000 metric; empirically set proficiency levels; trend items across cyclesFramework discipline (content × process × context), IRT scaling, empirically defined bands
CFPB Financial Well-Being ScaleAdult financial well-being; 10 items, one latent variableIRT development; common 0–100 metric across short and long forms; national normsA common metric across forms; the norming report as a public artifact
Balloon Analogue Risk Task (clinical psychology)Risk-taking propensity from a behavioral taskTask performance predicts self-reported real-world risk behavior in adults and adolescents, beyond demographics (Lejuez et al., 2002, 2003)The validation paradigm: a simulated behavior earns standing by predicting real behavior
Delay-discounting tasks (behavioral economics)Preference for smaller-sooner over larger-later rewardsPresent bias predicts credit-card debt; discount factors predict creditworthiness (Meier & Sprenger, 2010, 2012)A laboratory choice measure that predicts financial outcomes; a convergent measure for Saving and Borrowing
Stealth assessment / evidence-centered design (game-based measurement)Competencies inferred from gameplay telemetryMislevy's competency–evidence–task models with Bayesian scoring; Shute's program showed game-based assessment can be valid and reliableThe scoring architecture: FLIQ's domain × trait grid and confidence term are an ECD structure
SELweb (school-based social-emotional assessment)Children's social-emotional comprehension by direct assessmentBrowser-delivered; N ≈ 4,400; internal consistency .79–.88; temporal stability .55–.79; confirmed factor structureThe closest procedural sibling and a template validation sequence
NIH Toolbox (cognition)Executive function and other domains, ages 3–85Population norming (N ≈ 3,800); published scoring-algorithm comparisonsNorming plan structure; age-band handling
PROMIS (clinical outcomes)Patient-reported outcomes across domainsIRT-calibrated item banks; interchangeable short forms; adaptive testingCalibrate once, deliver many forms: how the same-instrument confound is removed
Lexile Framework (reading)Reader ability and text complexity on one scaleRasch-based scale licensed across publishers and assessmentsOne scale, many licensees, reported as a unit (1050L; 687 FLIQ)
FICO Score (consumer credit)Likelihood of credit defaultPredictive validity against realized default; continuous recalibrationValidate against later real outcomes. FLIQ differs in purpose: it is engineered to move with learning and is never used for selection

Deliberately absent: self-report character scales, because self-report is the weakness FLIQ exists to avoid; and the Iowa Gambling Task, a strong clinical analogue whose scoring reliability in children is a caution rather than a model.

Methods carried over

Eight practices, in the order they matter.

01
Document FLIQ as an evidence-centered design

Competency model (eight domains × four traits), evidence model (which observables bear on which competencies, with weights), task model (the scenarios). The one document that lets a measurement group evaluate the instrument without reading code.

02
Calibrate an item bank, then build alternate forms

PROMIS's logic. Removes the baseline/follow-up familiarity confound we currently disclose.

03
Norm on a stratified sample and publish the report

400 to 600 consented learners by age band and household income, as CFPB, NIH Toolbox and SELweb did.

04
Pre-register reliability gates

Omega ≥ .70; test-retest ≥ .65 at two to four weeks; standard errors by band.

05
Test measurement invariance early

DIF across income, gender, race and ethnicity, language. The specific threat: an instrument observing money behavior across incomes can encode circumstance as capability (Watts, Duncan & Quan, 2018).

06
Validate against real behavior, the BART way

Does a decision in simulation predict a decision in an applied setting, and persist at 3, 6, 12 months? This is the instrument's central claim.

07
Measure decay on purpose

Fernandes et al. found effects gone by 20 months; Kaiser et al. found them durable in better studies. A matched 3/6/12-month schedule is how FLIQ contributes to that argument.

08
Report to the standards funders use

What Works Clearinghouse design standards; ESSA evidence tiers; CONSORT-style attrition flow; pre-registration on OSF.

Study program

Ten phases, each supplying what the next requires.

Sample sizes are planning figures; a partner's power analysis governs. Phase 0 is the recommended first collaboration: it needs no data agreement and produces a paper.

PhaseStudyDesign and outputPlanning n
0Sector meta-analysisSystematic review of K–12 financial education restricted to behavioral outcomes; PRISMA; effect-size and decay estimates. A desk study: no data agreement, a publishable paper, and the priors for every later power calculation.
1ECD documentation and expert reviewCompetency, evidence and task models written; blind review of scenario–domain mapping.5–7 reviewers
2Cognitive interviewsThink-aloud across age bands during The $100 Week™.24–36 children
3Norming and internal structureStratified consented sample; CFA; general-factor test; IRT calibration.400–600
4Reliability and alternate formsTest-retest; parallel forms from the calibrated bank.150–250
5Measurement invarianceDIF and invariance on the norming sample and first cohorts.from 3
6Convergent and discriminant validityConcurrent delay-discounting, risk, executive-function and knowledge measures.200–300
7Program effect with comparisonCluster-randomized or matched comparison across classrooms; alternate forms; fidelity measured.20–30 classrooms
8Transfer and persistenceLearners with an applied or real-world node followed at 3, 6, 12 months.from 7
9Consequential validityHow reports are used by educators, districts and families.15–25 sites
References
  1. Kaiser, Lusardi, Menkhoff & Urban (2022). Financial education affects financial knowledge and downstream behaviors. Journal of Financial Economics, 145(2). www.sciencedirect.com/science/article/abs/pii/S0304405X21004281
  2. Fernandes, Lynch & Netemeyer (2014). Financial literacy, financial education, and downstream financial behaviors. Management Science, 60(8). ideas.repec.org/a/inm/ormnsc/v60y2014i8p1861-1883.html
  3. Bruhn et al. (2016). The impact of high school financial education: Evidence from a large-scale evaluation in Brazil. World Bank. documents1.worldbank.org/curated/en/753501468015879809/pdf/WPS6723.pdf
  4. OECD (2024). PISA 2022 Results, Volume IV: How Financially Smart Are Students? www.oecd.org/en/publications/pisa-2022-results-volume-iv_5a849c2a-en.html
  5. CFPB (2017). Financial Well-Being Scale: Scale development technical report. www.consumerfinance.gov/data-research/research-reports/financial-well-being-technical-report/
  6. Lejuez et al. (2003). Evaluation of the BART as a predictor of adolescent real-world risk-taking behaviours. Journal of Adolescence. pubmed.ncbi.nlm.nih.gov/12887935/
  7. Meier & Sprenger (2012). Time discounting predicts creditworthiness. Psychological Science. www.researchgate.net/publication/51867630_Time_Discounting_Predicts_Creditworthiness
  8. Shute, V. J. Stealth assessment primer; lessons learned and best practices. myweb.fsu.edu/vshute/pdf/SA_Primer.pdf
  9. McKown et al. (2016). Web-based assessment of children's social-emotional comprehension (SELweb). xsel-labs.com/wp-content/uploads/2023/03/McKown-et-al.-2016-SELweb-FINAL.pdf
  10. Fries et al. (2014). Item response theory, computerized adaptive testing, and PROMIS. Journal of Rheumatology, 41(1). www.jrheum.org/content/41/1/153
  11. NIH Toolbox executive function tasks in a U.S. norming sample (2024). pmc.ncbi.nlm.nih.gov/articles/PMC11841212/
  12. AERA, APA & NCME (2014). Standards for Educational and Psychological Testing. www.testingstandards.net/

Companion documents: the FLIQ Validation Brief v2.0 and the technical specification of the scoring model (under agreement). To take on a study, start at the Research Program.