foundations-measurement-theory
Installation
SKILL.md
Measurement Theory Foundations
Establish whether an observation supports its intended interpretation and decision. A repeatable score can consistently measure the wrong thing.
When to use
Use when a task asks whether a metric, questionnaire, benchmark, sensor, rating, or composite score measures its intended target, or whether scores can be compared across groups, instruments, or time.
Do not use when the task is only sampling uncertainty, causal identification, choosing an action, or implementing an evaluation pipeline. Those belong respectively to statistical inference, causal inference, decision theory, and ai-evals.
Workflow
- Define the intended interpretation, population, decision, and cost of measurement error. Identify what is observed and what remains latent.
- Choose the physical-measurement or psychometric branch in measurement primitives. Do not translate psychometric reliability into metrological traceability, or treat benchmark scores as physical quantities without justification.
- Audit the instrument with the eight primitives below. Use evidence actually available; mark missing evidence rather than supplying thresholds or validity claims.
- For score comparisons, read comparability and drift. Preserve instrument versions, scoring changes, administration conditions, and population differences.
- Complete the measurement audit template, using the synthetic example only as a format example.