Skip to main content

Audit · Psychometrics

Diagnostic data inventory.

Phase 0 of the psychometrics roadmap: count the response data we have before we fit any model. See docs/PSYCHOMETRICS.md for the publication thresholds and what is deliberately not yet done.

Response matrix

Distinct learners
2
Distinct items
20
Total responses
20
Matrix sparsity
50.0%
Overall correctness
75.0%

Sparsity is 1 − total_responses / (learners × items). A sparsity of 1.0 means an empty matrix; higher numbers near 1.0 mean most learner-item cells have no response. IRT fits are stable on sparse matrices only when individual items have enough independent observations.

Items by sample-size band

Each band corresponds to a publication status in docs/PSYCHOMETRICS.md. An item moves from author-tagged to irt_fit only after it crosses the n ≥ 100 floor.

data_insufficient
n=129

20 items

Item has been served but not enough times to be a candidate for any kind of calibration. Stays author-tagged.

calibration_ready
n=3099

0 items

Enough attempts to flag for review. Not enough to publish 2PL discrimination estimates with confidence — treat as a candidate, not a calibrated item.

irt_fit_candidate_1pl
n=100499

0 items

Eligible for Rasch / 1PL fit with confidence interval. The publication floor for difficulty estimates.

irt_fit_candidate_2pl
n=500

0 items

Eligible for 2PL fit with both difficulty and discrimination, with confidence intervals.

Top topics by attempt count

Where the diagnostic is actually being served. Topics with high attempt counts but low item counts are candidates for expanded item banks.

TopicItemsAttempts
conjugate-gradient-methods33
basic-logic-and-proof-techniques33
epsilon-nets-and-covering-numbers22
vectors-matrices-and-linear-maps22
attention-mechanism-theory11
asymptotic-statistics11
proximal-gradient-methods11
adam-optimizer11
feedforward-networks-and-backpropagation11
differentiation-in-rn11
cramer-rao-bound11
activation-functions11
reinforcement-learning-from-human-feedback-deep-dive11
vc-dimension11

Recommendations

  • No items have crossed the publication threshold (n ≥ 100). The largest item has 1 attempts. Stay at Phase 0; do not fit IRT yet.

Source rows: 20 AssessmentAttempt, 146 LearningEvent, 11 DiagnosticAttempt runs. Last write: Sun, 28 Jun 2026 06:41:16 GMT.

Read on