← All posts

hourlong learnings #7: SJT exams like Casper and PREview and cohort studies on their efficacy

In the following post I try to understand whether the Casper test is accurate, what the AAMC PREview exam is actually asking, how situational judgement tests (SJTs) influence medical-school admissions, and whether a number from either test means much about a future doctor

Casper and PREview are not miniature residency interviews and they are not personality scans. They are standardized tests of how closely an applicant’s answers match a professionally defined response to short, constrained situations. That can be useful information. It is much less than a direct measurement of empathy, ethics, bedside manner, or future clinical quality.

as usual, a basic thought map is below:

$$ \text{short scenario} \rightarrow \text{applicant judgement under time pressure} \rightarrow \text{standardized score} \rightarrow \text{admissions weighting at one school} \rightarrow \text{interview or acceptance decision} $$

The score is not the decision. The same Casper quartile or PREview 1–9 score can be a screen, a tiebreaker, a contextual data point, or nearly irrelevant depending on the school. That local weighting is usually the part applicants cannot see.


1. a situational judgement test asks for the “best next move”

An SJT presents a social, ethical, or professional situation with incomplete information, then asks the test taker to judge responses. Its basic premise is reasonable: academic metrics are not designed to capture how someone notices a conflict, handles a boundary, communicates uncertainty, or escalates a safety concern.

The important word is judgement. An SJT generally measures knowledge of what a professional response should look like in a stylized situation. It does not watch someone behave longitudinally with patients, teammates, sleep deprivation, hierarchy, money, or consequences. A person can recognize the right response and fail to enact it; another can be awkward in a timed assessment and excellent in practice.


2. casper and preview are related, but they are not interchangeable

FeatureCasperAAMC PREview
Response formatopen response: typed and videorate the effectiveness of proposed actions
Current standard format11 sections: 4 video-response and 7 typed-response scenariosselected-response professional-behavior ratings
Time pressure1 minute per video question; 3.5 minutes for each typed two-question set75-minute exam in the original operational format; details can change by testing year
What applicants receivetyped-section quartilescaled score from 1 to 9, percentile rank, and confidence band
Scoring logictrained raters score responses relative to scenario-specific standardscredit for alignment with a medical-educator consensus key; partial credit for a nearby rating on the same side of the scale
What a high score most directly supportsstronger performance on this open-response SJTstronger recognition of the consensus rating of professional behaviors

The format difference matters. Casper asks an applicant to generate language and, in part of the current format, speak on camera. That bundles situational judgement with reading speed, idea generation, typing speed, spoken fluency, comfort with recording, and the ability to organize an answer quickly. PREview instead supplies actions and asks whether each is very ineffective, ineffective, effective, or very effective. It reduces the generation burden but turns the task into calibrated agreement with an expert key.

Neither version is automatically more “real.” They sample different parts of a response

notice the issue
→ identify competing obligations
→ choose a proportionate action
→ explain it clearly
→ carry it out in a real setting

3. what a score actually contains

PREview is unusually explicit about this. Its score reflects how closely an applicant’s ratings align with consensus ratings created by medical-school representatives. Exact agreement receives full credit. A response one category away can receive half credit if it remains on the same effective/ineffective side. A higher score therefore means more agreement with that consensus—not a direct count of compassionate actions or ethical decisions.

Casper’s public result is even thinner. Applicants receive a typed-response quartile, meaning a relative position in that cycle’s comparison group, while programs receive more detailed score information. A fourth quartile is not 100% correct, and a first quartile is not a diagnosis of poor character. It is a rank category after a particular scoring and norming process.

This makes a common interpretation error easy to see:

$$ \text{reported score} = \text{measured judgement signal} + \text{test-format effects} + \text{random measurement error} $$

The inputs are the intended judgement signal, format-specific performance, and error; the output is the number the school receives. The equation is a causal accounting model, not a way to calculate anyone’s true professionalism. It says a low score can come from more than one source, and a high score is not pure access to a hidden trait.


4. reliability is not the same thing as accuracy

People often say an admissions test is “accurate” when they mean one of several different things:

QuestionTechnical nameWhat a good result would mean
similar scoring procedures give similar score?reliabilitythe number is not mostly scorer noise
does it add information after GPA, MCAT, experiences, interview?incremental impotanceit improves a decision model, not just correlates with existing filters
predict an outcome validityit is associated with later performance in a relevant setting
comparable across groups?fairness / measurement invariancescores and consequences do not create avoidable subgroup harm

Casper has published generalizability and inter-rater reliability estimates in a respectable range, including reports around .72–.83 for overall test reliability and .82–.95 for inter-rater reliability. That is evidence that its scoring operation can be consistent. It is not evidence that a one-quartile difference identifies who will be a better physician.


5. the numbers for preview are small, and that is informative

the largest early multi-school PREview study used 19,525 exam scores from 18,549 unique examinees in the 2020 and 2021 pilot administrations. PREview correlated $r=.29$ with MCAT total score and $r=.16$ with undergraduate GPA. Corrected correlations with interview ratings ranged from $r=.09$ to $.14$ across participating schools and interview formats.

Those are small positive correlations. They support the narrow claim that PREview is not identical to GPA, MCAT, or interview ratings. They do not mean it is unrelated to academic-selection processes, and they do not establish that a 7 will become a better doctor than a 5.

The official study’s correlation table is more useful than a marketing adjective because it shows the scale of the relationship. see published table here.


6. squared correlations help keep the magnitude honest

For a rough intuition, $r^2$ is the fraction of variation in one measured variable that is linearly shared with another in that sample. With $r=.29$ between PREview and MCAT, $r^2\approx .084$: about 8.4% shared variation. With $r=.14$ for one interview comparison, $r^2\approx .020$: about 2.0%.

This does not say PREview explains 8.4% of a student’s future quality, because MCAT is not future physician quality and $r^2$ is not a causal share. It only prevents a small correlation from being narrated as a large overlap.

Small effects can still matter in a large applicant pool. If 5,000 qualified people compete for 150 seats, a weak but nonredundant signal can move an applicant across an interview threshold. That is a decision effect, not proof of a large human difference.

small score difference
× unknown local weight
× crowded selection threshold
→ potentially large admission consequence

This is why applicants experience these exams as high-stakes even when the psychometric effect size is modest.


7. prediction gets distorted after a school selects people

Admissions research has a built-in geometry problem: schools only observe later outcomes for people they admitted. Those students already passed filters on GPA, MCAT, activities, essays, recommendations, geography, mission fit, and often interview performance. Their score range is compressed.

all applicants
→ several admissions filters
→ matriculants with a narrowed score range
→ later outcome study

This range restriction can make correlations inside a medical-school class look smaller than they would in the applicant pool. But it also means a correlation observed among already-selected students cannot simply be used as a causal validation of the original filter. The selection rule helped create the sample.

The better study design would prospectively retain all applicant data, pre-specify outcomes, correct transparently for selection where appropriate, and report uncertainty. Most importantly, it would test whether adding an SJT improves decisions compared with a realistic alternative—not merely whether a score has a statistically significant $p$ value.


8. subgroup differences are important

Fairness is not solved by calling a test noncognitive. In the PREview pilot study, mean scores were 5.0 for White examinees, 4.2 for Black or African American examinees, and 4.6 for Hispanic/Latino examinees; the White–Black standardized difference was $d=.43$ and the White–Hispanic difference was $d=.24$. The authors described overlap between distributions and noted that these gaps were smaller than many traditional academic measures. Both facts can be true.

The study’s reported group means and standardized differences show overlap and a nonzero average gap. The table and full methods are here.

For Casper, a single-school U.S. cohort of 1,237 interviewed applicants reported significantly lower scores among Black, Native American/Alaska Native, and Hispanic applicants than other applicants. The study’s figure compares the cohort’s Casper percentiles across gender, socioeconomic indicator, and UIM status.

Existing PLOS One figure comparing Casper percentile score patterns by demographic group in one U.S. medical-school applicant cohort

9. how schools can use the same score very differently

A school can require an SJT and use it in at least four ways:

UseOperational meaningRisk
threshold screenapplicants below a local cutoff are less likely to advancea noisy low score can exclude a strong applicant before richer review
holistic contextreviewers see it beside essays, experiences, academics, and mission fitvague weighting can become unaccountable weighting
interview triageit helps decide who receives scarce interview timethe test may reproduce what the interview could have measured directly
post-interview tiebreakerit is one small factor among close candidatesapplicants cannot estimate its effect from the public requirement alone

Programs have incentives to use scalable metrics: a 75-minute standardized measure can be available before interviews, while interviewing every applicant is impossible. That is an operations argument, not a validity argument. A tool can solve a queueing problem and still require careful validation.

One single-school study placed Casper beside MMI and traditional-interview scores rather than assuming that all three measured the same thing. Its comparison table is useful because it reports the actual outcome variables and subgroup patterns instead of calling the measure simply holistic.

Existing PLOS One table comparing Casper percentile scores with multiple-mini-interview and traditional-interview scores in a U.S. medical-school applicant cohort

The AAMC’s own admissions resources emphasize local validation and whether PREview adds value in a school’s process, which I think is smart


10. students practicing exam tactics doesn’t mean the score is any less meaningful

Applicants naturally prepare. They learn the format, practice reading scenarios under time, and become less likely to freeze. That can reduce irrelevant variance from surprise and interface unfamiliarity.

But intensive coaching can also move the measured construct. If an applicant learns a template that produces consensus-sounding answers without improved judgement, then score gains are mostly just test-taking gains. The same issue exists with every high-stakes assessment: preparation can teach the target skill, the test’s language, or can be a mix both (likely true here)

The evidence question is empirical: do preparation effects improve later relevant performance, merely score performance, or both?


11. the missing data are more important than another correlation

If meds schools want these tests to mean more than an opaque hurdle, each should publish a short validation report each cycle or every few cycles:

applicant pool and who completed the test
→ score distribution and subgroup analyses
→ how the score entered the admissions workflow
→ interview, acceptance, progression, and professionalism outcomes
→ incremental value beyond existing materials
→ uncertainty, limitations, and any changes to weighting

The best outcome is not necessarily one universal “professionalism” score. It may be a transparent, low-weight signal that helps structure review, paired with human evaluation of the actual applicant’s experiences and reasoning. some schools a school may find it adds no useful information when multiple interviews are already required and stop considering it.


12. the main conclusion

Casper is a structured, time-limited SJT with evidence of reliable scoring and some evidence of association with later professional-assessment outcomes. PREview is not an empathy exam either; it measures agreement with medical educators about the effectiveness of responses to professional situations. Both can contribute a small but still relevant nonacademic signal to the admissions process

Early PREview correlations with MCAT, GPA, and interview ratings were small. Some SJT work finds modest predictive value for outcomes such as later disciplinary action, but those outcomes are rare and not significant enough to support any real conclusions. Group score differences and the opacity of school-specific weighting mean that any consequential use needs local validation, reporting, and a way to detect harm.

So do the tests mean anything? Yes, but their meaning is narrower than student admissions culture makes them feel. They mean something about performance on a particular standardized judgement task. They may help a school sort a very large pool. They do not mean a 4th-quartile applicant is inherently more compassionate, or that a 1st-quartile applicant cannot become a great physician.

TLDR. My personal learning: I started this assuming Casper and PREview were trying to measure a hidden quality called “being a good doctor.” From what I understand, they are much more ordinary measurement tools: they sample how an applicant recognizes professionally preferred responses under a specific format and then produce a comparably scored number. That number can be reliable and still only modestly predictive; it can add a real signal and still be too weak, unequal, or opaque to deserve heavy admissions weight. The better question is not whether an SJT is perfectly accurate. It is whether a specific school can show that using it, at a specific weight, improves decisions without creating costs it has not measured.

Comments are closed for this post.