Methods and provenance Health systems science suite
Back

Where these numbers come from

Written for students and facilitators, and for anyone who wants to check a figure before repeating it. Every claim in the three sessions should be traceable from here.

The patients

Every patient is simulated

No individual in this suite is a real person, and no real record was copied, perturbed or resampled to produce one. Names, identifiers and dates are generated. Identifiers carry a SIM- prefix that could not be mistaken for a medical record number.

What is real is the shape of the population. The simulated cohort was built to reproduce the distributions and relationships of a de-identified primary care population of roughly 1.1 million patients, held in a secure data environment. That is why the patterns are worth taking seriously and the individuals are not.

How they were made

The real population was summarised into cells defined by campus region, age band, sex, comorbidity band, housing status, social vulnerability quartile and population density quartile. Cells containing fewer than 10 real patients were dropped entirely and the remaining proportions renormalised.

Synthetic patients were then allocated to surviving cells in proportion to the real counts, and every remaining attribute — exact age, race, individual conditions, encounter counts — was drawn stochastically within the cell.

Why this matters. Because cells are allocated and details are drawn, rather than records being copied or perturbed, no synthetic patient derives from any single real patient. Rare combinations cannot be reproduced, because the cells holding them were removed before anything was generated.

54,000 patients were generated, 6,000 per campus region. Generation is seeded, so the cohort is reproducible.

How a panel is chosen

The group code seeds a shuffle of the nine campus regions, which are then dealt to students in turn. Each student's region is shuffled once and split into disjoint halves, so two students sharing a region cannot share a patient. Non-overlap is arithmetic, not chance.

A panel depends only on the group code and the student number, never on how many students are enrolled — so joining late does not move anybody's patients, and the same code opens the same panel in Session 1 and Session 3.

Privacy

Real data left the secure environmentNo
Individual real records used to build a patientNo
Cells below 10 real patientsDropped before generation
Age released asDecade, 90+ aggregated
Groups below 11 patients in the appsSuppressed

Age is released in decades rather than single years. Screening eligibility could not be recovered from a decade — colorectal, breast and cervical criteria all straddle a decade boundary — so eligibility is carried as a separate flag computed before age was coarsened.

Money

Cost of care

Cost is built from encounter volumes priced at 2026 Medicare national rates: a blended office or outpatient encounter, an emergency visit including facility and professional components, and a flat amount per inpatient admission rather than a per-diem. Length of stay is not used, because it could not be derived reliably.

Encounter counting cannot see pharmacy, imaging, specialty or procedures, so the total is scaled by a multiplier anchored to published national per-person health expenditure. That multiplier sets the dollar labels on the axis and changes nothing else: capitation is budget-neutral, so margin scales proportionally and every relationship in the session is unaffected by it.

Capitation and risk adjustment

Capitation is risk-adjusted and budget-neutral: risk scores average exactly 1.0 and total payments equal total cost across the cohort, so margin sums to zero and every winner implies a loser.

The risk model uses age, sex and comorbidity. It deliberately excludes housing instability, social vulnerability and population density. That exclusion is the point of Session 3 — the gap it produces is something the data demonstrates rather than something a facilitator asserts.

Payer multiples

Payment by payer type is expressed as a multiple of the cost of care, taken from the Session 2 finance workbook, which was built from Indiana commercial-to- Medicare price ratios and 2026 plan parameters.

PayerPaysBasis
Commercial1.82× costIndiana facility prices run about 2.5× Medicare; professional about 1.3×
Medicare1.04× costClose to cost, and the benchmark others are quoted against
Medicaid0.73× costAbout 65% of Medicare rates
Uninsured0.15× costAfter self-pay discount and expected collection

Payer is not recorded in the cohort. The payer mix tool applies a single blended rate to every patient, which assumes payer does not depend on how sick someone is. It does in reality; without a payer variable there is no honest way to do better, and inventing one would manufacture the thing that is missing.

What "margin" means here

⚠ Not profit

Every margin figure in these sessions is contribution over the direct cost of care — what payers hand over, minus what the clinical care cost. It carries none of the overhead a health system runs on: buildings, administration, information systems, capital, teaching, and uncompensated care outside the panel. Those absorb most of it.

Real integrated health systems report operating margins in the low single digits as a share of revenue. Treat these figures as relative signals, never as money anyone banks.

Where a margin appears without a payment model named, it means the risk-adjusted capitated model.

Quality and infection

Screening

Screening rates are computed over eligible patients only, using age and sex criteria for each measure — colorectal 45–75, breast 40–74 female, cervical 21–65 female, preventive exam all adults. A rate over everybody would be meaningless.

Screening capture in the source data is uneven. One campus region records almost no screening at all — around 4–5% of the rate elsewhere, across four independent measures simultaneously. No clinical mechanism does that to four unrelated measures at once. It is a data capture failure, it is preserved rather than hidden, and it is one of the things the sessions are built to teach.

A second capture effect runs throughout: patients seen more often have more opportunities for a screening code to be recorded. Patients with housing instability appear better screened here, and they also average roughly twice the outpatient contact. A quality measure built from records partly measures contact frequency rather than care.

Composite quality score and the ranking game

The composite combines screening performance with absence of healthcare-associated infection, weighted 70/30. It is ours, not a Medicare score, and should not be quoted as one. Because it leans on screening data, it inherits the capture problem above.

The 42 comparison clinics in the ranking view are generated deterministically from the group code, and their spread is calibrated to the observed between-region variation in this cohort rather than to an arbitrary curve.

The penalty magnitude follows the Medicare Hospital Readmissions Reduction Program, whose maximum penalty is 3% of base operating payments. The bonus follows the Merit-based Incentive Payment System, which is budget neutral and has been small in practice. Those magnitudes are real; the clinics are not.

Infection

Surgical site infection, catheter-associated urinary infection, central line infection and C. difficile are modelled conditional on inpatient admission. A 3,000-patient panel holds roughly 28 infections of any type and about one central line infection. Numbers that small cannot support panel-level comparison, which is why infection is judged at clinic level.

Rules the apps enforce

Groups under 11 patientsSuppressed, and marked as withheld
Cost and margin figuresMedian and share, never an average
Screening ratesEligible denominators only
Housing instabilityClinic and cohort level only
RaceCohort level only
Unverified variable pairingsRefused, not drawn

Why medians. Cost in this cohort reaches several hundred thousand dollars for a single patient, and the most extreme margins belong to patients with no social risk flags at all. An average over a 3,000-patient panel can be set by one person, so the tools do not offer one.

Why some questions are refused. Only relationships checked against the cohort when it was built can be plotted. A chart of noise still looks like a finding, and students believe charts.

Known limitations

LimitationWhat it means for the session
Payer is not in the data The payer mix tool is a scenario layer over a blended rate, not a measurement. It cannot show how payer and illness interact.
Length of stay was dropped Bed-days cannot be reported. Admissions are counted instead.
One region's screening is missing Preserved deliberately and flagged in the interface, but it makes that region's quality composite uninterpretable as performance.
Two regions record no housing instability A zero there means the field was never completed, not that nobody is unstably housed. The interface says so rather than showing a count.
Two race groups are very small Native American and Asian patients number 29 and 137 in the whole cohort. Race is analysable only at cohort level, and even there descriptively.
Screening gradients are weak Only cervical screening shows a consistent social gradient. The sessions teach the measurement problem rather than claiming a clean disparity.
Associations, not causes Nothing here supports a causal claim, including the relationships that look strongest.

Sources

Comorbidity indexAHRQ Elixhauser, 38 conditions
Social vulnerabilityCDC/ATSDR Social Vulnerability Index
Unit prices2026 Medicare national rates
Payer multiplesSession 2 finance workbook
Penalty mechanicHospital Readmissions Reduction Program
Bonus mechanicMerit-based Incentive Payment System
Cost anchorPublished national per-person health expenditure

Screening benchmarks shown alongside observed rates are cited in the facilitator guide. Where a figure in these sessions has no source listed, it is derived from the cohort itself and can be reproduced from the generation scripts.