Skip to main content
Q6

How Well Did Our Assumptions Match Reality?

We built this study first on synthetic data calibrated to published NCES/IPEDS statistics, then re-ran it on real College Scorecard data. Here's what our assumptions got right — and what they got wrong.

Why This Matters

Every Model Starts with Assumptions

We generated 4,344 synthetic cases before we had real data. The synthetic data was calibrated to published NCES/IPEDS (federal postsecondary data) aggregate statistics — sector distributions, enrollment ranges, closure rates by institution type. But calibration to aggregates doesn't guarantee the distributions match at the institution level.

The retained public-data cohort contains 1,516 institutions. This check is itself an ADM demonstration: validation determines whether the lower-fidelity synthetic model was sufficient and which earlier claims must be withdrawn.

Summary Statistics

Synthetic vs. Real: The Numbers

High-level comparison of the two datasets. These summary statistics tell us whether our synthetic generation process produced a dataset of similar scale and closure prevalence.

Total Institutions
4,344
Synthetic
1,516
Real
Total Closures
244
Synthetic
6
Real
Overall Closure Rate
5.6%
Synthetic
0.4%
Real
Sector Distribution

Did We Get the Mix Right?

The most basic structural check: does the synthetic dataset have roughly the same proportion of institutions in each sector (Public 4-year, Private Nonprofit 4-year, For-Profit, etc.) as the real data? Getting this wrong would cascade into every downstream analysis.

Institution Count by Sector: Synthetic vs. Real
Closure Rates by Sector

Where Did Closures Actually Happen?

Sector-level closure rates are the core signal in this study. If the synthetic data nailed the overall rate but assigned closures to the wrong sectors, every investigation built on that data would inherit the bias.

Closure Rate by Sector: Synthetic vs. Real
Closure Rates by Region

Did Geography Hold Up?

Regional patterns in college closure are driven by demographics, state funding, and local economic conditions. Our synthetic data used regional closure rate multipliers from published NCES data. Here's how those calibrated rates compared to what we found in the real data.

Closure Rate by Region: Synthetic vs. Real
Enrollment Distribution

Did Synthetic Institutions Look Real?

Beyond counts and rates, did the synthetic institutions have realistic enrollment sizes? This matters because small institutions close at much higher rates, so getting the enrollment distribution wrong would distort the risk profile. We compare the enrollment distribution for Private Nonprofit 4-year institutions — the sector with the most closures.

Enrollment Distribution: Private NP 4-Year (Synthetic vs. Real)
Key Findings

What Matched and What Didn't

Finding 1
Sector composition did not match closely. Public four-year institutions were 16.1% of the synthetic cohort versus 33.0% of the real-data cohort; private nonprofit four-year institutions were 38.3% versus 58.0%.
Finding 2
Closure prevalence differed by more than an order of magnitude: 5.6% synthetic versus 0.4% real-data. The synthetic cohort therefore cannot validate real-world model performance.
Finding 3
Synthetic regional closure rates ranged from 4.2% to 7.8%. Real-data regional rates ranged from 0.0% to 1.1%. The calibration overstated closure levels throughout.
Finding 4
Enrollment distributions also shifted. Among private nonprofit four-year institutions, median enrollment was 1,031 in the synthetic cohort and 1,414 in the real-data cohort. Even when a variable looks plausible, the cohorts are not exchangeable.

The ADM takeaway: synthetic data was useful for building the workflow and exploring questions, but it failed numerical validation for closure prevalence and model performance. The real-data output replaces it for visitor-facing results, and its six observed closures are still too sparse for institution-level prediction.

Method Note

How We Built the Comparison

Synthetic data generated using numpy random distributions calibrated to published NCES/IPEDS aggregate statistics. Real data from College Scorecard API and IPEDS direct downloads. Comparison metrics computed on matching sector and region definitions. Enrollment distributions binned using consistent bin edges across both datasets.

Sectors follow IPEDS classification (Public 4-year, Private NP 4-year, For-Profit 4-year, Public 2-year, Private NP 2-year, For-Profit 2-year, For-Profit <2-year). Regions follow Census Bureau divisions. Closure defined as institution no longer operating as of 2023.