Madison Metropolitan School District · 2025–26 early literacy screener

The bar moves faster than the kids

Every grade in the district ended the year with fewer students at benchmark than it started — except third grade, which didn’t move at all.

The year in one picture

Percent of Madison kids meeting the reading target, start versus end of the 2025–26 school year, by grade 4K: 82 percent at the start of the year, 66 percent by the end. Kindergarten: 76 to 60 percent. Grade 1: 73 to 65 percent. Grade 2: 69 to 68 percent, nearly unchanged. Grade 3: 72 percent at both the start and the end of the year, no change. Percent of Madison Kids Meeting the Reading Target Beginning vs. end of the 2025–26 school year, by grade Start of year End of year 0% 25% 50% 75% 100% 4K 82% 66% Kindergarten 76% 60% Grade 1 73% 65% Grade 2 69% 68% Grade 3 no change 72% → 72% Four of five grades ended the year with fewer kids meeting the target than when it began. Only third grade held steady.
This is the introduction. The rest of the report explains the likely cause, what the numbers do and don’t prove, and how results differ by student group.
4K−16
Kindergarten−16
Grade 1−8
Grade 2−1
Grade 30

The short version

These numbers exist because of a law. 2023 Wisconsin Act 20 requires every public school district in the state to screen students in 4K through third grade for early literacy skills, and the Department of Public Instruction selected Pearson’s aimswebPlus as the statewide screener. Beginning in 2025–26, screening is required in fall and spring for 4K through third grade, and at midyear for 5K through third — which is why the 4K line in the chart below has a gap in it. Four-year-olds are screened twice a year, not three times. (2024–25 used a different, reduced schedule; see below.)

Statewide screening began the year before, in 2024–25, and Madison has results from that year too. What changed is the cadence: a 2024 amendment, Act 192, reduced the number of required administrations for 2024–25 only, so 2025–26 is the first year run on the full statutory schedule. The two years still shouldn’t be lined up against each other, for reasons that have nothing to do with missing data, covered in full below — including where each year’s numbers came from: the 2025–26 figures were obtained directly from the district, since DPI has not yet published a statewide report for that year. Five things stand out in the 2025–26 results.

The grade-by-grade slide

Read the change column below from top to bottom and the shape is unmistakable: the decline shrinks as the grades get older, until it disappears entirely.

Figure 1

Percent of students at or above benchmark, by grade and screening window

Line chart of benchmark rates by grade across three screening windows 4K falls from 82 percent in fall to 66 percent in spring with no winter screening. Kindergarten falls from 76 to 65 to 60 percent. Grade 1 falls from 73 to 68 to 65 percent. Grade 2 moves from 69 to 67 to 68 percent. Grade 3 stays at 72 percent at all three windows. 100% 90% 80% 70% 60% 50% Fall Winter Spring 4K not screened in winter 82% 76% 73% 72% 69% Grade 3 · 72% Grade 2 · 68% 4K · 66% Grade 1 · 65% Kindergarten · 60%
Benchmark targets rise at each window, so a falling line means fewer students clearing a higher bar. Act 20 requires only fall and spring screening for 4K; the dashed connector spans a window in which no data was collected, by design.
Percent at or above benchmark, 2025–26
GradeFallWinterSpringChange
4K82%not screened66%−16
Kindergarten76%65%60%−16
Grade 173%68%65%−8
Grade 269%67%68%−1
Grade 372%72%72%0

The temptation is to read a falling line as students getting worse. That is not what a falling line shows. Fall benchmarks for a four-year-old sit close to a floor — naming a handful of letters clears them, which is why 82% of 4K students met the fall mark. By spring, the same children face a higher cut score on partly different subtests.

Pearson says as much in its own guidance for Wisconsin educators. Asked how a student can post the same score in two seasons and receive a lower national percentile, the vendor answers that percentile rankings are adjusted because norming data shows scores rise from fall to winter to spring — so the same raw score almost always ranks lower later in the year. That is the mechanism, stated by the test publisher. What it doesn’t tell us is how much of Madison’s 16-point drop it accounts for.

Which is why third grade’s flatness deserves a second look. It is the one place where the fall picture and the spring picture agree, and the one grade sitting on the most consistent measure across the year. Its 72% is the most stable number in the file.

Enrollment isn’t the explanation

One obvious alternative story — that the numbers fell because a different set of students was tested — doesn’t hold. Working backward from the published counts, roughly 1,300 students were screened in 4K and between 1,680 and 1,780 in each of kindergarten through third grade, and those totals barely move across the three windows. The denominator is stable. The rate is what changed. (This rules out gross cohort turnover; it does not substitute for student-level matched scores.)

Who is below the line

District averages hide a great deal. The same five grades look very different depending on which students you follow. The five charts below show each grade on its own: six student groups, with a bar for every testing window that year. Spring is shown in solid color; fall and winter are shown in gray, since spring is the most complete picture of where the year landed.

Figure 2

4K: percent at benchmark by student group and testing window

4K: benchmark rates by student group and testing window 4K. All: fall 82, spring 66 percent; White: fall 93, spring 86 percent; Black: fall 79, spring 53 percent; Hispanic: fall 64, spring 45 percent; Econ. Disadv.: fall 72, spring 48 percent; IEP: fall 54, spring 34 percent. 100% 75% 50% 25% 0% 66 All 86 White 53 Black 45 Hispanic 48 Econ. Disadv. 34 IEP
Black students and economically disadvantaged students fell furthest of any group in 4K — 26 and 24 points, fall to spring.
Figure 3

Kindergarten: percent at benchmark by student group and testing window

Kindergarten: benchmark rates by student group and testing window Kindergarten. All: fall 76, winter 65, spring 60 percent; White: fall 92, winter 86, spring 83 percent; Black: fall 65, winter 47, spring 41 percent; Hispanic: fall 50, winter 37, spring 33 percent; Econ. Disadv.: fall 56, winter 43, spring 38 percent; IEP: fall 49, winter 44, spring 33 percent. 100% 75% 50% 25% 0% 60 All 83 White 41 Black 33 Hispanic 38 Econ. Disadv. 33 IEP
The White–Hispanic gap widens more here than anywhere else on this page: 42 points in fall, 50 by spring.
Figure 4

Grade 1: percent at benchmark by student group and testing window

Grade 1: benchmark rates by student group and testing window Grade 1. All: fall 73, winter 68, spring 65 percent; White: fall 91, winter 89, spring 87 percent; Black: fall 51, winter 47, spring 44 percent; Hispanic: fall 59, winter 46, spring 47 percent; Econ. Disadv.: fall 56, winter 47, spring 46 percent; IEP: fall 49, winter 41, spring 42 percent. 100% 75% 50% 25% 0% 65 All 87 White 44 Black 47 Hispanic 46 Econ. Disadv. 42 IEP
Hispanic students saw the largest drop of any group in first grade — 12 points — more than double the district average’s.
Figure 5

Grade 2: percent at benchmark by student group and testing window

Grade 2: benchmark rates by student group and testing window Grade 2. All: fall 69, winter 67, spring 68 percent; White: fall 85, winter 84, spring 84 percent; Black: fall 54, winter 49, spring 51 percent; Hispanic: fall 52, winter 48, spring 50 percent; Econ. Disadv.: fall 53, winter 49, spring 51 percent; IEP: fall 45, winter 38, spring 37 percent. 100% 75% 50% 25% 0% 68 All 84 White 51 Black 50 Hispanic 51 Econ. Disadv. 37 IEP
Students with IEPs fell furthest here, 8 points, even as the grade’s overall rate barely moved.
Figure 6

Grade 3: percent at benchmark by student group and testing window

Grade 3: benchmark rates by student group and testing window Grade 3. All: fall 72, winter 72, spring 72 percent; White: fall 87, winter 89, spring 89 percent; Black: fall 56, winter 54, spring 52 percent; Hispanic: fall 57, winter 55, spring 57 percent; Econ. Disadv.: fall 57, winter 55, spring 55 percent; IEP: fall 42, winter 39, spring 39 percent. 100% 75% 50% 25% 0% 72 All 89 White 52 Black 57 Hispanic 55 Econ. Disadv. 39 IEP
The only grade where a group’s rate rose over the year: White students went from 87% to 89%.
Spring 2026 — percent at or above benchmark
Student group4KKGGrade 1Grade 2Grade 3
All students66%60%65%68%72%
White86%83%87%84%89%
Asian60%75%75%77%80%
Multiracial79%62%62%68%74%
Black or African American53%41%44%51%52%
Hispanic or Latino45%33%47%50%57%
English learners39%*30%49%48%50%
Economically disadvantaged48%38%46%51%55%
Students with IEPs34%33%42%37%39%

White students’ rates hold up across the year — third grade actually improved, from 87% in fall to 89% in spring. Almost every other group’s did not. The result is that whatever is decreasing early-grade rates does so unevenly, and gaps that were already wide in September were wider in May.

Students with IEPs

The most consistent decline in the file. Every grade ended lower than it began: 4K from 54% to 34%, kindergarten from 49% to 33%, first grade from 49% to 42%, second from 45% to 37%, third from 42% to 39%. No grade finished above 42%.

English learners

A sharp fall in kindergarten, from 58% to 30%, but nearly flat in grades one through three — 58% to 49%, 50% to 48%, and 51% to 50%. The kindergarten drop tracks the point in the year when the screener’s language demands increase, which is worth knowing before drawing conclusions about instruction.

Asian students

The quiet exception to the pattern that high-performing groups stayed high. They fell from 89% to 75% in kindergarten and from 79% to 60% in 4K — declines closer to the district’s than to White students’. Whatever is decreasing early-grade scores is not simply a function of where a group started.


What the screener actually is

aimswebPlus, from NCS Pearson, is a PreK–12 reading and math screening and progress-monitoring system. Wisconsin’s Department of Administration selected it through a public bid, DPI contracted with Pearson in July 2024, and the early literacy and reading measures are provided to Wisconsin schools and districts at no cost. Madison, like every other public district in the state, did not choose this instrument.

The screener also measures different aspects of literacy at each grade. For Act 20 compliance in grades 2 and 3, only two measures count: Vocabulary and Oral Reading Fluency. In kindergarten, the Early Literacy Composite combines two of the six required measures; in first grade, three of seven. So “68% at benchmark in Grade 2” and “72% at benchmark in Grade 3” aren’t two scores on one ruler — they’re two grades being measured on different things. Comparisons within a single grade, across the school year, are more meaningful than comparisons across grade levels.

Screening is standardized. Diagnosis isn’t.

Act 20 has two stages, and only the first is uniform. Every district screens on aimswebPlus — that’s the fall/winter/spring benchmark data on this page. The second stage, a diagnostic assessment for students below the 25th percentile, is a local choice: Pearson’s own guidance for Wisconsin says diagnostic measures are “what you decide, locally, to use.” Madison’s 2024–25 filing names its diagnostic tool as FastBridge and Star by Renaissance Learning — not aimswebPlus — for every school and grade that reported one. A student’s screener and their diagnostic assessment can be, and in Madison’s case are, two different products.

That split is common statewide. Of the 429 Wisconsin districts that reported a diagnostic tool for 2024–25, aimswebPlus was the choice for most, but not all:

Diagnostic assessment tools in use statewide, 2024–25 (by district)
ToolDistricts
aimswebPlus by Pearson314
FastBridge and Star by Renaissance Learning152
i-Ready by Curriculum Associates71
MAP Fluency by NWEA12
HMH Amira by Houghton Mifflin Harcourt7

Of 430 Wisconsin districts and independent charter operators in the 2024–25 Act 20 report, 429 named at least one diagnostic tool; one reported none. Many districts named more than one — most often aimswebPlus paired with FastBridge/Star or with i-Ready — so each is counted once per tool named and the column does not sum to 429.

Why 2024–25 and 2025–26 don’t line up

Madison has a prior year of screener results, published in DPI’s 2025 Act 20 annual report. It shouldn’t be compared with 2025–26, for three reasons. Pearson re-normed aimswebPlus for 2025–26, replacing a 2015 reference sample, and DPI says results from before and after that change aren’t directly comparable. The schedule differed, too: 2024–25 began with the midyear window instead of fall, so 4K was screened just once that year, in spring, and 5K through third grade were screened twice, at midyear and in spring — not the fall-and-spring, three-times-a-year schedule described above, which took effect in 2025–26. And the two years aren’t even the same kind of source: the 2024–25 figures below are DPI’s own statewide compilation, while the 2025–26 figures on this page came directly from Madison, since DPI hasn’t yet published a 2025–26 report to check them against.

Madison’s 2024–25 numbers

DPI’s Act 20 annual report doesn’t publish a benchmark rate. It publishes one figure per grade: the share of students below the 25th percentile — the statutory threshold, detailed further down this page, that triggers a diagnostic assessment and personal reading plan in 5K through third grade. Madison’s district-wide figures for 2024–25, the first year of statewide screening (also viewable on an interactive map of DPI’s published data):

2024–25 district-wide results (Act 20 annual report)
GradeEnrolledAt/above 25th pct.Below 25th pct.
4K1,52787.0%13.0%
Kindergarten1,81445.6%54.4%
Grade 11,879∼41.6%†∼58.4%†
Grade 21,83153.9%46.1%
Grade 31,84656.9%43.1%

† DPI’s published file redacts Madison’s first-grade count directly. The figure above is estimated from the number of first graders who began a personal reading plan (1,098 of 1,879 enrolled, or 58.4%) — a close stand-in everywhere else in the file, where the personal-reading-plan count sits within a point of the below-25th-percentile count.

This is not the benchmark rate. A percentile ranks a student against other test-takers; a benchmark is different — it’s the score Pearson has found predicts roughly an 80% likelihood that a student will be reading at grade level by the end of the year. Publicly available Pearson technical documentation puts that cut near the 45th percentile for most reading measures, and closer to the 35th percentile for the early-literacy measures used in the youngest grades. The 25th percentile, by contrast, is just the bottom quarter of test-takers — a lower, simpler line that Act 20 uses because a statute needs one, not because it was calibrated to predict future reading success the way Pearson’s benchmark was. “At or above benchmark,” the number used everywhere else on this page, and “at or above the 25th percentile,” the only number DPI compiles statewide, are answering different questions. There is no published conversion between them. Setting the 2024–25 at-risk rate beside the 2025–26 benchmark rate would not show a year-over-year trend — it would show the distance between two different rulers, applied in different years, under different norms. Both problems compound rather than cancel.

How widely is it used?

Pearson does not publish adoption counts, and no reliable public tally of aimswebPlus states, districts, or schools exists — so any specific number should be treated with suspicion. What can be verified is narrower:

What “below benchmark” does and doesn’t mean

Everything on this page is a benchmark rate. Act 20 does not run on benchmark rates. The statute turns on a different threshold — the 25th percentile — and applies it differently by grade:

Act 20 “at-risk” determination
GradeHow a student is identified
4KBelow the 25th percentile on both required spring subtests. No student is identified from fall results. No diagnostic assessment or personal reading plan is required.
5K & Grade 1Below the 25th percentile on the specified composite.
Grades 2–3Below the 25th percentile on oral reading fluency.

For 5K through third grade, a student below that line must receive a diagnostic assessment — including a family history survey — and a personal reading plan, with notification to families and ongoing progress reporting. So behind every percentage on this page sits a set of individual obligations. But the two thresholds are not the same, and the number of students “below benchmark” is not the number legally identified as at risk. Nor can a page of district percentages establish whether Madison met its Act 20 duties: compliance is documented student by student, not in aggregate.

Context this data doesn’t contain

Madison purchased a new early literacy curriculum in 2022, following a 2021 task force report. These 2025–26 results fall in the fourth year of that adoption — which makes them a natural baseline and a poor before-and-after. No aimswebPlus series exists from before the purchase, because the statewide screener was not in use until 2024–25; and that one prior year was scored against different norms, so it cannot serve as a baseline either.

Where the data is thin

An honest read includes what the file cannot tell us. American Indian/Alaska Native and Native Hawaiian/Pacific Islander results are masked at every grade and every window, so those students are absent from this analysis entirely. Several 4K and kindergarten cells for English learners and Advanced Learners are masked as well.

A handful of published cells also look like reporting errors rather than findings. Winter first-grade results for Black students show 47% alongside a count of 47, where the count should be closer to 150. Spring 4K results for English learners show 39% alongside a count of 9. Anyone citing those specific cells should confirm them with the district first.

The question the data raises

If the early-grade declines are largely an artifact of an advancing bar, the spring snapshot is still the most complete one available — and it says that between a quarter and two-fifths of Madison’s youngest students finished the year below their grade’s benchmark, with that share climbing past half for Black, Hispanic, English-learning, and IEP students in most grades. That is a benchmark rate, not a legal determination; a student can be identified as at-risk under Act 20 from any window’s result, not only spring’s.

Screening data cannot say why. It can only say where to look, and it is pointing at kindergarten and first grade, where the floor drops out fastest and the gaps open widest.


Independent review

Three AI systems reviewed this page twice, a few weeks apart, against the same question: are the observations and conclusions correct, using the linked data and the Act 20 requirements? The first round’s dissent (from Perplexity) prompted several of the corrections already folded into the sections above — treating the rising-bar reading as an inference, separating “below benchmark” from the statutory 25th-percentile threshold, and noting that compliance can’t be judged from aggregate rates. This second round checked the revised page. Gemini and Grok found no remaining issues, including specifically validating the two additions made after round one — Madison’s diagnostic-tool choice and the district-direct provenance of the 2025–26 figures. Perplexity found three new, narrower problems, all corrected below.

Gemini · round 2

“Yes, the observations and conclusions presented in the linked report are correct.”

Gemini confirmed the statutory framework point by point — the aimswebPlus mandate, the twice-yearly/three-times-yearly cadence, and Act 192’s one-year reduction that makes 2025–26 the first full-schedule year — and called out the screening-versus-diagnosis distinction specifically: “the observation that Madison utilizes FastBridge and Star for diagnostics, rather than aimswebPlus, aligns perfectly with the flexibilities granted to districts under the law.”

On the data, it re-verified the fall-to-spring changes by grade, the widening kindergarten White–Hispanic gap, and the IEP and English-learner subgroup trends against the tables. It described the seasonal-norming mechanism as “a highly accurate psychometric conclusion” and endorsed the page’s own caveat that proving it would take student-level matched scores — the inference is sound, not the same as proof.

Grok · round 2

“Correct based on the linked district data and Act 20 requirements, with appropriate caveats already noted in the report itself.”

Grok specifically confirmed the provenance framing added after round one: that the 2025–26 figures come directly from the district because DPI has not yet published a statewide report for that year, and that comparing 2024–25 to 2025–26 is invalid on both re-norming and schedule grounds. It re-checked the fall-to-spring figures, stable enrollment counts, and spring demographic cross-tabs against the tables and found them consistent.

It credited the page for not overclaiming Act 20 compliance, since compliance “turns on individual diagnostics/plans/progress monitoring” rather than aggregate rates, and for correctly distinguishing the district’s reported “at or above benchmark” metric from the statute’s exact below-25th-percentile threshold.

Perplexity · round 2, dissenting in part

“Mostly accurate… but several conclusions overreach what these cross-sectional benchmark percentages can establish.”

Perplexity confirmed the schedule, the arithmetic, and the subgroup gaps, then raised three new, more specific objections than round one.

  • The age gradient isn’t necessarily an instructional finding. Calling it “the single most important pattern” implied a distinctive early-grade trend, when the same shape could come from differences in measures, norms, or administration across grades rather than anything happening in classrooms.
  • “Not yet reading at the level the screener expects” needs a narrower referent. Below-benchmark means below the aimswebPlus cut score for that grade’s composite — not below Wisconsin’s official measure of grade-level reading, which for third grade is the separate Forward Exam.
  • “The spring picture is the one that counts” misstates Act 20. A 5K–3 student is identified as at-risk the first time any window’s result falls below the 25th percentile — fall and winter results trigger a personal reading plan just as spring’s do. The law doesn’t treat spring as uniquely dispositive, even though it is the most complete snapshot available.

What changed as a result. The age-gradient finding now says “the largest pattern in the data” rather than the single most important one, with an explicit note that the data can’t distinguish an early-grade instructional story from a measurement one. The spring-focused finding is rewritten twice — here and in the closing section — to say a below-benchmark result is a screener cut score rather than a claim about grade-level reading, to link to the Forward Exam as Wisconsin’s actual grade-level measure, and to state plainly that Act 20 identification happens at any window, not spring exclusively. And a new sentence in the screener-description section flags that the benchmark composite differs by grade, so rates should be compared within a grade across time more confidently than across grades at a point in time.


Notes and sources

Data

The screener

Statute and guidance

Background