A necessary aside

“But doesn't the research say this hurts kids?”

Any time someone proposes an objective placement standard, this comes back. There is a body of research showing that pushing students into 8th grade algebra can backfire. The research is real. It is also, almost always, about something else.

Let us take it seriously first, because it deserves to be taken seriously.

1

Charlotte-Mecklenburg, and the harm is real

When the district pushed to get 60% of students proficient in Algebra I by 8th grade, the share of median students taking it rose from 51% to 85%, and in the second-lowest quintile from 18% to 63%. Algebra I scores fell 0.32 standard deviations, and passing Geometry by 11th grade fell about ten points. For the lowest performers the authors call the harm “unambiguous”.

Clotfelter, C. T., Ladd, H. F. and Vigdor, J. L., “The Aftermath of Accelerating Algebra: Evidence from District Policy Initiatives”, Journal of Human Resources 50(1), 2015, 159-188. Charlotte-Mecklenburg and Guilford County, cohorts entering 7th grade 2000-01 to 2004-05.

2

California, at the district level

As districts raised their 8th grade algebra rates, average 10th grade mathematics scores fell by about 0.07 standard deviations per standard deviation of expansion, concentrated in large districts. The authors attribute it to system strain: teachers reassigned, algebra classes made more heterogeneous, pre-algebra classes stripped of their strongest students.

Domina, T., McEachin, A., Penner, A. and Penner, E., “Aiming High and Falling Short: California's Eighth-Grade Algebra-for-All Effort”, Educational Evaluation and Policy Analysis 37(3), 2015, 275-295. A panel of 222 California unified districts, 2003-04 to 2009-10.

3

Chicago, when the remedial track was abolished

Requiring Algebra I of every 9th grader raised failure rates, lowered grades, did not raise test scores and did not increase college entry. A companion study found the policy harmed high achievers, because schools responded by dissolving differentiated classes and their peer skill levels fell.

Allensworth, E., Nomi, T., Montgomery, N. and Lee, V. E., “College Preparatory Curriculum for All”, EEPA 31(4), 2009, 367-391. The companion finding on high achievers is Nomi, T., “The Unintended Consequences of an Algebra-for-All Policy on High-Skill Students”, EEPA 34(4), 2012, 489-505.

Those are the findings. Now look at what each one actually did.

Charlotte-Mecklenburg had no objective standard. The researchers went looking for the policy and reported that they could find no written statement of it. It was administrative pressure on principals to raise a number. There was no test-score threshold, no published rule, and no guarantee to anybody. The study is not evidence about objective criteria. It is evidence about what happens when a district leans on people to move students up without one.

California and Chicago moved students downward from the threshold, or removed the threshold altogether. California's estimate is a district-level average of expanding enrolment broadly. Chicago abolished differentiation entirely. Neither is a policy that says: a student who has already demonstrated readiness on a state test shall not be excluded from the course.

And the Charlotte-Mecklenburg authors say so themselves, in print. “Our estimate of the effect of taking Algebra I by 8th grade applies primarily to students in the middle of the initial test score distribution given that students at the top of the distribution virtually always take Algebra I by 8th grade.” “Our identifying variation comes almost entirely from students at lower points in the achievement distribution. Assessing the impact of placing higher-achieving students in algebra in 8th grade would require observing policy variation within that group. Their own abstract: accelerating algebra “appears benign or beneficial for higher-performing students but unambiguously harmful to the lowest performers.” The paper is routinely cited against a policy the paper explicitly declines to evaluate.

Four things that get conflated

What gets saidWhat is actually proposed
Algebra for all Automatic placement only for students who have already demonstrated readiness
Lowering the bar Making the existing bar visible and applying it consistently
Eliminating teacher judgment Preventing teacher recommendation from being the only way in
Studying students at a cut score Asking whether equally qualified students get equal access

The last row is the methodological heart of it. A regression discontinuity study estimates what happens to students immediately above and below a threshold. That is a real and useful number, and it is a local one. As the methodologists put it, the estimate “is only identified for the very specific subset of the population whose scores are just above or below the cutoff, and is not necessarily informative or representative of what the treatment effect would be for units whose scores are far from the cutoff.” Studying the marginal student tells you about the margin. It does not tell you whether a student two standard deviations above the margin should have been let in.

There is one North Carolina study that does answer that question, because it was built the other way round. A district collected its fifth grade teachers' placement recommendations and then put every sixth grader into the same rigorous course regardless, which meant the students who would have been turned away could be observed anyway. Every student in its tables had already scored at the top level. It is the closest thing in this piece to the missing counterfactual. The one time a district collected the recommendations and then ignored them →

How “you are lowering the bar” gets manufactured

Watch the sequence, because every step in it is reasonable and the result is not.

1

A district has no written rule

Placement runs on recommendation. Its own data, if anyone looks, shows that a large number of top-scoring students are not being placed.

2

It adopts an objective threshold

Now there is a cutoff, and students just above it are treated differently from students just below it.

3

Researchers evaluate it at the cutoff

Correctly. A threshold is what makes a rigorous causal estimate possible at all, and the estimate you can get is the one at the threshold.

4

Nobody reports the other thing

How many high scorers the old system had been passing over, who they were, or how they did. Those students are nowhere near the cutoff, so the design cannot see them, and no separate study is commissioned to look.

What ends up in the published record, then, is a precise number about the marginal student and silence about the excluded one. People read this and infer that the bar is being lowered. They do not realise there was no bar previously. And if the estimate at the margin comes out negative, the recommendation offered is to go back to teacher judgment, even though the study never compared teacher judgment to anything. It compared a student just above a line to a student just below it, both under the new rule.

Here is what makes that reasoning fail, and it is a factual claim rather than a philosophical one. The students the discretionary system was missing were not marginal students. One district filled 291 gifted places with higher-income fourth graders who had average end-of-grade math scores, in a year when 228 low-income children with superior scores were left out. When another district replaced referral with universal screening at an unchanged standard, one in five of the newly identified students met the highest threshold. Children like that sit well above any cutoff. An estimate at the margin is blind to every one of them, by construction, no matter how well the study is executed. So the evidence base is asymmetric by design. We measure, with real precision, what a rule costs at its margin. We have never measured what discretion costs above it. And then the two are weighed against each other as though the second were zero.
And there is a study that makes the point empirically. Three of the four authors of the California paper later ran the individual-level design their district study could not. Across 510 California schools with a score threshold for 8th grade algebra, assignment raised 10th grade mathematics scores, raised English scores, and raised advanced mathematics enrolment by 30 points in 9th grade and 16 points in 11th. Effects were non-negative for every subgroup they examined. And the finding that matters here: the benefit was larger where the threshold was set higher. Schools with high cutoffs sustained a 20 to 25 point advanced-mathematics advantage into 11th grade; schools with low cutoffs sustained about 5. Where you set the bar changes the sign. That is the locality problem made concrete, and it points the opposite way from how the earlier work gets quoted. McEachin, A., Domina, T. and Penner, A. M., “Heterogeneous Effects of Early Algebra Across California Middle Schools”, Journal of Policy Analysis and Management, 2020. A multi-site fuzzy regression discontinuity across 510 schools and 753 school-years, 2007 to 2011.

What the evidence on discretion actually shows

The comparison is never “objective standard versus nothing”. It is objective standard versus the system we already have, which is discretion. That system has been studied too.

1

Referral misses qualified students at an unchanged bar

When one large district replaced parent and teacher referral with universal screening, without changing any eligibility threshold, gifted identification rose about 180% among disadvantaged students, 130% among Hispanic students and 80% among Black students. One in five of the newly identified students had an IQ of 130 or above, the district's highest threshold, the one applied to everybody. They had always qualified. Nobody had sent them to be tested.

Card, D. and Giuliano, L., “Universal Screening Increases the Representation of Low-Income and Minority Students in Gifted Education”, PNAS 113(48), 2016, 13678-13683. Broward County, Florida, 2004-05 and 2006-07 cohorts. The eligibility thresholds are set out in the working paper version, NBER 21519, 2015.

2

Discretion is where the gap enters

Holding prior test scores constant, Black students are identified as gifted at roughly half the odds of otherwise identical students. The gap concentrates at the referral stage, and it shrinks sharply when the teacher shares the student's race. This is not a story about children's ability. It is a story about who gets nominated.

Grissom, J. A. and Redding, C., “Discretion and Disproportionality”, AERA Open 2(1), 2016. 21,260 students in the Early Childhood Longitudinal Study, K to 5.

3

Demonstrated achievement beats ability testing

In the same district, students admitted to gifted classrooms on prior achievement rank gained 0.4 to 0.5 standard deviations in 4th grade reading and mathematics, with the largest gains among low-income and minority students. Students admitted on an IQ threshold gained essentially nothing. What a student has already done predicts better than what a test says they could do.

Card, D. and Giuliano, L., “Does Gifted Education Work? For Which Students?”, NBER Working Paper 20453, 2014. Regression discontinuity at the IQ and achievement-rank thresholds in the same district.

4

An objective rule closes the access gap. In Wake County.

When one North Carolina district replaced counsellor discretion with an EVAAS threshold, the income gap in acceleration among equally scoring students fell from 10.5 points to 2 or 3, and the Black and Hispanic gap from 7.4 points to 1 or 2, no longer statistically distinguishable from zero. Same standard for everyone. The gap was in who got asked.

Dougherty, S. M., Goodman, J. S., Hill, D. V., Litke, E. G. and Page, L. C., “Middle School Math Acceleration and Equitable Access to Eighth-Grade Algebra: Evidence from the Wake County Public School System”, EEPA 37(1 suppl), 2015, 80S-101S. This paper reports access, not achievement; the follow-up on outcomes was never published.

“But that district used a lower bar for minority students.” The district in the first of those four findings is Broward County, Florida, and the study is Card and Giuliano, “Universal Screening Increases the Representation of Low-Income and Minority Students in Gifted Education”, PNAS 113(48), 2016, with the eligibility rules set out in the working paper, NBER 21519, 2015. This objection is raised almost every time that study is cited, so here is the answer in full, because it is wrong twice. It is wrong about the mechanism. The district did have two thresholds. An IQ of 130 for students generally, and 116 for students who were on free or reduced price lunch, or who were learning English. Those categories are not race. They correlate with it, but a wealthy Black child faced the 130 threshold and a poor white child faced the 116 one. And the district had both thresholds before universal screening and kept both afterwards. The paper's own phrase for the change is “with no change in the standards for gifted certification”. What changed was who got measured at all. Before, a parent or a teacher had to put your name forward. After, every second grader was screened. The bar did not move. The queue to be tested against it stopped being a nomination. And it is wrong about what the result shows. One in five of the students newly identified had an IQ of 130 or above, against 25% of the students who had always been identified. So referral was not merely missing children who qualified under the lower threshold. It was missing children who cleared the highest threshold in the district, at almost the same rate those children occur among the students it did find. Which leaves the two-tier standard looking worse, not better. It exists on the assumption that disadvantaged children will not clear the top bar. Screening found plenty who had already cleared it, sitting in ordinary classrooms, waiting for somebody to say their name.
And then the district switched it off, which is as close to a controlled experiment as this subject ever gets. Screening began in 2005. In a budget crisis in 2007 the district cut the overtime budget for the staff doing the testing. In 2010 it suspended the screening programme altogether. By 2011 the gifted share of third graders was back to where it had been in 2004‑05, and the racial and ethnic gaps had reopened to their old size. Now notice what had actually been expensive, because it is not what people assume. The screen was cheap. A nonverbal test, under an hour, given by teachers in an ordinary classroom. What cost money was what the screen turned up: every child it flagged then needed an individual evaluation, about three hours of a psychologist's time, paid as overtime. Around 1,300 extra tests in the first year and about as many in the second. The district could afford to look. It could not afford to confirm what looking found. A modified version came back in 2012 and the district still screens today, but the results have never returned to what they were. The replacement leans more on verbal ability, and referral discretion came back with it. Susan Dynarski, “Why Talented Black and Hispanic Students Can Go Undiscovered”, New York Times, 10 April 2016: “Broward County suspended its universal screening program in 2010 in a spate of budget cutting after the Great Recession. Racial and ethnic disparities re-emerged, as large as they were before the policy change.” The funding and testing detail is from Card and Giuliano, NBER Working Paper 21519, 2015. Both accounts attribute the reversal to budget pressure in their own voice; no district official is quoted giving a reason, and we are not putting one in their mouths.
Now go back to the state's own slide. In 2017‑18, before the statute took effect, 31,170 eighth graders sat the NC Math 1 exam, and 13,991 of them had not scored at the top level in 7th grade. Nearly half the seats in 8th grade algebra were already going to students below the highest level, allocated by teacher recommendation, parent request and local judgment, with no rule and no report. So a statute that guarantees a seat to students who score at the very top cannot possibly be lowering a bar. The bar was already lower than that. It was just invisible, unevenly applied, and unaccountable. Writing it down did not move it. Writing it down made it checkable. And note when that slide was made. It is April 2019. Those are 2016‑17 and 2017‑18 results, scored on the old cut scores. Every “Level 5” on it is a Level 5 in the sense the word had when the law was written, when about one seventh grader in five reached it. Hold on to that, because it is about to matter.

“That was years ago. Things are different now.”

This is the second thing that always comes back, and it is a fair challenge. So show us. The usual answer is that the annual report to the General Assembly shows it: 92% of eligible students placed, 94% the year before. Look at what that number is a percentage of.

Who is sitting in 8th grade Math 1

And what share of those seats the annual report describes

Sources: NCDPI, Supporting the Implementation of HB 986, April 2019, for 2017‑18. NC DPI School Assessment and Other Indicator Data for 2024‑25, subject codes M1SEP and MA07, with the 2024-25 seat count matched to the 2023-24 grade 7 cohort that feeds it. On the 2019 slide, 13,101 of these students have a recorded 7th grade level below 5 and a further 890 have no recorded level, giving 13,991 who did not score at the highest level. The 2024‑25 bar is generous to the state: it assumes every top-scoring 7th grader went on to take Math 1 in 8th grade. They do not, so the real share who did not score at the highest level is higher than shown.

Be careful about what did and did not change here, because it is easy to get wrong. The number of 8th graders taking Math 1 went up, from 31,170 to 35,713. More students are getting into the course, not fewer. What fell is the number the rule reaches.

In 2017‑18, 45% of the seats in 8th grade algebra went to students who had not scored at the top level. In 2024‑25, at least 63% did. The share of that classroom allocated by somebody's judgment rather than by the rule did not shrink after the law. It grew by twenty points.

It grew because the state moved the cut score. Fewer students clear the bar, so the rule reaches fewer seats, so more of the room is filled the old way, by recommendation and request and whoever asked. The report to the General Assembly covers a third of that classroom. It says nothing about the other two thirds, which is now most of it: not who they are, not how they were chosen, not how they did.

So the honest answer to “things are different now” is: yes, they are. There are more discretionary seats in 8th grade algebra than there were in 2018, not fewer, and a smaller share of the decision is governed by any published rule. The report that is offered as evidence of improvement cannot show improvement in access, because the state redefined the denominator in the middle of the series. A placement rate among students who cleared a bar tells you nothing when the bar moved. Which points at the two numbers that would actually settle this argument, and neither of them is published by anybody. Are more students taking Honors mathematics in high school? Nobody counts it. Not the placement report, not the advanced courses report, not the school report cards. The closest published figure is Advanced Placement calculus enrolment, and that is down about 2,000 students since 2016‑17. Are fewer students needing remedial mathematics after high school? The most recent published North Carolina developmental placement rate is for students who entered in 2014‑15, more than a decade ago. The community college system's own implementation plan scheduled its first figures under the current framework for the summer of 2026. Those are the questions. Are there more students in Honors? Are there fewer students needing remedial mathematics? Everything else is a proxy, and the proxy currently on offer is one the state can move without telling anyone.

And this has been documented for thirty years

The other reason “that was a long time ago” does not land is that there is no decade in which it stopped being found. Different researchers, different states, different methods, the same result: hold measured achievement constant, and access still depends on who the student is.

1992

Same scores, ten times the odds

RAND examined transcripts and test scores in three California high schools and modelled the probability of reaching college-preparatory mathematics. At one school, Asian students were more than ten times as likely as their Latino classmates with the same mathematics and reading scores to be enrolled in it.

Oakes, Selvin, Karoly and Guiton, Educational Matchmaking, RAND R-4189-NSF/PC, 1992, Table 5.10.

1998

Three times more likely, at equal measured ability

Six thousand eighth graders in one large urban district, all of them scoring in the top quartile on a nationally normed mathematics test. Students from high socioeconomic backgrounds were three times more likely to be placed into algebra than low socioeconomic students with the same demonstrated ability.

Carolyn B. Stone, “Leveling the Playing Field”, The Urban Review 30(4), 1998. The three-times figure as such is quoted from Stone and Turba, 1999, summarising that study.

2011

North Carolina, and the same children a year apart

Seven middle schools put every sixth grader into one rigorous course. Fifth grade teachers had filled in their usual recommendation forms without knowing the forms would be set aside. Among top-scoring students, fifth grade teachers had recommended the advanced course for 71% of white children and 12% of non-Asian minority children. After one year in the same classroom, sixth grade teachers recommended the next advanced course for 68% and 61%.

Stiff, L. V., Johnson, J. L. and Akos, P., “Examining What We Know for Sure: Tracking in Middle Grades Mathematics”, in Tate, King and Anderson (eds), Disrupting Tradition, NCTM, 2011, Table 6.2.

2014

North Carolina again, with the performance controlled for

A longitudinal study followed North Carolina students from late elementary school into 8th grade. Black students had reduced odds of being placed in algebra by the time they entered 8th grade even after controlling for their performance in mathematics.

Faulkner, Stiff, Marshall, Nietfeld and Crossland, Journal for Research in Mathematics Education 45(3), 2014. Lee Stiff again.

2017

And the newspapers found it too

Seven years of state data, tracked student by student: 9,000 low-income North Carolina children with the scores, counted out of the classes those scores should have opened.

Neff, Helms and Raynor, Counted Out, News & Observer and Charlotte Observer, May 2017. Part five.

2019

Identical transcripts, different name

The cleanest test anyone has run. School counsellors were given the same transcript and asked whether to recommend the student for advanced coursework. Only the name on it changed. Black female students were less likely to be recommended for AP Calculus, and were rated least prepared.

Francis, de Oliveira and Dimmitt, B.E. Journal of Economic Analysis and Policy 19(4), 2019.

This is not a finding that keeps being made because nobody noticed the first time. It keeps being made because nothing in the system is designed to catch it. Each study is a one-off, funded by somebody, and when the funding ends the measurement ends. The 2009 EVAAS analysis that opens this piece was the most complete one North Carolina ever had, and it has not been repeated in sixteen years.

What we are not saying

We are not saying every student should take algebra in 8th grade. We are not saying tracking should end, and the Chicago evidence is a real warning about what happens to strong students when differentiation is dissolved without a plan. We are not saying teacher judgment is worthless; it catches things a test cannot, and it should keep opening doors. We are saying it should not be the only thing that can open one.

And we will be honest about the limit of our own case. North Carolina's law, Wake County's threshold and the automatic-enrolment policies in other states have all been shown to widen access and to close gaps between equally qualified students. None of them has yet produced a published evaluation of what happened to those students' achievement. That study has not been done. It is the same missing study this entire piece is about, and it is the first thing a serious accountability report would answer.