Out of five hundred firms, eight passed the test.
Not eight that said the right thing — eight that said it and could be shown to have kept the programs behind it. That number is the argument of this paper.
AUTHOR’S NOTE
For eighteen months the corporate inclusion landscape has been sorted into two bins. Companies that stood firm. Companies that retreated. The sorting is tidy, it is quotable, and it is close to useless as a description of what is actually happening inside institutions.
A study released this month gives us an occasion to say so with evidence rather than impression. Hanna Folsz and Jacob Grumbach asked whether firms that refused to dismantle their diversity programs after Executive Order 14173 paid a financial price. They found no penalty — in stock returns or in revenue, across four independent ways of measuring who maintained their programs.
That finding matters. But the part of their paper that should change how practitioners think is not in the abstract. It is in the sample size and the appendix.
To identify firms that genuinely maintained their programs, the authors refused to code on public statements alone. They required programmatic continuity as well — no internal memos, no quiet dissolutions. Applying that two-part test to the entire S&P 500, they found eight firms. And of those eight, their own appendix documents three that subsequently rebranded, renamed, or partially unwound what they had just publicly reaffirmed.
Two researchers set out to measure corporate resistance and discovered they first had to solve the problem of telling posture from structure. That is the whole subject of this paper, arrived at independently, from the other direction.
What follows reads their findings carefully, including the limits, and then applies the eight pillars of the Architecture of Inclusion to the landscape those findings expose.
I. What the Study Actually Is
The first thing to correct is the framing that has already begun circulating. This is not a study of the business case for diversity. It is a study of whether civil society organizations can refuse an instruction from the executive branch without being financially punished for it.
Folsz and Grumbach say so directly. Their significance statement opens on how authoritarian leaders consolidate power by intimidating businesses and other civil society organizations, and asks whether businesses can resist political pressure without paying a price. Diversity programs are the test case, not the subject. The subject is compliance.
That reframing matters for how the finding should be used. Read as a DEI study, it says something modest. Read as what it is — an empirical test of whether institutional intimidation actually carries the costs institutions fear — it says something considerably larger, and it speaks to universities, law firms, foundations, and professional associations facing the same calculus.
The design
Executive Order 14173, issued January 21, 2025, directed federal agencies to end diversity preferences and to identify up to nine civil compliance investigations of publicly traded corporations, large nonprofits, foundations, bar and medical associations, and universities. A Department of Justice memorandum followed on February 5.
The authors estimate cumulative abnormal returns using a Carhart four-factor model over a 250-trading-day pre-event window, then apply difference-in-differences — comparing the change in returns before and after the order for maintaining versus rolling-back firms.
Because abnormal returns already net out market, size, value, and momentum exposure, the identifying assumption is weaker than parallel trends in raw returns.
They test four independent measures of who maintained programs: public reaffirmation with programmatic continuity; the DEI Watch activist classification; intensity of diversity language in SEC 10-K filings; and rejection of anti-diversity shareholder resolutions.
They then repeat the analysis on revenue — eight quarters of 10-Q filings — because returns are forward-looking expectations while revenue is realized performance, and revenue is where an actual consumer boycott would show up.
The findings
Null across the board. Firms that maintained their programs performed no differently from firms that dismantled them, on returns and on revenue, under all four treatment measures. Short-run analysis over the seven trading days around the order shows no immediate reaction either. When President Trump publicly attacked Apple by name over its shareholder vote, the authors find no subsequent decline in Apple’s abnormal returns.
The authors’ own conclusion is the one practitioners should sit with: organized resistance by major civil society actors has been rare, and their findings imply that this scarcity may stem from firms overestimating what resistance would cost them.
The limits — stated plainly, because credibility depends on it
This is a null result on penalty. It is not evidence that diversity programs increase profitability, and it does not claim to be.
Statistical power is uneven, and honesty requires saying where. With only eight firms in the narrowest treatment group, the minimum detectable effect on returns runs as high as 0.60 control-group standard deviations — a wide net. The revenue results are far more precisely estimated, with minimum detectable effects as low as 0.03, and the broader treatment measures tighten the returns estimates considerably. The strongest claim the data supports is that a large penalty is unlikely, not that the true effect is exactly zero.
The authors raise the selection concern themselves: perhaps only firms that could absorb the cost chose to resist. They argue against it — firms routinely misjudge consumer reaction, and the administration’s behavior is widely characterized as unpredictable — but it remains a limit.
The sample is S&P 500. These are large, diversified, politically resourced firms. Nothing here establishes that a mid-cap manufacturer or a regional health system faces the same calculus.
It measures market and consumer response, not litigation exposure. Those remain different risks with different remedies, and conflating them is the most common analytical error in this debate.
I state the limits before the argument because the temptation in this moment is to seize a favorable finding and overclaim. Overclaiming is how the business case for inclusion got into trouble in the first place.
II. The Number That Should Stop You
Now to the part that changes how practitioners should read the entire landscape.
To build their narrowest treatment group, Folsz and Grumbach applied a deliberately two-sided test. A firm counted as having maintained its programs only if it showed programmatic continuity — no internal memos or public announcements dissolving the work — and external salience, meaning activists and media had actually identified it as holding the line. Executives also had to have said so publicly.
They explain why they built it that way. Prior research on what accounting scholars call diversity washing established that firms misrepresent their commitments, so statements alone could not be trusted as evidence.
The researchers could not use public posture as a proxy for structure. They knew it would not hold. So they built an instrument that tested both.
Applied across the S&P 500, that test returned eight firms: Apple, Cisco, Costco, Delta Air Lines, Dollar Tree, JPMorgan Chase, Microsoft, and Pfizer.
And a footnote reports something starker still. Searching for publicly traded firms outside the S&P 500 that reaffirmed commitment to their programs after the executive order yielded no results. Not few. None.
The instrument instability
Look at what happens to the population as you change how you define maintaining diversity programs:
Same universe of five hundred firms. Same question. Four instruments, and the answer ranges from eight to fifty-four — a factor of nearly seven.
For the authors’ purposes this variation is a strength: the null result holds across all four, which is exactly the robustness a careful design wants. For practitioners it is something else entirely. It is a measurement of how unstable the category has become. Whether a company counts as having stood firm depends almost entirely on which instrument you happen to pick up.
The public conversation assumes a substantial camp of companies that held the line.
Under the strictest available test — statements corroborated by evidence that the programs survived — that camp contains eight firms out of five hundred, and appears to be empty below the S&P 500 entirely.
The bin most commentary treats as populated is very nearly empty.
III. The Alibi That Just Expired
Consider what many leaders have actually been saying, in boardrooms and in private, across the past eighteen months. Rarely that they had stopped believing in the work. Something closer to: we had no choice. The market required it. Our fiduciary obligation left us no room.
That sentence has now been tested and it does not survive. No detectable penalty in returns. No detectable penalty in revenue — and revenue is where a consumer boycott would have to appear if one had materialized. No penalty even for the firm the President attacked by name.
Whatever was driving those decisions — legal caution, activist pressure, political calculation, exhaustion, the fear of becoming the next headline — it was not the observable price of holding firm.
A decision without a compulsion is a choice. And choices do not live in the finance function. They live in Pillar Two.
This is the architectural consequence, and it is why the study matters more to governance than to communications. It relocates the decision from the domain of necessity to the domain of leadership accountability. That relocation changes who owes an explanation, and to whom.
The supporting evidence points the same direction. Public support for corporate diversity commitments has declined only modestly — roughly seven in ten American adults still consider it important for business to support this work. A 2025 survey found that three-quarters of workers at large firms were more likely to stay if their employer maintained its programs. And the shareholder record is unambiguous: conservative groups placed anti-diversity resolutions at thirty-eight S&P 500 firms, every board recommended rejection, and shareholders rejected them by an average margin of ninety-eight percent.
Owners were not asking for this. Customers were not, on the evidence, punishing it. Employees were telling anyone who asked that it affected their intent to stay.
Where fairness is owed
Some of the caution has been genuine and well founded, and saying so is not a concession — it is a condition of being taken seriously. The legal environment did change. Federal contractor obligations did change. A general counsel who advised narrowing certain race-conscious program designs in 2025 was reading the law, and reading it correctly.
But legal risk is specific. It attaches to particular program mechanics, and it is answerable with redesign. It explains restructuring an eligibility rule. It does not explain withdrawing budget, personnel, and measurement from functions that carried no legal exposure at all. Redesign is a legal response. Withdrawal is a posture.
IV. The Binary Is a Communications Artifact
Six observations from the field, drawn from close work with organizations across this period, describe a reality the two-bin taxonomy cannot hold:
1. Many companies publicly framed as having stood firm have nonetheless made substantial adjustments to their diversity initiatives behind the scenes.
2. Many companies publicly framed as having retreated have nonetheless maintained a great deal of their diversity work.
3. Most companies, regardless of the bin they were sorted into, have significantly changed how they talk about this work — different terminology, less prominence, fewer public-facing references.
4. Many leaders, perhaps most, continue to hold the underlying values. Put them on a polygraph and ask whether they remain committed to fairness and dignity at work, and they would pass.
5. Many of those same leaders have nonetheless cut budgets and personnel or reconfigured programs — some from valid legal concern, some from overreaction, exhaustion, or confusion rather than careful analysis.
6. Many leaders are thinking one step ahead: how do we get through this period? The better ones are thinking further — whether the approach they are adopting now will still function after another cultural or administrative shift.
Until this month those six observations rested on practitioner experience. They now have documentary support, and it comes from inside the study itself.
The stood-firm classification measures disclosure behavior. It does not measure institutional design. It reads the press release and reports the finding as though it had read the architecture — letting hollow institutions collect credit and intact institutions absorb blame, and handing boards a scoreboard that tells them nothing about their own building.
V. The Posture–Structure Matrix
Separate the two variables the binary has fused. Public Posture is what the organization says: declared and visible, or muted and unnamed. Structural Integrity is what the organization has — the systems, budgets, data, accountabilities, and decision rights that actually produce outcomes: intact, or hollowed.
The diagonals are where the analysis lives. Aligned and Withdrawn are honest positions; you can tell what they are from outside. Performative and Submerged are the two that public classification cannot distinguish, because both have decoupled speech from structure — in opposite directions.
The quadrants, documented
Here is what makes this month different. The appendix to the Folsz–Grumbach paper contains case detail on the eight firms in the strictest treatment group, and three of those eight moved after their public reaffirmation. The authors describe them as partial, largely cosmetic retreats. Read as quadrant transitions, they are unusually clean illustrations — and because they are documented in the research record with dates and sources, they can be discussed by name.
Note what this means. Three of the eight firms that the strictest available test identified as having held the line had, within ten weeks, renamed, relabeled, or partially reconfigured the thing they had just publicly defended. None of these movements would be captured by a stood-firm list. All of them are visible to a posture–structure reading.
And note the direction of travel. Two of the three kept the machinery and dropped the language. That is not capitulation. It is submersion — and it is the dominant strategy of this period.
The specific hazard of the Submerged quadrant
Many thoughtful leaders have deliberately chosen submersion: keep the work, drop the label, lower the profile, wait out the weather. It is understandable and often defensible. But it has a failure mode its architects rarely price.
• What is unnamed cannot be governed. Board oversight requires a named object.
• What is unnamed cannot be budgeted. Line items that lose their name lose their defenders in the next cost cycle.
• What is unnamed cannot be inherited. A successor arriving in 2029 cannot maintain a system whose existence is undocumented and whose logic lives in the heads of people who may be gone.
• What is unnamed cannot be defended. When a challenge comes, an organization that cannot articulate what it does and why is not in a stronger legal position. It is in a more confused one.
• And what is unnamed cannot be counted. The eight-firm figure is itself partly an artifact of submersion: organizations that quietly kept their architecture are invisible to any instrument that reads public posture. Some unknown share of the four hundred and ninety-two are Submerged rather than Withdrawn — and we have no way to tell which.
Submersion is a survival strategy, not a design. Strategies expire. Designs are inherited.
A note on what actually protects
One further piece of evidence deserves attention, because it points the same way. Research on firms that suffered public controversies over race and identity found the expected negative effect on returns — but also found that the effect was offset when firms had made more meaningful investments in the underlying work.
Not statements. Investments. The protective factor was structural, and it was measurable. That is the empirical core of the argument this paper is making: posture is what gets scored, and structure is what does the work.
VI. The Eight-Pillar Inspection
Each pillar can be interrogated two ways. The disclosure question is what the public binary can see. The structural question is what actually determines whether people are treated fairly. The gap between the two columns is the diagnostic.
A note on Pillar Six
The 10-K measure in the study is worth pausing on, because it demonstrates the disclosure–structure gap inside a single instrument. Density of diversity language in an annual securities filing is a disclosure variable. It tells you what a company chose to put in front of regulators and investors. It tells you nothing about whether the underlying data is still collected, disaggregated, and reviewed.
Twenty-nine firms cleared that bar. Eight cleared the structural one. The gap between those numbers is the gap this paper is about.
Where the exposure actually sits
Pillar Eight deserves separate emphasis. While the public argument has been consumed by whether the word appears in an annual report, decision authority over hiring, scheduling, evaluation, and monitoring has been migrating into automated systems at speed. Those systems encode selection logic. They produce adverse impact or they do not. They are auditable or they are not.
An organization can retire every inclusion program it has and still be building a disparate-impact liability at machine scale in its applicant tracking, its scheduling algorithm, and its performance analytics. No instrument in the study reaches this pillar — not the filings, not the shareholder votes, not the activist classifications. That is not a criticism of the research design; it is a description of where the measurement frontier currently ends. This is the pillar where the next five years of exposure is accumulating, and it is being ignored because it never had a public posture to defend.
VII. The Reversibility Test
The most strategically serious leaders are already asking a further question: will the approach we are taking now still work if conditions shift again — after 2028, or sooner, or in the other direction? That instinct is correct, and it can be formalized. Four questions constitute the test.
1. Legal durability
Does what remains rest on ground that holds regardless of administration? Adverse-impact monitoring, accommodation processes, harassment prevention, pay-equity analysis, and contract compliance obligations are durable in a way that certain program designs of 2021 were not. Durable ground is not a retreat position. It is a foundation.
2. Reconstruction cost
If posture reverses, what would it cost to rebuild what was dismantled? Some cuts are cheap to reverse — a communications posture, a sponsorship, a webpage. Some are not: a broken longitudinal data series cannot be recovered retroactively, departed expertise does not return on request, and dissolved community relationships take years to rebuild. Track the irreversible ones separately. They are the real balance sheet.
3. Successor legibility
Could a leader arriving in 2029 find, read, and operate the system? Submerged architecture without documentation does not survive a leadership transition. Write down what you kept and why, even if you are not saying it publicly. Internal legibility and external quiet are compatible.
4. Whiplash cost
What does a second reversal cost in employee trust? Each cycle of announce, retreat, and re-announce depreciates the currency. Organizations that have now moved twice should understand that a third move will be read not as responsiveness but as the absence of any conviction at all — and that reading, once it settles, is very difficult to dislodge.
Design for the conditions that have not arrived yet. That sentence is the whole of Pillar One, and it is the difference between a strategy and a flinch.
VIII. Practitioner Instrument: The Posture–Structure Audit
Score each pillar twice, on a scale of 0 to 3. Posture: how visibly and explicitly the organization currently articulates this pillar publicly. Structure: how intact the underlying systems, budgets, data, and accountabilities actually are. Do the structural scoring with evidence, not impression — budget lines, meeting minutes, data dictionaries, compensation plans. This is the same discipline the researchers imposed on themselves, applied inward.
Interpretation
• Structure materially above Posture — Submerged. Apply the Successor Legibility and Reconstruction Cost tests immediately. Document internally what you have chosen not to say externally.
• Posture materially above Structure — Performative. Highest urgency. The gap is visible to your workforce whether or not it is visible to the market, and every additional statement widens it.
• Both high — Aligned. Verify the structural scores independently; self-assessment inflates in this quadrant more than any other. The evidence suggests this position is considerably rarer than organizations believe themselves to occupy it.
• Both low — Withdrawn. Begin with Pillar 4 and Pillar 8, the two with live legal exposure and the shortest path to defensible reconstruction.
One further instruction. Whatever the profile, run the Pillar Eight question this quarter regardless of the total score. Automated decision systems are the one exposure that grows while the organization is looking elsewhere.
IX. Closing
The study did not tell leaders what to do. It removed one of the things they had been saying to explain what they had already done — and then, almost incidentally, it measured how few institutions can demonstrate that they held their ground at all.
Eight firms out of five hundred. None below the S&P 500. And three of those eight quietly moved within ten weeks of saying they would not.
That is not a story about corporate cowardice. It is a story about a landscape in which speech and structure have come apart so thoroughly that we no longer have a reliable public instrument for telling institutions from one another. The researchers ran into that problem and solved it by building a two-sided test. Every board should do the same, pointed at itself.
Because the decisive variable was never the vocabulary. It was always the architecture. Whether the pillars are load-bearing. Whether anyone would know if they were not.
The bins were never the story.
Words can be withdrawn in an afternoon. Architecture takes years to build and years to rebuild, and it is what is standing when the weather changes again.
Build for the season after this one.
I am because we are.
Sources and Notes
• Hanna Folsz and Jacob M. Grumbach, “Markets Do Not Punish Firms for Maintaining DEI,” Democracy Policy Lab, University of California, Berkeley, August 14, 2026. Folsz is a doctoral candidate in political science at Stanford University; Grumbach is an associate professor at the Goldman School of Public Policy, UC Berkeley. Reported by The Guardian the same day. All treatment-group detail, sample sizes, case descriptions, and minimum-detectable-effect figures cited in this paper are drawn from that study and its appendix.
• Executive Order 14173, “Ending Illegal Discrimination and Restoring Merit-Based Opportunity,” January 21, 2025; and the Department of Justice memorandum of February 5, 2025.
• Andrew C. Baker, David F. Larcker, Charles G. McClure, Durgesh Saraph, and Edward M. Watts, “Diversity Washing,” Journal of Accounting Research 62(5), 2024 — the source of the diversity-washing construct underlying the Performative quadrant.
• David F. Larcker, Charles McClure, Shawn X. Shi, and Edward M. Watts, “The Limited Corporate Response to DEI Controversies,” Chicago Booth Research Paper 25-07, 2025 — the finding that controversy effects are offset by more meaningful underlying investment.
• Bentley–Gallup Business in Society Survey, 2022–2025, for public attitudes; Timothy Smith, “Corporate Support for DEI Continues Among Investors and Companies,” Harvard Law School Forum on Corporate Governance, August 2025, for the shareholder record.
• ISO 30415:2021, Human resource management — Diversity and inclusion. Clauses 5–6 (governance and leadership accountability) and Clause 8 (the HR management life cycle) anchor the structural questions in Pillars 2 and 4.
• Effenus Henderson, The Architecture of Inclusion: How Leaders Engineer Equity, Align Their Organizations, and Build Institutions That Last (2026). The eight-pillar framework applied throughout. The Honeysuckle Effect — surface inclusion masking structural rejection — is developed in the author’s prior work and applied here to the Performative quadrant.
A note on the named firms. The three cases in Section V are described exactly as the research record describes them, and are cited here as illustrations of framework categories — not as judgments of intent. Two of the three retained their programs while changing their language, which this paper treats as a defensible short-term strategy carrying a specific long-term cost.
Effenus Henderson is Founder, President & CEO of HenderWorks, Inc., and Convener of ISO/TC 260/WG 8, which produced ISO 30415:2021. He is the author of SPINE: The DEI Backbone for Agility and Adaptability in a VUCA World (2024), Diversity Is Not the Defect — Exclusion Is (2025), and The Architecture of Inclusion (2026).
Disclosure: This paper was drafted by the author and developed in collaboration with an AI writing assistant. The argument, the judgment, and the accountability are the author’s own.







