← Back to Log

Palantir Case Studies · AI-Native Transformation · Supply Chain

Palantir Case Study | 50 Million Decisions, Starting with 3,000 Orders a Night

A source-bounded General Mills case analysis: how Project ELF narrowed an estimated 50 million annual operating decisions into a repeatable loop across more than 3,000 plant-to-warehouse orders each night.

Author: Nick Zhu

General Mills estimated that its operating teams were making about 50 million supply-chain decisions a year. The useful part of the Palantir case is not that all 50 million decisions were automated. They were not.

The company started with a narrower operating loop. Project ELF examined more than 3,000 plant-to-warehouse orders overnight, combined them with current constraints, capacity, and network cost, and surfaced recommendations for disruption or cost opportunities. People still reviewed those recommendations; more than 70% were accepted in the publicly described deployment.

This is a source-bounded case analysis based on General Mills disclosures, Palantir customer materials, and reporting from public interviews. It is not an independent audit. The public record does not disclose ELF’s model mix, objective function, permissions, approval thresholds, contract cost, or the causal contribution of Palantir to General Mills’ wider enterprise savings.

Fifty million decisions is a striking number. It is also a poor implementation scope.

At Palantir’s AIPCon 3 in March 2024, General Mills described a North American supply network with 4,000 suppliers, more than 200 plants, and roughly 1.2 million customer orders a year. The company estimated that operating teams made about 50 million decisions annually across that system, influencing around $10 billion in cost of goods sold as well as service, quality, and greenhouse-gas goals.

The speaker’s ambition was to automate millions of “small but mighty” decisions. But the deployed example did not begin by turning an entire supply chain over to AI. It began with a bounded flow: movements from plants to warehouses.

That distinction is the reason this case is useful. “Transform 50 million decisions” is a strategy statement. “Review more than 3,000 plant-to-warehouse orders every night, identify disruptions and cost opportunities, prepare recommendations, and observe what people accept” is an operating design.

Scope funnel · public figures

The transformation became operable by narrowing the decision field

  1. Enterprise estimate≈50MOperating decisions each year
  2. Demand context≈1.2MCustomer orders each year
  3. Initial ELF loop3,000+Plant-to-warehouse orders examined overnight
  4. Interview exampleUp to 500Recommendations from the nightly assessment
  5. Human response>70%Recommendations accepted in the reported deployment

These figures describe different levels and sources; they are not one conversion funnel or a claim that all annual decisions enter ELF. The “up to 500” figure comes from a 2024 interview with General Mills' chief supply chain officer.

The practical starting point was not 50 million automated decisions. It was one repeatable order flow with a measurable human response.

A connected plan was not enough

General Mills said its journey began in 2019 with a connected data foundation and then a planning system. The sequence matters, but the lesson is not simply “clean all your data before using AI.”

A plan becomes stale as soon as operating conditions change. A labor shortage appears at a plant. Capacity moves. A customer order changes. Weather affects a route. A network that only produces a plan still leaves people reconciling the new reality through email, spreadsheets, and meetings.

Project ELF—short for end-to-end logistics flow—was described as an intelligent execution system built with Palantir. It consumed constraints, capacity, and network cost; examined thousands of orders; and surfaced recommendations where it found a disruption or a potential cost saving.

The important architectural shift, as far as the public evidence supports, was not from “no AI” to “AI.” It was from a periodically optimized plan to a recurring decision loop that could re-evaluate operating state.

Project ELF · source-bounded operating loop

From an optimized plan to a continuously re-evaluated decision

  1. 01Operating statePlant-to-warehouse orders and current network conditions
  2. 02ConstraintsCapacity, commitments, cost, and disruption signals
  3. 03AssessmentExamine more than 3,000 orders overnight
  4. 04RecommendationSurface a disruption response or cost opportunity
  5. 05Human judgmentAccept or do not accept the proposed change
  6. 06Visible impactObserve decisions made—and opportunities not taken

The next cycle starts from changed operating conditions, not from the original plan.

Public sources do not disclose the exact action writeback, rejection workflow, model-training loop, or approval configuration. This figure shows only the operating sequence supported by the customer description.

ELF made a plan revisable through a recurring assessment, recommendation, human-response, and impact loop.

More than 70% acceptance is a trust signal—not an accuracy score

The most interesting number in the case may not be the savings. It may be the reported recommendation acceptance rate.

At AIPCon, General Mills said people accepted more than 70% of ELF’s recommendations. The speaker interpreted this as the machine beating or matching what a person could do much of the time and said the organization was approaching a threshold where some decisions might be turned over directly to the machine.

That is the customer’s interpretation, not a validated measure of model accuracy.

Acceptance can be influenced by recommendation type, user role, available alternatives, financial threshold, confidence, time pressure, and the cost of overriding the system. Public materials do not show the denominator design, performance by decision class, false-positive rate, override outcomes, or whether an accepted recommendation later produced the expected result.

Still, an acceptance measure is operationally valuable. It records behavior at the point where a recommendation meets real work. It can reveal where the system has earned enough trust for greater delegation—and where human authority should remain.

Metric boundary · what 70% can and cannot mean

Acceptance, correctness, and autonomy are three different claims

Reported

Recommendation acceptance

>70%

People accepted more than 70% of recommendations in the described deployment.

Not disclosed

Outcome correctness

?

No public decision-class accuracy, counterfactual, or realized-outcome analysis.

Stated direction

Direct machine action

A future threshold was discussed; the public case did not show universal autonomy.

A recommendation can be accepted without being provably optimal, and a strong acceptance rate does not by itself define which actions may safely become autonomous.

The useful governance question is not “Is acceptance high?” but “For which decision class, under what limits, with what outcome evidence?”

The data foundation is real; its published size is inconsistent

Both General Mills and Palantir emphasize that the execution system depended on earlier connected-data work. But even the most repeated foundation number requires care.

In the official AIPCon video, the General Mills speaker said the company moved and connected 2,000 master and operational data tables in the cloud. Palantir’s later one-page impact study says 200 tables were integrated on the Palantir Ontology.

The sources may describe different subsets, or one may contain an error. Neither source explains the discrepancy, so I would not silently choose one number.

What is consistently supported is more important than the exact table count: the partnership began in 2019; General Mills connected previously fragmented master and operational data; the company treated that foundation as a single source of truth; and leaders credited it with accelerating later use cases.

What is not public is the customer-specific Ontology: its object types, properties, links, actions, functions, data-quality controls, and permissions. Palantir’s documentation explains what the platform can support, but platform capability should not be rewritten as General Mills’ disclosed configuration.

Source conflict · preserve the uncertainty

The public record agrees on the foundation, not the table count

AIPCon video · General Mills speaker

2,000 tables

Master and operational data tables moved to the cloud and connected.

Palantir impact study

200 tables

Master and operational data tables integrated on the Ontology.

Consistent across sources2019 start · connected data foundation · single source of truth · later use cases accelerated

The public sources do not reconcile 200 and 2,000. This analysis therefore treats the exact number as unresolved and does not infer General Mills' Ontology schema from generic product documentation.

Source discipline sometimes means keeping a conflict visible instead of manufacturing false precision.

The value claim evolved as the operating scope expanded

In March 2024, General Mills reported average savings of about $40,000 a day, or roughly $14 million annually, while ELF was deployed to only part of the network. Palantir repeats that customer statement in its impact materials.

The claim is meaningful, but bounded. Public sources do not disclose the baseline, calculation method, project cost, avoided-cost rules, or an independent audit. The savings should be attributed as a General Mills customer statement—not presented as a guaranteed platform result.

Later company disclosures show a broader operating evolution. In its fiscal 2024 results, General Mills said digital capabilities had reduced waste by 20% on manufacturing lines at some of its largest sites and that logistics optimization was expected to remove one million road miles annually. Those were company-wide digital initiatives; the disclosure did not attribute all of them to Palantir or ELF.

At General Mills’ October 2025 Investor Day, Paul Gallagher described ELF connecting system-to-system with a major retailer. He said the flow had removed 15,000 tons of carbon by reducing trucks and compressed the work of optimizing orders into truckloads from 18 hours to under 30 minutes. These are later customer-reported results for the evolving program, not figures from the original 2024 partial deployment.

In July 2026, General Mills announced a target of $3 billion in cumulative cost savings through fiscal 2030. Roughly $2 billion was assigned to its long-running Holistic Margin Management program; the remaining $1 billion was assigned to global transformation and other efficiency efforts, including supply-chain redesign and process streamlining. The company did not say that Palantir, ELF, or AI would deliver the full $3 billion.

Evolution · keep unlike claims separate

A bounded use case expanded; the enterprise savings target is a different claim

  1. 2019Connected foundationFragmented master and operating data connected
  2. 2023–24ELF decision loop3,000+ nightly orders; recommendations still reviewed by people
  3. Mar 2024Partial deploymentCustomer reported ≈$40K/day, ≈$14M annualized
  4. Oct 2025System-to-system flowCustomer reported 18h→<30m and 15K tons of carbon removed
  5. FY27–30Enterprise target$3B across HMM, transformation, and other efficiency programs

The 2026 $3 billion target is a forward-looking General Mills enterprise program. Public disclosures do not attribute it wholly—or by a disclosed share—to Palantir, ELF, or AI.

The case matured from a bounded logistics loop into broader connected operations, while enterprise transformation remained larger than any one platform or use case.

Human-in-the-loop was an operating state, not a slogan

The 2024 presentation was unusually direct about the destination: recommendations went to people “today,” while the future goal was to take them out of the loop for suitable decisions.

I would not turn that statement into a general rule that human review is merely temporary. It describes one company’s direction for high-volume supply-chain decisions. The right boundary depends on decision consequence, reversibility, confidence, contractual commitments, control requirements, and the evidence available after execution.

The more important clue is what General Mills said it was building around the system: standardized adoption metrics, upskilling and reskilling, change champions, and visibility into the impact of decisions made and decisions not made.

That is not a software feature list. It is an operating model for earning delegation.

A system can begin with recommendations and human review. As evidence accumulates, some decision classes may move to lighter review or bounded automatic action. Exceptions, high-consequence commitments, and unclear states can stay with identifiable human owners. The boundary should move because measured evidence supports it—not because an autonomy roadmap demands it.

Delegation ladder · Nick's operating interpretation

Earn autonomy one decision class at a time

  1. 01ObserveSystem surfaces state and opportunity; person decidesEvidence: can the state be trusted?
  2. 02RecommendSystem proposes a bounded change; person reviewsEvidence: acceptance and override reasons
  3. 03Act with approvalApproved recommendation changes the operating systemEvidence: realized outcome and recoverability
  4. 04Act within limitsLow-consequence classes execute under explicit boundariesEvidence: exceptions, drift, and control performance
  5. Human end stateOwn the boundaryPeople define intent, limits, escalation, override, and accountabilityNot every decision needs the same destination

This ladder is my transferable framework, not General Mills' disclosed approval design or a maturity model requiring every decision to become autonomous.

The goal is not maximum automation. It is the smallest safe human role for each decision class, supported by inspectable evidence.

What the public case supports—and what remains undisclosed

The strongest evidence is the shape of the operating loop and the customer-reported deployment figures. General Mills publicly described the decision scale, the nightly order scope, the inputs, the recommendations, the human acceptance rate, the partial-deployment savings, and later logistics outcomes.

The evidence becomes weaker when case-study metrics are converted into causal or universal claims. Public materials do not allow us to calculate Palantir’s standalone ROI, attribute General Mills’ wider cost programs to one platform, or prove that a 70% accepted recommendation was correct 70% of the time.

The exact implementation is also private. We do not know the customer-specific Ontology, models, optimization objective, data refresh by field, approval rules, action permissions, override mechanisms, incident recovery, or commercial terms.

That boundary does not make the case less valuable. It locates the value in what can actually be transferred: a large transformation ambition became concrete when the company selected a repeated decision flow, connected the minimum operating context, placed recommendations in front of real owners, recorded adoption, and exposed impact.

Build one Decision Loop Record

For a high-frequency operating decision in your own organization, capture:

  1. Outcome: What business state should this decision protect or improve?
  2. Decision class: What repeated choice is being made—not the entire process or department?
  3. Frequency and volume: How often does it occur, and how many instances enter each cycle?
  4. Operating scope: Which orders, products, locations, customers, or time window are included?
  5. Authoritative state: Which current facts and timestamps can be trusted?
  6. Constraints and options: What limits the decision, and which actions are permitted?
  7. Human authority: Who may accept, change, execute, override, and own the result?
  8. Evidence: How will acceptance, override, realized impact, failure, and opportunity cost be measured?
  9. Delegation rule: What evidence would justify changing the human role for this decision class?

This record is not a Palantir implementation specification or an instruction to build an enterprise-wide Ontology first. Its purpose is to convert an abstract automation ambition into one bounded decision loop that can be tested.

The previous Palantir case on Wendy’s QSCC followed one shortage from exception to placed orders. The AI-Native Operating Model explains how outcomes, decisions, context, actions, evidence, exceptions, and accountability fit into a Work Graph. General Mills adds another layer: how a company can use recommendation adoption and observed impact to move the human-machine boundary deliberately over time.

If a team cannot yet define the decision class, evidence, and authority boundary internally, FDE Delta Operating Partnership can be a restrained next step after the business scope is clear. It is not an offer of Palantir implementation, system selection, certified architecture, or guaranteed savings.

Sources and method

The primary case sources are General Mills’ AIPCon 3 presentation, Palantir’s General Mills impact study, General Mills’ fiscal 2024 fourth-quarter remarks, its 2025 Investor Day materials, and its July 2026 savings announcement. I used Supply Chain Dive’s interview with Paul Gallagher and CFO Brew’s interview with Dave Jackett as secondary operational context.

Public sources are mostly company, customer, vendor, and event materials. I retain attribution for performance statements, keep the 200-versus-2,000 table conflict visible, separate the 2024 partial-deployment claim from later results, and do not attribute General Mills’ $3 billion enterprise target to Palantir. All six diagrams are my source-bounded analytical views, not customer-supplied architecture or governance diagrams.

OPEN A CONVERSATION

Build something worth testing.

For AI-native products, global GTM, or independent projects, choose a channel below.

WECHAT

Scan to connect on WeChat

Scan to connect on WeChat