Two philosophies are fighting over how artificial intelligence should be built — one chases scale at any cost, the other asks what that cost actually is
Ujjwal K Chowdhury
Strapline: For a decade, AI research had one scoreboard: accuracy. A new one is forcing its way onto the field — energy, water, carbon and hardware. The contest between “Red AI” and “Green AI” is no longer academic; it is shaping how the world’s most powerful technology gets built.
The paper that named the problem
In 2020, a small group of computer scientists — Roy Schwartz, Jesse Dodge, Noah A. Smith and Oren Etzioni — published a short, blunt paper in the Communications of the ACM with a title that stuck: “Green AI.” It drew a line through the field. On one side sat what the authors called Red AI: research that chases state-of-the-art results by throwing ever more computation at a problem, treating accuracy as the only currency that matters. On the other side stood Green AI: research that treats efficiency — the resources spent per unit of result — as a first-class scientific goal, not an afterthought.
The label was provocative on purpose. Red AI was not, the authors were careful to say, morally wrong. It had produced genuine breakthroughs. But it had also quietly normalised an arms race in which each new record-setting model consumed dramatically more compute than the last, with the environmental bill rarely itemised in the paper’s appendix, let alone its abstract.
Six years on, that argument reads less like a provocation and more like a prophecy. Generative and agentic AI systems now sit inside search engines, office software, customer service lines and increasingly autonomous workflows that plan, browse, code and retry without a human in the loop. The scoreboard Schwartz and colleagues warned about has expanded from leaderboard rankings to gigawatts, litres and tonnes of carbon dioxide.
Two philosophies, one industry
Red AI, at its core, is a bet that more computation reliably buys more capability — bigger models, longer training runs, wider search over architectures, more parameters, more data, more reasoning steps at inference time. It is the logic behind scaling laws, and it has worked spectacularly well as a research strategy. But it has a hidden accounting problem: the “winning” run reported in a paper or press release is usually just the tip of an iceberg of failed experiments, architecture searches, ablations and evaluation runs that never make it into the final number. Recent lifecycle research — including a 2025 study led by Jacob Morrison that traced the full environmental cost of building a language-model family — found that model development contributed roughly half of the total training-related impact, not the celebrated final run alone.

Green AI, by contrast, asks a different question of every architectural choice, every training run and every product feature: what is the smallest, most efficient way to achieve an acceptable outcome? It treats efficiency — measured in floating-point operations, energy, water and, increasingly, successful outcomes per unit of resource — as an evaluation criterion sitting alongside accuracy, not subordinate to it.
Crucially, Green AI has matured past its original, somewhat narrow framing. In 2020 it was largely about training compute. Today, researchers describe it as the quality- and outcome-constrained minimisation of lifecycle environmental impact — a formulation that captures something Red-versus-Green rhetoric can miss: a computationally hungry model is not automatically the villain, and a lean one is not automatically virtuous. A large model solving a genuinely high-value problem in a handful of steps can outperform, environmentally, a small model that fails repeatedly and triggers costly retries. The real dividing line is not model size; it is whether computation is productive.
Why the contest matters now
The urgency comes from scale. According to the International Energy Agency’s most recent assessment, global data-centre electricity demand rose roughly 17% in 2025 alone — more than five times faster than overall global electricity growth — while electricity consumption specifically tied to AI-focused facilities surged around 50% in the same year. The IEA’s satellite-tracking programme, which watches construction of dedicated “AI factories” from orbit, found that their combined capacity has more than tripled in eighteen months. Data-centre electricity use worldwide, which stood at roughly 415–485 TWh depending on the estimate and year, is on a trajectory toward roughly 950 TWh to beyond 1,000 TWh by 2030 — comparable to the entire annual electricity consumption of Japan.
FAST FACTS > - Global data-centre electricity demand: ~485 TWh in 2025, heading toward ~950 TWh by 2030 (IEA) > - AI-focused data-centre demand: up ~50% in 2025 alone > - US data-centre share of national electricity: 4.4% in 2023, projected 6.7–12% by 2028 (LBNL) > - AI-rack power density: up roughly elevenfold, 2020–2025 (IEA) > - Ireland’s data centres already draw over a fifth of national electricity; Dublin’s local share runs close to 80%
This is precisely the terrain Red AI was warned about: growth compounding on growth, with local grids in Ireland, Northern Virginia and parts of the Netherlands already straining, and utilities in the United States requesting billions of dollars in rate increases partly attributable to data-centre load growth. Energy-policy academics have begun asking, pointedly, whether ordinary electricity customers should effectively subsidise the power appetite of trillion-dollar technology companies — a question with no comfortable answer for regulators.
Where the two camps actually clash
The Red AI/Green AI split is not simply “big model bad, small model good.” It shows up in concrete engineering and business decisions:
1. Model selection. Red-style practice defaults to the most capable, largest available model for every task, regardless of whether the task warrants it. Green practice builds a portfolio: small or domain-specific models for routine work, escalating to frontier models only when complexity demands it. Systems such as FrugalGPT, which learned to route easy queries to cheaper models and reserve expensive ones for hard cases, demonstrated cost reductions of up to 98% on selected benchmarks without materially sacrificing quality.
2. Reporting practice. Red AI habitually reports only the final training run’s cost. Green AI insists on lifecycle transparency — development experimentation, fine-tuning, evaluation, and the electricity, water and embodied-hardware cost of years of subsequent inference, which can dwarf the original training bill many times over.
3. Agentic design. This is the newest and sharpest fault line. An autonomous agent can quietly multiply a single user request into dozens or hundreds of model calls, tool invocations, retries and multi-agent “debates.” Early benchmark research has found up to a 9.4-fold energy difference between agent-framework designs solving the same software-engineering tasks, driven mostly by wasted loops and redundant verification. A 2026 preprint proposing a metric called Energy per Successful Goal (EpG) found that agentic workflows consumed, on average, 4.33 times more energy per completed goal than equivalent linear, non-agentic approaches. Red AI treats agent autonomy as an unqualified upgrade; Green AI treats it as a resource-management problem requiring budgets, loop detection and outcome-based evaluation.

4. The rebound trap. Perhaps the most uncomfortable insight from Green AI research is that efficiency gains alone do not guarantee lower total impact. If a model becomes twice as cheap to run, organisations often respond by running it far more than twice as often — generating more content, running more experiments, automating tasks nobody previously bothered to automate. This is a version of the century-old Jevons paradox, in which efficiency improvements in coal-fired steam engines led, historically, to more coal consumption, not less, because cheaper power expanded its uses. Green AI researchers now argue that intensity metrics (energy per task) must be paired with absolute-impact accounting (total annual energy, water and carbon) precisely to catch this rebound before it erases hard-won efficiency gains.
The measurement mess neither side can ignore
Part of what makes the Red/Green debate so combustible is that reliable, comparable numbers are still scarce. A landmark 2025 measurement of Google’s production systems found a median energy cost of just 0.24 watt-hours and 0.26 millilitres of water per text prompt — a strikingly small figure. Around the same time, a separate academic benchmark estimated that complex, long-context reasoning queries on certain models could consume more than 33 watt-hours — over a hundred times more. Both figures are credible within their own scope; they simply describe different systems, different tasks and different accounting boundaries. A 2025 peer-reviewed review of data-centre water use went further, finding that water consumption per workload can vary by more than 10,000-fold depending on cooling technology, grid water intensity, climate and utilisation.
This is why serious Green AI researchers are wary of single, universal “footprint per query” numbers circulating in the media — they tend to flatten an extraordinarily heterogeneous reality into a misleadingly precise soundbite. The more defensible approach, gaining traction in both research and emerging regulation such as the European Union’s data-centre reporting rules, is a layered hierarchy: from raw activity counts (tokens, model calls), up through compute energy, facility-adjusted energy, environmental impact (carbon and water, adjusted for time and place), full lifecycle impact including embodied hardware emissions, and finally outcome-normalised impact — energy and water per successfully completed task, not per token generated.
Not a morality play — a design discipline
It would be easy, and wrong, to read Red AI and Green AI as heroes and villains. Some of the most consequential AI applications — climate modelling, grid forecasting, drug discovery, materials science for batteries and solar cells — are legitimately compute-intensive, and restricting them to “small and frugal” would forfeit real value. The IEA itself estimates that mature AI applications could trim energy costs across several industries by 3 to 10 percentage points, and Google has reported enabling tens of millions of tonnes of avoided CO2-equivalent emissions through AI-optimised products in a single year. Green AI’s actual claim is narrower and more rigorous: that value should be measured against lifecycle cost, that claims of benefit require credible counterfactual evidence, and that scale should be earned by demonstrated necessity rather than assumed by default.
HIGHLIGHT > “A Green AI system is not simply smaller or faster. It is appropriately capable, transparently measured, powered and cooled responsibly, designed to avoid waste, and deployed where its verified value exceeds its environmental cost.”
What comes next
Expect the Red/Green fault line to move from academic papers into contracts and regulation. Procurement teams are beginning to demand model-level energy and water disclosures before signing cloud contracts. The EU’s AI Act ecosystem is developing standards for reporting the resource performance of general-purpose AI systems. Enterprises are experimenting with model-routing rules that default to the smallest sufficient model rather than the flashiest one. And a growing chorus of researchers argues that the next frontier metric will not be accuracy, or even energy per token, but energy per successful goal — a number that punishes both wasteful agents and models that fail so often they need constant escalation.

The Red AI era was not a mistake; it built the models the world now depends on. But the bill for that approach is now visible in gigawatts, litres and rising electricity tariffs, and it is arriving at a moment when climate constraints leave little room for waste. Green AI’s proposition is simple, if not easy: intelligence, at any scale, should have to justify its keep.
Reading the two camps side by side
| Red AI | Green AI | |
|---|---|---|
| Core metric | Accuracy / benchmark score | Quality-adjusted efficiency (energy, water, carbon per successful task) |
| Model choice | Biggest available, by default | Smallest sufficient model, escalate only when needed |
| Reporting | Final training run only | Full lifecycle: development, training, inference, hardware |
| Agents | Autonomy as unqualified upgrade | Autonomy as a budgeted, monitored resource |
| Risk | Rebound erases efficiency gains | Absolute-impact caps alongside intensity targets |
Framed this way, the contest is less a war between two tribes of researchers than a description of a choice every AI-building organisation now has to make, explicitly or by default, every time it ships a feature. The instinctive path — reach for the largest available model, let an agent iterate until it seems to have solved the problem, publish the headline benchmark and move on — is Red AI, whether or not anyone in the room uses the term. The alternative requires more upfront engineering discipline: measuring what a task actually needs, instrumenting the full resource cost, and being willing to report a less flattering number if that is the honest one.
Neither side of the debate disputes that AI can create enormous value. The disagreement is about method — whether that value is pursued by default at maximum scale, or earned deliberately at the scale a task actually requires. As electricity bills, water permits and carbon disclosures increasingly follow AI systems out of the lab and into public scrutiny, that distinction is starting to carry real financial and regulatory weight, not just scientific interest.
Add a Comment