AI Emissions Are Not One Number: A Better Way to Measure Corporate AI Use

Companies face a choice: wait for AI providers to publish perfect emissions data, or build a transparent estimate with what they have today. The first option avoids uncomfortable assumptions. It also leaves a rapidly growing source of electricity demand outside the management system.

That is not only a measurement problem. It is a governance decision.

Data centers accounted for roughly 5% of US electricity consumption in 2025. According to EPRI, this share could reach between 9% and 17% by 2030. Corporate AI use is expanding at the same time - through employee assistants, developer tools, application programming interfaces, and AI features embedded in software that procurement teams may not even classify as AI.

For many enterprises, the associated emissions remain small relative to the total corporate footprint. That may not last. More importantly, the systems needed to understand those emissions cannot be assembled retrospectively with much confidence.

Watershed's new framework offers a useful response. It does not promise a single, universally correct emissions factor. It proposes a method that can begin with limited information, disclose its assumptions, and become more precise as provider data improve.

Why AI emissions estimates disagree so dramatically

The popular question is: how much carbon does an AI query produce? It sounds straightforward. It is not.

Different estimates count different things. Some measure only the electricity used by an accelerator while it processes a request. Others include the host system, cooling and power conversion, idle capacity, model training, hardware manufacturing and supporting data-center infrastructure. Change the boundary and the answer changes with it.

The workload matters just as much. A short classification task, an image-generation request, and a multi-step agentic workflow are not comparable units of activity. The white paper finds electricity use varying by more than two orders of magnitude across common workload types. One apparent interaction can also trigger dozens of model calls through tool use, routing and self-correction.

Then there is provider efficiency. Production systems use batching, caching, specialized hardware and software optimization that external benchmarks often cannot observe. The research reviewed by Watershed suggests that non-production benchmarks can overstate inference energy by a factor of four to 20. Yet a narrow benchmark may also omit idle capacity, facility overhead or embodied emissions and understate the total boundary.

Both errors are possible. That is precisely the point. The problem is not that every estimate is useless; it is treating an estimate as comparable before checking its boundary, workload and data source.

A hypothetical company, three very different answers

Watershed illustrates the issue with a hypothetical financial-services company spending $100,000 a year on a frontier model for internal tools and a customer chatbot. Finance sees one supplier line. Engineering sees token volumes and application traffic. Sustainability sees an incomplete Scope 3 category.

Using a generic spend-based factor, the company's annual AI footprint is estimated at 13.4 tCO2e. With activity data, the estimate falls to between 3.7 and 5.4 tCO2e. Using provider-level information, it falls between 3.2 and 4.5 tCO2e.

The same use. A roughly fourfold spread.

This does not mean that spend-based methods always overstate AI emissions. A subsidized service, an inefficient model, or a carbon-intensive grid could reverse the result. The lesson is narrower and more useful: pricing is a weak proxy for compute, and compute is not enough without location and infrastructure data.

Three tiers, one direction of travel

The proposed framework recognizes that companies do not all have the same telemetry. Instead of requiring false precision from day one, it establishes three measurement tiers:

1. Spend Tier. AI-related expenditure is multiplied by an economic emissions factor for data processing and hosting. It is accessible and auditable, but it reflects an industry average rather than the model, workload, or electricity actually used. It works as a provisional backstop.

2. Activity Tier. Token volumes are combined with estimated energy intensity, data-center overhead, and grid carbon factors. A more detailed version separates input from output tokens, identifies the serving region, and adds estimated training and embodied emissions. This is materially more informative, but still depends on assumptions.

3. Provider Tier. The model or cloud provider supplies customer-relevant energy and emissions factors. In principle, this is the most precise tier because it can reflect the actual serving infrastructure, grid mix, and efficiency of the workload. In practice, the required disclosure is still rare.

The sensible direction is upward through the tiers as data become available. A company may also use different tiers in the same inventory: provider data for one API, activity estimates for another, and spend as a fallback for an embedded software assistant.

That is not methodological inconsistency. It is an honest description of current data quality - provided each calculation is labeled clearly.

Why tokens are useful - and imperfect

Watershed proposes kilograms of CO2 equivalent per million tokens as the functional unit for inference. Tokens are already recorded for billing and cost control, which makes them more operationally useful than a sustainability-only metric introduced at year-end.

The alignment matters. Reducing unnecessary context, caching repeated inputs, and routing simple tasks to smaller models can lower cost and electricity consumption at the same time.

Tokens are not a perfect unit. Tokenizers differ across models. Input and output tokens have different energy profiles: output generation is sequential and generally more energy-intensive. Long contexts do not always scale linearly. Nor does a token measure the quality or business value of the result. A system could become more "efficient" per token while producing more low-value output overall.

Still, the token is a practical bridge between carbon accounting and AI operations. The alternative - counting vaguely defined "queries" - often hides more than it reveals.

Inference is actionable; training remains uncertain

Training large models attracts attention because it creates a visible, upfront carbon cost. Allocating that cost to corporate users is considerably harder. Providers rarely disclose the model's total development compute, the emissions from unsuccessful runs, or the number of tokens the model will serve over its lifetime.

That denominator changes the answer dramatically. In Watershed's illustrative case, training represents between 32% and 56% of the footprint under the central assumptions. Across alternative lifetime-use assumptions, its share ranges from 5% to 83%.

This uncertainty should not be averaged away. The framework recommends reporting training as a separate line item, with the allocation method disclosed, rather than burying it inside an apparently precise per-token figure.

For widely deployed models, inference electricity may be the more useful near-term priority. It scales with corporate use and connects directly to decisions about model selection, prompt design, serving region, and electricity procurement. Training should remain inside the boundary. It simply should not prevent companies from improving the part they can measure and influence now.

Measurement should change decisions

A carbon inventory earns its keep when it changes a choice. The framework identifies two channels for doing that. The first is reducing electricity per task. Shorter context windows, fewer redundant calls, and task-aware model routing can all help. The white paper cites evidence that reasoning models can be roughly 30 times more energy-intensive than smaller production variants for equivalent queries. Reserving frontier capability for tasks that genuinely require it may therefore improve both cost and carbon performance.

The second is reducing emissions per unit of electricity. Where customers can choose a serving region, the carbon intensity of available US grid subregions can vary by more than fivefold. Provider efficiency and credible clean-electricity procurement matter even more, although customers can influence these mainly through vendor selection and contract requirements.

These choices involve trade-offs. A smaller model may reduce energy consumption but fail a quality threshold. A lower-carbon region may increase latency or create data-residency constraints. Additional telemetry may improve the inventory while increasing engineering effort. The framework does not eliminate those tensions. It makes them visible.

What a defensible first year could look like

For companies beginning this work, a credible sequence is relatively compact:

1. Map AI use by access channel. Hosted assistants, developer tools, direct APIs, cloud-mediated APIs, and self-hosted models expose different data. The inventory should reflect that difference rather than relying only on the vendor list.

2. Use the best-supported tier for each material source. Spend can cover blind spots. Token and regional data can upgrade important workloads. Provider figures should replace assumptions when they are relevant to the customer's actual model and traffic.

3. Ask providers for decision-useful information. Model family, energy per token, operational emissions, serving region, and reporting period are the essential starting fields. Training, embodied emissions, power-usage effectiveness, and third-party assurance can improve the picture further.

4. Report electricity and emissions separately. Electricity shows whether operational efficiency is improving. Emissions show the effect of the electricity mix and clean-energy procurement. Combining them into one figure can conceal which lever produced the change.

The supporting governance is equally important: named data owners, documented assumptions, version-controlled factors and a trigger for moving a source to a higher tier. Otherwise, the first estimate becomes a permanent workaround. We have all seen how that ends.

The better question

AI emissions may not yet be material for every company. But materiality can change faster than an annual reporting cycle, particularly as embedded tools and agentic workflows multiply activity that procurement data do not reveal.

The objective is not to manufacture certainty. It is to establish a measurement system that shows where uncertainty sits, improves when better evidence arrives, and connects the result to real operating decisions. That leaves senior teams with more useful questions:

- Which AI channels are visible, and which remain hidden inside software contracts?

- Which assumptions drive the reported footprint?

- Which model, workload, and region choices affect both cost and emissions?

- Who is responsible for upgrading the method when better data becomes available?

Those questions turn AI emissions from a headline into a decision system. That is a better place to start.


Next
Next

Europe’s Wildfire Reality: Resilience Is No Longer Optional