This is exactly the direction the conversation needs to go. We've spent a lot of time debating the energy cost of a single prompt, but agentic AI changes the unit of analysis. The better question may be: what outcome did that energy produce? Looking at prompts in isolation risks missing the much bigger systems story. Thanks for a thoughtful piece.
“have produced overwhelmingly positive outcomes for us all.” Like climate change? Black lung disease from coal mining? Asthma and heart disease from power plant emissions? Sure, I like my electricity as much as the next person but I don’t kid myself that it doesn’t come with a huge and often negative cost.
I think it's pretty damn clear that many people have suffered from the reliance on certain forms of energy. But I'm not one of those people, so I'll raise my flute of Dom Pérignon Basquiat and nod in agreement.
I think this is actually the same systems problem as the original post, just at a different scale. ‘Overwhelmingly positive’ and ‘huge negative cost’ are both true,they’re just true for different people. Energy access has done enormous good in aggregate, but who bears the cost and who gets the benefit has never been evenly distributed. That’s the part ‘positive outcomes for us all’ tends to flatten. The question isn’t really whether energy sources have net value, it’s whether we’re honest about who’s been paying for that value and who hasn’t.
"The amount of math an AI chip can do per joule of energy has grown roughly 150-fold since 2016, doubling about every two years, per Epoch AI."
I checked your chart and it looks like you're comparing the P100 with FP32 computation to the B300 with FP4 precision. The precision makes a huge difference! It's more computation with higher precision. So a more apples-to-apples comparison would use the same precision level, which would lead to a much lower but still big increase in efficiency from computer hardware. Epoch AI estimates 1.4x efficiency per year, https://epoch.ai/data-insights/ml-hardware-energy-efficiency, so over a decade that is about 30x more efficient.
The less precise computation like FP4 and software improvements can really help with efficiency too as you point out. I really like that you raised the Jevon's paradox. I agree that all the efficiency together is driving more usage.
Thanks for this - this is a really valuable contribution which I will be sharing with others, and broadly agree with your analysis. You are right that both energy-per-token and tokens-per-task are both rapidly moving downwards for both hardware and software reasons - it is a very young technology, and there are still a lot of low hanging fruit for efficiency gains. I suspect that already your estimate will be on the high side, for this reason. Things are moving so quickly.
I would partly disagree on Jevons, though - people are quick to claim Jevons whenever there are efficiency gains alongside absolute increases in energy use, and it can be for other reasons. Jevons, strictly, is that efficiency gains result in reduced use of a resource by a service (in this case electricity and hardware), which reduces the cost of the service, allowing more use of it and more innovative ways to use it.
It is certainly true that AI is being used more, and there is extensive experimentation finding new ways to use it - but it does not appear to me that this is a result of resource cost reduction but simply the natural consequence of a brand-new already cheap-to-provide technology disseminating globally at a historically exceptional rapid pace (via the global internet). The fact that the US companies charge so much more than the Chinese companies shows that energy usage is a relatively small factor in their decision about price setting. Jevons may play some role, but at the moment the main driver of increased energy use is anticipatory investment in data centre capacity for both realistic and hype-driven reasons.
Interestingly, efficiency improvements may actually result in fewer data centres being required even while Jevons results in expanded use. This scenario will happen if most of us (including coders) find that 'good enough' AI meets our needs. Already, laptops are in the pipeline which can run ChatGPT models roughly as powerful as the frontier models of 18 months ago. Some high-value uses of 'super intelligence' (eg cybersecurity) may justify frontier model development, but for most of us what we already have is nearly good enough. In this scenario, your laptop will be powerful enough to produce those axolotl games. It wont be completely guilt free though. At least initially, such laptops will move towards gamer-spec levels of power consumption when at work - so energy use will increase at home.
Thats a fair point. My friend Jon Koomey always likes to point out how projections of exponential energy use in the early days of the internet never panned out because efficiency ultimately outpaced demand. Its just hard to know where things will land this time around!
Excellent comment, Prof. Preist. I appreciate your take on Jevons Paradox, and I'm somewhat relieved to learn how comparatively trivial my own consumption of AI power is. On Dr. Hausfather's bar chart, I'm a sub-Wh LLM user. I for one don't really care where the energy is consumed. It if speeds things up, I'm willing to pay for additional electricity to run GPT engines on my own device. Microsoft is working with OpenAI to embed the technology in laptop hardware, but I'm OK with Linux on my desk and free online GPTs for now!
At age 73, I'm safely retired, but was distrustful of LLMs until last spring, when I experimented with ChatGPT to evaluate my own and others' online comments for fact, logic, and rhetoric. When I ran into OpenAI's free-use limits, I discovered Gemini. As a "free" (with browser ad blocker) Google account owner, I haven't encountered limitations on what I can ask it to do, but then I have no need for agentic interaction like Dr. Hausfather describes, nor for "axolotl games" (Ima steal that!) The first thing I learned is to always check the GPT's sources, because it really does hallucinate and fabricate facts frequently! Yet familiarity makes it less scary, especially when its output is ludicrous 8^D.
In online pro-climate-science fora like this one, I mostly reply to comments made by others. After a protracted formal scientific education and 38 years of amateur obsession with anthropogenic climate change, I'm accustomed to doing it without AI assistance, thank you very much! My new practice is to draft my responses organically, drawing on what I already know, then running both the other's comment and my reply through Gemini's chat interface, like I did with this one (nice job, by the way)! If Gemini tells me something I don't already know, I know enough to ask for its sources, and I often discover it's full of figurative feces! On the other hand, if both Gemini and Grok tell me the same thing, I'm more confident in it.
Unless I make it explicit in my comment, I never use more than a sentence fragment of AI-generated text. I do, however, find Gemini has improved my writing, in this format at least. It points out weaknesses in my organic drafts: I usually have to tell it "I meant to say that", but not always, and it has kept me from making a couple of egregious errors. Hey, somebody should, and it's easier to take criticism from a robot! It's a valuable personal service. Nobody likes making an ass of themselves in public.
Bottom line: I'm as skeptical of "progress" as the next septuagenarian, but it's a brave new world (h/t A. Huxley). I'm intrigued to see how it turns out by 2050!
That's true.. the speed of computation to cost has been accelerating blindingly quick. Wish they'd AI the solution to the current projected water and electrical use. Don't forget.. the more you get from chip brings more demand to do more. A 🐍 going for its tail
This would be a non-issue if our permitting systems let us build unlimited wind and solar (and storage for those). After all, wind and solar are more cost-effective per kWh than natural gas or coal.
Wherever you are, your town will allow you to install solar panels on your roof and a battery in your garage. If you haven't done it, it's not a permitting issue.
If solar/wind + storage is truly more cost-effective than using natural gas, why wouldn’t the hyperscalers building generation behind the meter (meaning no permitting issues, only cost issues) be using those instead? Do they not have accountants?
Behind the meter generation's problem is that you don't have enough land. Permitting reform would mean the power grids themselves build out a lot of solar/wind + storage.
Thanks - fun to see a real analysis by a serious user! Since you touch on the related question toward the end: Google’s 2.7 GW data center being built N of Fort Collins CO promises to use natural gas via fuel cells, then to pump the exhaust stream (presumably pretty pure CO2) into their rock formation for permanent disposal, after using the heat to cool the project via absorption chillers. Is this do-able, and assuming it will actually happen, is there much of a reason (besides well- and delivery-methane leakage which could be drone-monitored and controlled in a sane regulatory environment) to dislike it for green reasons? I write a substack on green energy (including inherently safe small nuclear like Kairos should provide), and think Jade might have a good compromise here. -Bob Meyer
Zeke.. it's also important to remember that direct reductions in CO2 emissions will take none of the CO2 already added out of the atmosphere to potentially lower global temperatures. It does leave carbon in the ground .But more importantly those reductions will increasingly make transportation fuels less and less available as we move forward with the transition to renewables and EVs. In other words we can't even build these AI data centers much less feed eight billion people using electric transportation. What that must mean is that more CO2 will be added and little or none taken out. Because Mauna Loa CO2 correlates almost perfectly with population we can expect atmospheric CO2 to keep rising with no viable way to stop it or remove it. We must move towards innovative infrastructure to help us survive and adapt to extreme weather. And even that will also add CO2 going forward.
Excellent article. I greatly admire your methodology and your commitment to measuring reality instead of relying on simplified assumptions.
Reading your analysis left me thinking that the next bottleneck of AI may no longer be computation itself, but orchestration. As systems become increasingly agentic, the greatest opportunities may lie in coordinating distributed intelligence—efficient synchronization, memory and context management, data access, communication across large GPU fabrics, and system-level coordination.
History repeatedly shows that civilization advances not simply by making individual components more powerful, but by enabling them to work together more intelligently. Perhaps AI is entering that stage as well.
Thank you for a fascinating and thought-provoking piece.
An AI replied that video streaming uses as much computing units as AI, and that videogames use three times more. In a climate emergency we could limit videogames.
"next-generation nuclear and restarts of retired reactors"
----
I see no plausible cost-effective new nuclear reactors in the US and Europe that can come anywhere near what the Chinese can accomplish with their supply chains, purchasing power and established experience (and expertise). And even though the carbon costs (concrete and steel) should be roughly equivalent in both environments, Chinese reactors come online faster, making their "carbon payback time" shorter.
this is really insightful - and as someone who looks into AI energy use a lot, I'm realy glad to see these topics put in this context.
but i want to add - whle you rightly point out we have no accurate idea of true energy use. what we do know is Epoch's estimates only count inference - and both openAI and Anthropic use about 60 per cent of their computre for non inference. and as inference is their only marketable output - we can fairly say the energy use is at least 60% more than those estimates.
currently there is virtually no accurate assessment on embedded carbon - i.e from manufactire of the GPU chips but that is likley to be highly significant - google estimated the footprint of one of its GPUs is 1,100kg - and that was before the latest HBM memory - which has a huge carbon impact.
also 1.1MWh might be only 10 per cent of a US household's consumption, but it is over 30% of a german household - and as there are 115 million US households - that could really start to add up...
i totally agree that the variance in token use tasks makes any estimates misleading - what you point to is that the more you use AI models - the more likely you are to enter token heavy prompts ...so energy use is likely go up far faster than the rise in AI model use - and of course the AI. companies are betting billions on is all becoming heavier and heavier users... - i wonder of there is way to make a simple of graph of that...
It’s impossible to account for everything. For example, you have to subtract the energy saved by the reduction in brewing that would have been used by all the coffee that coders and people like Zeke would have drank to produce a similar amount of intellectual work.
Great 👍 cost/query/ true cost post Zeke!I think they should build AI centers deep or under cool water ( like China in the sea). Also, have their own reservoirs to help cool. The data centers should be autonomous in water usage and electricity. Otherwise, we are f'd.
Or space. If the data centers use so much water for cooling, let’s also put them in Greenland where they can prevent the ice melt from flowing into the sea and rising ocean levels?
Correcting the precision baseline from FP32 to FP4 cuts the headline hardware efficiency narrative from 150x down to a modest 1.4x annual rate. That is a vital distinction. But Jevons' paradox in modern agentic loops runs far deeper than raw FLOPS per joule. 📈
When you crush compute precision down to FP4 to save nanojoules on matrix multiplies, the resulting mantissa truncation noise destabilizes attention logits across extended reasoning chains. The model loses linear identifiability in its planning state. It overthinks, trips over tool calls, and re-reads its own accumulated context 14,000 times for a single task. ⚡
In agentic workflows, 96% of the energy isn't spent doing fresh math in the Tensor Cores. It's spent dragging cached KV tokens out of DRAM across the von Neumann bus just to re-read prior state. Truncating precision saves a fraction of a watt at the ALU, but logit drift forces quadratic memory-context expansion that burns hundreds of watt-hours at the memory interface. ⚙️
Until agentic state latches and motor overrides are phase-locked directly into spatial zero-DRAM SRAM registers, aggressive quantization simply trades cheap arithmetic for expensive memory bus fires. Efficiency without structural memory gating just accelerates total energy collapse. 🤖
Can you compare your Claude energy use with the amount of time it saved you versus doing the work the old-fashioned way>....KWh used per human hour saved?
Agentic AI is dangerous, unreliabe and a vast waste of electricity. Time to call a halt or at least a pause. But that would bring the whole tech sector to a shuddering halt.
Thankyou for a thought provoking analysis which left me relieved that the carbon footprint delta was less than I'd been led to believe through my incidental media exposure. Of course 10% delta for a US data analyst might translate to 20% for their EU counterpart who has a 50% smaller carbon footprint to start with.
The 341gCO2e per kWh is a fair assumption as a starting point. However behind the meter might be 20-25% higher e.g. in the case of Permian shale gas and a simple gas turbine rather than combined cycle.
This is exactly the direction the conversation needs to go. We've spent a lot of time debating the energy cost of a single prompt, but agentic AI changes the unit of analysis. The better question may be: what outcome did that energy produce? Looking at prompts in isolation risks missing the much bigger systems story. Thanks for a thoughtful piece.
all energy sources have produced overwhelmingly positive outcomes for us all. let's not lose sight of it.
“have produced overwhelmingly positive outcomes for us all.” Like climate change? Black lung disease from coal mining? Asthma and heart disease from power plant emissions? Sure, I like my electricity as much as the next person but I don’t kid myself that it doesn’t come with a huge and often negative cost.
I like to point out the abomination that is ethanol corn subsidies in the US, too.
https://www.youtube.com/watch?v=xN1mEMFZlgw
"positive outcomes for us all"
----
I think it's pretty damn clear that many people have suffered from the reliance on certain forms of energy. But I'm not one of those people, so I'll raise my flute of Dom Pérignon Basquiat and nod in agreement.
I think this is actually the same systems problem as the original post, just at a different scale. ‘Overwhelmingly positive’ and ‘huge negative cost’ are both true,they’re just true for different people. Energy access has done enormous good in aggregate, but who bears the cost and who gets the benefit has never been evenly distributed. That’s the part ‘positive outcomes for us all’ tends to flatten. The question isn’t really whether energy sources have net value, it’s whether we’re honest about who’s been paying for that value and who hasn’t.
I was surprised by this quote:
"The amount of math an AI chip can do per joule of energy has grown roughly 150-fold since 2016, doubling about every two years, per Epoch AI."
I checked your chart and it looks like you're comparing the P100 with FP32 computation to the B300 with FP4 precision. The precision makes a huge difference! It's more computation with higher precision. So a more apples-to-apples comparison would use the same precision level, which would lead to a much lower but still big increase in efficiency from computer hardware. Epoch AI estimates 1.4x efficiency per year, https://epoch.ai/data-insights/ml-hardware-energy-efficiency, so over a decade that is about 30x more efficient.
The less precise computation like FP4 and software improvements can really help with efficiency too as you point out. I really like that you raised the Jevon's paradox. I agree that all the efficiency together is driving more usage.
Fair point, though less precise computation is also a form of efficiency (as you note).
Thanks for this - this is a really valuable contribution which I will be sharing with others, and broadly agree with your analysis. You are right that both energy-per-token and tokens-per-task are both rapidly moving downwards for both hardware and software reasons - it is a very young technology, and there are still a lot of low hanging fruit for efficiency gains. I suspect that already your estimate will be on the high side, for this reason. Things are moving so quickly.
I would partly disagree on Jevons, though - people are quick to claim Jevons whenever there are efficiency gains alongside absolute increases in energy use, and it can be for other reasons. Jevons, strictly, is that efficiency gains result in reduced use of a resource by a service (in this case electricity and hardware), which reduces the cost of the service, allowing more use of it and more innovative ways to use it.
It is certainly true that AI is being used more, and there is extensive experimentation finding new ways to use it - but it does not appear to me that this is a result of resource cost reduction but simply the natural consequence of a brand-new already cheap-to-provide technology disseminating globally at a historically exceptional rapid pace (via the global internet). The fact that the US companies charge so much more than the Chinese companies shows that energy usage is a relatively small factor in their decision about price setting. Jevons may play some role, but at the moment the main driver of increased energy use is anticipatory investment in data centre capacity for both realistic and hype-driven reasons.
Interestingly, efficiency improvements may actually result in fewer data centres being required even while Jevons results in expanded use. This scenario will happen if most of us (including coders) find that 'good enough' AI meets our needs. Already, laptops are in the pipeline which can run ChatGPT models roughly as powerful as the frontier models of 18 months ago. Some high-value uses of 'super intelligence' (eg cybersecurity) may justify frontier model development, but for most of us what we already have is nearly good enough. In this scenario, your laptop will be powerful enough to produce those axolotl games. It wont be completely guilt free though. At least initially, such laptops will move towards gamer-spec levels of power consumption when at work - so energy use will increase at home.
Thats a fair point. My friend Jon Koomey always likes to point out how projections of exponential energy use in the early days of the internet never panned out because efficiency ultimately outpaced demand. Its just hard to know where things will land this time around!
Excellent comment, Prof. Preist. I appreciate your take on Jevons Paradox, and I'm somewhat relieved to learn how comparatively trivial my own consumption of AI power is. On Dr. Hausfather's bar chart, I'm a sub-Wh LLM user. I for one don't really care where the energy is consumed. It if speeds things up, I'm willing to pay for additional electricity to run GPT engines on my own device. Microsoft is working with OpenAI to embed the technology in laptop hardware, but I'm OK with Linux on my desk and free online GPTs for now!
At age 73, I'm safely retired, but was distrustful of LLMs until last spring, when I experimented with ChatGPT to evaluate my own and others' online comments for fact, logic, and rhetoric. When I ran into OpenAI's free-use limits, I discovered Gemini. As a "free" (with browser ad blocker) Google account owner, I haven't encountered limitations on what I can ask it to do, but then I have no need for agentic interaction like Dr. Hausfather describes, nor for "axolotl games" (Ima steal that!) The first thing I learned is to always check the GPT's sources, because it really does hallucinate and fabricate facts frequently! Yet familiarity makes it less scary, especially when its output is ludicrous 8^D.
In online pro-climate-science fora like this one, I mostly reply to comments made by others. After a protracted formal scientific education and 38 years of amateur obsession with anthropogenic climate change, I'm accustomed to doing it without AI assistance, thank you very much! My new practice is to draft my responses organically, drawing on what I already know, then running both the other's comment and my reply through Gemini's chat interface, like I did with this one (nice job, by the way)! If Gemini tells me something I don't already know, I know enough to ask for its sources, and I often discover it's full of figurative feces! On the other hand, if both Gemini and Grok tell me the same thing, I'm more confident in it.
Unless I make it explicit in my comment, I never use more than a sentence fragment of AI-generated text. I do, however, find Gemini has improved my writing, in this format at least. It points out weaknesses in my organic drafts: I usually have to tell it "I meant to say that", but not always, and it has kept me from making a couple of egregious errors. Hey, somebody should, and it's easier to take criticism from a robot! It's a valuable personal service. Nobody likes making an ass of themselves in public.
Bottom line: I'm as skeptical of "progress" as the next septuagenarian, but it's a brave new world (h/t A. Huxley). I'm intrigued to see how it turns out by 2050!
That's true.. the speed of computation to cost has been accelerating blindingly quick. Wish they'd AI the solution to the current projected water and electrical use. Don't forget.. the more you get from chip brings more demand to do more. A 🐍 going for its tail
This would be a non-issue if our permitting systems let us build unlimited wind and solar (and storage for those). After all, wind and solar are more cost-effective per kWh than natural gas or coal.
Wherever you are, your town will allow you to install solar panels on your roof and a battery in your garage. If you haven't done it, it's not a permitting issue.
If solar/wind + storage is truly more cost-effective than using natural gas, why wouldn’t the hyperscalers building generation behind the meter (meaning no permitting issues, only cost issues) be using those instead? Do they not have accountants?
Behind the meter generation's problem is that you don't have enough land. Permitting reform would mean the power grids themselves build out a lot of solar/wind + storage.
Cost per kWh isn’t the top consideration for data centers. But uptime is and permitting is, which is why natural gas is more practical.
Thanks - fun to see a real analysis by a serious user! Since you touch on the related question toward the end: Google’s 2.7 GW data center being built N of Fort Collins CO promises to use natural gas via fuel cells, then to pump the exhaust stream (presumably pretty pure CO2) into their rock formation for permanent disposal, after using the heat to cool the project via absorption chillers. Is this do-able, and assuming it will actually happen, is there much of a reason (besides well- and delivery-methane leakage which could be drone-monitored and controlled in a sane regulatory environment) to dislike it for green reasons? I write a substack on green energy (including inherently safe small nuclear like Kairos should provide), and think Jade might have a good compromise here. -Bob Meyer
Zeke.. it's also important to remember that direct reductions in CO2 emissions will take none of the CO2 already added out of the atmosphere to potentially lower global temperatures. It does leave carbon in the ground .But more importantly those reductions will increasingly make transportation fuels less and less available as we move forward with the transition to renewables and EVs. In other words we can't even build these AI data centers much less feed eight billion people using electric transportation. What that must mean is that more CO2 will be added and little or none taken out. Because Mauna Loa CO2 correlates almost perfectly with population we can expect atmospheric CO2 to keep rising with no viable way to stop it or remove it. We must move towards innovative infrastructure to help us survive and adapt to extreme weather. And even that will also add CO2 going forward.
If global population growth is correlated with carbon levels then it will start falling soon.
Ironically, economic growth depends on an increase in the number of stakeholders. So...CO2 will keep rising.
Groan...
Agree. But what else can we do that is realistic?
Excellent article. I greatly admire your methodology and your commitment to measuring reality instead of relying on simplified assumptions.
Reading your analysis left me thinking that the next bottleneck of AI may no longer be computation itself, but orchestration. As systems become increasingly agentic, the greatest opportunities may lie in coordinating distributed intelligence—efficient synchronization, memory and context management, data access, communication across large GPU fabrics, and system-level coordination.
History repeatedly shows that civilization advances not simply by making individual components more powerful, but by enabling them to work together more intelligently. Perhaps AI is entering that stage as well.
Thank you for a fascinating and thought-provoking piece.
An AI replied that video streaming uses as much computing units as AI, and that videogames use three times more. In a climate emergency we could limit videogames.
"next-generation nuclear and restarts of retired reactors"
----
I see no plausible cost-effective new nuclear reactors in the US and Europe that can come anywhere near what the Chinese can accomplish with their supply chains, purchasing power and established experience (and expertise). And even though the carbon costs (concrete and steel) should be roughly equivalent in both environments, Chinese reactors come online faster, making their "carbon payback time" shorter.
"Every one of these methodologies gives an answer between roughly 70 and 330 kilowatt-hours over 8 weeks."
Just for comparison, my EV can drive ~270 miles in good conditions on 80kWh
this is really insightful - and as someone who looks into AI energy use a lot, I'm realy glad to see these topics put in this context.
but i want to add - whle you rightly point out we have no accurate idea of true energy use. what we do know is Epoch's estimates only count inference - and both openAI and Anthropic use about 60 per cent of their computre for non inference. and as inference is their only marketable output - we can fairly say the energy use is at least 60% more than those estimates.
currently there is virtually no accurate assessment on embedded carbon - i.e from manufactire of the GPU chips but that is likley to be highly significant - google estimated the footprint of one of its GPUs is 1,100kg - and that was before the latest HBM memory - which has a huge carbon impact.
also 1.1MWh might be only 10 per cent of a US household's consumption, but it is over 30% of a german household - and as there are 115 million US households - that could really start to add up...
i totally agree that the variance in token use tasks makes any estimates misleading - what you point to is that the more you use AI models - the more likely you are to enter token heavy prompts ...so energy use is likely go up far faster than the rise in AI model use - and of course the AI. companies are betting billions on is all becoming heavier and heavier users... - i wonder of there is way to make a simple of graph of that...
It’s impossible to account for everything. For example, you have to subtract the energy saved by the reduction in brewing that would have been used by all the coffee that coders and people like Zeke would have drank to produce a similar amount of intellectual work.
Great 👍 cost/query/ true cost post Zeke!I think they should build AI centers deep or under cool water ( like China in the sea). Also, have their own reservoirs to help cool. The data centers should be autonomous in water usage and electricity. Otherwise, we are f'd.
Or space. If the data centers use so much water for cooling, let’s also put them in Greenland where they can prevent the ice melt from flowing into the sea and rising ocean levels?
Problem is the heat will melt the ice. We need the AMOCS intact. I was thinking deploying them in some deep shelves.
Correcting the precision baseline from FP32 to FP4 cuts the headline hardware efficiency narrative from 150x down to a modest 1.4x annual rate. That is a vital distinction. But Jevons' paradox in modern agentic loops runs far deeper than raw FLOPS per joule. 📈
When you crush compute precision down to FP4 to save nanojoules on matrix multiplies, the resulting mantissa truncation noise destabilizes attention logits across extended reasoning chains. The model loses linear identifiability in its planning state. It overthinks, trips over tool calls, and re-reads its own accumulated context 14,000 times for a single task. ⚡
In agentic workflows, 96% of the energy isn't spent doing fresh math in the Tensor Cores. It's spent dragging cached KV tokens out of DRAM across the von Neumann bus just to re-read prior state. Truncating precision saves a fraction of a watt at the ALU, but logit drift forces quadratic memory-context expansion that burns hundreds of watt-hours at the memory interface. ⚙️
Until agentic state latches and motor overrides are phase-locked directly into spatial zero-DRAM SRAM registers, aggressive quantization simply trades cheap arithmetic for expensive memory bus fires. Efficiency without structural memory gating just accelerates total energy collapse. 🤖
(⊙_⊙)
Can you compare your Claude energy use with the amount of time it saved you versus doing the work the old-fashioned way>....KWh used per human hour saved?
Agentic AI is dangerous, unreliabe and a vast waste of electricity. Time to call a halt or at least a pause. But that would bring the whole tech sector to a shuddering halt.
Thankyou for a thought provoking analysis which left me relieved that the carbon footprint delta was less than I'd been led to believe through my incidental media exposure. Of course 10% delta for a US data analyst might translate to 20% for their EU counterpart who has a 50% smaller carbon footprint to start with.
The 341gCO2e per kWh is a fair assumption as a starting point. However behind the meter might be 20-25% higher e.g. in the case of Permian shale gas and a simple gas turbine rather than combined cycle.