Discussion about this post

User's avatar
Mehar U Nisa, PhD's avatar

This is exactly the direction the conversation needs to go. We've spent a lot of time debating the energy cost of a single prompt, but agentic AI changes the unit of analysis. The better question may be: what outcome did that energy produce? Looking at prompts in isolation risks missing the much bigger systems story. Thanks for a thoughtful piece.

David Ruffner's avatar

I was surprised by this quote:

"The amount of math an AI chip can do per joule of energy has grown roughly 150-fold since 2016, doubling about every two years, per Epoch AI."

I checked your chart and it looks like you're comparing the P100 with FP32 computation to the B300 with FP4 precision. The precision makes a huge difference! It's more computation with higher precision. So a more apples-to-apples comparison would use the same precision level, which would lead to a much lower but still big increase in efficiency from computer hardware. Epoch AI estimates 1.4x efficiency per year, https://epoch.ai/data-insights/ml-hardware-energy-efficiency, so over a decade that is about 30x more efficient.

The less precise computation like FP4 and software improvements can really help with efficiency too as you point out. I really like that you raised the Jevon's paradox. I agree that all the efficiency together is driving more usage.

32 more comments...

No posts

Ready for more?