Verify any claim · lenz.io
Claim analyzed
Tech“Generating a single image with generative artificial intelligence consumes more than ten times as much energy as processing a text prompt.”
Submitted by Bright Sparrow 1b63
The conclusion
Open in workbench →Available measurements generally support a greater-than-tenfold energy difference between AI image and text generation, with some studies reporting much larger ratios. However, the multiplier is not universal: model choice, hardware, output settings, and the definition of a text request can produce lower comparisons. The statement is therefore broadly accurate but too categorical.
Caveats
- The greater-than-tenfold ratio is not guaranteed for every model, configuration, or prompt.
- Comparisons may measure complete text responses rather than only processing the input prompt.
- Several supporting studies are preprints, and estimates use differing hardware and measurement boundaries.
Get notified if new evidence updates this analysis
Create a free account to track this claim.
Sources
Sources used in the analysis
In fact, generating an image using a powerful AI model takes as much energy as fully charging your smartphone, according to a new study by researchers at the AI startup Hugging Face and Carnegie Mellon University. However, they found that using an AI model to generate text is significantly less energy-intensive. Creating text 1,000 times only uses as much energy as 16% of a full smartphone charge.
Generating visual material emits more CO2 than other AI-based tasks, with image generation consuming over 30 times as much energy as generating text (Luccioni et al., 2024).
The result: depending on the configuration, a single image can consume up to ten times more energy than an average ChatGPT request, which according to OpenAI CEO Sam Altman requires about 0.34 watt-hours.
This report was also strictly limited to text prompts, so it doesn’t represent what’s needed to generate an image or a video. (Other analyses, including one in MIT Technology Review’s Power Hungry series earlier this year, show that these tasks can require much more energy.)
The greatest amount of energy was expended by Stability AI's Stable Diffusion XL, an image generator. Nearly 1,600 grams of carbon dioxide is produced during such a session. … On the lowest end of the scale, basic text generation tasks expended the equivalent of a car driving just 3/500 of a mile.
We identify orders-of-magnitude differences in the amount of energy required per inference across models, modalities and tasks and shine light on an important trade-off between the benefit of multi-purpose systems, their energy cost, and ensuing carbon emissions.
We identify orders-of-magnitude differences in the amount of energy required per inference across models, modalities and tasks and shine light on an important trade-off between the benefit of multi-purpose systems, their energy cost, and ensuing carbon emissions.
By some estimates, text classification might consume 0.002 kWh per thousand queries, while image generation can demand 2.9 kWh for the same number—a 1,450-fold difference. So when you ask AI to create a picture instead of answering a text question, you're potentially using thousands of times more energy.
Results show that image generation models vary drastically in terms of the energy they consume, with up to a 46x difference. … while prompt length and content have no statistically significant impact.
Video generation can consume one to two orders of magnitude more energy than image generation.
Each interaction consumes energy—about 0.34 watt-hours per prompt.
Although most of the AI environmental research have revolved around the energy consumption and carbon footprints during training LLMs9,10,11,12, 13, the energy concerns associated with LLM inference have been studied much less. Yet, the LLM inference demand is substantially larger 14, 15.
An informal online estimate for ChatGPT indicates that it produces 0.382 g CO2e per query 18, based on 3.82 metric tons CO2e per day divided by 10,000,000 queries per day.
The energy consumption of AI and especially individual AI tasks is complex to measure. A critical aspect of the energy evaluation of AI systems is the precise definition of both the scope and methodology.
However, they found that using an AI model to generate text is significantly less energy-intensive. Creating text 1,000 times only uses as much energy as 16% of a full smartphone charge. … In contrast, generating 1,000 images with a powerful AI model, such as Stable Diffusion XL, is responsible for roughly as much carbon dioxide as driving the equivalent of 4.1 miles in an average gasoline-powered car. In contrast, the least carbon-intensive text generation model they examined was responsible for as much CO2 as driving 0.0006 miles in a similar vehicle.
To the best of our knowledge, no prior work has systematically examined the energy consumption of image generation by analyzing the interplay between model architecture, quantization, resolution, and prompt length.
When we set out to write a story on the best available estimates for AI’s energy and emissions burden, we knew there would be caveats and uncertainties to these numbers.
However, substantial computational costs and energy footprint of prompt inferencing process remain critical challenges while building generative AI applications.
Power consumption per AI task spans five orders of magnitude, from 0.24 Wh for a Gemini chat prompt to several kWh for a full day of Claude Code.
Generative AI creates the ability to automatically generate new content: not only text, but also images, sound and video. Through a new generation of tools, such as ChatGPT, Gemini, Claude and Le Chat, simple instructions, known as prompts, in natural language, either written or spoken, make it possible to rapidly generate content that had previously required complex human tasks to produce.
A new study from Carnegie Mellon University and the French-American AI startup HuggingFace reveals that utilizing generative AI requires high quantities of electricity, creating a range of worrying implications for the energy grid and our changing climate. … The study, published Tuesday on the arXiv ahead of peer review, examines the energy consumption—and thus also the carbon emissions—associated with both AI chatbots and image generators.
Diffusion models. We begin with the relatively more straightforward case of diffusion models, which are used for text-to-image, text-to-video, and image-to-video generation.
To train such an R-DTCBF that is valid not only on sampled states but also across the entire region, we employ a verification algorithm iteratively in a counterexample-guided approach.
In this work, we propose a novel RL framework that integrates generative diffusion models with a kernel-based method—Gaussian Process Regression (GPR)—to serve as the policy.
We present FlowRL, a novel framework for online reinforcement learning that integrates flow-based policy representation with Wasserstein-2-regularized optimization.
What determines the energy consumption of generating one response? For Large Language Models (LLMs), a response is a complete answer to a prompt with all output tokens included. For diffusion models, a response is one generated image or video.
The primary objective of this study is to discover characteristics in the LLM inference energy usage, specifically focusing on the impact of different types of prompts and hyperparameters.
1. Read a 100-page PDF and generate a one-page summary 2. Generate a 2K photorealistic image from a text prompt (multiple iterations)
CO2 Inference - a Hugging Face Space by sasha
What do you think of the claim?
Your challenge will appear immediately.
Challenge submitted!
For developers
This same pipeline is available via API.
Verify your AI's output programmatically.
/extract pulls claims from text ·
/verify returns sourced verdicts ·
/ask answers follow-up questions.
Continue your research
Verify a related claim next.
Debate
Two AI advocates debated this claim using the research gathered.
Argument for
Source 2 (arxiv.org) directly establishes that image generation consumes over 30 times as much energy as generating text, while Source 1 (technologyreview.com) quantifies that one image equals a full smartphone charge versus 1,000 text generations equaling only 16% of a charge, implying thousands of times greater energy use per image. These findings are reinforced by Source 8 (libguides.usc.edu), which reports a 1,450-fold difference, confirming that a single generative AI image exceeds ten times the energy of a text prompt.
The Proponent cherry-picks extreme estimates from Sources 1, 2, and 8 while ignoring the massive variability across models, which can differ in energy consumption by up to 46 times according to Source 9. Furthermore, the Proponent conveniently omits Source 3, which explicitly caps the energy difference by stating a single image consumes 'up to' ten times more energy than a text request, proving the universal claim false.
Argument against
Although limited evidence points to high energy costs for visual AI, the absolute claim that image generation always consumes "more than ten times" the energy of text is demonstrably false, as Source 3 explicitly caps this metric by stating an image consumes "up to" ten times more energy depending on the configuration. Furthermore, because energy consumption varies drastically—by up to 46 times between different image models alone (Source 9)—and depends heavily on complex, highly variable measurement methodologies (Source 14), asserting a strict, universal multiplier fundamentally misrepresents the data.
The Opponent misreads Source 3 (heise.de) by treating its configuration-dependent “up to ten times” phrasing as a universal ceiling that falsifies the claim, while ignoring that this same source still places image generation well above text and that Source 2 (arxiv.org) and Source 8 (libguides.usc.edu) document multiples of over 30 times and 1,450-fold. Variation across models noted in Source 9 (arxiv.org) and methodological complexity in Source 14 (kspublisher.com) do not erase the consistent empirical finding that a single image exceeds ten times the energy of a text prompt; they merely describe the range within which that excess holds.
Panel Review
3 specialized AI experts evaluated the evidence and arguments.
Reviewer 1 — The Logic Examiner
Sources 1, 2, and 8 supply direct quantitative comparisons showing a single generative-AI image uses roughly 30× to thousands of times the energy of a text prompt, while Source 3's configuration-dependent “up to ten times” figure is consistent with the lower end of that range rather than a universal ceiling; the Opponent's reading of Source 3 as falsifying the claim therefore fails. The inferential chain from the measured energy ratios to the claim is sound, so the claim is true.
Reviewer 2 — The Source Auditor
Reliable sources, including MIT Technology Review (Source 1) and multiple arXiv studies (Sources 2, 6, 8), consistently report that generating an image with AI consumes significantly more energy than generating text, with estimates ranging from 30 times to over 1,000 times more energy per task. While one source (Source 3) mentions 'up to ten times' compared to an average ChatGPT request, the overwhelming weight of high-quality evidence confirms that image generation generally consumes orders of magnitude more energy than text generation, making the claim true.
Reviewer 3 — The Precision Analyst
Sources 1 and 2 support ratios well above ten for the measured image and text-generation cases, but Source 3 and Source 9 show that energy use varies materially by configuration and model. The claim is directionally supported for prominent measured cases, yet its unqualified "a single image" wording overstates a conditional comparison as a general rule.