Inside The AI Stack

Research project

GPU Cloud Price Index — Methodology

A planned recurring measurement of accelerator pricing across cloud providers, normalised so the comparison is meaningful. Methodology published before data collection begins.

foundationalGPUCloudAI Infrastructure
ByJames JoynerPublished Verified 2 min read
State: planned

The problem with published GPU pricing

Headline hourly rates are not comparable across providers, and the differences are large enough to reverse a ranking:

  • What is bundled. Some rates include local NVMe, high-bandwidth networking, and egress allowance. Others price each separately.
  • Interconnect. Two instances with the same accelerator can differ by an order of magnitude in inter-node bandwidth, which decides whether a distributed job scales at all.
  • Commitment. On-demand, reserved, spot, and negotiated rates differ by multiples, and the published rate is often the one nobody pays.
  • Availability. A price for capacity you cannot obtain is not a price.
  • Region. The same configuration varies substantially between regions.

A comparison that ignores these produces a table that looks authoritative and is misleading. That is the failure mode this method is designed around.

Intended method

What is measured. Published on-demand pricing for a defined set of accelerator configurations, collected from public pricing pages and APIs on a fixed schedule, in a fixed set of regions.

Normalisation. Prices normalised per accelerator-hour, with bundled resources itemised separately so the reader can see what is included rather than having it averaged away.

What is recorded alongside. Interconnect bandwidth, local storage, memory per accelerator, and the commitment terms the price requires. A price without these is not comparable and will not be published without them.

Availability. Where it can be observed programmatically, recorded. A configuration that is listed but never obtainable is marked as such.

Frequency. Recurring, with each edition retained. The value of an index is the series, not any single reading.

What will be published

  • The full dataset as CSV, under an open licence
  • The collection code, so the numbers can be reproduced or disputed with the same harness
  • The exact configurations compared, and the reasoning for choosing them
  • Explicit notes on what is not comparable, rather than a single ranked list

What this will not be

  • A recommendation. The cheapest accelerator-hour is frequently the wrong choice, because the fabric or the storage makes the job slower. Price is one input.
  • An affiliate-driven comparison. No provider in this index will be an affiliate relationship. If that ever changes, it will be disclosed on the page and the affected provider will be excluded from any ranking.
  • A claim about negotiated rates. We measure published pricing. Enterprise agreements are not public and we will not speculate about them.

Timeline

Not scheduled. The method is defined; collection has not begun.

Announcements go to Inside The AI Stack Brief first.

Verification status

This resource has not been executed end to end in a lab environment. Commands and configuration are reviewed by an engineer, but treat them as reference rather than as a tested procedure.

Author

James Joyner

Builds and operates the infrastructure layers underneath production AI systems.

James founded Inside The AI Stack to publish the kind of infrastructure and operations material he wanted while running production systems: specific, tested where it claims to be tested, and written by someone who has had to fix the thing at 3am. He works across AI infrastructure, private cloud, and platform engineering, and reviews every technical resource published here before it is marked as verified.

  • AI infrastructure
  • OpenStack operations
  • Kubernetes
  • Terraform
  • Linux systems engineering
  • Observability

Related resources chosen because they are the next thing you would actually need — not because they share a keyword.

Guide

GPU Infrastructure for AI

What determines accelerator performance in production: memory capacity versus bandwidth, interconnect topology, and the checks that find a misplaced workload.

advanced· 6 minGPUNVIDIA

Newsletter

Inside The AI Stack Brief

A practical weekly briefing on AI engineering, infrastructure, production operations, and the technologies powering the AI stack.

One email a week. No sponsorship placements inside the technical sections. Unsubscribe in one click.