T4 GPU Rental Prices & Live Availability
As of , no single T4 has a current in-stock signal. The lowest listed price is $0.20 per GPU-hour on Theta EdgeCloud, but that provider does not expose a current stock signal. GPU Finder compares 4 providers for this size and refreshes availability hourly. Methodology.
Structured offer basis: 132 current 1× T4 listings range from $0.20 to $6.70 USD per physical GPU-hour. Whole-node prices for other GPU counts are not mixed into this range.
Cheapest single T4 rental
verified- Is now a good time to rent an T4 cloud GPU?
- ▼ near its 6-month low — The 1× T4 on-demand floor is $0.20/hr ($0.20/GPU-hr) — 6-month range $0.20–$0.20/hr for 1× nodes. In short, a relatively good time to rent — the floor price is near its 6-month low.
- What is the cheapest T4 cloud GPU available right now?
- No single T4 is confirmed in stock across tracked providers at the moment.
Prices are the cheapest listed on-demand rates; “in stock” means the provider reported capacity on our last hourly check. How we compute this →
T4 rental prices by provider
Get notified when T4 drops in price or reliability changes
One email per change, max once a day. We send a confirmation link first; one-click unsubscribe in every email.
30-day reliability depth
GPU Finder does not treat availability as a static yes/no flag. For T4, current stock badges are paired with a 30-day reliability score based on our hourly stock checks — the share of tracked time each listing was reported available. How we compute this →
- 0
- listings with a 30-day score
- 0%
- of listings on this page have a score
- —
- longest stock history on this page
This page has live pricing and current stock rows, but not enough retained availability history yet to publish 30-day reliability percentages.
Caveat: scores stay hidden until a listing has at least 48 tracked hours, and some provider APIs expose coarse capacity levels instead of exact stock counts. Use the score as historical depth beside current availability, not as a guarantee that a GPU will still be allocatable when you click through.
Price History & Comparison
Full T4 price history (7 months) →Spot pricing decision guide
Treat spot as a risk-adjusted capacity decision, not just a cheaper number.
per GPU-hour
71% below on-demand floor
daily median for the same offers with complete history
median $0.40/GPU-hr
Cheapest trusted spot capacity is on Azure; 0 providers currently report spot or interruptible stock.
30-day reliability context is still sparse for the trusted spot rows on this GPU. See the availability section for the broader stock history.
The discount is visible, but availability, reliability, or volatility argues for more caution.
Good fit: checkpointed batch jobs, flexible training, and experiments that can restart. Avoid for production serving, deadline-bound runs, or jobs that cannot tolerate eviction.
Source caveat: Runpod exposes explicit spot fields; Vast is labeled interruptible/bid-floor. Lambda is treated as on-demand-only unless a verified spot field is added.
About the T4
The NVIDIA T4 is the original inference-optimized datacenter GPU and remains one of the most widely deployed accelerators in cloud infrastructure. Its 70W power draw and single-slot PCIe form factor make it trivially easy to provision at scale. While outperformed by every newer GPU, the T4 still handles production inference for INT8-quantized models efficiently, and its ubiquity means rock-bottom pricing.
Key Specifications
| Architecture | Turing (TU104) |
| GPU Memory | 16 GB GDDR6 |
| Memory Bandwidth | 300 GB/s |
| FP16 Tensor Core | 65 TFLOPS |
| TDP | 70W |
| Interconnect | PCIe Gen3 x16 |
| Release Year | 2018 |
Cloud Pricing Context
T4 pricing starts as low as $0.27/hr on budget providers and about $0.35/hr on Google Cloud. Its low TDP keeps per-instance costs minimal. The T4 is often the default GPU for inference endpoints on managed ML platforms like Vertex AI and SageMaker.
Best For
- Production inference for INT8/FP16 models under 7B parameters
- Cost-sensitive batch processing and embedding generation
- Always-on inference endpoints where low power draw reduces TCO
- Entry-level GPU workloads on managed ML platforms