Last updated ·Uses decimal GB/TB for procurement sizing

AI KV Cache Memory Tier Calculator

Size the memory tiers behind long-context AI inference. The calculator estimates how much KV cache your workload creates, how much can remain in HBM, how much shared DDR5 is needed, how much NVMe fallback remains and whether bandwidth becomes the practical bottleneck.

It is built for infrastructure planners comparing HBM-heavy servers, CXL memory pooling, DDR5-backed shared memory and enterprise NVMe storage. Live DDR5 RDIMM and enterprise SSD prices are used where DatacenterDisk has sufficient current market data; HBM and fabric costs remain manual because those are typically quote-based system costs.

How Do You Size KV Cache Across HBM, DDR5 and NVMe?

Start with concurrent sequences, active context tokens and KV cache bytes per token. Reserve HBM for model weights and runtime, place the remaining hot KV cache in available HBM, then model the overflow into a shared DDR5 tier. Anything that still does not fit becomes NVMe fallback. The calculator keeps raw KV cache, effective incremental cache after prompt reuse, capacity overflow and bandwidth demand separate so a large-memory design is not mistaken for a low-latency design.

Live DDR5 ECC RDIMM Pricing

Current in-stock DDR5 RDIMM listings used as market inputs for the shared-memory tier.

Memory CAPEX calculator
ManufacturerPart NumberCapacityDDRSpeedRankPrice$/GBSellerStatusCTA
Kingston
KSM56R46BD8PMI
Details
32GBDDR5DDR5-5600-$999.99$31.25AmazonIn stockBuy/View Price
SK Hynix
HMCG88AEBRA168N
Details
32GBDDR5DDR5-4800-$1,175.59$36.74AmazonIn stockBuy/View Price
A-Tech
M321R4GA0PB0-CWM
Details
32GBDDR5DDR5-5600-$1,194.00$37.31AmazonIn stockBuy/View Price
NEMIX
MTC20F2085S1RC48BR
Details
32GBDDR5DDR5-4800-$1,207.49$37.73AmazonIn stockBuy/View Price
Kingston
KSM56R46BD4PMI
Details
64GBDDR5DDR5-5600-$2,064.00$32.25AmazonIn stockBuy/View Price
Samsung
M321R8GA0BB0-CQK
Details
64GBDDR5DDR5-4800-$2,245.00$35.08AmazonIn stockBuy/View Price
Micron
MTC40F2046S1RC56BD2
Details
64GBDDR5DDR5-5600-$2,624.99$41.02AmazonIn stockBuy/View Price
A-Tech
B0DYR9JXRH
Details
64GBDDR5DDR5-6400-$2,699.00$42.17AmazonIn stockBuy/View Price
A-Tech
B0DMC1VFNS
Details
96GBDDR5DDR5-4800-$3,139.49$32.70AmazonIn stockBuy/View Price
Samsung
M321RYGA0BB0-CQK
Details
96GBDDR5DDR5-4800-$3,139.49$32.70AmazonIn stockBuy/View Price
NEMIX
MTC40F204WS1RC56BB2R
Details
96GBDDR5DDR5-5600-$3,779.99$39.37AmazonIn stockBuy/View Price
NEMIX
MTC40F204WS1RC64BC1R
Details
96GBDDR5DDR5-6400-$3,884.99$40.47AmazonIn stockBuy/View Price
NEMIX
MTC40F204WS1RC64BR
Details
96GBDDR5DDR5-6400-$3,884.99$40.47AmazonIn stockBuy/View Price
Samsung
M321R2GA3BB6-CQK
Details
16GBDDR5DDR5-4800-$472.90$29.56AmazonIn stockBuy/View Price
Kingston
KSM56E46BD8KM
Details
48GBDDR5DDR5-5600-$1,749.99$36.46AmazonIn stockBuy/View Price
NEMIX
MTC10F1084S1RC56BR
Details
16GBDDR5DDR5-5600-$629.99$39.37AmazonIn stockBuy/View Price

Live Enterprise SSD Pricing

Current enterprise NVMe and SAS SSD listings used for NVMe fallback capacity and cost planning.

Enterprise NVMe prices
ManufacturerModelPart NumberCapacityInterfaceForm FactorTechPrice$/TBStatusCTA
Micron
7400 PRO 3.84TB U.3
Details
MTFDKCB3T8TDZ
3.84TBNVMe-PCIe4U.3TLC$701.68$182.73In stockBuy/View Price
Dell
3.84TB SAS SSD Read Intensive 12G
Details
MZILT3T8HBLS
3.84TBSAS-12G2.5"TLC$799.99$208.33In stockBuy/View Price
Micron
7450 PRO 3.84TB U.3
Details
MTFDKCC3T8TFS
3.84TBNVMe-PCIe4U.3TLC$1,990.00$518.23In stockBuy/View Price
Kioxia
CD8 Series 3.84TB NVMe U.2
Details
KCD8XVUG3T84
3.84TBNVMe-PCIe4U.2TLC$2,461.20$640.94In stockBuy/View Price
Micron
7600 PRO 3.84TB U.2 PCIe5
Details
MTFDLAL3T8THG
3.84TBNVMe-PCIe5U.2TLC$2,599.00$676.82In stockBuy/View Price
Samsung
PM9A3 3.84TB U.2 NVMe
Details
MZQL23T8HCLS
3.84TBNVMe-PCIe4U.2TLC$3,499.99$911.46In stockBuy/View Price
Samsung
PM9A3 7.68TB U.2 NVMe
Details
MZQL27T6HBLA-00A07
7.68TBNVMe-PCIe4U.2TLC$2,715.99$353.64In stockBuy/View Price
Samsung
PM9A3 7.68TB U.2 NVMe
Details
MZQL27T6HBLA
7.68TBNVMe-PCIe4U.2TLC$3,999.00$520.70In stockBuy/View Price
WD
Ultrastar DC SN655 7.68TB U.3
Details
WUS5EA176ESP7E1
7.68TBNVMe-PCIe4U.3TLC$8,690.03$1,131.51In stockBuy/View Price
Seagate
Nytro 5060 E3.S 15.36TB
Details
XP15360SE70045
15.36TBNVMe-PCIe4E3.STLC$5,879.00$382.75In stockBuy/View Price
SanDisk
Ultrastar DC SN655 61.44TB U.3
Details
WUS5EC0C1ESP7Y3
61.44TBNVMe-PCIe4U.3-$68,559.43$1,115.88In stockBuy/View Price
Samsung
PM9A3 960GB U.2 NVMe
Details
MZQL2960HCJR
0.96TBNVMe-PCIe4U.2TLC$728.00$758.33In stockBuy/View Price
Kioxia
CD8 Series 3.2TB NVMe U.2
Details
KCD8XVUG3T20
3.2TBNVMe-PCIe4U.2TLC$2,570.97$803.43In stockBuy/View Price
Samsung
PM9A3 1.92TB U.2 NVMe
Details
MZQL21T9HCJR
1.92TBNVMe-PCIe4U.2TLC$2,099.00$1,093.23In stockBuy/View Price
Samsung
PM1643a 1.92TB SAS SSD
Details
MZILT1T9HBJR-00007
1.92TBSAS-12G2.5"TLC$2,195.00$1,143.23In stockBuy/View Price
Samsung
PM1643A 1.92TB SAS SSD
Details
MZILT1T9HBJR
1.92TBSAS-12G2.5"TLC$2,199.00$1,145.31In stockBuy/View Price

What Is KV Cache in AI Inference?

KV cache is the memory that stores attention keys and values for tokens already processed by a transformer model. During generation, the model can reuse that state instead of recomputing attention across the full prompt and generated history. That reuse is one reason modern inference servers can serve long context windows, multiple users and streaming token output at useful speeds.

The difficult part is scale. A single short prompt may not create a large memory problem, but thousands of concurrent sequences with long active contexts can create many terabytes of cache pressure. Once the working set outgrows local accelerator HBM, the system must choose between reducing concurrency, shortening context, adding more accelerators, adding a shared memory tier or falling back to storage.

How Much Memory Does KV Cache Use?

The calculator uses a transparent planning formula: concurrent sequences times context tokens times KV bytes per token times active context utilization. The bytes-per-token input is intentionally editable because the true value depends on the model architecture, layer count, hidden size, attention layout, quantization, runtime and cache format. Model-size labels are shown for planning context only; they do not silently change the formula.

Prompt cache reuse is shown separately. A high reuse rate can reduce incremental KV cache growth when many users share prefixes or system prompts, but it does not make raw memory pressure disappear. Seeing both raw and effective cache prevents teams from building a plan that only works when reuse is perfect.

Why KV Cache Outgrows HBM

HBM is the fastest and closest memory tier for accelerators, but a large share is usually consumed by model weights, runtime state, activation workspace and framework overhead. The calculator therefore starts with total HBM, subtracts the reserved percentage and models only the remaining capacity as available for KV cache. If the effective cache exceeds that amount, overflow becomes a memory-tier design question rather than a simple GPU count question.

Adding accelerators increases HBM capacity, but it can also increase total serving demand. The practical goal is not simply to maximize memory; it is to keep the hottest cache close to compute, place warm cache in a lower-cost shared memory tier and avoid letting too much of the active working set become storage-bound.

HBM vs DDR5 vs NVMe for KV Cache

TierTypical roleRelative latencyRelative bandwidthCapacityPersistenceCost tendencyBest use
HBMLocal accelerator memoryLowestHighestLowestNoHighestModel weights, hot KV cache and latency-sensitive compute state
Shared DDR5Shared memory expansionLow to moderateHighHighNoLower than HBMWarm KV cache, memory pooling and high-concurrency inference
NVMePersistent storage tierHighest of these tiersModerate to highVery highYesOften lowest per TBModel files, datasets, vector stores and limited fallback cache

What Is a Shared AI Memory Tier?

A shared AI memory tier gives several accelerators access to memory outside their local HBM. It can be implemented with CXL memory expansion, DDR5 memory appliances, optical links or other rack-scale fabrics. The goal is not to make DDR5 behave exactly like HBM; it is to provide a larger intermediate tier so warm KV cache does not immediately fall to persistent storage.

This is where architectures such as Marvell Photonic Fabric Memory become relevant. The announced approach combines DDR5-backed pooled memory, local HBM caching and optical connectivity. The calculator remains vendor-neutral because the sizing problem exists whether the shared tier is optical, electrical, CXL-based or built into a future platform design.

How CXL Memory Pooling Works

CXL is a coherency and interconnect protocol that can expose memory expansion resources to hosts and accelerators. In an AI inference design, CXL may help attach external DDR5 capacity without treating it like a block device. That matters because memory semantics are very different from reading and writing files or object chunks from storage.

CXL does not remove the need to model capacity and bandwidth. A system can have enough pooled DDR5 capacity and still bottleneck if the fabric cannot move cache data at the rate required by the active sequences. This calculator therefore checks both configured shared-memory capacity and configured shared-memory bandwidth.

Why Optical Memory Fabrics Matter

Electrical connectivity becomes harder as distances, lane counts and bandwidth demands grow. Optical fabrics are attractive for rack-scale or pod-scale memory sharing because they can target longer reach and high aggregate bandwidth with different power and cabling tradeoffs. Vendor claims should still be treated as implementation-specific until production systems are benchmarked under real workloads.

The practical planning question is simple: if optical fabric makes a larger DDR5 pool reachable by more accelerators, how much HBM can remain dedicated to the hottest cache and how much lower-cost memory can absorb the warm working set? The answer depends on context length, concurrency, cache reuse and data movement per generated token.

How to Size DDR5 for AI Inference

DDR5 sizing starts after HBM allocation. The calculator places the HBM overflow into shared DDR5, then compares that requirement against configured DDR5 capacity. It also calculates approximate DIMM count across 64GB, 96GB, 128GB and 256GB RDIMMs. Whole DIMMs are required, so the final installed capacity is rounded up and unused capacity is shown explicitly.

Live median $/GB is used where enough current DDR5 RDIMM products exist for a capacity. If the market has insufficient data for a DIMM size, the calculator says so rather than inventing a price. That is important because larger DIMMs may improve slot efficiency while still carrying a different cost per GB.

When KV Cache Falls Back to NVMe

NVMe fallback is the remaining active KV-cache requirement after HBM and shared DDR5 are allocated. Enterprise SSDs can provide large capacity at a much lower cost per TB than HBM, but they are still persistent storage devices rather than memory. A fallback tier can be useful for cold or less latency-sensitive state; it is risky when a large share of the active cache must be read repeatedly during token generation.

The enterprise SSD optimizer compares 3.84TB, 7.68TB, 15.36TB, 30.72TB and 61.44TB capacities. It uses whole-drive counts, installed capacity and excess capacity so a planner can see the acquisition implication of each capacity class.

How Memory Bandwidth Affects Token Throughput

Capacity is not sufficient by itself. A tier can be large enough and still too slow if aggregate demand exceeds the available bandwidth. The bandwidth model estimates data movement from active sequences, tokens per second and bytes moved per generated token, then divides demand across HBM, shared memory and NVMe according to the expected hit-rate assumptions.

The output is intentionally an estimate rather than a benchmark. It helps identify whether a design is likely to be HBM-capacity bound, shared-memory-capacity bound, shared-memory-bandwidth bound, NVMe-capacity bound, NVMe-bandwidth bound or reasonably balanced under the assumptions entered.

How to Reduce KV Cache Infrastructure Cost

Cost reduction usually starts with workload discipline. Shorter active context, better prompt reuse, batching policies, context eviction, cache quantization and routing similar requests together can reduce the amount of new KV cache created. Hardware planning then decides whether more HBM, more DDR5, more shared-memory bandwidth or more NVMe capacity is the better use of budget.

For capital planning, compare this tool with the AI memory CAPEX calculator. That calculator models fleet-level server RAM and enterprise SSD CAPEX, while this one focuses on the memory hierarchy behind active inference.

KV Cache Formula Table

MetricFormula
KV cacheSequences x Tokens x Bytes per token
Effective KV cacheRaw KV cache x active utilization x (1 - prompt cache reuse)
HBM availableTotal HBM - reserved HBM, unless manually overridden
HBM overflowKV cache - available HBM
Shared RAM requirementPortion of HBM overflow assigned to DDR5
NVMe fallbackKV cache - HBM allocation - DDR5 allocation
DIMM countceil(Shared DDR5 required / DIMM size)
SSD countceil(NVMe fallback required / SSD capacity)

Frequently Asked Questions

What is KV cache in AI inference?

KV cache stores attention key and value tensors so an inference server does not have to recompute the full prior context for every generated token. It improves throughput and latency, but long-context and high-concurrency workloads can consume large amounts of memory.

How much memory does KV cache use?

A planning estimate is concurrent sequences multiplied by active context tokens multiplied by KV bytes per token. The calculator also applies average context utilization and shows raw KV cache separately from the effective incremental cache after prompt reuse.

Can KV cache fit entirely in HBM?

Small batches and shorter contexts may fit in available HBM. Large concurrent inference, long context windows and multi-tenant serving often exceed the HBM that remains after model weights and runtime overhead are reserved.

What is a shared AI memory tier?

A shared AI memory tier is capacity outside local accelerator HBM that multiple accelerators can access. It may be built with DDR5 memory modules, CXL memory expansion, optical fabric links or vendor-specific memory appliances.

Should KV cache spill to NVMe?

NVMe can provide large capacity and high sequential bandwidth, but it is a lower memory tier than HBM or DDR5. A small fallback share may be acceptable for some workloads, while a large active KV-cache share on NVMe can become a throughput and latency bottleneck.

Does this calculator require Marvell Photonic Fabric Memory?

No. Marvell Photonic Fabric Memory is one example of a shared-memory architecture. The calculator is vendor-neutral and can also model CXL memory pools, DDR5 memory expansion nodes and NVMe fallback designs.

Does DatacenterDisk have live HBM pricing?

No. HBM pricing is entered manually because it is normally part of accelerator or system pricing. DDR5 RDIMM and enterprise SSD inputs can use DatacenterDisk live market data where enough current listings exist.

Share:WhatsAppFacebookPost on X