Last updated ·Vendor claims are labeled where specifications are announced rather than independently measured
AI memory architecture

Marvell Photonic Fabric Memory: AI KV Cache, DDR5 Pooling & Optical Memory Explained

Marvell Photonic Fabric Memory is part of a new generation of AI memory infrastructure aimed at long-context inference. Instead of treating every cache miss as a storage problem, the architecture adds a large DDR5-backed shared memory tier between local accelerator HBM and persistent NVMe storage.

The idea is especially relevant for KV cache, agentic workloads, retrieval-heavy inference and systems where accelerators need access to more memory than their local HBM can hold. The important distinction is that Marvell has announced architectural targets and product direction; buyers should still separate vendor-reported specifications from production measurements and purchase availability.

What Is Marvell Photonic Fabric Memory?

Marvell Photonic Fabric Memory is an announced shared-memory architecture for AI inference. It places a DDR5-backed memory pool behind an optical fabric so multiple accelerators can reach a larger warm memory tier, while HBM remains the fastest local cache for the hottest data. For long-context AI, that shared tier can hold KV cache that would otherwise spill to slower persistent storage. Marvell describes the architecture with optical connectivity, CXL where applicable, DDR5 memory modules and local HBM caching. The practical value depends on final systems, topology, workload behavior, bandwidth, latency and software integration, so the safest way to plan is to model HBM, shared DDR5 and NVMe fallback separately.

Architecture Diagram

Accelerators
GPU / XPU
GPU / XPU
GPU / XPU
Local HBM
Fastest tier for model weights, hot KV cache and runtime state
Photonic Fabric Memory
Shared warm memory tier reached across an optical fabric
DDR5 Pool
Capacity-oriented volatile memory behind memory modules or appliances
NVMe / Persistent Storage
Lower tier for model files, datasets, checkpoints and fallback state
Shared Access
CXL where applicable
Optical fabric
Rack or multi-rack scale

Key Specification Table

Primary roleShared AI inference memory tier for warm KV cache and memory pooling
Backing memoryDDR5, according to Marvell's announced architecture
Fast cacheHBM used as a faster cache tier in the announced memory-module design
FabricOptical connectivity using Marvell Photonic Fabric technology
Host/XPU connectivityCXL connectivity where applicable; Marvell also describes PCIe Gen6 optical NIC capability
Vendor-reported module detailsMarvell has described a Photonic Fabric Memory Module with up to eight DDR5 DIMMs, 72GB of HBM cache and 7.2 Tbps optical fabric bandwidth.
Vendor-reported scaleMarvell describes rack-scale and multi-rack memory sharing. Exact reach and topology depend on the product implementation.
Intended workloadKV cache, agentic AI inference, context-heavy serving and memory-bound accelerator utilization
Storage fallbackReduced where a DDR5 memory tier absorbs warm data that would otherwise move to NVMe
AvailabilityAnnounced architecture and product direction; broad production availability and measured system performance remain implementation-dependent.

What Problem Photonic Fabric Memory Solves

AI inference is becoming a memory hierarchy problem. Model weights occupy expensive local HBM, active attention state expands with longer contexts and concurrent sessions, and retrieval or agent workflows keep more intermediate state hot for longer periods. If the system cannot hold that state near the accelerators, it may fall back to persistent storage or reduce the amount of useful work each accelerator can do.

Marvell's stated framing is that inference needs more memory capacity without making every accelerator carry all of that capacity locally. DDR5 is much cheaper and more expandable than HBM, but ordinary DDR5 placement does not automatically make it shareable at rack scale. Photonic Fabric Memory addresses that gap by creating an optical path to pooled memory so multiple accelerators can reach a larger warm tier.

HBM vs DDR5 vs Shared Memory vs NVMe

The core comparison is not simply fast versus slow. HBM is fast, local and capacity constrained. DDR5 is larger and cheaper per GB, but it needs a coherent or well-managed interconnect to behave like useful shared memory. NVMe is persistent and extremely dense, but it remains a storage tier with higher access latency.

TierTypical roleRelative latencyRelative bandwidthCapacityPersistenceCost tendencyBest use
HBMAccelerator-local memoryLowestHighestMost constrainedNoHighestWeights, hot KV cache and latency-sensitive compute state
Photonic Fabric MemoryShared warm memory tierLow to moderateHigh, implementation-dependentHighNoBetween HBM and storageWarm KV cache and shared inference memory pools
DDR5 server memoryHost or pooled system memoryLowHighHighNoLower than HBMHost RAM, memory expansion and capacity-oriented pooling
NVMe SSDPersistent flash storageHighest of these tiersModerate to highVery highYesOften lowest per TBModel files, datasets, checkpoints, vector stores and limited fallback

KV Cache Explanation

KV cache stores key and value tensors for prior tokens in a transformer sequence. During generation, the model reuses those values rather than recomputing every prior token. That is useful for speed, but the cache grows with context length, number of layers, hidden representation, precision, concurrent sequences and active context utilization.

A warm KV-cache tier is attractive because not every cached token has the same access pattern. The hottest state should remain as close to compute as possible. Older or less frequently reused state may tolerate a lower tier if bandwidth and latency are sufficient. Photonic Fabric Memory is aimed at that middle region: larger than local HBM, faster and more memory-like than falling straight to NVMe.

Photonic Fabric vs Conventional Memory Architecture

A conventional accelerator server mostly relies on local HBM, host DRAM and local storage inside the chassis. That model is simple, but it can strand memory. One accelerator may have headroom while another is constrained, and a single server may not be the right boundary for a multi-rack inference service.

Photonic Fabric Memory changes the boundary by making memory a shared resource. The announced architecture puts DDR5 behind an optical fabric and uses HBM caching to keep the hottest data close. That does not remove locality concerns. It gives architects another tier to size, price and monitor.

Marvell Photonic Fabric Memory vs CXL Memory Pooling

Photonic Fabric Memory and CXL memory pooling are not necessarily competitors. CXL is a protocol family for coherent or memory-semantic connectivity between hosts, accelerators and memory devices. An optical fabric is a transport and scaling mechanism. Pooled memory is the resource exposed through those layers.

Marvell says its broader AI memory infrastructure includes CXL memory expansion and pooling alongside photonic fabric modules, NICs and chiplet technology. A real system could therefore use CXL semantics across part of the path while optical connectivity handles reach, radix or rack-scale connection density. The practical question is not whether CXL or optics "wins"; it is whether the final topology provides enough capacity, latency and bandwidth for the workload.

Performance and Bandwidth Discussion

Marvell has reported aggressive architecture targets, including a Photonic Fabric Memory Module concept with 7.2 Tbps optical fabric bandwidth and HBM caching. It has also discussed warm KV-cache offload at large capacities and token-throughput improvements in vendor materials. These are vendor-reported claims, not universal performance guarantees.

Real performance will depend on memory population, software stack, model, attention pattern, hit rate, scheduler behavior, fabric topology and contention. For planning, treat the bandwidth number as an input to test, not a marketing line to copy into a bill of materials. The KV Cache Memory Tier Calculator lets you model configured shared-memory bandwidth against the demand created by active sequences.

Power and Density Implications

HBM capacity is expensive and packaged close to accelerators. NVMe flash is dense and persistent but a poor substitute for memory when the active working set requires repeated low-latency access. DDR5 sits between them: cheaper and more expandable than HBM, but not free in power, board area or memory-channel design.

Optical memory fabrics can shift rack design by moving some capacity away from every individual accelerator tray and into shared modules. That may reduce duplicated memory capacity and improve accelerator utilization if the application has enough reuse and the shared tier is fast enough. It also introduces fabric cost, optics power, cabling, monitoring and failure-domain considerations that have to be modeled rather than assumed away.

Use Cases

The clearest use case is long-context AI inference where KV cache grows faster than local HBM budgets. Agentic systems, coding assistants, document analysis, multi-turn customer support, search-augmented generation and large prompt-cache deployments can all create warm memory pressure that is neither pure compute nor ordinary bulk storage.

A second use case is rack-scale memory pooling. If a service has uneven memory pressure across accelerators, shared DDR5 may improve utilization by making more total memory available to the fleet. A third use case is reducing storage fallback: keeping warm state in memory can avoid pushing too much active context into NVMe, where latency and throughput behavior are different.

Model Your KV Cache Tier

Use the calculator to estimate KV cache size, HBM overflow, shared DDR5 capacity, NVMe fallback, bandwidth demand, DIMM count and enterprise SSD count. It is vendor-neutral, so Marvell Photonic Fabric Memory can be modeled as one possible shared-memory tier rather than a required product choice.

Launch the KV Cache Memory Tier Calculator

Relevant Live DDR5 RDIMM Pricing

Current server-memory prices for planning shared DDR5 capacity. Market data refreshed in the Sep 1, 2026, 7:30 AM.

Server RAM price index
ManufacturerPart NumberCapacityDDRSpeedRankPrice$/GBSellerStatusCTA
Kingston
KSM56R46BD8PMI
Details
32GBDDR5DDR5-5600-$999.99$31.25AmazonIn stockBuy/View Price
SK Hynix
HMCG88AEBRA168N
Details
32GBDDR5DDR5-4800-$1,175.59$36.74AmazonIn stockBuy/View Price
A-Tech
M321R4GA0PB0-CWM
Details
32GBDDR5DDR5-5600-$1,194.00$37.31AmazonIn stockBuy/View Price
NEMIX
MTC20F2085S1RC48BR
Details
32GBDDR5DDR5-4800-$1,207.49$37.73AmazonIn stockBuy/View Price
Kingston
KSM56R46BD4PMI
Details
64GBDDR5DDR5-5600-$2,064.00$32.25AmazonIn stockBuy/View Price
Samsung
M321R8GA0BB0-CQK
Details
64GBDDR5DDR5-4800-$2,245.00$35.08AmazonIn stockBuy/View Price
Micron
MTC40F2046S1RC56BD2
Details
64GBDDR5DDR5-5600-$2,624.99$41.02AmazonIn stockBuy/View Price
A-Tech
B0DYR9JXRH
Details
64GBDDR5DDR5-6400-$2,699.00$42.17AmazonIn stockBuy/View Price
A-Tech
B0DMC1VFNS
Details
96GBDDR5DDR5-4800-$3,139.49$32.70AmazonIn stockBuy/View Price
Samsung
M321RYGA0BB0-CQK
Details
96GBDDR5DDR5-4800-$3,139.49$32.70AmazonIn stockBuy/View Price
NEMIX
MTC40F204WS1RC56BB2R
Details
96GBDDR5DDR5-5600-$3,779.99$39.37AmazonIn stockBuy/View Price
NEMIX
MTC40F204WS1RC64BC1R
Details
96GBDDR5DDR5-6400-$3,884.99$40.47AmazonIn stockBuy/View Price
NEMIX
MTC40F204WS1RC64BR
Details
96GBDDR5DDR5-6400-$3,884.99$40.47AmazonIn stockBuy/View Price
Samsung
M321R2GA3BB6-CQK
Details
16GBDDR5DDR5-4800-$472.90$29.56AmazonIn stockBuy/View Price
Kingston
KSM56E46BD8KM
Details
48GBDDR5DDR5-5600-$1,749.99$36.46AmazonIn stockBuy/View Price
NEMIX
MTC10F1084S1RC56BR
Details
16GBDDR5DDR5-5600-$629.99$39.37AmazonIn stockBuy/View Price

Relevant Live Enterprise SSD Pricing

Enterprise NVMe prices matter when KV cache, datasets or context storage fall back below the memory tier.

Enterprise NVMe prices
ManufacturerModelPart NumberCapacityInterfaceForm FactorTechPrice$/TBStatusCTA
Micron
7400 PRO 3.84TB U.3
Details
MTFDKCB3T8TDZ
3.84TBNVMe-PCIe4U.3TLC$701.68$182.73In stockBuy/View Price
Dell
3.84TB SAS SSD Read Intensive 12G
Details
MZILT3T8HBLS
3.84TBSAS-12G2.5"TLC$799.99$208.33In stockBuy/View Price
Micron
7450 PRO 3.84TB U.3
Details
MTFDKCC3T8TFS
3.84TBNVMe-PCIe4U.3TLC$1,990.00$518.23In stockBuy/View Price
Kioxia
CD8 Series 3.84TB NVMe U.2
Details
KCD8XVUG3T84
3.84TBNVMe-PCIe4U.2TLC$2,461.20$640.94In stockBuy/View Price
Micron
7600 PRO 3.84TB U.2 PCIe5
Details
MTFDLAL3T8THG
3.84TBNVMe-PCIe5U.2TLC$2,599.00$676.82In stockBuy/View Price
Samsung
PM9A3 3.84TB U.2 NVMe
Details
MZQL23T8HCLS
3.84TBNVMe-PCIe4U.2TLC$3,499.99$911.46In stockBuy/View Price
Samsung
PM9A3 7.68TB U.2 NVMe
Details
MZQL27T6HBLA-00A07
7.68TBNVMe-PCIe4U.2TLC$2,715.99$353.64In stockBuy/View Price
Samsung
PM9A3 7.68TB U.2 NVMe
Details
MZQL27T6HBLA
7.68TBNVMe-PCIe4U.2TLC$3,999.00$520.70In stockBuy/View Price
WD
Ultrastar DC SN655 7.68TB U.3
Details
WUS5EA176ESP7E1
7.68TBNVMe-PCIe4U.3TLC$8,690.03$1,131.51In stockBuy/View Price
Seagate
Nytro 5060 E3.S 15.36TB
Details
XP15360SE70045
15.36TBNVMe-PCIe4E3.STLC$5,879.00$382.75In stockBuy/View Price
SanDisk
Ultrastar DC SN655 61.44TB U.3
Details
WUS5EC0C1ESP7Y3
61.44TBNVMe-PCIe4U.3-$68,559.43$1,115.88In stockBuy/View Price
Samsung
PM9A3 960GB U.2 NVMe
Details
MZQL2960HCJR
0.96TBNVMe-PCIe4U.2TLC$728.00$758.33In stockBuy/View Price
Kioxia
CD8 Series 3.2TB NVMe U.2
Details
KCD8XVUG3T20
3.2TBNVMe-PCIe4U.2TLC$2,570.97$803.43In stockBuy/View Price
Samsung
PM9A3 1.92TB U.2 NVMe
Details
MZQL21T9HCJR
1.92TBNVMe-PCIe4U.2TLC$2,099.00$1,093.23In stockBuy/View Price
Samsung
PM1643a 1.92TB SAS SSD
Details
MZILT1T9HBJR-00007
1.92TBSAS-12G2.5"TLC$2,195.00$1,143.23In stockBuy/View Price
Samsung
PM1643A 1.92TB SAS SSD
Details
MZILT1T9HBJR
1.92TBSAS-12G2.5"TLC$2,199.00$1,145.31In stockBuy/View Price

Marvell Photonic Fabric vs NVIDIA Context Memory Storage

Marvell Photonic Fabric Memory and NVIDIA-oriented context-memory storage address related bottlenecks from different parts of the hierarchy. Marvell's framing is a shared-memory tier: DDR5-backed pooled memory with optical connectivity and HBM caching. Context-memory storage designs are more storage-oriented, using high-performance SSDs as a lower tier for context, cache or retrieval data.

The difference matters for sizing. A shared-memory tier is meant to behave more like memory, with volatile capacity and lower latency than storage. SSD-backed context storage can scale capacity and persistence, but it still sits below memory. In practice, a large inference environment may use both: HBM for hot data, shared memory for warm KV cache and enterprise SSDs for persistent model files, datasets and colder context.

ArchitectureOrientationMemory tierStorage roleLikely workloadPlanning caution
Marvell Photonic Fabric MemoryShared memoryDDR5 pool with HBM cachingNVMe remains below memoryWarm KV cache and pooled inference memoryVendor targets need production validation
Context-memory storageStorage-backed context tierUsually below HBM and DRAMEnterprise SSD capacity and bandwidth are centralContext datasets, retrieval, cache persistence and colder stateDo not model SSD latency as DRAM latency

How to Think About Cost

Photonic memory does not remove the need for careful procurement math. A planner still has to price the HBM already bundled with accelerators, the shared DDR5 pool, memory-node infrastructure, optics or fabric costs, and the enterprise SSD tier underneath. The cheapest design on paper can become expensive if it forces low accelerator utilization or turns active cache into storage traffic.

Start with capacity: how much KV cache is active after reuse? Then model placement: how much remains in HBM, how much goes to shared DDR5 and how much falls to NVMe? Finally check bandwidth. A capacity-rich shared tier that cannot feed the workload is not a balanced design.

Sources and Current Status

Marvell's public materials describe Photonic Fabric technology, memory modules, optical NICs, CXL-oriented memory infrastructure and AI inference memory use cases. Those materials are the basis for the architecture summary here, while DatacenterDisk's live tables provide current DDR5 and enterprise SSD pricing context. Exact deployment sizing should use the final vendor datasheet and your own workload measurements.

Frequently Asked Questions

What is Marvell Photonic Fabric Memory?

Marvell Photonic Fabric Memory is an announced AI memory architecture that combines DDR5-backed pooled memory, HBM caching and optical connectivity to give accelerators access to a larger shared memory tier for inference workloads such as KV cache.

Is Photonic Fabric Memory the same as CXL memory pooling?

No. CXL is a coherency and memory interconnect protocol, while the photonic fabric is an optical transport and scaling approach. Marvell describes CXL connectivity as part of the broader architecture, so the concepts can be complementary.

Does Photonic Fabric Memory replace HBM?

No. The architecture still uses HBM as a faster local cache or accelerator-local tier. The DDR5-backed shared tier is meant to increase reachable memory capacity, not make lower-cost memory identical to HBM.

How does this help KV cache?

Long-context inference can create a large warm KV-cache working set. A shared DDR5 tier can keep more of that working set in memory instead of pushing it directly to NVMe storage.

Is Marvell Photonic Fabric Memory commercially available now?

Marvell has announced the architecture and product direction. Broad production availability, exact system configurations and independently measured performance depend on implementation and customer deployments.

How should buyers model the cost?

Model HBM, shared DDR5 memory and NVMe fallback separately. HBM and fabric pricing are usually quote-based, while DDR5 RDIMM and enterprise SSD prices can be compared against current market data.

Share:WhatsAppFacebookPost on X