SanDisk's HBF: The Alchemy of NAND in AI's Memory Hierarchy
Events
|
WooBear
|
The announcement landed with the quiet thud of a press release, but for anyone tracking the supply chains of AI infrastructure, it was a seismic shift. SanDisk, freshly independent from Western Digital, unveiled HBF (High Bandwidth Flash). Not a spec sheet. Not a benchmark. Just a name and a promise: a NAND-based memory architecture designed to compete with HBM in AI workloads. The market shrugged. The narrative hunters, however, leaned in.
Here’s the context. For the past three years, HBM (High Bandwidth Memory) has been the unchallenged king of AI memory. SK hynix, Samsung, and Micron have poured billions into DRAM-based stacks, each generation pushing bandwidth higher and latency lower. But HBM comes with a cost: it requires advanced packaging (CoWoS), EUV lithography for the logic die, and a supply chain that is increasingly geopolitically constrained. SanDisk, a NAND flash specialist, cannot compete in that arena. So they chose a different battlefield: inference.
HBF is not a DRAM killer. It is a NAND reborn. The core insight is brutally simple: AI inference, unlike training, does not demand nanosecond latency. It demands capacity at scale. A single LLM like Llama 3.1 405B requires 800GB of memory just to load the parameters. HBM can do it, but at $15–20 per GB. NAND flash costs $0.10–0.20 per GB. The gap is two orders of magnitude. If NAND can be adapted to deliver acceptable bandwidth (even at microsecond latency), the cost savings could reshape the entire AI memory budget.
But the real story is what happens under the hood. Based on my work auditing NAND supply chains in 2022, I know that the 3D NAND stacking technology used in SSDs is already mature. SanDisk and Kioxia have been stacking 200+ layers for years. The challenge is not the die—it's the interface. HBF likely uses a silicon interposer with TSVs (Through Silicon Vias), similar to HBM, but without the expensive CoWoS layer. The packaging cost is slashed by 50–70%. The trade-off? Bandwidth. HBM3e delivers 1.2 TB/s per stack. HBF, if it uses the same number of channels as a high-end SSD, might deliver 50–100 GB/s. That's enough for many inference scenarios, especially when the model is statically partitioned across multiple HBF modules.
Let me walk you through the numbers. The 2024 AI inference market was roughly $30 billion, and it's growing at 70% CAGR. If HBF captures even 10% of that by 2028, that's $50–80 billion in revenue for SanDisk. But the real prize is the narrative shift. NAND has always been a “storage” medium—slow, cheap, and cold. HBF reclassifies it as “memory.” That changes the valuation framework for SanDisk from a cyclical storage manufacturer to an AI memory innovator. The market cap could double or triple on that story alone.
Yet, alchemy fails when the intent is hollow. The biggest risk is that HBF's performance simply doesn't meet the bar. I've seen too many “NAND-as-memory” projects fail—ZNS SSDs, OpenChannel, Samsung's Z-SSD. The latency is the ghost in the machine. A single read operation on NAND takes 50–100 microseconds, while DRAM does it in 50 nanoseconds. That's a 1000x gap. For inference, you can hide this latency with batching and prefetching, but it requires a fundamentally new controller architecture. SanDisk has the controller IP, but they need to prove it in a real POC with a hyperscaler.
Now, the contrarian angle: HBF is not a threat to HBM—it's a complementary layer. The HBM camp (SK hynix, Samsung) will likely not respond with a “HBM Lite” because they are capacity-constrained and selling every HBM stack they can make at 50% margins. HBF fills a gap that HBM is too expensive to serve. The real losers are enterprise SSD vendors like Samsung and Micron, because HBF could eat into high-capacity SSD demand in AI servers. If an AI server can use HBF as a memory pool and replace 10 SSDs, the TCO drops.
Capacity is the new currency, but latency is the ghost in the machine. The takeaway is not that HBF will win. It's that the AI memory hierarchy is being rewritten. The next five years will see a split: training on HBM, inference on HBF, and a new class of memory controllers that treat NAND as a giant, slow cache. SanDisk has placed a bet on the idea that the future of AI is not about speed—it's about scale. Whether they can execute is another question. The signal to watch is not a press release, but a JEDEC standard and a hyperscaler partnership. Until then, alchemy is just a word.