
Published on: September 18, 2026
SK hynix HBF, PIM and SALT-KV Expand AI Memory From HBM Into a Tiered Inference Architecture
SK hynix presented HBF, PIM/AiMX and SALT-KV at AI Infra Summit 2026. The portfolio maps long-context capacity, fast decoding and KV-cache placement to different hardware and software layers. Public specifications, platform support and production benchmarks remain the commercial gates.
Inference broadens the AI memory hierarchy
SK hynix used AI Infra Summit 2026 to present a portfolio built around HBF, PIM/AiMX and SALT-KV. The event ran from September 15 to 17 in Santa Clara and drew roughly 6,000 attendees, twice the prior year's count. The exhibits included an HBF structural model, an AiM chip, an AiMX accelerator card, a server fitted with AiMX, and a SALT-KV demonstration using enterprise SSDs. The portfolio addresses memory placement and data movement rather than promoting one faster component.
Training systems have made HBM bandwidth a primary design variable. Inference adds a different set of constraints: long contexts, persistent sessions, expanding KV cache and strict cost per request. Keeping every cache object in HBM is expensive, while moving too much data to conventional storage can add latency. SK hynix separates the problem into a high-bandwidth flash tier, processing close to memory, and software-directed movement across HBM, DRAM and SSD.
HBF targets the space between HBM and SSD
High Bandwidth Flash applies through-silicon-via stacking to NAND, using a structural approach associated with HBM while retaining flash capacity economics. Its intended role is a new tier between high-bandwidth, high-cost HBM and much larger but slower SSD storage. HBF is not a direct HBM replacement because the media differ in latency, endurance, retention, capacity and cost.
Long-context inference is the clearest use case. HBF could hold data that exceeds the economic capacity of HBM while offering more bandwidth than a conventional SSD path. Commercial evidence still requires stack configuration, interface details, controllers, thermal limits, endurance specifications and system support. A structural model demonstrates technical direction, but it does not establish orderable part numbers, volume pricing, lead times or interchangeability with current memory products.
PIM and AiMX reduce data movement
Processing-in-memory embeds computational functions alongside memory to reduce traffic between processors and data. SK hynix displayed its AiM chip, the AiMX card and a server using the card, accompanied by an LLM service demonstration. The company positions the approach for fast-decoding workloads where latency and power efficiency carry premium value.
PIM performance depends on software tools, operator coverage, numerical precision, memory consistency and platform integration. Moving selected functions closer to data can ease the memory wall, but it does not remove the need for general-purpose CPUs and GPUs. Production evaluation must compare end-to-end latency, throughput, power, programmability and serviceability rather than peak silicon metrics alone.
SALT-KV makes cache placement a software variable
Semantic-Aware Lifecycle Tiering for KV Cache divides cache data into context-based segments, evaluates reuse value and storage cost, and places each segment in HBM, DRAM or SSD. The method turns memory tiering from a static capacity choice into a workload-aware policy. High-value and frequently reused objects remain close to compute; colder data can move to lower-cost tiers.
This model gives enterprise SSDs a more direct role in inference architecture. Random access behavior, latency consistency, endurance and quality of service become relevant to token generation when cache data reaches SSD. HBM and DRAM demand does not disappear; it concentrates on data with the highest reuse and latency value. Actual component consumption will depend on context length, concurrency, cache-hit behavior and the effectiveness of the tiering policy.
Commercial gates remain ahead
The September 17 disclosure is a technology showcase, not a volume-supply announcement. HBF has no public orderable part number or mass-production date in this material. PIM/AiMX deployment depends on server and software integration. SALT-KV economics need validation with production models, realistic concurrency and specific enterprise SSD configurations. The exhibition therefore does not support a conclusion that HBM, NAND or enterprise SSD pricing will rise uniformly in the near term.
The useful milestones are public HBF specifications, sample delivery, server-platform support, independent SALT-KV benchmarks and cloud deployment. Those signals would mark progress from an architectural proposal to a qualified BOM option. The broader direction is already visible: AI memory competition is expanding beyond individual HBM stacks into DRAM, NAND, enterprise SSDs, advanced packaging and software orchestration.
Implications for OEM and EMS planning
The new hierarchy changes how capacity requirements may be specified. A future AI server BOM could include several memory service levels defined by latency, bandwidth, endurance and cost rather than a single aggregate capacity number. OEM qualification would then need to cover controllers, firmware, thermal behavior and failure recovery across tiers. EMS execution would depend on traceable components and validated configurations because a substitution at one tier can change system performance.
Supply constraints may also migrate between manufacturing stages. HBF combines NAND output with TSV formation, die stacking, package assembly and testing. PIM adds specialized silicon and accelerator integration. SALT-KV increases the importance of enterprise SSD firmware and predictable service quality. Sufficient wafer output alone may not guarantee system availability if packaging, controllers or software validation lag.
Procurement evidence should remain product-specific. Orderable identifiers, qualified-vendor listings, authorized-channel quotes and committed production dates carry more weight than exhibition interest. HBM, DRAM and SSD cannot be treated as interchangeable capacity pools. Each tier serves a different performance and endurance envelope, and supplier roadmaps will reach production on different schedules.
A measurable path from showcase to deployment
Four groups of evidence can establish commercial progress. First, HBF needs public interface, density, bandwidth, endurance and thermal specifications. Second, server and accelerator platforms need documented support. Third, SALT-KV needs repeatable benchmarks that disclose model size, context length, concurrency, cache-hit rate and SSD configuration. Fourth, suppliers need volume schedules, qualification status and lead-time guidance.
Until these gates are visible, HBF, PIM and SALT-KV should be modeled as architecture options rather than assured supply. Their importance lies in defining how inference systems may balance capacity and bandwidth. The portfolio indicates that the next memory cycle will be evaluated at system level, where hardware tiers and software policy determine the economic value of each component.