Samsung is moving toward a future where the line between a processor and memory disappears. At Hot Chips 2026, the company showed a three phase plan to integrate compute power directly into memory stacks. This move aims to fix the massive data bottlenecks slowing down modern artificial intelligence workloads.
By bringing calculation power into the memory itself, Samsung expects to slash energy use while boosting speed. The strategy addresses the growing memory wall slowing down high-end AI systems. These shifts signify a departure from traditional computer layouts used for decades.
The Three Phases of Memory Fusion
The first step involves adding custom logic layers to the base of High Bandwidth Memory (HBM). These layers perform basic operations before data moves to the main GPU. This reduces traffic on the wires connecting components.
Samsung plans to transition toward zHBM designs to maximize efficiency. This technology moves processing even closer to the storage cells. Such proximity allows for faster response times in large language models.
- Custom logic dies built into HBM4 to reduce latency.
- zHBM technology for enhanced compute in memory capabilities.
- Direct 3D stacking of DRAM on top of processors for total integration.
The transition to HBM4 marks a turning point for the industry. This version allows for customized logic dies tailored for specific AI tasks. Instead of a general purpose memory chip, clients receive a specialized component built for unique software needs.
| Roadmap Phase | Key Technology | Primary Goal |
| Phase 1 | Custom HBM4 Logic | Faster Data Transfer |
| Phase 2 | zHBM Architecture | In-Memory Compute |
| Phase 3 | 3D Die Stacking | Unified Chip Design |
Eliminating the Gap Between Logic and Storage
The final phase places memory directly onto the processor die. This eliminates traditional interconnects entirely. Such a design choice allows chips to behave as one unified brain instead of separate parts.
As shown in the
Samsung memory roadmap, this future design bypasses current physical limits. Building processors this way helps solve heat and power problems found in massive AI clusters. Samsung expects these chips to support the next generation of generative AI models.
Samsung aims to minimize the power used for data movement. Currently, moving bits between RAM and the GPU consumes most of the electricity in a data center. By processing data where storage lives, Samsung helps companies save millions in energy costs.
Merging storage and math creates a more streamlined data path. Engineers believe this will lead to machines capable of handling trillion parameter models with ease. The industry looks toward these advancements to keep up with skyrocketing AI demands.