A little over a year after it upended the tech industry, DeepSeek is back with another apparent breakthrough: a means to stop current large language models (LLMs) from wasting computational depth on simple tasks.
Detailed in a recently published technical paper, the Chinese startup’s Engram concept offloads static knowledge (simple information lookups) from the LLM's primary memory to host memory (CPU RAM) in order to relieve the model’s core computational network.
Essentially, DeepSeek’s Engram model bypasses GPU memory constraints to potentially allow firms running LLMs to scale parameters aggressively without hitting GPU memory walls.
“Engram effectively “deepen[s]” the network by relieving early layers from static reconstruction tasks, thereby freeing up attention capacity to focus on global context and complex reasoning,” the paper reads. “This architectural shift translates into substantial improvements in long-context capabilities.”
DeepSeek's solution to transformer inefficiency
DeepSeek’s paper dropped amid a growing sense of revisionism toward transformer-based LLMs, with AI developers wanting more from their models. Researchers have explored several alternatives to traditional LLM architectures. Diffusion-based models, predominantly used for image generation in tools like Stable Diffusion and Midjourney, have been adapted for large language models, with the startup Inception developing its own diffusion-based LLMs called Mercury.
Another approach includes liquid foundation models (LFMs), which draw on concepts from signal processing to create faster and more adaptable general-purpose AI. And former Meta chief AI scientist Yann LeCun has proposed joint embedding predictive architecture (JEPA) as an alternative to auto-regressive generative architectures, which are used by conventional LLMs – though some JEPA implementations still incorporate some underlying transformer modules.
DeepSeek’s paper looks to add to the transformer revisionism, arguing traditional systems “lack a native primitive for knowledge lookup” due to the fact that they’re designed for vector transformation, not memory retrieval. As a result, current LLMs are forced to simulate retrieval through computation – meaning that they utilize considerable attention capacity and feed-forward network computations across multiple layers just to retrieve simple information.
Engram aims to address this design inefficiency by introducing what the Chinese startup calls “conditional memory” – a complementary axis of sparsity that works alongside the conditional computation provided by mixture-of-experts (MoE) architectures.
“Whereas conditional computation sparsely activates parameters to process dynamic logic, conditional memory relies on sparse lookup operations to retrieve static embeddings for fixed knowledge,” the paper reads.
Engram works alongside an MoE system: the conditional memory component saves energy by only performing a lookup when needed for simple facts and information, while the latter conditional computation segment only utilizes the parts of the model needed for highly complex tasks.
In practical terms, DeepSeek’s paper contends it was able to offload a 100-billion parameter embedding table entirely to host memory, which incurred less than 3% throughput penalty.
The researchers were then able to further boost performance optimization by leveraging Zipfian distribution – where they strategically cache frequently accessed embeddings in faster GPU memory, reserving slower, high-capacity storage for the less common "long tail" of rare patterns.
“This stratification allows Engram to scale to massive memory capacities with minimal impact on effective latency,” the paper reads.
Scaling efficiency amid the memory constraint crisis
Alongside transformer revisionism, DeepSeek’s Engram paper dropped at a time when developers and enterprises deploying large-scale AI face considerable constraints on the memory portion of their stack.
LLMs' reliance on high-bandwidth memory (HBM) just to process simple tasks not only creates a bottleneck in terms of performance as outlined above, but also in terms of costs.
Prices for HBM are soaring amid elongated lead times for hardware, so Engram’s ability to free up memory capacity for more complex reasoning tasks means models can more easily handle far longer context inputs more efficiently, but its deterministic retrieval mechanism means memory capacity can be scaled linearly across multiple GPUs to boost scalability.
“Engram advocates for infrastructure-aware efficiency as a first-class design principle. Its deterministic addressing allows for the decoupling of storage and compute, enabling the offloading of massive parameter tables to host memory with negligible inference overhead,” DeepSeek’s researchers wrote. “We envision conditional memory functions as an indispensable modeling primitive for next-generation sparse models.”
The Chinese startup’s paper was the second in a week aiming to address memory imbalances impacting AI model performance, following research from Google engineers that LLM inference is being impacted due to fundamental memory and networking problems, rather than a lack of compute power.
For the team at DeepSeek, though, Engram looks to kick off 2026 as it did last January: by doubling down on architectural improvements to overcome limited access to high-end hardware.
The startup debuted DeepSeek V3.2 last December, another of its top AI models built using Nvidia’s China market exclusive H800 chips, but one that takes advantage of an “efficient attention mechanism” to achieve performance levels comparable with that of OpenAI’s GPT-5.
But any future AI development efforts may be further hampered, as U.S. lawmakers are moving to add memory hardware to chip export controls. The House Select Committee on the Chinese Communist Party (CCP) advocated for expanded regulations to include high-bandwidth memory 3 extended (HBM3E) to prevent vital memory resources from being diverted away from American companies.
Comments