Persistent State Machines: LLM Attention with INT4 In-Memory Cells
Researchers present a formal framework for Large Language Model attention operators using Persistent State Machines (PSMs) implemented on programmable logic fabric. The study demonstrates that this architecture achieves high efficiency and low power consumption, validated through FPGA implementations.
Why it matters
This research offers a potential pathway for more energy-efficient and scalable AI hardware, moving away from traditional GPU-heavy compute models.
Persistent State Machines: Complete Mathematical Proofs and Vivado Implementation Synthesis (Version 8.0)
ABSTRACT We present a formal discrete framework for attention operators in Large Language Models via Persistent State Machines (PSMs). Computation is broadcast to stationary in-memory cells that evaluate local deterministic state transitions. Complete mathematical proofs are given for quantization error bounds, a concrete multi-phase discrete Softmax construction under an explicit bounded-logits assumption, deterministic finite-automaton equivalence with spatial factorization, and membership in DSPACE(O(n)).
We validate the implementation feasibility of the architecture on contemporary programmable logic fabric through two distinct evaluation flows:
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in